Issue 2026-09-12 · Industry · 研究 · 安全
Distinct internal patterns for reasoning steps

A new study finds that models exhibit distinct internal patterns for different written reasoning steps—calculation, formula retrieval, and deduction—with these patterns strongest in middle layers. This shows models internally process more than their visible chain-of-thought, affecting interpretability and AI safety analysis.
Read original ↗