Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
A standard API tool call can make closed frontier models expose intermediate reasoning traces.
The paper tests whether those traces behave like real chain-of-thought rather than after-the-fact explanations by comparing them with native CoT in open-source models. The extracted reasoning matches native-reasoning performance and beats no-reasoning baselines on math, science, and code tasks. For GPT-6 Astra, the authors report a more compact pattern: it selects a correct path earlier, keeps elementary steps internal, and externalizes only key reasoning. Source: ArXiv · AI/CL/LG's note.
The paper tests whether those traces behave like real chain-of-thought rather than after-the-fact explanations by comparing them with native CoT in open-source models. The extracted reasoning matches native-reasoning performance and beats no-reasoning baselines on math, science, and code tasks. For GPT-6 Astra, the authors report a more compact pattern: it selects a correct path earlier, keeps elementary steps internal, and externalizes only key reasoning. Source: ArXiv · AI/CL/LG's note.
score 6