Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
Reasoning training made models sound more deliberative, but not more aligned with the behaviors most linked to correct answers.
The paper measures this with “Behavioral Lift,” comparing correctness when a reasoning behavior appears versus when it does not. Across 15 models and 6 benchmarks, the authors found thinking models amplified self-correction, hypothesis testing, and uncertainty acknowledgment. The strongest signals of correctness were instead confidence calibration, knowledge alignment, and self-awareness. Uncertainty acknowledgment rose 3–7x but was weakly or negatively associated with correctness. HF Daily Papers' note
The paper measures this with “Behavioral Lift,” comparing correctness when a reasoning behavior appears versus when it does not. Across 15 models and 6 benchmarks, the authors found thinking models amplified self-correction, hypothesis testing, and uncertainty acknowledgment. The strongest signals of correctness were instead confidence calibration, knowledge alignment, and self-awareness. Uncertainty acknowledgment rose 3–7x but was weakly or negatively associated with correctness. HF Daily Papers' note
score 5