Answer-Distribution Trajectories: A Stochastic-Dynamics View of LLM Reasoning
The paper argues that LLM reasoning should be judged by how answer probabilities move during a chain of thought, not just by the final answer.
The authors introduce “answer-distribution trajectories,” tracking the full predictive distribution over candidate answers as reasoning unfolds. They say this reveals exploration, revision, motion, and commitment patterns that endpoint accuracy or entropy summaries can miss. Across sixteen open-weight models and four reasoning benchmarks, traces with similar final answers and entropy profiles still showed different dynamics. Training and inference choices also changed those profiles. ArXiv · AI/CL/LG's note
The authors introduce “answer-distribution trajectories,” tracking the full predictive distribution over candidate answers as reasoning unfolds. They say this reveals exploration, revision, motion, and commitment patterns that endpoint accuracy or entropy summaries can miss. Across sixteen open-weight models and four reasoning benchmarks, traces with similar final answers and entropy profiles still showed different dynamics. Training and inference choices also changed those profiles. ArXiv · AI/CL/LG's note
score 4