Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs
The paper argues that counting where two models diverge can calibrate reasoning confidence better than token probabilities alone.
The authors introduce Divergent Token Confidence, which tracks tokens where next-token distributions sharply disagree along the same reasoning path. The count of those divergent tokens is reported as nearly negatively associated with answer accuracy. Across six math benchmarks, the count-only white-box estimator posts 13.0% average expected calibration error, versus 32.7%–42.4% for standard full-sequence confidence methods. In black-box tests on DeepSeek-V3.2, mean expected calibration error drops from 32.1%–40.2% to 13.7%–16.3%. ArXiv · AI/CL/LG's note
The authors introduce Divergent Token Confidence, which tracks tokens where next-token distributions sharply disagree along the same reasoning path. The count of those divergent tokens is reported as nearly negatively associated with answer accuracy. Across six math benchmarks, the count-only white-box estimator posts 13.0% average expected calibration error, versus 32.7%–42.4% for standard full-sequence confidence methods. In black-box tests on DeepSeek-V3.2, mean expected calibration error drops from 32.1%–40.2% to 13.7%–16.3%. ArXiv · AI/CL/LG's note
score 4