Megadose AI progress, ranked and analyzed.

Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

· ArXiv · AI/CL/LG ·
The paper argues that counting where two models diverge can calibrate reasoning confidence better than token probabilities alone.

The authors introduce Divergent Token Confidence, which tracks tokens where next-token distributions sharply disagree along the same reasoning path. The count of those divergent tokens is reported as nearly negatively associated with answer accuracy. Across six math benchmarks, the count-only white-box estimator posts 13.0% average expected calibration error, versus 32.7%–42.4% for standard full-sequence confidence methods. In black-box tests on DeepSeek-V3.2, mean expected calibration error drops from 32.1%–40.2% to 13.7%–16.3%. ArXiv · AI/CL/LG's note

score 4

Categories: Research