Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models
Calibration helped small models say what their confidence meant, but rarely certified them to answer autonomously.
The paper tested 11 instruction-tuned models from 0.5B to 14B parameters on ARC-Challenge and TruthfulQA across 25,168 local predictions. It gives theoretical limits on what confidence calibration can and cannot fix, including cases where temperature scaling cannot work. Platt scaling cut calibration error sharply, but at a 20% risk budget only three model-task pairs earned certified autonomy, and none did at 10%. The author also reports finding and repairing an answer-ordering artifact in TruthfulQA’s multiple-choice setup. ArXiv · AI/CL/LG's note
The paper tested 11 instruction-tuned models from 0.5B to 14B parameters on ARC-Challenge and TruthfulQA across 25,168 local predictions. It gives theoretical limits on what confidence calibration can and cannot fix, including cases where temperature scaling cannot work. Platt scaling cut calibration error sharply, but at a 20% risk budget only three model-task pairs earned certified autonomy, and none did at 10%. The author also reports finding and repairing an answer-ordering artifact in TruthfulQA’s multiple-choice setup. ArXiv · AI/CL/LG's note
score 4