Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
Instruction tuning changed how confidently models answered, even when accuracy barely moved.
The paper compares three matched base and instruction-tuned models on question-answering benchmarks. Instruction tuning consistently altered answer confidence while producing limited accuracy changes and worse likelihood-based calibration. The authors also found rationale diversity shifted unevenly: diversity across rationales fell consistently, while surface lexical diversity varied by model and benchmark. Those effects remained after controlling for answer choice and rationale length. ArXiv · AI/CL/LG's note
The paper compares three matched base and instruction-tuned models on question-answering benchmarks. Instruction tuning consistently altered answer confidence while producing limited accuracy changes and worse likelihood-based calibration. The authors also found rationale diversity shifted unevenly: diversity across rationales fell consistently, while surface lexical diversity varied by model and benchmark. Those effects remained after controlling for answer choice and rationale length. ArXiv · AI/CL/LG's note
score 4