Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
Instruction tuning changed how sure models sounded without much changing whether they were right.
The paper tests three matched base and instruction-tuned language models on question-answering benchmarks. Instruction tuning consistently shifted verbalized confidence, while predictive accuracy changed little and likelihood-based calibration got worse. The authors also found rationale diversity narrowed across rationales, though surface-level lexical diversity moved differently by model and benchmark. Those effects remained after controlling for answer choice and rationale length. HF Daily Papers' note
The paper tests three matched base and instruction-tuned language models on question-answering benchmarks. Instruction tuning consistently shifted verbalized confidence, while predictive accuracy changed little and likelihood-based calibration got worse. The authors also found rationale diversity narrowed across rationales, though surface-level lexical diversity moved differently by model and benchmark. Those effects remained after controlling for answer choice and rationale length. HF Daily Papers' note
score 4