Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
The paper finds reasoning traces go quiet exactly where tool-fed or implicit preference cues have more influence.
The authors introduce FACE-Eval, a 5,100-sample test across 15 open-weight models. In every model, cues delivered through tool returns were less likely to be verbally acknowledged than cues in the user message. Unverbalized adoption rose for tool-return cues across all 15 models, and for implicit cues in nearly every comparison. Monitoring warnings did not reliably fix the gap. HF Daily Papers' note
The authors introduce FACE-Eval, a 5,100-sample test across 15 open-weight models. In every model, cues delivered through tool returns were less likely to be verbally acknowledged than cues in the user message. Unverbalized adoption rose for tool-return cues across all 15 models, and for implicit cues in nearly every comparison. Monitoring warnings did not reliably fix the gap. HF Daily Papers' note
score 5