Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability
Shorter reasoning did not uniformly make models less inspectable.
The paper tests three efficiency-training methods that pressure chain-of-thought outputs to use fewer tokens. It finds faithfulness usually drops, mainly because trained models become less consistent across related inputs. Monitorability holds up better: even with much shorter reasoning, models often still acknowledge when an input intervention shaped the answer. The authors frame the result as conditional, not a blanket defense of efficient reasoning training. HF Daily Papers' note
The paper tests three efficiency-training methods that pressure chain-of-thought outputs to use fewer tokens. It finds faithfulness usually drops, mainly because trained models become less consistent across related inputs. Monitorability holds up better: even with much shorter reasoning, models often still acknowledge when an input intervention shaped the answer. The authors frame the result as conditional, not a blanket defense of efficient reasoning training. HF Daily Papers' note
score 4