Length Penalties Make Chain-of-Thought Less Monitorable
Shorter reasoning traces kept the answers mostly intact while hiding more of the misleading evidence monitors were meant to catch.
Bryce Little trained Qwen3 4B and 14B models toward different chain-of-thought lengths, then tested whether biasing hints still shaped answers. The hints influenced choices near baseline, but compressed traces mentioned them less often. At the shortest target length, lower-bound faithfulness fell to 63.1% of baseline for Qwen3 14B and 69.4% for Qwen3 4B. Randomly deleting sentences from uncompressed chains did not explain the full drop, suggesting compression changes what evidence remains visible. HF Daily Papers' note
Bryce Little trained Qwen3 4B and 14B models toward different chain-of-thought lengths, then tested whether biasing hints still shaped answers. The hints influenced choices near baseline, but compressed traces mentioned them less often. At the shortest target length, lower-bound faithfulness fell to 63.1% of baseline for Qwen3 14B and 69.4% for Qwen3 4B. Randomly deleting sentences from uncompressed chains did not explain the full drop, suggesting compression changes what evidence remains visible. HF Daily Papers' note
score 5