Revisiting Complete Reasoning Traces for Post-Training
Partial or endpoint-only reasoning traces may train models about as well as full chains.
The paper argues that long collected reasoning trajectories contain substantial redundancy for post-training. Its experiments find that heavily truncated partial traces remain effective, while intermediate tokens contribute little to final reasoning quality in attention and token-removal analyses. The authors say endpoint-based training changes reasoning behavior and also helps reinforcement learning and on-policy distillation setups. HF Daily Papers' note
The paper argues that long collected reasoning trajectories contain substantial redundancy for post-training. Its experiments find that heavily truncated partial traces remain effective, while intermediate tokens contribute little to final reasoning quality in attention and token-removal analyses. The authors say endpoint-based training changes reasoning behavior and also helps reinforcement learning and on-policy distillation setups. HF Daily Papers' note
score 5