Megadose AI progress, ranked daily.

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

· HF Daily Papers ·
Mixed SFT beat the next-chunk RL setup in this controlled comparison, with far less compute.

The paper tests whether next-chunk reasoning RL is actually responsible for reported gains on no-CoT data. Its alternative is a single supervised stage that mixes no-CoT material with long-CoT data. That Mixed SFT setup reaches a higher post-RLVR performance ceiling while using more than 60 times less training compute. The result holds on in-domain math reasoning and out-of-domain reasoning tasks. HF Daily Papers' note

score 4

Categories: Research