Megadose AI progress, ranked and analyzed.

When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models

· HF Daily Papers ·
When2Think trains a hybrid reasoning model to spend fewer tokens on easy problems and keep longer reasoning for harder ones.

The paper frames efficient reasoning as instance-by-instance compute allocation, not a fixed token penalty or rigid route. Its IDAC reward shaping uses pre-computed accuracy and token-use statistics, plus verifier rewards, to optimize without a learned reward model or online reference-model queries. On math benchmarks, the authors report a 10.0% Pass@3 gain and 27.9% lower token use on AIME24 versus the base model, and 40.0% Pass@3 on AIME25. HF Daily Papers' note

score 5

Categories: Research