Megadose Built for builders and researchers.

Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

· ArXiv · AI/CL/LG ·
A 1.5B model learned to pick shorter or longer reasoning modes per problem, cutting MATH500 output tokens by 41% with only a small accuracy drop.

The paper trains the mode choice inside GRPO, using `NoThink`, `Short`, or `Long` as the first response token rather than adding a separate router. Hard token caps and shaped rewards keep the modes from collapsing into one behavior. On held-out MATH500, average accuracy was 0.782 versus 0.796 for the base model, while mean response length fell from 4,796 to 2,811 tokens. The learned policy also transferred without retraining, including a 76% token reduction on GSM8K. ArXiv · AI/CL/LG's note

score 6

Categories: Research