Improving Test-Time Scaling with Adaptive Looped Transformers
TaH2 routes extra loop iterations only to tokens likely to benefit, improving compute scaling over fixed-depth looping.
The paper says existing looped transformers often gain accuracy faster as decoding compute rises, but still trail non-looped baselines at matched compute. Its proposed method, TaH2, trains an iteration decider with lookahead depth supervision so the model can spend extra computation selectively. On AIME benchmarks, TaH2 raises the accuracy-compute slope by 53% over the non-looped baseline and beats that baseline’s peak accuracy by about 3.4 points at matched test-time compute. The reported gains grow as maximum iteration depth increases, while prior looped models mostly plateau. HF Daily Papers' note
The paper says existing looped transformers often gain accuracy faster as decoding compute rises, but still trail non-looped baselines at matched compute. Its proposed method, TaH2, trains an iteration decider with lookahead depth supervision so the model can spend extra computation selectively. On AIME benchmarks, TaH2 raises the accuracy-compute slope by 53% over the non-looped baseline and beats that baseline’s peak accuracy by about 3.4 points at matched test-time compute. The reported gains grow as maximum iteration depth increases, while prior looped models mostly plateau. HF Daily Papers' note
score 5