Megadose Built for builders and researchers.

Decoding Looped Transformers Better for (Almost) Free

· HF Daily Papers ·
LoopCD uses earlier recurrent passes as a built-in weak signal to improve looped Transformer decoding without training.

The paper says looped Transformers already produce intermediate states aligned to the same next-token prediction, but standard decoding ignores them. LoopCD contrasts the final prediction with an earlier loop, either through logits with one extra output pass or hidden states with no output overhead. The authors report gains across four looped Transformer families, including AIME 2024 and HumanEval improvements, and say the method can halve recurrent loops while matching or beating unguided full-depth baselines. HF Daily Papers' note

score 4

Categories: Research