Megadose Built for builders and researchers.

Decoding Looped Transformers Better for (Almost) Free

· ArXiv · AI/CL/LG ·
LoopCD uses earlier recurrent passes as a built-in weak signal to improve looped Transformer decoding without training.

The paper says looped Transformers already produce intermediate states for the same next-token prediction, but normal decoding throws those away. LoopCD contrasts the final prediction against an earlier pass, either through logits with one extra output pass or hidden states with no output overhead. The authors report gains across four looped Transformer families, including AIME 2024 and HumanEval improvements, and say the method can cut recurrent loops while matching or beating unguided full-depth baselines. ArXiv · AI/CL/LG's note

score 4

Categories: Research