Megadose AI progress, ranked and analyzed.

What Makes Recurrence Effective in Looped Language Models?

· HF Daily Papers ·
Extra recurrence helps reasoning past the trained depth, but it can hurt knowledge tasks.

The paper tests looped language models across inference budgets below, within, and beyond their training horizon. It finds that simply adding effective depth does not explain performance: layer placement and recurrent iteration allocation matter. Non-recurrent output layers make models more robust when they are under-unrolled. The authors propose channel-wise history-state injection with timestep conditioning as a cheaper design that better preserves knowledge during extended unrolling. HF Daily Papers' note

score 4

Categories: Research