Megadose Built for builders and researchers.

Recurrent Looped Transformer

· HF Daily Papers ·
RLT adds token-by-token feedback so its computation path can grow with the sequence while keeping per-token cost fixed.

The paper splits an eight-layer model into a causal encoder and recurrent decoder, feeding each token’s final decoder state into the next. In the reported algorithmic tests, some RLT splits generalize far beyond training lengths where the baseline Transformer fails, including parity and swap-based permutation tracking. Ablations say the gains depend on that feedback loop: removing it drops key tasks to chance. Chunked feedback preserves parity performance but hurts permutation tracking, which the authors say needs per-token updates. HF Daily Papers' note

score 5

Categories: Research