Megadose AI progress, ranked daily.

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation in Transformers

· HF Daily Papers ·
A three-layer recurrent setup claims GPT-2 Small-level accuracy with far fewer parameters.

The paper proposes fixed prelude and coda blocks around one shared transformer core that is reused multiple times. A lightweight gate modulates each recurrent update using the hidden state, the prelude output, and resampled noise. Under matched FLOPs, the authors say it matches a 12-layer GPT-2 Small baseline and beats several depth-sharing alternatives across their tested scale-budget grid. At large scale, they report 63% fewer parameters and 59% less peak decoding memory, with a 10% increase in compiled generation latency. HF Daily Papers' note

score 5

Categories: Research