Megadose AI progress, ranked and analyzed.

Maglev: Sliding Recurrent Memory

· HF Daily Papers ·
Maglev trains a fixed-size recurrent memory to stand in for full-history context at inference.

The architecture pairs a stronger prefiller with full-history access against a decoder limited to sliding-window attention plus recurrent K/V injection. During training, a memory consistency loss pushes the decoder’s memory toward the prefiller’s targets, so inference can run with the decoder alone. The authors report better validation loss and downstream pretraining benchmark results than sliding-window and latent recurrent Transformer baselines. They also say sharing parameters between the two models keeps most gains while reducing parameter memory. HF Daily Papers' note

score 5

Categories: Research