Megadose AI progress, ranked and analyzed.

Loop the Loopies!

· HF Daily Papers ·
Loopie claims looped Transformers can beat larger vanilla baselines under the same compute budget.

The paper introduces two MoE models: 20B parameters with 2B active, and 6B with 0.6B active. Its central claim is that looping the model can outperform simply scaling parameter count, a weakness earlier looped Transformers struggled with. The authors say ablations, including against a vanilla 30B-A3B model, show substantial gains at matched compute. They also report that a post-training method gives Loopie strong, frontier-level reasoning performance. HF Daily Papers' note

score 6

Categories: Model Releases, Research