Loop the Loopies!
Loopie claims looped Transformers can beat same-compute vanilla baselines.
The paper introduces two MoE Loopie models: 20B parameters with 2B active, and 6B with 0.6B active. Its central claim is that looping can outperform simply scaling a vanilla Transformer under the same compute budget, including against a 30B-A3B baseline. The authors also report a post-training pipeline that gives Loopie strong reasoning results, including gold-medal performance on the 2025 IMO and IPhO without tools. ArXiv · AI/CL/LG's note
The paper introduces two MoE Loopie models: 20B parameters with 2B active, and 6B with 0.6B active. Its central claim is that looping can outperform simply scaling a vanilla Transformer under the same compute budget, including against a 30B-A3B baseline. The authors also report a post-training pipeline that gives Loopie strong reasoning results, including gold-medal performance on the 2025 IMO and IPhO without tools. ArXiv · AI/CL/LG's note
score 7