Loop the Loopies!
Loopie claims looped Transformers can beat larger vanilla baselines under the same compute budget.
The paper introduces two MoE models: 20B parameters with 2B active, and 6B with 0.6B active. Its central claim is that looping the model can outperform simply scaling parameter count, a weakness earlier looped Transformers struggled with. The authors say ablations, including against a vanilla 30B-A3B model, show substantial gains at matched compute. They also report that a post-training method gives Loopie strong, frontier-level reasoning performance. HF Daily Papers' note
The paper introduces two MoE models: 20B parameters with 2B active, and 6B with 0.6B active. Its central claim is that looping the model can outperform simply scaling parameter count, a weakness earlier looped Transformers struggled with. The authors say ablations, including against a vanilla 30B-A3B model, show substantial gains at matched compute. They also report that a post-training method gives Loopie strong, frontier-level reasoning performance. HF Daily Papers' note
score 6