Megadose Built for builders and researchers.

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

· HF Daily Papers ·
LOOM is pitched as a way to make looped MoE models keep improving past the usual two-pass limit.

The paper says looped MoE Transformers stall because repeated passes destabilize hidden states and often route to the same experts again. LOOM counters that with scaled residual updates, repeated input reinjection, per-loop routers, and a residual path carrying earlier loop outputs forward. In the reported runs, models scaled stably to 9-12 loops. A 700M near-iso-FLOP model did best at 5 loops, while a 1.7B model without FLOP matching peaked at 9 loops.

HF Daily Papers' note

score 5

Categories: Research