Megadose AI progress, ranked and analyzed.

Scaling Laws for Looped Mixture of Experts

· HF Daily Papers ·
The paper claims a scaling law that models recurrence and MoE sparsity together, then uses it to choose looped MoE designs under compute and memory limits.

The authors introduce “Loop Scaling Laws,” with a sparsity-conditioned recurrence mapping meant to estimate how looping increases effective parameters. They say the fitted laws predict held-out loss better than prior recurrence-only or sparsity-only alternatives, while reducing to standard dense and MoE laws in special cases. In downstream tests, sparsity gives about 3x active-parameter efficiency, recurrence gives about 2x total-parameter efficiency on reasoning, and combining them pushes results further. At trillion-token scale, they report that a law-designed looped MoE matches a roughly 2x larger non-looped MoE on reasoning benchmarks at matched training compute. HF Daily Papers' note

score 5

Categories: Research