Megadose AI progress, ranked and analyzed.

How to Loop MoE: Flatten the Experts, Untie the Attention

· ArXiv · AI/CL/LG ·
Foil improves looped MoE by widening the expert choice per routing step while keeping expert parameters and compute fixed.

The paper says Foil halves expert layers, doubles experts per layer, and doubles passes, then gives each pass separate attention parameters while sharing experts and routers. In experiments, every Foil model beat the unflattened looped baseline at 20B tokens. At 100B tokens, more flattening produced lower loss, with the most flattened model 0.012 nat below baseline at equal parameters and compute. The authors also report downstream accuracy at parity or better, and more balanced, confident routing from untying attention. ArXiv · AI/CL/LG's note

score 5

Categories: Research