How to Loop MoE: Flatten the Experts, Untie the Attention
Foil improves looped MoE by widening the expert choice per routing step and giving each loop pass separate attention.
The paper keeps expert parameters and expert compute per token fixed, then flattens expert layers into more experts per layer and more passes. In experiments, every Foil model beats the unflattened looped baseline on pretraining loss at 20B tokens. At 100B tokens, more flattening improves loss monotonically, with the most flattened version 0.012 nat below baseline at equal parameters and compute. Untied attention also makes routing more balanced and more confident. HF Daily Papers' note
The paper keeps expert parameters and expert compute per token fixed, then flattens expert layers into more experts per layer and more passes. In experiments, every Foil model beats the unflattened looped baseline on pretraining loss at 20B tokens. At 100B tokens, more flattening improves loss monotonically, with the most flattened version 0.012 nat below baseline at equal parameters and compute. Untied attention also makes routing more balanced and more confident. HF Daily Papers' note
score 5