Megadose AI progress, ranked and analyzed.

SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers

· HF Daily Papers ·
SMELT reports lower training compute at matched FLOPs, parameters, and KV cache by reusing the middle layers of an MoE Transformer.

The paper compares looped and unlooped architectures under matched per-token FLOPs, non-embedding parameters, and cache budget. Its SMELT recipe loops the middle half of layers twice, then scales the setup to 54B non-embedding parameters. The authors fit separate Chinchilla-style laws and report 6.8-18.0% training FLOP savings on the compute-optimal frontier. They also say gains carry to downstream benchmarks, especially code, and grow with longer samples and more in-context examples. Source: HF Daily Papers' note.

score 6

Categories: Research