Megadose AI progress, ranked and analyzed.

Motif 3: Technical Report

· HF Daily Papers ·
Motif 3 is a 314B-parameter sparse MoE model that activates 13.2B parameters per token.

The report says each sparse layer routes a token to eight of 384 experts, aiming for high capacity with limited compute. Its core architecture is Grouped Differential Latent Attention, paired with other changes for stability, specialization, and inference efficiency. The model was pretrained on about 12.5 trillion tokens and trained for contexts up to 256K tokens. Post-training combines specialist teachers, a software-engineering teacher, and multi-teacher on-policy distillation. HF Daily Papers' note

score 7

Categories: Model Releases, Research