Stepped MoE: Segment-Level Routing with Configurable Inference Complexity
One model is trained to run at 1B, 2B, 3B, or 4B active parameters depending on device limits and task needs.
The paper combines elastic sub-networks with sparse gating so inference can be adjusted for both hardware constraints and input difficulty. Its backbone conditions on context and a target efficiency setting, then routes work through task-relevant parameters. In experiments, the authors report 2-5% gains over dense counterparts on knowledge-heavy benchmarks, with latency similar to dense models. HF Daily Papers' note
The paper combines elastic sub-networks with sparse gating so inference can be adjusted for both hardware constraints and input difficulty. Its backbone conditions on context and a target efficiency setting, then routes work through task-relevant parameters. In experiments, the authors report 2-5% gains over dense counterparts on knowledge-heavy benchmarks, with latency similar to dense models. HF Daily Papers' note
score 4