Megadose Built for builders and researchers.

SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving

· HF Daily Papers ·
SlimWise keeps prefill unpruned and prunes only decode, aiming to cut MoE serving traffic without giving up much quality.

The paper says batched decoding can touch nearly the full expert pool, making expert weights a bottleneck even when each token activates only a few experts. SlimWise reuses the full-model prefill KV cache directly in a pruned decode model, with an optional small distillation step to reduce remaining accuracy and generation-length distortions. In vLLM, the authors report up to 1.81x higher decode throughput on Qwen3.6-35B-A3B at 50% expert pruning with minimal accuracy loss. HF Daily Papers' note

score 5

Categories: Research