SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving
SlimWise keeps prefill unpruned and prunes only decode, aiming to cut MoE serving traffic without giving up much quality.
The paper says batched decoding can touch nearly the full expert pool, making expert weights a bottleneck even when each token activates only a few experts. SlimWise reuses the full-model prefill KV cache directly in a pruned decode model, with an optional small distillation step to reduce remaining accuracy and generation-length distortions. In vLLM, the authors report up to 1.81x higher decode throughput on Qwen3.6-35B-A3B at 50% expert pruning with minimal accuracy loss. HF Daily Papers' note
The paper says batched decoding can touch nearly the full expert pool, making expert weights a bottleneck even when each token activates only a few experts. SlimWise reuses the full-model prefill KV cache directly in a pruned decode model, with an optional small distillation step to reduce remaining accuracy and generation-length distortions. In vLLM, the authors report up to 1.81x higher decode throughput on Qwen3.6-35B-A3B at 50% expert pruning with minimal accuracy loss. HF Daily Papers' note
score 5