Megadose AI progress, ranked and analyzed.

Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging

· ArXiv · AI/CL/LG ·
Mira claims a 5.71x average throughput speedup for single-GPU MoE inference under tight memory limits.

The paper says MoE inference is held back by expert parameters crowding VRAM and by token routing that is hard to predict. Mira tries to move expert handling from reactive offload-and-cache behavior to proactive staging, using per-layer predictors to anticipate expert use two layers ahead. It keeps frequent experts in a HOT cache while staging predicted ones, and uses a custom compressed expert format to reduce transfer cost with minimal accuracy loss. The reported gains include 11.71x faster time-to-first-token and 3.84x average speedup for beam search inference. ArXiv · AI/CL/LG's note

score 5

Categories: Research