Megadose Built for builders and researchers.

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

· HF Daily Papers ·
MoE-ViE claims fine-grained expert routing can scale CLIP-style vision encoders with lower latency than dense growth.

The paper studies mixture-of-experts designs for image and video encoders, then proposes an auxiliary-loss-free balancing method and a custom MoE kernel. Its largest model is reported to match a state-of-the-art encoder 1.7x its size while running at 76% of that model’s latency. The authors also add frame-level distillation and a freezing mechanism to improve video ability without losing image performance. When paired with an LLM, MoE-ViE beats the compared encoders on image and video benchmarks, including models with up to 5x more activated parameters. HF Daily Papers' note

score 5

Categories: Research