Megadose AI progress, ranked and analyzed.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

· HF Daily Papers ·
UltraViT targets the vision encoder itself as the on-device latency bottleneck in LVLMs.

The paper proposes a pyramidal vision encoder designed around measured device latency, with heterogeneous spatial mixers selected at the macro-block level. Its training uses dense distillation first, then generative supervision from a frozen, capacity-mixed LLM. The authors say this beats encoder-focused baselines while running nearly 1.7x faster on device. Accepted at ECCV 2026. HF Daily Papers' note

score 4

Categories: Research