UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
UltraViT targets the vision encoder itself as the on-device latency bottleneck in LVLMs.
The paper proposes a pyramidal vision encoder designed around measured device latency, with heterogeneous spatial mixers selected at the macro-block level. Its training uses dense distillation first, then generative supervision from a frozen, capacity-mixed LLM. The authors say this beats encoder-focused baselines while running nearly 1.7x faster on device. Accepted at ECCV 2026. HF Daily Papers' note
The paper proposes a pyramidal vision encoder designed around measured device latency, with heterogeneous spatial mixers selected at the macro-block level. Its training uses dense distillation first, then generative supervision from a frozen, capacity-mixed LLM. The authors say this beats encoder-focused baselines while running nearly 1.7x faster on device. Accepted at ECCV 2026. HF Daily Papers' note
score 4