Megadose AI progress, ranked and analyzed.

The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding

· ArXiv · AI/CL/LG ·
Sparse-frame video grounding failed mainly in the vision stack, not the language model.

The paper reports Qwen3-VL 8B falling from 56.0% to 22.3% temporal mIoU when video input is cut to 16 frames. Fine-tuning only the final three ViT layers, about 4% of parameters, reached 68.8% temporal mIoU. Language-model fine-tuning gave little or negative benefit in the study. With sparse inputs, a fine-tuned 2B model beat a zero-shot 8B model even when the larger model had dense-frame access. ArXiv · AI/CL/LG's note

score 5

Categories: Research