Megadose AI progress, ranked daily.

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

· HF Daily Papers ·
Current video MLLMs still struggle to obey complex user constraints, even when they understand the video.

The paper introduces Video-IFBench, a 1.5K-sample benchmark for testing instruction following in video understanding. It covers single-task, multi-task, selection, and nested instructions across 32 task types and 39 constraint categories. The authors evaluated more than 20 recent multimodal LLMs and found the hardest cases involve many constraints, semantic requirements, and conditional paths tied to video content. HF Daily Papers' note

score 4

Categories: Research