Megadose AI progress, ranked and analyzed.

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models

· HF Daily Papers ·
The paper tests whether video-language models can handle 90-second clips with full paragraph descriptions.

CLIP-CC-Bench is built from 5 hours of movie content, with expert-written paragraph references for each clip. The authors evaluate 17 video-language models using coarse and fine semantic matching. They use an ensemble of five LLM-based embedding models to reduce reliance on a single judge. The paper also reports ranking stability and inter-judge agreement, and releases scripts, outputs, and aggregation tools. HF Daily Papers' note

score 4

Categories: Research