Megadose AI progress, ranked and analyzed.

PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?

· HF Daily Papers ·
Frontier video-language judges still miss a quarter of paired decisions on PlaylistEval.

The paper introduces PlaylistEval, an automated benchmark for testing judges over 100-hour playlist collections without human annotation. Its 630 paired-answer tests are designed so transcript-only shortcuts should not settle the choice. On a 152-pair subset, the benchmark matched human judgments 93.0% of the time. Across 17 models, the best judges reached 75.4% pairwise accuracy, with performance falling as playlist sets grew. HF Daily Papers' note

score 4

Categories: Research