PlaylistEval: Can Video-Language Judges Be Trusted at Day Scale and Beyond?
Frontier video-language judges still miss a quarter of paired decisions on PlaylistEval.
The paper introduces PlaylistEval, an automated benchmark for testing judges over 100-hour playlist collections without human annotation. Its 630 paired-answer tests are designed so transcript-only shortcuts should not settle the choice. On a 152-pair subset, the benchmark matched human judgments 93.0% of the time. Across 17 models, the best judges reached 75.4% pairwise accuracy, with performance falling as playlist sets grew. HF Daily Papers' note
The paper introduces PlaylistEval, an automated benchmark for testing judges over 100-hour playlist collections without human annotation. Its 630 paired-answer tests are designed so transcript-only shortcuts should not settle the choice. On a 152-pair subset, the benchmark matched human judgments 93.0% of the time. Across 17 models, the best judges reached 75.4% pairwise accuracy, with performance falling as playlist sets grew. HF Daily Papers' note
score 4