Megadose AI progress, ranked daily.

FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling

· HF Daily Papers ·
FIRM-Video scores text-to-video outputs only after checklist items are verified against visible temporal evidence.

The paper frames the problem as unreliable supervision for video reward models, where holistic judges can miss details or justify scores loosely. Its framework builds separate checks for instruction following, world coherence, and perceptual quality, then aggregates verified decisions. The authors report a 90K-instance dataset from 29,348 videos and a 750-annotation benchmark. Their Qwen3-VL-based 8B model posts the best overall MAE on that benchmark and leads VBench scores in Best-of-8 sampling across three generators. HF Daily Papers' note

score 4

Categories: Research