FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling
FIRM-Video scores text-to-video outputs only after checklist items are verified against visible temporal evidence.
The paper frames the problem as unreliable supervision for video reward models, where holistic judges can miss details or justify scores loosely. Its framework builds separate checks for instruction following, world coherence, and perceptual quality, then aggregates verified decisions. The authors report a 90K-instance dataset from 29,348 videos and a 750-annotation benchmark. Their Qwen3-VL-based 8B model posts the best overall MAE on that benchmark and leads VBench scores in Best-of-8 sampling across three generators. HF Daily Papers' note
The paper frames the problem as unreliable supervision for video reward models, where holistic judges can miss details or justify scores loosely. Its framework builds separate checks for instruction following, world coherence, and perceptual quality, then aggregates verified decisions. The authors report a 90K-instance dataset from 29,348 videos and a 750-annotation benchmark. Their Qwen3-VL-based 8B model posts the best overall MAE on that benchmark and leads VBench scores in Best-of-8 sampling across three generators. HF Daily Papers' note
score 4