VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
The paper replaces stitched-together quality metrics with a reward model trained on human video-audio preferences.
VA-Judger is built around VAPref-10K, a dataset of 9K prompts and 10.3K paired human comparisons from open-source generation models. The authors argue that separate scores for audio, visuals, and sync can miss whether the prompt, video, and sound actually cohere for viewers. Their benchmark tests whether reward models match human preferences both in-domain and out-of-domain. They report that VA-Judger beats metric baselines and improves post-training quality for audio-video generation models. HF Daily Papers' note
VA-Judger is built around VAPref-10K, a dataset of 9K prompts and 10.3K paired human comparisons from open-source generation models. The authors argue that separate scores for audio, visuals, and sync can miss whether the prompt, video, and sound actually cohere for viewers. Their benchmark tests whether reward models match human preferences both in-domain and out-of-domain. They report that VA-Judger beats metric baselines and improves post-training quality for audio-video generation models. HF Daily Papers' note
score 4