MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
The paper introduces a 350-sample benchmark for judging whether MR2AV systems can keep multiple references bound correctly while generating synchronized sound and video.
MultiRef-Compass targets multi-view subject preservation, multi-entity binding, and human-object-scene composition. Its evaluation protocol covers Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following across 14 sub-metrics. The authors combine automatic metrics with a rejudging-enhanced MLLM-as-a-Judge setup. Tests on eight MR2AV systems found substantial room for improvement. HF Daily Papers' note
MultiRef-Compass targets multi-view subject preservation, multi-entity binding, and human-object-scene composition. Its evaluation protocol covers Basic Quality, Reference Consistency, Audio-Visual Consistency, and Instruction Following across 14 sub-metrics. The authors combine automatic metrics with a rejudging-enhanced MLLM-as-a-Judge setup. Tests on eight MR2AV systems found substantial room for improvement. HF Daily Papers' note
score 4