AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
AVE-Compass tests whether video editors can change sound and image together without breaking what should stay untouched.
The benchmark uses 145 curated videos, 196 coupled audio-video editing instructions, and 2,688 checklist items. It scores instruction following, fidelity, realism, and editing intent with MLLM judging, a realism rubric, and automated audio, video, and cross-modal metrics. The authors report that current state-of-the-art models still struggle with cross-modal instructions and preserving non-target content. They also introduce AVE-Agent, a modular agent that decomposes edits into subtasks and improves results through self-reflection and evaluator feedback. HF Daily Papers' note
The benchmark uses 145 curated videos, 196 coupled audio-video editing instructions, and 2,688 checklist items. It scores instruction following, fidelity, realism, and editing intent with MLLM judging, a realism rubric, and automated audio, video, and cross-modal metrics. The authors report that current state-of-the-art models still struggle with cross-modal instructions and preserving non-target content. They also introduce AVE-Agent, a modular agent that decomposes edits into subtasks and improves results through self-reflection and evaluator feedback. HF Daily Papers' note
score 4