OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
OmniCapBench turns audio-visual captioning into atomic checks for entities, shots, and sound events.
The benchmark uses 786 densely annotated videos to score fine-grained caption behavior with deterministic constraints and localized semantic comparisons. The authors say it can separate failures such as temporal grounding errors, identity drift, cross-modal mismatch, and hallucinated descriptions. In their evaluation, frontier MLLMs showed strong local perception but weaker long-horizon audio-visual reasoning. HF Daily Papers' note
The benchmark uses 786 densely annotated videos to score fine-grained caption behavior with deterministic constraints and localized semantic comparisons. The authors say it can separate failures such as temporal grounding errors, identity drift, cross-modal mismatch, and hallucinated descriptions. In their evaluation, frontier MLLMs showed strong local perception but weaker long-horizon audio-visual reasoning. HF Daily Papers' note
score 4