Megadose Built for builders and researchers.

OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning

· HF Daily Papers ·
OmniCapBench turns audio-visual captioning into atomic checks for entities, shots, and sound events.

The benchmark uses 786 densely annotated videos to score fine-grained caption behavior with deterministic constraints and localized semantic comparisons. The authors say it can separate failures such as temporal grounding errors, identity drift, cross-modal mismatch, and hallucinated descriptions. In their evaluation, frontier MLLMs showed strong local perception but weaker long-horizon audio-visual reasoning. HF Daily Papers' note

score 4

Categories: Research