MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
MMAC tests whether audio models cover the right details and avoid unsupported ones.
The benchmark uses 5,638 audio clips from more than 20 sources, split across 6 capability categories and 15 evaluation dimensions. It scores a generated caption by checking both whether the relevant dimension is mentioned and whether that mention matches the reference label. The authors report clear gaps between AudioLLMs across coverage and reliability, and say they will release the benchmark and evaluation code. ArXiv · AI/CL/LG's note
The benchmark uses 5,638 audio clips from more than 20 sources, split across 6 capability categories and 15 evaluation dimensions. It scores a generated caption by checking both whether the relevant dimension is mentioned and whether that mention matches the reference label. The authors report clear gaps between AudioLLMs across coverage and reliability, and say they will release the benchmark and evaluation code. ArXiv · AI/CL/LG's note
score 5