Do Audio LLMs Listen Before They Act? Diagnosing Acoustic-Context Gating in Voice Agents
The benchmark finds voice agents often know the right tool, but fail to stay silent when the command is no longer addressed to them.
VGBench tests whether Audio LLMs gate actions across side-talk, self-talk, and speaker-switch cases. In speaker-switch pairs, six raw Audio LLMs and three training-free adaptations rarely withheld action; the best raw switch mute rate was 14%. A post-training VoxGate case study muted 91.3% of switched commands while preserving correct tool choice for nearby wearer commands and text-only controls. The authors frame the task as multi-cue acoustic-context gating, not just speaker identification. HF Daily Papers' note
VGBench tests whether Audio LLMs gate actions across side-talk, self-talk, and speaker-switch cases. In speaker-switch pairs, six raw Audio LLMs and three training-free adaptations rarely withheld action; the best raw switch mute rate was 14%. A post-training VoxGate case study muted 91.3% of switched commands while preserving correct tool choice for nearby wearer commands and text-only controls. The authors frame the task as multi-cue acoustic-context gating, not just speaker identification. HF Daily Papers' note
score 5