OmniEcho: Spatial Audio Understanding for Embodied Agents
OmniEcho pairs a new spatial audio-visual benchmark with a model built to use first-order ambisonics for embodied navigation.
The paper introduces OmniEchoBench, covering six tasks across 197 real-world scenes, 2,972 QA pairs, and 900 navigation samples. Its rendering pipeline is meant to keep sound sources, visual observations, and agent paths geometrically consistent. OmniEcho adds an FOA spatial encoder beside a pretrained semantic audio pathway, and the authors report state-of-the-art spatial audio-visual perception results. Sound-guided navigation approaches traditional vision-language navigation performance, though fine-grained localization and distance estimation remain open problems. HF Daily Papers' note
The paper introduces OmniEchoBench, covering six tasks across 197 real-world scenes, 2,972 QA pairs, and 900 navigation samples. Its rendering pipeline is meant to keep sound sources, visual observations, and agent paths geometrically consistent. OmniEcho adds an FOA spatial encoder beside a pretrained semantic audio pathway, and the authors report state-of-the-art spatial audio-visual perception results. Sound-guided navigation approaches traditional vision-language navigation performance, though fine-grained localization and distance estimation remain open problems. HF Daily Papers' note
score 5