SceneBind: Binding What and Where Across Vision, Audio and Language
SceneBind adds explicit 3D “where” information to cross-modal scene representations.
The paper represents scenes with a global semantic embedding plus object-centered semantic-spatial slots. Its matching method combines overall scene similarity with object alignment for retrieval and grounding across vision, binaural audio, and language. The authors also introduce a real-world audio-visual dataset with structured semantic and spatial annotations. They report state-of-the-art scene and spatial retrieval, with zero-shot transfer to audio-visual localization. ArXiv · AI/CL/LG's note
The paper represents scenes with a global semantic embedding plus object-centered semantic-spatial slots. Its matching method combines overall scene similarity with object alignment for retrieval and grounding across vision, binaural audio, and language. The authors also introduce a real-world audio-visual dataset with structured semantic and spatial annotations. They report state-of-the-art scene and spatial retrieval, with zero-shot transfer to audio-visual localization. ArXiv · AI/CL/LG's note
score 5