Megadose AI progress, ranked and analyzed.

SceneBind: Binding What and Where Across Vision, Audio and Language

· ArXiv · AI/CL/LG ·
SceneBind adds explicit 3D “where” information to cross-modal scene representations.

The paper represents scenes with a global semantic embedding plus object-centered semantic-spatial slots. Its matching method combines overall scene similarity with object alignment for retrieval and grounding across vision, binaural audio, and language. The authors also introduce a real-world audio-visual dataset with structured semantic and spatial annotations. They report state-of-the-art scene and spatial retrieval, with zero-shot transfer to audio-visual localization. ArXiv · AI/CL/LG's note

score 5

Categories: Research