WorldSonus: Bringing Sound to Worlds
WorldSonus is built to generate spatial stereo audio fast enough to keep up with interactive world-model video.
The paper says the system uses a streaming causal autoregressive diffusion design, reporting a real-time factor of 0.41 for chunked audio generation. It adds chunk-indexed prompt scheduling so sound events can be changed mid-stream. The authors also train for spatial alignment using stereo and ambisonic data, aiming for audio that tracks scene geometry and camera motion. They report results competitive with or better than state-of-the-art bidirectional video-to-audio models on quality and spatial alignment. ArXiv · AI/CL/LG's note
The paper says the system uses a streaming causal autoregressive diffusion design, reporting a real-time factor of 0.41 for chunked audio generation. It adds chunk-indexed prompt scheduling so sound events can be changed mid-stream. The authors also train for spatial alignment using stereo and ambisonic data, aiming for audio that tracks scene geometry and camera motion. They report results competitive with or better than state-of-the-art bidirectional video-to-audio models on quality and spatial alignment. ArXiv · AI/CL/LG's note
score 6