PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation
PhysWave adds acoustic constraints to text-to-FOA generation so generated sound better matches source direction and distance.
The model combines natural-language prompts and trajectory control through a shared waypoint-caption representation. Its diffusion training uses spherical-harmonic direction consistency and inverse-square distance consistency as differentiable priors. The authors also built a 300K-clip FOA dataset covering varied sound categories and source trajectories. They report improved spatial consistency while keeping competitive audio quality, with the priors also usable at inference time for training-free refinement. ArXiv · AI/CL/LG's note
The model combines natural-language prompts and trajectory control through a shared waypoint-caption representation. Its diffusion training uses spherical-harmonic direction consistency and inverse-square distance consistency as differentiable priors. The authors also built a 300K-clip FOA dataset covering varied sound categories and source trajectories. They report improved spatial consistency while keeping competitive audio quality, with the priors also usable at inference time for training-free refinement. ArXiv · AI/CL/LG's note
score 5