EchoWM: Open and Enterable Omnimodal World Models
EchoWM is presented as a navigable world model that generates video and synchronized audio together.
The paper says it responds to continuous camera navigation while producing 720p video, environmental sound, music, and speech. It maps both discrete commands and continuous poses into a shared metric-scale 6-DoF trajectory. The authors report strong trajectory following, high visual quality, and support for both first- and third-person interaction. HF Daily Papers' note
The paper says it responds to continuous camera navigation while producing 720p video, environmental sound, music, and speech. It maps both discrete commands and continuous poses into a shared metric-scale 6-DoF trajectory. The authors report strong trajectory following, high visual quality, and support for both first- and third-person interaction. HF Daily Papers' note
score 6