Wonder: Video World Model Done Better
Wonder turns a single image or video into a real-time, camera-controllable scene.
The paper describes a video world model that lets users move through generated space, reveal unseen areas, and return to earlier views over long rollouts. Its control system uses dense camera-coordinate conditioning so motion commands are treated as spatial visual cues. A sparse attention memory mechanism keeps retrieval fast as context grows. The authors report minute-scale generation at 16 FPS with coherent geometry, appearance, and dynamics, including real-time re-shooting of existing video scenes. HF Daily Papers' note
The paper describes a video world model that lets users move through generated space, reveal unseen areas, and return to earlier views over long rollouts. Its control system uses dense camera-coordinate conditioning so motion commands are treated as spatial visual cues. A sparse attention memory mechanism keeps retrieval fast as context grows. The authors report minute-scale generation at 16 FPS with coherent geometry, appearance, and dynamics, including real-time re-shooting of existing video scenes. HF Daily Papers' note
score 6