Marionette: Predicting World States, Rendering Geometry, Painting Appearance
Marionette separates game-world prediction into explicit 3D state, fixed geometry, and neural appearance.
The paper describes a model that predicts a 276-dimensional state for articulated game characters, then uses a zero-parameter renderer to compute pose, geometry, and occlusion before a video-diffusion model paints RGB frames. Its tests show the explicit state can be controlled: mismatched action streams changed root-aligned joint error by 31% across 48 held-out segments. The authors also report that long-horizon failures can be corrected at the state level, with terrain and separation rules cutting ground penetration by 66% and keeping two characters closer together. They say routing generation through this structured state did not produce a detectable fidelity cost, with FVD 831 versus 799 for recorded pose. ArXiv · AI/CL/LG's note
The paper describes a model that predicts a 276-dimensional state for articulated game characters, then uses a zero-parameter renderer to compute pose, geometry, and occlusion before a video-diffusion model paints RGB frames. Its tests show the explicit state can be controlled: mismatched action streams changed root-aligned joint error by 31% across 48 held-out segments. The authors also report that long-horizon failures can be corrected at the state level, with terrain and separation rules cutting ground penetration by 66% and keeping two characters closer together. They say routing generation through this structured state did not produce a detectable fidelity cost, with FVD 831 versus 799 for recorded pose. ArXiv · AI/CL/LG's note
score 6