Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
A vision-language agent turns real robot-object recordings into runnable physics simulations.
The paper says Agentic Real2Sim recovers scene geometry, object states, physical parameters, cameras, poses, actors, and trajectories from real-world interaction footage. It targets a process the authors describe as still dependent on manual tuning, mesh cleanup, coordinate alignment, and fragile tool chains. They test it across rigid manipulation, deformable-object interaction, and humanoid motion scenes. The authors say open-weight VLMs can drive the agentic decisions at much lower cost than frontier models while reaching comparable conversion success. ArXiv · AI/CL/LG's note
The paper says Agentic Real2Sim recovers scene geometry, object states, physical parameters, cameras, poses, actors, and trajectories from real-world interaction footage. It targets a process the authors describe as still dependent on manual tuning, mesh cleanup, coordinate alignment, and fragile tool chains. They test it across rigid manipulation, deformable-object interaction, and humanoid motion scenes. The authors say open-weight VLMs can drive the agentic decisions at much lower cost than frontier models while reaching comparable conversion success. ArXiv · AI/CL/LG's note
score 5