AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction
AcrossVAM1.0 separates robot motion from appearance, but its language signal is still weak.
The paper describes a lightweight video prediction model that uses semantic particles for the robot, arm, and gripper, then decodes future frames from the last observed image. On the authors’ VRS benchmark, particle dynamics cut trajectory error by 21.0% versus persistence and raised PSNR/SSIM to 20.573/0.8004 across delivery-mask seeds. The delivered model still does not beat persistence on LPIPS. Shuffling the language instruction changes trajectory error by only 2.8–3.1%, which the authors identify as a remaining grounding problem. ArXiv · AI/CL/LG's note
The paper describes a lightweight video prediction model that uses semantic particles for the robot, arm, and gripper, then decodes future frames from the last observed image. On the authors’ VRS benchmark, particle dynamics cut trajectory error by 21.0% versus persistence and raised PSNR/SSIM to 20.573/0.8004 across delivery-mask seeds. The delivered model still does not beat persistence on LPIPS. Shuffling the language instruction changes trajectory error by only 2.8–3.1%, which the authors identify as a remaining grounding problem. ArXiv · AI/CL/LG's note
score 4