WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
The paper’s claim is that VLA critics need an explicit world-modeling objective to estimate value well under partial observability.
WCM adds future-latent prediction to the critic, instead of training it only to regress scalar returns from single-frame observations. The authors say it plugs into on-policy and off-policy RL pipelines and works with Pi0, Pi0.5, and OpenVLA-OFT. They report state-of-the-art results across 149 tasks on four benchmarks, with stronger gains out of distribution. They also tested it on seven real-world manipulation tasks using OpenVLA-OFT and Pi0.5. ArXiv · AI/CL/LG's note
WCM adds future-latent prediction to the critic, instead of training it only to regress scalar returns from single-frame observations. The authors say it plugs into on-policy and off-policy RL pipelines and works with Pi0, Pi0.5, and OpenVLA-OFT. They report state-of-the-art results across 149 tasks on four benchmarks, with stronger gains out of distribution. They also tested it on seven real-world manipulation tasks using OpenVLA-OFT and Pi0.5. ArXiv · AI/CL/LG's note
score 5