WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning
The paper’s critic is trained to model future latent robot states, not just regress scalar rewards.
WCM uses a lightweight LeJEPA-style architecture to give VLA reinforcement learning a critic that can track temporal dynamics under partial observability. The authors say it works with on-policy and off-policy pipelines and with Pi0, Pi0.5, and OpenVLA-OFT backbones. They report state-of-the-art results across 149 tasks on four benchmarks, including out-of-distribution gains, plus validation on seven real-world manipulation tasks. HF Daily Papers' note
WCM uses a lightweight LeJEPA-style architecture to give VLA reinforcement learning a critic that can track temporal dynamics under partial observability. The authors say it works with on-policy and off-policy pipelines and with Pi0, Pi0.5, and OpenVLA-OFT backbones. They report state-of-the-art results across 149 tasks on four benchmarks, including out-of-distribution gains, plus validation on seven real-world manipulation tasks. HF Daily Papers' note
score 5