AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control
The paper says planning models can miss the action differences that matter, even when their factual predictions look good.
AD-WM adds action-recovery regularization to a joint-embedding world model so counterfactual rollouts keep information about which action caused them. The auxiliary heads are removed at test time, so the MPC procedure itself is left unchanged. In OGBench-Cube, the authors report hard-start success rising from 3.7% to 52.0% against a matched LeWM baseline. They also report better zero-shot transfer on a Franka pick-and-place setup, from 42.2% to 71.1%, without lab-specific adaptation. ArXiv · AI/CL/LG's note
AD-WM adds action-recovery regularization to a joint-embedding world model so counterfactual rollouts keep information about which action caused them. The auxiliary heads are removed at test time, so the MPC procedure itself is left unchanged. In OGBench-Cube, the authors report hard-start success rising from 3.7% to 52.0% against a matched LeWM baseline. They also report better zero-shot transfer on a Franka pick-and-place setup, from 42.2% to 71.1%, without lab-specific adaptation. ArXiv · AI/CL/LG's note
score 5