BadWAM: When World-Action Models Dream Right but Act Wrong
Small visual perturbations can make a world-action model keep “imagining” the right future while executing the wrong action.
The paper introduces BadWAM, a framework for testing World-Action Drift Attacks against world-action models. One attack directly pushes the model toward task-failing actions; another tries to keep the predicted future close to the clean version while shifting the action. In closed-loop tests, the action-only attack cut success from 96.5% to 43.1%. The authors argue this exposes a WAM-specific failure mode: future prediction may look plausible even when control has been desynchronized. HF Daily Papers' note
The paper introduces BadWAM, a framework for testing World-Action Drift Attacks against world-action models. One attack directly pushes the model toward task-failing actions; another tries to keep the predicted future close to the clean version while shifting the action. In closed-loop tests, the action-only attack cut success from 96.5% to 43.1%. The authors argue this exposes a WAM-specific failure mode: future prediction may look plausible even when control has been desynchronized. HF Daily Papers' note
score 4