When the World Lies: Backdoor Attacks on Latent World Models for Downstream Control
A poisoned world-model checkpoint can make a clean downstream controller choose attacker-targeted actions when a trigger appears.
The paper describes a supply-chain backdoor in reused latent world models for control. The victim trains and evaluates on clean data, but the poisoned model redirects trigger-bearing observations into a manipulated latent region. In the strongest settings, the trigger hijacks 100% of triggered steps while clean-task success stays at least about 75%. The effect stops when the trigger is removed, and clean fine-tuning may not remove it unless pushed hard enough to damage control performance. ArXiv · AI/CL/LG's note
The paper describes a supply-chain backdoor in reused latent world models for control. The victim trains and evaluates on clean data, but the poisoned model redirects trigger-bearing observations into a manipulated latent region. In the strongest settings, the trigger hijacks 100% of triggered steps while clean-task success stays at least about 75%. The effect stops when the trigger is removed, and clean fine-tuning may not remove it unless pushed hard enough to damage control performance. ArXiv · AI/CL/LG's note
score 5