AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
AdvSim2Real trains a web agent in a frozen simulator where tasks, attacks, and the agent adapt together.
The paper targets prompt injection on web pages that agents still need to read to finish a task. Its setup rewards the curriculum for tasks the agent only partly solves, and rewards the adversary only when an injected instruction flips a success into a failure. The authors report that a 4B agent improves both normal task completion and robustness, including against a frontier-model adversary it did not train against. On 150 web tasks, completion under that unseen adversary rose 33.6% relative to the base agent. ArXiv · AI/CL/LG's note
The paper targets prompt injection on web pages that agents still need to read to finish a task. Its setup rewards the curriculum for tasks the agent only partly solves, and rewards the adversary only when an injected instruction flips a success into a failure. The authors report that a 4B agent improves both normal task completion and robustness, including against a frontier-model adversary it did not train against. On 150 web tasks, completion under that unseen adversary rose 33.6% relative to the base agent. ArXiv · AI/CL/LG's note
score 6