AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
AdvSim2Real trains the agent, the task generator, and the injection attacker together inside a frozen web simulator.
The paper says fixed prompt-injection defenses fail when attackers adapt to the trained model. Its setup rewards tasks that are neither too easy nor too hard, and rewards attacks only when they flip a successful run into failure. In tests on 150 web tasks, the method lifted completion under an unseen frontier-model adversary by 33.6% relative to the base agent. The authors also report that the capability gain carried over to a real browser. HF Daily Papers' note
The paper says fixed prompt-injection defenses fail when attackers adapt to the trained model. Its setup rewards tasks that are neither too easy nor too hard, and rewards attacks only when they flip a successful run into failure. In tests on 150 web tasks, the method lifted completion under an unseen frontier-model adversary by 33.6% relative to the base agent. The authors also report that the capability gain carried over to a real browser. HF Daily Papers' note
score 5