Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
AnTrap finds that leading Android GUI agents break down when realistic runtime traps are injected into their task flows.
The benchmark adds adversarial anomalies such as unexpected pop-ups, action misuse, and state deadlocks while keeping tasks solvable. Its taxonomy spans State, Thinking, Action, and Round layers, with ten subcategories. Tests on 16 GUI models showed broad performance degradation, including for the strongest systems. GRPO training helped with some single-step state and action traps, but deeper contextual failures remained unresolved. HF Daily Papers' note
The benchmark adds adversarial anomalies such as unexpected pop-ups, action misuse, and state deadlocks while keeping tasks solvable. Its taxonomy spans State, Thinking, Action, and Round layers, with ten subcategories. Tests on 16 GUI models showed broad performance degradation, including for the strongest systems. GRPO training helped with some single-step state and action traps, but deeper contextual failures remained unresolved. HF Daily Papers' note
score 4