Megadose Built for builders and researchers.

Can Agents Work for Everyone? Cross-User Reliability for Mobile GUI Agents in Personalized User Interfaces

· ArXiv · AI/CL/LG ·
Personalized mobile UIs measurably break GUI-agent reliability across users.

The paper introduces PAIR to test the same mobile task under different user-conditioned app states, then RePAIR to train on those cross-user differences. Across six agents, subgoal achievement fell by 6.98 to 15.4 percentage points in user-conditioned UI contexts, and by 8.77 to 22.0 points when targets came from a user’s own content. The common failure was choosing the wrong item, especially before the intended target was exposed. RePAIR improved success metrics on unseen users over its supervised fine-tuning parent, including +9.42 points in overall Task SR. ArXiv · AI/CL/LG's note

score 5

Categories: Research