Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
The paper’s evaluator checks GUI task success by interrogating the finished environment, not just screenshots.
IRA proposes completion conditions, then verifies them with system, application, and GUI tools after execution. The authors introduce GUI-RewardBench, covering 321 Ubuntu desktop task trajectories across 10 application categories. IRA reaches 86.9% accuracy on that benchmark and is used as a reward signal for GUI-agent reinforcement learning, where it reports a 34.0% OSWorld success rate. ArXiv · AI/CL/LG's note
IRA proposes completion conditions, then verifies them with system, application, and GUI tools after execution. The authors introduce GUI-RewardBench, covering 321 Ubuntu desktop task trajectories across 10 application categories. IRA reaches 86.9% accuracy on that benchmark and is used as a reward signal for GUI-agent reinforcement learning, where it reports a 34.0% OSWorld success rate. ArXiv · AI/CL/LG's note
score 5