OSWorld-Pro: Process-based Evaluation for Computer Use Agents
OSWorld-Pro grades agents on the steps they take, not just the finished task.
The benchmark adds more than 300 computer-use tasks, split into over 2,800 subgoals and backed by over 67,000 human annotations. Its LLM judges score whether agents complete those intermediate subgoals, exposing where a run breaks down. The paper says top systems still struggle: Claude Opus 5 reaches 75.7% on OSWorld-Pro, below its 83.4% on OSWorld. The authors point to failures such as irrelevant actions and click mistakes as targets for improving computer-use agents. HF Daily Papers' note
The benchmark adds more than 300 computer-use tasks, split into over 2,800 subgoals and backed by over 67,000 human annotations. Its LLM judges score whether agents complete those intermediate subgoals, exposing where a run breaks down. The paper says top systems still struggle: Claude Opus 5 reaches 75.7% on OSWorld-Pro, below its 83.4% on OSWorld. The authors point to failures such as irrelevant actions and click mistakes as targets for improving computer-use agents. HF Daily Papers' note
score 6