Megadose AI progress, ranked and analyzed.

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

· ArXiv · AI/CL/LG ·
The benchmark breaks computer-use tasks into subgoals so failures can be judged step by step, not just by the final output.

OSWorld-Pro includes more than 300 tasks, over 2,800 subgoals, and more than 67,000 human annotations. The authors use human-aligned LLM judges to assess whether each subgoal is fulfilled during a run. They report that Claude Opus 5 reaches 75.7% on OSWorld-Pro, compared with 83.4% on OSWorld. The paper says the process view exposes failures such as irrelevant actions and click-based mistakes. ArXiv · AI/CL/LG's note

score 6

Categories: Research