Megadose Built for builders and researchers.

OSWorld-Pro: Process-based Evaluation for Computer Use Agents

· HF Daily Papers ·
OSWorld-Pro grades agents on the steps they take, not just the finished task.

The benchmark adds more than 300 computer-use tasks, split into over 2,800 subgoals and backed by over 67,000 human annotations. Its LLM judges score whether agents complete those intermediate subgoals, exposing where a run breaks down. The paper says top systems still struggle: Claude Opus 5 reaches 75.7% on OSWorld-Pro, below its 83.4% on OSWorld. The authors point to failures such as irrelevant actions and click mistakes as targets for improving computer-use agents. HF Daily Papers' note

score 6

Categories: Research