OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
The paper says VLM judges for computer-use agents are often too lenient, marking failed trajectories as successful.
OSReward is a new benchmark for testing reward models that judge whether computer-use agents completed tasks across platforms. Its authors say strong commercial judges are more reliable but too costly for large-scale use, while cheaper open models lag. They also release OS-Shepherd-100K and train OS-Shepherd 9B and 35B models, which they report match commercial judges at 30-60% lower cost.
ArXiv · AI/CL/LG's note
OSReward is a new benchmark for testing reward models that judge whether computer-use agents completed tasks across platforms. Its authors say strong commercial judges are more reliable but too costly for large-scale use, while cheaper open models lag. They also release OS-Shepherd-100K and train OS-Shepherd 9B and 35B models, which they report match commercial judges at 30-60% lower cost.
ArXiv · AI/CL/LG's note
score 5