OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
VLM judges are still too lenient to reliably grade computer-use agents at scale.
OSReward tests vision-language models on human-verified computer-use trajectories across platforms. The authors find even top models often mark failed runs as successes. The models that are reliable enough are too costly for broad use, while cheaper open models lag. They also release OS-Shepherd-100K and train OS-Shepherd reward models meant to match commercial judges at far lower cost. HF Daily Papers' note
OSReward tests vision-language models on human-verified computer-use trajectories across platforms. The authors find even top models often mark failed runs as successes. The models that are reliable enough are too costly for broad use, while cheaper open models lag. They also release OS-Shepherd-100K and train OS-Shepherd reward models meant to match commercial judges at far lower cost. HF Daily Papers' note
score 5