Megadose AI progress, ranked and analyzed.

cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

· ArXiv · AI/CL/LG ·
The paper argues CUA speed benchmarks are too infrastructure-dependent to compare cleanly.

The authors introduce cua-speedrun, a standardized VM, execution pipeline, and agent interface for testing computer-use agents across benchmarks. They evaluate four CUA benchmarks and find tradeoffs among performance, speed, and cost, with no model family leading on all three. They also report counterintuitive effects: more reasoning can sometimes shorten task completion, and faster environment I/O can slow overall completion. The paper says benchmark task sets can often be reduced without losing statistical power. ArXiv · AI/CL/LG's note

score 5

Categories: Research