Megadose AI progress, ranked and analyzed.

Rethinking Inference-Time Scaling in Local Computer-Use Agents: Failure Modes and Compute Tradeoffs

· ArXiv · AI/CL/LG ·
More inference compute did not reliably make local computer-use agents better.

The paper tests several scaling routes on local CUAs, including more context, longer execution, task decomposition, and parallel runs. Extra context helped stabilize trajectories, but gains saturated as token cost rose and failures shifted toward premature false successes. Longer horizons reduced max-step stalls without meaningfully raising task success, suggesting bad trajectories were often just extended. Structural decomposition added planning and formatting overhead, while parallel scaling helped only at substantial compute cost. ArXiv · AI/CL/LG's note

score 5

Categories: Research