GUI-HARVEST: Self-Improving GUI Agents through Evidence-Driven Harness Evolution
The paper says GUI agents can improve without changing the model, by rewriting the harness around it.
GUI-HARVEST compares intended actions with before-and-after screenshots, then uses repeated task runs to isolate which behavior changed outcomes. It turns recurring failures into bounded harness code edits and tests the predicted effects through repeated execution. On OSWorld-Verified, the authors report held-out gains across six backbone models, including a 12.33-point gain for Qwen3-VL-32B-Instruct. They also report frozen-harness transfer improving GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps. HF Daily Papers' note
GUI-HARVEST compares intended actions with before-and-after screenshots, then uses repeated task runs to isolate which behavior changed outcomes. It turns recurring failures into bounded harness code edits and tests the predicted effects through repeated execution. On OSWorld-Verified, the authors report held-out gains across six backbone models, including a 12.33-point gain for Qwen3-VL-32B-Instruct. They also report frozen-harness transfer improving GPT-5 by 13.87 percentage points on WindowsAgentArena at 50 steps. HF Daily Papers' note
score 5