Megadose AI progress, ranked and analyzed.

Learning from Research: Toward Lifelong Agent Harness Evolution

· ArXiv · AI/CL/LG ·
ScholarEvolve uses research papers as input for improving an agent’s harness while leaving the model fixed.

The framework groups harness changes by functional module, identifies improvement strategies with topic modeling, then tests combinations of those strategies. It is built to absorb new publications over time so harness updates can be driven proactively rather than only by observed failures. In experiments, it improved AppWorld Challenge completion for Qwen3.5-27B from 49.6% to 63.6%, and Tau2-Bench Telecom pass@1 for GPT-5.4-mini from 72.7% to 81.9%. ArXiv · AI/CL/LG's note

score 5

Categories: Research