AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks
AREX-2 trains an agent to improve its own answers over many test-time rounds, using supervised trajectories from ML and programming tasks.
The paper defines self-improvement as iterative refinement, with reflection producing better solutions and long-horizon execution keeping the loop useful. Its agent is built on Qwen3.8-27B and is trained on synthetic improvement trajectories with verifiable feedback. The authors report strong scores across MLE-bench Lite, Frontier-CS, BrowseComp, HLE, GAIA, and DeepSearchQA, with performance continuing to rise as more rounds are allowed. Code and models are listed as forthcoming. HF Daily Papers' note
The paper defines self-improvement as iterative refinement, with reflection producing better solutions and long-horizon execution keeping the loop useful. Its agent is built on Qwen3.8-27B and is trained on synthetic improvement trajectories with verifiable feedback. The authors report strong scores across MLE-bench Lite, Frontier-CS, BrowseComp, HLE, GAIA, and DeepSearchQA, with performance continuing to rise as more rounds are allowed. Code and models are listed as forthcoming. HF Daily Papers' note
score 5