Megadose AI progress, ranked and analyzed.

AREX-2: Advancing Self-Improving Agents through Long-Horizon Reflective Tasks

· HF Daily Papers ·
AREX-2 trains an agent to improve its own answers over many test-time rounds, using supervised trajectories from ML and programming tasks.

The paper defines self-improvement as iterative refinement, with reflection producing better solutions and long-horizon execution keeping the loop useful. Its agent is built on Qwen3.8-27B and is trained on synthetic improvement trajectories with verifiable feedback. The authors report strong scores across MLE-bench Lite, Frontier-CS, BrowseComp, HLE, GAIA, and DeepSearchQA, with performance continuing to rise as more rounds are allowed. Code and models are listed as forthcoming. HF Daily Papers' note

score 5

Categories: Research