AREX: Towards a Recursively Self-Improving Agent for Deep Research
AREX uses verification as the driver for further research, not just as a final check.
The paper describes agents that build a provisional answer, audit it constraint by constraint, then launch targeted follow-up work on unresolved claims. Its context-update tool compresses long interaction history into a compact state that keeps verified evidence and open constraints. The authors trained dense 4B and 122B-A10B MoE versions with agentic mid-training and long-horizon reinforcement learning. They report gains over comparable-scale baselines across BrowseComp, WideSearch, DeepSearchQA, HLE, and other benchmarks. HF Daily Papers' note
The paper describes agents that build a provisional answer, audit it constraint by constraint, then launch targeted follow-up work on unresolved claims. Its context-update tool compresses long interaction history into a compact state that keeps verified evidence and open constraints. The authors trained dense 4B and 122B-A10B MoE versions with agentic mid-training and long-horizon reinforcement learning. They report gains over comparable-scale baselines across BrowseComp, WideSearch, DeepSearchQA, HLE, and other benchmarks. HF Daily Papers' note
score 5