DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment
A 9B deep-search agent trained from its own filtered trajectories reaches competitive benchmark scores without a stronger teacher model.
The paper introduces DeepSearch-World, a deterministic environment with reproducible search and page-reading tools for multi-hop QA. Its 420K-task pool is built from entity-level random walks and is designed to support verification, reflection, and recovery during agent training. DeepSearch-Evolve uses that setup to generate, filter, mix, and fine-tune trajectories iteratively. The authors report 31.2% on BrowseComp, 61.5% on GAIA, and 93.4% on HotpotQA, and say they will release the environment, data, model, and code. HF Daily Papers' note
The paper introduces DeepSearch-World, a deterministic environment with reproducible search and page-reading tools for multi-hop QA. Its 420K-task pool is built from entity-level random walks and is designed to support verification, reflection, and recovery during agent training. DeepSearch-Evolve uses that setup to generate, filter, mix, and fine-tune trajectories iteratively. The authors report 31.2% on BrowseComp, 61.5% on GAIA, and 93.4% on HotpotQA, and say they will release the environment, data, model, and code. HF Daily Papers' note
score 6