Megadose AI progress, ranked and analyzed.

Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

· ArXiv · AI/CL/LG ·
SESA turns failed self-play searches into reusable skills that alter both memory and later training.

The paper describes a challenger-solver loop in which the challenger poses problems and the solver retrieves procedural skills while searching. Failures judged informative are distilled into skills and written back to memory, changing what the solver can do next and what the challenger is rewarded for asking. Across seven open-domain and multi-hop QA benchmarks, SESA improves average accuracy over SSP by 1.2 to 3.2 points and beats SkillRL by 0.9 points under the authors' unified setup. The authors also report that Qwen3 runs retain gains without retrieval, with the final skill bank adding another 0.5 to 1.0 points at inference. ArXiv · AI/CL/LG's note

score 4

Categories: Research