RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments
The paper’s claim is that agents can improve in unfamiliar software environments by building reusable memory, without retraining the model.
RSIAgent uses separate curriculum, actor, and verifier agents to explore an environment, check outcomes, and store environment-specific causal knowledge. Its “broad-then-deep” strategy first maps diverse structures, then probes harder cases and hidden constraints. The authors report gains on OSWorld-v2 and Agent’s Last Exam, including open-source models outperforming GPT-6 in their experiments. HF Daily Papers' note
RSIAgent uses separate curriculum, actor, and verifier agents to explore an environment, check outcomes, and store environment-specific causal knowledge. Its “broad-then-deep” strategy first maps diverse structures, then probes harder cases and hidden constraints. The authors report gains on OSWorld-v2 and Agent’s Last Exam, including open-source models outperforming GPT-6 in their experiments. HF Daily Papers' note
score 5