Shockingly Simple Self-retrospection Improves Agentic Models Without RL
Training on an agent’s own explanations improved software-task performance without a reward update.
The paper tests Retrospection-Only Fine-Tuning, where an agent attempts a task, writes a retrospective explanation from feedback, and is fine-tuned only on those explanation tokens. On held-out SWE-bench Verified and Pro, Qwen3.5-4B reached 49.2% and 26.8% solve rates after 20 updates in the reported runs. The authors say it outperformed the evaluated GRPO runs while using no verifier and fewer updates. They also report gains on tasks where all 64 base-model attempts initially failed, suggesting the method can learn without successful starting trajectories. ArXiv · AI/CL/LG's note
The paper tests Retrospection-Only Fine-Tuning, where an agent attempts a task, writes a retrospective explanation from feedback, and is fine-tuned only on those explanation tokens. On held-out SWE-bench Verified and Pro, Qwen3.5-4B reached 49.2% and 26.8% solve rates after 20 updates in the reported runs. The authors say it outperformed the evaluated GRPO runs while using no verifier and fewer updates. They also report gains on tasks where all 64 base-model attempts initially failed, suggesting the method can learn without successful starting trajectories. ArXiv · AI/CL/LG's note
score 5