SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning
SRPO turns a model’s own failed trajectories into token-level training signals without an external critic or reward model.
The paper proposes “reflection patches” distilled from completed reasoning runs, then uses reflection-conditioned teacher scores to guide student rollouts. The authors report state-of-the-art results on math and long-horizon agent benchmarks, including 73.3% on AIME’24 with a Qwen3-8B base model. They also report gains on WebShop, ALFWorld, and SWE-Bench-Lite, while claiming much lower training FLOPs than scaled supervised fine-tuning. ArXiv · AI/CL/LG’s note
The paper proposes “reflection patches” distilled from completed reasoning runs, then uses reflection-conditioned teacher scores to guide student rollouts. The authors report state-of-the-art results on math and long-horizon agent benchmarks, including 73.3% on AIME’24 with a Qwen3-8B base model. They also report gains on WebShop, ALFWorld, and SWE-Bench-Lite, while claiming much lower training FLOPs than scaled supervised fine-tuning. ArXiv · AI/CL/LG’s note
score 5