Megadose Built for builders and researchers.

SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

· ArXiv · AI/CL/LG ·
SRPO turns a model’s own failed trajectories into token-level training signals without an external critic or reward model.

The paper proposes “reflection patches” distilled from completed reasoning runs, then uses reflection-conditioned teacher scores to guide student rollouts. The authors report state-of-the-art results on math and long-horizon agent benchmarks, including 73.3% on AIME’24 with a Qwen3-8B base model. They also report gains on WebShop, ALFWorld, and SWE-Bench-Lite, while claiming much lower training FLOPs than scaled supervised fine-tuning. ArXiv · AI/CL/LG’s note

score 5

Categories: Research