Megadose AI progress, ranked and analyzed.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

· ArXiv · AI/CL/LG ·
Failed expert reasoning traces become training signal instead of dead weight.

The paper introduces “Golden Negative Trajectories,” cases where stronger expert models attempt hard problems but fail. ReflectRL uses those flawed traces to prompt reflective reasoning, then transfers the learned behavior back into direct reasoning during on-policy training. The authors report consistent gains across 9 benchmarks, 4 LLM backbones, and 4 on-policy training methods, with minimal overhead. ArXiv · AI/CL/LG's note

score 5

Categories: Research