Megadose AI progress, ranked and analyzed.

ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

· HF Daily Papers ·
ReflectRL trains models to learn from expert failures instead of throwing them away.

The paper calls these failed expert attempts “Golden Negative Trajectories” and treats them as material for reflection, not imitation. Its premise is that, on hard problems, critiquing a flawed solution can be easier than solving from scratch. ReflectRL first elicits reflective reasoning from those failures, then transfers that behavior back into direct reasoning. Across 9 benchmarks, 4 model backbones, and 4 on-policy training methods, the authors report consistent reasoning gains with minimal overhead. HF Daily Papers' note

score 5

Categories: Research