Megadose Built for builders and researchers.

Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight

· ArXiv · AI/CL/LG ·
SRD trains a model to predict, before acting, the hindsight lessons revealed by its own completed trajectories.

The paper frames this as a fix for RLVR cases where scalar rewards stop being useful because rollout groups all succeed or all fail. SRD uses post-hoc trajectory information as supervision for “foresight” during training, without requiring that foresight to be generated at inference time. Across 10 tool-integrated reasoning and long-horizon agent tasks, the authors report gains up to 24.2 percentage points over RLVR and self-distillation baselines. In one 2B setting with 98% all-failure groups, RLVR ended at 0.0% success, while RLVR plus SRD reached 60.6%. ArXiv · AI/CL/LG's note

score 5

Categories: Research