Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
SRD trains a model to predict, before acting, the hindsight lessons revealed by its own completed trajectories.
The paper frames this as a fix for RLVR cases where scalar rewards stop being useful because rollout groups all succeed or all fail. SRD uses post-hoc trajectory information as supervision for “foresight” during training, without requiring that foresight to be generated at inference time. Across 10 tool-integrated reasoning and long-horizon agent tasks, the authors report gains up to 24.2 percentage points over RLVR and self-distillation baselines. In one 2B setting with 98% all-failure groups, RLVR ended at 0.0% success, while RLVR plus SRD reached 60.6%. ArXiv · AI/CL/LG's note
The paper frames this as a fix for RLVR cases where scalar rewards stop being useful because rollout groups all succeed or all fail. SRD uses post-hoc trajectory information as supervision for “foresight” during training, without requiring that foresight to be generated at inference time. Across 10 tool-integrated reasoning and long-horizon agent tasks, the authors report gains up to 24.2 percentage points over RLVR and self-distillation baselines. In one 2B setting with 98% all-failure groups, RLVR ended at 0.0% success, while RLVR plus SRD reached 60.6%. ArXiv · AI/CL/LG's note
score 5