Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
SRD trains an agent to predict useful foresight from completed trajectories, even when reward signals are flat.
The paper frames this as a supplement to RL with verifiable rewards, where group-relative training can lose signal when every rollout gets the same outcome. Its method, Self-Retrospection Distillation, uses hindsight from a finished trajectory to supervise what the same policy should have anticipated before acting. The authors report gains across 10 tool-integrated reasoning and long-horizon agent tasks, with improvements up to 24.2 percentage points. In a 2B model setting where 98% of groups were all-failure, RLVR alone ended at 0.0% success, while RLVR plus SRD reached 60.6% under the same rollout budget. HF Daily Papers' note
The paper frames this as a supplement to RL with verifiable rewards, where group-relative training can lose signal when every rollout gets the same outcome. Its method, Self-Retrospection Distillation, uses hindsight from a finished trajectory to supervise what the same policy should have anticipated before acting. The authors report gains across 10 tool-integrated reasoning and long-horizon agent tasks, with improvements up to 24.2 percentage points. In a 2B model setting where 98% of groups were all-failure, RLVR alone ended at 0.0% success, while RLVR plus SRD reached 60.6% under the same rollout budget. HF Daily Papers' note
score 5