Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents
HMED tests a meta-skill revision against the prior version from the same restored discovery state.
The paper argues that judging revisions by later branch outcomes mixes the revision’s effect with the luck of the starting state. Its method replays both the incumbent and revised meta-skills from that shared state, then distills the comparison into reusable “Meta-Experience.” That lets failed or discarded revisions still provide training signal. The authors report gains in skill discovery across three interactive agent benchmarks using both open- and closed-source models. ArXiv · AI/CL/LG's note
The paper argues that judging revisions by later branch outcomes mixes the revision’s effect with the luck of the starting state. Its method replays both the incumbent and revised meta-skills from that shared state, then distills the comparison into reusable “Meta-Experience.” That lets failed or discarded revisions still provide training signal. The authors report gains in skill discovery across three interactive agent benchmarks using both open- and closed-source models. ArXiv · AI/CL/LG's note
score 5