Megadose AI progress, ranked and analyzed.

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

· HF Daily Papers ·
AgentOPSD turns sparse task outcomes into turn-level credit without adding a critic or extra rollouts.

The method uses teacher-student log-probability gaps as evidence, aggregates them by turn, and recursively updates a Bayesian belief state. That produces weights meant to identify which turns were pivotal in long multi-turn agent tasks. The paper reports gains over GRPO and self-distillation baselines on ALFWorld, WebShop, and Search-QA, including 89.1% success on ALFWorld with Qwen2.5-7B. HF Daily Papers' note

score 5

Categories: Research