Megadose AI progress, ranked and analyzed.

Trajectory-Relative Hindsight Distillation for Agentic Reinforcement Learning

· ArXiv · AI/CL/LG ·
TRIAL reallocates hindsight supervision across each rollout’s turns, and the authors report consistent gains over GRPO on WebShop and ALFWorld.

The paper frames the problem as sparse outcome rewards leaving too many possible hindsight signals without a clear turn-by-turn assignment. TRIAL scores each decision under ordinary and hindsight-conditioned contexts, then uses the signed log-probability gap to set token-level supervision strength. Its turn-level weights are normalized across the realized trajectory so the average multiplier stays fixed. In the reported WebShop Qwen3-1.7B run, success rises from 56.4% to 75.2%, with task score moving from 78.7% to 85.7%. ArXiv · AI/CL/LG's note

score 4

Categories: Research