AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning
AdaStep tries to make long-horizon agent training less blunt by shrinking noisy step-level credit when later randomness overwhelms the action’s signal.
The paper frames step credit as a weighting problem layered onto trajectory-level rewards. Its per-state coefficient keeps local advantage estimates when return variation appears action-driven, and dampens them when later actions, environment transitions, or trajectory length dominate. The method adds only scalar computation, with no critic, extra rollouts, or additional model inference. The authors report consistent gains across three model backbones on ALFWorld, WebShop, and ScienceWorld. ArXiv · AI/CL/LG's note
The paper frames step credit as a weighting problem layered onto trajectory-level rewards. Its per-state coefficient keeps local advantage estimates when return variation appears action-driven, and dampens them when later actions, environment transitions, or trajectory length dominate. The method adds only scalar computation, with no critic, extra rollouts, or additional model inference. The authors report consistent gains across three model backbones on ALFWorld, WebShop, and ScienceWorld. ArXiv · AI/CL/LG's note
score 5