Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning
The paper argues hint depth works best as a sampled range, not a fixed cutoff.
Agent-G² draws each task’s retained expert-prefix length from an online-estimated Gaussian, using rollout data already collected for policy training. The method sets the distribution around global and cluster-level difficulty, with variance tracking differences inside each cluster. On ALFWorld and WebShop with Qwen2.5 Instruct models, the authors report stronger ALFWorld results than hint-based, hint-free, and Aux-RL baselines, while using under one-third the rollout cost of per-sample probing. HF Daily Papers' note
Agent-G² draws each task’s retained expert-prefix length from an online-estimated Gaussian, using rollout data already collected for policy training. The method sets the distribution around global and cluster-level difficulty, with variance tracking differences inside each cluster. On ALFWorld and WebShop with Qwen2.5 Instruct models, the authors report stronger ALFWorld results than hint-based, hint-free, and Aux-RL baselines, while using under one-third the rollout cost of per-sample probing. HF Daily Papers' note
score 4