Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
The paper proposes evolution strategies as a lighter way to fine-tune long-horizon LLM agents than RL.
Agentic ESOpt samples parameter perturbations, scores the resulting agents with rewards, and applies reward-weighted updates without a heavyweight backpropagation stack. The authors argue this makes full-parameter optimization possible with inference-level GPU memory and avoids step-by-step credit assignment over long trajectories. In reported tests, Qwen-3.5-27B improves over a no-skill baseline on WebArena-Lite by 6.69%, and prompt-parameter co-evolution improves a matched baseline in 28 of 36 settings. Source: HF Daily Papers' note.
Agentic ESOpt samples parameter perturbations, scores the resulting agents with rewards, and applies reward-weighted updates without a heavyweight backpropagation stack. The authors argue this makes full-parameter optimization possible with inference-level GPU memory and avoids step-by-step credit assignment over long trajectories. In reported tests, Qwen-3.5-27B improves over a no-skill baseline on WebArena-Lite by 6.69%, and prompt-parameter co-evolution improves a matched baseline in 28 of 36 settings. Source: HF Daily Papers' note.
score 5