EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
EnvACE trains an agent to rehearse the environment’s response itself instead of relying on external environment interaction.
The method has the policy alternate between making a tool call and generating the response that action would induce. Those acting and rehearsal roles are optimized together with task-success rewards, so the model internalizes action-response dynamics. The paper reports stronger overall results than environment-scaling baselines across BFCL-v4, tau^2-Bench, VitaBench, and FinMCP-Bench. At test time, the internal world model can rehearse privately before execution under a moderate budget. ArXiv · AI/CL/LG's note
The method has the policy alternate between making a tool call and generating the response that action would induce. Those acting and rehearsal roles are optimized together with task-success rewards, so the model internalizes action-response dynamics. The paper reports stronger overall results than environment-scaling baselines across BFCL-v4, tau^2-Bench, VitaBench, and FinMCP-Bench. At test time, the internal world model can rehearse privately before execution under a moderate budget. ArXiv · AI/CL/LG's note
score 6