SPADE: Self-Play in Adaptive Synthetic Executable Environments
The paper’s core claim is that training agents on learnable, executable environments beats fixed task pools.
SPADE has one LLM design Gym-style environments in code while another role learns to act inside them. The designer uses a regret signal based on performance with and without privileged hints, pushing tasks toward the agent’s current limits while keeping them solvable. In experiments up to 30B parameters, the authors report gains over the strongest fixed-environment baseline across held-out reasoning, coding, science, math, tool-use, and game settings. ArXiv · AI/CL/LG's note
SPADE has one LLM design Gym-style environments in code while another role learns to act inside them. The designer uses a regret signal based on performance with and without privileged hints, pushing tasks toward the agent’s current limits while keeping them solvable. In experiments up to 30B parameters, the authors report gains over the strongest fixed-environment baseline across held-out reasoning, coding, science, math, tool-use, and game settings. ArXiv · AI/CL/LG's note
score 6