Verifiable Hidden Dynamics Play: Generating Agentic RL Environments from Solved Mechanisms
VHD-Play builds RL training environments by solving the mechanism first, then turning it into stateful agent tasks with matching evaluation.
The paper says this avoids the post hoc mismatch between environment dynamics and outcome rules. Its pipeline generated 3,300 agentic environments at a few cents each. Training Qwen3.6-35B-A3B on three environment families raised its mean agentic score from 0.204 to 0.815 in a five-family diagnostic, with gains reported on held-out and unseen mechanism families. The authors also report transfer to external benchmarks, including E-Commerce Bench, where the trained checkpoint completed every run without bankruptcy and beat Qwen3.7-Max. HF Daily Papers' note
The paper says this avoids the post hoc mismatch between environment dynamics and outcome rules. Its pipeline generated 3,300 agentic environments at a few cents each. Training Qwen3.6-35B-A3B on three environment families raised its mean agentic score from 0.204 to 0.815 in a five-family diagnostic, with gains reported on held-out and unseen mechanism families. The authors also report transfer to external benchmarks, including E-Commerce Bench, where the trained checkpoint completed every run without bankruptcy and beat Qwen3.7-Max. HF Daily Papers' note
score 5