Recursive Synthesis for Long-Horizon Terminal Tasks
RST generated 37,484 verified terminal-agent tasks at about $0.05 each, with difficulty rising across 15 rounds.
The paper’s method starts from verified seed tasks, extends the reference solution, realigns the verifier and instruction, then validates each new task in a fresh sandbox. Median reference solutions grew from 67 to 374 lines, while DeepSeek-V4-Pro pass@4 fell from 90% in round 1 to 2.5% in round 15. Training Qwen3.5 models on rejection-sampled trajectories from the synthetic tasks improved results on three terminal-agent benchmarks. HF Daily Papers' note
The paper’s method starts from verified seed tasks, extends the reference solution, realigns the verifier and instruction, then validates each new task in a fresh sandbox. Median reference solutions grew from 67 to 374 lines, while DeepSeek-V4-Pro pass@4 fell from 90% in round 1 to 2.5% in round 15. Training Qwen3.5 models on rejection-sampled trajectories from the synthetic tasks improved results on three terminal-agent benchmarks. HF Daily Papers' note
score 5