Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis
The paper’s claim is that improving the data-generation harness itself makes synthetic reasoning tasks harder without changing model weights or verifiers.
The authors propose task-harness co-evolution, where solver failures become reusable generation skills and batch-level reviews revise prompts, skills, and workflows. Candidate changes are kept only if they produce harder valid tasks within bounded added cost. Across math, coding, and science, solver accuracy fell from 100.0% to 54.8% over fourteen rounds. The synthesized data also improved SFT and GRPO results, including a 27B student reaching 62.5% mean-16 accuracy on APEX after fine-tuning on 10K math examples. ArXiv · AI/CL/LG's note
The authors propose task-harness co-evolution, where solver failures become reusable generation skills and batch-level reviews revise prompts, skills, and workflows. Candidate changes are kept only if they produce harder valid tasks within bounded added cost. Across math, coding, and science, solver accuracy fell from 100.0% to 54.8% over fourteen rounds. The synthesized data also improved SFT and GRPO results, including a 27B student reaching 62.5% mean-16 accuracy on APEX after fine-tuning on 10K math examples. ArXiv · AI/CL/LG's note
score 6