Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
Training weak agents on expert traces made them worse under harnesses built for their own behavior.
The paper reports regressions of 4 to 30 points across seven enterprise agent tasks when Qwen3-Coder and Gemma 4 imitated full expert trajectories inside evolved harnesses. The authors argue the weaker models copied expert planning patterns they could not execute, breaking the fit between model and harness. Their fix is on-policy expert correction: find the failing turn in the weaker model’s own rollout and have the expert rewrite only that turn. HF Daily Papers' note
The paper reports regressions of 4 to 30 points across seven enterprise agent tasks when Qwen3-Coder and Gemma 4 imitated full expert trajectories inside evolved harnesses. The authors argue the weaker models copied expert planning patterns they could not execute, breaking the fit between model and harness. Their fix is on-policy expert correction: find the failing turn in the weaker model’s own rollout and have the expert rewrite only that turn. HF Daily Papers' note
score 5