Megadose AI progress, ranked and analyzed.

Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails

· ArXiv · AI/CL/LG ·
Full-trajectory imitation made weaker agents worse after the harness had been tuned around them.

Across seven enterprise agent tasks, the authors found that copying an expert model’s complete rollouts under an evolved harness cut performance by 4 to 30 points for Qwen3-Coder and Gemma 4. The failure came from a mismatch: weaker models adopted the expert’s planning style without the same execution ability, breaking fit with the harness built around their native behavior. Their proposed fix trains on the weaker model’s own rollout, asking the expert to rewrite only the failing turn. ArXiv · AI/CL/LG's note

score 5

Categories: Research