Megadose AI progress, ranked and analyzed.

One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents

· HF Daily Papers ·
The paper targets “category see-saw” failures in SWE agent training, where pooled RL improves some task types while hurting others.

The authors propose category-specific expert training, then merge those experts into one deployable policy through label-routed on-policy distillation. Their loop refreshes task mastery, reuses verified successful trajectories for repair SFT, and reselects tasks for more RL without relying on an external model for solution traces. The final policy reports 58.04% mean resolution on Pro-618 and 59.00% on SWE-bench Multilingual. HF Daily Papers' note

score 5

Categories: Research