NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
NeoHorse-1 turns routing logs into training data for the next model update.
The paper describes an agentic post-training loop that records capability demand, service-tier routing, and the resulting interaction for each user turn. Those records are filtered and labeled, then used for staged supervised fine-tuning and routing-guided on-policy distillation. In reported benchmarks, post-training lifts macro-average scores from 58.94 to 64.87 for the 4B model and from 65.60 to 69.04 for the 9B model. The authors present it as an early harness-mediated prototype for recursive self-improvement.
HF Daily Papers' note
The paper describes an agentic post-training loop that records capability demand, service-tier routing, and the resulting interaction for each user turn. Those records are filtered and labeled, then used for staged supervised fine-tuning and routing-guided on-policy distillation. In reported benchmarks, post-training lifts macro-average scores from 58.94 to 64.87 for the 4B model and from 65.60 to 69.04 for the 9B model. The authors present it as an early harness-mediated prototype for recursive self-improvement.
HF Daily Papers' note
score 5