TokenRhythm/NeoHorse
NeoHorse-1 packages 4B and 9B Qwen3.5-derived agent models with a routing harness meant to turn tool-use outcomes back into training signal.
The repo describes the loop as evaluation, selection, and update: tasks go to a heterogeneous model pool, tool traces are recorded, and capability feedback informs the next mixture. Both checkpoints are text-in/text-out weights, released under Apache 2.0, with GGUF, quantized, MLX, Hugging Face, and ModelScope options listed. In the reported ten-benchmark average, NeoHorse-1-4B scores 64.87, up 5.93 over Qwen3.5-4B; NeoHorse-1-9B scores 69.04, up 3.44 over Qwen3.5-9B. The stated next step is extending the harness across successive iterations toward recursive self-improvement. GitHub · LLM repos' note
The repo describes the loop as evaluation, selection, and update: tasks go to a heterogeneous model pool, tool traces are recorded, and capability feedback informs the next mixture. Both checkpoints are text-in/text-out weights, released under Apache 2.0, with GGUF, quantized, MLX, Hugging Face, and ModelScope options listed. In the reported ten-benchmark average, NeoHorse-1-4B scores 64.87, up 5.93 over Qwen3.5-4B; NeoHorse-1-9B scores 69.04, up 3.44 over Qwen3.5-9B. The stated next step is extending the harness across successive iterations toward recursive self-improvement. GitHub · LLM repos' note
score 4