Ornith-1.5 open models launch in 397B, 35B, and 9 B sizes.
Ornith-1.5 turns its scaffold system into a self-improving training loop.
The new release has the model propose harder tasks, build task-specific scaffolds, and generate rollouts for reinforcement learning. Reward is tied to verifiability, novelty, and difficulty near the model’s current frontier, with malformed tasks receiving no value. Ornith reports its 397B MoE model at 85.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, alongside smaller 35B MoE and 9B dense variants. A quantized Mobile build is also listed for iPhone and Android. TestingCatalog's note
The new release has the model propose harder tasks, build task-specific scaffolds, and generate rollouts for reinforcement learning. Reward is tied to verifiability, novelty, and difficulty near the model’s current frontier, with malformed tasks receiving no value. Ornith reports its 397B MoE model at 85.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, alongside smaller 35B MoE and 9B dense variants. A quantized Mobile build is also listed for iPhone and Android. TestingCatalog's note
score 7