NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
NexForge turns high-level capability requirements into executable agent-training tasks without building a custom pipeline for each domain.
The paper says the framework researches demand, builds task profiles, generates directives, then assembles files, dependencies, runtime configs, and expert rollouts for supervised fine-tuning. It reports 3.6K terminal tasks and 2K office tasks, lifting Qwen3.5-35B-A3B Base from 22.5% to 52.0% on Terminal-Bench 2.0 and from 813 to 1338 Elo on GDPval. At 43.2K terminal tasks, the model reaches 58.4%, which the authors say is on par with Claude Opus 4.6 using Claude Code. The scaled data also contributes to Nex-N2, a public agent-model family reporting 75.3% on Terminal-Bench 2.1 and 1585 Elo on GDPval. HF Daily Papers' note
The paper says the framework researches demand, builds task profiles, generates directives, then assembles files, dependencies, runtime configs, and expert rollouts for supervised fine-tuning. It reports 3.6K terminal tasks and 2K office tasks, lifting Qwen3.5-35B-A3B Base from 22.5% to 52.0% on Terminal-Bench 2.0 and from 813 to 1338 Elo on GDPval. At 43.2K terminal tasks, the model reaches 58.4%, which the authors say is on par with Claude Opus 4.6 using Claude Code. The scaled data also contributes to Nex-N2, a public agent-model family reporting 75.3% on Terminal-Bench 2.1 and 1585 Elo on GDPval. HF Daily Papers' note
score 5