Agensh: Scaling Organizational Intelligence to 1,024 Agents
Agensh replaces a central orchestrator with workers that claim, verify, and merge tasks asynchronously.
The paper says that setup improved ProgramBench results when scaled across many GPT-5.6-sol agents. On the five hardest tasks, moving from 1 to 128 agents raised mean final test-pass rate from 19.31% to 28.78%. On pandoc, scaling to 1,024 agents raised the final pass rate from 33.89% to 55.06%. The authors frame agent count as a scaling dimension for complex tasks under tight latency or time budgets. HF Daily Papers' note
The paper says that setup improved ProgramBench results when scaled across many GPT-5.6-sol agents. On the five hardest tasks, moving from 1 to 128 agents raised mean final test-pass rate from 19.31% to 28.78%. On pandoc, scaling to 1,024 agents raised the final pass rate from 33.89% to 55.06%. The authors frame agent count as a scaling dimension for complex tasks under tight latency or time budgets. HF Daily Papers' note
score 5