StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Even the strongest tested agent completed only about 30% of the benchmark’s real-world workflow tasks.
StartupBench builds its tasks from adopted AI startup products, using their workflows and users as evidence of market demand. The paper turns those workflows into complete, deliverable-oriented tasks with detailed grading rubrics. Under a unified agent harness, models often made partial progress but failed reliable end-to-end completion. The authors point to complex instruction following and domain-specific expertise as major failure points. HF Daily Papers' note
StartupBench builds its tasks from adopted AI startup products, using their workflows and users as evidence of market demand. The paper turns those workflows into complete, deliverable-oriented tasks with detailed grading rubrics. Under a unified agent harness, models often made partial progress but failed reliable end-to-end completion. The authors point to complex instruction following and domain-specific expertise as major failure points. HF Daily Papers' note
score 5