Megadose AI progress, ranked and analyzed.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

· HF Daily Papers ·
GDPevo tests whether agents can turn prior business-task experience into measurable gains on held-out workflows.

The benchmark builds CRM, ERP, finance, healthcare, legal, and data-centric tasks from atomic business rules, then recombines those rules so improvements can be tied to training experience. Its first release has 120 tasks across 12 groups, with an automated pipeline that can expand the set to 240 tasks within two days. In evaluations of four agent-model harnesses under four supervision types, self-evolution raised held-out accuracy by as much as 16.44 percentage points. The strongest evolved agents still trailed a 91.6% oracle ceiling by a wide margin.

Source: HF Daily Papers' note

score 5

Categories: Research