APEX-Accounting
Frontier models still fail most expert-built accounting tasks in APEX-Accounting.
The benchmark uses 160 private tasks across 10 accounting “worlds,” with spreadsheets, PDFs, accounting systems, and expert-written rubrics. Claude-Fable-5 (Max) leads the reported nine-model run at 56.4% Mean Criteria@3, while no model clears 2.6% Pass^8. The paper also reports that higher token budgets improved aggregate scores, though within a fixed harness, tasks where models spent more tokens scored lower. ArXiv · AI/CL/LG's note
The benchmark uses 160 private tasks across 10 accounting “worlds,” with spreadsheets, PDFs, accounting systems, and expert-written rubrics. Claude-Fable-5 (Max) leads the reported nine-model run at 56.4% Mean Criteria@3, while no model clears 2.6% Pass^8. The paper also reports that higher token budgets improved aggregate scores, though within a fixed harness, tasks where models spent more tokens scored lower. ArXiv · AI/CL/LG's note
score 5