Megadose AI progress, ranked and analyzed.

Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training

· HF Daily Papers ·
AED turns failed agent rollouts into 50,228 checked error-diagnosis examples for training and analysis.

The dataset spans 9,961 source tasks across 33 environments, with traces and metadata kept so failures can be re-examined without rerunning the original rollout. The paper’s AET pipeline diagnoses natural failures, proposes corrections, and checks them against recorded evidence. In 3,062 matched replay pairs, first-proposal corrections lifted verifier pass rates from 18.4% to 51.1%. Fine-tuning Qwen3-8B on a frozen diagnosis release also raised exact-step agreement with internal teacher labels from 47.2% to 63.6% on a 943-case holdout. HF Daily Papers' note

score 5

Categories: Research