Megadose Built for builders and researchers.

ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications

· ArXiv · AI/CL/LG ·
ChartBmkAgent is meant to generate targeted Chart QA test cases from sparse error categories, with a harness checking that each sample stays on target.

The paper says the system uses specialized agents under a central governance harness to build complete chart-question-answer samples from identified capability gaps. In 300 samples, tested MLLMs ranged from 32.7% to 84.3% accuracy, with different category profiles. Targeted follow-ups were harder than matched controls, scoring 50.0% versus 82.2% across three source-model comparisons. Evaluator models judged 86.4% of samples as testing their specified error category. ArXiv · AI/CL/LG's note

score 4

Categories: Research