Megadose AI progress, ranked and analyzed.

SABRE: Scalable and Automated Benchmarking of VLMs under Stress

· ArXiv · AI/CL/LG ·
The benchmark pipeline is meant to keep producing new VLM stress tests, not freeze one dataset in place.

SABRE turns a Markdown task design into specifications, images, and question-answer pairs, then filters out cases a VLM can already solve. The paper’s SABRE-Prior set tests whether models follow visual evidence over learned expectations, with 600 images and 1,000 questions. Across six VLMs, reported macro-average accuracy runs from 17.8% to 31.3%. The authors also describe counting and spatial pilots as signs the workflow can extend beyond the prior-reliance test. ArXiv · AI/CL/LG's note

score 5

Categories: Research