SABRE: Scalable and Automated Benchmarking of VLMs under Stress
The benchmark pipeline is meant to keep producing new VLM stress tests, not freeze one dataset in place.
SABRE turns a Markdown task design into specifications, images, and question-answer pairs, then filters out cases a VLM can already solve. The paper’s SABRE-Prior set tests whether models follow visual evidence over learned expectations, with 600 images and 1,000 questions. Across six VLMs, reported macro-average accuracy runs from 17.8% to 31.3%. The authors also describe counting and spatial pilots as signs the workflow can extend beyond the prior-reliance test. ArXiv · AI/CL/LG's note
SABRE turns a Markdown task design into specifications, images, and question-answer pairs, then filters out cases a VLM can already solve. The paper’s SABRE-Prior set tests whether models follow visual evidence over learned expectations, with 600 images and 1,000 questions. Across six VLMs, reported macro-average accuracy runs from 17.8% to 31.3%. The authors also describe counting and spatial pilots as signs the workflow can extend beyond the prior-reliance test. ArXiv · AI/CL/LG's note
score 5