VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences
VIALS tests whether vision-language models can read real biotech lab artifacts, and the paper says they still fall short.
The benchmark has 161 visual question-answering tasks built from professional life sciences workflows, including gel blots, microscopy images, plasmid maps, flow cytometry plots, and molecular structures. The authors say the artifacts reflect biotech work materials, not polished textbook or publication figures. Frontier models can describe ordinary images fluently, but the paper reports that they do not accurately interpret these scientific images. Scientists with relevant expertise found the tasks straightforward. ArXiv · AI/CL/LG's note
The benchmark has 161 visual question-answering tasks built from professional life sciences workflows, including gel blots, microscopy images, plasmid maps, flow cytometry plots, and molecular structures. The authors say the artifacts reflect biotech work materials, not polished textbook or publication figures. Frontier models can describe ordinary images fluently, but the paper reports that they do not accurately interpret these scientific images. Scientists with relevant expertise found the tasks straightforward. ArXiv · AI/CL/LG's note
score 5