BIABench: Evaluating AI agents on real-world bioimage analysis tasks
The benchmark found current agents breaking down on 3D and time-lapse bioimage analysis, with some tasks scoring under 0.19.
BIABench reconstructs 16 analysis tasks from published biological studies, pairing raw imaging data with peer-reviewed ground truth. It scores both the output files and the agent’s process, including method choice and quality control. Routine 2D tasks were handled relatively well, but added spatial or temporal complexity exposed a sharp gap. Stronger models, biology-specific agents and expert instructions did not close it. HF Daily Papers' note
BIABench reconstructs 16 analysis tasks from published biological studies, pairing raw imaging data with peer-reviewed ground truth. It scores both the output files and the agent’s process, including method choice and quality control. Routine 2D tasks were handled relatively well, but added spatial or temporal complexity exposed a sharp gap. Stronger models, biology-specific agents and expert instructions did not close it. HF Daily Papers' note
score 5