BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
The benchmark found top AI-agent setups solved only about half of pathogen-surveillance evaluations.
BioSecBench-Surveillance tests whether agents can choose and execute the right genomic-analysis path from raw sequencing data and analyst-style context. Its 100 evaluations cover seven task categories, including taxonomic classification and genetic-engineering detection. Across 3,962 gradable attempts from sixteen model-harness pairs, the best reported configurations reached 50.2 percent. The paper says failures often came not from missing the broad workflow, but from surrounding choices like references, thresholds, filters, and normalization. ArXiv · AI/CL/LG's note
BioSecBench-Surveillance tests whether agents can choose and execute the right genomic-analysis path from raw sequencing data and analyst-style context. Its 100 evaluations cover seven task categories, including taxonomic classification and genetic-engineering detection. Across 3,962 gradable attempts from sixteen model-harness pairs, the best reported configurations reached 50.2 percent. The paper says failures often came not from missing the broad workflow, but from surrounding choices like references, thresholds, filters, and normalization. ArXiv · AI/CL/LG's note
score 6