Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
SDABench tests LLMs by the kinds of scientific claims their analyses are supposed to support.
The benchmark covers six capabilities across biology, chemistry, environment, geography, and physics. It includes 527 real-data instances and 6,000 synthetic ones, with both multiple-choice and open-ended formats. In tests of 15 LLMs, models did well on descriptive analysis but fell off on assumption selection, latent-process modeling, and mechanistic reasoning. The authors say stronger models identify scope and variables more reliably, while still struggling with procedures, relationships, and valid conclusions. HF Daily Papers' note
The benchmark covers six capabilities across biology, chemistry, environment, geography, and physics. It includes 527 real-data instances and 6,000 synthetic ones, with both multiple-choice and open-ended formats. In tests of 15 LLMs, models did well on descriptive analysis but fell off on assumption selection, latent-process modeling, and mechanistic reasoning. The authors say stronger models identify scope and variables more reliably, while still struggling with procedures, relationships, and valid conclusions. HF Daily Papers' note
score 5