Megadose AI progress, ranked and analyzed.

Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists

· HF Daily Papers ·
SDABench tests LLMs by the kinds of scientific claims their analyses are supposed to support.

The benchmark covers six capabilities across biology, chemistry, environment, geography, and physics. It includes 527 real-data instances and 6,000 synthetic ones, with both multiple-choice and open-ended formats. In tests of 15 LLMs, models did well on descriptive analysis but fell off on assumption selection, latent-process modeling, and mechanistic reasoning. The authors say stronger models identify scope and variables more reliably, while still struggling with procedures, relationships, and valid conclusions. HF Daily Papers' note

score 5

Categories: Research