Can LLMs Discover Scientific Laws in Real and Parallel Worlds?
SCILAWS-BENCH tests whether LLMs can infer scientific equations from real observations and controlled “parallel” worlds.
The benchmark has 118 problems from 381 papers, spanning 291 candidate laws and about 8 million real data points across six disciplines. Its real setting scores proposed laws on held-out predictive fit and scientific validity from the literature. Its parallel setting lets models actively query calibrated synthetic worlds to recover hidden laws derived from published forms. The authors report that strong prediction and scientific validity can come apart, and that memorization and candidate selection both affect results. ArXiv · AI/CL/LG's note
The benchmark has 118 problems from 381 papers, spanning 291 candidate laws and about 8 million real data points across six disciplines. Its real setting scores proposed laws on held-out predictive fit and scientific validity from the literature. Its parallel setting lets models actively query calibrated synthetic worlds to recover hidden laws derived from published forms. The authors report that strong prediction and scientific validity can come apart, and that memorization and candidate selection both affect results. ArXiv · AI/CL/LG's note
score 6