Megadose AI progress, ranked and analyzed.

onepot-Bench 0: towards lab-aware in silico chemistry benchmarks

· ArXiv · AI/CL/LG ·
The benchmark is built around wet-lab chemistry decisions, including private reaction data from the authors’ own lab.

onepot-Bench 0 tests language models on synthetic chemistry tasks the authors say existing evaluations miss. Its three parts cover cheminformatics and numerical reasoning, refusal behavior around benign and drug-related targets, and reaction outcome or catalyst prediction. The paper frames the suite as a way to measure basic competence, reliability, and deeper chemistry knowledge needed for lab use. ArXiv · AI/CL/LG's note

score 5

Categories: Research