onepot-Bench 0: towards lab-aware in silico chemistry benchmarks
The benchmark is built around wet-lab chemistry decisions, including private reaction data from the authors’ own lab.
onepot-Bench 0 tests language models on synthetic chemistry tasks the authors say existing evaluations miss. Its three parts cover cheminformatics and numerical reasoning, refusal behavior around benign and drug-related targets, and reaction outcome or catalyst prediction. The paper frames the suite as a way to measure basic competence, reliability, and deeper chemistry knowledge needed for lab use. ArXiv · AI/CL/LG's note
onepot-Bench 0 tests language models on synthetic chemistry tasks the authors say existing evaluations miss. Its three parts cover cheminformatics and numerical reasoning, refusal behavior around benign and drug-related targets, and reaction outcome or catalyst prediction. The paper frames the suite as a way to measure basic competence, reliability, and deeper chemistry knowledge needed for lab use. ArXiv · AI/CL/LG's note
score 5