Dynamic language model representations for multi-objective reaction optimisation
The paper’s core claim is that reaction conditions can be optimized faster when the model learns its chemistry representation from text.
The authors fine-tune a language model with Gaussian process surrogates inside a multi-objective Bayesian optimization loop. In nickel- and palladium-catalysed cross-coupling tests, it reached convergence in fewer experiments than descriptor libraries or one-hot encoding. In prospective high-throughput runs, two rounds covering 192 reactions and under 3% of each design space produced gram-scale conditions at 94% and 84% isolated yield, with the hydrogenation example reaching 99.6% enantiomeric excess. ArXiv · AI/CL/LG's note
The authors fine-tune a language model with Gaussian process surrogates inside a multi-objective Bayesian optimization loop. In nickel- and palladium-catalysed cross-coupling tests, it reached convergence in fewer experiments than descriptor libraries or one-hot encoding. In prospective high-throughput runs, two rounds covering 192 reactions and under 3% of each design space produced gram-scale conditions at 94% and 84% isolated yield, with the hydrogenation example reaching 99.6% enantiomeric excess. ArXiv · AI/CL/LG's note
score 4