Megadose Built for builders and researchers.

Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis

· HF Daily Papers ·
The paper’s core claim is that retrosynthesis models need to produce several chemically plausible answers, not chase one “correct” route.

The authors introduce Top-K prompting for training and inference in single-step retrosynthesis. They train C3LM on about 45.6 million verified reactions, adding ChemCensor-based and novelty-oriented rewards. The model reports state-of-the-art results on the OOD URSA-expert-2026 benchmark. Their analysis also says LLMs and conventional models cover partly different reaction spaces, making ensembles useful. HF Daily Papers' note

score 5

Categories: Research