HuatuoGPT-3: RL-Only Domain Adaptation from Base Models
OnePO trains a medical model with RL alone, using teacher outputs only until the policy can beat them.
The paper argues that standard SFT+RL can narrow exploration and add avoidable training complexity. Its proposed One-stage Policy Optimization method targets two RL-only failure modes: slow learning from useful teacher tokens and over-attachment to stale teacher outputs. In medical adaptation, OnePO scores 67.2 on HealthBench Total with 20K samples, ahead of the authors’ SFT+RL and pure-RL baselines. Scaled into HuatuoGPT-3, the 27B model reaches 70.1 on HealthBench Total and 71.4 on HealthBench Professional, with models and code released. HF Daily Papers' note
The paper argues that standard SFT+RL can narrow exploration and add avoidable training complexity. Its proposed One-stage Policy Optimization method targets two RL-only failure modes: slow learning from useful teacher tokens and over-attachment to stale teacher outputs. In medical adaptation, OnePO scores 67.2 on HealthBench Total with 20K samples, ahead of the authors’ SFT+RL and pure-RL baselines. Scaled into HuatuoGPT-3, the 27B model reaches 70.1 on HealthBench Total and 71.4 on HealthBench Professional, with models and code released. HF Daily Papers' note
score 5