SLPO: Scaling Latent Reasoning via a Surrogate Policy
The paper proposes outcome-reward RL for latent reasoners without decoding every reasoning step into tokens.
SLPO adds a surrogate policy density over latent transitions so trajectory-level credit can be assigned in continuous reasoning space. It also trains a stopping head that can turn fixed thinking budgets into variable-horizon computation. The authors report gains in Pass@k under parallel sampling and higher deterministic accuracy by spending more latent computation on harder cases. HF Daily Papers' note
SLPO adds a surrogate policy density over latent transitions so trajectory-level credit can be assigned in continuous reasoning space. It also trains a stopping head that can turn fixed thinking budgets into variable-horizon computation. The authors report gains in Pass@k under parallel sampling and higher deterministic accuracy by spending more latent computation on harder cases. HF Daily Papers' note
score 5