Megadose AI progress, ranked and analyzed.

SLPO: Scaling Latent Reasoning via a Surrogate Policy

· HF Daily Papers ·
The paper proposes outcome-reward RL for latent reasoners without decoding every reasoning step into tokens.

SLPO adds a surrogate policy density over latent transitions so trajectory-level credit can be assigned in continuous reasoning space. It also trains a stopping head that can turn fixed thinking budgets into variable-horizon computation. The authors report gains in Pass@k under parallel sampling and higher deterministic accuracy by spending more latent computation on harder cases. HF Daily Papers' note

score 5

Categories: Research