Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
The paper argues PPO-Clip’s exploration failure comes from using the wrong geometry for policy updates.
The authors say PPO-Clip treats policy discrepancy with a Euclidean metric, while the policy space is better understood as a Riemannian manifold. That mismatch makes updates too cautious in low-probability regions and too aggressive in high-probability regions, which they identify as the driver of exploration collapse. Their proposed method, RIPO, is designed to keep policy updates isometric on that manifold and balance exploration with exploitation. They report gains across seven competition-level benchmarks, including up to 60% over GRPO on AIME24. HF Daily Papers' note
The authors say PPO-Clip treats policy discrepancy with a Euclidean metric, while the policy space is better understood as a Riemannian manifold. That mismatch makes updates too cautious in low-probability regions and too aggressive in high-probability regions, which they identify as the driver of exploration collapse. Their proposed method, RIPO, is designed to keep policy updates isometric on that manifold and balance exploration with exploitation. They report gains across seven competition-level benchmarks, including up to 60% over GRPO on AIME24. HF Daily Papers' note
score 5