ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
ReSPO targets a clipping failure that can mute useful positive rollouts while letting bad high-weight rollouts dominate.
The paper says rollout reuse in RLVR can widen the gap between the current policy and the policy that generated the data. Its proposed fix replaces clipping with a smooth sequence-level kernel, keeping gradient signal for under-generated positive responses and damping heavily over-generated negative ones. The authors report faster early optimization and better held-out benchmark performance on dense and MoE Qwen3 models under rollout reuse. HF Daily Papers' note
The paper says rollout reuse in RLVR can widen the gap between the current policy and the policy that generated the data. Its proposed fix replaces clipping with a smooth sequence-level kernel, keeping gradient signal for under-generated positive responses and damping heavily over-generated negative ones. The authors report faster early optimization and better held-out benchmark performance on dense and MoE Qwen3 models under rollout reuse. HF Daily Papers' note
score 4