WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation
WarpSAC argues that off-policy RL stabilizers should change with the data regime, not stay fixed.
The paper says massively parallel simulation makes some data-limited training tricks less useful or even restrictive. Its experiments find parameter normalization helps with narrow replay coverage, while clipped double-Q can be relaxed in high-throughput manipulation settings. WarpSAC uses Sample Weight Decay and splits into a CPU-scale variant and a GPU-parallel variant. The authors report gains over FlashSAC, including 23.1% higher normalized score-step AUC across fourteen GPU-parallel environments and a UnitreeG1TransportBox-v1 success-rate jump from 19.8% to 96.4%. HF Daily Papers' note
The paper says massively parallel simulation makes some data-limited training tricks less useful or even restrictive. Its experiments find parameter normalization helps with narrow replay coverage, while clipped double-Q can be relaxed in high-throughput manipulation settings. WarpSAC uses Sample Weight Decay and splits into a CPU-scale variant and a GPU-parallel variant. The authors report gains over FlashSAC, including 23.1% higher normalized score-step AUC across fourteen GPU-parallel environments and a UnitreeG1TransportBox-v1 success-rate jump from 19.8% to 96.4%. HF Daily Papers' note
score 4