Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training
Rollplex reports up to 2.24x faster VLM RL post-training on the same GPU budget.
The paper targets wasted GPU time in on-policy RL pipelines for vision-language models, where rollout, reference scoring, and actor training usually run as serial phases. Rollplex moves prefix computation for reference and training into the rollout decode window while keeping synchronous RL semantics intact. It adds memory management and weight-sharing mechanisms to handle high HBM use and different tensor-parallel layouts. On 32 H800 GPUs, the authors report 1.23x-1.30x speedup over serial colocation and 1.57x-2.24x over disaggregation. ArXiv · AI/CL/LG's note
The paper targets wasted GPU time in on-policy RL pipelines for vision-language models, where rollout, reference scoring, and actor training usually run as serial phases. Rollplex moves prefix computation for reference and training into the rollout decode window while keeping synchronous RL semantics intact. It adds memory management and weight-sharing mechanisms to handle high HBM use and different tensor-parallel layouts. On 32 H800 GPUs, the authors report 1.23x-1.30x speedup over serial colocation and 1.57x-2.24x over disaggregation. ArXiv · AI/CL/LG's note
score 5