Megadose AI progress, ranked and analyzed.

Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training

· ArXiv · AI/CL/LG ·
Rollplex reports up to 2.24x faster VLM RL post-training on the same GPU budget.

The paper targets wasted GPU time in on-policy RL pipelines for vision-language models, where rollout, reference scoring, and actor training usually run as serial phases. Rollplex moves prefix computation for reference and training into the rollout decode window while keeping synchronous RL semantics intact. It adds memory management and weight-sharing mechanisms to handle high HBM use and different tensor-parallel layouts. On 32 H800 GPUs, the authors report 1.23x-1.30x speedup over serial colocation and 1.57x-2.24x over disaggregation. ArXiv · AI/CL/LG's note

score 5

Categories: Research