Nereus: Adaptive Parallelism for LLM Post-Training
Nereus changes the parallelism plan during RL post-training instead of staying with a layout that may become slow or infeasible.
The paper frames the problem as a moving target across generation, inference, and training on shared GPU clusters. Nereus uses a low-overhead controller, Elastic Model Units, and a global transition graph to decide and carry out plan changes. In reported tests, online TP/PP adaptation cut average step latency by 27.7% versus an initial fixed layout, while six transitions in a 1,000-step run took 0.079% of runtime. The authors also report 8B PPO throughput gains over OpenRLHF and Verl. HF Daily Papers' note
The paper frames the problem as a moving target across generation, inference, and training on shared GPU clusters. Nereus uses a low-overhead controller, Elastic Model Units, and a global transition graph to decide and carry out plan changes. In reported tests, online TP/PP adaptation cut average step latency by 27.7% versus an initial fixed layout, while six transitions in a 1,000-step run took 0.079% of runtime. The authors also report 8B PPO throughput gains over OpenRLHF and Verl. HF Daily Papers' note
score 5