NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale
A 1T-model refit that took 87.5 minutes fell to 150 seconds in the paper’s relay-tree setup.
NeMo-DCR sends only changed weights between training and rollout clusters while preserving bit-exact checkpoint state. The authors say BF16 training changes about 1% of stored weight values per step, making full-checkpoint transfer wasteful for agentic RL. Their method uses XOR masks and overwrites, applies updates in place, and supports retries after partial writes. At 3% to 5% change rates, they report 12-40x faster refits for 30B-1T models than a transport-only full-checkpoint reference. HF Daily Papers' note
NeMo-DCR sends only changed weights between training and rollout clusters while preserving bit-exact checkpoint state. The authors say BF16 training changes about 1% of stored weight values per step, making full-checkpoint transfer wasteful for agentic RL. Their method uses XOR masks and overwrites, applies updates in place, and supports retries after partial writes. At 3% to 5% change rates, they report 12-40x faster refits for 30B-1T models than a transport-only full-checkpoint reference. HF Daily Papers' note
score 5