Megadose AI progress, ranked and analyzed.

SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs

· HF Daily Papers ·
SlideDP reports up to 2.64x higher fine-tuning throughput than existing host-offload baselines in matched-batch tests.

The paper describes a synchronous data-parallel runtime for multi-GPU systems where model state lives on the host. It keeps one authoritative host copy, then pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. On four H100s, it comes close to GPU-resident FSDP2 for Qwen3-14B at smaller batch size, and beats FSDP2’s measured peak throughput by 11.2% at larger batch size. It also reports 256K-token sequence support for Qwen3-14B and Qwen2.5-72B fine-tuning on four RTX 4090s. HF Daily Papers' note

score 4

Categories: Research