Megadose Built for builders and researchers.

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

· HF Daily Papers ·
LoGRA claims up to 45.7% lower average RL training memory without losing performance.

The paper compresses RL post-training updates into low-rank gradient sketches, using them for model updates and policy synchronization. It adds predicted-KL step control to estimate policy movement before each update and scale it back when needed. The authors say this let them train a 27B-parameter model for more than 1,100 steps on one eight-GPU node, while dense Adam ran out of memory. HF Daily Papers' note

score 5

Categories: Research