Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization
The paper says 4-bit AdamW state quantization works better when rounding is designed around the preconditioner, not just the stored moment values.
The authors introduce ZIP-SR and ZE-EDEN, two 4-bit recipes aimed at reducing second-moment distortion in AdamW. Both keep NF4 quantization for the first moment, with targeted stochastic rounding on the LM-head first moment late in training. In GPT- and Llama-style pretraining from 130M to 2.7B parameters, both methods narrow TorchAO 4-bit AdamW’s validation-loss gap to 32-bit AdamW at every tested size. The largest reported gap reduction is 70%, and supervised fine-tuning results also beat TorchAO while staying close to 32-bit AdamW. ArXiv · AI/CL/LG's note
The authors introduce ZIP-SR and ZE-EDEN, two 4-bit recipes aimed at reducing second-moment distortion in AdamW. Both keep NF4 quantization for the first moment, with targeted stochastic rounding on the LM-head first moment late in training. In GPT- and Llama-style pretraining from 130M to 2.7B parameters, both methods narrow TorchAO 4-bit AdamW’s validation-loss gap to 32-bit AdamW at every tested size. The largest reported gap reduction is 70%, and supervised fine-tuning results also beat TorchAO while staying close to 32-bit AdamW. ArXiv · AI/CL/LG's note
score 4