G^2PTQ: Improving LLM Post-Training Quantization with Generalized Gradient Compensation
G²PTQ updates its gradient and Hessian guidance block by block, aiming to keep quantization aligned with the full-precision model.
The paper targets post-training quantization for LLMs, where GPTQ-style methods can lose accuracy from either local-only objectives or stale global estimates. Its generalized gradient compensation combines first- and second-order information under a globally supervised block-wise objective. A trust-region scaling step is added to keep exact gradient compensation from producing unstable weight updates. The authors report stronger results than state-of-the-art baselines across model families and bit widths. HF Daily Papers' note
The paper targets post-training quantization for LLMs, where GPTQ-style methods can lose accuracy from either local-only objectives or stale global estimates. Its generalized gradient compensation combines first- and second-order information under a globally supervised block-wise objective. A trust-region scaling step is added to keep exact gradient compensation from producing unstable weight updates. The authors report stronger results than state-of-the-art baselines across model families and bit widths. HF Daily Papers' note
score 4