Softmax Reparameterization for Output-Head Quantization
The paper targets the output head as a quantization bottleneck in small language models.
The authors propose a post-training softmax reparameterization that changes the output head before quantization while preserving the full-precision softmax distribution. It searches one scalar coefficient by validation KL for several quantizers, including RTN, activation-weighted MSE, and GPTQ. Reported W4 gains are largest where baseline quantization badly shifts predictions, including Phi-4-mini with AW-MSE KL dropping from 0.936 to 0.256. For compatible heads, the method adds no inference operation and preserves packed W4 execution, with a reported 10.8% batch-one latency reduction versus a BF16 output head. HF Daily Papers' note
The authors propose a post-training softmax reparameterization that changes the output head before quantization while preserving the full-precision softmax distribution. It searches one scalar coefficient by validation KL for several quantizers, including RTN, activation-weighted MSE, and GPTQ. Reported W4 gains are largest where baseline quantization badly shifts predictions, including Phi-4-mini with AW-MSE KL dropping from 0.936 to 0.256. For compatible heads, the method adds no inference operation and preserves packed W4 execution, with a reported 10.8% batch-one latency reduction versus a BF16 output head. HF Daily Papers' note
score 4