Pretraining Transformers with Quantized Softmax in Attention
The paper tests where quantized softmax actually breaks during pretraining, and finds the backward path matters as much as the approximation.
The authors study K-interval attention, an approximate softmax that replaces exponentials with grid values, under matched pretraining settings. Detaching row extrema keeps the forward pass the same but later raises validation loss. At K=4, hard rounding with min-max calibration and a pre-normalization surrogate performs poorly, while changing the calibration or surrogate placement narrows the gap. In their 124M-parameter, 2.5B-token run, fixed-window calibration gets within +0.019 nats at K=4 and +0.004 nats at K=16 versus standard softmax. HF Daily Papers' note
The authors study K-interval attention, an approximate softmax that replaces exponentials with grid values, under matched pretraining settings. Detaching row extrema keeps the forward pass the same but later raises validation loss. At K=4, hard rounding with min-max calibration and a pre-normalization surrogate performs poorly, while changing the calibration or surrogate placement narrows the gap. In their 124M-parameter, 2.5B-token run, fixed-window calibration gets within +0.019 nats at K=4 and +0.004 nats at K=16 versus standard softmax. HF Daily Papers' note
score 4