Megadose AI progress, ranked and analyzed.

Pretraining Transformers with Quantized Softmax in Attention

· HF Daily Papers ·
The paper tests where quantized softmax actually breaks during pretraining, and finds the backward path matters as much as the approximation.

The authors study K-interval attention, an approximate softmax that replaces exponentials with grid values, under matched pretraining settings. Detaching row extrema keeps the forward pass the same but later raises validation loss. At K=4, hard rounding with min-max calibration and a pre-normalization surrogate performs poorly, while changing the calibration or surrogate placement narrows the gap. In their 124M-parameter, 2.5B-token run, fixed-window calibration gets within +0.019 nats at K=4 and +0.004 nats at K=16 versus standard softmax. HF Daily Papers' note

score 4

Categories: Research