Megadose AI progress, ranked and analyzed.

LeapQuant: Efficient Linear Attention with Accurate Recurrent State Quantization

· ArXiv · AI/CL/LG ·
LeapQuant targets the recurrent-state bottleneck in linear-attention LLM inference with near-lossless 8-bit quantization.

The paper says hybrid models such as Gated DeltaNet and Kimi Delta Attention still spend heavily on repeatedly reading and updating their fixed-size recurrent state. LeapQuant reduces accumulated rounding error by quantizing only at token-window boundaries while keeping buffered updates in higher precision inside the window. It also preserves the largest state outliers as high-precision “Compensator Tokens,” then smooths the remaining residual before quantization. Across Qwen, Kimi, and GLM model families, the authors report FP32-comparable accuracy, 2.05–3.70x kernel speedups, and 1.47x end-to-end inference speedup on tested NVIDIA GPUs. ArXiv · AI/CL/LG's note

score 4

Categories: Research