Megadose AI progress, ranked and analyzed.

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

· ArXiv · AI/CL/LG ·
STEPQuant claims low-bit recurrent-state quantization can cut serving memory without giving up FP32-level accuracy.

The paper says linear-attention models still face a memory bottleneck because their fixed recurrent states persist during concurrent serving. Its method assigns precision by both how long an error will survive and where that error occurs in the state. In tests on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct, the authors report near-FP32-state accuracy at a nominal 6-bit budget, and better results than uniform INT8 in a 4-bit setup. With SGLang kernels, they report more than 5x recurrent-state compression and up to 68.7% lower total serving memory. ArXiv · AI/CL/LG's note

score 4

Categories: Research