ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals
ResidualQuant cuts looped-transformer KV storage by 80.7% while staying near BF16 accuracy in mixed precision.
The paper targets a memory problem specific to looped Transformers: KV cache size still grows with each recurrent loop. Its method stores final-loop KV states as the reference and encodes earlier loops as low-precision residuals, reaching INT2 with added scaling, rotations, and loop-wise mixed precision. The authors report up to 13.0% higher accuracy than a rotation-based KV-quantization baseline at the same memory budget. On an RTX 5090, they claim decode throughput gains up to 2.73x at fixed batch size and up to 4.15x peak throughput from larger batches. ArXiv · AI/CL/LG's note
The paper targets a memory problem specific to looped Transformers: KV cache size still grows with each recurrent loop. Its method stores final-loop KV states as the reference and encodes earlier loops as low-precision residuals, reaching INT2 with added scaling, rotations, and loop-wise mixed precision. The authors report up to 13.0% higher accuracy than a rotation-based KV-quantization baseline at the same memory budget. On an RTX 5090, they claim decode throughput gains up to 2.73x at fixed batch size and up to 4.15x peak throughput from larger batches. ArXiv · AI/CL/LG's note
score 5