WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
WUSH-KV targets the KV-cache bottleneck by learning separate transforms for keys and values before low-bit quantization.
The paper says KV cache memory and bandwidth rise with context length and batch size, making long-context inference harder to run efficiently. WUSH-KV builds data-adaptive transforms from calibration data, folding the value transform into model weights and applying the key transform after RoPE. In the authors’ tests, the method reduces layerwise reconstruction error and gives the lowest end-to-end perplexity among the compared transforms. Integrated into SGLang at 2-bit, it performs comparably to or better than OSCAR across the evaluated models and tasks. ArXiv · AI/CL/LG's note
The paper says KV cache memory and bandwidth rise with context length and batch size, making long-context inference harder to run efficiently. WUSH-KV builds data-adaptive transforms from calibration data, folding the value transform into model weights and applying the key transform after RoPE. In the authors’ tests, the method reduces layerwise reconstruction error and gives the lowest end-to-end perplexity among the compared transforms. Integrated into SGLang at 2-bit, it performs comparably to or better than OSCAR across the evaluated models and tasks. ArXiv · AI/CL/LG's note
score 5