Megadose AI progress, ranked and analyzed.

WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms

· HF Daily Papers ·
WUSH-KV targets the KV-cache cost of long-context inference by learning separate data-adaptive transforms for keys and values before low-bit quantization.

The method builds its transforms from calibration data, folding the value transform into model weights and applying the key transform after RoPE. The paper says pairing WUSH-KV with clipped quantizers reduces layerwise reconstruction error and gives the lowest tested end-to-end perplexity among compared transforms. In SGLang, using OSCAR-style percentile-clipped affine quantization, its 2-bit version matched or beat OSCAR across the evaluated models and downstream tasks. Source: HF Daily Papers' note.

score 5

Categories: Research