Megadose Built for builders and researchers.

SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference

· HF Daily Papers ·
SparseDecoding prunes against the tokens the model actually generates during decoding, then pairs that with a sparse matrix-vector kernel built for the decoding path.

The paper argues that pruning calibrated on fixed natural text misses the activation distribution seen when an LLM is generating its own tokens. SparseDecoding instead collects layer activations from dense-model autoregressive generation, excluding prefill, and uses those for calibration. The authors also add an optimized N:M sparse matrix-vector kernel using bitmask indexing and fixed-step traversal. On Llama, Qwen, and other tested models, they report better long-form generation results than fixed-text calibration and up to 1.48x end-to-end decoding speedup on A100 GPUs. HF Daily Papers' note

score 5

Categories: Research