Megadose AI progress, ranked and analyzed.

Disaggregated Quantization: Specializing LLM Prefill and Decode

· HF Daily Papers ·
The paper argues that quantization should be split by inference phase, not treated as one format for the whole run.

Its “disaggregated quantization” approach uses different compute formats, weights, and placement for prefill and decode. On Qwen 3 and Gemma 3, removing activation quantization during decode improved accuracy on decode-heavy tasks without adding inference cost. The authors also report that a separate NVFP4 prefiller improved 1-bit accuracy on MMLU-Pro and MMMU-Pro for a Qwen3.8-27B GGUF decoder, without changing the decode checkpoint. Their offloaded prefill setup streamed the extra weights from SSD and showed a 1.78x time-to-first-token speedup at an 8K prompt length. HF Daily Papers' note

score 5

Categories: Research