FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
FlashPrefill V2 claims production-oriented sparse prefill attention with very large speedups at 128K context.
The paper says the new version adds mean correction to control approximation error at high sparsity. It also rebuilds the sparse attention operator around PackGQA access, warp specialization, pingpong pipelining, and FP8 inference support. The authors say it can work with paged KV cache and continuous batching, including as an SGLang attention backend. On NVIDIA H20 GPUs, they report up to 47.26x FP8 and 27.19x BF16 speedups over FlashAttention-2 at 128K context. HF Daily Papers' note
The paper says the new version adds mean correction to control approximation error at high sparsity. It also rebuilds the sparse attention operator around PackGQA access, warp specialization, pingpong pipelining, and FP8 inference support. The authors say it can work with paged KV cache and continuous batching, including as an SGLang attention backend. On NVIDIA H20 GPUs, they report up to 47.26x FP8 and 27.19x BF16 speedups over FlashAttention-2 at 128K context. HF Daily Papers' note
score 5