TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration
TileMix routes attention tiles through FP16 or INT8 during prefill while keeping dense attention intact.
The paper frames precision as a per-tile execution choice inside fused dense attention, using compact bitmasks to decide which score-tile groups use each path. Both paths update the same online-softmax state, and the method keeps token connectivity instead of pruning interactions. The authors say it needs no training and supports grouped-query attention, variable-length batches, and INT8 key/value caches. In LongEval, LV-Eval, and A100 prefill tests on LLaMA, Qwen, and Vicuna, it recovered quality lost with uniform INT8 while improving prefill throughput over FP16. HF Daily Papers' note
The paper frames precision as a per-tile execution choice inside fused dense attention, using compact bitmasks to decide which score-tile groups use each path. Both paths update the same online-softmax state, and the method keeps token connectivity instead of pruning interactions. The authors say it needs no training and supports grouped-query attention, variable-length batches, and INT8 key/value caches. In LongEval, LV-Eval, and A100 prefill tests on LLaMA, Qwen, and Vicuna, it recovered quality lost with uniform INT8 while improving prefill throughput over FP16. HF Daily Papers' note
score 5