Megadose Built for builders and researchers.

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

· HF Daily Papers ·
TileMix routes attention tiles through FP16 or INT8 during prefill while keeping dense attention intact.

The paper frames precision as a per-tile execution choice inside fused dense attention, using compact bitmasks to decide which score-tile groups use each path. Both paths update the same online-softmax state, and the method keeps token connectivity instead of pruning interactions. The authors say it needs no training and supports grouped-query attention, variable-length batches, and INT8 key/value caches. In LongEval, LV-Eval, and A100 prefill tests on LLaMA, Qwen, and Vicuna, it recovered quality lost with uniform INT8 while improving prefill throughput over FP16. HF Daily Papers' note

score 5

Categories: Research