MassAlloc Attention: Let Attention Allocate Its Own Compute
MALA keeps full causal score access but spends less work after softmax on low-contribution attention mass.
The paper frames the method as a fused attention primitive using normalized contribution to decide what post-score computation to retain. In matched-work 8K tests, it nearly reaches a per-instance reference-mass oracle, with mean omitted mass of 0.0188% versus 0.0182%. At 128K tokens, the authors report 2.2x forward, 3.0x backward, and 1.6x decoding latency reductions versus FullAttn. Scaling runs from 0.6B to 14B parameters, plus continued 14B and 32B models, are reported as closely tracking FullAttn on perplexity and evaluated capabilities while using fewer training FLOPs. HF Daily Papers' note
The paper frames the method as a fused attention primitive using normalized contribution to decide what post-score computation to retain. In matched-work 8K tests, it nearly reaches a per-instance reference-mass oracle, with mean omitted mass of 0.0188% versus 0.0182%. At 128K tokens, the authors report 2.2x forward, 3.0x backward, and 1.6x decoding latency reductions versus FullAttn. Scaling runs from 0.6B to 14B parameters, plus continued 14B and 32B models, are reported as closely tracking FullAttn on perplexity and evaluated capabilities while using fewer training FLOPs. HF Daily Papers' note
score 5