Retrieval Capacity of Self-Attention Under Competition
The paper estimates how many attention-selected tokens a model needs before its loss starts to rise.
The authors prune self-attention to the highest-weight tokens for each head, layer, and query, then measure how much NLL worsens. Small retained sets can stay close to full attention, but the needed size changes by model and grows as context is extended. Extra background can push a fixed supporting fact lower in attention rank and reduce its attention mass. Renormalizing the kept weights lowers the required set size, suggesting retrieval capacity depends on both token selection and aggregation. HF Daily Papers' note
The authors prune self-attention to the highest-weight tokens for each head, layer, and query, then measure how much NLL worsens. Small retained sets can stay close to full attention, but the needed size changes by model and grows as context is extended. Extra background can push a fixed supporting fact lower in attention rank and reduce its attention mass. Renormalizing the kept weights lowers the required set size, suggesting retrieval capacity depends on both token selection and aggregation. HF Daily Papers' note
score 4