SANTA++: Sampling Attention through Representative Keys
SANTA++ cuts KV-cache reads by sampling representative key teams, then corrects the estimate with importance weighting.
The paper says the method is training-free and avoids scanning the full cache before choosing what to attend to. With 32 or 64 sampled teams, it uses 16% to 22% of dense attention’s KV reads while keeping 94% to 99% of baseline scores on LongBench v2 and HELMET’s RAG subset. On RULER, retention is lower, at 85% to 91%. A GPU implementation with 31 sampled teams reports a 1.69x attention speedup over dense FlashAttention at 32K context. ArXiv · AI/CL/LG's note
The paper says the method is training-free and avoids scanning the full cache before choosing what to attend to. With 32 or 64 sampled teams, it uses 16% to 22% of dense attention’s KV reads while keeping 94% to 99% of baseline scores on LongBench v2 and HELMET’s RAG subset. On RULER, retention is lower, at 85% to 91%. A GPU implementation with 31 sampled teams reports a 1.69x attention speedup over dense FlashAttention at 32K context. ArXiv · AI/CL/LG's note
score 6