LatentIndex: Cross-Layer Sharing with Layer-Specific Selection for Sparse Attention
LatentIndex shares continuous indexer caches across layers while still letting each layer choose its own tokens.
The paper targets sparse-attention overhead from repeated index selection and per-layer key-cache storage. Its method builds a shared latent cache for each layer group, then uses layer-specific scoring instead of forcing all layers onto the same selected token set. In reported tests on DeepSeek-V3.2 and GLM-5, the training-free version keeps long-context benchmark performance close to native DSA while improving attention-mass recall over IndexCache. A hierarchical-selection variant trades some recall for 2.30-2.72x decode indexer speedups across 8K-128K contexts. ArXiv · AI/CL/LG's note
The paper targets sparse-attention overhead from repeated index selection and per-layer key-cache storage. Its method builds a shared latent cache for each layer group, then uses layer-specific scoring instead of forcing all layers onto the same selected token set. In reported tests on DeepSeek-V3.2 and GLM-5, the training-free version keeps long-context benchmark performance close to native DSA while improving attention-mass recall over IndexCache. A hierarchical-selection variant trades some recall for 2.30-2.72x decode indexer speedups across 8K-128K contexts. ArXiv · AI/CL/LG's note
score 5