Megadose Built for builders and researchers.

VFold: Symmetry-Aware Cross-Layer Value Cache Compression

· ArXiv · AI/CL/LG ·
VFold targets LLM decoding memory by merging similar value-cache states across layers without changing the model architecture.

The paper says KV caches can dominate memory at long context lengths. VFold focuses on the value cache, using symmetry-aware merging to cut memory while limiting performance loss and decoding overhead. The authors also report that it can be combined with quantization or key-cache pruning for higher compression than either approach alone. Source: ArXiv · AI/CL/LG's note

score 5

Categories: Research