Real Long-Term Memory for AI: A 50-Million-Token Window That Is Faster and Cheaper Than Recompute
The test claims byte-exact KV-state reuse across a 50-million-token stream on one H100.
The paper tests `galahad-kv`, which saves roughly 16,000-token KV blocks to encrypted local NVMe storage and reloads them without recomputation. In the reported probes, both Gemma 4 12B and 31B reloaded 100 of 100 blocks from depths up to 50M tokens. Loads were 2.8x to 4.3x faster than recompute and used 8.8x to 12.3x less GPU energy, with flat GPU memory. The author cautions that this is stored-state reuse, not a broader attention window: one block is loaded at a time, and the store requires terabytes of local NVMe disk. HF Daily Papers' note
The paper tests `galahad-kv`, which saves roughly 16,000-token KV blocks to encrypted local NVMe storage and reloads them without recomputation. In the reported probes, both Gemma 4 12B and 31B reloaded 100 of 100 blocks from depths up to 50M tokens. Loads were 2.8x to 4.3x faster than recompute and used 8.8x to 12.3x less GPU energy, with flat GPU memory. The author cautions that this is stored-state reuse, not a broader attention window: one block is loaded at a time, and the store requires terabytes of local NVMe disk. HF Daily Papers' note
score 6