Megadose AI progress, ranked and analyzed.

Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost

· HF Daily Papers ·
Galahad caches byte-identical KV state so repeated document reads can be skipped instead of recomputed.

The paper says 98.7% of prompt tokens across seven real-world datasets were text the model had already read. Its Taliesin layer saves and reloads KV state for matching text, while Blaise stores documents and sends only the needed section. In a 97,000-token recall test, Galahad raised performance from 10/100 answers without it to 98/100 with Taliesin alone, and 100/100 with Blaise added. The authors report bit-identical restored logits and say failed cache loads fall back to recomputation. HF Daily Papers' note

score 6

Categories: OSS & Tools, Research