LatentPress: Context Compression Beyond Text and Vision
LatentPress compresses context into soft memory tokens a frozen language model can read directly.
The paper reports 4-16x compression using a small writer adapter, without reconstructing text at inference. On LongMemEval, it reaches 0.504 accuracy at 7.70x compression, slightly above the uncompressed evidence score of 0.490. It also beats text summaries and OCR-style compression in the reported tests. Writing is listed at 43ms per conversation, with reading 5-9x faster than raw context or cached OCR. HF Daily Papers' note
The paper reports 4-16x compression using a small writer adapter, without reconstructing text at inference. On LongMemEval, it reaches 0.504 accuracy at 7.70x compression, slightly above the uncompressed evidence score of 0.490. It also beats text summaries and OCR-style compression in the reported tests. Writing is listed at 43ms per conversation, with reading 5-9x faster than raw context or cached OCR. HF Daily Papers' note
score 5