ProxyFormer: A Dual-Stream Proxy Architecture for Ultra-Long Context and High-Resolution Generation
ProxyFormer reports million-token retrieval from models trained on far shorter windows.
The paper proposes proxy tokens that compress local features, run global attention in the smaller proxy space, then inject the result back into the local stream. It says this avoids some one-shot compression loss because the local stream remains available across layers. In the reported setup, a 16GB GPU goes from about 20K trainable tokens for a standard decoder-only model to about 0.7M with a 64x compression ratio. The author also reports 92%-95% multi-needle retrieval accuracy at 1,048,576 tokens after 64K-window training, plus preliminary pixel- and latent-space image-generation tests. ArXiv · AI/CL/LG's note
The paper proposes proxy tokens that compress local features, run global attention in the smaller proxy space, then inject the result back into the local stream. It says this avoids some one-shot compression loss because the local stream remains available across layers. In the reported setup, a 16GB GPU goes from about 20K trainable tokens for a standard decoder-only model to about 0.7M with a 64x compression ratio. The author also reports 92%-95% multi-needle retrieval accuracy at 1,048,576 tokens after 64K-window training, plus preliminary pixel- and latent-space image-generation tests. ArXiv · AI/CL/LG's note
score 5