LOCI: Spatial Linear Memory for Streaming World Models
LOCI keeps revisited scenes accessible without letting memory grow with the full video.
The paper proposes a hybrid memory design for streaming video world models: some transformer blocks keep key-value caches, while others use current-chunk attention plus recurrent linear-attention memory. That recurrent memory is conditioned on projective camera geometry, so viewpoint affects both where memory is read and what is stored. On MIND and held-out recorded trajectories, LOCI is reported to reproduce revisited content more faithfully than representative world models and a same-recipe full-softmax model. With full history, it cuts peak memory by about 30% versus full softmax at equal length; with a bounded observation bank, it streams long videos at constant memory.
Source: HF Daily Papers' note
The paper proposes a hybrid memory design for streaming video world models: some transformer blocks keep key-value caches, while others use current-chunk attention plus recurrent linear-attention memory. That recurrent memory is conditioned on projective camera geometry, so viewpoint affects both where memory is read and what is stored. On MIND and held-out recorded trajectories, LOCI is reported to reproduce revisited content more faithfully than representative world models and a same-recipe full-softmax model. With full history, it cuts peak memory by about 30% versus full softmax at equal length; with a bounded observation bank, it streams long videos at constant memory.
Source: HF Daily Papers' note
score 5