Megadose Built for builders and researchers.

Memorizon: Training World Models Beyond Their Context Window

· HF Daily Papers ·
Memorizon trains revisit consistency over long spans while keeping attention bounded.

The paper separates supervision length from attention length: samples can span any duration, but loss is applied only to the last `k` chunks. Each scored chunk retrieves top-`K` latent memories by camera co-visibility, forming a shared bank capped at `kK`. The authors report that extending span from 100 to 400 seconds adds 12% step time, while longer spans improve revisit consistency versus a sliding-window baseline. They also find that using a bank from another episode cuts revisit correlation by 83%, evidence the model is relying on retrieved memories. HF Daily Papers' note

score 5

Categories: Research