HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models
HLA-WM tries to keep long video rollouts consistent by retrieving only the past scene chunks that matter.
The paper says recurrent linear attention is efficient but can forget distant, relevant scenes as new state updates overwrite old information. HLA-WM adds a training-free hybrid layer that uses camera geometry to retrieve compact historical chunks, then rebuilds query-specific recurrent states. On 60-second SANA-WM-Bench runs, it improved all six aggregate revisit-consistency and camera-control metrics, including a 0.74 dB PSNR gain and 28.5% lower rotation error. The authors report a 12x reduction in historical-state memory versus full KV caching, with at most a 1.6% inference-throughput drop. HF Daily Papers' note
The paper says recurrent linear attention is efficient but can forget distant, relevant scenes as new state updates overwrite old information. HLA-WM adds a training-free hybrid layer that uses camera geometry to retrieve compact historical chunks, then rebuilds query-specific recurrent states. On 60-second SANA-WM-Bench runs, it improved all six aggregate revisit-consistency and camera-control metrics, including a 0.74 dB PSNR gain and 28.5% lower rotation error. The authors report a 12x reduction in historical-state memory versus full KV caching, with at most a 1.6% inference-throughput drop. HF Daily Papers' note
score 5