Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation
GEAR uses geometry to point attention at the right past video latents, instead of building a persistent 3D scene memory.
The paper frames long-horizon camera-controlled video as a memory problem: the model must recover content it has already seen as the camera moves. GEAR keeps historical observations as frame latents, then uses per-frame geometry to form token-level correspondences with the target view. Its Geometric Correspondence Attention injects matched historical features during denoising, while an Invisible Octree filters out correspondences that should be occluded. The authors report state-of-the-art visual quality, camera control, and revisit consistency for minute-long generated videos. Source: HF Daily Papers' note.
The paper frames long-horizon camera-controlled video as a memory problem: the model must recover content it has already seen as the camera moves. GEAR keeps historical observations as frame latents, then uses per-frame geometry to form token-level correspondences with the target view. Its Geometric Correspondence Attention injects matched historical features during denoising, while an Invisible Octree filters out correspondences that should be occluded. The authors report state-of-the-art visual quality, camera control, and revisit consistency for minute-long generated videos. Source: HF Daily Papers' note.
score 5