Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering
The paper proposes a training-free way to stop autoregressive video renderers from changing a scene when the camera returns to it.
The method uses 3D-engine correspondences already available from pose and depth. It retrieves earlier latent chunks into the KV cache as loop-closure memory, then biases attention toward geometrically matching regions. The authors test on loop-closure trajectories from TartanAir and TartanGround and report better revisit consistency than other training-free baselines without reducing overall video quality. HF Daily Papers' note
The method uses 3D-engine correspondences already available from pose and depth. It retrieves earlier latent chunks into the KV cache as loop-closure memory, then biases attention toward geometrically matching regions. The authors test on loop-closure trajectories from TartanAir and TartanGround and report better revisit consistency than other training-free baselines without reducing overall video quality. HF Daily Papers' note
score 4