Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
Scal3R attacks long-video collapse by changing pose estimation, not the depth backbone.
The paper says online 3D reconstruction fails on long videos because poses are tied to a fixed first-frame anchor, pushing the model outside its training range. Scal3R instead queries relative poses against multiple past keyframes using small learnable tokens added to a frozen backbone. It adds online pose-graph optimization with loop closure to limit drift. The authors report convergence in 8 hours on one GPU and more than 60% lower average ATE on KITTI versus the online baseline. HF Daily Papers' note
The paper says online 3D reconstruction fails on long videos because poses are tied to a fixed first-frame anchor, pushing the model outside its training range. Scal3R instead queries relative poses against multiple past keyframes using small learnable tokens added to a frozen backbone. It adds online pose-graph optimization with loop closure to limit drift. The authors report convergence in 8 hours on one GPU and more than 60% lower average ATE on KITTI versus the online baseline. HF Daily Papers' note
score 4