Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
ABot-Recon argues that strictly local memory can beat long-range memory for very long streaming 3D reconstruction.
The model caches KV features from only the previous 11 frames, then predicts a current-camera point map and adjacent-frame relative pose. It recovers global motion and geometry by sequential composition, with a lightweight refiner to reduce rotation drift. On Oxford Spires, the paper reports 4.35 m ATE and 0.12° RPE-R, about 40% lower than the best prior results cited. HF Daily Papers' note
The model caches KV features from only the previous 11 frames, then predicts a current-camera point map and adjacent-frame relative pose. It recovers global motion and geometry by sequential composition, with a lightweight refiner to reduce rotation drift. On Oxford Spires, the paper reports 4.35 m ATE and 0.12° RPE-R, about 40% lower than the best prior results cited. HF Daily Papers' note
score 5