ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
ReViV claims a single egocentric RGB video is enough to recover both the wearer’s motion and the surrounding 4D scene.
The paper presents a unified feed-forward model for reconstructing camera trajectory, gaze, full-body and hand motion, depth, and RGB dynamics together. It targets a gap in prior methods that depend on extra trajectory inputs or split scene and ego-motion into separate tasks. The authors report state-of-the-art accuracy and efficiency across several egocentric benchmarks, with competitive depth estimation and open-sourced code and models. HF Daily Papers' note
The paper presents a unified feed-forward model for reconstructing camera trajectory, gaze, full-body and hand motion, depth, and RGB dynamics together. It targets a gap in prior methods that depend on extra trajectory inputs or split scene and ego-motion into separate tasks. The authors report state-of-the-art accuracy and efficiency across several egocentric benchmarks, with competitive depth estimation and open-sourced code and models. HF Daily Papers' note
score 5