InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
InfiniHand estimates hand pose and camera motion together from uncalibrated egocentric video, aiming to cut drift and pipeline overhead.
The paper presents a streaming feed-forward system that predicts MANO hand parameters, camera trajectories, and hand locations in one architecture. It uses persistent spatiotemporal memory and hand-centered visual features to connect local hand geometry with camera motion. The authors say it was trained in two stages, backed by about 5,000 hours of egocentric data from public datasets. In evaluations, it reports a 21.4% ARCTIC PA-p reduction versus ViDiHand and runs at 11.19 FPS, more than twice HaWoR’s throughput. HF Daily Papers' note
The paper presents a streaming feed-forward system that predicts MANO hand parameters, camera trajectories, and hand locations in one architecture. It uses persistent spatiotemporal memory and hand-centered visual features to connect local hand geometry with camera motion. The authors say it was trained in two stages, backed by about 5,000 hours of egocentric data from public datasets. In evaluations, it reports a 21.4% ARCTIC PA-p reduction versus ViDiHand and runs at 11.19 FPS, more than twice HaWoR’s throughput. HF Daily Papers' note
score 4