RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning
RLHND uses a video diffusion backbone to estimate both hand pose and tactile signals from egocentric video.
The paper says the model adapts the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor for hand tracking. It predicts anatomically plausible hand poses and can condition on shape parameters to keep hand shape consistent across clips or actors. A separate tactile stream estimates dense contact and force over the hand surface. The authors report state-of-the-art results on pose, contact, and force benchmarks, plus retargeting and real-world robot experiments. HF Daily Papers' note
The paper says the model adapts the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor for hand tracking. It predicts anatomically plausible hand poses and can condition on shape parameters to keep hand shape consistent across clips or actors. A separate tactile stream estimates dense contact and force over the hand surface. The authors report state-of-the-art results on pose, contact, and force benchmarks, plus retargeting and real-world robot experiments. HF Daily Papers' note
score 5