Megadose Built for builders and researchers.

RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

· HF Daily Papers ·
RLHND uses a video diffusion backbone to estimate both hand pose and tactile signals from egocentric video.

The paper says the model adapts the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor for hand tracking. It predicts anatomically plausible hand poses and can condition on shape parameters to keep hand shape consistent across clips or actors. A separate tactile stream estimates dense contact and force over the hand surface. The authors report state-of-the-art results on pose, contact, and force benchmarks, plus retargeting and real-world robot experiments. HF Daily Papers' note

score 5

Categories: Research