A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models
The paper offers a 2D input layer that lets 3D-pretrained motion language models work on monocular video without retraining.
The authors say accurate 3D motion recovery from single-camera video is a practical bottleneck for real-world MoLM use. Their plug-and-play interface lets existing models accept 2D motion inputs while keeping the original models unchanged. On public datasets, it performs close to 3D-input setups across multiple MoLMs and beats training MoLMs from scratch on 2D motion. They also add a monocular real-video evaluation set and report that, under their pose-estimation setting, 2D motions were more useful than 3D motions. HF Daily Papers' note
The authors say accurate 3D motion recovery from single-camera video is a practical bottleneck for real-world MoLM use. Their plug-and-play interface lets existing models accept 2D motion inputs while keeping the original models unchanged. On public datasets, it performs close to 3D-input setups across multiple MoLMs and beats training MoLMs from scratch on 2D motion. They also add a monocular real-video evaluation set and report that, under their pose-estimation setting, 2D motions were more useful than 3D motions. HF Daily Papers' note
score 4