ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation
ReMoMask-2 moves retrieval into the generator’s own latent space to narrow the gap between evidence and motion synthesis.
The paper says existing retrieval-augmented text-to-motion systems miss motion’s part-level, spatial-temporal structure and retrieve evidence in a separate semantic space. Its first framework, ReMoMask, adds hierarchical contrastive learning, topology-aware fusion, and structured masking for stronger grounding. ReMoMask-2 then rebuilds the retrieval database in the generator’s pre-quantization latent space and uses a lightweight distilled projector for text queries. In tests on HumanML3D, KIT-ML, and SnapMoGen, the authors report state-of-the-art retriever accuracy, lowest FID on KIT-ML and SnapMoGen, and faster inference from a single mask-transformer stage. HF Daily Papers' note
The paper says existing retrieval-augmented text-to-motion systems miss motion’s part-level, spatial-temporal structure and retrieve evidence in a separate semantic space. Its first framework, ReMoMask, adds hierarchical contrastive learning, topology-aware fusion, and structured masking for stronger grounding. ReMoMask-2 then rebuilds the retrieval database in the generator’s pre-quantization latent space and uses a lightweight distilled projector for text queries. In tests on HumanML3D, KIT-ML, and SnapMoGen, the authors report state-of-the-art retriever accuracy, lowest FID on KIT-ML and SnapMoGen, and faster inference from a single mask-transformer stage. HF Daily Papers' note
score 4