Megadose AI progress, ranked and analyzed.

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

· HF Daily Papers ·
Motion-Omni generates speech and full-body avatar motion in one model pass, instead of cascading audio into a separate motion system.

The paper says the model outputs facial expression, hand, upper-body, and lower-body motion from the same hidden states used to produce speech. Its training pipeline pseudo-labels 422,856 speech-motion pairs, totaling 1,402 hours. With a Qwen2.5-7B-Instruct backbone, Motion-Omni-Q7 comes within 2% of the same-audio teacher cascade on reference-free motion metrics while running 5.4x faster. The authors also introduce SwDA-500 and a public evaluation protocol for open-ended full-body spoken dialogue. Source: HF Daily Papers' note.

score 5

Categories: Research