Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Ex-Omni-2D tries to make voice agents visibly present by generating synchronized avatar video alongside text and speech.
The framework uses a Visual Thought Plan to specify scene, emotion, and motion before producing response text and speech units. Those speech units are decoded into audio and aligned with video frames so the audio and avatar systems share timing. The paper also describes a streaming student model meant to reduce startup latency versus the full-sequence teacher at 400x720 or 720x400 output. HF Daily Papers' note
The framework uses a Visual Thought Plan to specify scene, emotion, and motion before producing response text and speech units. Those speech units are decoded into audio and aligned with video frames so the audio and avatar systems share timing. The paper also describes a streaming student model meant to reduce startup latency versus the full-sequence teacher at 400x720 or 720x400 output. HF Daily Papers' note
score 5