Megadose AI progress, ranked and analyzed.

REALM: A Coarse-to-Fine Generative Framework for Embodied Reactive Listening

· HF Daily Papers ·
REALM models a robot listener’s face by aligning delayed speaker cues with ongoing listener motion, then adding stochastic expression details.

The paper targets responsive facial motion for embodied conversational AI, including brief expressions and blinks that are hard to predict deterministically. Its framework fuses listener history with speaker audio using a delay-centered attention prior and adaptive gating, then refines a coarse motion path with audio-conditioned residuals. Tests on ViCo and L2L beat the evaluated baselines across multiple motion-quality metrics. The authors also report deployment on an Ameca humanoid robot and a perceptual user study. HF Daily Papers' note

score 4

Categories: Research