HelloWorld: Enabling Socially Interactive Characters in Video World Models
The paper’s core move is making video-world characters respond on cue to the viewer.
HelloWorld lets a user press a button and trigger an on-screen character to turn toward the camera, wave, nod, or speak a short greeting. The authors train it with a self-distillation pipeline using model-synthesized clips that combine social interactions with camera motion. At inference, a training-free module localizes the response to the button-press window by adjusting DiT cross-attention masks. They also introduce HelloWorldBench, a 400-sample benchmark for measuring interaction quality alongside standard video metrics. HF Daily Papers' note
HelloWorld lets a user press a button and trigger an on-screen character to turn toward the camera, wave, nod, or speak a short greeting. The authors train it with a self-distillation pipeline using model-synthesized clips that combine social interactions with camera motion. At inference, a training-free module localizes the response to the button-press window by adjusting DiT cross-attention masks. They also introduce HelloWorldBench, a 400-sample benchmark for measuring interaction quality alongside standard video metrics. HF Daily Papers' note
score 5