Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
JoyAI-Echo-1.5 is built to keep characters, voices, and camera-controlled worlds coherent over longer generated sequences.
The paper describes two variants: one for long videos using cross-shot memory and speaker cues, and one for interactive worlds using calibrated 6-DoF camera trajectories. Its training setup turns a bidirectional audio-visual backbone into a causal few-step generator for extended rollouts. The authors report gains over long-video baselines and a first-place WBench score of 81.7 for the world-model variant. Source: HF Daily Papers' note.
The paper describes two variants: one for long videos using cross-shot memory and speaker cues, and one for interactive worlds using calibrated 6-DoF camera trajectories. Its training setup turns a bidirectional audio-visual backbone into a causal few-step generator for extended rollouts. The authors report gains over long-video baselines and a first-place WBench score of 81.7 for the world-model variant. Source: HF Daily Papers' note.
score 6