DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
DreamX-Creator pairs video and audio generation inside one 7B model, then adds a one-step 2K refinement stage.
The paper says the system takes a first frame and text prompt, then jointly denoises separate audio and video streams before coupling them with gated cross-modal attention. Its training stack includes curated audio-video data, progressive joint training, high-quality finetuning, and reinforcement learning with modality-aware feedback. The authors claim performance competitive with state-of-the-art open-source systems, and say they are releasing both the compact generator and 2K refiner. Source: HF Daily Papers' note.
The paper says the system takes a first frame and text prompt, then jointly denoises separate audio and video streams before coupling them with gated cross-modal attention. Its training stack includes curated audio-video data, progressive joint training, high-quality finetuning, and reinforcement learning with modality-aware feedback. The authors claim performance competitive with state-of-the-art open-source systems, and say they are releasing both the compact generator and 2K refiner. Source: HF Daily Papers' note.
score 6