Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
Salt++ targets the context mismatch that hurts few-step causal audio-video generation.
The paper proposes a two-stage post-training setup: Causal Self-Flow, then context-aligned autoregressive DMD. CSF trains a noisy-history student against a clean-history EMA teacher to improve semantic extraction and cross-modal alignment. The DMD stage keeps the causal mask and prefix consistent across sampling, fake-score training, and real-score evaluation. The authors report gains over OmniForcing at 480p under a 4-step causal setting, and a scale-wise stage for 4-step 1664x960 generation. HF Daily Papers' note
The paper proposes a two-stage post-training setup: Causal Self-Flow, then context-aligned autoregressive DMD. CSF trains a noisy-history student against a clean-history EMA teacher to improve semantic extraction and cross-modal alignment. The DMD stage keeps the causal mask and prefix consistent across sampling, fake-score training, and real-score evaluation. The authors report gains over OmniForcing at 480p under a 4-step causal setting, and a scale-wise stage for 4-step 1664x960 generation. HF Daily Papers' note
score 4