Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training
Prism claims a 2.5x training speedup over full attention while improving generation quality for native 2K video-audio models.
The paper targets the cost of full attention when jointly training high-resolution video and audio generation systems. Prism groups tokens into spatiotemporal macro-zones, then adapts each zone’s attention blocks based on visual feature variation and audio-to-video cross-attention strength. The method is designed to keep blocks semantically coherent, especially around regions where sound and image content are coupled. Its experiments report faster training than full attention and better generated output quality. HF Daily Papers' note
The paper targets the cost of full attention when jointly training high-resolution video and audio generation systems. Prism groups tokens into spatiotemporal macro-zones, then adapts each zone’s attention blocks based on visual feature variation and audio-to-video cross-attention strength. The method is designed to keep blocks semantically coherent, especially around regions where sound and image content are coupled. Its experiments report faster training than full attention and better generated output quality. HF Daily Papers' note
score 5