All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation
RecCAR targets a one-way weakness in multimodal diffusion models: video influences companion streams more strongly than they influence video.
The paper defines that mismatch as a “reciprocal correspondence gap” across cross-attention over video tokens. Its proposed regularizer, RecCAR, aligns weaker modality-to-video attention with the stronger video-to-modality correspondence. In video-motion and video-audio generation tests, it raises the Human Anatomy score from 0.69 to 0.75 and reduces audio-video desynchronization from 0.804 to 0.752. Source: HF Daily Papers' note.
The paper defines that mismatch as a “reciprocal correspondence gap” across cross-attention over video tokens. Its proposed regularizer, RecCAR, aligns weaker modality-to-video attention with the stronger video-to-modality correspondence. In video-motion and video-audio generation tests, it raises the Human Anatomy score from 0.69 to 0.75 and reduces audio-video desynchronization from 0.804 to 0.752. Source: HF Daily Papers' note.
score 5