Megadose AI progress, ranked and analyzed.

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

· HF Daily Papers ·
OmniVAE trains audio and video latents together so synchronization is built into the representation, not left entirely to the generator.

The paper argues that separate audio and video VAEs leave downstream models to learn cross-modal timing and semantics from scratch. OmniVAE adds segment-level audio-video contrastive learning to align the two latent spaces, while distilling features from pretrained modality-specific encoders. The authors report that these objectives improve latent learnability, generation quality, and audio-video synchronization in downstream text-to-audio-video generation. HF Daily Papers' note

score 5

Categories: Research