KVAE: Family of Tokenizers for Multimodal Generative Models
KVAE packages audio, image, and video tokenizers for later text-conditioned generation.
The report describes KVAE-Audio at 48 kHz, two causal KVAE-3D video tokenizers, and a KVAE-2D image tokenizer. The authors say their reconstruction and generation metrics match or beat open-source tokenizer baselines including Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio, and MMAudio. They also include training details, model selection methods, and ablations on design choices. Code is listed as publicly available. HF Daily Papers' note
The report describes KVAE-Audio at 48 kHz, two causal KVAE-3D video tokenizers, and a KVAE-2D image tokenizer. The authors say their reconstruction and generation metrics match or beat open-source tokenizer baselines including Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio, and MMAudio. They also include training details, model selection methods, and ablations on design choices. Code is listed as publicly available. HF Daily Papers' note
score 5