Megadose AI progress, ranked and analyzed.

Training-Free Speech-Centric Omni Understanding with Frozen VLMs

· HF Daily Papers ·
The paper claims frozen vision-language models can gain strong speech-centered audio-visual ability by routing Whisper transcripts through their existing text interface.

TFO keeps the VLM architecture and visual pathway unchanged, using confidence-filtered, timestamped speech transcripts instead of new audio encoders. In matched tests against native omni models, it was competitive across 56 benchmarks and 21 languages. The authors report better average audio-only performance in all five model settings and stronger multilingual speech results. They also say freezing the backbone generally preserved more of the original model’s image, video, grounding, coding, math, and medical QA ability. HF Daily Papers' note

score 5

Categories: Research