Megadose AI progress, ranked and analyzed.

GigaAM Multilingual: Foundation Model for Underrepresented Languages

· HF Daily Papers ·
GigaAM Multilingual targets Kazakh, Kyrgyz, and Uzbek ASR with balancing methods meant to keep larger languages from dominating training.

The paper presents a Conformer encoder pre-trained on 2 million hours of audio with a HuBERT-style objective. Its pre-training uses cluster-level data balancing, followed by domain-aware sampling during fine-tuning. In controlled tests, the authors say it beats Whisper Large v3 and Omnilingual-1B on the target languages, with stronger gains on spontaneous speech. The foundation encoder and ASR model are being released. HF Daily Papers' note

score 5

Categories: Model Releases, Research