GigaAM Multilingual: Foundation Model for Underrepresented Languages
GigaAM Multilingual targets Kazakh, Kyrgyz, and Uzbek ASR with balancing methods meant to keep larger languages from dominating training.
The paper presents a Conformer encoder pre-trained on 2 million hours of audio with a HuBERT-style objective. Its pre-training uses cluster-level data balancing, followed by domain-aware sampling during fine-tuning. In controlled tests, the authors say it beats Whisper Large v3 and Omnilingual-1B on the target languages, with stronger gains on spontaneous speech. The foundation encoder and ASR model are being released. HF Daily Papers' note
The paper presents a Conformer encoder pre-trained on 2 million hours of audio with a HuBERT-style objective. Its pre-training uses cluster-level data balancing, followed by domain-aware sampling during fine-tuning. In controlled tests, the authors say it beats Whisper Large v3 and Omnilingual-1B on the target languages, with stronger gains on spontaneous speech. The foundation encoder and ASR model are being released. HF Daily Papers' note
score 5