Building a Production Greek-English Speech Recognizer
The team’s single-model runs could not satisfy all nine production gates at once.
Sophea was tested across Greek and English error rates, language ID, and hallucinations on non-speech audio. The paper says Greek noisy-traffic performance needed much longer dense domain exposure than English language-ID preservation could tolerate. A three-model ROVER ensemble passed all nine gates and cut overlapping-speech WER from 53.35% to 37.87%. The authors release methodology and results, but no model weights or training data. HF Daily Papers' note
Sophea was tested across Greek and English error rates, language ID, and hallucinations on non-speech audio. The paper says Greek noisy-traffic performance needed much longer dense domain exposure than English language-ID preservation could tolerate. A three-model ROVER ensemble passed all nine gates and cut overlapping-speech WER from 53.35% to 37.87%. The authors release methodology and results, but no model weights or training data. HF Daily Papers' note
score 4