Do speech foundation models really learn words?
HuBERT and wav2vec 2.0 appear to encode word identity beyond just phoneme patterns in later layers.
The paper tests whether strong word discrimination in speech foundation models is only a byproduct of encoding word form. Huo and Dunbar use residualization to partial out phoneme information, then check what remains. They report that later-layer representations still preserve words with reasonable fidelity, independent of local phonetic content. The same disentangling step can improve word discovery tasks by making higher-order linguistic information clearer. ArXiv · AI/CL/LG's note
The paper tests whether strong word discrimination in speech foundation models is only a byproduct of encoding word form. Huo and Dunbar use residualization to partial out phoneme information, then check what remains. They report that later-layer representations still preserve words with reasonable fidelity, independent of local phonetic content. The same disentangling step can improve word discovery tasks by making higher-order linguistic information clearer. ArXiv · AI/CL/LG's note
score 4