SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation
SepRQ targets speech models that have to understand overlapping speakers, and reports benchmark gains without relying on masked prediction.
The paper presents an open-source self-supervised framework built around pseudo-source separation over frozen random-projection codebooks. Its authors say SepRQ reaches state-of-the-art results on SUPERB speaker diarization and speech separation, beating WavLM and other cocktail-party SSL systems at Base and Large scales. They also report strong results on target-speaker tasks, DIHARD 3 diarization, and three-speaker WSJ0-3Mix separation. ArXiv · AI/CL/LG's note
The paper presents an open-source self-supervised framework built around pseudo-source separation over frozen random-projection codebooks. Its authors say SepRQ reaches state-of-the-art results on SUPERB speaker diarization and speech separation, beating WavLM and other cocktail-party SSL systems at Base and Large scales. They also report strong results on target-speaker tasks, DIHARD 3 diarization, and three-speaker WSJ0-3Mix separation. ArXiv · AI/CL/LG's note
score 6