Megadose Built for builders and researchers.

SepRQ : Self-Supervised Speech Mixture Representation Learning via Mask-Free, Multi-Scale Source Separation

· ArXiv · AI/CL/LG ·
SepRQ targets speech models that have to understand overlapping speakers, and reports benchmark gains without relying on masked prediction.

The paper presents an open-source self-supervised framework built around pseudo-source separation over frozen random-projection codebooks. Its authors say SepRQ reaches state-of-the-art results on SUPERB speaker diarization and speech separation, beating WavLM and other cocktail-party SSL systems at Base and Large scales. They also report strong results on target-speaker tasks, DIHARD 3 diarization, and three-speaker WSJ0-3Mix separation. ArXiv · AI/CL/LG's note

score 6

Categories: OSS & Tools, Research