DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization
The paper’s core claim is that speech-quality models improve when they learn where the audible damage is, not just the final MOS.
The authors introduce a partially distorted speech dataset with frame-level distortion annotations. They train a localization model to produce distortion cues, then feed those cues into DAMOS for MOS prediction. On multiple public benchmarks, they report consistent gains over existing methods and stronger cross-dataset generalization. Source: ArXiv · AI/CL/LG's note.
The authors introduce a partially distorted speech dataset with frame-level distortion annotations. They train a localization model to produce distortion cues, then feed those cues into DAMOS for MOS prediction. On multiple public benchmarks, they report consistent gains over existing methods and stronger cross-dataset generalization. Source: ArXiv · AI/CL/LG's note.
score 4