Megadose AI progress, ranked and analyzed.

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis

· ArXiv · AI/CL/LG ·
GrainSpeech argues that tiny speech models do not need attention over long phoneme spans to preserve output quality.

The paper reports no consistent pitch, energy, or duration gains once self-attention extends beyond 15 phonemes. Its fixed-receptive-field convolutional encoder cuts those prediction errors by 36.0%, 17.3%, and 3.4%, respectively. The authors also replace a transferred image-domain gradient-variance loss with a Mel-specific version to recover fine detail without hurting predicted quality. GrainSpeech has 264.8K parameters and runs Mel generation at 17.9x real time on an MCU, with UTMOS scores comparable to much larger models.

ArXiv · AI/CL/LG's note

score 4

Categories: Research