Paradee: Distilling Kokoro-82M into an 8M-Parameter Single-Voice Text-to-Speech Model
An 8.5 MB int8 TTS model matches much of its 82M-parameter teacher’s output for a single voice.
Paradee is distilled from Kokoro-82M by narrowing the same architecture and training its text side and decoder separately against the frozen teacher.
The paper says it uses 10x fewer parameters, 15x less compute, and runs 25x faster than real time on one CPU thread.
Its reported UTMOS score is 4.41, compared with Kokoro’s 4.52.
A post-synthesis phase-locking filter is used to reduce a slight buzz traced to voiced speech between 2 and 8 kHz.
ArXiv · AI/CL/LG's note
Paradee is distilled from Kokoro-82M by narrowing the same architecture and training its text side and decoder separately against the frozen teacher.
The paper says it uses 10x fewer parameters, 15x less compute, and runs 25x faster than real time on one CPU thread.
Its reported UTMOS score is 4.41, compared with Kokoro’s 4.52.
A post-synthesis phase-locking filter is used to reduce a slight buzz traced to voiced speech between 2 and 8 kHz.
ArXiv · AI/CL/LG's note
score 5