EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation
A frozen TTS model can be steered more reliably by splitting emotion control into shared and emotion-specific directions.
The paper says EmoRES separates an emotion vector into a component that moves speech away from neutral and a residual that pushes toward the requested emotion. It does this without retraining the backbone model. On IEMOCAP, it beats CoCoEmo across four objective emotion metrics on IndexTTS-2 and CosyVoice2. Human tests also favored EmoRES on dominant-emotion recognition, fidelity, and naturalness in reported comparisons. ArXiv · AI/CL/LG's note
The paper says EmoRES separates an emotion vector into a component that moves speech away from neutral and a residual that pushes toward the requested emotion. It does this without retraining the backbone model. On IEMOCAP, it beats CoCoEmo across four objective emotion metrics on IndexTTS-2 and CosyVoice2. Human tests also favored EmoRES on dominant-emotion recognition, fidelity, and naturalness in reported comparisons. ArXiv · AI/CL/LG's note
score 4