Steerspeech: Activation Steering For Emotion Control In Generated Speech
SteerSpeech controls generated speech emotion at inference time by steering hidden activations while leaving the TTS model frozen.
The paper trains lightweight low-rank transforms for target emotions, aiming to change emotional intensity without losing speaker identity or linguistic content. It uses a two-pass generation-and-replay setup so supervision can pass through sampled speech tokens. Tests with Qwen3-TTS report stronger continuous emotion control across seen, unseen, and accented speakers, with limited degradation. The submission is under review at IEEE ICASSP 2027. ArXiv · AI/CL/LG's note
The paper trains lightweight low-rank transforms for target emotions, aiming to change emotional intensity without losing speaker identity or linguistic content. It uses a two-pass generation-and-replay setup so supervision can pass through sampled speech tokens. Tests with Qwen3-TTS report stronger continuous emotion control across seen, unseen, and accented speakers, with limited degradation. The submission is under review at IEEE ICASSP 2027. ArXiv · AI/CL/LG's note
score 5