Steering Speech-Language Models: Training-Free Task Specialization via Contrastive Activation Addition
A training-free steering method improved SpeechLLM task behavior by adding contrastive activation vectors at inference time.
The paper derives those vectors from a small set of labeled utterances for tasks such as transcription.
The authors report that the approach improves task enforcement and, when paired with prompts, usually beats prompting alone on evaluated tasks including ASR and emotion recognition.
They also say the steering transfers to out-of-domain data and can enforce a target script through script-normalization directions.
ArXiv · AI/CL/LG's note
The paper derives those vectors from a small set of labeled utterances for tasks such as transcription.
The authors report that the approach improves task enforcement and, when paired with prompts, usually beats prompting alone on evaluated tasks including ASR and emotion recognition.
They also say the steering transfers to out-of-domain data and can enforce a target script through script-normalization directions.
ArXiv · AI/CL/LG's note
score 4