Megadose Built for builders and researchers.

DirectSpeech2LLM: A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLMs

· ArXiv · AI/CL/LG ·
A speech-LLM trained only on ASR data still handled unseen translation and emotion instructions zero-shot.

DirectSpeech2LLM is meant to reduce prompt overfitting, where a speech-conditioned model keeps acting like an ASR system even when given a different instruction. The framework aligns speech embeddings to a frozen LLM embedding space using a distance-based CTC loss and greedy CTC labels. On 960 hours of LibriSpeech ASR training, it beat a cascaded system on ASR and closely matched that system’s upper bound on the two unseen tasks. The authors also report that explicit regression for geometric alignment was less important than expected. ArXiv · AI/CL/LG's note

score 5

Categories: Research