Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs
Pruning speech-LLMs can make error rates less equal across demographic groups, even when aggregate WER looks acceptable.
The paper tests audio-encoder pruning on SLAM-ASR using Fair-Speech and Common Voice. It finds the spread between best- and worst-performing demographic groups grows after pruning, across three encoder scales. LoRA adaptation lowers WER for every group, but can help already stronger groups more and widen some gaps. The authors argue deployment checks should include per-group WER and treat the worst-performing group’s error rate as an explicit criterion. ArXiv · AI/CL/LG's note
The paper tests audio-encoder pruning on SLAM-ASR using Fair-Speech and Common Voice. It finds the spread between best- and worst-performing demographic groups grows after pruning, across three encoder scales. LoRA adaptation lowers WER for every group, but can help already stronger groups more and widen some gaps. The authors argue deployment checks should include per-group WER and treat the worst-performing group’s error rate as an explicit criterion. ArXiv · AI/CL/LG's note
score 4