Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Pruned CTC cuts ASR training memory by limiting alignment work to the batch’s target tokens plus blank, while preserving full-vocabulary normalization.
The paper says this reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. With a Zipformer-M encoder and 180K vocabulary, it reports a 5.1x full-step memory reduction at 17% step-time overhead. Across three corpora, accuracy matches standard CTC. The authors also use it for LLM-CTC, reporting much faster recognition than LLM-CE with modest WER gaps. HF Daily Papers' note
The paper says this reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. With a Zipformer-M encoder and 180K vocabulary, it reports a 5.1x full-step memory reduction at 17% step-time overhead. Across three corpora, accuracy matches standard CTC. The authors also use it for LLM-CTC, reporting much faster recognition than LLM-CE with modest WER gaps. HF Daily Papers' note
score 4