Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Abstract
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1times with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10times faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.
Community
Memory-efficient CTC loss with exact loss and first-order gradient equivalence under vocabulary reduction, and activation memory that no longer scales linearly with vocabulary size.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Alignment-Path Distillation from Non-streaming ASR-LLMs for Streaming Speech Recognition (2026)
- X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation (2026)
- TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context (2026)
- A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition (2026)
- FastE: Readout-Triggered Token Compression for LLM Embedding Inference (2026)
- StreamAlign: Streaming Text-Aligned Speech Tokenization (2026)
- StreamHear: Domain-Adapted Pseudo-Labeling for Semi-Supervised Streaming Speech Recognition (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper