arXiv · 2609.28877
Learning New Words from Unlabeled Test Data in Automatic Speech Recognition
Abstract
New words are invented every day. A human listener can learn a new word by hearing it clearly once and inferring its usage from sentence context. This paper proposes granting ASR a similar ability to learn the contextual representations and spellings of new words from unlabeled test data at test time. A frozen CTC acoustic model provides spellings, a frozen language model provides contextual evidence for out-of-vocabulary (OOV) word detection, and an adaptation module expands the vocabulary by learning the lexical token representations with distributions over CTC-generated candidates. The spelling model of each token is optimized by minimizing a Kullback-Leibler divergence (KLD) objective. We demonstrate that the CTC-weighted language model log likelihood ratio can be interpreted as the KLD between the unknown correct ASR and the unsupervised learned ASR, and that, using a Pinsker bound, the square root of KLD can be interpreted as an upper bound on the total variation distance between the true and estimated spelling of the unknown word. Experiments show relative OOV character-error-rate reductions of up to 14.97% on LibriSpeech and 6.67% on dysarthric Speech Accessibility Project data for recurring OOV words, relative to the corresponding rescoring system.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mengqi Wang, Mark A. Hasegawa-Johnson, Haolong Zheng, Chang D. Yoo. 2026-09-24. Learning New Words from Unlabeled Test Data in Automatic Speech Recognition. https://arxiv.org/abs/2609.28877
Cite the original work for its findings. Save a collection to share your selection of sources.