arXiv · 2610.01182
Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR
Abstract
Recent ASR development has placed growing emphasis on generalization across diverse domains and acoustic conditions. Existing approaches typically adapt pretrained ASR models to front-end functions such as wake-up word (WuW) detection through additional training or task-specific modules. In this work, we explore the use of a shared pretrained ASR backbone for WuW detection without gradient-based fine-tuning and examine whether a compact encoder can be extracted using the PCA-based structured pruning approach of SliceGPT. Experiments with Parakeet-TDT-0.6B-v3 and Moonshine-base show that WuW detection performance remains relatively stable when the encoder channel dimension is reduced by 50%. These results suggest that task-relevant compact encoders can be derived from pretrained ASR models without fine-tuning.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Hwayeon Kim, Youngwon Choi, Hyeonyu Kim. 2026-10-01. Toward Elastic Speech Inference: Training-Free Wake-Word Detection from Pretrained ASR. https://arxiv.org/abs/2610.01182
Cite the original work for its findings. Save a collection to share your selection of sources.