arXiv · 2603.18024
ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody
Abstract
Current keyword spotting systems primarily use phoneme-level matching to distinguish confusable words but ignore user-specific pronunciation traits like prosody (intonation, stress, rhythm). This paper presents ProKWS, a novel framework integrating fine-grained phoneme learning with personalized prosody modeling. We design a dual-stream encoder where one stream derives robust phonemic representations through contrastive learning, while the other extracts speaker-specific prosodic patterns. A collaborative fusion module dynamically combines phonemic and prosodic information, enhancing adaptability across acoustic environments. Experiments show ProKWS delivers highly competitive performance, comparable to state-of-the-art models on standard benchmarks and demonstrates strong robustness for personalized keywords with tone and intent variations.
Explore related subjects
Keep this discovery
Jianan Pan, Yuanming Zhang, Kejie Huang. 2026-03-05. ProKWS: Personalized Keyword Spotting via Collaborative Learning of Phonemes and Prosody. https://arxiv.org/abs/2603.18024
Cite the original work for its findings. Save a collection to share your selection of sources.