arXiv · 2610.03017
Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations
Abstract
This work presents an automatic speech recognition (ASR) system personalized for a Czech speaker with a permanent tracheal stoma and severe dysarthria rendering their speech unintelligible to untrained listeners. We release a public dataset containing 33 annotated hours of the speaker's speech, collected using a novel "artificial conversation" protocol designed for high engagement and dialogue realism. We propose a multi-stage training pipeline based on Whisper Base: fine-tuning on standard Czech speech, acoustically simulated tracheostomic speech, and the speaker's data. We evaluate the system across three near real-time scenarios: scripted conversations, question answering, and spontaneous dialogue, achieving a 50\% relative reduction in Character Error Rate compared to Whisper Base baseline and surpassing the average recognition accuracy of their assistants in acoustic recognition of isolated utterances. We demonstrate that even for severely impeded speech, a helpful ASR is achievable, as evidenced by the quantitative results and the feedback from the speaker.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
David Nadrchal, Monorama Swain, Florian Schmid, Gerhard Widmer, Paul Primus. 2026-10-02. Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations. https://arxiv.org/abs/2610.03017
Cite the original work for its findings. Save a collection to share your selection of sources.