arXiv · 2308.12490
MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
Abstract
Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open-response pronunciation assessment models have predominantly focused on a single pronunciation task, such as sentence-level accuracy, rather than offering a comprehensive assessment in various aspects. We propose MultiPA, a Multitask Pronunciation Assessment model that provides sentence-level accuracy, fluency, prosody, and word-level accuracy assessment for open responses. We examined the correlation between different pronunciation tasks and showed the benefits of multi-task learning. Our model reached the state-of-the-art performance on existing in-domain data sets and effectively generalized to an out-of-domain dataset that we newly collected. The experimental results demonstrate the practical utility of our model in real-world applications.
Explore related subjects
Keep this discovery
Yu-Wen Chen, Zhou Yu, Julia Hirschberg. 2023-08-24. MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios. https://arxiv.org/abs/2308.12490
Cite the original work for its findings. Save a collection to share your selection of sources.