arXiv · 2307.05796
Improved POS tagging for spontaneous, clinical speech using data augmentation
Abstract
This paper addresses the problem of improving POS tagging of transcripts of speech from clinical populations. In contrast to prior work on parsing and POS tagging of transcribed speech, we do not make use of an in domain treebank for training. Instead, we train on an out of domain treebank of newswire using data augmentation techniques to make these structures resemble natural, spontaneous speech. We trained a parser with and without the augmented data and tested its performance using manually validated POS tags in clinical speech produced by patients with various types of neurodegenerative conditions.
Explore related subjects
Keep this discovery
Seth Kulick, Neville Ryant, David J. Irwin, Naomi Nevler, Sunghye Cho. 2023-07-11. Improved POS tagging for spontaneous, clinical speech using data augmentation. https://arxiv.org/abs/2307.05796
Cite the original work for its findings. Save a collection to share your selection of sources.