arXiv · cmp-lg/9405008
A Stochastic Finite-State Word-Segmentation Algorithm for Chinese
Abstract
We present a stochastic finite-state model for segmenting Chinese text into dictionary entries and productively derived words, and providing pronunciations for these words; the method incorporates a class-based model in its treatment of personal names. We also evaluate the system's performance, taking into account the fact that people often do not agree on a single segmentation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Richard Sproat, Chilin Shih, William Gale, Nancy Chang. 1994-05-05. A Stochastic Finite-State Word-Segmentation Algorithm for Chinese. https://arxiv.org/abs/cmp-lg/9405008
Cite the original work for its findings. Save a collection to share your selection of sources.