arXiv · cmp-lg/9409003
A Probabilistic Model of Compound Nouns
Abstract
Compound nouns such as example noun compound are becoming more common in natural language and pose a number of difficult problems for NLP systems, notably increasing the complexity of parsing. In this paper we develop a probabilistic model for syntactically analysing such compounds. The model predicts compound noun structures based on knowledge of affinities between nouns, which can be acquired from a corpus. Problems inherent in this corpus-based approach are addressed: data sparseness is overcome by the use of semantically motivated word classes and sense ambiguity is explicitly handled in the model. An implementation based on this model is described in Lauer (1994) and correctly parses 77% of the test set.
Explore related subjects
Keep this discovery
Mark Lauer, Mark Dras. 1994-09-06. A Probabilistic Model of Compound Nouns. https://arxiv.org/abs/cmp-lg/9409003
Cite the original work for its findings. Save a collection to share your selection of sources.