arXiv · cmp-lg/9605012
A New Statistical Parser Based on Bigram Lexical Dependencies
Abstract
This paper describes a new statistical parser which is based on probabilities of dependencies between head-words in the parse tree. Standard bigram probability estimation techniques are extended to calculate probabilities of dependencies between pairs of words. Tests using Wall Street Journal data show that the method performs at least as well as SPATTER (Magerman 95, Jelinek et al 94), which has the best published results for a statistical parser on this task. The simplicity of the approach means the model trains on 40,000 sentences in under 15 minutes. With a beam search strategy parsing speed can be improved to over 200 sentences a minute with negligible loss in accuracy.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Michael Collins. 1996-05-06. A New Statistical Parser Based on Bigram Lexical Dependencies. https://arxiv.org/abs/cmp-lg/9605012
Cite the original work for its findings. Save a collection to share your selection of sources.