arXiv · cmp-lg/9408011
Distributional Clustering of English Words
Abstract
We describe and experimentally evaluate a method for automatically clustering words according to their distribution in particular syntactic contexts. Deterministic annealing is used to find lowest distortion sets of clusters. As the annealing parameter increases, existing clusters become unstable and subdivide, yielding a hierarchical ``soft'' clustering of the data. Clusters are used as the basis for class models of word coocurrence, and the models evaluated with respect to held-out test data.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Fernando Pereira, Naftali Tishby, Lillian Lee. 1994-08-22. Distributional Clustering of English Words. https://arxiv.org/abs/cmp-lg/9408011
Cite the original work for its findings. Save a collection to share your selection of sources.