Distributional clustering of English words
1993 · 994 citations · 10 references
EngineeringNeurolinguisticsSemanticsCorpus LinguisticsText MiningParticular Syntactic ContextsApplied LinguisticsNatural Language ProcessingData ScienceComputational LinguisticsLanguage StudiesDocument ClusteringComputational LexicologyKnowledge DiscoveryDistributional SemanticsDeterministic AnnealingDistributional ClusteringCluster MembershipLexical Complexity PredictionLinguisticsSemantic Similarity
We describe and evaluate experimentally a method for clustering words according to their distribution in particular syntactic contexts. Words are represented by the relative frequency distributions of contexts in which they appear, and relative entropy between those distributions is used as the similarity measure for clustering. Clusters are represented by average context distributions derived from the given words according to their probabilities of cluster membership. In many cases, the clusters can be thought of as encoding coarse sense distinctions. Deterministic annealing is used to find lowest distortion sets of clusters: as the annealing parameter increases, existing clusters become unstable and subdivide, yielding a hierarchical "soft" clustering of the data. Clusters are used as the basis for class models of word coocurrence, and the models evaluated with respect to held-out test data.
10
Maximum Likelihood from Incomplete Data Via the <i>EM</i> Algorithm
A. P. Dempster, N. M. Laird, Donald B. Rubin · Journal of the Royal Statistical Society Series B (Statistical Methodology) · 1977
Statistical Signal ProcessingMixture DistributionEngineering+13
49.2K citations
Pattern Classification and Scene Analysis
Michael Thompson, Richard O. Duda, Peter E. Hart · Leonardo · 1974
4.5K citations
Class-based n -gram models of natural language
Peter F. Brown, P.V. deSouza, Robert L. Mercer et al. · Computational Linguistics · 1992
2.9K citations
A stochastic parts program and noun phrase parser for unrestricted text
Kenneth Church · 1988
974 citations
Noun classification from predicate-argument structures
Donald Hindle · 1990
543 citations