2020 · 37 citations · 20 references
EngineeringVector SpaceCorpus LinguisticsText MiningWord EmbeddingsNatural Language ProcessingInformation RetrievalData ScienceData MiningPattern RecognitionDocument ClassificationSpherical K-meansVector ModelDocument ClusteringKnowledge DiscoveryComputer ScienceDimensionality ReductionVector Space ModelTf-idf Model
The TF-IDF model is the most common way of representing documents in the vector space. However, its results are highly dimensional, posing problems to the classic clustering algorithms due to the curse of dimensionality. Recent word embeddings based techniques can reduce the documents representations dimensionality while also preserving the semantic relationships between words. In this paper, we analyze the accuracy of four different classical clustering algorithms (K-Means, Spherical K-Means, LDA, and DBSCAN) in combination with the Document to Vector model.
20
A density-based algorithm for discovering clusters in large spatial Databases with Noise
Martin Ester, Hans‐Peter Kriegel, Jörg Sander et al. · 1996 · 19.1K citations
Efficient Estimation of Word Representations in Vector Space
Tomáš Mikolov, Kai Chen, Greg S. Corrado · arXiv (Cornell University) · 2013 · 18.1K citations · Full text
George A. Miller · Communications of the ACM · 1995 · 14K citations · Full text
Natural Language Processing, Meaningful Words, Engineering +14