2015 · 17 citations · 22 references
Cluster ComputingEngineeringCommunity MiningCommunity DiscoveryUnsupervised Machine LearningText MiningNatural Language ProcessingComputational Social ScienceSocial MediaInformation RetrievalData ScienceData MiningText SegmentationTextual FeaturesRough ApproximationsCommunity DetectionSocial Network AnalysisSocial Medium MiningDocument ClusteringKnowledge DiscoveryComputer ScienceScalable K-nnBusiness
Clustering items using textual features is an important problem with many applications, such as root-cause analysis of spam campaigns, as well as identifying common topics in social media. Due to the sheer size of such data, algorithmic scalability becomes a major concern. In this work, we present our approach for text clustering that builds an approximate k-NN graph, which is then used to compute connected components representing clusters. Our focus is to understand the scalability / accuracy tradeoff that underlies our method: we do so through an extensive experimental campaign, where we use real-life datasets, and show that even rough approximations of k-NN graphs are sufficient to identify valid clusters. Our method is scalable and can be easily tuned to meet requirements stemming from different application domains.
22
Distributed Representations of Words and Phrases and their Compositionality
Tomáš Mikolov, Ilya Sutskever, Kai Chen et al. · arXiv (Cornell University) · 2013 · 18.1K citations · Full text
Least squares quantization in PCM
Sheelagh Lloyd · IEEE Transactions on Information Theory · 1982 · 15.1K citations · Full text
Santo Fortunato · Physics Reports · 2009 · 11.1K citations · Full text
Finding community structure in networks using the eigenvectors of matrices
M. E. J. Newman · Physical Review E · 2006 · 4.8K citations · Full text