2003 · 38 citations · 6 references
Cluster ComputingDocument ClusteringBootstrap ResamplingEngineeringMachine LearningData ScienceData MiningPattern RecognitionBootstrap SampleData Stream MiningKnowledge DiscoveryComputer ScienceK-means ClusteringStatisticsBig DataOptimization-based Data Mining
K-means clustering is one of the most popular clustering algorithms used in data mining. However, clustering is a time consuming task, particularly with the large data sets found in data mining. In this paper we show how bootstrap averaging with k-means can produce results comparable to clustering all of the data but in much less time. The approach of bootstrap (sampling with replacement) averaging consists of running k-means clustering to convergence on small bootstrap samples of the training data and averaging similar cluster centroids to obtain a single model. We show why our approach should take less computation time and empirically illustrate its benefits. We show that the performance of our approach is a monotonic function of the size of the bootstrap sample. However, knowing the size of the bootstrap sample that yields as good results as clustering the entire data set remains an open and important question.
6
Data mining: concepts and techniques
Jiawei Han, Micheline Kamber · Choice Reviews Online · 2012 · 28.8K citations
Leo Breiman · Machine Learning · 1996 · 16.6K citations · Full text
Quantizing for minimum distortion
J. Max · IEEE Transactions on Information Theory · 1960 · 2.1K citations
Input Signal, Output Levels, Statistical Signal Processing +11
Applied Physics Letters · 2000 · 1.1K citations · Full text
Refining Initial Points for K-Means Clustering
Paul S. Bradley, Usama M. Fayyad · 1998 · 999 citations