Proceedings of the Institution of Mechanical Engineers Part C Journal of Mechanical Engineering Science · 2004 · 22 citations · 3 references
Cluster ComputingEngineeringMultiset Data AnalysisUnsupervised Machine LearningOptimization-based Data MiningData ScienceData MiningPattern RecognitionParallel ComputingStatisticsKnowledge DiscoveryComputer ScienceTwo-phase K-means AlgorithmNew AlgorithmSignal ProcessingComputational ScienceTwo-phase K-meansBig DataK-means Algorithm
One of the drawbacks of the K-means algorithm is the need for several iterations over datasets before it converges on a solution. Therefore, its application is limited to relatively small datasets. This paper presents a scalable version of the K-means algorithm that employs a buffering technique. The new algorithm, Two-Phase K-means, can robustly find a good solution in only one iteration.
3
Scaling clustering algorithms to large databases
Patricia Bradley, Usama M. Fayyad, Cory Reina · 1998 · 707 citations