Publication | Closed Access
A Comparison of External Clustering Evaluation Indices in the Context of Imbalanced Data Sets
46
Citations
11
References
2012
Year
Unknown Venue
EngineeringMachine LearningUnsupervised Machine LearningText MiningImbalanced Data SetsClassification MethodInformation RetrievalData ScienceData MiningPattern RecognitionClass ImbalanceClass Size ImbalanceRand IndexStatisticsClustering (Nuclear Physics)Knowledge DiscoveryData ClassificationCluster DevelopmentEvaluation MeasureMutual InformationClustering (Data Mining)
For highly imbalanced data sets, almost all the instances are labeled as one class, whereas far fewer examples are labeled as the other classes. In this paper, we present an empirical comparison of seven different clustering evaluation indices when used to assess partitions generated from highly imbalanced data sets. Some of the metrics are based on matching of sets (F-measure), information theory (normalized mutual information and adjusted mutual information), and pair of objects counting (Rand and adjusted Rand indices). We also investigate the BCubed metric, which takes into account the concepts of recall, precision, as well as counting pairs. Furthermore, in order to avoid the class size imbalance effect, we propose a modification to the Rand index, referred to as the normalized class size Rand (NCR) index. In terms of results, apart from NCR, our experiments indicate that all the other analyzed indices are not able to deal properly with the problem of class size imbalance.
| Year | Citations | |
|---|---|---|
Page 1
Page 1