European Scientific Journal ESJ · 2017 · 20 citations · 37 references
EngineeringMachine LearningRare ObservationsMining MethodsCredit ScoreClassification MethodData ScienceData MiningPattern RecognitionClass ImbalanceReal World DatasetsClass Imbalanced LearningCredit ScoringStatisticsAlternative DataCredit Scoring DatasetsPredictive AnalyticsKnowledge DiscoveryIntelligent ClassificationData ClassificationCase Study
Classification of datasets is one of the major issues encountered by the data mining community. This problem heightens when the real world datasets is also imbalanced in nature. A dataset happens to be imbalanced when the numbers of observations belonging to rare class are greatly outnumbered by the observations of another class. Class with greater number of observation is called the majority or the negative class, while the other with rare observations is referred to as the minority or the positive class. Literature represents number of resampling techniques that address the problem of class imbalance. One of the most important strategies is to resample the datasets that aim to balance the number of minority or majority observations by over-sampling or under-sampling respectively. This paper aims to investigates and analyze the performance of most widely used oversampling procedure Synthetic Minority Oversampling Technique (SMOTE) for different thresholds of oversampling using four classifiers for three credit scoring datasets.
37
Leo Breiman · Machine Learning · 2001 · 119.3K citations · Full text
SMOTE: Synthetic Minority Over-sampling Technique
Nitesh V. Chawla, Kevin W. Bowyer, Lawrence Hall et al. · Journal of Artificial Intelligence Research · 2002 · 29.6K citations · Full text
Data mining: concepts and techniques
Jiawei Han, Micheline Kamber · Choice Reviews Online · 2012 · 28.8K citations
UCI Machine Learning Repository
Arthur Asuncion · Medical Entomology and Zoology · 2007 · 24.3K citations
J. R. Quinlan · Machine Learning · 1986 · 14.5K citations · Full text