Computer Science and Information Systems · 2017 · 13 citations · 33 references
EngineeringMachine LearningImbalanced Data ClassificationSupport Vector MachineClassification MethodData ScienceData MiningPattern RecognitionHybrid ResamplingClass ImbalanceManagementStatisticsRelative BalanceImbalanced DatasetsPredictive AnalyticsComputer ScienceData ClassificationClassifier SystemCost-sensitive Learning
Imbalanced datasets exist widely in real life. The identification of the minority class in imbalanced datasets tends to be the focus of classification. As a variant of enhanced support vector machine (SVM), the twin support vector machine (TWSVM) provides an effective technique for data classification. TWSVM is based on a relative balance in the training sample dataset and distribution to improve the classification accuracy of the whole dataset, however, it is not effective in dealing with imbalanced data classification problems. In this paper, we propose to combine a re-sampling technique, which utilizes oversampling and under-sampling to balance the training data, with TWSVM to deal with imbalanced data classification. Experimental results show that our proposed approach outperforms other state-of-art methods.
33
SMOTE: Synthetic Minority Over-sampling Technique
Nitesh V. Chawla, Kevin W. Bowyer, Lawrence Hall et al. · Journal of Artificial Intelligence Research · 2002 · 29.6K citations · Full text
Proceedings of the 24th international conference on Machine learning
2007 · 11.7K citations
Haibo He, Edwardo A. Garcia · IEEE Transactions on Knowledge and Data Engineering · 2009 · 9.2K citations
ADASYN: Adaptive synthetic sampling approach for imbalanced learning
Haibo He, Yang Bai, Edwardo A. Garcia et al. · 2008 · 4.3K citations
Artificial Intelligence, Data Classification, Classification Method +15