Concepedia

Publication | Open Access

A Survey on Methods for Solving Data Imbalance Problem for Classification

45

Citations

7

References

2015

Year

Abstract

The term "data imbalance" in classification is a well established phenomenon in which data set contains unbalanced class distributions. Dataset is called unbalanced if it contains at least one class which is presented by very few examples. A range of solutions have been proposed for the problem of data imbalance including data sampling, cost evaluation of model, bagging, boosting, Genetic Programming (GP) based methods etc. This paper presents a survey of various methods introduced by researchers to handle data imbalance problem in order to improve classification performance and further the comparison between the methods on the basis of their advantages and disadvantages is done.

References

YearCitations

Page 1