Data preprocessing techniques for classification without discrimination
Knowledge and Information Systems · 2011 · 1.2K citations · 22 references
Artificial IntelligenceEngineeringMachine LearningBiometricsDiscriminationClassification MethodData ScienceData MiningPattern RecognitionManagementSensitive AttributesBiased Decision ProcessAlgorithmic BiasKnowledge DiscoveryDisparate ImpactComputer ScienceData ClassificationData Preprocessing TechniquesAlgorithmic FairnessData TreatmentClassificationBinary Sensitive AttributeData Modeling
The Discrimination‑Aware Classification Problem was introduced to address training data that exhibit unlawful discrimination toward sensitive attributes such as gender or ethnicity, and it is relevant when data are generated by a biased process or when the sensitive attribute proxies unobserved features. The study aims to learn a classifier that maximizes accuracy while eliminating discrimination in test‑data predictions, focusing on a single binary sensitive attribute and a two‑class classification setting. We analyze the optimal trade‑off between accuracy and non‑discrimination for pure classifiers and propose algorithmic solutions that preprocess data—by suppressing the sensitive attribute, massaging class labels, and reweighing or resampling—to remove discrimination before learning a classifier. The preprocessing techniques were implemented in a modified version of Weka, and experiments on real‑life data demonstrate their effectiveness.
Recently, the following Discrimination-Aware Classification Problem was introduced: Suppose we are given training data that exhibit unlawful discrimination; e.g., toward sensitive attributes such as gender or ethnicity. The task is to learn a classifier that optimizes accuracy, but does not have this discrimination in its predictions on test data. This problem is relevant in many settings, such as when the data are generated by a biased decision process or when the sensitive attribute serves as a proxy for unobserved features. In this paper, we concentrate on the case with only one binary sensitive attribute and a two-class classification problem. We first study the theoretically optimal trade-off between accuracy and non-discrimination for pure classifiers. Then, we look at algorithmic solutions that preprocess the data to remove discrimination before a classifier is learned. We survey and extend our existing data preprocessing techniques, being suppression of the sensitive attribute, massaging the dataset by changing class labels, and reweighing or resampling the data to remove discrimination without relabeling instances. These preprocessing techniques have been implemented in a modified version of Weka and we present the results of experiments on real-life data.
22
SMOTE: Synthetic Minority Over-sampling Technique
Nitesh V. Chawla, Kevin W. Bowyer, Lawrence Hall et al. · Journal of Artificial Intelligence Research · 2002
29.6K citations
UCI Machine Learning Repository
Arthur Asuncion · Medical Entomology and Zoology · 2007
24.3K citations
Wrappers for feature subset selection
Ron Kohavi, George H. John · Artificial Intelligence · 1997
8.8K citations
The foundations of cost-sensitive learning
Charles Elkan · 2001
1.8K citations
1.3K citations