Journal of the Royal Statistical Society Series A (Statistics in Society) · 2019 · 11 citations · 41 references
Summary Rejection inference aims to reduce sample bias and to improve model performance in credit scoring. We propose a semisupervised clustering approach as a new rejection inference technique. K-prototype clustering can deal with mixed types of numeric and categorical characteristics, which are common in consumer credit data. We identify homogeneous acceptances and rejections and assign labels to part of the rejections according to the label of acceptances. We test the performance of various rejection inference methods in logit, support vector machine and random-forests models based on data sets of real consumer loans. The predictions of clustering rejection inference show advantages over other traditional rejection inference methods. Inferring the label of the rejection from semisupervised clustering is found to help to mitigate the sample bias problem and to improve the predictive accuracy.
41
Leo Breiman · Machine Learning · 2001 · 119.3K citations · Full text
Corinna Cortes, Vladimir Vapnik · Machine Learning · 1995 · 39.8K citations · Full text
Corinna Cortes, Vladimir Vapnik · Machine Learning · 1995 · 31.8K citations · Full text