Classification models for heart disease prediction using feature selection and PCA

Anna Karen Gárate-Escamilla, Amir Hajjam El Hassani, Emmanuel Andrès

Informatics in Medicine Unlocked · 2020 · 413 citations · 49 references

DOIFull text

Open access

Concepts

TL;DR

Cardiac disease prediction improves clinical decision‑making, and machine learning offers a means to reduce and interpret heart‑disease symptoms. This study proposes a dimensionality‑reduction approach that combines feature selection with PCA to identify key predictors of heart disease. Using the 74‑feature UCI Heart Disease dataset, the authors applied chi‑square feature selection and PCA, then evaluated six classifiers, including random forests. Chi‑square plus PCA with random forests achieved 98.7–99.4% accuracy across Cleveland, Hungarian, and combined datasets, identified clinically relevant features, and outperformed other methods, while PCA alone produced lower results.

Abstract

The prediction of cardiac disease helps practitioners make more accurate decisions regarding patients' health. Therefore, the use of machine learning (ML) is a solution to reduce and understand the symptoms related to heart disease. The aim of this work is the proposal of a dimensionality reduction method and finding features of heart disease by applying a feature selection technique. The information used for this analysis was obtained from the UCI Machine Learning Repository called Heart Disease. The dataset contains 74 features and a label that we validated by six ML classifiers. Chi-square and principal component analysis (CHI-PCA) with random forests (RF) had the highest accuracy, with 98.7% for Cleveland, 99.0% for Hungarian, and 99.4% for Cleveland-Hungarian (CH) datasets. From the analysis, ChiSqSelector derived features of anatomical and physiological relevance, such as cholesterol, highest heart rate, chest pain, features related to ST depression, and heart vessels. The experimental results proved that the combination of chi-square with PCA obtains greater performance in most classifiers. The usage of PCA directly from the raw data computed lower results and would require greater dimensionality to improve the results.

References

49