Research on the Influencing Factors of Heart Disease based on Logistic Regression and XGBoost

Authors

  • Zitong Zhou

DOI:

https://doi.org/10.61173/qeache10

Keywords:

Heart disease, influencing factors, logistic regression, XGBoost

Abstract

This study leverages the UCI Cleveland Heart Disease dataset (297 complete records, 14 features) to present a two-stage pipeline that couples rigorous feature engineering with hybrid modeling and compares Logistic Regression (LR) against Gradient Boosting Decision Trees (GBDT/XGBoost) for cardiac risk prediction. Mutual information and recursive feature elimination first isolate the most informative variables-chest-pain type, maximum heart rate, exercise-induced ST depression, ST slope, and number of major vessels-whose clinical relevance is well established. GBDT then markedly outperforms LR across all evaluated metrics: accuracy 89.2% vs 83.7%, recall 86.7% vs 80.0%, and AUC 0.925 vs 0.874, demonstrating the value of capturing non-linear interactions among risk factors. Feature-importance analysis corroborates these predictors’ medical interpretability. The authors acknowledge limitations arising from the modest sample size and the ongoing need for transparent models; they recommend expanding the data, integrating SHAP explanations, and incorporating real-time monitoring. Overall, explainable GBDT-based tools can support clinicians in early identification of high-risk individuals, enabling personalized interventions and improved patient outcomes.

References

[1] Detrano R, Janosi A, Steinbrunn W, et al. International application of a new probability algorithm for the diagnosis of coronary artery disease. The American Journal of Cardiology, 1989, 64(5): 304-310.

[2] Liaw A, Wiener M. Classification and Regression by randomForest. R News, 2002, 2(3): 18-22.

[3] Vapnik V N. Statistical Learning Theory. Wiley, 1998.

[4] Khosla A, Cao Y, Lin C C Y, et al. An integrated machine learning approach to stroke prediction. Proceedings of the ACM SIGKDD, 2010, 183-192.

[5] Rajpurkar P, et al. CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning. Working paper, 2017.

[6] Dietterich T G. Ensemble methods in machine learning. International Workshop on Multiple Classifier Systems, 2000, 1-15.

[7] Ribeiro M T, Singh S, Guestrin C. “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. ACM SIGKDD, 2016, 1135-1144.

[8] Babyak M A. What you see may not be what you get: a brief, nontechnical introduction to overfitting in regression-type models. Psychosomatic Medicine, 2004, 66(3): 411-421.

[9] Potdar K, Pardawala T S, Pai C D. A comparative study of categorical variable encoding techniques for neural network classifiers. International Journal of Computer Applications, 2017, 175(4): 7-9.

[10] Han J, Kamber M, Pei J. Data Mining: Concepts and Techniques. Elsevier, 2011.

[11] Vergara J R, Estévez P A. A review of feature selection methods based on mutual information. Neural Computing and Applications, 2014, 24(1): 175-186.

[12] Guyon I, Weston J, Barnhill S, Vapnik V. Gene selection for cancer classification using support vector machines. Machine Learning, 2002, 46(1-3): 389-422.

[13] Hosmer D W, Lemeshow S, Sturdivant R X. Applied Logistic Regression. John Wiley & Sons, 2013.

[14] Friedman J H. Greedy function approximation: a gradient boosting machine. Annals of Statistics, 2011, 29(5): 1189-1232.

[15] Chen T, Guestrin C. XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD, 2019, 785-794.

[16] Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 2015, 10(3).

Downloads

Published

2025-10-23