Research on the Influencing Factors of Heart Disease based on Logistic Regression and XGBoost
DOI:
https://doi.org/10.61173/qeache10Keywords:
Heart disease, influencing factors, logistic regression, XGBoostAbstract
This study leverages the UCI Cleveland Heart Disease dataset (297 complete records, 14 features) to present a two-stage pipeline that couples rigorous feature engineering with hybrid modeling and compares Logistic Regression (LR) against Gradient Boosting Decision Trees (GBDT/XGBoost) for cardiac risk prediction. Mutual information and recursive feature elimination first isolate the most informative variables-chest-pain type, maximum heart rate, exercise-induced ST depression, ST slope, and number of major vessels-whose clinical relevance is well established. GBDT then markedly outperforms LR across all evaluated metrics: accuracy 89.2% vs 83.7%, recall 86.7% vs 80.0%, and AUC 0.925 vs 0.874, demonstrating the value of capturing non-linear interactions among risk factors. Feature-importance analysis corroborates these predictors’ medical interpretability. The authors acknowledge limitations arising from the modest sample size and the ongoing need for transparent models; they recommend expanding the data, integrating SHAP explanations, and incorporating real-time monitoring. Overall, explainable GBDT-based tools can support clinicians in early identification of high-risk individuals, enabling personalized interventions and improved patient outcomes.
References
[1] Detrano R, Janosi A, Steinbrunn W, et al. International application of a new probability algorithm for the diagnosis of coronary artery disease. The American Journal of Cardiology, 1989, 64(5): 304-310.
[2] Liaw A, Wiener M. Classification and Regression by randomForest. R News, 2002, 2(3): 18-22.
[3] Vapnik V N. Statistical Learning Theory. Wiley, 1998.
[4] Khosla A, Cao Y, Lin C C Y, et al. An integrated machine learning approach to stroke prediction. Proceedings of the ACM SIGKDD, 2010, 183-192.
[5] Rajpurkar P, et al. CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning. Working paper, 2017.
[6] Dietterich T G. Ensemble methods in machine learning. International Workshop on Multiple Classifier Systems, 2000, 1-15.
[7] Ribeiro M T, Singh S, Guestrin C. “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. ACM SIGKDD, 2016, 1135-1144.
[8] Babyak M A. What you see may not be what you get: a brief, nontechnical introduction to overfitting in regression-type models. Psychosomatic Medicine, 2004, 66(3): 411-421.
[9] Potdar K, Pardawala T S, Pai C D. A comparative study of categorical variable encoding techniques for neural network classifiers. International Journal of Computer Applications, 2017, 175(4): 7-9.
[10] Han J, Kamber M, Pei J. Data Mining: Concepts and Techniques. Elsevier, 2011.
[11] Vergara J R, Estévez P A. A review of feature selection methods based on mutual information. Neural Computing and Applications, 2014, 24(1): 175-186.
[12] Guyon I, Weston J, Barnhill S, Vapnik V. Gene selection for cancer classification using support vector machines. Machine Learning, 2002, 46(1-3): 389-422.
[13] Hosmer D W, Lemeshow S, Sturdivant R X. Applied Logistic Regression. John Wiley & Sons, 2013.
[14] Friedman J H. Greedy function approximation: a gradient boosting machine. Annals of Statistics, 2011, 29(5): 1189-1232.
[15] Chen T, Guestrin C. XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD, 2019, 785-794.
[16] Saito T, Rehmsmeier M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 2015, 10(3).
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
