Performance Investigation of Feature Selection Based on Random Forest in Heart Disease Prediction Using KNN Model

Authors

  • Shun Kong

DOI:

https://doi.org/10.61173/97ar1w87

Keywords:

Heart disease, K-Nearest Neighbors (KNN), random forest, feature selection

Abstract

Heart disease is one of the leading causes of death worldwide, claiming millions of lives each year. To address this serious public health challenge, early prediction of heart disease using machine learning techniques has become a hot topic of research. This study explores the impact of different numbers of features on the performance of the K-Nearest Neighbors (KNN) model in predicting heart disease. Initially, a random forest algorithm was employed to rank the importance of a large set of features and identify the key factors most influential in predicting heart disease. Subsequently, starting with the most important features, the study incrementally increased the number of features applied to the KNN model, comparing the model’s accuracy and recall across different feature combinations. The results show that as the number of features increases, the model’s predictive performance does not consistently improve. When the number of features is initially increased, accuracy experiences a sharp decline; although it slightly recovers later, the overall performance does not return to the high level observed with fewer features. Meanwhile, recall significantly improves when the number of features first increases but then starts to fluctuate and noticeably decreases when a certain number of features is reached. This study demonstrates that simply increasing the number of features does not guarantee improved model performance; instead, it may introduce redundant information or noise, weakening the model’s effectiveness.

References

[1] WHO, Cardiovascular diseases (CVDs), https://www.who. int/zh/news-room/fact-sheets/detail/cardiovascular-diseases- (cvds), 2021.

[2] Jindal H, Agrawal S, Khera R, et al. Heart disease prediction using machine learning algorithms. IOP conference series: materials science and engineering. IOP Publishing, 2021, 1022(1): 012072.

[3] Srinivas K, Rani B K, Govrdhan A. Applications of data mining techniques in healthcare and prediction of heart attacks. International Journal on Computer Science and Engineering (IJCSE), 2010, 2(02): 250-255.

[4] Masethe H D, Masethe M A. Prediction of heart disease using classification algorithms. Proceedings of the world Congress on Engineering and computer Science. 2014, 2(1): 25- 29.

[5] Du YY. Analysis of heart disease prediction performance using different classification models (in Chinese). Modeling and Simulation, 2023, 12: 5600.

[6] Wang LL, Fu ZL, Tao P, et al. Heart disease classification based on an imbalanced multi-class AdaBoost algorithm using active learning (in Chinese). Computer Applications, 2017, 37(7): 1994-1998.

[7] CDC, Behavioral Risk Factor Surveillance System. https:// www.cdc.gov/brfss/annual_data/annual_data.htm, 2024

[8] Rigatti SJ. Random forest. Journal of Insurance Medicine. 2017 Jan 1;47(1):31-9.

[9] Biau G, Scornet E. A random forest guided tour. Test. 2016 Jun;25:197-227.

[10] Guo G, Wang H, Bell D, Bi Y, Greer K. KNN modelbased approach in classification. InOn The Move to Meaningful Internet Systems 2003: CoopIS, DOA, and ODBASE: OTM Confederated International Conferences, CoopIS, DOA, and ODBASE 2003, Catania, Sicily, Italy, November 3-7, 2003. Proceedings 2003 (pp. 986-996). Springer Berlin Heidelberg.

[11] Zhang S, Li X, Zong M, Zhu X, Cheng D. Learning k for knn classification. ACM Transactions on Intelligent Systems and Technology (TIST). 2017 Jan 12;8(3):1-9.

Downloads

Published

2024-12-31