Prediction of Diabetes Based on Machine Learning Algorithm

Authors

  • Zijian Zhou

DOI:

https://doi.org/10.61173/6793rf84

Keywords:

Diabetes, Missing value, Feature visualization, Machine learning

Abstract

Diabetes is a well-known chronic disease that includes a range of metabolic disorders characterized by persistently elevated blood sugar levels over an extended period of time. Early and precise prediction of diabetes is essential to reduce risk factors and minimize potential complications associated with the disease. However, there are significant challenges in creating reliable predictive models due to factors such as limited labeling data, the presence of outliers, and the absence of information in diabetes-related datasets. To address these barriers, this paper proposes a comprehensive framework aimed at improving diabetes prediction through data preprocessing and machine learning techniques. The framework combines methods for dealing with missing values, data standardization, and feature visualization to extract meaningful insights. In addition, various machine learning classifiers - including support vector machine (SVM), decision tree, logistic regression, and naive Baye - are implemented to improve prediction accuracy and support early diagnosis of diabetes. Among these models, SVM shows better comprehensive performance.

References

[1] Maniruzzaman M, Rahman M J, Al-MehediHasan M, et al. Accurate diabetes risk stratification using machine learning: role of missing value and outliers. Journal of medical systems, 2018, 42: 1-17.

[2] McLachlan G J. Discriminant analysis and statistical pattern recognition. John Wiley & Sons, 2005.

[3] Cover T M. Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition. IEEE transactions on electronic computers, 1965 (3): 326-334.

[4] Webb G I, Boughton J R, Wang Z. Not so naive Bayes: aggregating one-dependence estimators. Machine learning, 2005, 58: 5-24.

[5] Brahim-Belhouari S, Bermak A. Gaussian process for nonstationary time series prediction. Computational Statistics & Data Analysis, 2004, 47(4): 705-712.

[6] Cortes C. Support-Vector Networks. Machine Learning, 1995.

[7] Reinhardt A, Hubbard T. Using neural networks for prediction of the subcellular location of proteins. Nucleic acids research, 1998, 26(9): 2230-2236.

[8] Kégl B. The return of AdaBoost. MH: multi-class Hamming trees. arxiv preprint arxiv:1312.6086, 2013.

[9] Jenhani I, Amor N B, Elouedi Z. Decision trees as possibilistic classifiers. International journal of approximate reasoning, 2008, 48(3): 784-807.

[10] Breiman L. Random forests. Machine learning, 2001, 45: 5-32

Downloads

Published

2024-12-31