The Data Preprocessing Technique in Machine Learning

Authors

  • Jiale Deng

DOI:

https://doi.org/10.61173/0hfqqz07

Keywords:

Data Preprocessing, Feature engineering, Machine learning, Deep learning

Abstract

Data preprocessing has a vital impact on the performance of traditional machine learning. However, with the continuous development of deep learning technology and the powerful representation learning ability of neural networks, deep learning models can easily convert raw data into continuous feature representation and have made remarkable achievements in many downstream tasks. Recently, with the application of deep learning technology in more fields that require robustness and stability, its model bias caused by data quality problems are gradually exposed, which makes the data preprocessing technology regain the attention of researchers. This paper systematically expounds on the primary data preprocessing technology in machine learning and discusses the model bias caused by data deviation and the robustness in the face of attacks. In addition, this paper also illustrates the possibility of a data preprocessing foundation in solving the defects of deep learning foundation based on glove word vector and image classification task based on convolutional neural network. The research in this paper can provide a valuable reference for researchers in related fields.

References

[1] Chunlong Xia, Xinliang Wang, Feng Lv, et al. Vit-comer: Vision transformer with convolutional multi-scale feature interaction for dense predictions//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 5493-5502.

[2] S. Joshua Johnson, M. Ramakrishna Murty, I. Navakanth. A Dean&Francis ISSN 2959-6157 detailed review on word embedding techniques with emphasis on word2vec. Multimedia Tools and Applications, 2024, 83(13): 37979-38007.

[3] Lijie Fan, Dilip Krishnan, Phillip Isola, et al. Improving clip training with language rewrites. Advances in Neural Information Processing Systems, 2024, 36.

[4] Andrea Vallebueno, Cassandra Handan-Nader, Christopher D. Manning, et al. Statistical Uncertainty in Word Embeddings: GloVe-V. arXiv preprint arXiv:2406.12165, 2024.

[5] Linyu Tang, Lei Zhang. Robust Overfitting Does Matter: Test-Time Adversarial Purification With FGSM[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024: 24347-24356.

[6] Farhan Ullah, Irfan Ullah, Rehan Ullah Khan, et al. Conventional to deep ensemble methods for hyperspectral image classification: A comprehensive survey. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 2024.

[7] Ashish Bajaj, Dinesh Kumar Vishwakarma. A state-of-the-art review on adversarial machine learning in image classification. Multimedia Tools and Applications, 2024, 83(3): 9351-9416.

Downloads

Published

2024-12-31