Application of Random Forest Technique in Data Cleaning

Authors

  • Wenhao Bai

DOI:

https://doi.org/10.61173/0az4my98

Keywords:

Random forest, Data cleaning, Missing values

Abstract

Data cleaning is an important part of data preprocessing. Its role is to ensure the accuracy of data analysis and improve model performance. Its main tasks are to identify and resolve various problems in raw data, such as missing values, outliers, and duplicate records, thereby providing strong support for subsequent data mining and modeling. In the era of big data, traditional data cleaning methods suffer from poor adaptability and low accuracy when dealing with high-dimensional, nonlinear, and massive datasets. By contrast, random forest has been widely applied in data cleaning because of its resistance to overfitting, strong ability to handle high-dimensional data, low requirement for complex preprocessing, and relatively strong interpretability. This paper reviews recent domestic and international studies on the principles, methods, and research status of random forest in the three core scenarios of data cleaning. It mainly discusses its advantages over traditional methods, its existing limitations, and future research directions. At the same time, visualized content is incorporated to improve the readability of this review and to provide references for future research and applications.

References

[1] Breiman L. Random forests. Machine Learning, 2001, 45(1): 5-32.

[2] Xiao Yang, Wang Xinzhang, Peng Cheng, Chen Junfeng, Jiang Tao. Data cleaning method for operational data of offshore oil and gas production equipment based on improved random forest. Modern Chemical Research, 2023, (12): 155-157.

[3] Han Honggui, Zhao Zifan, Wu Xiaolong, Yang Shiheng, He Zheng, Zhao Nan. Data cleaning method for urban wastewater treatment process operation data based on improved random forest. Journal of Beijing University of Technology, 2021, 47(05): 421-430.

[4] Wei Tai, He Shaoxiong, Hu Ziwu, Cao Lixin. Cleaning of abnormal data from wind turbines based on an improved isolation forest algorithm. Science Technology and Engineering, 2024, 24(09): 3691-3699.

[5] Cao Ying. Research on industrial sensing data cleaning and compression method based on elastic edge computing. Wuhan University of Technology, 2022.

[6] Liu Yue, Hao Shuxin, Liu Jie, Xu Dongqun. Research Dean&Francis Wenhao Bai on a data cleaning framework for environmental health risk assessment data. Journal of Environmental Hygiene, 2025, 15(10): 878-884+907.

[7] Hou Dengyun, Nan Xinyuan, Li Hailong. Data cleaning method for wastewater treatment process based on adaptive DBSCAN-LOF. Journal of Northeast Normal University (Natural Science Edition), 2025, 57(03): 47-55.

[8] Wang Jing. Research on pollutant discharge data cleaning method for catalytic cracking units based on isolation forest and neural network. China University of Petroleum (Beijing), 2020.

[9] Chen Hongqiao. Big data cleaning algorithm for the Internet of Things based on isolation forest. Information Recording Materials, 2025, 26(01): 154-156+165.

[10] Sun Ruizao, Wei Lu. Cleaning of abnormal wind power data based on isolation forest and standard deviation detection. Henan Science, 2023, 41(03): 313-320.

Downloads

Published

2026-06-24