Research on sales forecasting in e-commerce industry for imbalanced classification data
DOI:
https://doi.org/10.61173/349hq709Keywords:
Random Over Sampler, Extra Trees Regressor, Sales Forecast, imbalanced classification dataAbstract
The background of this study is that with the advent of the big data era, e-commerce sales forecasting has become a key factor in improving the market competitiveness and economic benefits of enterprises. To solve this problem, we used machine learning technology to build a comprehensive sales forecasting system. By processing massive sales data, including data cleaning, label encoding, outlier processing and other steps, we established a complete data set. In terms of model selection, we tried multiple regression models, such as RandomForestRegressor [1], ExtraTreesRegressor, etc., and evaluated their performance through cross-validation. In order to solve the problem of data imbalance, a combination of oversampling technology (RandomOverSampler)[2] and normalization processing is used. Finally, we selected ExtraTreesRegressor as the best model and evaluated it on the training set. The research results show that the accuracy and reliability of sales forecasts can be improved by comprehensively processing sales data and selecting appropriate machine learning models. The contribution of this study in the field of e-commerce sales forecasting is to provide a comprehensive and practical solution, which provides important decision-making support for enterprises in market competition. Combining machine learning technology and data processing methods, we provide e-commerce companies with an effective sales forecasting strategy that is expected to have a positive impact in improving market competitiveness, reducing risk costs, and accelerating revenue growth [3]. Future research directions can be carried out in deeply exploring the characteristics of sales data, optimizing model parameter adjustment, and combining professional knowledge in more fields. Introducing more emerging machine learning algorithms and technologies to adapt to the changing market demands in the e-commerce field is expected to further improve the performance and adaptability of the sales forecasting system.
References
[1] Li Xinhai. (2013). Application of random forest model in classification and regression analysis. Journal of Applied Entomology (04), 1190-1197.
[2] Fang Yu, Zheng Huyu, Cao Xuemei. Three-way oversampling imbalanced data classification method [J]. Journal of Shandong University (Science Edition), 2023(012):058.
[3] Ou Jiequan. Data mining analysis based on e-commerce platform product information [J]. Electronic Technology and Software Engineering, 2016, No. 93(19): 209.
[4] Lin Muxing. Research on commodity sales forecast model using large-scale data Gaussian process regression under demand uncertainty [D]. Jinan University, 2020.
[5] Jiang Yanmei, Bu Qingkai. Supermarket Commodity Sales Forecast Based on Data Mining[J]. 2018.
[6] Pu Jiapeng. Application of machine learning in commodity sales forecast[J]. Electronic Production, 2018(22):3.DOI:CNKI: SUN:DZZZ.0.2018-22-039.
[7] Sun Puyang, Zhang Yan, Huang Jiuli. Export behavior, marginal cost and sales fluctuation - a study based on Chinese industrial enterprise data [J]. Financial Research, 2015(9):15. DOI:CNKI:SUN:JRYJ.0.2015 -09-011.
[8] Chen Yun, Wang Huanchen, Shen Huizhang. Research on price competition between e-commerce retailers and traditional retailers [J]. Systems Engineering Theory and Practice, 2006, 26(1):7.DOI:10.3321/j.issn:1000- 6788.2006.01.005.
[9] Hu Bowen. (2022). Research on e-commerce sales prediction based on deep learning (Master’s thesis, Qingdao University). https://kns.cnki.net/KCMS/detail/detail.aspx?dbname=CMFD20 2301&filename= 1022773119.nh.
[10] Yang Yuxin. (2021). Accurate market description and forecast based on big data analysis (Master’s thesis, Beijing Jiaotong University). https://link.cnki.net/doi/10.26944/d.cnki. gbfju.2021.000991doi:10.26944/d.cnki.gbfju.2021.000991.
[11] Yin Chunwu. Application of GM(1,1) in commodity sales forecast[J]. China Business and Trade, 2010(28):2.DOI:10.3969/ j.issn.1005-5800.2010.28.160.
[12] Raju V N G , Lakshmi K P , Jain V M ,et al.Study the Influence of Normalization/Transformation process on the Accuracy of Supervised Classification[C]//2020 Third International Conference on Smart Systems and Inventive Technology (ICSSIT).2020.DOI:10.1109/ ICSSIT48917.2020.9214160.
[13] Zhou Yuduan Yongrui. Retail product sales forecast based on clustering and machine learning [J]. Computer System Applications, 2021, 30(11):188-194.
[14] Liu Ying, Wei Gong, LIUYing, et al. Improvement of GRUBBS method in outlier detection [J]. Henan Science, 2006, 24(5):641-644.DOI:10.3969/j.issn.1004-3918.2006 .05.006.
[15] Tang Liang, Duan Jianguo, Xu Hongbo, et al. Feature selection algorithm and application based on mutual information maximization [J]. Computer Engineering and Applications, 2008, 44(13):4.DOI:10.3778/j.issn .1002-8331.2008.13.039.
[16] Duan Lili. Research on the impact of order flow imbalance on the yield and volatility of agricultural product futures market [D]. Harbin Institute of Technology, 2016. DOI: 10.7666/ d.D01099911.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
