Analysis of the Apple Quality Dataset
DOI:
https://doi.org/10.61173/ts88st51Keywords:
Data analysis, Machine learning, StatisticsAbstract
This article introduces a dataset containing apple features, with 4000 rows, including apple identifiers, size, weight, sweetness, crispness, juiciness, ripeness, acidity, and other characteristics. The data set can support classification and regression tasks, where quality features can be used as classification targets or converted to numerical values for regression. In addition, the data set’s features are evenly distributed, which is beneficial to model training. The experiment used two algorithms, decision tree, and random forest, for classification tasks. The results showed that the accuracy of the random forest reached 90.625%, which was better than the 80.625% of the decision tree. This confirms the effectiveness of the dataset in classification tasks and the superior performance of the random forest model.
References
Cramer, G. M., Ford, R. A., & Hall, R. L. (1976). Estimation of toxic hazard—a decision tree approach. Food and cosmetics toxicology, 16(3), 255-276.
Greenhalgh, T. (1997). How to read a paper: Statistics for the non-statistician. II: “Significant” relations and their pitfalls. BMJ, 315(7105), 422-425.
Jordan, M. I., & Mitchell, T. M. (2015). Machine learning: Trends, perspectives, and prospects. Science, 349(6245), 255- 260.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
