A Comparative Study of Inception V3 and InceptionResNetV2 for Bathroom Item Classification with and without ImageNet Pretraining
DOI:
https://doi.org/10.61173/3c99b006Keywords:
Inception V3, InceptionResNetV2, ImageN-et pretraining, bathroom item datasetAbstract
Accurate classification of bathroom items presents notable challenges due to the objects’ small sizes, high visual similarity, and frequent background clutter. The investigation focuses on the influence of convolutional neural network architecture and pretraining strategy on the performance of image classification models in such fine-grained scenarios. Two widely used architectures, InceptionV3 and InceptionResNetV2, were selected for evaluation under two training regimes: training from scratch and transfer learning via ImageNet pretraining. A curated dataset containing ten categories of common bathroom items was used for training and testing. Model performance was quantitatively assessed using overall accuracy, macro-averaged precision, recall, and F1-score, alongside qualitative analysis through confusion matrices. Experimental results demonstrate that ImageNet pretraining can significantly enhance model performance across all metrics. InceptionResNetV2 with ImageNet weights achieved the highest accuracy of 96.19%, while models trained from random initialization showed unstable convergence and poor generalization, often collapsing into predicting a dominant class. The superior performance of pretrained models is attributed to the reuse of domain-invariant features learned from large-scale datasets, which serve as effective initializations for downstream tasks with limited labeled data. These findings confirm the effectiveness of transfer learning in small-sample visual classification and highlight the additional benefit of residual connections in deeper architectures when fine-tuning on domain-specific tasks.
References
[1] Hou X, Yang Y. Research and application of image classification based on convolutional neural networks. Electronic Components and Information Technology, 2022, 6(11): 93-97.
[2] Xiao Z, Wang X, Yang B, et al. Research on painting image classification based on convolutional neural networks. Journal of China University of Metrology, 2017, 28(2): 226-233.
[3] Revathi K, Kumar S. V. Development of medical image retrieval and classification using YOLOv7 segmentation and Inception V3 classifier. Proceedings of the 2024 9th International Conference on Communication and Electronics Systems (ICCES), Coimbatore, India, 2024: 1169-1174.
[4] Deng G, et al. Image classification and detection of cigarette combustion cone based on Inception Resnet V2. Proceedings of the 2020 5th International Conference on Computer and Communication Systems (ICCCS), Shanghai, China, 2020: 395- 399.
[5] Yulita I. N, Ardiansyah F, Sholahuddin A, Rosadi R, Trisanto A, Ramdhani M. R. Garbage classification using Inception V3 as image embedding and extreme gradient boosting. Proceedings of the 2024 ASU International Conference in Emerging Technologies for Sustainability and Intelligent Systems (ICETSIS), Manama, Bahrain, 2024: 1394-1398.
[6] Kaggle. Common Objects in Bathroom Dataset. [EB/OL]. https://www.kaggle.com/datasets/mehantkammakomati/cobcommon-objects-in-bathroom?select=sink
[7] Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z. Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016: 2818-2826.
[8] Szegedy C, Ioffe S, Vanhoucke V, Alemi A. Inception-v4, inception-resnet and the impact of residual connections on learning. Proceedings of the AAAI Conference on Artificial Intelligence, 2017, 31(1).
[9] Huh M, Agrawal P, Efros A. A. What makes ImageNet good for transfer learning? arXiv preprint arXiv:1608.08614, 2016.
[10] He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016: 770- 778.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
