Research and Analysis of Image-based Food Identification Technology in Complex Scenarios
DOI:
https://doi.org/10.61173/q0jxm119Keywords:
Food image recognition, complex scenarios, deep learningAbstract
Food image recognition, as a vital branch of finegrained visual analysis, holds significant promise for revolutionizing health management and the food industry. While deep learning models have achieved remarkable accuracy in controlled settings, their performance often degrades in real-world environments due to complex challenges. This paper presents a comprehensive review of food image recognition under these complex scenarios. This paper systematically analyzes the primary obstacles, including environmental disturbances (e.g., lighting and background variations), perspective and structural changes (e.g., occlusion and viewpoint diversity), intrinsic food variations (e.g., non-rigid deformation and inter-class similarity), and stringent system constraints (e.g., realtime and computational limits). Correspondingly, this paper surveys and discusses representative technical solutions, such as attention mechanisms, multimodal fusion, geometric transformation, and lightweight network architectures, highlighting their strengths and limitations. The review concludes that while existing methods have made substantial progress, critical issues like the accuracyefficiency trade-off, limited model generalization, and inadequate robustness in dynamic extremes remain unresolved. Future research should prioritize enhancing model adaptability in open environments, integrating semantic reasoning with perception, and developing comprehensive evaluation benchmarks to bridge the gap between laboratory research and practical deployment.
References
[1] Li Z, Liu F, Yang W, Peng S, Zhou J. A survey of convolutional neural networks: analysis, applications, and prospects. IEEE Trans Neural Netw Learn Syst. 2022;33(12):6999–7019. doi:10.1109/TNNLS.2021.3084827.
[2] Guo M-H, et al. Attention mechanisms in computer vision: a survey. Comput Vis Media. 2022;8(3):331–368. doi:10.1007/ s41095-022-0271-y.
[3] Zhang Y, Yang Q. An overview of multi-task learning. Natl Sci Rev. 2018;5(1):30–43. doi:10.1093/nsr/nwx105.
[4] Zhang C, Yang Z, He X, Deng L. Multimodal intelligence: representation learning, information fusion, and applications. IEEE J Sel Top Signal Process. 2020;14(3):478–493. doi:10.1109/JSTSP.2020.2987728.
[5] Zhao D, Wu R, Liu X, Li Y. Localization of apple picking robot in complex background based on YOLO deep convolutional neural network. Trans Chin Soc Agric Eng. 2019;35(3):164–173. Available from: CNKI:SUN:NYGU.0.2019-03-021.
[6] Zhao X, Liu P, Tang X, Liu Y. A background modeling and object detection method adaptive to outdoor illumination changes. Acta Autom Sin. 2011;37(8):915–922. Available from: CNKI:SUN:MOTO.0.2011-08-004.
[7] Su H, Zhou J, Zhang Z. A survey on super-resolution image reconstruction. Acta Autom Sin. 2013;39(8):1202–1213. Available from: CNKI:SUN:MOTO.0.2013-08-005.
[8] Zeng Q, Chen Y, Wang Y, Liu J. A fast large-view image matching algorithm based on ORB. Control Decis. 2017;32(12):2233–2239. doi:10.13195/j.kzyjc.2016.1521.
[9] Li X, Liang R. A survey on occluded face recognition: from subspace regression to deep learning. Chin J Comput. Dean&Francis Guanbo Liu 2018;41(1):177–207.
[10] Liang H, Wen X, Liang D, Zhang L. Fine-grained food image recognition using multi-level convolutional feature pyramid. J Image Graph. 2019;24(6):870–881.
[11] Liang S, Gu Y. A coarse-to-fine feature aggregation neural network with a boundary-aware module for accurate food recognition. Foods. 2025;14(3):383. doi:10.3390/ foods14030383.
[12] Sarafis I, Papadopoulos A, Delopoulos A. Weakly supervised food image segmentation using vision transformers and segment anything model. arXiv preprint arXiv:2509.19028. 2025.
[13] Mijwil MM, Doshi R, Hiran KK, Al-Mistarehi AA. MobileNetV1-based deep learning model for accurate brain tumor classification. Mesopotam J Comput Sci. 2023;2023:29– 38.
[14] Ma N, Zhang X, Zheng H-T, Sun J. ShuffleNet v2: practical guidelines for efficient CNN architecture design. In: Proc Eur Conf Comput Vis (ECCV); 2018. p.116–131.
[15] Redmon J, Farhadi A. YOLOv3: an incremental improvement. arXiv preprint arXiv:1804.02767. 2018.
[16] Iandola FN, Han S, Moskewicz MW, Ashraf K, Dally WJ, Keutzer K. SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5 MB model size. arXiv preprint arXiv:1602.07360. 2016.
[17] Sandler M, Howard A, Zhu M, Zhmoginov A, Chen L-C. MobileNetV2: inverted residuals and linear bottlenecks. In: Proc IEEE Conf Comput Vis Pattern Recognit (CVPR); 2018. p.4510– 4520.
[18] Gao H, Tian Y, Xu F, Zhang L. A survey of model compression and acceleration for deep networks. J Softw. 2021;32(1):68–92. doi:10.13328/j.cnki.jos.006096.
[19] Wu Y, Chen Y, Wang L, Ye Y, Liu Z, Guo Y. Large scale incremental learning. In: Proc IEEE/CVF Conf Comput Vis Pattern Recognit (CVPR); 2019. p.374–382.
[20] Zhang Y, Yang Q. A survey on multi-task learning. IEEE Trans Knowl Data Eng. 2021;34(12):5586–5609.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
