The Evolution of Deep Learning in Medical Imaging
DOI:
https://doi.org/10.61173/vhq2zk70Keywords:
Medical imaging, Convolutional neural networks, ResNet, TransformerAbstract
In the past 10 years, the evolution path of medical imaging AI in terms of model architecture is very clear: it is evolving from CNNs for end-to-end feature learning, to ResNets for alleviating deep degradation issues, to EfficientNet balancing accuracy and efficiency with compound scaling, to Transformers for global modeling with self-attention. It is systematic and clear, but it is also important to map it systematically, which will help to be clearer about technology selection and further optimize clinical translation. This paper focuses on four major models/captures CNN, ResNet, EfficientNet, and Transformer, summarizes their basic designs and innovation points.
References
[1] Goodfellow I, Bengio Y, Courville A, Bengio Y. Deep learning, vol. 1. Cambridge: MIT Press, 2016.
[2] Alzubaidi L, Zhang J, Humaidi A J, et al. Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. Journal of Big Data, 2021, 8(1): 53.
[3] He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016: 770– 778.
[4] Tan M, Le Q V. EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning (ICML), 2019: 6105–6114.
[5] Azim M A, Jahan M U, Almotiri S H, et al. DU-Net: A dual-encoder U-Net architecture for scalable medical image segmentation. Journal of Healthcare Engineering, 2024, Article ID 3583612:,pages 1–20.
[6] Chowdhury M E H, Rahman T, Khandakar A, et al. Can AI help in screening viral and COVID-19 pneumonia? IEEE Access, 2020, 8: 132665–132676.
[7] Ghafoor S S M, Munir K, Al Farraj O, et al. A comprehensive CNN approach for Alzheimer’s disease classification using MRI scans. IEEE Access, 2022, 10: 42613–42626.
[8] Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision. Advances in Neural Information Processing Systems (NeurIPS), 2021, 34: 1–25.
[9] Kirillov A, Mintun E, Ravi N, et al. Segment anything. arXiv:2304.02643, 2023. [Online]. Available: https://arxiv.org/ abs/2304.02643
[10] d’ Ascoli S, Touvron H, Leavitt M L, et al. ConViT: Improving vision transformers with soft convolutional inductive biases. arXiv:2103.10697, 2021.[Online]. Available: https:// arxiv.org/abs/2103.10697
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
