The Evolution of Deep Learning in Medical Imaging

Authors

  • Yuanao Ye

DOI:

https://doi.org/10.61173/vhq2zk70

Keywords:

Medical imaging, Convolutional neural networks, ResNet, Transformer

Abstract

In the past 10 years, the evolution path of medical imaging AI in terms of model architecture is very clear: it is evolving from CNNs for end-to-end feature learning, to ResNets for alleviating deep degradation issues, to EfficientNet balancing accuracy and efficiency with compound scaling, to Transformers for global modeling with self-attention. It is systematic and clear, but it is also important to map it systematically, which will help to be clearer about technology selection and further optimize clinical translation. This paper focuses on four major models/captures CNN, ResNet, EfficientNet, and Transformer, summarizes their basic designs and innovation points.

References

[1] Goodfellow I, Bengio Y, Courville A, Bengio Y. Deep learning, vol. 1. Cambridge: MIT Press, 2016.

[2] Alzubaidi L, Zhang J, Humaidi A J, et al. Review of deep learning: concepts, CNN architectures, challenges, applications, future directions. Journal of Big Data, 2021, 8(1): 53.

[3] He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016: 770– 778.

[4] Tan M, Le Q V. EfficientNet: Rethinking model scaling for convolutional neural networks. Proceedings of the 36th International Conference on Machine Learning (ICML), 2019: 6105–6114.

[5] Azim M A, Jahan M U, Almotiri S H, et al. DU-Net: A dual-encoder U-Net architecture for scalable medical image segmentation. Journal of Healthcare Engineering, 2024, Article ID 3583612:,pages 1–20.

[6] Chowdhury M E H, Rahman T, Khandakar A, et al. Can AI help in screening viral and COVID-19 pneumonia? IEEE Access, 2020, 8: 132665–132676.

[7] Ghafoor S S M, Munir K, Al Farraj O, et al. A comprehensive CNN approach for Alzheimer’s disease classification using MRI scans. IEEE Access, 2022, 10: 42613–42626.

[8] Radford A, Kim J W, Hallacy C, et al. Learning transferable visual models from natural language supervision. Advances in Neural Information Processing Systems (NeurIPS), 2021, 34: 1–25.

[9] Kirillov A, Mintun E, Ravi N, et al. Segment anything. arXiv:2304.02643, 2023. [Online]. Available: https://arxiv.org/ abs/2304.02643

[10] d’ Ascoli S, Touvron H, Leavitt M L, et al. ConViT: Improving vision transformers with soft convolutional inductive biases. arXiv:2103.10697, 2021.[Online]. Available: https:// arxiv.org/abs/2103.10697

Downloads

Published

2026-02-28