CNN-Transformer Hybrid Models for Object Detection: A Comprehensive Review

Authors

  • Lyuyang Gao

DOI:

https://doi.org/10.61173/19bc6r78

Keywords:

CNN-Transformer Hybrid Model, Serial Ar-chitecture Fusion Approach, Parallel Architecture Fusion Method

Abstract

Initially, conventional convolutional neural networks were the primary approach for object detection, a core computer vision task. However, the emergence of Transformer architecture has significantly enhanced detection accuracy and generalization capabilities, playing a pivotal role in advancing intelligent systems across various domains. Recently, the integration of CNN and Transformer architectures has emerged as a key area of investigation for detecting objects. By combining the complementary advantages of CNNs and Transformers, these hybrid architectures enhance accuracy in various object recognition scenarios. This study commences with a concise overview of CNNs and Transformers, critically analyzing their respective advantages and limitations. Subsequently, we conduct a systematic examination of state-of-the-art hybrid architectures and their optimization strategies. Finally, a comprehensive comparison and summary are presented in tabular form to facilitate clear performance evaluation. These approaches are designed to harness CNNs’ superiority in local feature extraction while leveraging Transformers’ capacity for global context modeling. At the end of the paper, the prospects of hybrid models in object detection and the insights to guide further research have been discussed.

References

[1] Krizhevsky A, Sutskever I, Hinton G E. ImageNet classification with deep convolutional neural networks. Advances in Neural Information Processing Systems, 2012, 25: 1097-1105.

[2] Fangfang. He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016: 770-778.

[3] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in Neural Information Processing Systems, 2017, 30: 5998-6008.

[4] Carion N, Massa F, Synnaeve G, et al. End-to-end object detection with transformers. European Conference on Computer Vision, 2020: 213-229.

[5] Beal J, Kim E, Tzeng E, et al. Toward transformer-based object detection. arXiv preprint arXiv:2012.09958, 2020.

[6] Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations, 2021.

[7] Lin T Y, Maire M, Belongie S, et al. Microsoft COCO: Common objects in context. European Conference on Computer Vision, 2014: 740-755.

[8] Yan S, Xiong X, Arnab A, et al. ConTNet: Why not use convolution and transformer at the same time? arXiv preprint arXiv:2104.13497, 2021.

[9] Deng J, Dong W, Socher R, et al. ImageNet: A largescale hierarchical image database. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2009: 248-255.

[10] Peng Z, Huang W, Gu S, et al. ConFormer: Local features coupling global representations for visual recognition. Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021: 367-376.

[11] Wang X, Zhang X D, Wang G, et al. TransFusionNet: Semantic and spatial features fusion framework for liver tumor and vessel segmentation under Jetson TX2. IEEE Transactions on Medical Imaging, 2022, 41(5): 1123-1135.

[12] Chen Y, Dai X, Chen D, et al. Mobile-Former: Bridging MobileNet and transformer. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 5270-5279.

[13] Zhu X, Su W, Lu L, et al. Deformable DETR: Deformable transformers for end-to-end object detection. International Conference on Learning Representations, 2021.

[14] Bai X, Hu Z, Zhu X, et al. TransFusion: Robust LiDAR- camera fusion for 3D object detection with transformers. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 1718-1727.

[15] Liu Ze, Lin Yutong, Cao Yue, Hu Han, et al. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. arXiv:2103.14030, 2021.

Downloads

Published

2025-08-26