Exploring Backbone Network Choices in FCOS3D: Performance and Efficiency Analysis
DOI:
https://doi.org/10.61173/15b9t728Keywords:
3D object detection, FCOS3D, FCOS-Swin, FCOS-ConvNeXt, nuScenes dataset, performance com-parisonAbstract
This study presents a comparative analysis of the performance of two modified object detection models, FCOS-Swin and FCOS-ConvNeXt, against the FCOS3D baseline model using the nuScenes dataset. The study evaluates the models based on their classification results for various categories of objects and on multiple evaluation metrics. We compare FCOS-Swin and FCOS-ConvNeXt, which utilize different backbone architectures, to evaluate their effectiveness in 3D object detection. Results show that the modified models exhibit comparable performance with slight variations in all metrics compared to the baseline, but fall short of the fine-tuned FCOS3D model. Potential reasons for this performance gap, including model parameter size, data augmentation methods, learning rate settings, and training epochs, are discussed. This study also explores possible improvements and future work, such as switching to larger backbone models, utilizing stronger data augmentation techniques, adjusting the learning rate method, increasing training epochs, and incorporating temporal and spatial logic to optimize model performance.
References
[1] Mao, J., Shi, S., Wang, X., & Li, H. (2023). 3D object detection for autonomous driving: A comprehensive survey. International Journal of Computer Vision, 131(8), 1909-1963.
[2] Singh, A. (2023). Surround-view vision-based 3d detection for autonomous driving: A survey. In 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW).pp. 3235-3244. IEEE.
[3] Zheng, C., Wang, F., Wang, N., Cui, S., & Li, Z. (2024). Towards Flexible 3D Perception: Object-Centric Occupancy Completion Augments 3D Object Detection. https://arxiv.org/ abs/2412.05154.
[4] Wang, T., Zhu, X., Pang, J., & Lin, D. (2021). Fcos3d: Fully convolutional one-stage monocular 3d object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 913-922.
[5] Tian, Z., Chu, X., Wang, X., Wei, X., & Shen, C. (2022). Fully convolutional one-stage 3d object detection on lidar range images. Advances in Neural Information Processing Systems, 35, 34899-34911.
[6] Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., ... & Beijbom, O. (2020). nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/ CVF conference on computer vision and pattern recognition. pp. 11621-11631.
[7] He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770- 778.
[8] Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., ... & Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/ CVF international conference on computer vision. pp. 10012- 10022.
[9] Liu, Z., Mao, H., Wu, C. Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022). A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 11976-11986.
[10] Girshick, R., Donahue, J., Darrell, T., & Malik, J. (2014). Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 580-587.
[11] Sapkota, R., Qureshi, R., Flores-Calero, M., Badgujar, C., Nepal, U., Poulose, A., ... & Karkee, M. (2024). Yolov10 to its genesis: A decadal and comprehensive review of the you only look once series. https://arxiv.org/abs/2406.19407.
[12] Dosovitskiy, A. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. https://arxiv.org/ abs/2010.11929.
[13] Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., & Zagoruyko, S. (2020). End-to-end object detection with transformers. In European conference on computer vision. pp. 213-229. Cham: Springer International Publishing.
[14] Zhao, Y., Lv, W., Xu, S., Wei, J., Wang, G., Dang, Q., ... & Chen, J. (2024). Detrs beat yolos on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 16965-16974.
[15] Shi, S., Wang, X., & Li, H. (2019). Pointrcnn: 3d object proposal generation and detection from point cloud. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 770-779.
[16] Lang, A. H., Vora, S., Caesar, H., Zhou, L., Yang, J., & Beijbom, O. (2019). Pointpillars: Fast encoders for object detection from point clouds. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. Dean&Francis ISSN 2959-6157 12697-12705.
[17] Yan, Y., Mao, Y., & Li, B. (2018). Second: Sparsely embedded convolutional detection. Sensors, 18(10), 3337.
[18] Chen, Y., Liu, S., Shen, X., & Jia, J. (2020). Dsgn: Deep stereo geometry network for 3d object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 12536-12545.
[19] Li, Z., Wang, W., Li, H., Xie, E., Sima, C., Lu, T., ... & Dai, J. (2022). Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal transformers. In European conference on computer vision. pp. 1-18. Cham: Springer Nature Switzerland.
[20] Duan, K., Bai, S., Xie, L., Qi, H., Huang, Q., & Tian, Q. (2019). Centernet: Keypoint triplets for object detection. In Proceedings of the IEEE/CVF international conference on computer vision. pp. 6569-6578.
[21] Liu, Z., Wu, Z., & Tóth, R. (2020). Smoke: Single-stage monocular 3d object detection via keypoint estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops. pp. 996-997.
[22] Wang, T., Xinge, Z. H. U., Pang, J., & Lin, D. (2022). Probabilistic and geometric depth: Detecting objects in perspective. In Conference on Robot Learning. pp. 1475-1485. PMLR.
[23] Li, P., Chen, X., & Shen, S. (2019). Stereo r-cnn based 3d object detection for autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 7644-7652.
[24] Philion, J., & Fidler, S. (2020). Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XIV 16. pp. 194-210. Springer International Publishing.
[25] Huang, J., Huang, G., Zhu, Z., Ye, Y., & Du, D. (2021). Bevdet: High-performance multi-camera 3d object detection in bird-eye-view. https://arxiv.org/abs/2112.11790.
[26] Wang, Y., Guizilini, V. C., Zhang, T., Wang, Y., Zhao, H., & Solomon, J. (2022). Detr3d: 3d object detection from multi-view images via 3d-to-2d queries. In Conference on Robot Learning. pp. 180-191. PMLR.
[27] Vora, S., Lang, A. H., Helou, B., & Beijbom, O. (2020). Pointpainting: Sequential fusion for 3d object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 4604-4612.
[28] Huang, T., Liu, Z., Chen, X., & Bai, X. (2020). Epnet: Enhancing point features with image semantics for 3d object detection. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XV 16. pp. 35-52. Springer International Publishing.
[29] Pang, S., Morris, D., & Radha, H. (2020). CLOCs: Camera- LiDAR object candidates fusion for 3D object detection. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). pp. 10386-10393. IEEE.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
