Investigating the Impact of Residual Connections and the Integration of VGG19 Architecture on U-Net for Car View Segmentation

Authors

  • Yunyang Wang

DOI:

https://doi.org/10.61173/fn01a143

Keywords:

Deep learning, Sematic segmentation, Computer vision, U-Net

Abstract

This study investigates the performance of three computer vision neural networks architecture, which are the standard U-Net, Deep Residual U-Net(ResU-Net), and VGG19 Integrated U-Net (VGG19U-Net) on car view segmentation. The models are trained with 4000 images and their masks and are tested at different stages of training. The validation criterion includes training and validation loss, Intersection of unions and Dice coefficient. The results demonstrate that ResU-Net outperforms the other models in segmentation accuracy while maintaining competitive prediction speeds. The VGG19U-Net shows improved performance over the standard U-Net, highlighting the benefits of deeper architectures in semantic segmentation tasks. Additionally, the research underlines the importance of architectural modifications like residual connections and deeper convolutional layers for enhancing segmentation accuracy. This study offers valuable insights into optimizing U-Net variants for vehicle segmentation, which can be extended to other real-world applications, including autonomous driving. These findings provide a view for future improvement in real-time image segmentation for complex environments.

References

[1] Van Brummelen J, O’Brien M, Gruyer D, et al. Autonomous vehicle perception: The technology of today and tomorrow. Transportation Research Part C: Emerging Technologies, 2018, 89: 384-406. https://doi.org/10.1016/j.trc.2018.02.012 Dean&Francis Yunyang Wang

[2] Ren S, He K, Girshick R, et al. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(6): 1137-1149.

[3] Shelhamer E, Long J, Darrell T. Fully convolutional networks for semantic segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017, 39(4): 640- 651. https://doi.org/10.1109/tpami.2016.2572683

[4] Ronneberger O, Fischer P, Brox T. U-Net: Convolutional networks for biomedical image segmentation. Lecture Notes in Computer Science, 2015: 234-241. https://doi.org/10.1007/978- 3-319-24574-4_28

[5] Simonyan K, Zisserman A. Very deep convolutional networks for large-scale image recognition. Computer Vision and Pattern Recognition, 2014. http://export.arxiv.org/pdf/1409.1556

[6] He K, Zhang X, Ren S, et al. Deep residual learning for image recognition. CVPR, 2016. https://doi.org/10.1109/ cvpr.2016.90

[7] Nawaz A, Akram U, Salam AA, et al. VGG-UNET for brain tumor segmentation and ensemble model for survival prediction. 2021 International Conference on Robotics and Automation in Industry (ICRAI), 2021: 1-6. https://doi.org/10.1109/ ICRAI54018.2021.9651367

[8] Zhang Z, Liu Q, Wang Y. Road extraction by deep residual U-Net. IEEE Geoscience and Remote Sensing Letters, 2018, 15(5): 749-753. https://doi.org/10.1109/lgrs.2018.2802944

[9] Mascarenhas S, Agarwal M. A comparison between VGG16, VGG19 and ResNet50 architecture frameworks for image classification. 2021 International Conference on Disruptive Technologies for Multi-Disciplinary Research and Applications (CENTCON), 2021: 96-99.

[10] Kingma DP, Ba J. Adam: A method for stochastic optimization. arXiv.org, 2014. https://doi.org/10.48550/ arXiv.1412.6980

[11] Balduzzi D, Frean M, Leary L, et al. The shattered gradients problem: If resnets are the answer, then what is the question? International Conference on Machine Learning, 2017: 342-350.

Downloads

Published

2024-12-31