A Comparative Visualization Analysis of Neural Network Models Using Grad-CAM
DOI:
https://doi.org/10.61173/yzp9wt79Keywords:
Grad-CAM, ViTs, model decision, visualization of models, deep learningAbstract
As the area of deep learning is advancing at the speed of light, the issue of model interpretability is now a priority for a better understanding and enhancement of the decision-making processes implemented by sophisticated neural networks. So, as deeper learning models create, it becomes vital to guarantee greater transparency and interpretability, in particular, in applications as medical image analytics, Auto-Driving, and Security Systems. The process of visualization of these decisions is assisted by Grad-CAM which is a powerful visualization tool. The rationale behind this work stems from the growing concern on the interpretability of deep learning models with the hope of systematically evaluating how different models attend to certain regions of images during classification. In this study, the three deep neural network models were Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), and Swin Transformer, which were used to classify the images with the help of Grad-CAM in visualizing the heatmaps of the important regions of the input images important for the models’ decision-making. The conclusively featured experimental outcomes prove that Grad-CAM can help to improve the interpretability of these deep networks irrespective of the used architecture or type. This work also extends Grad-CAM, demonstrating its capability to offer insights of model processes toward achieving better and more interpretable AI models.References
[1] Simonyan, K., Vedaldi, A., Zisserman, A. Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034,2014.
[2] Simonyan, K., Zisserman, A. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations ,2015,1-11
[3] He, K., Zhang, X., Ren, S., Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016. 770-778.
[4] Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Houlsby, N.. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2021,1-9
[5] Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D.Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision, 2017, 618- 626.
[6] Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., Torralba, A. Learning deep features for discriminative localization. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, 2921-2929.
[7] Samek, W., Wiegand, T., Müller, K. R. Explainable artificial intelligence: Understanding, visualizing and interpreting deep learning models. arXiv preprint arXiv:1708.08296, 2017.
[8] Simonyan, K., Zisserman, A. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations ,2015, 1-13.
[9] Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Guo, B. Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, 10012-10022.
[10] Kim, W., Son, B., Kim, I. ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision. arXiv preprint arXiv:2102.03334,2021.
[11] Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization. Proceedings of the IEEE International Conference on Computer Vision, 2017, 618-626.
[12] Simonyan, K., & Zisserman, A. Very deep convolutional networks for large-scale image recognition. In Proceedings of the International Conference on Learning Representations ,2015
[13] Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Guo, B.. Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021. 10012-10022.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
