Counterfactual Causal Attention Learning: Enhancing Fine-Grained Visual Recognition via Indirect Effect Optimization
DOI:
https://doi.org/10.61173/h9ah2p88Keywords:
fine-grained visual recognition, attention mechanism, counterfactual attention learning, causal inference, indirect EffectAbstract
Fine-grained visual recognition (FGVR) aims to distinguish subtle differences among visually similar categories. However, conventional attention mechanisms lack quantitative approaches to evaluate the quality of the learned attention during training, which limits their effectiveness. To address this limitation, we propose a novel Counterfactual Causal Attention Learning (CCAL) framework for fine-grained image classification and person re-identification. In our approach, the attention map is modeled as a confounding variable within a causal graph, and counterfactual interventions are employed to assess its impact on model predictions. By optimizing the indirect effect (IE), CCAL enhances the reliability of attention and improves overall recognition performance. Extensive experiments on multiple FGVR benchmarks demonstrate consistent improvements, including a 1.3% Top-1 accuracy gain on the CUB-200-2011 dataset.
References
[1] Y. Rao, G. Chen, J. Lu, and J. Zhou, “Counterfactual attention learning for fine-grained visual categorization and reidentification,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Virtual, 2021.
[2] M. H. Guo, T. X. Xu, J. J. Liu, Z. N. Liu, P. T. Jiang, T. M. Mu, S. M. Martin, M. X. Ma, H. C. Zhang, Y. Q. Cui, Y. R. Xu, Z. P. Zhou, S. S. Zhou, R. B. Liang, B. F. Ding, J. H. Li, and Figure 2 The heatmap results shown above M. M. Cheng, “Attention mechanisms in computer vision: A indicate that the model has effectively survey,” IEEE Transactions on Pattern Analysis and Machine captured the head features of the Yellow- Intelligence, vol. 44, no. 3, pp. 1489-1513, 2022 headed Blackbird [3] M. H. Guo, C. Z. Lu, Z. N. Liu, M. M. Cheng, and S. M. Martin, “Visual attention network,” IEEE Transactions on E. Conclusion Pattern Analysis and Machine Intelligence, vol. 44, no. 10, This study aims to address the limitations of conventional pp. 7832-7839, 2022.K. Elissa, “Title of paper if known,” attention mechanisms in fine-grained visual recognition unpublished. (FGVR), specifically the insufficient evaluation of atten- [4] M. H. Guo et al., “Attention mechanisms in computer vision: tion quality and weak supervisory signals. Inspired by pre- A survey,” J. Adv. Res., vol. 38, pp. 215-249, 2022, doi: 10.1016/ vious work, we propose and implement a counterfactual j.jare.2021.11.006Y. Yorozu, M. Hirano, K. Oka, and Y. Tagawa, causal attention learning method designed to enhance the “Electron spectroscopy studies on magneto-optical media and model’s focus on discriminative regions while mitigating plastic substrate interface,” IEEE Transl. J. Magn. Japan, vol. 2, reliance on spurious features. pp. 740–741, August 1987 [Digests 9th Annual Conf. Magnetics By modeling the attention map (A) as a confounding vari- Japan, p. 301, 1982]. able in the causal path from input to prediction (X → Y) [5] D. Zhang, H. Zhang, J. Tang et al., “Causal intervention for and employing counterfactual interventions, our frame- weakly-supervised semantic segmentation,” in Adv. Neural Inf. work effectively evaluates attention quality and generates Process. Syst., vol. 33, 2020, pp. 2225-2236. robust supervisory signals through optimization of the [6] H. Dong, F. Han, L. Si, W. Qiang, and L. Zhang, “Background indirect effect (IE). We explore various intervention strat- Debiased SAR Target Recognition via Causal Interventional egies and attention methods to improve this approach. Regularizer,” arXiv preprint arXiv:2308.06606, 2023. Experimental results on multiple FGVR tasks, including [7] P. T. Jiang, L. H. Han, Q. Hou, M. M. Cheng et al., “Online fine-grained bird classification (CUB-200-2011), demon- attention accumulation for weakly supervised semantic strate the effectiveness of our method. Compared to segmentation,” IEEE Trans. Image Process., vol. 32, pp. 586- traditional baselines, our method significantly improves 599, 2022 Dean&Francis Hangyu Peng
[8] B. Zhang, J. Xiao, Y. Wei, M. Sun, K. Huang, “Reliability generation: A survey on causal generative modeling,” arXiv does matter: An end-to-end weakly supervised semantic preprint arXiv:2310.11011, 2023. segmentation approach,” in Proc. AAAI Conf. Artif. Intell., [15] L. De Lara, A. González-Sanz, N. Asher, and L. Risser, 2020, vol. 34, no. 7, pp. 12711–12718. “Transport-based counterfactual models,” J. Mach. Learn. Res.,
[9] Y. Zhang, Z. Zhou, Y. Cao, G. Li, S. Wu, and X. Zhang, vol. 25, pp. 1-62, 2024. “MAMC-Optimal on Accuracy and Efficiency for Automatic [16] H. Li, X. Wang, Z. Zhang, and W. Zhu, “Out-of- Modulation Classification with Extended Signal Length,” 2024 distribution generalization on graphs: A survey,” arXiv preprint International Conference on Machine Learning and arXiv:2202.07987, 2022. Cybernetics (ICMLC), vol. 1, pp. 272–276, 2024, doi: 10.1109/ [17] D. Wu, M. Ye, G. Lin, and X. Gao, “Person re-identification ICMLC61633.2024.10705364. by context-aware part attention and multi-head collaborative
[10] Z. Yang, Z. Wang, L. Luo, H. Gan, and T. Zhang, learning,” IEEE Trans. Image Process., vol. 30, pp. 4843–4856, “SWS-DAN: Subtler WS-DAN for fine-grained image 2021. classification,” Journal of Visual Communication and Image [18] M. Liu, J. Zhao, Y. Zhou, H. Zhu, and R. Yao, “Survey Representation, vol. 79, p. 103233, Aug. 2021. for person re-identification based on coarse-to-fine feature
[11] J. Ruan, G. Liang, J. Zhao, H. Zhao, J. Qiu, et al., learning,” Multimed. Tools Appl., vol. 81, no. 12, pp. 17099– “Deep learning for cybersecurity in smart grids: Review and 17124, 2022. perspectives,” IET Smart Grid, vol. 5, no. 1, pp. 2–16, 2022. [19] Peiqin Zhuang, Yali Wang, and Yu Qiao. Learning attentive
[12] B. Zhang, J. Xiao, Y. Wei, M. Sun, K. Huang, “Reliability pairwise interaction for fine-grained classification. arXiv preprint does matter: An end-to-end weakly supervised semantic arXiv:2002.10191, 2020 segmentation approach,” in Proc. AAAI Conf. Artif. Intell., [20] Jianlong Fu, Heliang Zheng, and Tao Mei. Look closer to 2020, vol. 34, no. 7, pp. 12711–12718. see better: Recurrent attention convolutional neural network for
[13] P. Sermanet, A. Frome, E. Real, “Attention for fine-grained f ine-grained image recognition. In CVPR, 2017. categorization,” arXiv preprint arXiv:1412.7054, 2014.1 [21] Yue Chen, Yalong Bai, Wei Zhang, and Tao Mei.
[14] A. Komanduri, X. Wu, Y. Wu, and F. Chen, “From Destruction and construction learning for fine-grained image identifiable causal representations to controllable counterfactual recognition. In CVPR, pages 5157–5166, 2019.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
