The Evolution of Knowledge Distillation in Image Classification Tasks
DOI:
https://doi.org/10.61173/c18pw152Keywords:
Knowledge distillation, Computer vision, Model compressionAbstract
In today’s digital age, image classification plays a crucial role as a key task in the field of computer vision. Image classification tasks aim to accurately assign images to predefined categories, but training efficient models for high-precision classification operations using large-scale datasets remains challenging. To this end, researchers have discovered knowledge distillation strategies to achieve model performance compression. Knowledge distillation aims to transform complex models into lightweight ones through parameter optimization, enabling lightweight models to learn the capabilities of complex models from limited data without altering the original model structure, achieving excellent scalability. Knowledge distillation primarily involves the transfer of knowledge through three forms: model outputs, feature map matching, and structural knowledge. This paper primarily analyzes and discusses three aspects: decoupled knowledge distillation and decision boundary interpretation structures with an outputoriented focus, knowledge distillation based on classical feature-based attention transfer and Wasserstein distance, and relationship-based virtual distillation techniques and knowledge distillation that preserves Lipschitz continuity.
References
[1] Hinton G, Vinyals O, Dean J. Distilling the Knowledge in a Neural Network. arXiv Preprint arXiv:1503.02531, 2015.
[2] Si Zhaofeng, Qi Honggang. A Review of the Research and Application of Knowledge Distillation Methods. Journal of Image and Graphics, 2023, 28(09): 2817-2832.
[3] Zhao B, Cui Q, Song R, Qiu Y, Liang J. Decoupled Knowledge Distillation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022: 11953-11962.
[4] Son W, Na J, Choi J, Hwang W. Densely Guided Knowledge Distillation Using Multiple Teacher Assistants. Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021: 9395-9404.
[5] Ojha U, Li Y, Sundara Rajan A, Liang Y, Lee Y J. What Knowledge Gets Distilled in Knowledge Distillation? Advances in Neural Information Processing Systems, 2023, 36: 11037- 11048.
[6] Romero A, Ballas N, Kahou S E, Chassang A, Gatta C, Bengio Y. FitNets: Hints for Thin Deep Nets. arXiv Preprint arXiv:1412.6550, 2014.
[7] Guo Z, Yan H, Li H, Lin X. Class Attention Transfer Based Knowledge Distillation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023: 11868-11877.
[8] Lv J, Yang H, Li P. Wasserstein Distance Rivals Kullback- Leibler Divergence for Knowledge Distillation. Advances in Neural Information Processing Systems, 2024, 37: 65445-65475.
[9] Cortes C, Mohri M, Rostamizadeh A. Algorithms for Learning Kernels Based on Centered Alignment. Journal of Machine Learning Research, 2012, 13(1): 795-828.
[10] Yang C, Yu X, An Z, Xu Y. Categories of Response-Based, Feature-Based, and Relation-Based Knowledge Distillation. In: Advancements in Knowledge Distillation: Towards New Horizons of Intelligent Systems. Cham: Springer International Publishing, 2023: 1-32.
[11] Zhang W, Xie F, Cai W, Ma C. VRM: Knowledge Distillation via Virtual Relation Matching. arXiv Preprint arXiv:2502.20760, 2025.
[12] Shang Y, Duan B, Zong Z, Nie L, Yan Y. Lipschitz Continuity Guided Knowledge Distillation. Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021: 10675-10684.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
