Application of Knowledge Distillation in Natural Language Processing

Authors

  • Xinyi Xu

DOI:

https://doi.org/10.61173/jkr5zx03

Keywords:

Knowledge distillation, Natural language processing, Teacher-student framework, Pre-trained language models

Abstract

As the technology of AI progresses steadily, natural language processing (NLP) has become an essential field of study in computer science and AI. It includes technologies that allow computers to comprehend and analyze as well as create human language. The introduction of huge pre-trained language models like Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformer (GPT) has boosted the performance of models to a great extent. However, these models face challenges such as their massive parameter counts, high computational costs, and incompatibility with resource-constrained embedded devices. As one of the effective model compression methods, knowledge distillation (KD) is a method of knowledge transfer between large models and lightweight models in the form of a teacher-student paradigm that has shown considerable performance-to-efficiency benefits. This paper is a review of the fundamental technical strategies in output layer distillation, feature layer distillation and multi-teacher assisted distillation. It synthesizes materials of pertinent research papers; therefore, providing a systematic account of the current situation in this technology. It briefly describes the common application cases and experimental findings of knowledge distillation in natural language processing such as sentiment analysis, text classification, multilingual processing, named entity recognition, and web filtering. This paper ends by summarizing existing knowledge distillation applications on natural language processing and future development perspectives.

References

[1] Hinton G, Vinyals O, Dean J. Distilling the knowledge in a neural network. Computer Science, 2015, 14(7): 38–39.

[2] Vakili Y Z, Fallah A, Sajedi H. Distilled BERT model in natural language processing. 14th International Conference on Computer and Knowledge Engineering (ICCKE), Mashhad, Iran, 2024: 243–250.

[3] Mei T, Zi Y, Cheng X, Gao Z, Wang Q, Yang H. Efficiency optimization of large-scale language models based on deep learning in natural language processing tasks. 2nd International Conference on Sensors, Electronics and Computer Engineering (ICSECE), Jinzhou, China, 2024: 1231–1237.

[4] Salmony M Y A, Faridi A R. BERT distillation to enhance the performance of machine learning models for sentiment analysis on movie review data. 9th International Conference on Computing for Sustainable Global Development (INDIACom), New Delhi, India, 2022: 400–405.

[5] Jiao X, Yin Y, Shang L, Jiang X, Chen X, Liu L. TinyBERT: Distilling BERT for natural language understanding. arXiv:1909.10351, 2019.

[6] Dong X, Huang O, Thulasiraman P, Mahanti A. Improved knowledge distillation via teacher assistants for sentiment Dean&Francis Xinyi Xu analysis. IEEE Symposium Series on Computational Intelligence (SSCI), Mexico City, Mexico, 2023: 300–305.

[7] Jin C, Yang S. Named entity recognition method based on multi-teacher collaborative cyclical knowledge distillation. 27th International Conference on Computer Supported Cooperative Work in Design (CSCWD), Tianjin, China, 2024: 230–235.

[8] Chen Z, Hu T, Chen C, Ge J, Wu C, Cheng W. An adaptive knowledge distillation algorithm for text classification. IEEE International Conference on Emergency Science and Information Technology (ICESIT), Chongqing, China, 2021: 439–442.

[9] Lin L. Multilingual text classification based on deep learning models. 11th Joint International Information Technology and Artificial Intelligence Conference (ITAIC), Chongqing, China, 2023: 1202–1205.

[10] C. Ye and A. A. Hernandez, “Knowledge Distillation Scheme for Named Entity Recognition Model Based on BERT,” 2023 5th International Conference on Machine Learning, Big Data and Business Intelligence,Hangzhou, China, 2023, pp. 10-17, doi: 10.1109/MLBDBI60823.2023.10482264.

[11] Vörös T, Bergeron S P, Berlin K. Web content filtering through knowledge distillation of large language models. IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology (WI-IAT), Venice, Italy, 2023: 357–361.

Downloads

Published

2026-02-28