A review of sentiment analysis research based on BERT and its improved models
DOI:
https://doi.org/10.61173/68j5ae23Keywords:
BERT, Pre-trained Language Model, Senti-ment Analysis, Natural Language Processing, Model Im-provementAbstract
In recent years, the rapid evolution of social media and online reviews has exposed limitations in traditional sentiment analysis methods, which rely on sentiment lexicons and machine learning classifiers, particularly in terms of contextual modeling and generalization capabilities. In past few years, Deep learning, particularly convolutional neural networks (CNNs), recurrent neural networks (RNNs), and their variants, has demonstrated superior feature learning capabilities in sentiment recognition. However, these methods are still limited by the long-distance dependency modeling and the complexity of Chinese semantics. The introduction of BERT has opened a new chapter for pre-trained language models in sentiment analysis. The model’s semantic understanding capabilities have been significantly enhanced through a bidirectional Transformer architecture and large-scale corpus pre-training. Subsequently, not only ALBERT, but also improved models such as RoBERTa played a significant role in English sentiment analysis tasks after fine-tuning according to task requirements. A series of improved models such as RoBERTa-wwm-ext, ERNIE, ALBERT-zh, and MacBERT continued to refresh performance records in Chinese sentiment analysis tasks. This paper systematically reviews the research progress of sentiment analysis based on BERT and its improved models in recent years. It focuses on comparing the performance of different models in text sentiment classification, implicit sentiment recognition, and fine-grained sentiment analysis, and reveals their advantages and challenges in semantic modeling, cross-domain transfer, and multi-granularity information fusion. Finally, this paper discusses the current bottlenecks faced by the research and looks forward to the future development direction of sentiment analysis.
References
[1] Pang, B., Lee, L., & Vaithyanathan, S. (2002). Thumbs up? Sentiment classification using machine learning techniques. In Proceedings of the ACL-02 Conference on Empirical Methods in Natural Language Processing (pp. 79–86). Association for Computational Linguistics. https://doi. org/10.3115/1118693.1118704
[2] Liddy, E. D. (2001). Natural language processing. In M. A. Drake (Ed.), Encyclopedia of library and information science (Vol. 69, pp. 212–232). Marcel Dekker.
[3] Noble W S. What is a support vector machine?[J]. Nature Biotechnology, 2006, 24(12): 1565-1567.
[4] McCallum, A., & Nigam, K. (1998). A comparison of event models for naive Bayes text classification. In Proceedings of the AAAI-98 Workshop on Learning for Text Categorization (Vol. 752, pp. 41–48). AAAI Press.
[5] LeCun Y, Bottou L, Bengio Y, et al. Gradient-based learning applied to document recognition[J]. Proceedings of the IEEE, 1998, 86(11): 2278-2324.
[6] Elman J L. Finding structure in time[J]. Cognitive Science, 1990, 14(2): 179-211.
[7] Hochreiter S, Schmidhuber J. Long short-term memory[J]. Neural Computation, 1997, 9(8): 1735-1780.
[8] Bahdanau, D., Cho, K., & Bengio, Y. (2015). Neural machine translation by jointly learning to align and translate. In 3rd International Conference on Learning Representations (ICLR 2015). arXiv. https://arxiv.org/abs/1409.0473
[9] Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186). Association for Computational Linguistics. https:// doi.org/10.48550/arXiv.1810.04805
[10] Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., … & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
[11] Sun, Y., Wang, S., Li, Y., Feng, S., Chen, X., Zhang, H., … & Tian, H. (2019). ERNIE: Enhanced representation through knowledge integration. arXiv preprint arXiv:1904.09223.
[12] Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., & Soricut, R. (2019). ALBERT: A lite BERT for selfsupervised learning of language representations. arXiv preprint arXiv:1909.11942.
[13] Cui, Y., Che, W., Liu, T., Qin, B., & Yang, Z. (2020). Revisiting pre-trained models for Chinese natural language processing. Proceedings of the 2020 Conference on Empirical Methods inNaturalLanguageProcessingFindings ,657–668. https://doi.org/10.18653/v1/2020.findings-emnlp.58
[14] Clark, K., Luong, M.-T., Le, Q. V., & Manning, C. D. (2020). ELECTRA: Pre-training text encoders as discriminators rather than generators. arXiv preprint arXiv:2003.10555.
[15] Hu, M., & Liu, B. (2004). Mining and summarizing customer reviews. Proceedings of the 10th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 168–177. https://doi.org/10.1145/1014052.1014073
[16] Kim, Y. (2014). Convolutional neural networks for sentence classification. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 1746– Dean&Francis Jingxuan Chen 1751. https://doi.org/10.3115/v1/D14-1181
[17] Sun, Z., Wang, S., Li, Y., Feng, S., Chen, X., Zhang, H., … & Tian, H. (2021). ChineseBERT: Chinese pretraining enhanced by glyph and pinyin information. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics (ACL), 2065–2075. https://doi.org/10.18653/v1/2021.acl-long.164
[18] Zhang, Y., Chen, H., Li, Y., & Deng, Y. (2022). A review of Chinese sentiment analysis: Subjects, methods, and trends. Information Processing & Management, 59(2), 102805. https:// doi.org/10.1016/j.ipm.2021.102805
[19] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., … & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems (NeurIPS), 30, 5998–6008. https://doi.org/10.48550/ arXiv.1706.03762
[20] Xu, H., Li, Y., Xia, R., & Huang, W. (2020). COTE: A benchmark for Chinese implicit sentiment analysis.
[21] He, P., Liu, X., Gao, J., & Chen, W. (2021). DeBERTa: Decoding-enhanced BERT with disentangled attention. International Conference on Learning Representations (ICLR). https://doi.org/10.48550/arXiv.2006.03654
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
