LSTM-Based Forum Topic Classification Model
DOI:
https://doi.org/10.61173/g3xexb16Keywords:
Deep Learning, Natural Language Processing, Long Short-Term Memory Network, Forum, Text Multi-classificationAbstract
With the rapid development of information technology, forums have become an important platform for information exchange. However, the manual classification of forum topics consumes a significant amount of human resources and is prone to classification errors. To address this issue, this study proposes a forum topic classification model based on Long Short-Term Memory (LSTM) networks. By leveraging LSTM’s capability in text processing, the accuracy and efficiency of topic classification are significantly improved. This study used approximately 68,000 entries from 29 topic categories, scraped from Zhihu, for the experiments. Preprocessing steps such as text cleaning, tokenization, and word vectorization were performed, and a classification model with LSTM and Dropout layers was designed. The experimental results indicate that the model performs well in most topic classifications, although overfitting remains an issue for certain categories. The paper concludes by summarizing the advantages and limitations of the model and discusses the potential for improving classification accuracy through increasing data volume and optimizing the model in the future.
References
[1] Han Mei. Research on Sentiment Analysis of Chinese Bullet Screen Text Based on Deep Learning [D]. Nanchang University, 2024.
[2] Liang Dengyu. Application Research of Chinese Text Multi-Classification Based on LSTM [J]. Journal of Shanghai University of Electric Power, 2020.
[3] Xin S .Multi-classification application of Chinese news text based on deep learning[J].Journal of Physics Conference Series, 2020, 1549:022011.DOI:10.1088/1742-6596/1549/2/022011.
[4] Hua, Zhang, Jiawei Qin, Yan Wang, Yuan Ma, L. Yao and Jun Lei. “Research on Android Multi-classification Based on Text.” Journal of Physics: Conference Series 1828 (2021): n. pag.
[5] Lei, Tao, Regina Barzilay and T. Jaakkola. “Molding CNNs for text: non-linear, non-consecutive convolutions.” Conference on Empirical Methods in Natural Language Processing (2015).
[6] Feng, Guozhong, Shaoting Li, Tieli Sun and Bangzuo Zhang. “A probabilistic model derived term weighting scheme for text classification.” Pattern Recognit. Lett. 110 (2018): 23-29.
[7] Machová, Kristína, Martin Mikula, Xiaoying Gao and Marián Mach. “Lexicon-based Sentiment Analysis Using the Particle Swarm Optimization.” Electronics (2020): n. pag.
[8] Zhu Lili. Research on Chinese Text Classification Based on Attention Mechanism and LSTM-CNN [D]. Chongqing University of Technology, 2023.
[9] Kong Weize, Liu Yiqun, Zhang Min, et al. Research on the Evaluation Method of Answer Quality in Question and Answer Communities [J]. Journal of Chinese Information Processing, 2011.
[10] Chang Lei, Wang Yilun, Chen Yanping. Application Research of Text Multi-Classification Based on Bert Model [J]. Dean&Francis Computer Knowledge and Technology, 2023.
[11] Zhang Xin, Zhai Zhengli, Yao Luyao. Chinese News Text Classification Based on CNN and LSTM Hybrid Model [J]. Computer and Digital Engineering, 2023.
[12] Xu Peng. Research on News Text Classification Methods Based on Deep Learning [D]. Nanjing University of Information Science and Technology, 2024. DOI:10.27248/d.cnki. gnjqc.2023.000230.
[13] Shi, Xingjian, Zhourong Chen, Hao Wang, D. Y. Yeung, Wai-Kin Wong and Wang-chun Woo. “Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting.” Neural Information Processing Systems (2015).
[14] Behera, Ranjan Kumar, Monalisa Jena, Santanu Kumar Rath and Sanjay Misra. “Co-LSTM: Convolutional LSTM model for sentiment analysis in social big data.” Inf. Process. Manag. 58 (2021): 102435.
[15] Xue Jincheng, Jiang Di, Wu Jiande. Research on Automatic Patent Text Classification Based on Word2Vec [J]. Information Technology, 2020, 44(02): 73-77. DOI:10.13274/j.cnki. hdzj.2020.02.015.
Downloads
Published
Issue
Section
License
Copyright (c) 2024 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
