Multimodal Speech Emotion Recognition Model: A Dynamic Feature Fusion Approach Based on DeBERTa and Wav2Vec2.0
DOI:
https://doi.org/10.61173/cth55r34Keywords:
Speech emotion recognition, Multimodal fusion, Temporal modeling, Cross-modal attention, DeBERTaAbstract
Speech Emotion Recognition (SER) is a core research direction in affective computing and human-computer interaction, with its primary challenge lying in effectively fusing complementary information from speech signals and textual content. This study proposes a dynamic multimodal fusion model based on Decoding-enhanced BERT (DeBERTa) and Wav2Vec2.0. By leveraging bidirectional LSTM to model audio temporal features, fine-tuning DeBERTa to optimize text representations, and incorporating a cross-modal attention mechanism for feature alignment, the model significantly enhances emotion classification performance. Experiments on the IEMOCAP dataset demonstrate that the improved model achieves an accuracy of 83.5% on the test set, representing a 19.1% improvement over baseline models. This research provides a novel technical framework for multimodal emotion understanding in complex scenarios.
References
[1] Devlin, J., Chang, M.-W., Lee, K., Toutanova, K. (2019) BERT: Pre-training of Deep Bidirectional Transformers Dean&Francis ISSN 2959-6157 for Language Understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Minneapolis, MN, pp. 4171–4186.
[2] He, P., Liu, X., Gao, J., Chen, W. (2020) DeBERTa: Decoding-enhanced BERT with Disentangled Attention. arXiv preprint, arXiv:2006.03654.
[3] Baevski, A., Zhou, H., Mohamed, A.-R., Auli, M. (2020) wav2vec 2.0: A Framework for Self-supervised Learning of Speech Representations. Advances in Neural Information Processing Systems, 33: 12449–12460.
[4] Busso, C., Bulut, M., Lee, C.C., Kazemzadeh, A., Mower, E., Kim, S., Chang, J.N., Lee, S., Narayanan, S.S. (2008) IEMOCAP: Interactive Emotional Dyadic Motion Capture Database. Journal of Language Resources and Evaluation, 42(4): 335–359.
[5] Yue, L., Chen, W., Li, X., Zuo, W., Yin, M. (2019) A Survey of Sentiment Analysis on Social Media. Knowledge and Information Systems, 60(2): 617–663.
[6] Du, K.-L., Swamy, M.N.S. (2013) Neural Networks and Statistical Learning. Springer Science and Business Media, Berlin.
[7] Sundermeyer, M., Schlüter, R., Ney, H. (2012) LSTM Neural Networks for Language Modeling. In: 13th Annual Conference of the International Speech Communication Association (INTERSPEECH). Portland, OR, pp. 194–197.
[8] Achenkun, F.A., Wenyu, C., Nunoo-Mensah, H. (2020) Text-based Emotion Detection: Progress, Challenges, and Opportunities. Engineering Reports, 2(12): e12189.
[9] Yoon, S., Byun, S., Jung, K. (2018) Multimodal Speech Emotion Recognition Using Audio and Text. In: IEEE Workshop on Spoken Language Technology (SLT). Athens, Greece, pp. 112–118.
[10] Chapuis, E., Colombo, P., Manica, M., Labeau, M., Clavel, C. (2020) Hierarchical Pre-training for Sequence Labeling in Spoken Dialog. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). pp. 2636–2648.
[11] Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Brew, J. (2019) Hugging Face’s Transformers: State-of-the-art Natural Language Processing. arXiv preprint, arXiv:1910.03771.
[12] He, P., Gao, J., Chen, W. (2021) DeBERTaV3: Improving DeBERTa using ELECTRA-style Pre-training with Gradient-disentangled Embedding Sharing. arXiv preprint, arXiv:2111.09543.
[13] Geiping, J., Goldblum, M., Pope, P., Moeller, M., Goldstein, T. (2021) Stochastic Training is Not Necessary for Generalization. arXiv preprint, arXiv:2109.14119.
[14] Minsky, M. (2007) The Emotion Machine: Commonsense Thinking, Artificial Intelligence, and the Future of the Human Mind. Simon & Schuster, New York.
[15] Bhangale, K., Kothandaraman, M. (2023) Speech Emotion Recognition Using Multi-acoustic Features and Deep Convolutional Neural Networks. Electronics, 12(3): 689–702.
[16] Michalis, P., Spyrou, E., Giannakopoulos, T., Stanitskos, G., Spouropoulos, D., Mylonas, P., Makedon, F. (2017) Comparing Deep Visual Attributes and Handcrafted Audio Features in Cross-domain Speech Emotion Recognition. Computation, 5(4): 52.
[17] Majid, W.T., Gunawan, T.S., Qadri, S.A.A., Kartiwi, M., Ambikairajah, E. (2021) A Comprehensive Review of Speech Emotion Recognition Systems. IEEE Access, 9: 47795–47814.
[18] Rashid, J., WaliTeh, Y., Hanif, F., Mujtaba, G. (2021) Deep Learning Approaches for Speech Emotion Recognition: Current Trends and Challenges. Multimedia Tools and Applications, 80(5): 8057–8083.
[19] Joshi, A., Bhat, A., Jain, A., Singh, A., Modi, A. (2022) COGMEN: COntextualized GNN based Multimodal Emotion Recognition. arXiv preprint, arXiv:2205.02455.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
