Artificial Intelligence Approaches in Music Audio Analysis
DOI:
https://doi.org/10.61173/60g9eb06Keywords:
Mel Frequency Cepstral Coefficients (MFCC), Machine Learning Methods, Audio Analysis, Feature ExtractionAbstract
As people’s pursuit of art gradually increases, the forms and types of music have become increasingly diverse. In the process of exploring different types of music, the demand for music audio analysis has gradually increased. With the continuous development of artificial intelligence technology, machine learning, deep learning and other methods have gradually entered the public eye with their efficient data processing capabilities. Therefore, artificial intelligence methods in music audio analysis have gradually become the focus of research. At present, the processing mode based on traditional features such as Mel frequency cepstral coefficients (MFCC) combined with machine learning methods has been widely adopted, but there are still certain limitations in feature expression ability and classification accuracy. In recent years, endto-end deep learning methods have shown stronger adaptability and accuracy by automatically extracting features for classification and recognition, promoting the advancement of pure music audio analysis technology. This article aims to provide theoretical support and practical guidance for researchers in related fields by organizing the application of artificial intelligence methods in pure music audio analysis, comparing and analyzing the advantages and disadvantages of various methods, and promoting the sustainable development and technological innovation of this field.
References
In the task of recognizing individual techniques of guitar Machine Learning, 1995, 20(3): 273–297. playing, such as in identifying guitar sounds, the amount [3] Lu Lifei. Research and Implementation of a Music Analysis of spectral features that are extracted to a nearby set of and Retrieval Platform Based on Score Generation. Shanghai: lightweight CNNs, such as MobileNetV2, InceptionV3, Shanghai Jiao Tong University, 2019. and ResNet50 has enhanced the recognition of nine types [4] Chellamani G K, N A, C C, et al. SpectroFusionNet: A CNN of guitar sounds [4]. In spite of these, the generalization of Approach Utilizing Spectrogram Fusion for Electric Guitar Play
the model of real-world data is no more than 70.9 percent, Recognition. Scientific Reports, 2025, 15: 16842. and the performance of the model has shown a downgrade [5] V K, S S P. Hybrid Machine Learning Classification Scheme on the presence of more elaborate background noise and for Speaker Identification. Journal of Forensic Sciences, 2022, variation in performance. 67: 1033–1048. Multi-class audio classification therefore faces intercon- [6] Costantini G, Cesarini V, Brenna E. High-Level CNN and nected problems such as complexity of high-dimensional Machine Learning Methods for Speaker Recognition. Sensors,
features, scarcity of data, and noise in the environment, so 2023, 23(7): 3461. that accuracy-generalization trade-offs are especially hard [7] Hosseinzadeh M, Haider A, Malik M H, Adeli M, Mzoughi to achieve in the practice. To overcome them, researchers O, et al. Enhanced Heart Sound Classification Using Melare trying to consider them with their diversified feature Frequency Cepstral Coefficients and Comparative Analysis of fusion, transfer learning, and reinforcement learning to be Single vs. Ensemble Classifier Strategies. PLOS ONE, 2024, more robust and applicable, setting the key research prior- 19(12): e0316645. ities to be used in the future within this sphere. [8] Zhantleuova A K, Makashev Y K, Duzbayev N T. Optimizing MFCC Parameters for Breathing Phase Detection. Sensors, 2025, 25(16): 5002. 6. Conclusion [9] Wei J-Q, Wang X-Y, Zheng X-L, Tong X. Stridulatory The application of artificial intelligence methods in the Organs and Sound Recognition of Three Species of Longhorn
field of music audio analysis has become increasingly di- Beetles (Coleoptera: Cerambycidae). Insects, 2024, 15(11): 849. verse with the development of the times. From traditional [10] Li, J., Han, L., Li, X. et al. An evaluation of deep neural machine learning methods such as combining Mel fre- network models for music classification using spectrograms.
quency cepstral coefficients with support vector machines, Multimed Tools Appl 81, 4621–4647 (2022). it has gradually evolved into an end-to-end automatic ex- [11] Pandeya Y R, Bhattarai B, Lee J. Deep-Learning-Based traction of audio features for classification and recognition Multimodal Emotion Classification for Music Videos. Sensors,
in deep learning. Its application scope has gradually ex- 2021, 21(14): 4927. panded from a single music audio processing to medicine, [12] Jiang S, Shi N, Liu C. Analysis of Artificial Intelligence society, education and teaching, and so on. The accuracy, Knowledge Graphs for Online Music Learning Platform Under
robustness, and generalization ability of audio analysis Deep Learning. Scientific Reports, 2025, 15: 16481. have also become the main directions of development and [13] Oguike O E, Primus M. Multimodal Music Genre research in related fields. Classification of Sotho-Tswana Musical Videos. IEEE Access,
The continuous development of artificial intelligence 2025, 13: 28799–28808. methods in the field of pure music audio analysis not only
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
