Evolution, Evaluation, and Challenges of Automatic Music Style Clustering Techniques: A Review from Handcrafted Features to Self-Supervised Learning

Authors

  • Hongyu Hu

DOI:

https://doi.org/10.61173/axq15e97

Keywords:

Music style clustering, Self-supervised learning, Unsupervised representation, Music information retrieval, Custer evaluation

Abstract

Automatic discovery and clustering of music styles are fundamental to music information retrieval, recommendation systems, and computational musicology. Recent advances in deep learning and selfsupervised representation learning have substantially improved audio embeddings, enabling more effective unsupervised clustering based on musical style and acoustic characteristics. This paper reviews representative approaches to music style clustering, including traditional feature-based methods, two-stage frameworks combining deep embeddings with clustering, and emerging selfsupervised online clustering models that jointly optimize representation learning and clustering structure. Furthermore, this paper summarizes commonly used datasets and evaluation metrics in this field, reviews experimental findings from representative studies, and analyzes the advantages and limitations of various methods. Finally, this paper points out the current open challenges, including inconsistent evaluation standards, difficulties in cross-cultural and multi-label style recognition, insufficient interpretability of clustering results, and scalability issues in large-scale streaming scenarios. This paper aims to provide researchers with a clear and structured technical roadmap to facilitate the development and evaluation of future music style clustering systems.

References

Regarding experimental evaluation, deep learning meth- Tools and Applications, 2022, 81(4): 4621-4647. ods have established significant advantages on multiple [4] Kang W H, Alam J, Fathan A. An analytic study on benchmarks. However, issues such as inconsistent evalu- clustering-based pseudo-labels for self-supervised deep speaker ation standards, dataset bias, and the gap between internal verification//International Conference on Speech and Computer. Dean&Francis Hongyu Hu

Cham: Springer International Publishing, 2022: 338-348. arXiv:2208.12415, 2022 [5] Chen K, Wichern G, Germain F G, et al. Paᗧ-HuBERT: Self- [10] Sturm B L. The GTZAN dataset: Its contents, its faults, Supervised Music Source Separation Via Primitive Auditory their effects on evaluation, and its future use. arXiv preprint Clustering And Hidden-Unit Bert//2023 IEEE International arXiv:1306.1461, 2013. Conference on Acoustics, Speech, and Signal Processing [ 11 ] Wo l ff D , We y d e T. A d a p t i n g s i m i l a r i t y o n t h e

Workshops (ICASSPW). IEEE, 2023: 1-5. magnatagatune database: effects of model and feature choices// [6] Chen S, Wang C, Chen Z, et al. Wavlm: Large-scale self- Proceedings of the 21st international conference on world wide

supervised pre-training for full stack speech processing. IEEE web. 2012: 931-936.

Journal of Selected Topics in Signal Processing, 2022, 16(6): [12] Parvathi S S, Chandrasekar D. Feature separation of music 1505-1518 across diverse dataset: a comparative perspective. Bulletin of [7] Ding Y, Lerch A. Audio embeddings as teachers for music Electrical Engineering and Informatics, 2025, 14(5): 3903-3912.

classification. arXiv preprint arXiv:2306.17424, 2023. [13] Sun Y, Xu Q, Su Y, et al. AudioSet-R: A Refined AudioSet [8] Niizumi D, Takeuchi D, Ohishi Y, et al. Byol for with Multi-Stage LLM Label Reannotation//Proceedings of audio: Self-supervised learning for general-purpose audio the 33rd ACM International Conference on Multimedia. 2025: representation//2021 International Joint Conference on Neural 13089-13096.

Networks (IJCNN). IEEE, 2021: 1-8. [14] Stewart G, Al-Khassaweneh M. An implementation of the [9] Huang Q, Jansen A, Lee J, et al. Mulan: A joint embedding HDBSCAN* clustering algorithm. Applied Sciences, 2022, of music audio and natural language. arXiv preprint 12(5): 2405.

Downloads

Published

2026-02-28