Generation of Dance Movement Sequences Based on Audio Information

Authors

  • Ziyao Meng

DOI:

https://doi.org/10.61173/41xw3p78

Keywords:

Action sequence generation, neural net-works, deep learning, multi-modal generation, dance movement generation

Abstract

The task of action sequence generation based on audio is a cross-modal generation task, which automatically generates continuous action sequences with similar or consistent time, semantics or emotion with the input information through the input audio signal. With advances in computer vision, as well as digital entertainment, methods that link human speech to digital body movements have made rapid progress. At present, the main methods are based on neural networks and deep learning for multi-modal generation. Through the multi-modal generation method based on computer network, it often faces the problems of low correlation between the generated content and the input content, low overall fluency, and unclear emotional expression. Based on the systematic review of the existing literature, this paper analyzes the mainstream methods in the task of dance movement generation. Finally, the key problems to be solved in this field are discussed, and the future research content and direction are prospected.

References

[1] Fan Rukun, Xu Songhua, Geng Weidong. Example-based automatic music-driven conventional dance motion synthesis. IEEE transactions on visualization and computer graphics, 2011, 18(3): 501-515.

[2] Ferstl Y, Neff M, McDonnell R. Multi-objective adversarial gesture generation, Proceedings of the 12th ACM SIGGRAPH Conference on Motion, Interaction and Games. 2019: 1-10.

[3] Hu Hao, Liu Changhong, Chen Yong, et al. Multi-scale cascaded generator for music-driven dance synthesis, 2022 International Joint Conference on Neural Networks (IJCNN). IEEE, 2022: 1-7.

[4] Fukayama S, Goto M. Music content driven automated choreography with beat-wise motion connectivity constraints. Proceedings of SMC, 2015: 177-183.

[5] Luka C, Louise C. Generative choreography using deep learning. arXiv preprint arXiv:1605.06921, 2016.

[6] Tang Taoran, Jia Jia, Mao Hanyang. Dance with melody: An lstm-autoencoder approach to music-oriented dance synthesis, Proceedings of the 26th ACM international conference on Multimedia. 2018: 1598-1606.

[7] Aristidou A, Yiannakidis A, Aberman K, et al. Rhythm is a dancer: Music-driven motion synthesis with global structure. IEEE transactions on visualization and computer graphics, 2022, 29(8): 3519-3534.,

[8] Ferreira J P, Coutinho T M, Gomes T L, et al. Learning to dance: A graph convolutional adversarial network to generate realistic dance motions from audio. Computers & Graphics, 2021, 94: 11-21.

[9] Alexanderson S, Nagy R, Beskow J, et al. Listen, denoise, action! audio-driven motion synthesis with diffusion models. ACM Transactions on Graphics (TOG), 2023, 42(4): 1-20.

[10] Zhang Yue, Research on the Intelligent Generation Method of Dance Motions Based on Deep Learning, Shanghai: Shanghai University, 2023.

Downloads

Published

2025-08-26