Character Pose Generation and Editing Method Based on Large Model and Multimodal Prompt Words

Authors

  • Boxiao Xu

DOI:

https://doi.org/10.61173/8wq6d604

Keywords:

Pose editing, Large model, Multi-modal

Abstract

 

Character pose generation and editing is an important technology for character shaping in the fields of film, animation, and game character development. The current mainstream posture editing methods usually focus on text and image input, but overlook the equally critical audio modality; Meanwhile, existing technologies generally cannot run on consumer grade graphics cards. Therefore, this paper proposes a new technical approach: using a large model as a multimodal processor, a stable diffusion model as a transition device to generate four views after attitude changes, and finally using 3D Gaussian multi view generation models to generate the final model. Then, regarding the fusion of audio modalities, this paper introduces the “persona mask” mechanism to set a unified audio analysis method for the large model in advance, achieving accurate recognition of character emotions from speech semantics, and synchronously mapping emotional features to character actions. During this process, all models are called using interfaces. Through experiments, it can generate poses that match the character design very well on the 4070 graphics card. Hope this technology can be integrated with skeletal animation in the future, so as to complete the editing of character action animations.

References

[1] Li Y H, Hou R B, Chang H, Shan S G, Chen X L, UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing, arXiv preprint arXiv:2411.16781, 2024.

[2] Feng Y, Lin J, Dwivedi S K, Sun Y, Patel P, Black M J, ChatPose: Chatting about 3D Human Pose, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.

[3] Cha J, Kim J, CoT-Pose: Feng Y, Lin J, Dwivedi S K, Sun Y, Patel P, Black M J, ChatPose: Chatting about 3D Human Pose, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.

[4] Feng Y, Lin J, Dwivedi S K, Sun Y, Patel P, Black M J, ChatPose: Chatting about 3D Human Pose, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.

[5] Ta C K, et al., Multi-modal Pose Diffuser: A Multimodal Generative Conditional Pose Prior (MOPED), arXiv preprint arXiv:2410.14540, 2024.

[6] Lu J Z, Lin J, Dou H K, Zeng A L, Deng Y, Liu X, Cai Z G, Yang L, Zhang Y L, Wang H Q, Liu Z W, DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior, arXiv preprint arXiv:2508.00599, 2025.

[7] Delmas G, Weinzaepfel P, Moreno-Noguer F, Rogez G, PoseEmbroider: Towards a 3D, Visual, Semantic-aware Human Pose Representation, arXiv preprint arXiv:2409.06535, 2024.

[8] Shen Y T, Eum S, Lee D, Shete R, Wang C Y, Kwon H, Bhattacharyya S S, AutoComPose: Automatic Generation of Pose Transition Descriptions for Composed Pose Retrieval Using Multimodal LLMs, arXiv preprint arXiv:2503.22884, 2025.

[9] Jin Z D, Xia G Y, Yang P K, Wang M X, Sun Y B, Liu Q S, Text-driven Human Image Generation with Texture and Pose Control, Neurocomputing, 2025.

[10] Feng D, Guo P, Peng E, Zhu M, Yu W, Wang P, PoseLLaVA: Pose Centric Multimodal LLM for Fine-Grained 3D Pose Manipulation, AAAI Conference on Artificial Intelligence, vol. 39, no. 3, 2025.

Downloads

Published

2025-12-19