Research and Prospects of Multimodal Technology of AIGC in NPC Dialogue Generation

Authors

  • Yuanman Li

DOI:

https://doi.org/10.61173/cph6gs49

Keywords:

Multimodal Learning, NPC Dialogue Gen-eration, AI-Generated Content

Abstract

In today’s technologically advanced world, in-game NPC dialogue with players is a crucial component of modern game production, directly impacting player immersion and engagement. The creation of large language models such as GPT-4V and VIMA, along with other advances in multimodal learning methods, has significantly enhanced the ability of game NPC systems to recognize and respond to diverse signals and complex scenarios. This paper examines the current state of the art in multimodal dialogue production in great detail, covering common methods, widely used datasets, and evaluation criteria. This work proposes solutions to these issues, including leveraging lightweight model structures, developing effective methods for aligning data from various modalities, and creating improved multimodal datasets specifically designed for gaming environments. This paper aims to inform the next generation of NPC dialogue systems using multimodal AIGC by examining the strengths and weaknesses of current approaches. This will help in-game dialogue become more meaningful, characters more aware of their surroundings, and interactions feel more realistic.

References

[1] Ammanabrolu P, et al. Learning to speak and act in a fantasy text adventure game. ACL. 2021.

[2] Radford A, et al. Learning transferable visual models from natural language supervision. ICML. 2021.

[3] Alayrac J B, et al. Flamingo: a visual language model for few-shot learning. NeurIPS. 2022.

[4] Bubeck S, et al. Sparks of artificial general intelligence: early experiments with GPT-4. arXiv: 2303.12712. 2023.

[5] OpenAI. GPT-4 technical report. arXiv.2303.08774. 2023.

[6] Driess D, et al. PaLM-E: an embodied multimodal language model. ICML. 2023.

[7] Hendricks L A, et al. Measuring bias in multimodal models. CVPR. 2021.

[8] Sutskever I, Vinyals O, Le QV. Sequence to sequence learning with neural networks. NeurIPS. 2014.

[9] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. NeurIPS. 2017.

[10] Radford A, Wu J, Child R, et al. Language models are unsupervised multitask learners. OpenAI Technical Report. 2019.

[11] Brown T B, Mann B, Ryder N, et al. Language models are few-shot learners. NeurIPS. 2020.

[12] Ouyang L, et al. Training language models to follow instructions with human feedback. NeurIPS. 2022.

[13] Touvron H, et al. LLaMA: open and efficient foundation language models. arXiv.2302.13971. 2023.

[14] Li Y, et al. LLaVA: large language and vision assistant. arXiv.2304.08485. 2023.

Downloads

Published

2025-12-19