Deep Reinforcement Learning in Video Games
DOI:
https://doi.org/10.61173/bct3hx20Keywords:
Deep Reinforcement Learning, Game AI, Reinforcement Learning, Multi-Agent Reinforcement Learning, Large Language ModelsAbstract
The high interactivity and instant feedback characteristics of games are highly compatible with the trial-and-error learning mechanism of Deep Reinforcement Learning (DRL). In recent years, DRL has achieved a series of landmark breakthroughs in Game AI, from Atari games to superhuman levels in complex multi-agent game environments. With the integration of Large Language Models (LLMs) with DRL, especially Multi-Agent Reinforcement Learning (MARL), DRL applications scenarios are no longer limited to performance score, but have gradually evolved into more complex situations requiring social reasoning and human-machine collaboration. This paper reviews the application of DRL methods in game AI with some famous games, which can be classified into three categories of games according to the type of games, such as classical Arcade games, First- Person 3D games and Multiplayer Competitive games. This paper conducts a detailed review of the evolution of DRL approaches. In addition, there are some challenges existing in this field, like AI agents’ generalization and human-machine collaboration capabilities. This paper also discusses current multi-agent games and some emerging methods that combine LLM with MARL, while outlining future research directions. This review aims to provide researchers with a clear technical overview.
References
[1] Shaheen A., Badr A., Abohendy A., et al. Reinforcement learning in strategy-based and Atari games: a review of Google DeepMind’s innovations. arXiv preprint arXiv:2502.10303: 2025.
[2] Mnih V., Kavukcuoglu K., Silver D., et al. Playing Atari with Deep Reinforcement Learning. arXiv preprint arXiv:1312.5602, 2013.
[3] Silver D., Huang A., Maddison C.J., et al. Mastering the game of Go with deep neural networks and tree search. Nature, 2016, 529(7587): 484–489.
[4] Silver D., Schrittwieser J., Simonyan K., et al. Mastering the game of Go without human knowledge. Nature, 2017, 550(7676): 354-359.
[5] Berner C., Brockman G., Chan B., et al. Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:1912.06680, 2019.
[6] Vinyals O., Babuschkin I., Czarnecki W.M., et al. Grandmaster level in StarCraft II using multi-agent reinforcement learning. Nature, 2019, 575(7782): 350–354.
[7] Sarkar B., Xia W., Liu C.K., et al. Training language models for social deduction with multi-agent reinforcement learning. arXiv preprint arXiv:2502.06060, 2025.
[8] Shao K., Tang Z., Zhu Y., et al. A survey of deep reinforcement learning in video games. arXiv preprint arXiv:1912.10944, 2019.
[9] Grondman I., Busoniu L., Lopes G.A., Babuska R. A survey of actor-critic reinforcement learning: Standard and natural policy gradients. IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), 2012, 42(6): 1291-1307.
[10] Foerster J.N., Chen R.Y., Al-Shedivat M., et al. Learning with Opponent-Learning Awareness. arXiv preprint arXiv:1709.04326, 2017.
[11] Claus C., Boutilier C. The dynamics of reinforcement learning in cooperative multiagent systems. AAAI/IAAI, 1998(746-752): 2.
[12] Tan M. Multi-agent reinforcement learning: Independent vs. cooperative agents. Proceedings of the Tenth International Conference on Machine Learning, 1993: 330-337.
[13] Van Hasselt H., Guez A., Silver D. Deep reinforcement learning with double Q-learning. Proceedings of the AAAI Conference on Artificial Intelligence, 2016, 30(1).
[14] Wang Z., Schaul T., Hessel M., et al. Dueling network architectures for deep reinforcement learning. International Conference on Machine Learning. PMLR, 2016: 1995-2003.
[15] Hessel M., Modayil J., Van Hasselt H., et al. Rainbow: Combining improvements in deep reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence, 2018, 32(1).
[16] Ratcliffe D.S., Devlin S., Kruschwitz U., et al. Clyde: A Deep Reinforcement Learning DOOM Playing Agent. AAAI Workshops, 2017.
[17] Jaderberg M., Mnih V., Czarnecki W.M., et al. Reinforcement Learning with Unsupervised Auxiliary Tasks. arXiv preprint arXiv:1611.05397, 2016.
[18] Bard N., Foerster J.N., Chandar S., et al. The Hanabi Challenge: A New Frontier for AI Research. Artificial Intelligence, 2020, 280: 103216.
[19] Fuchs A., Walton M., Chadwick T., et al. Theory of Mind for Deep Reinforcement Learning in Hanabi. arXiv preprint arXiv:2101.09328, 2021.
[20] Sudhakar A.V., Nekoei H., Reymond M., et al. A Generalist Hanabi Agent. arXiv preprint arXiv:2503.14555, 2025.
[21] Chen L., Lu K., Rajeswaran A., et al. Decision Transformer: Reinforcement learning via sequence modeling. Advances in Neural Information Processing Systems, 2021, 34: 15084-15097.
[22] Siu H.C., Peña J., Chen E., et al. Evaluation of human-AI teams for learned and rule-based agents in Hanabi. Advances in Neural Information Processing Systems, 2021, 34: 16183-16195.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by the authors.

This work is licensed under a Creative Commons Attribution 4.0 International License.
