Research on Obstacle Avoidance and Path of an Intelligent Robot Based on Reinforcement Learning

Authors

  • Sinan Qi

DOI:

https://doi.org/10.61173/ra97mx26

Keywords:

Reinforcement learning, AI, robot, obstacle avoidance

Abstract

This paper focuses on the application and verification of the Q-learning algorithm and the Sarsa algorithm in a robot obstacle avoidance scenario. With the increasing application of intelligent robots, the complex dynamic environment puts forward higher requirements for their obstacle avoidance ability. Traditional obstacle avoidance algorithms are difficult to adapt to a changing environment. Reinforcement learning shows strong adaptability and obstacle avoidance effects through the interaction between robots and the environment, which has become a current research hotspot. In this paper, based on Q-learning and the Sarsa algorithm, a Python program is used to build the experimental environment, and the test scene is processed graphically to facilitate the observation of the obstacle avoidance path of the robot. Both Q-learning and the State-Action-Reward-State-Action (SARSA) algorithm avoid conventional obstacles and reach the end point by the shortest path in the experiment. In the dangerous obstacle scene, the Q-learning algorithm can still avoid obstacles and find the shortest path, while the Sarsa algorithm selects a longer route. The verification results show that the two algorithms have their advantages and disadvantages, which provides a reference for the selection and optimization of robot obstacle avoidance algorithms and has important practical significance and theoretical value. This study aims to promote the development of robot obstacle avoidance technology and provide a useful reference for research and application in related fields.

References

vironments, the reward mechanism is applied after a series Syst. 2025, 36(6): 11399-11413. of actions rather than after each action, which makes the [2] Dongbin Z, Haitao W, Kun S, Yuanheng Z. Deep logic of selection of algorithms in different environments reinforcement learning with experience replay based on SARSA. unclear. Thus, the desired results cannot be achieved [4]. IEEE Symposium Series on Computational Intelligence. 2016. Or refer to the articles of Sreyas Ramesh and Shamima [3] Li F, Qu H, Zhang L, Fu M, Chen W, Yi Z. Q-ADER: An Najnin to combine the two algorithms for special scenar- effective q-learning for recommendation with diminishing action

ios, such as cross-context noun learning and microgrid space. IEEE Trans Neural Netw Learn Syst. 2025, 36(5): 8510- energy management, and develop and upgrade strategies 8524. on the original ability to improve the effectiveness of the [4] Bonyadi MR, Wang R, Ziaei M. Self-punishment and reward algorithm in specific scenarios. Improved solutions for backfill for deep q-learning. IEEE Trans Neural Netw Learn

different scenarios [5, 6]. Syst. 2023, 34(10): 8086-8093. [5] Ramesh S, N SB, Sathyavarapu SJ, Sharma V, A A NK, Khanna M. Comparative analysis of Q-learning, SARSA, and 7. Conclusion deep Q-network for microgrid energy management. Sci Rep.

According to the above experimental results, it can be 2025, 15(1): 694. found that both algorithms have advantages and disadvan- [6] Najnin S, Banerjee B. Pragmatically framed cross-situational tages, but the application scenarios of robots in reality are noun learning using computational reinforcement models. Front

more complex: for example, more complex routes, dif- Psychol. 2018, 9: 5.

Downloads

Published

2025-08-26