Research on the Application of Deep Reinforcement Learning in Logistics Scheduling

Authors

  • Changsheng Lu

DOI:

https://doi.org/10.61173/a4n15a63

Keywords:

Deep Reinforcement Learning, Logistics Scheduling, Order Picking, PPO Algorithm, A* Algorithm

Abstract

Traditionally, the approaches of logistics scheduling are hard to adapt to the rapid changes of the operational conditions and many of them do not have self-adaptive optimization capability, which will leads to the unsmooths and high operational expenses. In this research, we use Deep Reinforcement Learning (DRL) to improve the logistics scheduling. We take the warehouse order picking as an application example, we build the simulated warehouse environment on a grid system and design the intelligent scheduling based on Proximal Policy Optimization (PPO) algorithm. We compare the performance of it with the usual A* path planning method. The results show that the DRL model built on PPO has better results than the A* algorithm for several important measures, for example, the average steps to finish, the rate of order finish and the total reward accumulation. This investigation proves the possibility of DRL to make flexible and smart logistics scheduling, which is both theoretical reference and .a low-cost simulation tool for the intelligent transformation of logistics enterprises.

References

[1] Schulman, J., Wolski, F., Dhariwal, P., et al. (2017) Proximal Policy Optimization Algorithms. arXiv preprint arXiv:1707.06347.

[2] Yu, Y., Zhou, Z.H. (2020) Research on the Theory and Methods of Multi-agent Reinforcement Learning. Science Press, Beijing.

[3] Feng, G.Z., Liu, Y.W. (2022) Dynamic embedding deep reinforcement learning for heterogeneous vehicle routing problems with unloading time constraints. Journal of Management Sciences in China, 25(5): 1-16.

[4] Kochenderfer, M.J., Wheeler, T.A., Wray, K.H. (2019) Algorithms for Optimization. MIT Press, Cambridge.

[5] Bengio, Y., Lecun, Y., Hinton, G. (2021) Deep Learning for AI. Communications of the ACM, 64(7): 58-65.

[6] Zhang H, Li M. AGV path optimization based on improved DQN algorithm in intelligent warehouse [J]. Computer Integrated Manufacturing Systems, 2021, 27 (8):2312-2321.

[7] Liu S, Wang T. Multi-objective order picking scheduling with deep reinforcement learning under dynamic order demand [J]. Journal of Systems Engineering, 2023, 38 (2):245-254.

[8] Han X, Chen L. A review of deep reinforcement learning applications in logistics operation optimization [J]. Control and Decision, 2022, 37 (11):2421-2432.

[9] Huang J, Zhao Y. Real-time scheduling optimization of warehouse picking robots based on DDPG [J]. Computer Engineering and Applications, 2023, 59 (14):267-275.

[10] Li W, Sun Q. Dynamic warehouse layout optimization combining PPO and digital twin technology [J]. Journal of Industrial Engineering and Engineering Management, 2024, 38 (1):102-110.

[11] Gao R, Wu J. Heuristic algorithm versus reinforcement learning for order picking path planning: A comparative study [J]. International Journal of Logistics Research and Applications, 2022, 25 (9):987-1004.

[12] Wang Z. Research progress of intelligent logistics scheduling driven by deep reinforcement learning [J]. Systems Engineering Theory & Practice, 2023, 43 (5):1289-1301.

Downloads

Published

2026-08-13