A Decision-Making Framework for Task Allocation in Single-Pilot Operations: Synergistic Empowerment Mechanisms of Reinforcement Learning and Deep Learning

Authors

  • Qingyun Hou
  • Yanjun Zhou

DOI:

https://doi.org/10.61173/acyezv20

Keywords:

Single-pilot operation, reinforcement learning, deep learning, collaborative empowerment, task allocation decision-making

Abstract

The task allocation decision-making architecture for Single Pilot Operation (SPO) represents a critical development direction in aviation to address soaring operational costs and the global pilot shortage. This paper systematically reviews the synergistic enabling mechanisms of Reinforcement Learning (RL) and Deep Learning (DL) within this architecture. DL serves as the perceptual foundation, processing visual information via CNNs and optimizing human-machine interaction through Transformers to achieve efficient multimodal data comprehension and situational awareness. RL functions as the decision core, leveraging methods such as multi-agent Proximal Policy Optimization (PPO) and Deep Q-Network (DQN) to model complex task allocation problems as Markov Decision Processes, enabling dynamic resource scheduling and multi-constraint optimization. Through deep integration in a “perception-decision-optimization” closed-loop, this dual approach significantly enhances the SPO system’s responsiveness, robustness, and safety in high-real-time, high-uncertainty environments. This collaborative mechanism provides critical theoretical foundations and technical pathways for developing next-generation intelligent aviation systems compliant with airworthiness standards and enabling efficient human-machine collaboration.

References

Reinforcement learning employs algorithms such as Civil Aviation University of China, 2020, 38(6): 12-17. multi-agent near-term policy optimisation and deep Q-net- [4] Bilimoria K D, Johnson W W, Schutte P C. Conceptual works to model task allocation as a Markov decision pro- Framework for Single Pilot Operations[C]. International cess. In simulations, this approach significantly enhances Conference on Human-Computer Interaction in Aerospace,

task allocation efficiency, reduces response latency, and 2014: 1-8. demonstrates exceptional sequential decision-making [5] Liu W, Yang J Z. Review of New Human-Computer capabilities. Deep learning serves as the foundational sup- Interaction Methods in Civil Aircraft Cockpits[J]. Chinese

port for perception and representation, employing models Journal of Ergonomics, 2024, 30(3): 81-86. such as CNNs, SRCCNs, and Transformers to process [6] Yang L. Research on Multi-Agent Aircraft Path Planning multimodal data. This enhances situational awareness, Based on Deep Reinforcement Learning[D]. Harbin Engineering natural interaction, and cognitive state recognition, pro- University, 2024. viding high-quality environmental state representations [7] Liu H F. Research on Aircraft Approach Decision-Making for decision-making. RL and DL do not operate in isola- in Terminal Area Based on Deep Reinforcement Learning[D]. tion but form a deep collaborative mechanism through a Sichuan University, 2024. ‘perception-decision-optimisation’ closed loop. DL sup- [8] Dong L, Liu J, Sun Z, et al. Resource Allocation Approach ports RL’s decision inputs through front-end perception of Avionics System in SPO Mode Based on Proximal Policy

and state representation, reducing interaction latency and Optimization[J]. Aerospace, 2024, 11(10): 812-812. state uncertainty. RL, in turn, dynamically generates task [9] Shan S Z, Zhang W W. Air Combat Intelligent Decisionallocation strategies via reward-driven policy optimisa- Making Method Based on Self-Play and Deep Reinforcement tion, while leveraging techniques such as policy extraction Learning[J]. Acta Aeronautica et Astronautica Sinica, 2024, to influence DL model training and refinement. Together, 45(4): 206-218. they construct an adaptable, responsive intelligent deci- [10] Ye K W, Bao H, Wei S D. Layout Optimization for Aircraft sion system, providing a viable pathway for achieving Cockpit Man-Machine Interface Based on Visual Attention efficient collaboration among ‘pilot-airborne agent-ground Distribution[J]. Journal of Nanjing University of Aeronautics

station’. Nevertheless, integrating RL and DL within SPO and Astronautics, 2018, 50(3): 416-421. systems remains fraught with challenges. Future research [11] Razzaghi P, Tabrizian A, Guo W, et al. A Survey on must first address the strategic mismatch and insufficient Reinforcement Learning in Aviation Applications[J]. Engineering

adaptability of a single system, algorithm, or model when Applications of Artificial Intelligence, 2024, 136: 108911. confronted with various uncertainties, disturbances, or [12] Dong L, Chen H B, Chen X, et al. Distributed Multi-Agent Dean&Francis ISSN 2959-6157 Coalition Task Allocation Strategy for Single Pilot Operations 183. Based on DQN[J]. Acta Aeronautica et Astronautica Sinica, [15] Tian C W, Song M J, Zuo W M, et al. Application of 2023, 44(13): 180-195. Convolutional Neural Networks in Image Super-Resolution[J]. [13] Cheng W Y. Research on Anomaly Identification of Network CAAI Transactions on Intelligent Systems, 2025, 20(3): 719- Traffic Data Based on Deep Reinforcement Learning[J]. Journal 749.

of Intelligent Internet of Things Technology, 2025, 57(4): 87-91. [16] Wang Y Z, Li X, Mou R, et al. Air Traffic Control Speech [14] Li C, He S Q, Wei X, et al. Automatic Identification Method Enhancement Algorithm Based on Improved SEGAN[J]. Journal for Anomaly Traffic in Elastic Optical Networks Based on of Beijing University of Aeronautics and Astronautics, 2024,

Isolation Forest Algorithm[J]. Laser Journal, 2024, 45(1): 179- 50(12): 3930-3939.

Downloads

Published

2025-12-19