Transformer-Augmented MCTS for Aircraft Landing Problem
Abstract
1. Introduction
- Light (L): MTOW less than 7 t;
- Medium (M): MTOW between 7 and 136 t;
- Heavy (H): MTOW greater than 136 t.
- Satisfy all prescribed wake time separation constraints;
- Minimize the total completion time for takeoff and landing operations;
- Reduce associated landing costs.
- Time Window Constraint: Each flight must land within a specified time interval bounded by an Earliest Landing Time (ELT) and a Latest Landing Time (LLT).
- Wake separation constraint: Successive landings must adhere to the minimum wake turbulence separation intervals.
- Model Formulation: A reinforcement learning-based optimization model is established for aircraft sequencing, framing the scheduling task as a sequential decision-making process.
- Algorithm Design: A novel flight scheduling methodology is proposed, which integrates the structural strengths of the Transformer architecture with the strategic search capability of MCTS.
- Architecture Innovation: A specialized two-head output module is designed, comprising a policy head, and a value head, to effectively capture the unique dependencies and priority constraints inherent in flight scheduling.
2. Methodology
2.1. Algorithm Framework
- Transformer-based Scheduling Architecture: Flight data are initially encoded into feature vectors. These vectors are then processed through multiple encoder layers, each consisting of multi-head attention mechanisms and feed forward neural networks, to simultaneously capture the constraints of paired wake separation and the relationship between delay time and cost.
- 2.
- MCTS-Guided Strategy Optimization: Leveraging the prior probabilities provided by the Transformer, MCTS iteratively refines flight scheduling strategies through a cyclic process of selection, expansion, simulation, and backpropagation. At each decision node, the improved policy and corresponding value derived from MCTS serve as supervisory signals to train the Transformer network.
- 3.
- Actor–Critic Reinforcement Learning with Transformer Core: An Actor–Critic architecture, integrated with policy gradient methods, employs the Transformer as its central model. Decision-making is guided by MCTS during exploration, while collected trajectory data are used to update the Transformer parameters, thereby facilitating continuous policy optimization for the scheduling agent.
2.2. Aircraft Landing Issues and Reinforcement Learning Models
2.2.1. Mathematical Model for ALP
- Wake separation interval constraint:
- 2.
- Earliest Landing Time constraint:
- 3.
- Latest Landing Time constraint:
2.2.2. Markov Decision Process (MDP) Modeling
- State space
- 2.
- Action Space
- 3.
- Reward function
- Cost penalty: A negative reward proportional to the cumulative cost incurred at the current scheduling step, with dynamically adjusted weighting throughout the process.
- Completion reward: A positive reward granted upon successful scheduling of all flights. The global reward is computed using the method described in the previous section for fixed-order ALP, yielding the total cost under the constructed sequence.
- Deviation Bonus: An additional reward provided when the final scheduling cost outperforms the historical best recorded value.
- 4.
- State transition.
2.3. Solving for Aircraft Landing Times Under Fixed Sequence
2.4. Solving for the Optimal Landing Sequence
2.4.1. Monte Carlo Tree Search
2.4.2. Transformer Architecture
3. Model Training
3.1. Training Framework
3.2. Experimental Data
3.3. Parameter Settings
3.4. Training Methods
4. Results and Discussion
4.1. Results
4.2. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- 2023 Statistical Bulletin on the Development of the Civil Aviation Industry. Available online: http://www.caac.gov.cn/XXGK/XXGK/TJSJ/202505/t20250515_227513.html (accessed on 1 January 2026).
- Li, Y.Y. Analysis and Prediction of Airport Delay Based on Big Data Mining. Master’s Thesis, Nanjing University of Aeronautics and Astronautics, Nanjing, China, 2019. [Google Scholar]
- Balakrishnan, H.; Chandran, B.G. Algorithms for Scheduling Runway Operations Under Constrained Position Shifting. Oper. Res. 2010, 58, 1650–1665. [Google Scholar] [CrossRef] [Scilit]
- Xu, B. An Efficient Ant Colony Algorithm Based on Wake-Vortex Modeling Method for Aircraft Scheduling Problem. J. Comput. Appl. Math. 2017, 317, 157–170. [Google Scholar] [CrossRef] [Scilit]
- Girish, B.S. An Efficient Hybrid Particle Swarm Optimization Algorithm in a Rolling Horizon Framework for the Aircraft Landing Problem. Appl. Soft Comput. 2016, 44, 200–221. [Google Scholar] [CrossRef] [Scilit]
- Feng, X.R.; Gao, Z.D.; Wang, J.; Wang, X.L.; Hui, K.H. Research on Flight Landing Scheduling Problem Based on Compact Subsequences. J. Beijing Univ. Aeronaut. Astronaut. 2024, 50, 2421–2431. [Google Scholar] [CrossRef]
- Beasley, J.E.; Krishnamoorthy, M.; Sharaiha, Y.M.; Abramson, D. Scheduling Aircraft Landings—The Static Case. Transp. Sci. 2000, 34, 180–197. [Google Scholar] [CrossRef] [Scilit]
- Beasley, J.E.; Krishnamoorthy, M.; Sharaiha, Y.M.; Abramson, D. Displacement Problem and Dynamically Scheduling Aircraft Landings. J. Oper. Res. Soc. 2004, 55, 54–64. [Google Scholar] [CrossRef] [Scilit]
- Pinol, H.; Beasley, J.E. Scatter Search and Bionomic Algorithms for the Aircraft Landing Problem. Eur. J. Oper. Res. 2006, 171, 439–462. [Google Scholar] [CrossRef] [Scilit]
- Yu, S.-P.; Cao, X.-B.; Zhang, J. A Real-Time Schedule Method for Aircraft Landing Scheduling Problem Based on Cellular Automation. Appl. Soft Comput. 2011, 11, 3485–3493. [Google Scholar] [CrossRef] [Scilit]
- Gui, D.; Le, M.; Luo, X.; Huang, Z. A Metaheuristic Algorithm for Efficient Aircraft Sequencing and Scheduling in Terminal Maneuvering Areas. Optim. Lett. 2025, 19, 579–604. [Google Scholar] [CrossRef] [Scilit]
- Pamplona, D.A.; Alves, C.J.P. A Fast Heuristic for Aircraft Landing Scheduling with Time Windows: Application to Guarulhos Airport. Aerospace 2025, 12, 1008. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Yu, H.; Cao, K.; Zhou, J.; Wei, T.; Hu, S. Uncertainty-Aware Flight Scheduling for Airport Throughput and Flight Delay Optimization. IEEE Trans. Aerosp. Electron. Syst. 2020, 56, 853–862. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; You, L.; Zhou, M.; Yang, C.; Kang, B. Multi-objective arrival sequencing and scheduling based on point merge system. J. Beijing Univ. Aeronaut. Astronaut. 2021, 49, 66–73. [Google Scholar] [CrossRef]
- Wang, J.; Ding, X.; Wang, S. Collaborative Sequencing of Arrival and Departure Aircraft Considering Potential Conflicts. J. Transp. Syst. Eng. Inf. Technol. 2023, 23, 312. [Google Scholar] [CrossRef]
- Chen, K.; Situ, T.; Fang, Y. An Improved Multi-Objective Restart Variable Neighborhood Search Algorithm for Aircraft Sequencing Problem with Complex Interdependent Runways. J. Air Transp. Manag. 2025, 127, 102807. [Google Scholar] [CrossRef] [Scilit]
- Jiang, H.; Liu, J.X.; Zhou, W.S. Bi-level Programming Model for Joint Scheduling of Arrival and Departure Flights Based on Traffic Scenario. Trans. Nanjing Univ. Aeronaut. Astronaut. 2021, 38, 671–684. [Google Scholar] [CrossRef]
- Zhou, D.K. Terminal Area Arrival and Departure Flight Sequencing Based on Reinforcement Learning. Master’s Thesis, Civil Aviation Flight University of China, Guanghan, China, 2024. [Google Scholar]
- Kang, R.; Yang, M.; Lin, Z.Y.; Yang, Z.Y. Optimization of Arrival Flight Sequencing Based on EoR Operations. Sci. Technol. Eng. 2025, 25, 8289–8296. [Google Scholar] [CrossRef]
- Zhang, C.; Jin, Z.; Ng, K.K.; Tang, T.; Tang, Q. Distributionally robust optimisation approach for aircraft sequencing and scheduling with learning-driven arrival and departure time predictions. Omega 2026, 138, 103415. [Google Scholar] [CrossRef] [Scilit]
- Dönmez, K.; Bakır, M.; Cecen, R.K. A Comprehensive Data-Driven MCDM Approach to Determine the Best Single Objective Function for the Aircraft Sequencing and Scheduling Problem. Expert Syst. Appl. 2026, 296, 129172. [Google Scholar] [CrossRef] [Scilit]
- Coulom, R. Efficient Selectivity and Backup Operators in Monte-Carlo Tree Search. In Computers and Games; Van Den Herik, H.J., Ciancarini, P., Donkers, H.H.L.M., Eds.; Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2007; Volume 4630, pp. 72–83. ISBN 978-3-540-75537-1. [Google Scholar]
- Silver, D.; Huang, A.; Maddison, C.J.; Guez, A.; Sifre, L.; Van Den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. Mastering the Game of Go with Deep Neural Networks and Tree Search. Nature 2016, 529, 484–489. [Google Scholar] [CrossRef] [Scilit]
- Song, W.S.; Ren, B.Y.; Guan, T. Research on Arch Dam Placement Sequencing Based on Deep Monte Carlo Tree Search. J. Hydroelectr. Eng. 2024, 43, 120–130. [Google Scholar] [CrossRef]
- Wang, G.; Pu, H.; Song, T.; Li, W.; Zhang, H.; Hu, G.; Wang, G.; Pu, H.; Song, T.; Li, W.; et al. Integrated Method Based on Reinforcement Learning and Monte Carlo Tree Search for Railway Alignment Optimization. Tiedao Xuebao/J. China Railw. Soc. 2025, 47, 103–110. [Google Scholar]
- Peng, J.; Zhu, G.L.; Wu, Q.S.; Li, Y.F.; He, S.; Jin, Y.Y.; Xu, M.L. Scheduling Method for Carrier Aircraft Support Operations Based on Monte Carlo Tree Search. Acta Aeronaut. Astronaut. Sin. 2026, 47, 332444. [Google Scholar] [CrossRef]
- Yu, Z.; Huo, M.; Wang, Y.; Wang, S.; Li, Z.; Zhao, Y.; Qi, B.; Qi, N. A UAV Mission Planning Method Based on Improved Monte Carlo Tree Search. J. Astronaut. 2025, 46, 874–883. [Google Scholar] [CrossRef]
- Pang, Y.; Zhao, P.; Hu, J.; Liu, Y. Machine Learning-Enhanced Aircraft Landing Scheduling under Uncertainties. Transp. Res. Part C Emerg. Technol. 2024, 158, 104444. [Google Scholar] [CrossRef] [Scilit]
- Feng, X.R.; Zhang, S.; Qiu, D.L.; Wang, X.L. A Bounding Optimization Method for Solving Flight Landing Scheduling Problems. J. Nanjing Univ. Aeronaut. Astronaut. 2024, 56, 1024–1035. [Google Scholar] [CrossRef]
- Sabar, N.R.; Kendall, G. An Iterated Local Search with Multiple Perturbation Operators and Time Varying Perturbation Strength for the Aircraft Landing Problem. Omega 2015, 56, 88–98. [Google Scholar] [CrossRef] [Scilit]












| The Aircraft Ahead | The Aircraft Behind | ||
|---|---|---|---|
| Heavy | Medium | Light | |
| Heavy | 96 | 157 | 196 |
| Medium | 60 | 69 | 131 |
| Light | 60 | 69 | 82 |
| Model | Parameter | Value |
|---|---|---|
| MCTS | Number of simulations | 128 |
| Exploration constant | 1.414 | |
| Batch search quantity | 32 | |
| Transformer | Basic feature dimensions | 256 |
| Number of attention heads | 8 | |
| Number of encoder stacking layers | 6 | |
| Dropout | 0.1 | |
| Initial learning rate | 0.0003 | |
| Experience replay buffer capacity | 10,000 | |
| Number of training batches | 128 |
| Group | BKV | TMCTS | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Num = 32 | Num = 64 | Num = 128 | Num = 256 | ||||||
| TC | TC | t (s) | TC | t (s) | TC | t (s) | TC | t (s) | |
| 1 | 700 | 700 | 0.75 | 700 | 0.977 | 700 | 1.573 | 700 | 2.517 |
| 2 | 1480 | 1480 | 0.351 | 1480 | 0.683 | 1480 | 1.511 | 1480 | 2.804 |
| 3 | 820 | 820 | 0.457 | 820 | 0.962 | 820 | 1.919 | 820 | 3.677 |
| 4 | 2520 | 2520 | 0.536 | 2520 | 0.884 | 2520 | 2.032 | 2520 | 3.675 |
| 5 | 3100 | 6260 | 0.451 | 3100 | 0.889 | 3100 | 2.111 | 3100 | 3.656 |
| 6 | 24,442 | 24,442 | 3.114 | 24,442 | 6.301 | 24,442 | 12.844 | 24,440 | 25.559 |
| 7 | 1550 | 1550 | 1.569 | 1550 | 3.214 | 1550 | 6.686 | 1550 | 13.819 |
| 8 | 1950 | 2285 | 1.221 | 2025 | 2.499 | 1860 | 5.196 | 1860 | 9.972 |
| 9 | 5611 | 5907.85 | 5.821 | 5861.51 | 11.423 | 5748.73 | 24.811 | 5706.44 | 58.409 |
| 10 | 12,329 | 14,142.58 | 11.902 | 13,696.76 | 22.52 | 12,703.58 | 53.382 | 12,793.52 | 119.577 |
| 11 | 12,418 | 12,698.92 | 20.691 | 12,675.54 | 39.035 | 12,503.92 | 85.725 | 12,503.92 | 172.669 |
| 12 | 16,209 | 17,192.69 | 31.114 | 17,318.36 | 66.428 | 16,725.44 | 121.36 | 16,803.22 | 256.906 |
| 13 | 44,832 | 38,134.56 | 145.918 | 38,549.12 | 302.504 | 38,629.45 | 644.307 | 38,444.89 | 1221.15 |
| Average difference | —— | 13.277 | —— | −247.9 | —— | −398.298 | —— | −403 | —— |
| Average time | —— | —— | 17.22 | —— | 35.255 | —— | 74.112 | —— | 145.722 |
| Average time for the first 8 groups | —— | —— | 1.056 | —— | 2.051 | —— | 4.234 | —— | 8.21 |
| Average time for the last 5 groups | —— | —— | 43.089 | —— | 88.382 | —— | 185.917 | —— | 365.742 |
| Group | T-Only | TMCTS | Improvements |
|---|---|---|---|
| 9 | 6892.78 | 5748.73 | 16.60% |
| 10 | 18,060.00 | 12,703.58 | 29.66% |
| 11 | 13,417.16 | 12,503.58 | 6.81% |
| 12 | 17,594.72 | 16,725.44 | 4.94% |
| 13 | 39,435.04 | 38,629.45 | 2.04% |
| Average | 19,079.94 | 17,262.22 | 12.01% |
| Group | BKV | FCFS | TMCTS | DPALO+GA | DPALO+PSO | CPLEX | ||||
|---|---|---|---|---|---|---|---|---|---|---|
| TC | t (s) | TC | t (s) | TC | t (s) | TC | t (s) | |||
| 1 | 700 | 1280 | 700 | 1.573 | 700 | 0.004 | 700 | 0.006 | 700 | 0.8 |
| 2 | 1480 | 1790 | 1480 | 1.511 | 1480 | 0.028 | 1482 | 0.046 | 1480 | 0.8 |
| 3 | 820 | 1790 | 820 | 1.919 | 820 | 0.025 | 820 | 0.058 | 820 | 1.3 |
| 4 | 2520 | 4890 | 2520 | 2.032 | 2520 | 0.004 | 2520 | 0.006 | 2520 | 2.3 |
| 5 | 3100 | 6470 | 3100 | 2.111 | 3100 | 0.043 | 3244 | 0.159 | 3100 | 3.7 |
| 6 | 24,442 | 24,442 | 24,442 | 12.844 | 24,442 | 0.004 | 24,442 | 0.006 | 24,442 | 0.7 |
| 7 | 1550 | 1550 | 1550 | 6.686 | 1550 | 0.005 | 1550 | 0.007 | 1550 | 1.2 |
| 8 | 1950 | 18,870 | 1860 | 5.196 | 1863.75 | 0.056 | 1863.25 | 0.185 | 1950 | 2.2 |
| 9 | 5611 | 18,937.16 | 5748.73 | 24.811 | 5694.52 | 2.541 | 5675.01 | 3.223 | 5611.7 | 932.2 |
| 10 | 12,329 | 27,660 | 12,703.58 | 53.382 | 13,826.42 | 5.727 | 14,799.94 | 4.615 | 12,292.2 | 3600 |
| 11 | 12,418 | 35,532.91 | 12,503.92 | 85.725 | 12,671.36 | 4.706 | 12,922.42 | 5.506 | 12,418.3 | 3600 |
| 12 | 16,209 | 46,471.72 | 16,725.44 | 121.36 | 17,249.11 | 6.72 | 17,626.09 | 6.849 | 16,122.2 | 3600 |
| 13 | 44,832 | 99,776.58 | 38,629.45 | 644.307 | 38,724.21 | 21.91 | 41,323.23 | 12.052 | 37,848.9 | 3600 |
| Average difference | —— | 12,423.03 | −398.298 | —— | −255.36 | —— | 77.457 | —— | −546.6 | —— |
| Average time | —— | —— | —— | 74.112 | —— | 3.214 | —— | 2.517 | —— | 1180.4 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hu, J.; Zhang, S.; Feng, X.; Wang, X. Transformer-Augmented MCTS for Aircraft Landing Problem. Aerospace 2026, 13, 438. https://doi.org/10.3390/aerospace13050438
Hu J, Zhang S, Feng X, Wang X. Transformer-Augmented MCTS for Aircraft Landing Problem. Aerospace. 2026; 13(5):438. https://doi.org/10.3390/aerospace13050438
Chicago/Turabian StyleHu, Jie, Shuai Zhang, Xiaorong Feng, and Xinglong Wang. 2026. "Transformer-Augmented MCTS for Aircraft Landing Problem" Aerospace 13, no. 5: 438. https://doi.org/10.3390/aerospace13050438
APA StyleHu, J., Zhang, S., Feng, X., & Wang, X. (2026). Transformer-Augmented MCTS for Aircraft Landing Problem. Aerospace, 13(5), 438. https://doi.org/10.3390/aerospace13050438
