Behavioral Analysis of Transfer Learning in DQN-Based Pedestrian Agents Using Single-Agent Pretraining
Abstract
1. Introduction
2. Background and Related Work
2.1. Pedestrian Simulation
2.2. DQN and Conservative Q-Learning
2.3. Transfer Learning in Reinforcement Learning
3. Two-Stage Transfer-Learning Framework
3.1. Overview
3.2. Adaptive Multi-Task Pretraining
3.3. Transfer to Multi-Agent Learning
4. Simulation Experiments
4.1. Experimental Setup
4.1.1. Pedestrian Agent and Learning Parameters
4.1.2. Source and Target Tasks
4.2. Source-Task Pretraining
4.3. Target-Task Learning Performance
4.4. Target-Task Training Time
4.5. Behavioral Analysis of Transferred Knowledge
4.5.1. Spatial Exploration During Early Learning
4.5.2. Post-Training Pedestrian Trajectories
4.6. Discussion
4.6.1. Mechanism of Learning Acceleration and Transferable Knowledge
4.6.2. Limitation in Interaction-Dominated Environments
4.6.3. Interpretability, Limitations, and Future Work
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CQL | Conservative Q-Learning |
| DQN | Deep Q-Network |
| MDP | Markov Decision Process |
| ORCA | Optimal Reciprocal Collision Avoidance |
| RL | Reinforcement Learning |
| SFM | Social Force Model |
References
- Paduraru, C.; Paduraru, M. Pedestrian motion in simulation applications using deep learning. In Proceedings of the 2022 IEEE/ACM 6th International Workshop on Games and Software Engineering (GAS); IEEE: Pittsburgh, PA, USA, 2022; pp. 1–8. [Google Scholar]
- Zhang, L.; Liu, M.; Wu, X.; AbouRizk, S.M. Simulation-based route planning for pedestrian evacuation in metro stations: A case study. Autom. Constr. 2016, 71, 430–442. [Google Scholar] [CrossRef] [Scilit]
- Lei, W.; Li, A.; Gao, R.; Hao, X.; Deng, B. Simulation of pedestrian crowds’ evacuation in a huge transit terminal subway station. Phys. A Stat. Mech. Its Appl. 2012, 391, 5355–5365. [Google Scholar] [CrossRef] [Scilit]
- Xu, J.; Peng, Y.; Ye, C.; Gao, S.; Cheng, M. Hospital flow simulation and space layout planning based on low-trust social force model. IEEE Access 2024, 12, 90135–90144. [Google Scholar] [CrossRef] [Scilit]
- Jaros, M.; Di Angelo, M.; Ferschin, P. Modeling and simulation of pedestrian behaviour: As planning support for building design. In Proceedings of the 6th International Conference on Simulation and Modeling Methodologies, Technologies and Applications (SIMULTECH 2016), Lisbon, Portugal, 29–31 July 2016; SciTePress: Setúbal, Portugal, 2016; pp. 149–156. [Google Scholar]
- Mandal, T.; Rao, K.R.; Tiwari, G. Evacuation of metro stations: A review. Tunn. Undergr. Space Technol. 2023, 140, 105304. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Shi, T.; Li, N. Pedestrian evacuation simulation in indoor emergency situations: Approaches, models and tools. Saf. Sci. 2021, 142, 105378. [Google Scholar] [CrossRef] [Scilit]
- Senanayake, G.P.D.P.; Kieu, M.; Zou, Y.; Dirks, K. Agent-based simulation for pedestrian evacuation: A systematic literature review. Int. J. Disaster Risk Reduct. 2024, 111, 104705. [Google Scholar] [CrossRef] [Scilit]
- Felemban, E.; Hammad, M.; Ur Rehman, F. Crowd Simulation: A Multi-Dimensional Systematic Mapping Study and Taxonomy. ISPRS Int. J. Geo-Inf. 2026, 15, 223. [Google Scholar] [CrossRef] [Scilit]
- Helbing, D.; Molnár, P. Social force model for pedestrian dynamics. Phys. Rev. E 1995, 51, 4282–4286. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- van den Berg, J.; Guy, S.J.; Lin, M.; Manocha, D. Reciprocal n-body collision avoidance. In Robotics Research: The 14th International Symposium ISRR; Springer: Berlin/Heidelberg, Germany, 2011; pp. 3–19. [Google Scholar]
- Martinez-Gil, F.; Lozano, M.; Fernández, F. Strategies for simulating pedestrian navigation with multiple reinforcement learning agents. Auton. Agents Multi-Agent Syst. 2015, 29, 98–130. [Google Scholar] [CrossRef] [Scilit]
- Ravichandran, N.B.; Yang, F.; Peters, C.; Lansner, A.; Herman, P. Pedestrian simulation as multi-objective reinforcement learning. In Proceedings of the 18th ACM International Conference on Intelligent Virtual Agents; Association for Computing Machinery: New York, NY, USA, 2018; pp. 307–312. [Google Scholar]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, Y.; Chai, Z.; Lykotrafitis, G. Deep reinforcement learning with a particle dynamics environment applied to emergency evacuation of a room with obstacles. Phys. A Stat. Mech. Its Appl. 2021, 571, 125845. [Google Scholar] [CrossRef] [Scilit]
- Huang, Z.; Liang, R.; Xiao, Y.; Fang, Z.; Li, X.; Ye, R. Simulation of pedestrian evacuation with reinforcement learning based on a dynamic scanning algorithm. Phys. A Stat. Mech. Its Appl. 2023, 625, 129011. [Google Scholar] [CrossRef] [Scilit]
- Vizzari, G.; Cecconello, T. Pedestrian simulation with reinforcement learning: A curriculum-based approach. Future Internet 2023, 15, 12. [Google Scholar] [CrossRef] [Scilit]
- Vizzari, G.; Falbo, A.; Tenderini, R.; Briola, D. RL-Godot pedestrian simulation: Curriculum-based reinforcement learning for pedestrian simulation. Acta Polytech. CTU Proc. 2026, 57, 323–329. [Google Scholar] [CrossRef] [Scilit]
- Sharma, J.; Andersen, P.-A.; Granmo, O.-C.; Goodwin, M. Deep Q-Learning with Q-Matrix Transfer Learning for Novel Fire Evacuation Environment. IEEE Trans. Syst. Man Cybern. Syst. 2021, 51, 7363–7381. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Z.; Lin, K.; Jain, A.K.; Zhou, J. Transfer Learning in Deep Reinforcement Learning: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 13344–13362. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Taylor, M.E.; Stone, P. Transfer learning for reinforcement learning domains: A survey. J. Mach. Learn. Res. 2009, 10, 1633–1685. [Google Scholar]
- Hughes, R.L. A continuum theory for the flow of pedestrians. Transp. Res. Part B Methodol. 2002, 36, 507–535. [Google Scholar] [CrossRef] [Scilit]
- Tordeux, A.; Lämmel, G.; Hänseler, F.S.; Steffen, B. A mesoscopic model for large-scale simulation of pedestrian dynamics. Transp. Res. Part C Emerg. Technol. 2018, 93, 128–147. [Google Scholar] [CrossRef] [Scilit]
- Yang, J.; Zang, X.; Chen, W.; Luo, Q.; Wang, R.; Liu, Y. Improved social force model based on pedestrian collision avoidance behavior in counterflow. Phys. A Stat. Mech. Its Appl. 2024, 642, 129762. [Google Scholar] [CrossRef] [Scilit]
- Siddharth, S.M.P.; Perumal, V. Development of the Social Force Model Considering Pedestrian Characteristics and Behavior. Transp. Res. Rec. J. Transp. Res. Board 2024, 2678, 436–450. [Google Scholar] [CrossRef] [Scilit]
- Zhao, X.; Li, W.; Mo, Z.; Xue, Y.; Wu, H. Simulation of Pedestrian Grouping and Avoidance Behavior Using an Enhanced Social Force Model. Sustainability 2026, 18, 746. [Google Scholar] [CrossRef] [Scilit]
- Kumar, A.; Zhou, A.; Tucker, G.; Levine, S. Conservative Q-Learning for offline reinforcement learning. In Advances in Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 1179–1191. [Google Scholar]
- Gronauer, S.; Diepold, K. Multi-agent deep reinforcement learning: A survey. Artif. Intell. Rev. 2022, 55, 895–943. [Google Scholar] [CrossRef] [Scilit]
- Foerster, J.; Nardelli, N.; Farquhar, G.; Afouras, T.; Torr, P.H.S.; Kohli, P.; Whiteson, S. Stabilising experience replay for deep multi-agent reinforcement learning. In Proceedings of the 34th International Conference on Machine Learning; JMLR: Cambridge, MA, USA, 2017; Volume 70, pp. 1146–1155. [Google Scholar]








| Parameter | Value |
|---|---|
| Maximum pretraining iterations | 100 |
| New experiences collected per pretraining iteration (N) | 2000 |
| CQL regularization parameter () | 1.0 |
| Gradient updates per pretraining iteration (K) | 100 |
| Evaluation rollouts per source task () | 100 |
| Source-task weighting parameter () | 0.5 |
| Pretraining termination threshold () | 0.8 |
| Maximum target-task episodes | 2000 |
| Exploration rate () | 0.1 |
| Discount factor () | 0.98 |
| Learning rate | 0.0001 |
| Initial random exploration | 1000 steps |
| Replay-buffer size | 500,000 |
| Minibatch size | 128 |
| Target-network update frequency | 100 steps |
| Source Task | Goal Progress | Wall-Collision Rate |
|---|---|---|
| O-shaped | 0.991 | 0.049 |
| S-shaped | 0.980 | 0.055 |
| U-shaped | 1.016 | 0.070 |
| Multi-pillar | 0.963 | 0.125 |
| Target Task | Transfer Efficacy |
|---|---|
| Intersection | 0.0231 |
| Bi-flow | 0.0198 |
| Bi-door | 0.0285 |
| Bottleneck | 0.0233 |
| Target Task | No Pretraining | Proposed Method | p-Value |
|---|---|---|---|
| Intersection | |||
| Bi-flow | |||
| Bi-door | |||
| Bottleneck |
| Metric | Episodes 1–100 | Episodes 101–200 | ||
|---|---|---|---|---|
| Pretraining | No Pretraining | Pretraining | No Pretraining | |
| Goal-arrival ratio | ||||
| Steps to goal | ||||
| Recorded trajectory length | ||||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hayashida, T.; Sekizaki, S.; Kato, N. Behavioral Analysis of Transfer Learning in DQN-Based Pedestrian Agents Using Single-Agent Pretraining. Mach. Learn. Knowl. Extr. 2026, 8, 288. https://doi.org/10.3390/make8090288
Hayashida T, Sekizaki S, Kato N. Behavioral Analysis of Transfer Learning in DQN-Based Pedestrian Agents Using Single-Agent Pretraining. Machine Learning and Knowledge Extraction. 2026; 8(9):288. https://doi.org/10.3390/make8090288
Chicago/Turabian StyleHayashida, Tomohiro, Shinya Sekizaki, and Natsuki Kato. 2026. "Behavioral Analysis of Transfer Learning in DQN-Based Pedestrian Agents Using Single-Agent Pretraining" Machine Learning and Knowledge Extraction 8, no. 9: 288. https://doi.org/10.3390/make8090288
APA StyleHayashida, T., Sekizaki, S., & Kato, N. (2026). Behavioral Analysis of Transfer Learning in DQN-Based Pedestrian Agents Using Single-Agent Pretraining. Machine Learning and Knowledge Extraction, 8(9), 288. https://doi.org/10.3390/make8090288

