Learning System-Optimal and Individual-Optimal Collision Avoidance Behaviors by Autonomous Mobile Agents
Abstract
1. Introduction
2. Related Work
3. Distributed Stochastic Search Algorithm
3.1. Framework
3.2. DSSA+
| Algorithm 1 DSSA+ |
|
4. Deep Reinforcement Learning
4.1. Q-Learning
4.2. Deep Q Network
5. Approach
5.1. Basic Idea
5.2. Distributed Stochastic Search Algorithm with Deep Q-Network (DSSQ)
5.2.1. Overview of DSSQ
| Algorithm 2 DSSQ |
|
5.2.2. State Space
5.2.3. Action Space
5.2.4. Reward Function
5.2.5. DQN Architecture
6. Experiment
- para2
- As shown in Figure 3a, there are two fully homogeneous agents that have their individual destinations located in front of the other agent. They go in parallel at first, but need to intersect somewhere in the middle of their routes. The reference speed of any agent is set to 25, which is equal to . The coordinates of the origins and destinations of all agents are shown in Table 1. Given these settings, the lower bound of the time step for every agent to reach its destination is 21.1.
- overtake3
- As shown in Figure 3b, there are three agents in a row going forward in the same direction that have their destinations located on the same straight line with an opposite order. The reference speed is 5 for agent 1, 12 for agent 2, and 25 for agent 3. This means the latter agents need to overtake the former agents somewhere in the middle of their courses. The coordinates of the origins and destinations of all agents are shown in Table 1. Given these settings, the average lower bound of the time step for each agent to reach its destination is 46.3.
- face4
- As shown in Figure 3c, there are four agents placed on the vertices of a square, each of which aims to reach the destination behind another agent on its diagonal. Clearly, their shortest paths intersect in the center of the square. The reference speed of any agent is set to 25. The coordinates of the origins and destinations of all agents are shown in Table 1. Given these settings, the lower bound of the time step for every agent to reach its destination is 37.7.
- para4
- As shown in Figure 3d, there are four agents that require similar but more complex crossing behavior than para2. The reference speed of any agent is set to 25. The coordinates of the origins and destinations of all agents are shown in Table 1. Given these settings, the lower bound of the time step for every agent to reach its destination is 22.4.
- cross16
- As shown in Figure 3e, there are 16 agents with all four groups of four agents crossing each other in the middle and heading toward the other sides of their starting points. The reference speed of any agent is set to 25. The coordinates of the origins and destinations of all agents are shown in Table 1. Given these settings, the lower bound of the time step for every agent to reach its destination is 23.0.

| Scenario | Coordinates of Origins and Destinations |
|---|---|
| para2 | agent 1: o[368, 700] d[432, 200], agent 2: o[432, 700] d[368, 200] |
| overtake3 | agent 1: o[300, 500] d[500, 300], agent 2: o[200, 600] d[600, 200], agent 3: o[100, 700] d[700, 100] |
| face4 | agent 1: o[100, 100] d[750, 750], agent 2: o[100, 700] d[750, 50], agent 3: o[700, 100] d[50, 750], agent 4: o[700, 700] d[50, 50] |
| para4 | agent 1: o[368, 700] d[432, 200], agent 2: o[432, 700] d[368, 200], agent 3: o[304, 700] d[496, 200], agent 4: o[496, 700] d[304, 200] |
| cross16 | agent 1: o[150, 250] d[700, 250], agent 2: o[150, 350] d[700, 350], agent 3: o[150, 450] d[700, 450], agent 4: o[150, 550] d[700, 550] |
| agent 5: o[250, 650] d[250, 100], agent 6: o[350, 650] d[350, 100], agent 7: o[450, 650] d[450, 100], agent 8: o[550, 650] d[550, 100] | |
| agent 9: o[650, 250] d[100, 250], agent 10: o[650, 350] d[100, 350], agent 11: o[650, 450] d[100, 450], agent 12: o[650, 550] d[100, 550] | |
| agent 13: o[250, 150] d[250, 700], agent 14: o[350, 150] d[350, 700], agent 15: o[450, 150] d[450, 700], agent 16: o[550, 150] d[550, 700] |
7. Conclusions and Future Work
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Hennes, D.; Claes, D.; Meeussen, W.; Tuyls, K. Multi-Robot Collision Avoidance with Localization Uncertainty. In Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems (AAMAS-2012), Valencia, Spain, 4–8 June 2012; pp. 147–154. [Google Scholar]
- Huang, Y.; Chen, L.; Chen, P.; Negenborn, R.R.; van Gelder, P. Ship collision avoidance methods: State-of-the-art. Saf. Sci. 2020, 121, 451–473. [Google Scholar] [CrossRef] [Scilit]
- Wang, Q.; Phillips, C. Cooperative collision avoidance for multi-vehicle systems using reinforcement learning. In Proceedings of the 2013 18th International Conference on Methods Models in Automation Robotics (MMAR), Miedzyzdroje, Poland, 26–29 August 2013; pp. 98–102. [Google Scholar]
- Chung, S.J.; Paranjape, A.A.; Dames, P.; Shen, S.; Kumar, V. A Survey on Aerial Swarm Robotics. IEEE Trans. Robot. 2018, 34, 837–855. [Google Scholar] [CrossRef] [Scilit]
- Alonso-Mora, J.; Breitenmoser, A.; Rufli, M.; Beardsley, P.; Siegwart, R. Optimal reciprocal collision avoidance for multiple non-holonomic robots. In Distributed Autonomous Robotic Systems: The 10th International Symposium; Springer: Berlin/Heidelberg, Germany, 2013; pp. 203–216. [Google Scholar]
- Kuderer, M.; Kretzschmar, H.; Sprunk, C.; Burgard, W. Feature-Based Prediction of Trajectories for Socially Compliant Navigation. In Robotics: Science and Systems VIII; The MIT Press: Cambridge, MA, USA, 2013; pp. 193–200. [Google Scholar]
- Phillips, M.; Likhachev, M. SIPP: Safe interval path planning for dynamic environments. In Proceedings of the 2011 IEEE International Conference on Robotics and Automation (ICRA-2011), Shanghai, China, 9–13 May 2011; pp. 5628–5635. [Google Scholar]
- van den Berg, J.; Guy, S.J.; Lin, M.; Manocha, D. Reciprocal n-Body Collision Avoidance. In Robotics Research; Pradalier, C., Siegwart, R., Hirzinger, G., Eds.; Springer: Berlin/Heidelberg, Germany, 2011; pp. 3–19. [Google Scholar]
- Chen, Y.F.; Liu, M.; Everett, M.; How, J.P. Decentralized non-communicating multiagent collision avoidance with deep reinforcement learning. In Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA-2017), Singapore, 29 May–3 June 2017; pp. 285–292. [Google Scholar]
- Everett, M.; Chen, Y.F.; How, J.P. Motion planning among dynamic, decision-making agents with deep reinforcement learning. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS-2018), Madrid, Spain, 1–5 October 2018; pp. 3052–3059. [Google Scholar]
- Fan, T.; Long, P.; Liu, W.; Pan, J. Distributed multi-robot collision avoidance via deep reinforcement learning for navigation in complex scenarios. Int. J. Robot. Res. 2020, 39, 856–892. [Google Scholar] [CrossRef] [Scilit]
- Felner, A.; Stern, R.; Shimony, S.E.; Boyarski, E.; Goldenberg, M.; Sharon, G.; Sturtevant, N.; Wagner, G.; Surynek, P. Search-Based Optimal Solvers for the Multi-Agent Pathfinding Problem: Summary and Challenges. In Proceedings of the Tenth International Symposium on Combinatorial Search (SoCS-2017), Pittsburgh, PA, USA, 16–17 June 2017; pp. 29–37. [Google Scholar]
- Tang, S.; Thomas, J.; Kumar, V. Hold or take Optimal Plan (HOOP): A quadratic programming approach to multi-robot trajectory generation. Int. J. Robot. Res. 2018, 37, 1062–1084. [Google Scholar] [CrossRef] [Scilit]
- Kim, D.G.; Hirayama, K.; Park, G.K. Collision Avoidance in Multiple-Ship Situations by Distributed Local Search. J. Adv. Comput. Intell. Intell. Informatics 2014, 18, 839–848. [Google Scholar] [CrossRef] [Scilit]
- Kim, D.; Hirayama, K.; Okimoto, T. Ship Collision Avoidance by Distributed Tabu Search. TransNav Int. J. Mar. Navig. Saf. Sea Transp. 2015, 9, 23–29. [Google Scholar] [CrossRef] [Scilit]
- Kim, D.; Hirayama, K.; Okimoto, T. Distributed Stochastic Search Algorithm for Multi-ship Encounter Situations. J. Navig. 2017, 70, 699–718. [Google Scholar] [CrossRef] [Scilit]
- Zheng, H.; Negenborn, R.R.; Lodewijks, G. Fast ADMM for Distributed Model Predictive Control of Cooperative Waterborne AGVs. IEEE Trans. Control Syst. Technol. 2017, 25, 1406–1413. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.; Hopman, H.; Negenborn, R.R. Distributed model predictive control for vessel train formations of cooperative multi-vessel systems. Transp. Res. Part C Emerg. Technol. 2018, 92, 101–118. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.; Negenborn, R.R.; Hopman, H. Intersection Crossing of Cooperative Multi-vessel Systems. IFAC-PapersOnLine 2018, 51, 379–385. [Google Scholar] [CrossRef] [Scilit]
- Ferranti, L.; Negenborn, R.R.; Keviczky, T.; Alonso-Mora, J. Coordination of Multiple Vessels Via Distributed Nonlinear Model Predictive Control. In Proceedings of the 2018 European Control Conference (ECC), Limassol, Cyprus, 12–15 June 2018; pp. 2523–2528. [Google Scholar]
- Hirayama, K.; Miyake, K.; Shiota, T.; Okimoto, T. DSSA+: Distributed Collision Avoidance Algorithm in an Environment where Both Course and Speed Changes are Allowed. TransNav Int. J. Mar. Navig. Saf. Sea Transp. 2019, 13, 117–124. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Liu, J.; Negenborn, R.R. Distributed coordination for collision avoidance of multiple ships considering ship maneuverability. Ocean Eng. 2019, 181, 212–226. [Google Scholar] [CrossRef] [Scilit]
- Akdağ, M.; Fossen, T.I.; Johansen, T.A. Collaborative Collision Avoidance for Autonomous Ships Using Informed Scenario-Based Model Predictive Control. IFAC-PapersOnLine 2022, 55, 249–256. [Google Scholar] [CrossRef] [Scilit]
- Tran, H.A.; Johansen, T.A.; Negenborn, R.R. Parallel distributed collision avoidance with intention consensus based on ADMM. IFAC-PapersOnLine 2024, 58, 302–309. [Google Scholar] [CrossRef] [Scilit]
- Zhang, W.; Wang, G.; Xing, Z.; Wittenburg, L. Distributed stochastic search and distributed breakout: Properties, comparison and applications to constraint optimization problems in sensor networks. Artif. Intell. 2005, 161, 55–87. [Google Scholar] [CrossRef] [Scilit]
- Boyd, S.; Parikh, N.; Chu, E.; Peleato, B.; Eckstein, J. Distributed Optimization and Statistical Learning via the Alternating Direction Method of Multipliers. Found. Trends®Mach. Learn. 2011, 3, 1–122. [Google Scholar]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- CBSMornings. Synchronized Walking Becomes Staple at Japanese University. 2021. Available online: https://www.youtube.com/watch?v=uDgEQGsh7Qs (accessed on 22 August 2025).
- Shen, H.; Hashimoto, H.; Matsuda, A.; Taniguchi, Y.; Terada, D.; Guo, C. Automatic collision avoidance of multiple ships based on deep Q-learning. Appl. Ocean Res. 2019, 86, 268–288. [Google Scholar] [CrossRef] [Scilit]
- Godoy, J.E.; Karamouzas, I.; Guy, S.J.; Gini, M. Implicit coordination in crowded multi-agent navigation. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (AAAI-2016), Phoenix, AZ, USA, 12–17 February 2016; pp. 2487–2493. [Google Scholar]
- Purwin, O.; D’Andrea, R.; Lee, J.W. Theory and implementation of path planning by negotiation for decentralized agents. Robot. Auton. Syst. 2008, 56, 422–436. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Y.; Li, S.E.; Li, K.; Borrelli, F.; Hedrick, J.K. Distributed Model Predictive Control for Heterogeneous Vehicle Platoons Under Unidirectional Topologies. IEEE Trans. Control Syst. Technol. 2017, 25, 899–910. [Google Scholar] [CrossRef] [Scilit]
- Wang, P.; Deng, H.; Zhang, J.; Wang, L.; Zhang, M.; Li, Y. Model Predictive Control for Connected Vehicle Platoon Under Switching Communication Topology. IEEE Trans. Intell. Transp. Syst. 2022, 23, 7817–7830. [Google Scholar] [CrossRef] [Scilit]
- Qiang, Z.; Dai, L.; Chen, B.; Xia, Y. Distributed Model Predictive Control for Heterogeneous Vehicle Platoon With Inter-Vehicular Spacing Constraints. IEEE Trans. Intell. Transp. Syst. 2023, 24, 3339–3351. [Google Scholar] [CrossRef] [Scilit]
- Du, G.; Zou, Y.; Zhang, X.; Fan, J.; Sun, W.; Li, Z. Efficient Motion Control for Heterogeneous Autonomous Vehicle Platoon Using Multilayer Predictive Control Framework. IEEE Internet Things J. 2024, 11, 38273–38290. [Google Scholar] [CrossRef] [Scilit]
- Hirayama, K.; Yokoo, M. Distributed Partial Constraint Satisfaction Problem. In Proceedings of the Third International Conference on Principles and Practice of Constraint Programming (CP-1997), Linz, Austria, 29 October–1 November 1997; pp. 222–236. [Google Scholar]
- Petcu, A.; Faltings, B. A Scalable Method for Multiagent Constraint Optimization. In Proceedings of the 19th International Joint Conference on Artificial Intelligence (IJCAI-2005), Edinburgh, UK, 30 July–5 August 2005; pp. 266–271. [Google Scholar]
- Gershman, A.; Meisels, A.; Zivan, R. Asynchronous Forward Bounding for Distributed COPs. J. Artif. Intell. Res. 2009, 34, 61–88. [Google Scholar] [CrossRef] [Scilit]
- Pertzovskiy, A.; Zivan, R.; Agmon, N. Collision Avoiding Max-Sum for Mobile Sensor Teams. J. Artif. Intell. Res. 2024, 79, 1281–1311. [Google Scholar] [CrossRef] [Scilit]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; The MIT Press: Cambridge, MA, USA, 2018. [Google Scholar]
- Watkins, C.J.; Dayan, P. Technical Note: Q-Learning. Mach. Learn. 1992, 8, 279–292. [Google Scholar] [CrossRef] [Scilit]
- Van Hasselt, H.; Guez, A.; Silver, D. Deep reinforcement learning with double q-learning. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (AAAI-2016), Phoenix, AZ, USA, 12–17 February 2016; pp. 2094–2100. [Google Scholar]
- Schaul, T.; Quan, J.; Antonoglou, I.; Silver, D. Prioritized experience replay. arXiv 2015, arXiv:1511.05952. [Google Scholar]
- De Asis, K.; Hernandez-Garcia, J.F.; Holland, G.Z.; Sutton, R.S. Multi-step reinforcement learning: A unifying algorithm. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence (AAAI-2018), New Orleans, LA, USA, 2–7 February 2018; pp. 2902–2909. [Google Scholar]
- Wang, Z.; Schaul, T.; Hessel, M.; Hasselt, H.; Lanctot, M.; Freitas, N. Dueling network architectures for deep reinforcement learning. In Proceedings of the 33nd International Conference on Machine Learning (ICML-2016), New York, NY, USA, 19–24 June 2016; pp. 1995–2003. [Google Scholar]
- Hessel, M.; Modayil, J.; Van Hasselt, H.; Schaul, T.; Ostrovski, G.; Dabney, W.; Horgan, D.; Piot, B.; Azar, M.; Silver, D. Rainbow: Combining improvements in deep reinforcement learning. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence (AAAI-2018), New Orleans, LA, USA, 2–7 February 2018; pp. 3215–3222. [Google Scholar]
- Fan, J.; Zhang, X.; Zheng, K.; Zou, Y.; Zhou, N. Hierarchical path planner combining probabilistic roadmap and deep deterministic policy gradient for unmanned ground vehicles with non-holonomic constraints. J. Frankl. Inst. 2024, 361, 106821. [Google Scholar] [CrossRef] [Scilit]
- Wen, M.; Kuba, J.; Lin, R.; Zhang, W.; Wen, Y.; Wang, J.; Yang, Y. Multi-agent reinforcement learning is a sequence modeling problem. Adv. Neural Inf. Process. Syst. 2022, 35, 16509–16521. [Google Scholar]



| Para2 | Overtake3 | Face4 | Para4 | Cross16 | |
|---|---|---|---|---|---|
| 25.0 | 51.4 | 38.1 | 39.1 | 27.2 | |
| 22.8 | 59.3 | 39.6 | 25.0 | 33.2 | |
| 25.0 | 50.0 | 38.2 | 34.3 | 26.2 | |
| DSSQ | 22.1 | 50.2 | 38.7 | 25.2 | 26.3 |
| Lower bound | 21.1 | 46.3 | 37.7 | 22.4 | 23.0 |
| Para2 | Overtake3 | Face4 | Para4 | Cross16 | |
|---|---|---|---|---|---|
| 20.33 | 141.56 | 0.19 | 6704.19 | 487.43 | |
| 1.00 | 1170.67 | 4.25 | 4.69 | 437.18 | |
| 20.25 | 169.56 | 0.19 | 1529.69 | 280.73 | |
| DSSQ | 1.00 | 169.56 | 3.00 | 5.69 | 40.06 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Hirayama, K.; Gohara, K.; Koue, J.; Okimoto, T.; Kim, D. Learning System-Optimal and Individual-Optimal Collision Avoidance Behaviors by Autonomous Mobile Agents. Algorithms 2025, 18, 671. https://doi.org/10.3390/a18110671
Hirayama K, Gohara K, Koue J, Okimoto T, Kim D. Learning System-Optimal and Individual-Optimal Collision Avoidance Behaviors by Autonomous Mobile Agents. Algorithms. 2025; 18(11):671. https://doi.org/10.3390/a18110671
Chicago/Turabian StyleHirayama, Katsutoshi, Kazuma Gohara, Jinichi Koue, Tenda Okimoto, and Donggyun Kim. 2025. "Learning System-Optimal and Individual-Optimal Collision Avoidance Behaviors by Autonomous Mobile Agents" Algorithms 18, no. 11: 671. https://doi.org/10.3390/a18110671
APA StyleHirayama, K., Gohara, K., Koue, J., Okimoto, T., & Kim, D. (2025). Learning System-Optimal and Individual-Optimal Collision Avoidance Behaviors by Autonomous Mobile Agents. Algorithms, 18(11), 671. https://doi.org/10.3390/a18110671

