TAP-DDQN: Multiplicative Potential-Based Reward Shaping Framework for Tactical Decision-Making of Unmanned Surface Vehicles in Adversarial Maritime Engagements
Abstract
1. Introduction
- Multiplicative Tactical Approach Potential (TAP): We formulate a multiplicative TAP that couples engagement-distance regulation and aspect-aware tactical positioning within a potential-based reward-shaping framework.
- Explainable Reward Decomposition: We introduce an explainable four-component reward decomposition consisting of spatio-temporal approach, operational constraints, weapon-employment reward, and tactical potential, enabling direct interpretation and component-wise ablation.
- GRU-enhanced Dueling DQN with Dynamic Target Curriculum: We combine a dueling deep Q-network with GRU memory and a dynamic target curriculum to improve temporal reasoning, sample efficiency, and coverage across short and long-range engagement regimes.
- TRNLE-Based Maritime Simulator Integration and USV Tactical Validation: We integrate the proposed method with a legacy maritime combat simulator through the TRNLE wrapper and validate its effectiveness for discrete USV tactical decision-making, with naval training-system adversary generation demonstrated as a dual-use application.
2. Research Background
2.1. Unmanned Surface Vehicles and Maritime Tactical Decision-Making
2.2. On-Board Training System for Naval Combat System
2.3. Deep Q-Network and Dueling Architecture
2.4. Maneuver Generation
2.5. Potential-Based Reward Shaping
2.6. Maritime Regulatory Landscapes and Dual-Use Maneuvering Constraints
- (i)
- COLREGs Rule 13 (Overtaking) for Compliant Escort: In standard security shadowing or escort operations, the USV behaves as the “give-way” vessel. It is legally restricted from crossing the bow of the target ship. Instead, it must approach and hold a stable position abaft the target’s beam (relative bearing angle to , representing the stern light sector). Our framework translates this regulatory obligation into an aspect-aware angular constraint, ensuring commercial safety acceptability.
- (ii)
- COLREGs Rule 2 (General Prudential Rule) for Emergency Evasion: When a non-cooperative intruder exhibits erratic maneuvers or deploys dynamic high-speed hazards (modeled as incoming projectile-like threats in the simulator), passive adherence to standard stand-on rules is superseded. Under Rule 2, mariners are mandated to depart from standard rules to avoid immediate danger. The proposed framework rigorously implements Rule 2, allowing the USV to trigger non-linear, highly agile evasive maneuvers to preserve survivability under extreme hazard circumstances.
3. TAP-DDQN: Dueling DQN with a Multiplicative Tactical Potential Field
3.1. Tactical Maneuver Generation in Naval Engagements
3.2. Scenario and Tactical Geometry
3.3. Observation and Action
3.4. Reward Design with a Multiplicative Tactical Approach Potential
3.4.1. Spatio-Temporal Reward (STR)
3.4.2. Constraint Reward
3.4.3. Weapon-Firing Reward
3.4.4. Tactical Approach Potential (TAP)
| Algorithm 1 TAP-DDQN training with Tactical Approach Potential shaping and Dynamic Target Curriculum |
|
3.5. Curriculum Learning Strategy via Dynamic Target Switching
4. Learning Framework for Intelligent Simulation Target Objects
TRNLE Framework
5. Experiment and Results
5.1. Experimental Setup and Ablation Design
- (A)
- Rule-based CGF (Baseline): deterministic rules for maneuvering and firing; lower-bound reference representing the conventional, non-adaptive adversary behavior found in existing OBTS.
- (B)
- DQN+MLP w/o TAP: standard DQN with an MLP encoder and STR + constraint + weapon rewards, without TAP shaping or Dueling decomposition.
- (C)
- PPO+MLP w/o TAP: policy-gradient counterpart of (B), evaluating whether on-policy methods offer advantages in this discrete-action tactical domain.
- (D)
- DDQN+MLP w/o TAP: Dueling DQN with MLP encoder and STR but without TAP shaping.
- (E)
- DDQN+MLP w/ TAP (Additive): Variant (D) augmented with the additive TAP shaping. This variant serves as a controlled internal additive surrogate for additive tactical-shaping families [47,48]: it shares identical distance and aspect potentials with the multiplicative variants and differs only in the coupling structure, reproducing the range-independent angular gradient characteristic of additive schemes.
- (F)
- DDQN+MLP w/ TAP (Multiplicative) and w/ DTC: Variant (D) augmented with the multiplicative TAP shaping (Equation (24)) and DTC.
- (G)
- DDQN+RNN w/ TAP (Multiplicative) and w/ DTC (Proposed): Full framework combining GRU-augmented Dueling DQN with multiplicative TAP and DTC.
5.2. Training Efficiency
5.3. Mission-Level Quantitative Performance
- Distance Goal Achievement Rate (DGAR). The fraction of steps within an evaluation episode during which the agent’s range to the ownship lies inside the engagement band, averaged across evaluation episodes:where is the step count of episode e, is the agent–ownship range in episode e at step t, is the target distance, and is the indicator function.
- Behavioral Entropy (BE). The Shannon entropy of the empirical action distribution over the discrete action set :with the empirical action probability over an episode of length T. BE is bounded by : values close to the upper bound indicate near-uniform action use, while values close to 0 indicate a degenerate, near-deterministic policy.
- Mean Tactical Potential (). The time-averaged TAP score (Equation (24)) over an episode:Because is multiplicatively coupled, rewards sustained occupancy of the engagement band together with a favorable aspect; a brief instant of either condition alone does not raise appreciably.
- Jerk Cost (). The root-mean-square (RMS) of the per-step change in the agent’s longitudinal acceleration computed over the episode:where is the scalar agent acceleration at step t. By penalizing abrupt acceleration changes, serves as a smoothness surrogate: a high combined with a low shows that the tactical advantage is achieved with executable maneuvers rather than high-frequency control chatter infeasible on a real USV.
Statistical Significance of the Ablation Contrasts
5.4. Reward-Aligned Verification of the Multiplicative TAP Coupling
5.5. Qualitative Analysis of Agent Behavior
6. Discussion
6.1. Mechanism of the Multiplicative TAP Coupling
6.2. Role of Recurrent Temporal Memory
6.3. Implications for USV Tactical Autonomy and Naval Training System
6.4. Limitations
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| BE | Behavioral Entropy |
| CGF | Computer Generated Force |
| CMS | Combat Management System |
| DDQN | Dueling Deep Q-Network |
| DGAR | Distance Goal Achievement Rate |
| DNN | Deep Neural Networks |
| DQN | Deep Q-Network |
| DRL | Deep Reinforcement Learning |
| eGAR | Effective Goal Achievement Rate |
| GRU | Gated Recurrent Unit |
| LVC | Live, Virtual, Constructive |
| MASS | Maritime Autonomous Surface Ships |
| MLP | Multi-Layer Perceptron |
| OBTS | On-Board Training System |
| PBRS | Potential-Based Reward Shaping |
| PPO | Proximal Policy Optimization |
| RL | Reinforcement Learning |
| RNN | Recurrent Neural Network |
| STR | Spatio-Temporal Reward |
| TAP | Tactical Approach Potential |
| TD | Temporal-Difference |
| TJAR | Tactical Joint Achievement Rate |
| TMG | Tactical Maneuver Generation |
| TRNLE | TRaiNing Learning Environment |
| USV | Unmanned Surface Vehicle |
References
- Wang, H.; Zhou, Z.; Jiang, J.; Deng, W.; Chen, X. Autonomous Air Combat Maneuver Decision-Making Based on PPO-BWDA. IEEE Access 2024, 12, 119116–119132. [Google Scholar] [CrossRef] [Scilit]
- Wang, N.; Li, Z.; Liang, X.; Hou, Y.; Yang, A. A review of deep reinforcement learning methods and military application research. Math. Probl. Eng. 2023, 2023, 7678382. [Google Scholar] [CrossRef] [Scilit]
- Brockman, G. OpenAI Gym. arXiv 2016, arXiv:1606.01540. [Google Scholar]
- Kurach, K.; Raichuk, A.; Stańczyk, P.; Zając, M.; Bachem, O.; Espeholt, L.; Riquelme, C.; Vincent, D.; Michalski, M.; Bousquet, O.; et al. Google Research Football: A Novel Reinforcement Learning Environment. arXiv 2020, arXiv:1907.11180. [Google Scholar]
- Samvelyan, M.; Rashid, T.; De Witt, C.S.; Farquhar, G.; Nardelli, N.; Rudner, T.G.; Hung, C.M.; Torr, P.H.; Foerster, J.; Whiteson, S. The starcraft multi-agent challenge. arXiv 2019, arXiv:1902.04043. [Google Scholar]
- Bingham, B.; Agüero, C.; McCarrin, M.; Klamo, J.; Malia, J.; Allen, K.; Lum, T.; Rawson, M.; Waqar, R. Toward Maritime Robotic Simulation in Gazebo. In Proceedings of the MTS/IEEE OCEANS Conference, Seattle, WA, USA, 27–31 October 2019. [Google Scholar] [CrossRef] [Scilit]
- Paravisi, M.; Santos, D.H.; Jorge, V.; Heck, G.; Gonçalves, L.M.; Amory, A. Unmanned Surface Vehicle Simulator with Realistic Environmental Disturbances. Sensors 2019, 19, 1068. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bae, J.H.; Jung, H.; Kim, S.; Kim, S.; Kim, Y.D. Deep reinforcement learning-based air-to-air combat maneuver generation in a realistic environment. IEEE Access 2023, 11, 26427–26440. [Google Scholar] [CrossRef] [Scilit]
- Hu, D.; Yang, R.; Zuo, J.; Zhang, Z.; Wu, J.; Wang, Y. Application of deep reinforcement learning in maneuver planning of beyond-visual-range air combat. IEEE Access 2021, 9, 32282–32297. [Google Scholar] [CrossRef] [Scilit]
- Piao, H.Y.; Yang, S.; Chen, H.; Li, J.; Yu, J.; Peng, X.; Yang, X.; Yang, Z.; Sun, Z.; Chang, Y. Discovering Expert-Level Air Combat Knowledge via Deep Excitatory-Inhibitory Factorized Reinforcement Learning. ACM Trans. Intell. Syst. Technol. 2024, 15, 1–28. [Google Scholar] [CrossRef] [Scilit]
- Roessingh, J.J.; Toubman, A.; van Oijen, J.; Poppinga, G.; Hou, M.; Luotsinen, L. Machine learning techniques for autonomous agents in military simulations—Multum in Parvo. In Proceedings of the 2017 IEEE International Conference on Systems, Man, and Cybernetics (SMC); IEEE: Piscataway, NJ, USA, 2017; pp. 3445–3450. [Google Scholar] [CrossRef] [Scilit]
- Boron, J.; Darken, C. Developing combat behavior through reinforcement learning in wargames and simulations. In Proceedings of the 2020 IEEE Conference on Games (CoG); IEEE: Piscataway, NJ, USA, 2020; pp. 728–731. [Google Scholar] [CrossRef] [Scilit]
- Dimitriu, A.; Michaletzky, T.V.; Remeli, V.; Tihanyi, V.R. A Reinforcement Learning Approach to Military Simulations in Command: Modern Operations. IEEE Access 2024, 12, 77501–77513. [Google Scholar] [CrossRef] [Scilit]
- Narvekar, S.; Peng, B.; Leonetti, M.; Sinapov, J.; Taylor, M.E.; Stone, P. Curriculum learning for reinforcement learning domains: A framework and survey. J. Mach. Learn. Res. 2020, 21, 1–50. [Google Scholar] [CrossRef] [Scilit]
- Hutsebaut-Buysse, M.; Mets, K.; Latré, S. Hierarchical reinforcement learning: A survey and open research challenges. Mach. Learn. Knowl. Extr. 2022, 4, 172–221. [Google Scholar] [CrossRef] [Scilit]
- Ma, H.; Vo, T.V.; Leong, T.Y. Mixed-Initiative Bayesian Sub-Goal Optimization in Hierarchical RL. In Proceedings of the 23rd International Conference on Autonomous Agents and MultiAgent Systems (AAMAS); IFAAMAS: Richland, SC, USA, 2024; pp. 1328–1336. [Google Scholar]
- Ng, A.Y.; Harada, D.; Russell, S. Policy invariance under reward transformations: Theory and application to reward shaping. Proc. ICML 1999, 99, 278–287. [Google Scholar]
- Cooley, T.; Oswalt, I.; Oxford, M.; Park, L. Operationalizing Artificial Intelligence in Simulation Based Training. In Proceedings of the 2021 Interservice/Industry Training, Simulation, and Education Conference (I/ITSEC), Orlando, FL, USA, 29 November–3 December 2021. [Google Scholar]
- Black, S.; Darken, C. Scaling Intelligent Agents in Combat Simulations for Wargaming. arXiv 2024, arXiv:2402.06694. [Google Scholar]
- Qiao, Y.; Yin, J.; Wang, W.; Duarte, F.; Yang, J.; Ratti, C. Survey of Deep Learning for Autonomous Surface Vehicles in Marine Environments. IEEE Trans. Intell. Transp. Syst. 2023, 24, 3678–3701. [Google Scholar] [CrossRef] [Scilit]
- Hong, L.; Liu, L.; Peng, Z.; Zhang, F. Control of Marine Robots in the Era of Data-Driven Intelligence. Annu. Rev. Control. Robot. Auton. Syst. 2025, 9, 243–272. [Google Scholar] [CrossRef] [Scilit]
- International Maritime Organization. Interim Guidelines for MASS Trials; Msc.1/circ.1604; Adopted at MSC 101; International Maritime Organization (IMO): London, UK, 2019. [Google Scholar]
- Wu, C.; Yu, W.; Li, G.; Liao, W. Deep reinforcement learning with dynamic window approach based collision avoidance path planning for maritime autonomous surfaces. Ocean Eng. 2023, 284, 115208. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Zhang, X.; Yang, Z.; Bashir, M.; Lee, K. Collision Avoidance for Autonomous Ship Using Deep Reinforcement Learning and Prior-Knowledge-Based Approximate Representation. Front. Mar. Sci. 2023, 9, 1084763. [Google Scholar] [CrossRef] [Scilit]
- Tam, C.; Bucknall, R.; Greig, A. Review of Collision Avoidance and Path Planning Methods for Ships in Close Range Encounters. J. Navig. 2009, 62, 455–476. [Google Scholar] [CrossRef] [Scilit]
- Jung, Y.R. A study on multi sensor track fusion algorithm for naval combat system. J. Korea Inst. Mil. Sci. Technol. 2007, 10, 34–42. [Google Scholar]
- Go, Y. A Study on Bottom-Up Update of TPR-Tree for Target Indexing in Naval Combat Systems. J. Korea Inst. Mil. Sci. Technol. 2019, 22, 266–277. [Google Scholar]
- Navy, U. Navy’s Newest Combat Simulator Trains Its First Ships; Technical Report; United States Navy: Washington, DC, USA, 2020. [Google Scholar]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, USA, 2018. [Google Scholar]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.; Graves, A.; Riedmiller, M.; Fidjeland, A.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, Z.; Schaul, T.; Hessel, M.; van Hasselt, H.; Lanctot, M.; de Freitas, N. Dueling Network Architectures for Deep Reinforcement Learning. In Proceedings of the 33rd International Conference on Machine Learning (ICML), New York, NY, USA, 19–24 June 2016; pp. 1995–2003. [Google Scholar]
- Cho, K.; van Merriënboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. arXiv 2014, arXiv:1406.1078. [Google Scholar]
- Coble, J.; Barton, A.; Darken, C.; Black, S. Optimizing Naval Movement Using Deep Reinforcement Learning. In Proceedings of the 2023 International Conference on Machine Learning and Applications (ICMLA); IEEE: Piscataway, NJ, USA, 2023; pp. 400–407. [Google Scholar]
- Markowitz, J.; Sheffield, R.; Mullins, G. Maritime Platform Defense with Deep Reinforcement Learning. In Proceedings of the Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications IV; SPIE: Bellingham, WA, USA, 2022; Volume 12113, pp. 423–429. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Zhou, H.; Wei, Y.; Huang, C. Autonomous maneuver decision-making method based on reinforcement learning and Monte Carlo tree search. Front. Neurorobot. 2022, 16, 996412. [Google Scholar] [CrossRef] [Scilit]
- Mei, J.; Li, G.; Huang, H. Deep Reinforcement-Learning-Based Air-Combat-Maneuver Generation Framework. Mathematics 2024, 12, 3020. [Google Scholar] [CrossRef] [Scilit]
- Bae, J.H.; Kang, Y.; Yoon, S.; Kim, Y.; Kim, S. Aircraft Reinforcement Learning using Curriculum Learning. J. Korean Inst. Inf. Sci. Eng. (KIISE) 2021, 48, 707–712. [Google Scholar] [CrossRef] [Scilit]
- Zhu, J.; Kuang, M.; Zhou, W.; Shi, H.; Zhu, J.; Han, X. Mastering air combat game with deep reinforcement learning. Def. Technol. 2024, 34, 295–312. [Google Scholar] [CrossRef] [Scilit]
- Gorton, P.R.; Strand, A.; Brathen, K. A survey of air combat behavior modeling using machine learning. arXiv 2024, arXiv:2404.13954. [Google Scholar]
- Bartsiokas, I.A.; Ntakolia, C.; Avdikos, G.; Lyridis, D.V. Intelligent Multi-Objective Path Planning for Unmanned Surface Vehicles via Deep and Fuzzy Reinforcement Learning. J. Mar. Sci. Eng. 2025, 13, 2285. [Google Scholar] [CrossRef] [Scilit]
- Ntakolia, C.; Lyridis, D.V. A Swarm Intelligence Graph-Based Pathfinding Algorithm Based on Fuzzy Logic (SIGPAF): A Case Study on Unmanned Surface Vehicle Multi-Objective Path Planning. J. Mar. Sci. Eng. 2021, 9, 1243. [Google Scholar] [CrossRef] [Scilit]
- Ntakolia, C.; Lyridis, D.V. A Comparative Study on Ant Colony Optimization Algorithm Approaches for Solving Multi-Objective Path Planning Problems in Case of Unmanned Surface Vehicles. Ocean Eng. 2022, 255, 111418. [Google Scholar] [CrossRef] [Scilit]
- Wiewiora, E.; Cottrell, G.; Elkan, C. Principled methods for advising reinforcement learning agents. In Proceedings of the 20th International Conference on Machine Learning (ICML-03), Washington, DC, USA, 12–24 August 2003. [Google Scholar]
- Gao, Y.; Toni, F. Potential-based reward shaping for hierarchical reinforcement learning. In Proceedings of the IJCAI, Buenos Aires, Argentina, 25–31 July 2015. [Google Scholar]
- Asmuth, J.; Littman, M.L.; Zinkov, R. Potential-based shaping in model-based reinforcement learning. In Proceedings of the AAAI, Chicago, IL, USA, 13–17 July 2008. [Google Scholar]
- Devlin, S.; Kudenko, D. Dynamic Potential-Based Reward Shaping. In Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems (AAMAS), Valencia, Spain, 4–8 June 2012; pp. 433–440. [Google Scholar]
- Zhang, P.; Wang, X.; Wang, Y.; Ma, Z.; Lu, J. Combat game strategy for unmanned surface vessels based on situational field PPO (SF-PPO). Ocean Eng. 2025, 316, 119853. [Google Scholar] [CrossRef] [Scilit]
- Yang, K.; Dong, W.; Cai, M.; Jia, S.; Liu, R. UCAV Air Combat Maneuver Decisions Based on a Proximal Policy Optimization Algorithm with Situation Reward Shaping. Electronics 2022, 11, 2602. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Yuan, Y.; Cheng, Y.; Hua, L. Predictive Air Combat Decision Model with Segmented Reward Allocation. Complex Intell. Syst. 2024, 10, 7513–7530. [Google Scholar] [CrossRef] [Scilit]
- Kun, Y.; Ao, S.; Nengwei, X.; Fang, D.; Maobin, L.; Chen, C. A review of reinforcement learning approaches for pursuit-evasion games. Chin. J. Aeronaut. 2025, 39, 103940. [Google Scholar] [CrossRef] [Scilit]
- Wang, N.; Hou, Y.; Qiu, C.; You, Z. Underactuated Navigation Actor-Critic Deep Reinforcement Learning Framework for Holistic Path Planning of Uncrewed Surface Vehicles. IEEE Trans. Intell. Transp. Syst. 2025, 26, 21245–21256. [Google Scholar] [CrossRef] [Scilit]
- Shin, H.; Ahn, J.; Ahn, J.H.; Kim, M.; Kim, J.; Kim, S. Constructing a Reinforcement Learning Environment Based on a Legacy Simulator for Training a Prototype Intelligent Simulation Agent for Surface Vessels. J. Korean Inst. Intell. Syst. 2025, 35, 163–170. [Google Scholar] [CrossRef] [Scilit]
- International Maritime Organization. Convention on the International Regulations for Preventing Collisions at Sea, 1972 (COLREGs); International Maritime Organization: London, UK, 2003. [Google Scholar]
- Mou, J.; Shi, B.; Wang, B.; Yu, C.; Wang, Y.; Zhong, F.; Zheng, L.; Wang, J.; Li, J. A novel reinforcement learning framework-based path planning algorithm for unmanned surface vehicle. Front. Mar. Sci. 2025, 12, 1641093. [Google Scholar] [CrossRef] [Scilit]














| Category | Observation | Description | Dim |
|---|---|---|---|
| Agent | Agent health (durability) | 1 | |
| Agent course angle | 1 | ||
| Agent speed | 1 | ||
| Incoming-projectile presence | 2 | ||
| Distance error to ownship | 1 | ||
| Closing speed | 1 | ||
| Closing acceleration | 1 | ||
| Target-distance direction | 2 | ||
| Remaining weapons | 1 | ||
| Launch availability/firable step | 2 | ||
| Ownship | Relative position X (ownship − agent) | 1 | |
| Relative position Y (ownship − agent) | 1 | ||
| Ownship course | 1 | ||
| Ownship velocity | 1 | ||
| Agent–ownship distance | 1 | ||
| Bearing from ownship to agent | 1 | ||
| Relative bearing | 1 | ||
| Bearing-quadrant indicator (one-hot) | 2 | ||
| Weapon | Orthogonal distance to incoming projectile | 1 | |
| First derivative of orthogonal distance | 1 | ||
| Second derivative of orthogonal distance | 1 | ||
| History | Previous action | 6 | |
| Total | 31 |
| Action Types | Description |
|---|---|
| No operation | Maintain current maneuvering |
| Acceleration | Increase speed by |
| Deceleration | Decrease speed by |
| Turn left | Decrease course by |
| Turn right | Increase course by |
| Weapon firing | Fire a weapon at the ownship |
| Component | Symbol | Weight | Range | Primary Role |
|---|---|---|---|---|
| Spatio-Temporal | Smooth distance approach | |||
| Constraint | Speed & boundary safety | |||
| Weapon-Firing | – | Timely weapon use | ||
| Tactical Potential | Aspect & range advantage |
| Category | Ownship | Agent |
|---|---|---|
| Starting point | Random area (within bounds) | Within detection area |
| Initial course | Random | Random |
| Initial speed | 9.7 knots (18 km/h) | Random knots (≈ km/h) |
| Radar detection range () | 8 km | 8 km |
| Weapon range () | 5 km | 5 km |
| Weapon | Yes | Yes |
| Action | Course, Velocity, Weapon Firing | Course, Velocity, Weapon Firing |
| Module | Class Name | Description |
|---|---|---|
| Environment | TRNLE | Data generation and object control for training. |
| Com-TRNLE | Data exchange and conversion between TRNLE and Client Agent across platforms. | |
| Client Agent | Agent Mng | Management of training objects by platform. |
| Preprocess | Observation data pre-processing. | |
| Reward | Compute the reward function of Equation (17). | |
| Com-Client | Send and receive data. | |
| Sensor Mng | Sensor attribute data of the simulation target. | |
| Weapon Mng | Weapon attribute data of the simulation target. | |
| Server Agent | Trainer | Manage and control server-agent training settings. |
| Com-Server | Send and receive data. | |
| Algorithm | Management of RL algorithm configuration. | |
| Neural Network | Manage neural networks for training. | |
| Environment | Processing of environment information for RL. | |
| NNModel | Save trained models and weight values. |
| Parameter | DQN-MLP | Dueling DQN-MLP | Dueling DQN-GRU | PPO-MLP |
|---|---|---|---|---|
| MLP hidden dim | 128 | 128 | – | 128 |
| RNN hidden dim | – | – | 128 | – |
| Input dimension | 31 | |||
| Output dimension | 6 | |||
| Learning rate () | Actor: Critic: | |||
| Discount factor () | 0.99 | |||
| Update cycle | 50 | |||
| Episode limit (steps) | 3000 | |||
| Batch size (episodes) | 5 | 10 | ||
| start/min/decay | 1.0/0.05/0.99 | – | ||
| PPO clip/K epochs | – | 0.2/10 | ||
| Architecture | DGAR (%) | BE | Potential () | Jerk |
|---|---|---|---|---|
| (A) Rule-based Agent | ||||
| (B) DQN+MLP w/o TAP | ||||
| (C) PPO+MLP w/o TAP | ||||
| (D) DDQN+MLP w/o TAP | ||||
| (E) DDQN+MLP w/ TAP (Add.) | ||||
| (F) DDQN+MLP w/ TAP (Mult.) | ||||
| (G) DDQN+RNN w/ TAP (Mult.) [Proposed] |
| Contrast | ΔDGAR (pp) | Welch t | p | Cohen’s d | 95% CI | |
|---|---|---|---|---|---|---|
| (G) Proposed vs. (D) DDQN+MLP w/o TAP | ||||||
| (G) Proposed vs. (B) DQN+MLP (strongest baseline) | ||||||
| (F) Multiplicative vs. (E) Additive TAP | ||||||
| (G) Recurrent vs. (F) Feedforward | ||||||
| (F) TAP-Mult. vs. (D) no-TAP |
| (A) Marginal | (B) Aggregate of | (C) Joint Conditions | |||
|---|---|---|---|---|---|
| Model | DGAR (%) | Potential () | P( > 1.3) (%) | eGAR (%) | TJAR (%) |
| DQN+MLP w/o TAP | |||||
| PPO+MLP w/o TAP | |||||
| DDQN+MLP w/o TAP | |||||
| DDQN+MLP w/TAP (Add.) | |||||
| DDQN+MLP w/TAP (Multi.) | |||||
| DDQN+RNN w/ TAP (Multi.) | |||||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Shin, H.; Ahn, J.; Kim, D.; Ahn, J.H.; Kim, J.; Kim, J.; Kim, S. TAP-DDQN: Multiplicative Potential-Based Reward Shaping Framework for Tactical Decision-Making of Unmanned Surface Vehicles in Adversarial Maritime Engagements. J. Mar. Sci. Eng. 2026, 14, 1569. https://doi.org/10.3390/jmse14171569
Shin H, Ahn J, Kim D, Ahn JH, Kim J, Kim J, Kim S. TAP-DDQN: Multiplicative Potential-Based Reward Shaping Framework for Tactical Decision-Making of Unmanned Surface Vehicles in Adversarial Maritime Engagements. Journal of Marine Science and Engineering. 2026; 14(17):1569. https://doi.org/10.3390/jmse14171569
Chicago/Turabian StyleShin, Hunyong, Jinsu Ahn, Dongyoung Kim, Jin Ho Ahn, Jinyong Kim, Jonggeun Kim, and Sungshin Kim. 2026. "TAP-DDQN: Multiplicative Potential-Based Reward Shaping Framework for Tactical Decision-Making of Unmanned Surface Vehicles in Adversarial Maritime Engagements" Journal of Marine Science and Engineering 14, no. 17: 1569. https://doi.org/10.3390/jmse14171569
APA StyleShin, H., Ahn, J., Kim, D., Ahn, J. H., Kim, J., Kim, J., & Kim, S. (2026). TAP-DDQN: Multiplicative Potential-Based Reward Shaping Framework for Tactical Decision-Making of Unmanned Surface Vehicles in Adversarial Maritime Engagements. Journal of Marine Science and Engineering, 14(17), 1569. https://doi.org/10.3390/jmse14171569

