Deep Reinforcement Learning-Based Dual-Loop Adaptive Control Method and Simulation for Loitering Munition Fuze
Abstract
1. Introduction
- (1)
- A dual-loop fuze control architecture is developed, in which the outer loop handles mode-level decision and safety rollback, and the inner loop performs continuous online tuning of fuzzy scaling factors.
- (2)
- A fuze-oriented Markov Decision Process is formulated by incorporating engagement deviation, interference intensity, readiness-related information, and mode-dependent control objectives.
- (3)
- A co-simulation verification framework is established to evaluate the proposed method under dynamic interference and task-switching conditions.
2. Dynamic Reconfiguration Requirements and Technical Route of Loitering Munition Fuze
2.1. Dynamic Reconfiguration Requirements
2.2. Hierarchical Dual-Loop Control Architecture
2.2.1. Overall Architecture
2.2.2. Fuzzy Logic as the Execution Layer
2.2.3. TD3-Based Online Tuning Layer
2.2.4. Multi-Modal Reconfiguration and Safety Modeling Based on MDP
2.2.5. Technical Route
3. Multi-Modal Reconfiguration and Fuzzy Safety Control Modeling Based on Markov Decision Process
3.1. Physical State Vector Normalization and Observation State Space Construction
3.1.1. Physical State Vector and Execution Parameter Definition
3.1.2. State Normalization Processing and Observation Space Construction
- Quantization of interference intensity
- 2.
- Quantization of target confidence
- 3.
- Quantization of relative target velocity
- 4.
- Quantization of loitering munition acceleration
- 5.
- Quantization of relative target distance
- 6.
- Construction of deviation-related observation variables
3.1.3. Practical Setting Rationale for Threshold Parameters
3.2. Continuous Reconfiguration Optimization Modeling Based on MDP
3.2.1. Action Space
3.2.2. Reward Function Design
- Precision reward : Real-time encouragement to minimize burst height error, adopting the form of a Gaussian kernel function:
- 2.
- Safety constraint reward : If a false alarm (false safety release) occurs in a non-initiation zone (e.g., outside the safe distance), a massive penalty is given:
- 3.
- Action smoothness penalty : Penalizes abrupt changes in the reconfiguration gain, so as to reduce high-frequency parameter jitter and improve the practical executability of the control signal:
3.3. Fuzzy Inference Model for Loitering Munition Fuze Initiation Parameter Reconfiguration
3.3.1. Fuzzy Subsets and Membership Function Design
3.3.2. Dynamic Fuzzy Inference
3.3.3. Defuzzification and Parameter Output
3.4. Multi-Modal Task and Safety-State Reconfiguration Decision Model
3.4.1. Multi-Modal Task Reconfiguration Logic
3.4.2. Safety-State Reconfiguration Logic
4. Design of Initiation Parameter Reconfiguration Optimization Algorithm Based on Deep Reinforcement Learning
4.1. Principle and Optimization Mechanism of TD3 Algorithm
4.1.1. Clipped Double Q-Learning Mechanism
4.1.2. Target Policy Smoothing
4.1.3. Delayed Policy Updates
4.2. Policy Network Architecture Design
4.2.1. Actor–Critic Network Topology
- Actor network : Responsible for outputting the reconfiguration gain action .
- 2.
- Critic network (): Adopts a dual-network structure, responsible for evaluating action values.
4.2.2. Optimizer Configuration and Parameter Settings
4.2.3. Hyperparameter Sensitivity and Baseline Algorithm Comparison
4.3. Co-Simulation Training Platform Setup
4.3.1. Training Environment Interface Design
4.3.2. Algorithm Online Execution Pseudocode
| Algorithm 1: TD3 Pseudocode |
| 1. Initialize the Actor policy network and two Critic value networks with random parameters. Initialize the corresponding target networks and copy the main network parameters to the target networks. Establish an experience replay buffer with capacity . 2. For each training episode , reset the MATLAB/Simulink simulation environment and obtain the initial fuze observation state . 3. Within each simulation time step , execute the following operations in sequence: (1) Action Generation: Output the reconfiguration action based on the current policy network, and add Gaussian exploration noise to enhance exploration capability: . (2) Environment Response: Pass into the MATLAB environment to correct the fuzzy scaling factors in real time and execute one simulation step. If the safety truncation mechanism is triggered (FRI < 0.4), forcibly abort the episode and impose a penalty; otherwise, observe the immediate reward and the next state . (3) Experience Storage: Store the state transition tuple () into the experience replay buffer . If the number of samples in exceeds the capacity , overwrite the oldest data. 4. When the data volume in the replay buffer meets the batch size B requirement, randomly sample a batch of data. Add clipped noise when calculating the target action to achieve policy smoothing and calculate the target value through the clipped double Q-learning mechanism, taking the minimum of the two target Critic networks. Update the two Critic network parameters using the mean squared error minimization criterion. 5. If the current update step meets the delay frequency (i.e., ), perform one Actor network update by maximizing the Q-value via the deterministic policy gradient. Simultaneously, use the soft update method to synchronously update all target network parameters. 6. Update the current state . If burst height convergence or a timeout is not triggered, the current episode is not over, so loop back to step 3 to continue execution. Otherwise, end this episode; if the convergence conditions are met, save the current network parameters as the optimal parameters and end the training. |
5. Implementation and Simulation of Optimization Algorithm Based on Deep Reinforcement Learning
5.1. Simulation Environment Setup and Signal Characteristics
5.1.1. Virtual Battlefield Environment Model
5.1.2. Evaluation Metrics Definition
- Root Mean Square Error (RMSE): Quantifies the robustness of the model under extreme working conditions.
- 2.
- Mean Absolute Error (MAE): Measures how much the algorithm deviates from the true distance on average throughout the entire combat process.
- 3.
- False Alarm Rate (FAR): Measures the probability of the fuze triggering falsely in a non-initiation zone, characterizing safety.
- 4.
- Reconfiguration Response Time (): Measures the adaptation speed of the algorithm to sudden interference.
5.2. Empirical Training Stability Analysis
5.3. Full-Process Task Decision and Safety Logic Simulation
- Initial State (T = 0~1.5 s): The system is in the target approach phase. Input conditions include: high target confidence , target identified as personnel, and fuze readiness .
- Tactical Change (T = 1.5 s): The seeker secondarily confirms that the target nature has changed to a reinforced bunker. The system needs to respond to this tactical input by performing task modal reconfiguration, changing the initiation mode to penetration mode.
- Safety Fuse (T = 2.5 s): This simulates a sudden mission abort command, causing a Fuze Readiness Index , falling below the safety threshold. The system must immediately execute reverse safety-state reconfiguration, returning the fuze to a safe state.
5.4. Reconfiguration Performance Verification Under Typical Strike Scenarios
5.5. Comparative Analysis of Performance of Different Control Strategies
5.5.1. Strategy Response Comparative Analysis
5.5.2. Statistical Performance Metrics Evaluation
5.6. Robustness Analysis
5.6.1. Robustness Analysis Under Different Interference Noise Ratios
5.6.2. Robustness Analysis Under Intermittent Sensor Signal Loss
5.6.3. Robustness Analysis Under Different Sensor Noise Levels
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Wang, P.; Li, H.J.; Yu, H.; Zhang, C. Calculation of recoverable failure rate for loitering munition fuze electronic safety system. J. Detect. Control 2025, 47, 57–63. [Google Scholar] [CrossRef]
- Zhang, C.H.; Li, H.J.; Gong, X.F.; Chen, Z.P.; Yu, H. Design and verification of multi-state safety logic control method for loitering munition fuze based on electronic safety system. Acta Armamentarii 2023, 44, 3079–3090. [Google Scholar] [CrossRef]
- Li, S.Q.; Peng, Z.L.; Zhao, H.M.; Yang, Y.; Xia, Y.; Wang, Y. Program design and simulation of loitering munition electronic safety system. J. Ordnance Equip. Eng. 2022, 43, 303–308. (In Chinese) [Google Scholar] [CrossRef]
- Zhang, X.; Zhou, Y.Y.; Bai, F.; Yang, Y. Development status and key technology analysis of loitering munition. Electro-Mech. Eng. 2025, 41, 34–39. [Google Scholar] [CrossRef]
- Jia, J.W.; Gao, M.; Han, Z.Z.; Dao, X. Review of anti-informational jamming techniques for radio proximity fuze. AIP Adv. 2025, 15, 100704. [Google Scholar] [CrossRef] [Scilit]
- Cui, Y.Y. Multi-mode fusion of fuze based on hierarchical information fusion. Telecommun. Eng. 2016, 56, 670–674. [Google Scholar] [CrossRef]
- Liu, P.; Li, J.; Yu, H.; Zhang, H. Innovation of control technology for smart fuzes: Precise detonation and efficient damage via a ternary cascade controller. Front. Phys. 2024, 11, 1309373. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Sun, H.; Zhou, X.; Zhao, H. Inherent anti-jamming performance evaluation of cyclic modulation continuous wave radio fuzes based on ambiguity function incision. Arab. J. Sci. Eng. 2018, 43, 3037–3047. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Z.J.; Gong, P.; Zhang, J.H.; Li, W.Y.; Ren, K.; Gao, X.; Zhang, G.W. Simulation method of anti-information jamming effectiveness for FMCW fuze. Acta Armamentarii, 2025; Epub ahead of printing. (In Chinese). [CrossRef]
- Wang, J.W.; Shi, K.; Xu, G.T.; Qian, C.R.; Yan, J. Roll estimation of high rotation speed correction fuze based on extended Kalman filter. J. Northwestern Polytech. Univ. 2016, 34, 938–944. Available online: https://kns.cnki.net/kcms2/article/abstract?v=y_SiIdm5mqvrG8sJBxNfaGwD9sXVLS_HpRH7wQnur2ATIbrcC2no2uktqbbrxS3lNQUFnl2OfjPpx7t4UZeYN5NzhM290Idt9UBp17MU8Am0iXPgxBZsyDSikuBovwWnc0WWdeATFW4k_1fo3mDleQys44IWAuIsx0IJfdhbD-DG_0u-_IoyqQ==&uniplatform=NZKPT&language=CHS (accessed on 4 March 2026). (In Chinese)
- Anagnostou, D.E.; Chryssomallis, M.T.; Goudos, S. Reconfigurable antennas. Electronics 2021, 10, 897. [Google Scholar] [CrossRef] [Scilit]
- Woźniak, M.; Zielonka, A.; Sikora, A. Driving support by type-2 fuzzy logic control model. Expert Syst. Appl. 2022, 207, 117798. [Google Scholar] [CrossRef] [Scilit]
- Di Renzo, M.; Zappone, A.; Debbah, M.; Alouini, M.-S.; Yuen, C.; de Rosny, J.; Tretyakov, S. Smart radio environments empowered by reconfigurable intelligent surfaces: How it works, state of research, and road ahead. IEEE J. Sel. Areas Commun. 2020, 38, 2450–2525. [Google Scholar] [CrossRef] [Scilit]
- Aivaliotis-Apostolopoulos, P.; Loukidis, D. Swarming genetic algorithm: A nested fully coupled hybrid of genetic algorithm and particle swarm optimization. PLoS ONE 2022, 17, e0275094. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.-Q.; Guo, J.-L.; Yang, H.; Wang, Z.-Q.; Liu, H.-L. Research on navigation of bidirectional A* algorithm based on ant colony algorithm. J. Supercomput. 2021, 77, 1958–1975. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Chen, Y.; Zhao, X.; Huang, J. An improved DQN path planning algorithm. J. Supercomput. 2022, 78, 616–639. [Google Scholar] [CrossRef] [Scilit]
- He, R.; Lv, H.; Zhang, S.; Zhang, D.; Zhang, H. Lane following method based on improved DDPG algorithm. Sensors 2021, 21, 4827. [Google Scholar] [CrossRef] [Scilit]
- Li, T.J. Research on Rate Optimization of UAV Communication System Assisted by Reconfigurable Intelligent Surface. Master’s Thesis, Beijing University of Posts and Telecommunications, Beijing, China, 2024. (In Chinese) [Google Scholar] [CrossRef]
- Fujimoto, S.; van Hoof, H.; Meger, D. Addressing function approximation error in actor-critic methods. In Proceedings of the 35th International Conference on Machine Learning (PMLR), Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 1587–1596. [Google Scholar] [CrossRef] [Scilit]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning (PMLR), Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 1861–1870. [Google Scholar] [CrossRef] [Scilit]
- Shi, Q.; Lam, H.K.; Xuan, C.; Chen, M. Adaptive neuro-fuzzy PID controller based on twin delayed deep deterministic policy gradient algorithm. Neurocomputing 2020, 402, 183–194. [Google Scholar] [CrossRef] [Scilit]
- Wachi, A.; Shen, X.; Sui, Y. A survey of constraint formulations in safe reinforcement learning. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI-24), Jeju, Republic of Korea, 3–9 August 2024. [Google Scholar]
- Yang, D.; Liu, Z.; Yi, P. Computational efficiency of accelerated particle swarm optimization combined with different chaotic maps for global optimization. Neural Comput. Appl. 2017, 28, 1245–1264. [Google Scholar] [CrossRef] [Scilit]
- Gu, X.P.; Shi, X.J. Reconfigurability analysis and design of UAV based on structure analysis and tree seed algorithm. Syst. Eng. Electron. 2025, 48, 932–945. Available online: https://link.cnki.net/urlid/11.2422.TN.20250611.0906.014 (accessed on 3 January 2026). (In Chinese)
- Luo, J.S.; Huang, P.; Luo, X.Y.; Lai, H. Aerial reconfigurable intelligent surface-assisted full-duplex UAV secure communication. Telecommun. Sci. 2024, 40, 34–50. [Google Scholar] [CrossRef]
- Krayani, A.; Alam, A.S.; Marcenaro, L.; Nallanathan, A.; Regazzoni, C. Automatic jamming signal classification in cognitive UAV radios. IEEE Trans. Veh. Technol. 2022, 71, 12972–12988. [Google Scholar] [CrossRef] [Scilit]
- Jiang, H.; Li, S.; Lin, C.; Wang, C.; Tian, G.; Wang, S.; Zhong, K.; Li, J. Research on target assignment method based on ant colony-fish group algorithm. J. Phys. Conf. Ser. 2019, 1419, 012002. [Google Scholar] [CrossRef] [Scilit]
- Jia, J.W.; Gao, M.; Han, Z.Z.; Liu, L.M.; Yin, Y.W. Overviewofanti-informationinterferencetechnologyforradioproximityfuze. Syst. Eng. Electron. 2025, 47, 1074–1107. Available online: https://link.cnki.net/urlid/11.2422.tn.20241217.1546.002 (accessed on 11 January 2026). (In Chinese)











| Target Type | Firing Mode (M) | Core Control Parameter |
|---|---|---|
| Personnel/Radar | Proximity Mode M = 00 | Burst Height Threshold |
| Trenches | Airburst Mode M = 01 | Delay Time |
| Light Vehicles/General Fortifications | Impact Mode M = 10 | Impact Acceleration |
| Tanks/Bunkers | Delay Mode M = 11 | Penetration Delay |
| Combat Phase Name | Phase Parameter | Operational Meaning |
|---|---|---|
| Safe/Loitering Mode | 1 | Firing prohibited, parameters cannot be regulated |
| Approach/Judgment Mode | 2 | Parameter fine-tuning, trigger judgment begins |
| Attack/Firing Mode | 3 | Enters final firing window |
| Parameter Name | Sign | Set Value | Meaning |
|---|---|---|---|
| Actor Learning Rate | A small learning rate ensures stable policy evolution | ||
| Critic Learning Rate | A larger learning rate accelerates the convergence of value evaluation | ||
| Discount Factor | Focuses on long-term cumulative returns | ||
| Experience Replay Buffer Capacity | Stores a large number of historical samples to eliminate correlation | ||
| Batch Size | B | The number of samples sampled from the buffer each time | |
| Soft Update Coefficient | The moving average update rate for target network parameters | ||
| Exploration Noise | Gaussian noise applied to actions in the early stage of training |
| Noise Ratio | RMSE | MAE |
|---|---|---|
| 30% | 0.542 | 0.431 |
| 60% | 0.589 | 0.475 |
| 80% | 0.512 | 0.410 |
| 100% | 0.685 | 0.544 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhang, L.; Li, H.; Zhang, C.; Zhao, Y.; Qiao, S.; Yu, H. Deep Reinforcement Learning-Based Dual-Loop Adaptive Control Method and Simulation for Loitering Munition Fuze. Technologies 2026, 14, 239. https://doi.org/10.3390/technologies14040239
Zhang L, Li H, Zhang C, Zhao Y, Qiao S, Yu H. Deep Reinforcement Learning-Based Dual-Loop Adaptive Control Method and Simulation for Loitering Munition Fuze. Technologies. 2026; 14(4):239. https://doi.org/10.3390/technologies14040239
Chicago/Turabian StyleZhang, Lingyun, Haojie Li, Chuanhao Zhang, Yuan Zhao, Shixiang Qiao, and Hang Yu. 2026. "Deep Reinforcement Learning-Based Dual-Loop Adaptive Control Method and Simulation for Loitering Munition Fuze" Technologies 14, no. 4: 239. https://doi.org/10.3390/technologies14040239
APA StyleZhang, L., Li, H., Zhang, C., Zhao, Y., Qiao, S., & Yu, H. (2026). Deep Reinforcement Learning-Based Dual-Loop Adaptive Control Method and Simulation for Loitering Munition Fuze. Technologies, 14(4), 239. https://doi.org/10.3390/technologies14040239

