Long-Horizon Constraint-Aware Collaborative Scheduling for Multiple Phased-Array Radars Using Mamba Temporal Encoding and Structured Hybrid Actions
Abstract
1. Introduction
- (1)
- We formulate collaborative scheduling of multiple homogeneous phased-array radars as a finite-horizon constraint-aware decision problem with structured hybrid actions. The formulation explicitly distinguishes instantaneous hard constraints, long-term resource constraints, soft mission-level requirements, and terminal performance metrics.
- (2)
- We develop a radar-specific structured hybrid action-generation mechanism. The neural scheduler first predicts task priorities and radar–task compatibility scores; these scores are then converted into executable radar–task assignments by graph-edge masking and constrained maximum-weight matching.
- (3)
- We introduce a feasible transmit-power projection layer and an action-dependent radar performance model, which links transmit power to effective signal-to-noise ratio, detection probability, measurement noise, and tracking covariance. This establishes a physical causal path from continuous power actions to scheduling utility.
- (4)
- We combine Mamba temporal encoding with two-stage policy optimization. Behavioral cloning provides a stable initialization from constraint-aware heuristic trajectories, while actor–critic fine-tuning improves beyond the heuristic expert under the same feasible action-generation pipeline.
- (5)
- We construct a reproducible MRSched-Bench simulation protocol and report not only normalized cost-effectiveness score and latency but also hard-constraint violation rate, soft-constraint satisfaction rate, task completion rate, power utilization, and tracking error.
2. Related Work
2.1. Radar Resource Management and Cooperative Radar Scheduling
2.2. Constrained and Hybrid-Action Reinforcement Learning
2.3. Sequence Models for Decision Making
2.4. Positioning of This Work
3. Method
3.1. Overall Framework
3.2. Problem Formulation
3.2.1. Radar Network and Task Model
3.2.2. State Variables and Transition Dynamics
3.2.3. Constraint Classification
3.3. Reward Function and Evaluation Metric
3.4. Mamba-Based Temporal Encoder
3.5. Decision-Mamba Scheduler
- Stage 1: Task Priority Prediction.
- Stage 2: Constraint-Aware Radar–Task Matching.
- Stage 3: Feasible Power Allocation.
- Two-Stage Policy Optimization.
4. Experiments
4.1. Experimental Setup
4.2. Baselines and Implementation Details
4.3. Constraint and Task-Level Metrics
4.4. Results on MRSched-Bench
4.5. Results on Public Scheduling Benchmarks
4.6. Complexity and Real-Time Analysis
4.7. Ablation Study
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Ota, K. Decision Mamba: Reinforcement Learning via Sequence Modeling with Selective State Spaces. arXiv 2024, arXiv:2403.19925. [Google Scholar]
- Chen, H.; Luo, T.; Yang, B.; Sun, L.; Huang, S.; Hu, J.; Yang, Z.; Yang, L. Mamba-based Reinforcement Learning for Long-Horizon Decision Making. arXiv 2024, arXiv:2406.00079. [Google Scholar]
- Dao, T.; Gu, A. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality. arXiv 2024, arXiv:2405.21060. [Google Scholar]
- Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. Adv. Neural Inf. Process. Syst. 2017, 30, 6379–6390. [Google Scholar]
- Rashid, T.; Samvelyan, M.; Schroeder, C.; Farquhar, G.; Foerster, J.; Whiteson, S. QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning. In Proceedings of the International Conference on Machine Learning (ICML), Stockholm, Sweden, 10–15 July 2018. [Google Scholar]
- Zhang, H.; Liu, W.; Zhang, L.; Meng, Y.; Han, W.; Song, T.; Yang, R. An allocation strategy integrated power, bandwidth, and subchannel in a RCC network. Def. Technol. 2025, 60, 138–154. [Google Scholar] [CrossRef]
- Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.; Wu, Y. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. Adv. Neural Inf. Process. Syst. 2022, 35, 24611–24624. [Google Scholar] [CrossRef]
- Zhang, H.; Liu, W.; Zhang, L.; Meng, Y.; Song, T.; Xu, H.X. Joint resource and trajectory optimization in a UAV-enabled dual-function radar-communication network. Def. Technol. 2025, 60, 30–49. [Google Scholar] [CrossRef]
- Sun, M.; Zhang, Q.; Chen, G. Adaptive scheduling algorithm for phased array radar under dynamic time windows. J. Radar 2018, 7, 303–312. [Google Scholar]
- Wang, X.; Yi, W.; Kong, L. Joint beam and dwell time allocation method for phased array radar based on multi-objective tracking. J. Radar 2017, 6, 602–610. [Google Scholar]
- Wang, X. Research on Beam and Time Resource Management Algorithms for Phased Array Radar Tracking Modes. Master’s Thesis, University of Electronic Science and Technology of China, Chengdu, China, 2018. [Google Scholar]
- Xie, M.; Yi, W.; Kong, L.; Kirubarajan, T. Receive beam resource allocation for multiple target tracking with distributed MIMO radars. IEEE Trans. Aerosp. Electron. Syst. 2018, 54, 2421–2436. [Google Scholar] [CrossRef]
- Tang, J. Research on Resource Scheduling Method for Multi-Function Radar Networking. Master’s Thesis, University of Electronic Science and Technology of China, Chengdu, China, 2021. [Google Scholar] [CrossRef]
- Li, S.; Xu, G.; Lin, H. Research on Automatic Radar Deployment Based on Improved Ant Colony Algorithm. Foreign Electron. Meas. Technol. 2021, 40, 41–47. [Google Scholar] [CrossRef]
- Hu, B.; Zhu, Y.; Zhou, Y. Simulated annealing whale radar resource scheduling algorithm for Koei variation. J. Northwest. Polytech. Univ. 2022, 40, 796–803. [Google Scholar] [CrossRef]
- Zhao, L.; Shi, X. Simulation of adaptive resource scheduling for multifunctional phased array radar. Fire Control Radar Technol. 2022, 51, 102–108. [Google Scholar] [CrossRef]
- Feng, L.-W.; Liu, S.-T.; Xu, H.-Z. Multifunctional Radar Cognitive Jamming Decision Based on Dueling Double Deep Q-Network. IEEE Access 2022, 10, 112150–112157. [Google Scholar] [CrossRef]
- Ye, Y.; Wei, Y.; Lingjiang, K. Joint tracking sequence and dwell time allocation for multi-target tracking with phased array radar. Signal Process. 2022, 192, 108374. [Google Scholar] [CrossRef]
- Yi, W.; Yuan, Y.; Liu, G. Research progress on multi-radar cooperative detection technology: Cognitive tracking and resource scheduling algorithms. J. Radar 2023, 12, 471–499. [Google Scholar]
- Zhou, Q.; Yi, W.; Yuan, Y.; Ding, J.; Kong, L.; Yang, J. Distributed multi-base passive radar signal-level cooperative target localization technology. Mod. Radar 2024, 46, 64–78. [Google Scholar] [CrossRef]
- Ding, J. Research on Complexity Mechanisms and Applications of Human-Machine Decision-Making Integration. Mod. Radar 2024, 46, 1–8. [Google Scholar] [CrossRef]
- Li, Z.; Chen, X.; Huang, H. Optimization of Multi-Park Integrated Energy Systems Based on Hierarchical Reinforcement Learning for Multi-Agent Systems. Zhejiang Electr. Power 2025, 44, 46–57. [Google Scholar] [CrossRef]
- Huang, Y.; Zhang, X.; Yue, D.; Hu, S.; Wang, J.; Li, Z. Active voltage regulation strategy for distribution networks based on multi-agent deep reinforcement learning. Power Syst. Autom. 2025, 49, 65–73. [Google Scholar]
- Wang, T.; Dou, L.; Li, Z. Prescribed performance tracking control for nonlinear multi-agent systems. Control Theory Appl. 2026, 43, 79–89. [Google Scholar] [CrossRef]
- Lu, X. Research on Radar Resource Scheduling Method for Target Tracking Under Low Capture Constraints. Ph.D. Thesis, University of Electronic Science and Technology of China, Chengdu, China, 2023. [Google Scholar] [CrossRef]
- Wang, P. Selection of Cognitive Radar Waveform Parameters and Resource Scheduling for Collaborative Tracking Optimization. Master’s Thesis, Harbin Institute of Technology, Harbin, China, 2021. [Google Scholar]
- Bai, H. Research on Resource Scheduling Method for Multi-Station Radar Collaborative Detection. Master’s Thesis, Xi’an University of Electronic Science and Technology, Xi’an, China, 2023. [Google Scholar] [CrossRef]
- Guo, Z.; Liu, Y.; Wang, Y.; Meng, Y.; Liu, B. Joint Communication-Motion Planning for UAV Swarm against Jamming with Multi-Agent Deep Reinforcement Learning. In 2024 IEEE 35th International Symposium on Personal, Indoor and Mobile Radio Communications (PIMRC); IEEE: Piscataway, NJ, USA, 2024; pp. 1–7. [Google Scholar] [CrossRef]
- Zhan, Y.; Zhang, J.; Lin, Z.; Xiao, L. Energy-Efficient Anti-Jamming Metaverse Resource Allocation Based on Reinforcement Learning. In 2024 IEEE Wireless Communications and Networking Conference (WCNC); IEEE: Piscataway, NJ, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef]
- Godrich, H.; Petropulu, A.; Poor, H.V. A combinatorial optimization framework for subset selection in distributed multiple-radar architectures. In 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: Piscataway, NJ, USA, 2011; pp. 2796–2799. [Google Scholar] [CrossRef]
- Xie, M.; Yi, W.; Kirubarajan, T.; Kong, L. Joint Node Selection and Power Allocation Strategy for Multitarget Tracking in Decentralized Radar Networks. IEEE Trans. Signal Process. 2018, 66, 729–743. [Google Scholar] [CrossRef]
- Wu, Y.; Fioranelli, F.; Gao, C. RadMamba: Efficient Human Activity Recognition Through a Radar-Based Micro-Doppler-Oriented Mamba State-Space Model. IEEE Trans. Radar Syst. 2026, 4, 261–272. [Google Scholar] [CrossRef]
- Kozy, M.; Yu, J.; Buehrer, R.M.; Martone, A.; Sherbondy, K. Applying Deep-Q Networks to Target Tracking to Improve Cognitive Radar. In 2019 IEEE Radar Conference (RadarConf); IEEE: Piscataway, NJ, USA, 2019; pp. 1–6. [Google Scholar] [CrossRef]
- Zhang, Z.; Yuan, Y.; Sun, J.; Han, K.; Li, H.; Yi, W. Reinforcement-Learning-Based Agile Transmission Strategy for Networked Radar System Anti-Jamming. IEEE Trans. Aerosp. Electron. Syst. 2025, 61, 11543–11557. [Google Scholar] [CrossRef]
- Shen, S.; Petropulu, A.P.; Poor, H.V. Hierarchical Reinforcement Learning for Joint Task Assignment and Power Allocation in Cognitive Radar Networks. IEEE Trans. Signal Process. 2023, 71, 1245–1260. [Google Scholar]
- Sun, S.; Petropulu, A.P. A Sparse Linear Array Approach in Automotive Radars Using Matrix Completion. In ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: Piscataway, NJ, USA, 2020; pp. 8614–8618. [Google Scholar] [CrossRef]
- Feng, L.; Liu, S.; Xu, H. Constrained Reinforcement Learning for Hard-Constrained Radar Resource Management with Action Masking. IEEE Trans. Cybern. 2023, 53, 7012–7025. [Google Scholar]
- Ahmed, A.M.; Ahmad, A.A.; Fortunati, S.; Sezgin, A.; Greco, M.S.; Gini, F. A Reinforcement Learning Based Approach for Multitarget Detection in Massive MIMO Radar. IEEE Trans. Aerosp. Electron. Syst. 2021, 57, 2622–2636. [Google Scholar] [CrossRef]
- Gongguo, X.; Ganlin, S.; Xiusheng, D. Sensor scheduling for ground maneuvering target tracking in presence of detection blind zone. J. Syst. Eng. Electron. 2020, 31, 692–702. [Google Scholar] [CrossRef]
- Lu, X.; Xu, Z.; Ren, H.; Yi, W. LPI-based Resource Allocation Strategy for Target Tracking in the Moving Airborne Radar Network. In 2022 IEEE Radar Conference (RadarConf22); IEEE: Piscataway, NJ, USA, 2022; pp. 1–6. [Google Scholar] [CrossRef]
- Yan, J.; Pu, W.; Zhou, S.; Liu, H.; Greco, M.S. Optimal Resource Allocation for Asynchronous Multiple Targets Tracking in Heterogeneous Radar Networks. IEEE Trans. Signal Process. 2020, 68, 4055–4068. [Google Scholar] [CrossRef]
- Li, Z.-J.; Zhang, H.-W.; Xie, J.-W.; Xiang, H.-H.; Ge, J.-A.; Wang, B. An Optimized Resource Allocation Algorithm in Cognitive C-MIMO Radar for Multiple Maneuvering Target Tracking. In 2021 CIE International Conference on Radar (Radar); IEEE: Piscataway, NJ, USA, 2021; pp. 1691–1694. [Google Scholar] [CrossRef]
- Su, Y.; He, Z.; Deng, M.; Wang, J. Collaborative Resource Allocation and Beampattern Optimization for Maneuvering Targets Tracking with Distributed Radar Network. In IGARSS 2022–2022 IEEE International Geoscience and Remote Sensing Symposium; IEEE: Piscataway, NJ, USA, 2022; pp. 7669–7672. [Google Scholar] [CrossRef]
- Li, Z.; Xie, J.; Liu, W.; Zhang, H.; Xiang, H. Joint Strategy of Power and Bandwidth Allocation for Multiple Maneuvering Target Tracking in Cognitive MIMO Radar With Collocated Antennas. IEEE Trans. Veh. Technol. 2023, 72, 190–204. [Google Scholar] [CrossRef]
- Zhao, L.; Wang, F.; Hu, W. Two-Stage Imitation Learning with Reinforcement Fine-Tuning for Real-Time Radar Task Scheduling. IEEE Trans. Veh. Technol. 2024, 73, 13682–13693. [Google Scholar]
- Yang, S.X.; Deb, K. A survey of dynamic multi-objective optimization. Swarm Evol. Comput. 2021, 62, 100859. [Google Scholar]
- Ahmed, A.; Zhang, Y.D. Optimized Resource Allocation for Distributed Joint Radar-Communication System. IEEE Trans. Veh. Technol. 2024, 73, 3872–3885. [Google Scholar] [CrossRef]
- Lu, Z.; Kalia, S.; Gursoy, M.C.; Mohan, C.K.; Varshney, P.K. Multi-Objective Reinforcement Learning for Cognitive Radar Resource Management. In ICASSP 2025—2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: Piscataway, NJ, USA, 2025; pp. 1–5. [Google Scholar] [CrossRef]
- Wang, Y.; Liang, Y.; Wang, Z. Hierarchical Reinforcement-Learning-Based Joint Allocation of Jamming Task and Power for Countering Networked Radar. IEEE Trans. Aerosp. Electron. Syst. 2025, 61, 2149–2167. [Google Scholar] [CrossRef]
- Lu, Z.; Gursoy, M.C.; Mohan, C.K.; Varshney, P.K. Adaptive Resource Management in Cognitive Radar via Deep Deterministic Policy Gradient. In 2025 IEEE International Radar Conference (RADAR); IEEE: Piscataway, NJ, USA, 2025; pp. 1–6. [Google Scholar] [CrossRef]
- Yan, J.; Liu, H.; Greco, M.S. CTDE-MARL for Coherent Multi-Phased Array Radar Coordination with Long-Term Resource Constraints. IEEE Trans. Signal Process. 2025, 73, 890–905. [Google Scholar] [CrossRef]




| Method Family | Long Horizon | Hybrid Action | Hard Feasibility | Radar Physical Model | Online Fine-Tuning |
|---|---|---|---|---|---|
| MILP/ILP | Partial | Yes | Yes | Simplified | No |
| CMDP/CPO | Yes | Limited | Penalty-based | No | Yes |
| P-DQN | Partial | Yes | Usually no | No | Yes |
| MAPPO/QMIX | Partial | Limited | Mask or penalty | Usually no | Yes |
| Decision Transformer | Yes | Limited | No | No | Limited |
| Decision Mamba | Yes | Limited | No | No | Limited |
| Classical radar allocation | Limited | Yes | Yes | Yes | No |
| CS-Mamba | Yes | Yes | Matching + projection | Yes | Yes |
| Component | Symbol | Dimension | Meaning |
|---|---|---|---|
| Residual power | Remaining radar energy budget | ||
| Beam availability | Whether radar r can execute a task | ||
| Radar load | Historical utilization of radar nodes | ||
| Target position | or | Target spatial location | |
| Target velocity | or | Target motion state | |
| Tracking covariance trace | Tracking uncertainty | ||
| Task priority | Mission priority | ||
| Deadline | Remaining task deadline | ||
| Refresh interval | Slots since last execution | ||
| Jamming level | Jammer-to-noise ratio | ||
| Clutter level | Clutter-to-noise ratio | ||
| Geometric accessibility | Beam visibility and coverage |
| Variable or Setting | Dimension or Value |
|---|---|
| Input state | |
| Embedding vector | |
| Hidden state | |
| Output representation | |
| Number of Mamba blocks | 3 |
| State-space dimension | 16 |
| Sequence length in Medium scenario | 200 |
| Parameter | Small | Medium | Large |
|---|---|---|---|
| Number of radars | 2 | 4 | 8 |
| Number of targets | 4–6 | 8–12 | 16–24 |
| Candidate task categories | 10 | 20 | 40 |
| Scheduling horizon K | 100 | 200 | 400 |
| Slot length | 10 ms | 10 ms | 10 ms |
| JNR range | 0–15 dB | 0–20 dB | 5–25 dB |
| CNR range | 0–10 dB | 0–15 dB | 5–20 dB |
| Training scenarios | 1000 | 2000 | 3000 |
| Test scenarios | 100 | 100 | 100 |
| Random seeds | 5 | 5 | 5 |
| Coefficient | Value | Meaning |
|---|---|---|
| 0.35 | Resource-cost weight in CE | |
| 0.25 | Penalty weight in CE | |
| 0.10 | Transmit-power cost coefficient | |
| 0.50 | Cooperative tracking penalty coefficient | |
| 0.50 | Periodic refresh penalty coefficient | |
| 1.00 | Detection-probability slope | |
| 10 dB | Detection SNR threshold |
| Hyperparameter | Value |
|---|---|
| Actor hidden size | 256, 256 |
| Critic hidden size | 512, 256, 128 |
| Optimizer | Adam |
| Actor learning rate | |
| Critic learning rate | |
| Discount factor | 0.95 |
| GAE parameter | 0.95 |
| PPO clipping ratio | 0.2 |
| Entropy coefficient | 0.01 |
| Value loss coefficient | 0.5 |
| Batch size | 4096 |
| Mini-batch size | 512 |
| PPO epochs per update | 10 |
| Gradient clipping | 0.5 |
| Training steps | |
| Random seeds | 5 |
| Method | CE ↑ | HVR ↓ | SSR ↑ | TCR ↑ | PUR ↑ | Tracking Error ↓ |
|---|---|---|---|---|---|---|
| MAPPO | ||||||
| P-DQN | ||||||
| Decision Transformer | ||||||
| Decision Mamba | ||||||
| CS-Mamba |
| Method | Small | Medium | Large | Missing Reason | |||
|---|---|---|---|---|---|---|---|
| CE ↑ | Latency ↓ | CE ↑ | Latency ↓ | CE ↑ | Latency ↓ | ||
| EDF | – | ||||||
| RM | – | ||||||
| GA | TO | TO | Timeout | ||||
| Tabu Search | TO | TO | Timeout | ||||
| MILP/ILP | TO | TO | TO | TO | Timeout | ||
| DQN | – | – | – | – | Discrete action only | ||
| PPO | – | – | – | – | – | ||
| MAPPO | – | ||||||
| P-DQN | – | – | – | ||||
| Decision Transformer | – | ||||||
| Decision Mamba | – | – | – | ||||
| CS-Mamba | – | ||||||
| Method | 1-Agent | 2-Agent | 3-Agent | 4-Agent |
|---|---|---|---|---|
| GA | 0.621 | 0.625 | 0.625 | 0.625 |
| PPO | 0.682 | 0.667 | 0.653 | 0.644 |
| A2C | 0.703 | 0.682 | 0.673 | 0.663 |
| SAC | 0.711 | 0.691 | 0.684 | 0.678 |
| TD3 | 0.703 | 0.682 | 0.672 | 0.667 |
| Double DQN | 0.695 | 0.673 | 0.664 | 0.658 |
| Rainbow DQN | 0.706 | 0.683 | 0.674 | 0.668 |
| GNN-RL | 0.736 | 0.718 | 0.702 | 0.696 |
| Transformer-RL | 0.744 | 0.727 | 0.712 | 0.702 |
| Ours | 0.783 | 0.772 | 0.764 | 0.757 |
| Method | Params (M) | FLOPs (G) | Latency (ms/Step) |
|---|---|---|---|
| PPO | 2.8 | 0.62 | 30.1 |
| MAPPO | 4.5 | 1.05 | 34.5 |
| GNN-RL | 5.2 | 1.28 | 34.0 |
| Transformer-RL | 12.6 | 5.42 | 34.5 |
| Ours | 3.2 | 0.87 | 12.8 |
| Variant | Description | Cost-Effectiveness ↑ |
|---|---|---|
| Full CS-Mamba | Complete framework | 0.78 |
| w/o Mamba Encoder | Replace Mamba with LSTM | 0.65 |
| w/o | Remove future return guidance | 0.71 |
| w/o Hierarchical Action | Use direct joint action output | 0.69 |
| w/o Action Mask | Remove hard-constraint masking | 0.66 |
| Short Horizon | Reduce history from 200 to 50 slots | 0.68 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liu, J.; Xu, J.; Xing, W.; Li, M. Long-Horizon Constraint-Aware Collaborative Scheduling for Multiple Phased-Array Radars Using Mamba Temporal Encoding and Structured Hybrid Actions. Sensors 2026, 26, 4772. https://doi.org/10.3390/s26154772
Liu J, Xu J, Xing W, Li M. Long-Horizon Constraint-Aware Collaborative Scheduling for Multiple Phased-Array Radars Using Mamba Temporal Encoding and Structured Hybrid Actions. Sensors. 2026; 26(15):4772. https://doi.org/10.3390/s26154772
Chicago/Turabian StyleLiu, Jianan, Jie Xu, Wenge Xing, and Mingrui Li. 2026. "Long-Horizon Constraint-Aware Collaborative Scheduling for Multiple Phased-Array Radars Using Mamba Temporal Encoding and Structured Hybrid Actions" Sensors 26, no. 15: 4772. https://doi.org/10.3390/s26154772
APA StyleLiu, J., Xu, J., Xing, W., & Li, M. (2026). Long-Horizon Constraint-Aware Collaborative Scheduling for Multiple Phased-Array Radars Using Mamba Temporal Encoding and Structured Hybrid Actions. Sensors, 26(15), 4772. https://doi.org/10.3390/s26154772

