Scarcity-Coefficient Gated Projection Reinforcement Learning for Planning-Layer Capacity Activation in Emergency Wireless Networks
Abstract
1. Introduction
- We formulate the joint planning-layer action and distinguish zero-residual per-period feasibility constraints from explicitly weighted soft planning-risk indicators.
- We develop a compact nine-control action mechanism that couples scarcity-conditioned total activation with a deterministic regional score transformation, scalar projection, capped-simplex projection, and final joint energy rescaling.
- We evaluate constrained performance, component effects, robustness, parameter sensitivity, and computational cost under a common experimental protocol.
2. Related Work
3. Task Modeling and Problem Formulation
3.1. Task Definition
3.2. Regional Wargame-Style Planning-Map Modeling
3.3. Constrained MDP Formulation
4. SCGP-RL
4.1. Capacity Availability and Risk Evidence
4.2. Activation and Upper-Bound Controls
4.3. Coupled Projection and Policy Learning
| Algorithm 1 Scarcity-Coefficient Gated Projection Reinforcement Learning (SCGP-RL) Capacity Activation Procedure. |
| Input: State , policy actor , capacity-bound operators, scarcity evidence functions, and projection operators. Output: Executed feasible capacity action and feedback on urgent-demand relief, unmet demand, and remaining power margin.
|
5. Experiments
5.1. Experimental Setup
5.2. Implementation and Reproducibility
5.3. Comparison Methods and Evaluation Metrics
5.4. Overall Performance Comparison
5.5. Learned Capacity Activation Behavior
5.6. Component Contribution Analysis
5.7. Robustness Under Resource Stress
5.8. Parameter and Map-Size Sensitivity
6. Conclusions and Discussion
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Wang, Q.; Li, W.; Yu, Z.; Abbasi, Q.; Imran, M.; Ansari, S.; Sambo, Y.; Wu, L.; Li, Q.; Zhu, T. An Overview of Emergency Communication Networks. Remote Sens. 2023, 15, 1595. [Google Scholar] [CrossRef] [Scilit]
- Yang, Z.; Barroca, B.; Mebarki, A.; Laffréchine, K.; Dolidon, H.; Lilas, L. Critical Infrastructure Resilience: A Guide for Building Indicator Systems Based on a Multi-Criteria Framework with a Focus on Implementable Actions. Nat. Hazards Earth Syst. Sci. 2024, 24, 3723–3753. [Google Scholar] [CrossRef] [Scilit]
- Zhu, C.; Shi, Y.; Zhao, H.; Chen, K.; Zhang, T.; Bao, C. A Fairness-Enhanced Federated Learning Scheduling Mechanism for UAV-Assisted Emergency Communication. Sensors 2024, 24, 1599. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- 3GPP. NR; Study on Integrated Access and Backhaul; Version 16.0.0, Release 16; Technical Report TR 38.874; 3rd Generation Partnership Project: Sophia Antipolis, France, 2020. [Google Scholar]
- Polese, M.; Giordani, M.; Zugno, T.; Roy, A.; Goyal, S.; Castor, D.; Zorzi, M. Integrated Access and Backhaul in 5G mmWave Networks: Potential and Challenges. IEEE Commun. Mag. 2020, 58, 62–68. [Google Scholar] [CrossRef] [Scilit]
- Madapatha, C.; Makki, B.; Fang, C.; Teyeb, O.; Dahlman, E.; Alouini, M.S.; Svensson, T. On Integrated Access and Backhaul Networks: Current Status and Potentials. IEEE Open J. Commun. Soc. 2020, 1, 1374–1389. [Google Scholar] [CrossRef] [Scilit]
- Auer, G.; Giannini, V.; Desset, C.; Godor, I.; Skillermark, P.; Olsson, M.; Imran, M.A.; Sabella, D.; Gonzalez, M.J.; Blume, O.; et al. How Much Energy Is Needed to Run a Wireless Network? IEEE Wirel. Commun. 2011, 18, 40–49. [Google Scholar] [CrossRef] [Scilit]
- Tan, R.; Shi, Y.; Fan, Y.; Zhu, W.; Wu, T. Energy Saving Technologies and Best Practices for 5G Radio Access Network. IEEE Access 2022, 10, 51747–51756. [Google Scholar] [CrossRef] [Scilit]
- Achiam, J.; Held, D.; Tamar, A.; Abbeel, P. Constrained Policy Optimization. In Proceedings of the International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; Volume 70, pp. 22–31. [Google Scholar]
- Liu, Y.; Ding, J.; Zhang, Z.L.; Liu, X. CLARA: Constrained Reinforcement Learning Based Resource Allocation for Network Slicing. arXiv 2021, arXiv:2111.08397. [Google Scholar] [CrossRef] [Scilit]
- Zhang, L.; Chen, Y.; Takisaka, T.; Zhao, K.; Li, W.; Liu, J. Situational-Constrained Sequential Resources Allocation via Reinforcement Learning. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, Montreal, QC, Canada, 16–22 August 2025; pp. 9121–9129. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Duchi, J.; Shalev-Shwartz, S.; Singer, Y.; Chandra, T. Efficient Projections onto the ℓ1-Ball for Learning in High Dimensions. In Proceedings of the 25th International Conference on Machine Learning, Helsinki, Finland, 5–9 July 2008; pp. 272–279. [Google Scholar] [CrossRef] [Scilit]
- Kim, J.; Jeon, Y.; Lee, J.; Lee, M.S.; Kwon, T. Joint Scheduling and Resource Allocation Based on Reinforcement Learning in Integrated Access and Backhaul Networks. ICT Express 2025, 11, 536–541. [Google Scholar] [CrossRef] [Scilit]
- Perla, P. Wargaming and the Cycle of Research and Learning. Scand. J. Mil. Stud. 2022, 5, 197–208. [Google Scholar] [CrossRef] [Scilit]
- Davis, P.K.; Bracken, P. Artificial Intelligence for Wargaming and Modeling. J. Def. Model. Simul. Appl. Methodol. Technol. 2022, 22, 25–40. [Google Scholar] [CrossRef] [Scilit]
- Banks, D.E. The Methodological Machinery of Wargaming: A Path toward Discovering Wargaming’s Epistemological Foundations. Int. Stud. Rev. 2024, 26, viae002. [Google Scholar] [CrossRef] [Scilit]
- Feng, D.; Jiang, C.; Lim, G.; Cimini, L.J.; Feng, G.; Li, G.Y. A Survey of Energy-Efficient Wireless Communications. IEEE Commun. Surv. Tutor. 2013, 15, 167–178. [Google Scholar] [CrossRef] [Scilit]
- Wu, J.; Zhang, Y.; Zukerman, M.; Yung, E.K.N. Energy-Efficient Base-Stations Sleep-Mode Techniques in Green Cellular Networks: A Survey. IEEE Commun. Surv. Tutor. 2015, 17, 803–826. [Google Scholar] [CrossRef] [Scilit]
- Kaur, P.; Garg, R.; Kukreja, V. Energy-Efficiency Schemes for Base Stations in 5G Heterogeneous Networks: A Systematic Literature Review. Telecommun. Syst. 2023, 84, 115–151. [Google Scholar] [CrossRef] [Scilit]
- Piovesan, N.; López-Pérez, D.; De Domenico, A.; Geng, X.; Bao, H.; Debbah, M. Machine Learning and Analytical Power Consumption Models for 5G Base Stations. IEEE Commun. Mag. 2022, 60, 56–62. [Google Scholar] [CrossRef] [Scilit]
- Ariyoshi, R.; Li, A.; Hasegawa, M.; Ohtsuki, T. Energy-Efficient Resource Allocation Scheme Based on Reinforcement Learning in Distributed LoRa Networks. Sensors 2025, 25, 4996. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Saleh, V.; Eslami, M.; Kazemi, K. DDPG-Based Energy Efficiency Optimization for ABS-Assisted Beyond-5G Cellular Networks with Sleep Mode Management. Front. Commun. Netw. 2026, 6, 1764320. [Google Scholar] [CrossRef] [Scilit]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
- Zhao, W.; He, T.; Chen, R.; Wei, T.; Liu, C. State-Wise Safe Reinforcement Learning: A Survey. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, Macao, China, 19–25 August 2023; pp. 6814–6822. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wachi, A.; Shen, X.; Sui, Y. A Survey of Constraint Formulations in Safe Reinforcement Learning. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, Jeju, Republic of Korea, 3–9 August 2024; pp. 8262–8271. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gu, S.; Yang, L.; Du, Y.; Chen, G.; Walter, F.; Wang, J.; Yang, Y.; Knoll, A. A Review of Safe Reinforcement Learning: Methods, Theories, and Applications. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 11216–11235. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ji, J.; Zhou, J.; Zhang, B.; Dai, J.; Pan, X.; Sun, R.; Huang, W.; Geng, Y.; Liu, M.; Yang, Y. OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research. J. Mach. Learn. Res. 2024, 25, 1–6. [Google Scholar]
- Xu, S.; Liu, Q.; Gong, C.; Wen, X. Energy-Efficient Multi-Agent Deep Reinforcement Learning Task Offloading and Resource Allocation for UAV Edge Computing. Sensors 2025, 25, 3403. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Amos, B.; Kolter, J.Z. OptNet: Differentiable Optimization as a Layer in Neural Networks. In Proceedings of the 34th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2017; Volume 70, pp. 136–145. [Google Scholar]
- Kotary, J.; Fioretto, F.; Van Hentenryck, P.; Wilder, B. End-to-End Constrained Optimization Learning: A Survey. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, Montreal, QC, Canada, 19–27 August 2021; pp. 4475–4482. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2018; Volume 80, pp. 1861–1870. [Google Scholar]
- Fujimoto, S.; van Hoof, H.; Meger, D. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the 35th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2018; Volume 80, pp. 1587–1596. [Google Scholar]
- Song, M.; Zhang, W.; Bai, J. Resource Allocation and Trajectory Planning in Integrated Sensing and Communication Enabled UAV-Assisted Vehicular Network. Sensors 2025, 25, 7295. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, S.; Kishk, M.A.; Alouini, M.S. A Survey on Integrated Access and Backhaul Networks. Front. Commun. Netw. 2021, 2, 647284. [Google Scholar] [CrossRef] [Scilit]
- Abbasalizadeh, M.; Narain, S. Joint Scheduling and Resource Allocation in mmWave IAB Networks Using Deep Reinforcement Learning. arXiv 2025, arXiv:2508.07604. [Google Scholar] [CrossRef] [Scilit]
- Amponis, G.; Lagkas, T.; Zevgara, M.; Katsikas, G.; Xirofotos, T.; Moscholios, I.; Sarigiannidis, P. Drones in B5G/6G Networks as Flying Base Stations. Drones 2022, 6, 39. [Google Scholar] [CrossRef] [Scilit]
- Ghasemi Alavicheh, R.; Razavizadeh, S.M.; Yanikomeroglu, H. Integrated Access and Backhaul (IAB) in Low Altitude Platforms. IEEE Open J. Commun. Soc. 2024, 5, 5890–5904. [Google Scholar] [CrossRef] [Scilit]
- Sarkar, M.; Sahoo, P.K. Leveraging Edge Computing for Video Data Streaming in UAV-Based Emergency Response Systems. Sensors 2024, 24, 5076. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lu, X.; Wang, P.; Niyato, D.; Kim, D.I.; Han, Z. Wireless Networks with RF Energy Harvesting: A Contemporary Survey. IEEE Commun. Surv. Tutor. 2015, 17, 757–789. [Google Scholar] [CrossRef] [Scilit]
- Alsaedi, W.; Ahmadi, H.; Khan, Z.; Grace, D. Spectrum Options and Allocations for 6G: A Regulatory and Standardization Review. IEEE Open J. Commun. Soc. 2023, 4, 1787–1812. [Google Scholar] [CrossRef] [Scilit]








| Item | Setting |
|---|---|
| Observation/executed action | 3811 state entries; common executed with |
| SCGP raw action/network | 9 controls; actor 3811–256–256–9, reward critic 3811–256–256–1; tanh |
| Projected CPO raw action/network | 271 controls (one scalar + 270 logits); actor 3811–256–256–271; separate 256–256 reward and cost critics; tanh |
| Trainable parameters | SCGP-RL 2,085,907; Projected CPO 3,195,424 |
| Training budget | 120,000 environment steps; four parallel environments |
| Learning rate | Linear schedule, |
| PPO rollout/batch/epochs | 256 per environment/128/10 |
| PPO discount/GAE | / |
| PPO clip/target KL | 0.20/0.03 |
| Entropy/value coefficients | 0.001/0.5 |
| Maximum gradient norm | 0.5 |
| CPO rollout/critic batch | 800 per environment (10 complete 80-step episodes); 3200 samples total; batch 400 |
| CPO cost/update | Mean cost budget 0.45; ; cost GAE 0.95; critic learning rate |
| CPO trust region | target KL 0.01; CG max 15; damping 0.10; line-search shrink 0.80/max 10; 10 reward- and 10 cost-critic passes |
| CPO initialization | scalar bias ; regional-logit biases 0; output weights scaled by 0.10 |
| Checkpoint evaluation interval | 5000 environment steps |
| Structured-prior warm-up | 60,000 environment steps |
| Evaluation action | Deterministic Gaussian mean followed by deterministic decoder |
| Reward and hard thresholds | Equation (22); all hard residual thresholds equal zero |
| Item | Setting |
|---|---|
| Software provenance | Python 3.11.10; NumPy 2.1.2; Pandas 3.0.2; Gymnasium 1.2.3; Stable-Baselines3/SB3-Contrib 2.8.0; PyTorch 2.4.0+cu121; Linux 5.4 |
| Environment split | Train 0–4; validation 1101–1110; test 2101–2130 |
| Training environment seeds | For outer seed s, vector environments use |
| Hardware | Dual-socket AMD EPYC 7542 host (128 logical CPUs, approximately 2.0 TiB RAM); NVIDIA GeForce RTX 3090, 24,576 MiB, driver 535.216.03; all training on physical GPU 0 |
| Model records | Architecture-derived parameter counts; checkpoint SHA-256 values and measured file sizes recorded in the run manifests |
| Hardware platform | AMD EPYC 7542 32-Core Processor; 64 physical/128 logical CPU cores; 2004 GiB host RAM; NVIDIA GeForce RTX 3090 (24 GiB), driver 535.216.03; Python 3.11.10 (main, 7 September 2024, 18:35:41) [GCC 11.4.0], PyTorch 2.4.0+cu121 |
| Training time and memory | 120,000 steps in 48.24 min; throughput 41.46 step/s; peak GPU-memory increment 383.0 MiB |
| Batch-one deployment profile | Policy inference 0.425 ms on average and 0.441 ms at p95; complete prediction and decoding 14.617 ms on average and 14.833 ms at p95; peak RSS 795.4 MiB; peak CUDA allocation 33.7 MiB |
| Source code | GitHub repository (v1.0.0) https://github.com/MurrayMa0816/SCGP-RL/tree/v1.0.0 (accessed on 16 August 2026) |
| Scenario | Energy Process | H Mean/Std | High-Priority Demand Setting | ||
|---|---|---|---|---|---|
| Planned charging | deterministic_wave | 12.0 | 2.0 | 3.0/1.0 | base |
| Random recharge | stochastic | 12.0 | 2.0 | 3.0/1.0 | base |
| Power outage | outage | 12.0 | 2.0 | 3.0/1.0 | base |
| Regime shift | regime_shift | 12.0 | 2.0 | state vector | base |
| Low initial energy | outage | 9.0 | 2.0 | 2.6/1.0 | base |
| High energy floor | regime_shift | 12.0 | 4.0 | state vector | base |
| Urgent-demand surge | outage | 11.0 | 3.0 | 2.5/1.1 | min 0.72/width 0.032 |
| Backhaul capacity contraction | regime_shift | 12.0 | 3.0 | state vector | min 0.74/width 0.026 |
| Spectrum contraction/node degradation | stochastic | 8.5 | 3.0 | 2.4/1.3 | min 0.76/width 0.030 |
| Type | Method | Unmet Score | Backlog | Remaining Power | Activated PCU | Low Power Rate |
|---|---|---|---|---|---|---|
| Main | SCGP-RL | |||||
| Learn. | Fixed-Budget PPO | |||||
| Learn. | Direct-Proj. PPO | |||||
| Learn. | Conservative-Proj. PPO | |||||
| Learn. | Projected CPO | |||||
| Rule | Power Margin Guard | |||||
| Rule | Demand-Gated Act. | |||||
| Rule | Power-Urgent Bound | |||||
| Rule | Power-Region Bound | |||||
| Rule | Rolling-Marginal | |||||
| Rule | Lyapunov DPP | |||||
| Rule | Robust Power Guard | |||||
| Rule | High-Margin Guard | |||||
| Rule | Fixed-Act. Rule | |||||
| Rule | Full-Act. Rule | |||||
| Rule | Uniform Bound Rule |
| Comparator to SCGP-RL | Metric | Effect (95% CI) | Raw p | Holm p |
|---|---|---|---|---|
| Projected CPO | Unmet urgent-demand cost | <0.001 | <0.001 | |
| Direct-Projected PPO | Unmet urgent-demand cost | <0.001 | <0.001 | |
| Rolling-Marginal Greedy | Unmet urgent-demand cost | <0.001 | <0.001 |
| Removal | Unmet Score | Loss Exceed. | Low Power Rate | Infeasible Proposal | Executed/Proposed | Relief Gain |
|---|---|---|---|---|---|---|
| w/o Power Margin Planning | ||||||
| w/o Urgent-Demand Gate | ||||||
| w/o Marginal-Relief Probe | ||||||
| w/o Urgency Floor | ||||||
| w/o Planning Backhaul Cap | ||||||
| w/o Node-Health Awareness |
| Disabled Component | Unmet Score | Loss Exceedance | Demand-Relief Gain |
|---|---|---|---|
| Urgent-demand gate | 0.08190 | 0.16877 | |
| Marginal demand-relief estimate | 0.00543 | 0.00725 | |
| Urgency safeguard | 0.00000 | 0.00000 | 0.00000 |
| Method | Service-Window Loss | Backlog | Low Power Rate | Relief Ratio | Safe Scenarios |
|---|---|---|---|---|---|
| SCGP-RL | 0.361 | 0.878 | 0.017 | 98.8% | 100% |
| Projected CPO | 0.719 | 0.963 | 0.073 | 61.5% | 56% |
| Conservative-Projection PPO | 0.407 | 0.920 | 0.066 | 94.7% | 44% |
| Direct-Projected PPO | 0.408 | 0.920 | 0.076 | 94.6% | 44% |
| Fixed-Budget PPO | 0.386 | 0.886 | 0.986 | 100.0% | 0% |
| Demand-Gated Activation | 0.391 | 0.904 | 0.100 | 98.0% | 44% |
| Power Margin Guard | 0.524 | 0.946 | 0.399 | 92.1% | 0% |
| Robust Power Guard | 0.550 | 0.952 | 0.114 | 89.7% | 22% |
| Full-Activation Rule | 0.548 | 0.940 | 0.985 | 92.1% | 0% |
| Uniform Bound Rule | 0.691 | 0.960 | 0.985 | 76.9% | 0% |
| Setting | Method | Unmet-Demand Cost | Low Power Rate | Mean Activated PCU | Cumulative Service Gain |
|---|---|---|---|---|---|
| Canonical control | SCGP-RL | ||||
| Projected CPO | |||||
| Direct-Projected PPO | |||||
| SCGP-RL | |||||
| Projected CPO | |||||
| Direct-Projected PPO | |||||
| SCGP-RL | |||||
| Projected CPO | |||||
| Direct-Projected PPO | |||||
| SCGP-RL | |||||
| Projected CPO | |||||
| Direct-Projected PPO | |||||
| SCGP-RL | |||||
| Projected CPO | |||||
| Direct-Projected PPO | |||||
| Horizon | SCGP-RL | ||||
| Projected CPO | |||||
| Direct-Projected PPO | |||||
| Horizon | SCGP-RL | ||||
| Projected CPO | |||||
| Direct-Projected PPO | |||||
| Activation cap PCU | SCGP-RL | ||||
| Projected CPO | |||||
| Direct-Projected PPO | |||||
| Activation cap PCU | SCGP-RL | ||||
| Projected CPO | |||||
| Direct-Projected PPO | |||||
| Regime replenishment scale | SCGP-RL | ||||
| Projected CPO | |||||
| Direct-Projected PPO | |||||
| Regime replenishment scale | SCGP-RL | ||||
| Projected CPO | |||||
| Direct-Projected PPO |
| Planning Map | Unmet-Demand Cost | Low Power Rate | Mean Activated PCU | Cumulative Service Gain |
|---|---|---|---|---|
| 12 × 15 | ||||
| 15 × 18 | ||||
| 18 × 21 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ma, J.; Liu, P.; Ma, H.; Lu, G.; Zhang, Y. Scarcity-Coefficient Gated Projection Reinforcement Learning for Planning-Layer Capacity Activation in Emergency Wireless Networks. Sensors 2026, 26, 5248. https://doi.org/10.3390/s26165248
Ma J, Liu P, Ma H, Lu G, Zhang Y. Scarcity-Coefficient Gated Projection Reinforcement Learning for Planning-Layer Capacity Activation in Emergency Wireless Networks. Sensors. 2026; 26(16):5248. https://doi.org/10.3390/s26165248
Chicago/Turabian StyleMa, Jingxiang, Ping Liu, Hongbin Ma, Guiping Lu, and Youzhi Zhang. 2026. "Scarcity-Coefficient Gated Projection Reinforcement Learning for Planning-Layer Capacity Activation in Emergency Wireless Networks" Sensors 26, no. 16: 5248. https://doi.org/10.3390/s26165248
APA StyleMa, J., Liu, P., Ma, H., Lu, G., & Zhang, Y. (2026). Scarcity-Coefficient Gated Projection Reinforcement Learning for Planning-Layer Capacity Activation in Emergency Wireless Networks. Sensors, 26(16), 5248. https://doi.org/10.3390/s26165248

