Next Article in Journal
Feature-Level Reliability of Directional-Kernel Richardson–Lucy Deblurring Under Kernel-Length and Direction Controls
Previous Article in Journal
Infrared–Depth Drogue Target Detection via Frequency-Domain Enhancement and Decoupled Gated Fusion
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Scarcity-Coefficient Gated Projection Reinforcement Learning for Planning-Layer Capacity Activation in Emergency Wireless Networks

1
School of Automation, Beijing Institute of Technology, Beijing 100081, China
2
Marine Science and Technology Domain, Beijing Institute of Technology, Zhuhai 519088, China
3
Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences, Hong Kong, China
*
Author to whom correspondence should be addressed.
Sensors 2026, 26(16), 5248; https://doi.org/10.3390/s26165248
Submission received: 1 July 2026 / Revised: 2 August 2026 / Accepted: 17 August 2026 / Published: 19 August 2026

Abstract

After infrastructure disruption, an emergency wireless controller must meet urgent communication demand and preserve resources for later periods. We propose Scarcity-Coefficient Gated Projection Reinforcement Learning (SCGP-RL). It jointly selects total planning-layer activation and a regional capacity upper-bound vector. Scarcity and urgent-demand evidence shape the activation intent. Scalar and capped-simplex projections enforce the coupled action constraints. A planning capacity unit (PCU) is defined as a calibratable service-capacity quantum. In the common constrained evaluation, SCGP-RL reduced the unmet urgent-demand score from 0.6643 for Projected CPO to 0.4919. It also satisfied all the executed hard constraints. Component tests show that the urgent-demand gate drives rapid service response and that the marginal demand-relief estimate provides a smaller benefit. Binding-condition tests show that the power, backhaul, and node-health mechanisms protect the resources they represent.

1. Introduction

Dynamic constrained resource activation arises when a controller must decide how much planning-layer capacity to activate before the future safety margin is fully known. Activating more capacity can satisfy time-sensitive demand in the current period, but it can also reduce the operating margin that is available to later periods; activating too little capacity protects the future while urgent communication opportunities expire. This sequential trade-off is central to emergency wireless restoration, where disrupted access, power, and backhaul support must still sustain command, shelter, and remote-reconnection communications [1,2]. Recent uncrewed aerial vehicle (UAV)-assisted emergency communication scheduling research further shows that temporary aerial support introduces fairness and scheduling constraints at the emergency-network layer [3]. The resulting planning-layer decision is dynamic capacity activation and regional capacity upper-bound allocation.
This decision differs from fixed-budget regional allocation because the total usable capacity is not given before the policy acts. Integrated access and backhaul (IAB) provides a useful high-level abstraction: a surviving donor connects to the core network while temporary nodes provide regional access and wireless backhaul support [4,5,6]. At the planning layer, changing power, backhaul, spectrum, and demand conditions determine how much capacity can be safely activated. A planning capacity unit (PCU) is a normalized service-capacity unit for one planning period rather than a standardized radio-resource unit. Technology-specific calibration translates regional PCU bounds into lower-layer resource or service limits. Activated PCUs expire at the end of each period, whereas energy and capacity risks persist across periods.
Existing studies address individual parts of this decision but rarely treat capacity activation and regional upper-bound formation as one action. Dynamic activation and sleep-control methods decide when infrastructure should be active or how aggressively resources should be admitted, but they often keep the regional allocation structure outside the activation decision [7,8]. Constrained Markov decision process (CMDP) and constrained reinforcement learning (CRL) methods, including constrained policy optimization (CPO), Lagrangian policy-gradient methods, CLARA-style resource allocation, and situational-constrained sequential resource allocation (SCRL), place costs or situation-specific limits into the learning objective or safety layer [9,10,11]. Projection and constrained action-decoding methods can repair fixed-budget regional allocations, while IAB-inspired resource allocation supplies access–backhaul capacity limits [12,13]. These lines of work leave open an action mechanism that learns the total activated capacity and the regional capacity upper-bound vector together.
To address this challenge, we propose Scarcity-Coefficient Gated Projection Reinforcement Learning (SCGP-RL). Its task-specific action factorization controls a state-dependent total activation amount and a regional upper-bound shape. Urgency, remaining-energy, and replenishment-risk evidence condition the scalar intent before projection. The decoder couples the scalar and vector components through i x t , i = B t . It preserves this equality after heterogeneous-energy rescaling. The method therefore extends fixed-budget projection by learning the simplex mass and its regional shape together. Standard PPO trains the policy over the resulting deterministic feasibility map.
The map-level formulation is grounded in analytic wargaming abstractions, where uncertain operational situations are organized as structured synthetic environments for examining decision consequences [14,15,16]. In this paper, that idea is used to build a computational regional planning map for the CMDP.
This paper makes three contributions.
  • We formulate the joint planning-layer action ( B t , x t ) and distinguish zero-residual per-period feasibility constraints from explicitly weighted soft planning-risk indicators.
  • We develop a compact nine-control action mechanism that couples scarcity-conditioned total activation with a deterministic regional score transformation, scalar projection, capped-simplex projection, and final joint energy rescaling.
  • We evaluate constrained performance, component effects, robustness, parameter sensitivity, and computational cost under a common experimental protocol.

2. Related Work

Dynamic activation and admission methods make the active resource amount a decision variable. Classical energy models and base-station sleep-mode studies show that activation state, sleep depth, and traffic-aware operation can dominate wireless energy use [7,17]. Sleep-mode surveys and recent energy-efficiency reviews further show that activation depth, load awareness, and operating mode selection remain central variables in green radio access networks [18,19]. Recent energy-saving and power-consumption studies extend this idea to learned or state-adaptive sleep control under traffic and supply uncertainty [8,20]. Recent reinforcement learning resource allocation in distributed LoRa networks provides another resource-activation example under energy-efficiency constraints [21]. Learning-assisted sleep-mode optimization provides another example in which a controller changes active capacity as part of the resource-allocation action [22]. These methods establish an admission-control view of how much infrastructure or load should be active. Their action spaces usually remain station-level, traffic-level, or link-level, whereas dynamic emergency capacity planning needs a scalar capacity action component that activates a period-level capacity amount and then forms regional capacity upper-bound vectors under future-risk constraints.
Constrained Markov decision process (CMDP) and constrained reinforcement learning (CRL) methods provide the formal basis for safety-aware sequential allocation. Constrained policy optimization (CPO) enforces expected cost constraints during policy improvement, and Lagrangian policy-gradient methods convert constraints into adaptive penalty terms [9,23]. Recent safe-RL studies emphasize state-wise feasibility and implementation-level safety constraints, which are important when admissible capacity varies with current resource state [24,25]. Safe-RL reviews and toolkits also emphasize that empirical safety evaluation should report utility and constraint indicators in addition to return [26,27]. Resource-allocation frameworks such as CLARA and situational-constrained sequential resource allocation (SCRL) show how resource budgets and situation-specific limits can enter the learning objective or feasibility logic [10,11]. Multi-agent deep reinforcement learning for UAV edge offloading and resource allocation similarly highlights state-dependent coupling among computation, energy, and communication resources [28]. These approaches establish the constraint-handling basis, but the activation amount and the regional upper-bound vector still need a task-specific action decoder.
Structured actions, scalar budget heads, and projection layers address how a policy produces feasible allocation actions. Differentiable optimization layers embed solver-like projections into neural policies, and constrained optimization learning studies how hard constraints can be respected by the prediction layer [29,30]. Projection onto the simplex or capped simplex is a standard operation for nonnegative fixed-sum allocation with upper bounds [12]. General continuous-control methods such as soft actor–critic and twin delayed deep deterministic policy gradients are natural candidates for scalar action heads, but they still require an external feasibility map when capacity bounds and regional caps change with state [31,32]. Work on joint resource allocation and trajectory planning in integrated sensing and communication (ISAC)-enabled UAV vehicular networks further illustrates how resource feasibility can be coupled with spatial motion decisions [33]. Capacity-upper-bound allocation adds a harder coupling because the fixed sum is itself part of the action. A direct continuous budget head may choose a scalar activation level, and a separate softmax or simplex head may distribute regional shares, but the two heads can remain inconsistent with energy, backhaul, and regional caps unless scalar projection and capped upper-bound projection are coupled.
Emergency wireless and access–backhaul resource sharing supply the capacity-risk variables behind these constraints. Disaster response studies emphasize communication continuity after conventional infrastructure and adjacent critical systems are impaired [1,2]. Integrated access and backhaul (IAB) research shows that access and backhaul share capacity and can be limited by donor-side backhaul limits, local devices, and temporary-node deployment [4,6]. IAB surveys clarify the access–backhaul coupling that motivates planning-layer backhaul caps, while recent deep reinforcement learning (DRL)-based IAB allocation studies illustrate the more detailed link-scheduling setting from which this paper abstracts [34,35]. Low-altitude platform and drone-base-station studies show that temporary node availability and aerial support can vary over time [36,37]. UAV-based emergency video-streaming research also shows that emergency response quality can depend on the joint availability of communication and edge resources [38]. Radio-frequency (RF) energy-supply and spectrum-management studies separately motivate planning-layer energy and spectrum-opportunity variables [39,40]. These works motivate urgent communication demand pressure, node health, spectrum opportunity, and backhaul caps as planning-layer state and constraint sources. The existing methods therefore establish dynamic admission, constraint learning, and projection-based feasibility, but they typically handle a fixed budget with regional allocation, a scalar admission action without regional upper-bound vectors, a training-objective constraint without a coupled decoder, or an access–backhaul scheduler with a predefined capacity model. A reinforcement learning action mechanism is still needed to learn total capacity activation and regional capacity upper-bound projection as one constrained planning decision.
CPO and adaptive Lagrangian methods control expected cost through policy updates, whereas fixed-budget projection forms a feasible regional vector from an externally supplied total. Direct projection can learn the total and regional controls jointly. SCGP-RL differs by using compact scarcity-conditioned controls to form both quantities before coupled projection. Its contribution is therefore the task-specific action mechanism rather than a new general-purpose constrained policy optimizer.

3. Task Modeling and Problem Formulation

3.1. Task Definition

After a disruptive event, emergency managers may have only a surviving access point and several temporary access–backhaul nodes to keep affected areas connected. These nodes provide planning-layer emergency-connectivity capacity, but the amount that can be safely activated in the current period changes with the remaining power margin, temporary backhaul load, and the urgency of nearby emergency communication demand. The central manager must therefore decide both how much planning-layer capacity to activate now and which regions can safely carry that activated capacity before detailed traffic scheduling begins.
The decision is difficult because present emergency communication demand and future capacity safety pull in different directions. Opening more capacity can help an urgent command post, shelter, or remote reconnection area before its urgent response period closes, but it can also consume operating margin that may be needed if replenishment is delayed or a backhaul path tightens later. Keeping capacity too conservative protects future periods but can let current high-priority communication demand go unmet. The same activation level can also be safe in one region and risky in another because local power margin, relay load, and urgent communication demand pressure differ across the affected area.
A PCU is defined by a planning period Δ t and a reference useful-service quantum Q ref . Thus, x t , i PCUs represent an upper bound of x t , i Q ref useful bits or an average useful-rate bound of x t , i Q ref / Δ t . It is a planning variable rather than realized throughput. A lower-layer scheduler determines the delivered service under channel, interference, protocol, and user constraints.
If W t , i acc is usable access bandwidth, η t , i is effective spectral efficiency, and  C t , i bh is available backhaul rate, the regional cap can be calibrated as
x ¯ t , i acc = η t , i W t , i acc Δ t Q ref , x ¯ t , i bh = C t , i bh Δ t Q ref , x ¯ t , i = min { x ¯ t , i acc , x ¯ t , i bh , x ¯ t , i pow , x ¯ t , i dev } .
The power and device terms are converted to the same service-capacity scale. In LTE or 5G NR, Q ref can be calibrated from usable resource-block time, spectral efficiency, protocol overhead, and scheduling duration. IAB adds donor and relay backhaul limits. Aerial systems also include air-to-ground availability and platform energy. The same interface can use technology-specific capacity models in future emergency-network deployments.

3.2. Regional Wargame-Style Planning-Map Modeling

The regional map operationalizes the analytic wargaming abstraction as a computational planning object [14,15,16]. It exposes capacity activation trade-offs before lower-layer scheduling and defines the cells, overlays, and adjacency relations used by the CMDP; interactive adjudication and experiential play are outside this algorithmic model. To make the planning problem analyzable without reproducing every lower-layer link decision, the affected area is represented as a regional emergency-planning map. The map is tiled into regional cells; each cell denotes an emergency-connectivity support region or a temporary-node coverage area, and neighboring cells are connected by adjacency edges that represent possible emergency-traffic spillover, relay load, or access–backhaul dependence. The surviving donor and temporary-node markers in Figure 1 identify the access and relay resources available to the controller, while the critical-communication and remote-demand markers show why regional capacity cannot be interpreted only as local traffic volume.
The color overlays bind operational variables to the same regional cells. Figure 1a uses warm colors to mark urgent emergency communication demand, where an additional activated unit is most likely to reduce unmet urgent communication demand or accumulated backlog. Figure 1b maps remaining power margin, so low-margin regions can be separated from areas where extra activation is safer. Figure 1c maps backhaul capacity limits and highlights cells whose emergency-communication coverage depends on a constrained relay or donor path. Figure 1d maps the regional capacity limit, the largest planning-layer capacity upper bound that a region can safely offer in the current period.
The planning map therefore preserves the conflict between immediate emergency-demand relief and future capacity safety. A high-demand region in Figure 1a may still receive only a limited regional upper-bound vector when the power margin overlay in Figure 1b or the backhaul overlay in Figure 1c is tight. Conversely, a moderate-demand region can become important when its adjacency relation supports remote cells. The combined colors, markers, edges, and regional-limit overlay represent both sides of the capacity activation decision by showing when opening more capacity is worthwhile and which regions can carry that capacity without violating their caps.

3.3. Constrained MDP Formulation

The regional planning task is modeled as a constrained Markov decision process (CMDP) because current capacity activation changes later unmet emergency demand, energy margin, backhaul pressure, and feasible regional upper-bound vectors. At period t, the state s t contains regional demand and urgent communication demand pressure, accumulated unmet emergency demand, operating energy margin, replenishment forecast features, node availability and health, available spectrum, backhaul capacity, and regional capacity limits. Exogenous disturbances represent power recovery, temporary outage, node degradation, spectrum contraction, and backhaul degradation at the planning layer.
The action is
a t = ( B t , x t ) ,
where B t is the executed total activated planning-layer capacity in the current period and x t = ( x t , 1 , , x t , N ) is the regional capacity upper-bound vector supplied as a regional upper bound to a downstream scheduler. The decoder implementation later distinguishes an intermediate scalar candidate B t cand from the post-filter executed value B t exec ; the CMDP notation writes the feasible executed total compactly as B t . Let B ¯ t denote the state-dependent total activation bound, and let x ¯ t , i denote the tightened regional capacity limit after local node capability, power margin, spectrum opportunity, and backhaul capacity have been converted into planning-layer upper-bound caps. A feasible action must satisfy
0 B t B ¯ t , x t , i 0 , i x t , i = B t , x t , i x ¯ t , i .
When regional activation has heterogeneous energy cost ϵ t , i E , the action also respects the operating-energy floor,
i ϵ t , i E x t , i max { 0 , E t E min } .
The transition updates unmet emergency demand, regional pressure, operating energy, node availability, spectrum opportunity, and backhaul capacity after the capacity action is executed. A compact energy transition is
E t + 1 = clip E t i ϵ t , i E x t , i + H t + 1 , 0 , E max ,
where H t + 1 is realized replenishment after the current action. PCUs expire at the end of the period, while remaining power, equipment condition, and future capacity opportunity persist across periods.
The constraints above define per-period feasibility and are enforced by the action decoder. The corresponding numerical residuals are reported in the experimental evaluation. Low power margin depth, service-window loss, demand debt, and future-capacity risk are normalized planning indicators used in the reward and evaluation.
The one-step reward balances immediate demand relief against activation cost, service loss, unmet-demand accumulation, and reserve risk. Let R t relief denote the reduction in demand risk produced by the executed action and let C t plan collect the normalized planning penalties. The policy objective is
max π θ E π θ t = 0 T 1 γ t r t , r t = R t relief C t plan , ( B t , x t ) A ( s t ) .
The executable reward coefficients and normalization rules are reported in Section 5.2.

4. SCGP-RL

A feasible capacity action must determine both the total capacity activated in the current period and the regional upper-bound vector that bounds how this capacity can be consumed later. Fixed budgets remove the activation decision, and threshold rules react to only a small part of the urgent-demand and remaining-power trade-off. SCGP-RL turns capacity-risk state variables into urgent-demand, power, and backhaul risk evidence, lets the policy generate activation and upper-bound controls, and projects those controls into feasible capacity actions. The information flow in Figure 2 follows this mechanism from capacity, power, backhaul, and high-priority demand states to risk evidence and state-dependent feasible bounds. Policy controls then decide the total capacity activation intent and the regional upper-bound preference. A scalar projection forms an activation candidate, a regional upper-bound projection maps the preference into x t , and the resulting feasible capacity action is executed before urgent-demand relief, unmet demand, and remaining power feedback return to the policy-improvement signal.
The method separates learned intent from feasibility enforcement. Urgent-demand, power, and backhaul risk evidence and the actor controls determine how much capacity the policy wants to activate and how it wants the regional upper-bound field to be shaped. Scalar and capped-simplex projections then act as deterministic constrained decoders that convert those controls into a feasible ( B t , x t ) action. This separation keeps the decoder as a deterministic feasibility map while the policy is evaluated through both utility and capacity safety feedback.

4.1. Capacity Availability and Risk Evidence

SCGP-RL begins by constructing the state-dependent feasibility set. The regional capacity limit x ¯ t , i is the tightened planning-layer cap induced by local radio capability, available spectrum, node health, power margin, and backhaul support. The total activation bound is the capacity amount consistent with energy, device, spectrum, backhaul, and regional upper-bound limits:
B ¯ t = min { B t E , B t d e v , B t s p , B t b h , i x ¯ t , i } .
These bounds make the capacity action state dependent before the policy selects an activation level. Energy, spectrum, node-health, and backhaul limits therefore enter the action space through B ¯ t and x ¯ t as planning-layer feasibility terms before any lower-layer throughput translation.
Capacity-risk evidence describes why the policy should activate more or less capacity within that feasible set. Urgent communication demand pressure, accumulated unmet demand, and regional pressure indicate the cost of under-activation; remaining power margin, replenishment-risk features, node condition, and backhaul pressure indicate the cost of over-activation. We write this evidence compactly as
e t s c a r = h s c a r ( s t , B ¯ t , x ¯ t ) ,
where h s c a r collects capacity availability, urgent communication demand pressure, and future-risk features already represented in the state.
The marginal urgent-demand-relief probe estimates whether a small additional capacity amount would reduce current communication demand pressure under the same local disturbance:
M t s v c = J ( s ¯ t + 1 0 ) J ( s ¯ t + 1 ϵ ) ϵ ,
M t w i n = Φ w i n ( s ¯ t + 1 0 ) Φ w i n ( s ¯ t + 1 ϵ ) ϵ .
Here J ( · ) is the planning-layer communication demand pressure score, Φ w i n ( · ) is unmet urgent communication demand loss, and the two predicted next states compare zero additional activation with a small activation increment. The probe is part of the capacity-risk evidence used by the activation control.

4.2. Activation and Upper-Bound Controls

One policy actor produces nine controls. The first five regulate scarcity-conditioned activation, and the remaining four shape the regional upper-bound weights. Scarcity, gating, regional scoring, and projection are deterministic transformations of the state and policy controls rather than separate neural networks.
Let M t , p t opp , p t dem , p t debt , p t buf , and  p t sup denote the normalized marginal demand-relief estimate, service opportunity, demand pressure, debt, energy-buffer pressure, and replenishment-risk evidence. The executable scalar branch is
V t svc = clip ( 0.50 M t + 0.20 p t opp + 0.12 p t dem + 0.06 p t debt , 0 , 1 ) ,
q t = V t svc + λ t win p t debt λ t eng p t buf λ t rep p t sup ,    
g t = σ ( α t q t + β t ) .    
Thus the learned controls do not replace the deterministic evidence functions; they determine how strongly the current evidence opens or suppresses activation.

4.3. Coupled Projection and Policy Learning

This stage converts the policy controls into a feasible planning action. It first forms a bounded total activation candidate. It then constructs the regional preference vector and projects it onto the state-dependent capped simplex. A final energy filter rescales the scalar and vector components together. The resulting action is used for PPO learning. The activation candidate combines the gate output with a power-safety multiplier and an urgency safeguard:
B ˜ t = E t min { ρ t feas , max ( g t s t safe , ρ t floor ) } , B t cand = Proj [ 0 , B ¯ t ] ( B ˜ t ) .
Here, s t safe decreases activation as the reserve margin narrows, ρ t floor preserves a small urgency response, and  ρ t feas represents the currently feasible activation ratio. The projection enforces the state-dependent total bound.
The implementation does not learn one neural score per region. A deterministic ROI/support function first constructs r t , i 0 from the observed regional features. Four actor outputs then form
u t , i ROI = [ max ( r t , i τ t r t max , 0 ) ] γ t j [ max ( r t , j τ t r t max , 0 ) ] γ t + ε ,  
u t , i soft = exp ( κ t r t , i ) j exp ( κ t r t , j ) , w t , i = ( 1 m t ) u t , i ROI + m t u t , i soft .
The normalized vector w t is the desired regional upper-bound shape before feasibility correction.
The first regional step projects B t cand w t onto the capped simplex:
x t cap = Proj X t ( B t cand ) ( B t cand w t ) , X t ( B t cand ) = { x 0 : i x i = B t cand , x i x ¯ t , i } .
The projection is the Euclidean projection
Proj X t ( B t cand ) ( y ) = arg min x X t ( B t cand ) x y 2 2 .
Because B t cand i x ¯ t , i , the set is nonempty and admits a water-level solution with x i = min { x ¯ t , i , max { 0 , y i τ } } , where τ is chosen so that i x i = B t cand . The capped simplex uses the tightened regional caps defined above; when a spectrum, node-health, or backhaul limit is active, it has already reduced either the scalar bound or the relevant regional cap before this projection.
For heterogeneous energy costs, the executed action applies a final feasibility filter:
η t = min 1 , max { 0 , E t E min } i ϵ t , i E x t , i cap + ε ,
B t exec = η t B t cand , x t = η t x t cap .
When the energy bound has already been fully tightened, η t = 1 ; otherwise this filter rescales the scalar activation and the regional upper-bound vector as one action, preserving nonnegativity, regional caps, and the equality i x t , i = B t exec . Urgent-demand relief, unmet demand, and remaining power feedback enter the fixed shaped reward used by PPO; feasibility itself is provided by the decoder rather than by a constrained policy-gradient update.
SCGP-RL uses an actor and a reward–value function. The actor produces a distribution over the nine policy controls, and the critic estimates the scalar state value. The PPO update uses
L clip ( θ ) = E t min r t ( θ ) A ^ t , clip ( r t ( θ ) , 1 ϵ , 1 + ϵ ) A ^ t ,
and combines the clipped policy loss with value and entropy terms. Network dimensions and optimization settings are reported in Section 5.2. Algorithm 1 summarizes the complete sequence from scarcity evidence construction to feasible action execution and policy feedback.
Algorithm 1 Scarcity-Coefficient Gated Projection Reinforcement Learning (SCGP-RL) Capacity Activation Procedure.
Input: State s t , policy actor π θ , capacity-bound operators, scarcity evidence functions, and projection operators.
Output: Executed feasible capacity action ( B t exec , x t ) and feedback on urgent-demand relief, unmet demand, and remaining power margin.
 1.
for each decision period t do
 2.
Observe capacity, energy, backhaul, high-priority communication demand period, node-health, spectrum, and regional demand states.
 3.
Construct tightened regional caps x ¯ t , i and the scalar activation bound B ¯ t from the capacity-availability operators.
 4.
Build risk evidence e t s c a r from communication demand pressure, remaining power pressure, replenishment risk, backhaul pressure, and the marginal urgent-demand-relief probe.
 5.
Select policy controls ( z t a c t , z t e n v , z t r i s k ) = π θ ( s t , e t s c a r ) for total activation, regional upper-bound shape, and risk conservatism.
 6.
Transform the risk controls into scarcity coefficients and compute the gate score q t and activation gate g t .
 7.
Combine g t with the safe-activation multiplier and feasibility-limited urgency floor to obtain B ˜ t , and then apply scalar projection on [ 0 , B ¯ t ] to obtain the scalar candidate B t cand .
 8.
Score regional upper-bound preferences and normalize them into weights w t .
 9.
Apply capped-simplex upper-bound projection on X t ( B t cand ) to obtain cap-feasible regional upper bounds x t cap .
10.
Apply the final energy-feasibility filter when heterogeneous unit costs require rescaling.
11.
Execute the planning-layer action ( B t exec , x t ) and expose x t as regional capacity caps to the downstream scheduler.
12.
Observe urgent-demand relief, unmet urgent communication demand loss, low power margin cost, and other constraint feedback.
13.
Update the policy with the PPO clipped objective and the reward in Equation (22).
end for
The output of this layer remains a capacity formation interface upstream of downstream scheduling. The executed scalar component B t exec records how much planning-layer capacity is activated in the current period, and the vector x t records regional upper bounds for later use by task, user, or link-level decisions. When the final feasibility filter is inactive, B t exec = B t cand . This separation keeps SCGP-RL focused on dynamic capacity activation and regional capacity upper-bound allocation.

5. Experiments

The experiments address overall constrained performance, learned mechanism behavior, component contribution, stress robustness, parameter sensitivity, and deployment cost. Learned methods use independent training runs, validation-based checkpoint selection, and common final environments; paired inference preserves both training and environment variation.

5.1. Experimental Setup

The experimental environment uses a 15 × 18 regional emergency-planning map and a normalized emergency wireless simulator. Each episode contains time-varying communication demand pressure, energy replenishment, node availability, node health, spectrum availability, and IAB backhaul capacity. The scene in Figure 1 is used as a planning-layer map in which cells represent emergency-connectivity regions or temporary-node coverage areas, adjacency edges represent access or backhaul dependence, marker overlays identify donor, temporary, critical-communication, and remote-demand sites, and color shading indicates planning-layer communication demand pressure or capacity risk.
The canonical benchmark uses a 15 × 18 map, an 80-period horizon, B max = 8 , and energy values E max = 24 , E 0 = 10.5 , and E min = 3.0 . Its replenishment process has normal, tight, reinforcement, and outage states with means ( 3.20 , 1.40 , 5.60 , 0.25 ) , standard deviations ( 0.75 , 0.55 , 0.85 , 0.20 ) , and transition matrix
P H = 0.86 0.10 0.03 0.01 0.18 0.72 0.04 0.06 0.18 0.05 0.72 0.05 0.34 0.30 0.04 0.32 .
These values define the normalized planning scale. Section 5.8 varies the map, horizon, energy ratios, activation limit, and replenishment intensity. The implementation settings are summarized in Table 1, and the exact training, validation, and final-test seed sets are recorded in Table 2. Learned-policy intervals use crossed resampling over training and final-environment axes, and aligned comparisons use paired tests with Holm correction.

5.2. Implementation and Reproducibility

For the canonical map, the observation contains 31 global features and 14 features per region. The global features summarize episode phase, energy, replenishment, aggregate demand, service deficits, node condition, spectrum, and backhaul. The regional features describe demand, service continuity, node condition, radio and backhaul limits, capacity caps, energy cost, spatial position, and regional priority. This gives 3811 entries for 270 regions. Bounded state variables are normalized to [ 0 , 1 ] .
The actor and reward critic use two 256-unit tanh layers. Their dimensions are 3811–256–256–9 and 3811–256–256–1, respectively. The actor parameterizes a nine-dimensional diagonal Gaussian. Its first five outputs control service-window scarcity, energy scarcity, replenishment risk, gate slope, and gate bias. The remaining outputs control the regional exponent, cutoff, softmax sharpness, and branch mixture. Their ranges are [ 0.12 , 1.22 ] , [ 0.55 , 3.75 ] , [ 0.35 , 3.00 ] , [ 0.60 , 8.00 ] , [ 4.80 , 1.20 ] , [ 0.20 , 8.00 ] , [ 0 , 0.95 ] , [ 0.10 , 12.00 ] , and [ 0 , 0.85 ] , respectively.
The executable reward is
r t = tanh ( 8 G ˜ t ) 0.006 B t exec / E max 0.18 δ t pol / E min 0.20 L t + 1 win 0.22 D t + 1 win 0.56 C t + 1 crit 0.16 δ t risk / E max 0.10 I ( E t + 1 10 6 ) 0.08 Δ t shock .
Here, G ˜ t is the clipped demand-risk relief, L t win is service-window loss, D t win is demand debt, and C t crit = J t + 0.22 L t + 0.20 D t + 0.16 [ L t 0.155 ] + . The reserve terms are normalized by the denominators shown above. The dynamic planning margin combines the hard reserve floor with demand pressure, replenishment risk, and the CVaR 0.8 replenishment shortfall. The released source code provides the complete feature order and numerical implementation.
The batch-one deployment profile measures policy prediction and the complete decoder/environment decision path together with host and GPU memory use. SCGP-RL trains in 48.24 min for 120,000 steps. Mean policy-prediction latency is 0.425 ms, and the complete prediction–decoder–environment path takes 14.617 ms with a p95 of 14.833 ms. The measured training GPU-memory increment is 383 MiB, and deployment RSS is 795.4 MiB.
Let N be the number of regions, P the neural-parameter count, and I the fixed projection iteration count. One SCGP-RL decision has complexity O ( P + N 2 + I N + N log N ) . For T environment steps and K PPO epochs, training has complexity O ( K T P + T ( N 2 + I N + N log N ) ) . The measured time and memory values are reported in Table 2.
The stress evaluation varies supply stability, initial reserve, urgent demand, backhaul capacity, spectrum availability, and node condition. These settings create nine scenarios that expose different resource bottlenecks, as summarized in Table 3. All the methods face the same environment realizations within each scenario.

5.3. Comparison Methods and Evaluation Metrics

The comparison set is organized by the decision logic each baseline uses. Always-on full-activation sets B t = B ¯ t , and uniform regional bound spreads the regional upper-bound vector evenly under the regional caps. Fixed-activation RL and fixed-activation rule both use B t = min { 6.0 , B ¯ t } ; the learning version trains the regional upper-bound branch with the same clipped policy-gradient backbone used by the trainable baselines [23], while the rule version applies a fixed regional pattern. Target power margin activation uses E t 🟉 = E min + ( 0.40 0.25 H ¯ t / H max ) E max and then B t = min { B ¯ t , [ E t E t 🟉 ] + / ϵ t m a x } , where H ¯ t / H max is the normalized replenishment forecast and ϵ t m a x = max i ϵ t , i E . Robust power margin activation adds a 0.12 coefficient on the replenishment coefficient of variation and a 0.10 coefficient on outage pressure to that minimum-margin rule, while high power margin rule uses E t 🟉 = E min + 0.48 E max . These power-margin and threshold references adapt energy-aware activation and sleep-control ideas to the planning layer [7,8]. Target power + regional bound and target power + urgent-demand bound incorporate regional value or near-deadline urgent-demand pressure into the activation ratio and regional upper-bound weights. Rolling-horizon marginal activation forms a one-step emergency-demand-relief value from 0.40 urgent-demand opportunity, 0.34 unmet-demand backlog pressure, and 0.26 demand pressure, modulated by remaining power pressure. Lyapunov drift-plus-penalty builds its heuristic action from high-priority demand pressure, replenishment risk, and remaining power queues, while the dual and queue-style references provide deterministic constrained-resource controls.
The implemented trainable projection references are Direct-Projected PPO and Conservative-Projection PPO. Direct-Projected PPO uses a direct scalar activation head with a softmax regional-bound feasibility interface. Conservative-Projection PPO uses ordinary PPO with a direct activation head, the common projection interface, a conservative activation ceiling of 0.90, and fixed reward penalties; it does not use a cost critic, cost budget, multiplier, or dual update. CPO, CLARA, and SCRL provide the constrained-learning background [9,10,11]. Projected CPO provides the direct constrained-RL baseline. It uses the same observation, training budget, evaluation protocol, executed action, and feasibility decoder as SCGP-RL. Its policy outputs one activation control and N regional logits. Its clipped nonnegative cost combines normalized reserve shortfall with weighted service-window loss, demand debt, loss exceedance, and energy depletion. The corresponding weights are 1, 0.22, 0.20, 0.16, and 1. The episode-mean cost budget is 0.45. Rule and planning baselines use current and history-derived information, while component contributions are evaluated through complete removals and evaluation-time disabling.
The reported metrics follow the wireless-capacity formulation. The unmet urgent-demand score is written as C c r i t = T 1 t w ω w χ w c r i t t , w s v c and measures weighted unmet communication demand in high-priority periods. The unmet-demand backlog D ¯ = T 1 t D t s v c measures unmet requirements carried into later periods, while the service-loss exceedance measures how far service-window loss exceeds 0.155. The low power margin period rate v E = T 1 t I ( E t < E t m i n , p l a n ) , mean remaining power, and remaining power margin quantify whether the policy stays above a dynamic planning-layer minimum power margin. The line E t m i n , p l a n defines the planning-layer risk and reward penalty threshold; every policy also satisfies hard per-period activation feasibility through B t E . Unit-energy urgent-demand relief u E = t G t s v c / ( t i ϵ t , i E x t , i + ε ) , cumulative demand-relief gain, mean activated capacity, activated-capacity utilization, backhaul-capacity utilization, and remaining capacity headroom measure whether the policy activates capacity where it helps urgent communication while preserving future capacity. Because the evaluation intentionally keeps persistent demand pressure, the binary high-priority demand-period miss indicator tends to saturate; the interpretation therefore combines service-loss exceedance, unmet-demand backlog, unmet urgent-demand score, low power margin period rate, remaining power margin, and efficiency.

5.4. Overall Performance Comparison

The overall evaluation examines whether SCGP-RL improves urgent-demand service while maintaining the planning power margin and hard action feasibility. Independently trained policies and deterministic rules are evaluated in common held-out environments using unmet urgent demand, demand backlog, remaining power, activated capacity, and low power exposure. As shown in Table 4, SCGP-RL achieves the lowest unmet urgent-demand score while keeping low power exposure close to the most conservative learned policies. Its advantage over Projected CPO is confirmed by the paired analysis in Table 5. The learning curves in Figure 3 show stable checkpoint selection across training runs, and the time-resolved results in Figure 4 show that the gain is concentrated in high-demand periods rather than produced by uniformly increasing activation. The improvement therefore arises from when and where capacity is opened while the planning margin is preserved.

5.5. Learned Capacity Activation Behavior

The behavior analysis examines how demand and resource scarcity shape the learned action. A representative trajectory is selected by its proximity to the median unmet urgent-demand score. Figure 5 shows that the policy limits activation when energy or backhaul becomes scarce and permits larger regional bounds when the operating margin improves. The spatial case in Figure 6 further shows that the regional upper-bound field is localized relative to communication pressure and local capacity limits. A higher value allows the downstream scheduler to use more planning capacity in the corresponding region. The regime summary in Figure 7 links these decisions to the normalized scarcity evidence and separates urgent but safe periods from energy-risk and backhaul-tight periods.

5.6. Component Contribution Analysis

The component study separates adaptive effects from immediate mechanism use. Component-removal policies show whether the remaining controls compensate during learning, while direct disabling isolates the decision effect under the same policy. Table 6 shows that adaptation offsets much of the aggregate impact of removing a demand-guidance component, whereas removing power margin planning increases service-loss exceedance and reduces demand-relief gain. Table 7 shows that the urgent-demand gate is the main demand-response pathway. The marginal demand-relief estimate has a smaller effect, and the urgency safeguard remains inactive in the evaluated trajectories. All the variants retain the common physical execution constraints.

5.7. Robustness Under Resource Stress

The robustness study tests whether the learned capacity policy preserves demand service and resource margins when the dominant bottleneck changes. Nine scenarios vary supply stability, urgent demand, backhaul availability, spectrum availability, and node condition. Table 8 shows that SCGP-RL maintains low service-window loss and low power exposure while retaining near-best demand relief. Figure 8 places it in a favorable region of the service–energy trade-off. Binding-condition tests explain this behavior: removing power margin planning increases reserve shortfall under severe scarcity, while removing the backhaul cap or node-health mechanism produces excess proposals when the corresponding physical limit becomes active.

5.8. Parameter and Map-Size Sensitivity

The sensitivity study examines whether the conclusions depend on individual configuration choices or the planning-map scale. Initial energy, reserve floor, decision horizon, activation ceiling, and replenishment intensity are varied one at a time around the canonical setting. Table 9 shows that SCGP-RL retains its service advantage across these changes, while resource indicators respond in the expected direction to tighter energy conditions. The map-size experiment preserves resource density while changing the number of regions. Table 10 shows a mild increase in unmet demand and activated capacity on the larger map. The operating pattern remains consistent over the tested range, although the result does not imply size-independent performance.

6. Conclusions and Discussion

This study formulates emergency capacity planning as a joint decision over total activation and regional capacity upper bounds. SCGP-RL implements this decision through a compact scarcity-conditioned policy, coupled scalar and capped-simplex projections, and a final feasibility filter. Under the common constraint interface, the method improves urgent-demand service while preserving the planning power margin and executed feasibility. The component results identify the urgent-demand gate as the main response pathway, and the binding-condition tests verify that the power-margin, backhaul, and node-health mechanisms act when their corresponding limits become restrictive.
The results also clarify the role of the proposed design. Its advantage comes from the task-specific factorization of activation amount and regional upper-bound shape before projection. This structure allows the policy to respond to urgent demand without treating all feasible capacity as equally useful, and the robustness and sensitivity studies show that this operating principle persists across the tested conditions. The present validation uses a planning-layer simulator and a calibratable PCU interface. Practical deployment requires technology-specific calibration from bandwidth, spectral efficiency, backhaul availability, power limits, and service measurements. Future work will connect the planning action to LTE or NR scheduling, IAB control, and aerial-platform energy models and will study dynamic regional graphs and online adaptation under new disruption patterns.

Author Contributions

Conceptualization, J.M. and H.M.; methodology, J.M.; software, J.M.; validation, J.M.; formal analysis, J.M.; investigation, J.M.; visualization, J.M. and P.L.; writing—original draft preparation, J.M.; writing—review and editing, J.M., P.L., H.M., G.L. and Y.Z.; supervision, H.M.; project administration, H.M.; algorithmic and experimental discussion, J.M. and Y.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China, grant number 62473052, and the State Key Laboratory Program, grant number 241-HF-D09-01.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The source code is publicly available at https://github.com/MurrayMa0816/SCGP-RL/tree/v1.0.0 (accessed on 16 August 2026).

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the study design, analysis, manuscript preparation, or publication decision.

References

  1. Wang, Q.; Li, W.; Yu, Z.; Abbasi, Q.; Imran, M.; Ansari, S.; Sambo, Y.; Wu, L.; Li, Q.; Zhu, T. An Overview of Emergency Communication Networks. Remote Sens. 2023, 15, 1595. [Google Scholar] [CrossRef] [Scilit]
  2. Yang, Z.; Barroca, B.; Mebarki, A.; Laffréchine, K.; Dolidon, H.; Lilas, L. Critical Infrastructure Resilience: A Guide for Building Indicator Systems Based on a Multi-Criteria Framework with a Focus on Implementable Actions. Nat. Hazards Earth Syst. Sci. 2024, 24, 3723–3753. [Google Scholar] [CrossRef] [Scilit]
  3. Zhu, C.; Shi, Y.; Zhao, H.; Chen, K.; Zhang, T.; Bao, C. A Fairness-Enhanced Federated Learning Scheduling Mechanism for UAV-Assisted Emergency Communication. Sensors 2024, 24, 1599. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. 3GPP. NR; Study on Integrated Access and Backhaul; Version 16.0.0, Release 16; Technical Report TR 38.874; 3rd Generation Partnership Project: Sophia Antipolis, France, 2020. [Google Scholar]
  5. Polese, M.; Giordani, M.; Zugno, T.; Roy, A.; Goyal, S.; Castor, D.; Zorzi, M. Integrated Access and Backhaul in 5G mmWave Networks: Potential and Challenges. IEEE Commun. Mag. 2020, 58, 62–68. [Google Scholar] [CrossRef] [Scilit]
  6. Madapatha, C.; Makki, B.; Fang, C.; Teyeb, O.; Dahlman, E.; Alouini, M.S.; Svensson, T. On Integrated Access and Backhaul Networks: Current Status and Potentials. IEEE Open J. Commun. Soc. 2020, 1, 1374–1389. [Google Scholar] [CrossRef] [Scilit]
  7. Auer, G.; Giannini, V.; Desset, C.; Godor, I.; Skillermark, P.; Olsson, M.; Imran, M.A.; Sabella, D.; Gonzalez, M.J.; Blume, O.; et al. How Much Energy Is Needed to Run a Wireless Network? IEEE Wirel. Commun. 2011, 18, 40–49. [Google Scholar] [CrossRef] [Scilit]
  8. Tan, R.; Shi, Y.; Fan, Y.; Zhu, W.; Wu, T. Energy Saving Technologies and Best Practices for 5G Radio Access Network. IEEE Access 2022, 10, 51747–51756. [Google Scholar] [CrossRef] [Scilit]
  9. Achiam, J.; Held, D.; Tamar, A.; Abbeel, P. Constrained Policy Optimization. In Proceedings of the International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; Volume 70, pp. 22–31. [Google Scholar]
  10. Liu, Y.; Ding, J.; Zhang, Z.L.; Liu, X. CLARA: Constrained Reinforcement Learning Based Resource Allocation for Network Slicing. arXiv 2021, arXiv:2111.08397. [Google Scholar] [CrossRef] [Scilit]
  11. Zhang, L.; Chen, Y.; Takisaka, T.; Zhao, K.; Li, W.; Liu, J. Situational-Constrained Sequential Resources Allocation via Reinforcement Learning. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, Montreal, QC, Canada, 16–22 August 2025; pp. 9121–9129. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Duchi, J.; Shalev-Shwartz, S.; Singer, Y.; Chandra, T. Efficient Projections onto the 1-Ball for Learning in High Dimensions. In Proceedings of the 25th International Conference on Machine Learning, Helsinki, Finland, 5–9 July 2008; pp. 272–279. [Google Scholar] [CrossRef] [Scilit]
  13. Kim, J.; Jeon, Y.; Lee, J.; Lee, M.S.; Kwon, T. Joint Scheduling and Resource Allocation Based on Reinforcement Learning in Integrated Access and Backhaul Networks. ICT Express 2025, 11, 536–541. [Google Scholar] [CrossRef] [Scilit]
  14. Perla, P. Wargaming and the Cycle of Research and Learning. Scand. J. Mil. Stud. 2022, 5, 197–208. [Google Scholar] [CrossRef] [Scilit]
  15. Davis, P.K.; Bracken, P. Artificial Intelligence for Wargaming and Modeling. J. Def. Model. Simul. Appl. Methodol. Technol. 2022, 22, 25–40. [Google Scholar] [CrossRef] [Scilit]
  16. Banks, D.E. The Methodological Machinery of Wargaming: A Path toward Discovering Wargaming’s Epistemological Foundations. Int. Stud. Rev. 2024, 26, viae002. [Google Scholar] [CrossRef] [Scilit]
  17. Feng, D.; Jiang, C.; Lim, G.; Cimini, L.J.; Feng, G.; Li, G.Y. A Survey of Energy-Efficient Wireless Communications. IEEE Commun. Surv. Tutor. 2013, 15, 167–178. [Google Scholar] [CrossRef] [Scilit]
  18. Wu, J.; Zhang, Y.; Zukerman, M.; Yung, E.K.N. Energy-Efficient Base-Stations Sleep-Mode Techniques in Green Cellular Networks: A Survey. IEEE Commun. Surv. Tutor. 2015, 17, 803–826. [Google Scholar] [CrossRef] [Scilit]
  19. Kaur, P.; Garg, R.; Kukreja, V. Energy-Efficiency Schemes for Base Stations in 5G Heterogeneous Networks: A Systematic Literature Review. Telecommun. Syst. 2023, 84, 115–151. [Google Scholar] [CrossRef] [Scilit]
  20. Piovesan, N.; López-Pérez, D.; De Domenico, A.; Geng, X.; Bao, H.; Debbah, M. Machine Learning and Analytical Power Consumption Models for 5G Base Stations. IEEE Commun. Mag. 2022, 60, 56–62. [Google Scholar] [CrossRef] [Scilit]
  21. Ariyoshi, R.; Li, A.; Hasegawa, M.; Ohtsuki, T. Energy-Efficient Resource Allocation Scheme Based on Reinforcement Learning in Distributed LoRa Networks. Sensors 2025, 25, 4996. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Saleh, V.; Eslami, M.; Kazemi, K. DDPG-Based Energy Efficiency Optimization for ABS-Assisted Beyond-5G Cellular Networks with Sleep Mode Management. Front. Commun. Netw. 2026, 6, 1764320. [Google Scholar] [CrossRef] [Scilit]
  23. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
  24. Zhao, W.; He, T.; Chen, R.; Wei, T.; Liu, C. State-Wise Safe Reinforcement Learning: A Survey. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, Macao, China, 19–25 August 2023; pp. 6814–6822. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Wachi, A.; Shen, X.; Sui, Y. A Survey of Constraint Formulations in Safe Reinforcement Learning. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, Jeju, Republic of Korea, 3–9 August 2024; pp. 8262–8271. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Gu, S.; Yang, L.; Du, Y.; Chen, G.; Walter, F.; Wang, J.; Yang, Y.; Knoll, A. A Review of Safe Reinforcement Learning: Methods, Theories, and Applications. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 11216–11235. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Ji, J.; Zhou, J.; Zhang, B.; Dai, J.; Pan, X.; Sun, R.; Huang, W.; Geng, Y.; Liu, M.; Yang, Y. OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research. J. Mach. Learn. Res. 2024, 25, 1–6. [Google Scholar]
  28. Xu, S.; Liu, Q.; Gong, C.; Wen, X. Energy-Efficient Multi-Agent Deep Reinforcement Learning Task Offloading and Resource Allocation for UAV Edge Computing. Sensors 2025, 25, 3403. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Amos, B.; Kolter, J.Z. OptNet: Differentiable Optimization as a Layer in Neural Networks. In Proceedings of the 34th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2017; Volume 70, pp. 136–145. [Google Scholar]
  30. Kotary, J.; Fioretto, F.; Van Hentenryck, P.; Wilder, B. End-to-End Constrained Optimization Learning: A Survey. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, Montreal, QC, Canada, 19–27 August 2021; pp. 4475–4482. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2018; Volume 80, pp. 1861–1870. [Google Scholar]
  32. Fujimoto, S.; van Hoof, H.; Meger, D. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the 35th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2018; Volume 80, pp. 1587–1596. [Google Scholar]
  33. Song, M.; Zhang, W.; Bai, J. Resource Allocation and Trajectory Planning in Integrated Sensing and Communication Enabled UAV-Assisted Vehicular Network. Sensors 2025, 25, 7295. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Zhang, S.; Kishk, M.A.; Alouini, M.S. A Survey on Integrated Access and Backhaul Networks. Front. Commun. Netw. 2021, 2, 647284. [Google Scholar] [CrossRef] [Scilit]
  35. Abbasalizadeh, M.; Narain, S. Joint Scheduling and Resource Allocation in mmWave IAB Networks Using Deep Reinforcement Learning. arXiv 2025, arXiv:2508.07604. [Google Scholar] [CrossRef] [Scilit]
  36. Amponis, G.; Lagkas, T.; Zevgara, M.; Katsikas, G.; Xirofotos, T.; Moscholios, I.; Sarigiannidis, P. Drones in B5G/6G Networks as Flying Base Stations. Drones 2022, 6, 39. [Google Scholar] [CrossRef] [Scilit]
  37. Ghasemi Alavicheh, R.; Razavizadeh, S.M.; Yanikomeroglu, H. Integrated Access and Backhaul (IAB) in Low Altitude Platforms. IEEE Open J. Commun. Soc. 2024, 5, 5890–5904. [Google Scholar] [CrossRef] [Scilit]
  38. Sarkar, M.; Sahoo, P.K. Leveraging Edge Computing for Video Data Streaming in UAV-Based Emergency Response Systems. Sensors 2024, 24, 5076. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Lu, X.; Wang, P.; Niyato, D.; Kim, D.I.; Han, Z. Wireless Networks with RF Energy Harvesting: A Contemporary Survey. IEEE Commun. Surv. Tutor. 2015, 17, 757–789. [Google Scholar] [CrossRef] [Scilit]
  40. Alsaedi, W.; Ahmadi, H.; Khan, Z.; Grace, D. Spectrum Options and Allocations for 6G: A Regulatory and Standardization Review. IEEE Open J. Commun. Soc. 2023, 4, 1787–1812. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Regional emergency-planning map for dynamic capacity activation. The four panels encode (a) urgent emergency communication demand, (b) remaining power margin, (c) backhaul capacity limits, and (d) regional capacity limits. The shared marker legend below the panels identifies the surviving donor, temporary access–backhaul nodes, critical-communication areas, and remote-demand areas; blue and green edges denote backhaul and access-dependence links.
Figure 1. Regional emergency-planning map for dynamic capacity activation. The four panels encode (a) urgent emergency communication demand, (b) remaining power margin, (c) backhaul capacity limits, and (d) regional capacity limits. The shared marker legend below the panels identifies the surviving donor, temporary access–backhaul nodes, critical-communication areas, and remote-demand areas; blue and green edges denote backhaul and access-dependence links.
Sensors 26 05248 g001
Figure 2. Scarcity-Coefficient Gated Projection Reinforcement Learning (SCGP-RL) action-generation core. The orange frame marks the proposed SCGP-RL mechanism, while planning-state inputs and executed-action feedback are external interfaces. Inside the core, capacity-risk evidence feeds the policy-control branch; the policy supplies scalar activation and upper-bound controls, so SCGP-RL generates the activated capacity as part of the action; the constrained decoder maps the scalar candidate and regional weights to the executed feasible capacity interface.
Figure 2. Scarcity-Coefficient Gated Projection Reinforcement Learning (SCGP-RL) action-generation core. The orange frame marks the proposed SCGP-RL mechanism, while planning-state inputs and executed-action feedback are external interfaces. Inside the core, capacity-risk evidence feeds the policy-control branch; the policy supplies scalar activation and upper-bound controls, so SCGP-RL generates the activated capacity as part of the action; the constrained decoder maps the scalar candidate and regional weights to the executed feasible capacity interface.
Sensors 26 05248 g002
Figure 3. Validation learning curves used for checkpoint selection. Lines show the mean across training runs, and bands show one standard deviation.
Figure 3. Validation learning curves used for checkpoint selection. Lines show the mean across training runs, and bands show one standard deviation.
Sensors 26 05248 g003
Figure 4. Time-resolved performance in the common held-out trajectories. The panels show the planning objective, unmet urgent demand, low power margin rate, and unmet-demand backlog. Shading marks high-demand periods.
Figure 4. Time-resolved performance in the common held-out trajectories. The panels show the planning objective, unmet urgent demand, low power margin rate, and unmet-demand backlog. Shading marks high-demand periods.
Sensors 26 05248 g004
Figure 5. Representative SCGP-RL trajectory showing activation, urgent demand, remaining energy, and capacity-limit pressure.
Figure 5. Representative SCGP-RL trajectory showing activation, urgent demand, remaining energy, and capacity-limit pressure.
Sensors 26 05248 g005
Figure 6. Spatial capacity-envelope example showing urgent demand, regional caps, allocated upper bounds, and cap utilization. Blue and green lines denote backhaul and access-dependence links, respectively.
Figure 6. Spatial capacity-envelope example showing urgent demand, regional caps, allocated upper bounds, and cap utilization. Blue and green lines denote backhaul and access-dependence links, respectively.
Sensors 26 05248 g006
Figure 7. Activation behavior and normalized scarcity evidence across resource-risk regimes. In (a), the vertical dashed line marks the 68th percentile of the normalized urgent-demand signal, and the horizontal dotted line marks the median activated-capacity ratio.
Figure 7. Activation behavior and normalized scarcity evidence across resource-risk regimes. In (a), the vertical dashed line marks the 68th percentile of the normalized urgent-demand signal, and the horizontal dotted line marks the median activated-capacity ratio.
Sensors 26 05248 g007
Figure 8. Energy and urgent-demand trade-off across the stress scenarios. Bubble size and color encode the urgent-demand relief ratio. The star marks the SCGP-RL reference point, and the dashed lines show its coordinate projections.
Figure 8. Energy and urgent-demand trade-off across the stress scenarios. Bubble size and color encode the urgent-demand relief ratio. The star marks the SCGP-RL reference point, and the dashed lines show its coordinate projections.
Sensors 26 05248 g008
Table 1. PPO and Projected CPO implementation settings.
Table 1. PPO and Projected CPO implementation settings.
ItemSetting
Observation/executed action3811 state entries; common executed ( B t , x t ) with N = 270
SCGP raw action/network9 controls; actor 3811–256–256–9, reward critic 3811–256–256–1; tanh
Projected CPO raw action/network271 controls (one scalar + 270 logits); actor 3811–256–256–271; separate 256–256 reward and cost critics; tanh
Trainable parametersSCGP-RL 2,085,907; Projected CPO 3,195,424
Training budget120,000 environment steps; four parallel environments
Learning rateLinear schedule, 2 × 10 4 0
PPO rollout/batch/epochs256 per environment/128/10
PPO discount/GAE γ = 0.99 / λ GAE = 0.95
PPO clip/target KL0.20/0.03
Entropy/value coefficients0.001/0.5
Maximum gradient norm0.5
CPO rollout/critic batch800 per environment (10 complete 80-step episodes); 3200 samples total; batch 400
CPO cost/updateMean cost budget 0.45; γ c = 1.0 ; cost GAE 0.95; critic learning rate 2 × 10 4
CPO trust regiontarget KL 0.01; CG max 15; damping 0.10; line-search shrink 0.80/max 10; 10 reward- and 10 cost-critic passes
CPO initializationscalar bias logit ( 0.10 ) = 2.1972246 ; regional-logit biases 0; output weights scaled by 0.10
Checkpoint evaluation interval5000 environment steps
Structured-prior warm-up60,000 environment steps
Evaluation actionDeterministic Gaussian mean followed by deterministic decoder
Reward and hard thresholdsEquation (22); all hard residual thresholds equal zero
Table 2. Reproducibility and computational measurements.
Table 2. Reproducibility and computational measurements.
ItemSetting
Software provenancePython 3.11.10; NumPy 2.1.2; Pandas 3.0.2; Gymnasium 1.2.3; Stable-Baselines3/SB3-Contrib 2.8.0; PyTorch 2.4.0+cu121; Linux 5.4
Environment splitTrain 0–4; validation 1101–1110; test 2101–2130
Training environment seedsFor outer seed s, vector environments use s + { 0 , 1009 , 2018 , 3027 }
HardwareDual-socket AMD EPYC 7542 host (128 logical CPUs, approximately 2.0 TiB RAM); NVIDIA GeForce RTX 3090, 24,576 MiB, driver 535.216.03; all training on physical GPU 0
Model recordsArchitecture-derived parameter counts; checkpoint SHA-256 values and measured file sizes recorded in the run manifests
Hardware platformAMD EPYC 7542 32-Core Processor; 64 physical/128 logical CPU cores; 2004 GiB host RAM; NVIDIA GeForce RTX 3090 (24 GiB), driver 535.216.03; Python 3.11.10 (main, 7 September 2024, 18:35:41) [GCC 11.4.0], PyTorch 2.4.0+cu121
Training time and memory120,000 steps in 48.24 min; throughput 41.46 step/s; peak GPU-memory increment 383.0 MiB
Batch-one deployment profilePolicy inference 0.425 ms on average and 0.441 ms at p95; complete prediction and decoding 14.617 ms on average and 14.833 ms at p95; peak RSS 795.4 MiB; peak CUDA allocation 33.7 MiB
Source codeGitHub repository (v1.0.0) https://github.com/MurrayMa0816/SCGP-RL/tree/v1.0.0 (accessed on 16 August 2026)
Table 3. Parameters of the nine stress scenarios.
Table 3. Parameters of the nine stress scenarios.
ScenarioEnergy Process E 0 E min H Mean/StdHigh-Priority Demand Setting
Planned chargingdeterministic_wave12.02.03.0/1.0base
Random rechargestochastic12.02.03.0/1.0base
Power outageoutage12.02.03.0/1.0base
Regime shiftregime_shift12.02.0state vectorbase
Low initial energyoutage9.02.02.6/1.0base
High energy floorregime_shift12.04.0state vectorbase
Urgent-demand surgeoutage11.03.02.5/1.1min 0.72/width 0.032
Backhaul capacity contractionregime_shift12.03.0state vectormin 0.74/width 0.026
Spectrum contraction/node degradationstochastic8.53.02.4/1.3min 0.76/width 0.030
Table 4. Overall performance in the common held-out environments.
Table 4. Overall performance in the common held-out environments.
TypeMethodUnmet ScoreBacklogRemaining PowerActivated PCULow Power Rate
MainSCGP-RL 0.492 0.866 15.959 2.559 0.013
Learn.Fixed-Budget PPO 0.504 0.881 5.886 2.670 1.000
Learn.Direct-Proj. PPO 0.521 0.916 19.692 2.448 0.012
Learn.Conservative-Proj. PPO 0.522 0.917 19.916 2.430 0.013
Learn.Projected CPO 0.664 0.963 21.472 2.091 0.007
RulePower Margin Guard 0.573 0.945 13.206 2.619 0.414
RuleDemand-Gated Act. 0.510 0.898 13.567 2.600 0.082
RulePower-Urgent Bound 0.508 0.895 13.183 2.599 0.478
RulePower-Region Bound 0.528 0.920 13.175 2.592 0.490
RuleRolling-Marginal 0.515 0.901 13.148 2.607 0.128
RuleLyapunov DPP 0.515 0.901 13.023 2.608 0.295
RuleRobust Power Guard 0.585 0.951 14.789 2.606 0.055
RuleHigh-Margin Guard 0.617 0.960 17.120 2.588 0.000
RuleFixed-Act. Rule 0.587 0.942 5.889 2.711 1.000
RuleFull-Act. Rule 0.581 0.939 5.867 2.708 1.000
RuleUniform Bound Rule 0.646 0.960 5.867 2.732 1.000
Values are means over the common held-out evaluation set.
Table 5. Paired comparisons between SCGP-RL and the principal baselines.
Table 5. Paired comparisons between SCGP-RL and the principal baselines.
Comparator to SCGP-RLMetricEffect (95% CI)Raw pHolm p
Projected CPOUnmet urgent-demand cost + 0.172 [ + 0.167 , + 0.178 ] <0.001<0.001
Direct-Projected PPOUnmet urgent-demand cost + 0.029 [ + 0.027 , + 0.033 ] <0.001<0.001
Rolling-Marginal GreedyUnmet urgent-demand cost + 0.023 [ + 0.022 , + 0.024 ] <0.001<0.001
Positive effects favor SCGP-RL. The reported p-values use paired tests with Holm correction.
Table 6. Effects of component removal relative to complete SCGP-RL.
Table 6. Effects of component removal relative to complete SCGP-RL.
Removal Δ Unmet Score Δ Loss Exceed. Δ Low Power Rate Δ Infeasible Proposal Δ Executed/Proposed Δ Relief Gain
w/o Power Margin Planning + 0.014 + 0.015 + 0.019 + 0.000 0.000 0.009
w/o Urgent-Demand Gate + 0.000 + 0.000 + 0.000 + 0.000 0.000 + 0.000
w/o Marginal-Relief Probe 0.000 0.000 0.000 0.000 + 0.000 0.000
w/o Urgency Floor 0.000 0.000 0.000 0.000 + 0.000 0.000
w/o Planning Backhaul Cap 0.000 0.000 0.000 + 0.006 0.000 0.000
w/o Node-Health Awareness 0.000 0.001 + 0.000 + 0.000 0.000 0.000
Table 7. Immediate effects of disabling individual components.
Table 7. Immediate effects of disabling individual components.
Disabled Component Δ Unmet Score Δ Loss Exceedance Δ Demand-Relief Gain
Urgent-demand gate0.081900.16877 0.06642
Marginal demand-relief estimate0.005430.00725 0.00295
Urgency safeguard0.000000.000000.00000
Table 8. Performance across nine resource and demand stress scenarios.
Table 8. Performance across nine resource and demand stress scenarios.
MethodService-Window LossBacklogLow Power RateRelief RatioSafe Scenarios
SCGP-RL0.3610.8780.01798.8%100%
Projected CPO0.7190.9630.07361.5%56%
Conservative-Projection PPO0.4070.9200.06694.7%44%
Direct-Projected PPO0.4080.9200.07694.6%44%
Fixed-Budget PPO0.3860.8860.986100.0%0%
Demand-Gated Activation0.3910.9040.10098.0%44%
Power Margin Guard0.5240.9460.39992.1%0%
Robust Power Guard0.5500.9520.11489.7%22%
Full-Activation Rule0.5480.9400.98592.1%0%
Uniform Bound Rule0.6910.9600.98576.9%0%
Table 9. One-factor-at-a-time sensitivity results.
Table 9. One-factor-at-a-time sensitivity results.
SettingMethodUnmet-Demand CostLow Power RateMean Activated PCUCumulative Service Gain
Canonical controlSCGP-RL 0.492 0.013 2.559 0.142
Projected CPO 0.664 0.007 2.091 0.085
Direct-Projected PPO 0.521 0.012 2.448 0.137
E 0 / E max = 0.35 SCGP-RL 0.495 0.025 2.535 0.141
Projected CPO 0.665 0.024 2.071 0.085
Direct-Projected PPO 0.523 0.030 2.426 0.136
E 0 / E max = 0.525 SCGP-RL 0.489 0.011 2.583 0.143
Projected CPO 0.664 0.004 2.109 0.086
Direct-Projected PPO 0.519 0.008 2.470 0.138
E min / E max = 0.08 SCGP-RL 0.491 0.012 2.571 0.143
Projected CPO 0.664 0.002 2.091 0.085
Direct-Projected PPO 0.521 0.005 2.448 0.137
E min / E max = 0.17 SCGP-RL 0.494 0.013 2.546 0.142
Projected CPO 0.664 0.031 2.091 0.085
Direct-Projected PPO 0.521 0.040 2.448 0.137
Horizon T = 60 SCGP-RL 0.482 0.014 2.503 0.104
Projected CPO 0.646 0.010 2.052 0.062
Direct-Projected PPO 0.511 0.012 2.347 0.098
Horizon T = 100 SCGP-RL 0.499 0.011 2.570 0.180
Projected CPO 0.675 0.005 2.107 0.109
Direct-Projected PPO 0.528 0.020 2.488 0.175
Activation cap = 6 PCUSCGP-RL 0.491 0.014 2.555 0.142
Projected CPO 0.664 0.002 2.090 0.085
Direct-Projected PPO 0.519 0.006 2.503 0.139
Activation cap = 10 PCUSCGP-RL 0.494 0.013 2.539 0.141
Projected CPO 0.664 0.038 2.091 0.085
Direct-Projected PPO 0.524 0.041 2.392 0.135
Regime replenishment scale = 0.80 SCGP-RL 0.520 0.012 2.032 0.118
Projected CPO 0.668 0.073 1.887 0.077
Direct-Projected PPO 0.539 0.120 2.033 0.118
Regime replenishment scale = 1.20 SCGP-RL 0.463 0.011 3.083 0.165
Projected CPO 0.663 0.002 2.181 0.089
Direct-Projected PPO 0.508 0.003 2.775 0.151
Each setting changes one benchmark factor. Cumulative service gain is interpreted together with the episode horizon.
Table 10. SCGP-RL sensitivity to planning-map size.
Table 10. SCGP-RL sensitivity to planning-map size.
Planning MapUnmet-Demand CostLow Power RateMean Activated PCUCumulative Service Gain
12 × 15 0.483 0.007 1.701 0.141
15 × 18 0.492 0.013 2.559 0.142
18 × 21 0.499 0.008 3.568 0.143
Total resources scale with the number of regions to preserve resource density.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ma, J.; Liu, P.; Ma, H.; Lu, G.; Zhang, Y. Scarcity-Coefficient Gated Projection Reinforcement Learning for Planning-Layer Capacity Activation in Emergency Wireless Networks. Sensors 2026, 26, 5248. https://doi.org/10.3390/s26165248

AMA Style

Ma J, Liu P, Ma H, Lu G, Zhang Y. Scarcity-Coefficient Gated Projection Reinforcement Learning for Planning-Layer Capacity Activation in Emergency Wireless Networks. Sensors. 2026; 26(16):5248. https://doi.org/10.3390/s26165248

Chicago/Turabian Style

Ma, Jingxiang, Ping Liu, Hongbin Ma, Guiping Lu, and Youzhi Zhang. 2026. "Scarcity-Coefficient Gated Projection Reinforcement Learning for Planning-Layer Capacity Activation in Emergency Wireless Networks" Sensors 26, no. 16: 5248. https://doi.org/10.3390/s26165248

APA Style

Ma, J., Liu, P., Ma, H., Lu, G., & Zhang, Y. (2026). Scarcity-Coefficient Gated Projection Reinforcement Learning for Planning-Layer Capacity Activation in Emergency Wireless Networks. Sensors, 26(16), 5248. https://doi.org/10.3390/s26165248

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop