1. Introduction
In recent years, the combination of high proportions of renewable energy being fed into the grid, a decline in system inertia and frequent extreme weather events has led to a continuous narrowing of the safety margins for power system operations. Local disturbances are now more likely to escalate into widespread blackouts under the combined effects of power flow shifts, protection system operations and deteriorating operational conditions. The 2023 Pakistan blackout, the 2021 Texas cold snap blackout, the 2021 European grid separation, and the 2019 United Kingdom power outage demonstrate that modern power grids exhibit greater vulnerability under complex disturbances [
1,
2,
3,
4]. Existing accident analyses and review studies have indicated that when the system is under heavy load, weak support, or extreme external forcing conditions, initial faults are more likely to propagate along critical sections and trigger cascading consequences [
5,
6]. Consequently, research into the mechanisms of cascading fault propagation under extreme scenarios, the identification of critical vulnerabilities, and proactive defense methods has become a key issue in the safe operation of new-generation power systems.
Existing research on the mechanisms of cascading fault propagation and vulnerability identification can be broadly categorized into three types. The first category of methods is based on complex network theory and utilizes metrics such as node centrality, line degree and network connectivity to identify system vulnerabilities [
7,
8]. These methods are computationally efficient and suitable for the rapid screening of large-scale systems; however, they do not adequately account for operational states or the temporal progression of fault propagation. The second category of methods takes an electrical perspective, employing cascade failure models, N-k safety analysis and fault scenario simulation to characterize the evolution of faults under power flow redistribution [
9,
10,
11,
12]. Compared with topological analysis, these methods better capture fault propagation paths and the influence of critical components; however, many models still rely on static out-of-limit criteria and do not sufficiently account for thermal accumulation in lines or the sequential approach of faults under extreme conditions. Overall, existing research has laid the groundwork for the identification of cascading faults, but there remains a lack of analytical frameworks capable of simultaneously reflecting extreme operating scenarios, thermal accumulation effects, and the dynamic migration characteristics of critical vulnerable branches.
Existing research on the role of energy storage in enhancing grid resilience and the allocation of defense resources has largely focused on two aspects. One category of studies aims to optimize economic operation and the integration of renewable energy, analyzing the role of energy storage in peak shaving and valley filling, improving the utilization rate of renewable energy, and reducing operational costs [
13,
14]. The other category of research discusses energy storage deployment from the perspectives of disturbance resilience and system security, highlighting that energy storage can utilize its rapid response capabilities to enhance short-term system security and play a role in supporting vulnerable areas [
15,
16,
17,
18]. These studies demonstrate that energy storage is no longer merely a conventional regulating resource, but is increasingly being utilized to enhance system resilience. However, most existing methods focus on single functions, with greater attention paid to active power compensation, frequency support or conventional configuration optimization. Consideration of scenarios where local disconnection, power imbalance and insufficient reactive power support coexist under cascading fault conditions remains inadequate, and there is a relative lack of regionalized configuration approaches that account for the evolution of faults.
Beyond the above resilience-oriented planning and support studies, recent research has also emphasized converter-level coordination and distributed control in hybrid AC/DC architectures. For instance, regarding distributed secondary control and communication-aware coordination, Espina et al. [
19] and Yang et al. [
20] developed consensus-based frameworks to achieve voltage/frequency restoration while reducing communication burdens. In parallel, from the perspective of coordination-based power management, Salman et al. [
21] and Mahmoudian et al. [
22] proposed adaptive strategies centered on interlinking converters, enabling coordinated active/reactive power sharing across AC and DC subgrids under stressed operating conditions. Although these studies are mainly oriented to hybrid AC/DC microgrids and converter-level coordination, they provide useful insights into distributed power sharing, coordinated voltage/frequency regulation, and communication-aware control design. Nevertheless, they do not directly address the coupled problem of dynamic vulnerability identification, ESS allocation, and coordinated online defense against cascading failures in transmission grids. Motivated by these ideas, the present study further investigates the coordinated control of spatially distributed ESSs at the transmission-grid level for cascading-failure defense.
Furthermore, existing studies on energy storage systems for grid resilience predominantly focus on pre-event static capacity planning or post-event passive restoration [
23,
24]. While these methods effectively allocate emergency backup resources, they inherently decouple static capacity planning from millisecond-level online topological dynamic control [
25,
26]. Consequently, ESSs are often treated as passive power sources rather than dynamic, topology-aware regulating units, restricting their capability to suppress the rapid propagation of cascading failures. To bridge this structural gap, there is an urgent need to establish an integrated closed-loop defense framework that links offline vulnerability screening and static resource planning tightly with online dynamic defense.
In recent years, deep reinforcement learning has shown great promise for application in the emergency control of power systems, particularly in the context of artificial intelligence and collaborative defense control [
27,
28]. Relevant research indicates that reinforcement learning methods can enhance online decision-making efficiency through interaction with the environment and possess strong adaptability in high-dimensional non-linear scenarios. However, from the perspective of defense against cascading faults, existing methods still have shortcomings: on the one hand, traditional reinforcement learning typically simplifies the system state into vector inputs, offering limited expression of graph-structural features such as topological reconstruction, island formation and risk migration; on the other hand, existing graph learning research largely remains at the perception and evaluation stage, failing to form a tight closed-loop with continuous control. Particularly during the evolution of cascading faults, system risks continuously change with line disconnection, power flow migration and local disconnection; defense control requires dynamic adjustment as the process progresses, and such multi-stage collaborative characteristics remain inadequately addressed in current research. Specifically, many existing DRL-based emergency control approaches primarily focus on discrete actions such as load shedding, generator redispatch, or topology adjustment, while the role of multiple spatially distributed ESSs is either absent or treated only as an auxiliary control variable without reactive power considerations [
29,
30]. Traditional vector-based DRL models struggle to capture the spatial topological mutations during cascading propagation. Although some studies have explored topology-aware models for generation redispatch [
31], the implementation of multi-ESS coordinated active and reactive dynamic support under varying topologies remains largely unexplored.
In summary, substantial progress has been made in cascading-failure identification, ESS planning, and intelligent control. However, these three aspects remain insufficiently integrated. Existing fault-identification studies pay limited attention to the continuous evolution of thermal accumulation under extreme scenarios. ESS planning still lacks a three-dimensional coordinated configuration approach for local autonomy and fault defense. Intelligent control has not yet fully adapted to rapid topological changes and multi-stage risk migration. To address these gaps, this paper investigates coordinated defense against cascading failures in power grids under extreme scenarios, as outlined below:
A thermal-accumulation-aware vulnerability identification framework is developed for extreme operating scenarios, enabling the tracking of critical-branch migration during fault evolution.
A three-dimensional coordinated ESS allocation model is established by jointly considering active power, energy capacity, and reactive power support, so as to support zoned siting and coordinated sizing in vulnerable areas.
A graph-structured multi-ESS coordinated defense framework is proposed, in which GNN encodes topology-varying system states and PPO generates coordinated active/reactive control actions, thereby linking offline identification and allocation with online dynamic defense.
2. The Evolutionary Mechanisms of Cascading Failures in Extreme Scenarios and the Need for Coordinated Defense
2.1. Temporal Evolution Mechanism of Cascading Failures
A cascading failure in a power grid is not a widespread blackout caused directly by the failure of a single component, but rather a dynamic process that gradually escalates following the triggering of an initial disturbance through a sequence of power flow redistribution, local stress accumulation, protection operation and network reconfiguration. In grids with a high proportion of renewable energy, this process often exhibits greater uncertainty. On the one hand, fluctuations in generation and load, as well as power flow reversals, can cause certain critical sections to operate under high stress for extended periods; on the other hand, extreme weather conditions can alter the external thermal boundary of power lines and increase the risk of equipment failure, making it more likely for local disturbances to occur at the system’s weakest points. Following an initial fault, power is redistributed along adjacent pathways; while some lines may not trip immediately, their conductor temperatures continue to rise, and thermal safety margins are progressively eroded, causing the system to enter a sensitive phase prior to fault propagation. If the power flow at critical sections cannot be promptly suppressed and local voltage levels maintained during this phase, subsequent events such as the successive withdrawal of lines, local disconnection, and worsening power imbalance may occur, ultimately evolving into a large-scale power outage. Therefore, between the occurrence of a fault and the system entering irreversible expansion, there is in fact an intervention window dominated by thermal accumulation. Accurately identifying high-risk branches and rapidly mobilizing defense resources within this window is key to preventing cascading faults under extreme scenarios.
This paper focuses on the initial stage of a typical fault sequence, namely the continuous propagation process following an initial disturbance, which is driven by the interaction among power-flow changes, thermal-stress accumulation, and protection actions. It characterizes the risk-propagation pathways that can be mitigated through early ESS intervention. The temporal logic of this process is illustrated in
Figure 1.
During normal operation, the system remains balanced, but under extreme conditions it often operates close to its security limits. The initial disturbance phase is typically triggered by external disturbances or equipment failures and lasts for only a short time. The subsequent accident-propagation phase is the critical stage in fault spread. Because components such as transmission lines and transformers exhibit thermal inertia, a period of thermal accumulation exists between the onset of overload and the actual operation of protection devices. This period may last for several minutes and represents the transition from a local fault to a system-wide collapse. If the fault is not effectively contained during this interval, the system may enter the collapse phase, leading to voltage or frequency instability and widespread outages.
2.2. Mechanisms of Fault Initiation and Propagation Under Extreme Conditions
The impact of extreme scenarios on cascading-failure propagation is reflected not only in the probability of an initial failure but also in simultaneous changes in the system operating state and in the pathways along which failures propagate.
In scenarios such as maximum net load and maximum gradient, the backbone sections are subjected to prolonged heavy loads, significantly reducing the regulation margin of conventional units; should any local lines be taken out of service, power flows are more likely to concentrate on adjacent routes. In scenarios such as maximum renewable energy output and worst-case voltage conditions, changes in power flow direction, insufficient local reactive power support and the vulnerability of receiving-end voltages combine to further exacerbate the consequences of a fault.
Unlike traditional static out-of-limit criteria, the risk associated with line overloading is not resolved instantaneously; there is a continuous evolutionary process between the rise in conductor temperature and the activation of protection. This process constitutes a critical transitional phase in which a cascading fault progresses from a local disturbance to topological disconnection. During this phase, the local response of a single device often struggles to simultaneously address peak shaving at the fault section, regional power support and voltage maintenance at vulnerable nodes. Energy storage, however, possesses rapid active power regulation and local reactive power support capabilities, enabling it to alter the subsequent power flow distribution and the direction of risk evolution during the early stages of fault propagation. When multiple ESS units act in coordination according to network topology and real-time operating status, they can cover multiple critical areas spatially, provide both overload suppression and voltage support functionally, and intervene continuously from the early stage of a fault through its propagation. Therefore, preventing cascading failures under extreme scenarios is not a single-point control problem but a coordinated defense problem centered on critical sections, vulnerable areas, and dynamic evolution.
Based on the above analysis, this paper breaks down the problem of preventing and controlling cascading faults under extreme scenarios into three interrelated stages: firstly, identifying key vulnerable branches that undergo dynamic migration under the influence of source-load fluctuations in extreme conditions; secondly, completing the zonal allocation and capacity matching of energy storage resources to meet the requirements for local disconnection and support of vulnerable areas; and thirdly, organizing coordinated control of multiple energy storage systems in response to topological changes and risk states during the evolution of faults. Focusing on these three stages, the subsequent sections will explore dynamic vulnerability identification, three-dimensional energy storage configuration, and online coordinated defense, respectively.
5. Multi-ESS Coordinated Defense Control Strategy Based on Graph Reinforcement Learning
Section 4 completed the screening of candidate energy storage nodes and the three-dimensional coordinated capacity determination of active power, energy, and reactive power. However, the planning results can enhance system resilience only if they are implemented promptly and in a coordinated manner during fault evolution. Under extreme scenarios, cascading failures involve time-varying topology, strong state nonlinearity, and tightly coupled constraints. During fault propagation, line outages, power-flow redistribution, and local islanding continuously reshape the system risk distribution. Conventional control methods based on fixed rules or linear feedback cannot effectively balance overload suppression, voltage support, and ESS energy management. Therefore, this paper develops a multi-ESS coordinated defense framework that integrates graph neural networks with proximal policy optimization. The framework captures changes in network topology and operational risk in a graph-structured state space and outputs coordinated active and reactive power commands for ESS units. Its overall architecture consists of an environment layer, a perception layer, a decision layer, and an execution layer, as shown in
Figure 2.
Communication assumptions: To focus on the multi-ESS cooperative control mechanism itself, this study assumes an ideal and reliable communication link between the control center and the energy storage units. Communication constraints such as delays, packet loss, and measurement asynchrony are not explicitly modeled at the current stage.
5.1. Intelligent Collaborative Control Framework Based on Graph Neural Networks
The defense capability of ESSs depends not only on planning-level capacity allocation but also on whether the operational layer can rapidly coordinate multi-node responses based on real-time conditions. Based on the thermal-accumulation-aware cascading-failure model in
Section 3 and the ESS allocation results in
Section 4, this paper formulates the power-grid environment, ESS units, and fault-propagation process as a graph-structured decision-making problem. Specifically, grid nodes and transmission lines form a time-varying graph, while quantities such as nodal power, voltage, ESS state of charge, and line thermal risk constitute the state features. The graph neural network extracts topology-related representations, and the PPO agent generates coordinated control actions for multiple ESSs, thereby jointly suppressing overload in critical sections and voltage risk in weak areas on a second-level timescale.
5.1.1. Graph-Structured State Space
Considering that during the cascading-failure process, there may be line withdrawals, network decoupling and local topology reconfiguration, the traditional fixed-dimensional vector-based state representation struggles to accurately characterize the connection relationships between nodes and the risk migration process. This paper represents the power grid state at time
as a time-varying graph
, where
is the node set and
is the edge set that continuously changes with the fault evolution. The state space
is composed of the node feature matrix
and the edge feature matrix
. For each node
,
, the following Equation (41) is constructed:
In the formula,
and
represent the amplitude and phase angle of the node voltage;
and
represent the active and reactive power injection at the node;
is the line load rate;
is the state of charge of each energy storage unit;
is the current adjacent topology matrix;
is the node degree;
is the energy storage installation identifier.
For each line
, the edge feature matrix
is constructed as shown in Equation (42):
In the formula,
represents the line load rate;
represents resistance and reactance;
represents the real-time temperature of the line.
5.1.2. Multidimensional Discrete Action Space
Considering that PPO is more suitable for stable training in discrete or low-dimensional action spaces and that the actual control instructions of the energy storage system involve the continuous regulation of active and reactive power of multiple units, this paper adopts a multi-dimensional discretization method to encode the control actions of the energy storage system. Assuming there are
controllable energy storage units in the system, the action
at time
is composed of the active and reactive power adjustment quantities of each energy storage unit, as shown in Equation (43):
In the formula,
and
represent the regulation actions for active power and reactive power, respectively.
5.1.3. Reward Function Design
In order to guide the intelligent agent to reduce unnecessary control costs while ensuring system security, this paper constructs a composite reward function including thermal safety, voltage stability, control smoothness and system survivability, as shown in Equation (44):
In the formula,
represents the thermal safety reward for the line;
is the temperature warning threshold;
is the voltage stability reward;
is the nominal value of the node voltage;
is the dead zone for the allowable voltage deviation;
is the penalty for action smoothness and energy consumption;
is the survival reward; and
,
,
and
are the weighting coefficients for the line thermal accumulation penalty, voltage violation penalty, control smoothness penalty, and SOC limit penalty, respectively. To balance the multiple objectives during the dynamic cascading defense, these weights are empirically tuned and set as follows:
,
,
, and
. Additionally, a positive survival reward of 0.1 is granted for each step the system operates without triggering blackout conditions, thereby encouraging the agent to prolong the secure operational margin.
5.1.4. Network Architecture and Proximal Policy Optimization
To effectively extract the state features of the graph structure and ensure the stability of the training process, this paper constructs a perception network based on Graph Attention Network v2 (GATv2) and Layer Norm and combines it with the PPO algorithm to complete the policy update.
At the perception layer, GATv2 is first utilized to aggregate node features and edge features. Compared to traditional graph attention networks, GATv2 has a more flexible dynamic attention mechanism, which can more sensitively capture the impact of key routes and high-risk nodes on the global state after a fault. Subsequently, Layer Norm is introduced after the convolution layer output to normalize features of different dimensions such as voltage, power, and temperature, in order to alleviate the problems of numerical instability and gradient fluctuations. Finally, through global pooling operations, the updated node embeddings are aggregated into graph-level representations, serving as the shared input for the Actor network and the Critic network.
At the decision-making level, PPO is used to update the parameters of the policy network. PPO introduces a clipping objective function to limit the update range between the old and new policies. The objective function is shown in Equation (45), thereby avoiding performance fluctuations during the training process due to excessive policy updates.
In the formula,
represents the parameters of the policy network;
indicates the expected value at time step
;
is the probability ratio of the old and new policies;
is the generalized advantage estimation;
is the clipping hyperparameter.
Taking into account strategy optimization, value function fitting and exploration ability, the overall loss function is defined as shown in Equation (46). By minimizing this loss, the Actor and Critic networks can simultaneously complete parameter updates, and gradually learn to acquire multi-energy storage collaborative control strategies applicable to complex extreme scenarios.
In the formula,
represents the mean square error loss of the value network;
represents the policy entropy;
represents the weight coefficient for balancing the losses of all items.
5.2. Agent-Driven Collaborative Defense Process for Cascading Failures
After strategy training is completed, an offline-training/online-defense closed-loop process is established, as shown in
Figure 3. The process first generates extreme scenarios based on the source-load uncertainty model and uses the vulnerability-identification results from
Section 3 to determine the high-risk fault set. During training, the complexity of fault scenarios is gradually increased through curriculum learning, and the agent continuously collects measurement data from the power grid environment to construct a graph-structured state and output energy storage active and reactive control instructions; the environment layer then performs power flow calculation, thermal stability verification, and voltage safety verification based on this, and feeds back the rewards to PPO for strategy update.
During the online defense stage, when the system encounters initial disturbances in critical lines or cascading failures, the intelligent agent quickly generates the optimal energy storage response strategy based on the real-time graph state. This strategy, on the one hand, suppresses the continuous overload and temperature rise accumulation at the critical sections through active power regulation, and on the other hand, stabilizes the voltage in the local weak areas through reactive power support, thereby delaying the fault propagation and reducing the risks of system separation and load loss. As the fault evolves, the intelligent agent continues to output control actions iteratively based on the updated topology and state until the system returns to stability or the fault process ends.
Compared with traditional static defense methods, this closed-loop process achieves the integration of the risk identification in
Section 3, the energy storage configuration in
Section 4, and the online collaborative control in this section. Its core lies not in the local optimization at a single moment but in the dynamic defense for the entire fault process. This enables the energy storage resources to continuously respond over time, collaboratively cover in space, and simultaneously perform overload suppression and voltage support functions. Ultimately, it forms a collaborative defense chain for extreme scenario cascading failures.
6. Calculation Examples
To verify the effectiveness of the proposed framework for dynamic vulnerability identification, hierarchical ESS planning, and intelligent coordinated control, case studies were conducted on a modified IEEE 39-bus system, as shown in
Figure 4. In
Figure 4, the red numbers denote line numbers, and the black numbers denote bus numbers. To reflect a high level of renewable penetration, buses 34 and 37 were replaced by an equal-capacity photovoltaic plant and a wind farm, respectively, and an additional 300 MW wind farm was connected to bus 20. To maintain system power balance, the outputs of conventional units G09 and G10 were reduced accordingly. The model was implemented in Python 3.8. Grid modeling and power-flow calculation were carried out in pandapower 2.14.11, and the policy network was implemented in PyTorch 2.4.1+cu124. The total training horizon of the GNN-PPO agent was 3 × 10
6 timesteps. The discount factor was 0.99, the learning rate was 3 × 10
−4, the number of sampling steps per update was 1024, the batch size was 64, and each update was repeated 10 times. The entropy regularization coefficient was set to 0.01. Offline training and online evaluation were performed on a computer equipped with an AMD Ryzen 7940H CPU and an NVIDIA RTX 4060 GPU. The offline training process required approximately 10 h. During online execution, each dynamic fault scenario was completed within 0.35 s, which satisfies the real-time requirements of emergency cascading-failure defense.
6.1. Identification Results of Grid Vulnerability
Based on the source-load uncertainty model constructed in
Section 3.1.1, time-series curves for wind power, photovoltaic power, and dynamic loads were generated, as shown in
Figure 5 and
Figure 6, respectively. The corresponding samples contain random fluctuation characteristics and can be used as input for extracting subsequent extreme scenarios.
Figure 5 and
Figure 6 are mainly used to illustrate the randomness and correlation of the source-load sample library. They are not intended as the final comparative results of control performance. Based on these time-series samples, typical operating points are extracted according to seven indicators: maximum net load, minimum net load, maximum renewable output, maximum wind output, maximum photovoltaic output, maximum ramp rate, and worst-case voltage.
The simulation selects 7 typical scenarios covering extreme source-grid-load states, specifically including the maximum net load (Scenario 1), minimum net load (Scenario 2), maximum renewable generation output (Scenario 3), maximum wind power output (Scenario 4), maximum photovoltaic power output (Scenario 5), maximum ramping rate (Scenario 6), and worst-case voltage condition (Scenario 7).
Under these scenarios, this paper constructs a fault sample set, performs batch cascading-failure simulations, and statistically analyzes the key fault sequences and their dynamic vulnerability scores.
Table 2 presents the chain fault sequences with the highest threat ranking and the corresponding proximity in each scenario. The results show that, in multi-stage fault sequences, the vulnerability proximity of subsequently disconnected lines generally increases. This indicates that power-flow redistribution caused by the initial disturbance continuously erodes the remaining safety margin of the system and drives the fault chain toward more vulnerable topological links. Thus, critical vulnerable branches are not fixed; they vary with both the operating scenario and the stage of fault evolution.
Figure 7 compares the decline in residual load ratio under continuous line removal for different identification methods. When the line sequence identified by the proposed method is removed, the residual load ratio decreases most rapidly, eventually reaching 46.1%, which is lower than that obtained with PageRank and CEI. This result indicates that the proposed method more accurately identifies branches that have a decisive effect on system supply capability.
Figure 8 further shows that the proposed method leads to the formation of eight electrical islands, indicating a higher degree of topological fragmentation and a stronger ability to uncover deep vulnerable structures.
6.2. Energy Storage Configuration Results
After identifying the key vulnerable branches and their temporal–spatial distribution, this paper further maps line risks to candidate ESS deployment areas and installation nodes using the hierarchical planning method proposed in
Section 4. First, the Louvain algorithm partitions the modified IEEE 39-bus system into communities with strong electrical coupling. Then, key hub nodes are selected from each zone based on electrical betweenness to form the candidate set, as summarized in
Table 3.
The final set of candidate nodes is: {2, 3, 4, 6, 8, 10, 14, 16, 17, 21, 24, 25, 26, 39}. This result ensures that the energy storage layout covers the key areas spatially and avoids excessive concentration of defense resources.
Based on this, the NSGA-II algorithm was used to solve the three-dimensional collaborative capacity allocation model constructed in
Section 4, and the energy storage configuration scheme shown in
Table 4 was obtained.
The results show clear functional differences among nodes in terms of active power, energy, and reactive power configuration. Nodes 3, 4, and 21 are located near critical power-flow corridors and therefore receive larger active-power ratings, indicating their role in peak shaving and load relief during the early stage of faults. Nodes 16, 21, and 24 are located near renewable-rich areas or weak-voltage areas and therefore receive a higher share of reactive-power capacity, reflecting the model’s emphasis on local voltage support. Some nodes are assigned larger energy capacities to ensure sustained support during prolonged fault evolution and local islanding. Overall, the configuration demonstrates complementary allocation of active power, energy, and reactive power, and shows that three-dimensional coordinated planning is better suited to cascading-failure defense than single-capacity planning.
6.3. Analysis of the Cooperative Defense Effect of Chain Failures Under Energy Storage Participation
As shown by the training-loss curve in
Figure 9, the horizontal axis denotes the training timesteps, and the vertical axis denotes the loss value. The loss decreases rapidly during the initial exploration phase and converges to a stable region by approximately 3 × 10
6 timesteps, demonstrating the stable learning capability of the graph-structured policy under complex fault scenarios.
After the physical ESS configuration was completed, the online defense performance of the GNN-PPO-based coordinated control strategy was further evaluated during fault evolution. The outage of line L7 was selected as the initial triggering event. Under this scenario, the system operated at a relatively high load level. Without effective intervention, the outage of L7 would rapidly redistribute power flow, causing continuous overload in key lines such as L27 and driving the system into a thermal-accumulation risk region.
Table 5 summarizes the cascading-failure evolution process with ESS participation.
In the first stage, following the tripping of the initial faulty line L7, the loading of L27 increased rapidly. The intelligent agent identified the energy storage nodes that had a significant impact on this section based on the current graph state, and organized multi-node coordinated output to implement active power peak shaving for the critical channel, thereby reducing the load rate of L27 from the high-risk zone to close to the safety boundary. This result indicates that energy storage can utilize the intervention window formed by thermal inertia in the early stage of the fault to provide rapid support for the critical lines, delaying subsequent tripping.
During the second to fourth stages, although the critical section risks from the previous stage were alleviated, as the system topology weakened, the overload risks gradually shifted to other lines. The system successively experienced the withdrawal of L13, L19, and L3. In the face of the constantly changing topology and overload locations, the intelligent agent did not maintain a fixed control mode. Instead, it continuously adjusted the main supporting nodes based on the real-time status, allowing the center of energy storage output to dynamically shift along with the migration of risk sections. This shows that the proposed method does not aim to rigidly suppress all local disturbances. Instead, it dynamically coordinates limited energy storage resources to prioritize maintaining the backbone grid and key power supply capabilities.
In the fifth stage, after line L8 was disconnected, L25 experienced further overload. At this point, multiple ESS units became capacity-constrained, and the overall system support capability approached its rated limit. Although the extreme overload was not completely eliminated, the system remained in a critically stable state and did not immediately evolve into an uncontrollable collapse. This provided additional time for subsequent breaker actions or load-shedding measures. These results indicate that the proposed strategy remains robust even in highly complex scenarios and can shift the defense objective from completely blocking all faults to prioritizing system survivability and reducing outage consequences.
To further illustrate the spatiotemporal characteristics of coordinated defense,
Figure 10 presents the three-dimensional distribution of active-power output for the 14 ESS nodes throughout the cascading-failure process. The ESS units do not respond simultaneously or proportionally. Instead, their output exhibits clear spatial shifts and intensity redistribution as the failure stage progresses. Key nodes assume a dominant supporting role during high-risk periods, while other nodes provide background support and local compensation. This behavior demonstrates the spatial hierarchy, temporal continuity, and functional coordination of the proposed method.
6.4. Baseline Comparison, Statistical Evaluation, and Sensitivity Analysis
To validate the superiority and robustness of the proposed framework, a comprehensive statistical comparison is conducted between the proposed GNN-PPO strategy and a conventional Sensitivity-Based Control (SBC) baseline. The conventional SBC strictly relies on static numerical sensitivities (e.g., PTDF) to dispatch local ESSs, fundamentally lacking global graph-topology awareness and multi-layer coordination.
To evaluate performance under highly stochastic operating conditions, a Monte Carlo simulation was implemented. Gaussian white noise (±10% fluctuation) was superimposed on the source-load profiles to generate N = 10 randomized extreme operational seeds. The statistical comparison results are comprehensively summarized in
Table 6.
As demonstrated in
Table 6, under equivalent extreme cascading fault scenarios, the proposed GNN-PPO framework significantly outperforms the conventional SBC method across all critical metrics. By introducing global topological awareness, the proposed method reduces the maximum line thermal stress by 25.33% and narrows its standard deviation (from 26.53 to 19.88), indicating much more stable flow control under stochastic noise. Furthermore, the GNN-PPO framework’s dynamic active/reactive coordination successfully mitigates local voltage dips, increasing the average number of cascading stages survived from 3.80 to 4.80, which corresponds to a 26.32% improvement in system survivability.
- 2.
Sensitivity analysis
A sensitivity analysis was conducted with respect to renewable-energy penetration. The renewable-penetration factor was increased from the baseline case (Scale = 1.0) to more extreme low-inertia conditions (Scale = 1.66 and 2.0). Under the baseline scale, the system survived four cascading stages, with a maximum loading of 162.25%. When renewable penetration doubled (Scale = 2.0), overall system inertia decreased substantially, accelerating fault propagation. Nevertheless, under the continuous spatial regulation of the GNN-PPO agent, dynamic reactive-power support effectively compensated for voltage instability. Even at Scale = 2.0, the system survived four critical cascading stages without total blackout, although the maximum line loading increased to 225.58%. These results confirm the adaptability and robustness of the proposed multi-ESS framework under extremely high renewable penetration.