1. Introduction
The increasing penetration of distributed energy resources and the progressive decentralization of electrical distribution systems have motivated the development of advanced coordination and control strategies for networked microgrids (NMGs) [
1,
2]. In these settings, local subsystems must continuously optimize their internal operation while simultaneously coordinating energy exchanges with neighboring microgrids. This naturally leads to hierarchical or bilevel decision-making structures in which local operational objectives interact with system-level coordination goals [
3]. Achieving fast, stable, and scalable solutions to such optimization problems is essential for enabling real-time energy management, particularly under the variability and heterogeneity characteristic of modern microgrids [
4,
5].
A common approach to solving bilevel problems in distributed settings relies on singular perturbation arguments or time-scale separation: fast dynamics handle local optimization, while slower coordination dynamics ensure consistency across the network. The consensus problem has emerged as a compelling paradigm for orchestrating complex interactions in interconnected multi-agent systems [
6]. Depending on the communication structure and control objectives, consensus can be approached from cooperative, distributed, or decentralized perspectives. By organizing agents in a decentralized manner, each one is in charge of making decisions individually [
7]. This accommodates the inherent complexity of such systems, offering a structured framework to balance global objectives and local behaviors. One notable avenue to explore this concept from an optimization perspective is through the framework of bilevel optimization. Bilevel optimization extends traditional optimization paradigms by introducing two optimization layers: an upper-level optimization task that seeks to optimize a global objective, subject to constraints defined by lower-level optimization tasks [
8,
9]. This hierarchical formulation offers a natural means of integrating coordination and autonomy within complex systems.
In the context of multi-agent systems, the bilevel structure aligns well with the consensus methodology: the upper level represents a global coordination goal, while the lower levels correspond to the optimization of individual agent behaviors. Such an integration enables a unified framework for addressing coordination and consensus problems [
10]. While effective under idealized assumptions, these methods encounter significant limitations in practice. Enforcing strict time-scale separation becomes increasingly difficult when subsystem dynamics operate at heterogeneous speeds, communication delays are present, or the system experiences rapid fluctuations in renewable generation and load. Moreover, singular perturbation-based algorithms often display degraded convergence properties or require conservative tuning to maintain stability, reducing their suitability for real-time operation [
11]. To overcome these limitations, recent work has explored predictive or sensitivity-enhanced update rules that improve the behavior of distributed optimization algorithms without relying on idealized temporal structure [
12]. However, most existing formulations either impose strong regularity assumptions, lack a systematic conditioning mechanism, or do not explicitly address the bilevel nature of microgrid coordination problems [
13]. These gaps highlight the need for distributed methods that remain stable and efficient under heterogeneous subsystem speeds, reduce convergence time, and scale naturally with network size [
14].
In the context of power systems, this hierarchical framework can be naturally extended to networks of networked microgrids, where each microgrid acts as an intelligent agent optimizing its local operational objectives—such as power balance or cost minimization while contributing to the overall network coordination goal [
1,
15]. At the upper level of this framework, the leader’s dynamics are formulated within a bilevel optimization structure, representing the global coordination layer of the networked microgrid system [
10,
16,
17,
18]. At this level, the objective is to minimize a global cost function, typically reflecting overall operational efficiency, economic dispatch, or power balance, defined as the aggregation of local microgrid costs while satisfying the corresponding global constraints. Complementing this, the lower level encompasses the individual optimization problems of each microgrid, where each entity seeks to optimize its own operational objectives, such as local generation–demand balance or energy trading, while accounting for its impact on the overall coordination goal [
18,
19].
To ensure effective coordination between the optimization levels, it is essential to characterize how agents are interconnected within the multi-agent network. These interconnections can be static or dynamic, homogeneous or heterogeneous, and may involve linear or nonlinear coupling depending on the communication topology and the nature of the agents’ interactions [
20]. In the context of networked microgrids, such interconnections represent power-flow exchanges or information links that enable coordination among distributed generation units and local controllers [
21,
22]. In many practical settings, these interactions occur over different time-scales, where global coordination among microgrids evolves more slowly than the fast electrical and control dynamics within each unit [
23]. Such multi-scale behavior complicates the relationship between local decision variables and the global coordination objective, often introducing sensitivity to parameter variations and dynamic conditions [
24]. To systematically capture these effects, sensitivity analysis becomes a fundamental tool, enabling the quantification of how variations at the local or lower level influence the upper-level optimization outcomes [
10,
17,
25]. The sensitivity-conditioning concept was proposed as an alternative to singular-perturbation/time-scale separation to accelerate nested dynamics while preserving stability [
14]. Prior predictive-sensitivity work applied this idea to leader–follower dynamics and hierarchical decision problems [
10,
17]. Traditional microgrid coordination and distributed control approaches often rely on singular perturbation/time-scale separation for stability analysis, which can be restrictive under heterogeneous subsystem speeds. Finally, bilevel formulations for microgrid and distribution-grid interactions are an active area of research and provide a natural application domain for the proposed sensitivity-conditioning approach [
26].
The main contribution of this paper is the introduction of a distributed sensitivity-conditioning approach for bilevel optimization in networked microgrids. The proposed work is methodological, supported by theoretical and application-driven advances. A distributed sensitivity-conditioned bilevel dynamic framework that relaxes restrictive time-scale separation assumptions inherent to classical hierarchical control is developed, while preserving convergence guarantees and enabling scalable real-time implementation in networked microgrids. The key idea is to embed sensitivity-based predictive terms into the dynamic updates of each subsystem, effectively conditioning the distributed dynamics to mitigate the impact of heterogeneity and reduce the need for strict time-scale separation. Unlike conventional singular perturbation techniques, the proposed formulation provides a principled means to accelerate convergence while preserving robustness in the presence of subsystem differences and real-world operational variability [
17]. In the microgrid setting, these matrices capture how adjustments in local control actions, such as generation scheduling or voltage regulation, affect global synchronization metrics like frequency or power balance. The optimization levels interact dynamically, enabling the system to adapt to variations in both global objectives and local operating conditions. The proposed method ensures that the sensitivity matrices are computed efficiently, resulting in minimal additional computational burden. The effectiveness of the approach is demonstrated through simulations on a networked microgrid scenario. The results show that sensitivity-conditioning leads to faster convergence, improved stability characteristics, and efficient energy exchange coordination across diverse operating conditions. These properties make the proposed method a promising candidate for real-time distributed energy management, offering a scalable and computationally efficient alternative for bilevel control in emerging power network architectures.
Limitations of Existing Hierarchical and Bilevel Coordination Methods
Hierarchical and bilevel optimization frameworks have been widely investigated for the coordinated control and energy management of networked microgrids. Recent surveys highlight that most existing approaches can be broadly categorized into hierarchical control schemes, distributed optimization methods, and bilevel or Stackelberg-based formulations, each with distinct advantages and limitations [
27,
28]. A dominant class of bilevel control approaches relies on singular perturbation (SP) theory, where a strict time-scale separation between upper-level coordination and lower-level energy management is assumed. While analytically convenient, this assumption is difficult to enforce in practical microgrids, where pricing mechanisms, economic dispatch, and local EMS decisions often evolve on comparable time-scales. Violations of time-scale separation have been reported to cause slow transients and degraded performance, especially under fast demand fluctuations and renewable intermittency [
28].
Alternative coordination strategies based on distributed optimization and ADMM have been proposed for economic dispatch and frequency regulation in microgrids. These methods typically require multiple inner iterations per control update and frequent communication among agents, leading to non-negligible latency and convergence-dependent runtime [
29,
30]. Such characteristics limit their applicability in real-time energy management scenarios with tight computational and communication constraints. Bilevel and Stackelberg game-based formulations have also been applied to microgrid operation, particularly for leader–follower pricing, demand response, and hybrid energy storage coordination [
31,
32]. While these methods effectively capture hierarchical decision-making structures, they often assume static lower-level optimal responses and embed the follower’s solution via Karush–Kuhn–Tucker (KKT) conditions or equilibrium reconstruction. This results in centralized implementations, increased numerical complexity, and limited robustness to modeling uncertainty and dynamic coupling effects. More recently, sensitivity-based leader–follower dynamics have been explored as a means to accelerate convergence by incorporating information on the followers’ response into the leader’s update law [
17,
18]. However, existing sensitivity-conditioned approaches are typically limited to static or centralized settings, rely on exact lower-level optimality, or lack rigorous convergence guarantees under distributed communication constraints and networked interactions.
The remainder of this paper is organized as follows.
Section 2 introduces the Energy Management System (EMS) formulation for networked microgrids, describing the structure of the upper- and lower-level optimization problems.
Section 3 presents the proposed sensitivity-conditioning framework for bilevel optimization and its analytical foundations, including assumptions, convergence properties, and algorithmic formulation.
Section 4 validates the proposed dynamics using a benchmark classification problem and analyzes the convergence improvement introduced by the sensitivity-conditioning term. Finally,
Section 5 summarizes the main findings and outlines future research directions.
Notation. The set of real numbers is denoted as . and x denote matrices and vectors, respectively. We write , , for the transpose of a matrix, a vector, or a function, respectively. The optimal values in the optimization problems are denoted as . variables with overlines and underlines denote upper and lower limits, respectively (e.g., and represent the maximum and minimum power of unit i). In graph theory, the Laplacian matrix of a graph is defined as , where is the degree matrix, and is the adjacency matrix. Based on the structure of , at least one of its eigenvalues is zero and the rest have non-negative real parts. A digraph G is weight-balanced if and only if .
3. A Sensitivity-Conditioning for Bilevel Optimization
In this section, we show how decentralized interconnected systems with two-time-scale decentralized interconnected systems can be reformulated into a single-time-scale framework by means of a predictive-sensitivity matrix in continuous time. Initially, we introduce the concept of the sensitivity matrix that relates the dynamics of the systems to a bilevel optimization problem.
Building upon the bilevel formulation introduced in the previous section, we now express the general structure of leader–follower optimization problems in a compact form. This abstraction allows extending the NMG coordination problem to a broader class of decentralized multi-agent systems. Traditionally, we define a set of leader agents as the upper-level problem, and the followers’ optimization problem is referred to as the lower-level problem. Each level has its own set of variables, constraints, and objective functions, as defined by
where
is the upper-level variable, and
is the lower-level variable,
is the objective function of the leader agents, and
is the objective function of the followers. We can assume that each leader has an objective function
with
, and each follower
, where
N is the number of agents. The predictive-sensitivity matrix captures how changes in the leader’s variables
x affect the followers’ responses
, thus linking the upper- and lower-level problems in continuous time. The following assumptions will be used, which are standard in bilevel optimization and decentralized optimization literature [
34].
Assumption 1. is assumed twice continuously differentiable, and is three times continuously differentiable, both with Lipschitz continuous partial derivatives.
Assumption 2. Function is assumed μ-strongly convex in y for all i.
In microgrid applications, generation and exchange costs are commonly modeled using quadratic or piecewise-quadratic functions, reflecting fuel costs, wear-and-tear, and efficiency losses. When convex operational constraints are enforced (generation limits, storage bounds), the resulting EMS problems are strongly convex in most operating regions.
Assumption 3. For every x, the lower-level problem admits at least one solution , and the Hessian is nonsingular.
For differentiable EMS cost functions with strictly convex structure, sensitivity matrices naturally arise from the implicit function theorem and can be computed analytically or numerically with low computational effort. In real-time microgrid operation, these sensitivities can be updated either analytically (for quadratic models), or via local linearization around the current operating point. Thus, this assumption is practically reasonable for EMS models deployed in real-time controllers. For differentiable EMS cost functions with strictly convex structure, sensitivity matrices naturally arise from the implicit function theorem and can be computed analytically or numerically with low computational effort.
Assumption 4. The communication graph among agents is strongly connected and weight-balanced. Moreover, it contains a spanning tree with the leader as the root node.
In NMGs, communication networks are typically designed to ensure at least weak connectivity, often via supervisory controllers or distribution system operators (DSOs). Weight-balanced or approximately balanced graphs naturally arise in peer-to-peer communication protocols and consensus-based coordination mechanisms used in energy management. Temporary communication failures, packet losses, or topology changes may violate this assumption. While the current analysis ensures convergence under fixed or slowly switching connected topologies, abrupt or prolonged disconnections may degrade performance. Nevertheless, the distributed nature of the algorithm allows microgrids to continue operating locally, preserving safety even when global optimality is temporarily lost.
Under Assumption 1, the invertibility assumption defines the sensitivity of
. Then, the implicit-function theorem [
35] guarantees the local existence of the map
, and gives an expression for its derivative as
For a given
x, Assumption 1 ensures that the lower-level problem in (
8) admits at most one optimal solution, denoted by
. The mapping
allows for defining a reduced objective function
, whose gradient leads to the definition of the total derivative:
for points where
. Under Assumptions 1–3, the Hessian
is invertible for
y in a neighborhood of
. Since it is generally not possible to compute
directly,
can be approximated using the sensitivity matrix, defined as
where
quantifies how variations in the upper-level variables
x affect the lower-level responses
y.
Geometrically, it represents the local Jacobian of the optimal-response manifold , describing how the follower equilibrium deforms as the leader dynamics evolve. From a dynamical perspective, the inclusion of introduces a predictive feed-forward correction that anticipates the impact of leader updates on follower equilibria, thereby reducing inter-level lag and improving damping. In the microgrid context, this matrix admits a physical interpretation as a price–response sensitivity, quantifying how local power exchange decisions react to marginal price variations imposed by the DSO.
From a numerical standpoint, the sensitivity matrix is computed by solving the linear system , rather than explicitly forming the inverse of , which improves numerical robustness.
In the economic dispatch formulation considered in this work, the lower-level objective is quadratic and separable across microgrids, which implies that is diagonal and positive definite. As a result, the sensitivity matrix admits a closed-form expression that can be computed locally by each agent, significantly reducing the computational burden compared to general bilevel optimization settings and mitigating potential conditioning issues.
These quantities satisfy the following consistency conditions:
To guarantee that a solution for the optimization problem (14) exists, we propose the following definition.
Definition 1. A pair is said to be a local solution to the bilevel optimization problem (14) if
- 1.
is a local minimizer of the lower-level problem ;
- 2.
there exists a neighborhood of such thatwhere denotes a local solution of the lower-level problem associated with x.
The sensitivity-conditioned bilevel optimization framework considered in this work is inspired by the predictive sensitivity-based interconnection originally proposed in [
14]. While that framework addresses generic bilevel optimization problems, the present work specializes the approach to the economic dispatch problem in networked microgrids. In particular, the quadratic structure of the lower-level cost functions and the power balance constraints enable an explicit and decentralized construction of the sensitivity matrix, as well as a physically interpretable coupling with consensus-based coordination dynamics.
The dynamics to obtain the solution of the system can be defined based on [
36] as a dynamic consensus problem of the form
where
are design parameters. In this case,
and
represent the primal and dual variables associated with agent
i, respectively, while
denotes the local cost function. The term
represents a local gradient-descent update for the follower variable
, driving it toward the optimal solution
of the lower-level problem.
Let
denote the stacked vector of local functions, and accordingly
the vector of local gradients. In a compact network form, these distributed dynamics can be written as
where
is the Laplacian matrix of the communication graph.
Traditionally, dynamical systems (
22) are analyzed through singular perturbation representations [
36], where a parameter
typically separates the fast follower dynamics
from the slower leader updates
. However, since the second subsystem cannot be made arbitrarily fast in practice, the design choice of
necessarily slows down the first subsystem, limiting the convergence rate and deteriorating overall performance. To further clarify the conceptual differences between the proposed sensitivity-conditioning framework and traditional singular perturbation (SP) approaches commonly used in bilevel and multi-time-scale optimization, we include a structured comparison in
Table 1.
Building upon the distributed dynamics defined in (
22), the overall behavior of the system can be described as an interconnected dynamical structure representing the coupled evolution of all agents. The solution of the optimization problem (14) can thus be interpreted as the equilibrium of this interconnected system.
where
M introduces the coupling between the leader and follower dynamics through the sensitivity matrix
. Then, considering the solution (
21), we can rewrite the interconnection as
The last row modifies the follower update by incorporating the influence of both
and
, ensuring coordinated evolution without requiring explicit time-scale separation. The matrix
M remains nonsingular under bounded sensitivity conditions, and leads to the algorithm as
With this sensitivity-conditioning approach, instead of accelerating the second subsystem as in the singular perturbation case, the conditioning term
modifies the follower dynamics without explicit time-scale separation. Given Assumption 1, we ensure local existence and uniqueness of solutions
to (
27) if the following holds:
Assumption 5. The vector fieldis locally Lipschitz in a neighborhood of the equilibrium. Under this condition, the right-hand side of (
27) is locally Lipschitz, and hence the solution
exists and is unique locally in time [
37]. Based on this definition, for analytical purposes it is important to consider the Lagrangian function associated with the reduced optimization problem. Using the Lagrange multipliers
associated with the consensus constraint
, where
denotes the vector of Lagrange multipliers associated with the consensus constraint (not to be confused with the strong convexity constant
introduced in Assumption 2), the Lagrangian is defined as
where
denotes the Laplacian matrix of the communication graph. The notion of a local solution, together with suitable first- and second-order optimality conditions, provides a foundation for analyzing the equilibrium points of the bilevel optimization dynamics. We then state the following result.
Proposition 1 (First-order conditions)
. If is a local solution of problem (14), then it is a stationary point satisfying the first-order Karush–Kuhn–Tucker (KKT) conditions:where μ denotes the Lagrange multipliers associated with the consensus constraint. The stacked vectors of local variables are defined aswith denoting the column-wise concatenation of all agents’ local variables. In the same way, the second-order necessary KKT conditions can be stated as follows.
Proposition 2 (Second-order optimality conditions)
. Let be a stationary point of problem (8) with Lagrange multipliers μ. Then, for every feasible direction satisfyingthe Hessian of the Lagrangiansatisfies the condition Likewise, the second-order sufficient optimality condition can be stated as follows:
Proposition 3 (Second-order sufficient condition)
. Let be a stationary point of problem (8) with multipliers μ. If the Hessian of the Lagrangianis positive definite on the feasible tangent space, i.e.,then is a strict local minimum. The first- and second-order optimality conditions derived above characterize the local behavior of the bilevel optimization problem around a stationary point. In particular, they establish that the equilibrium of the distributed dynamics corresponds to a local minimum of the reduced problem satisfying both the consensus and feasibility constraints.
Consequently, under the differentiability and Lipschitz assumptions stated earlier, the continuous-time distributed algorithm (
22) can be shown to converge locally toward such equilibrium points. The continuous-time sensitivity-conditioned bilevel dynamics considered in this work are inspired by the sensitivity-based interconnection framework originally proposed in [
14]. In that framework, a predictive feed-forward term is introduced to overcome time-scale separation requirements in bilevel optimization. In contrast, the present work specializes the sensitivity-conditioned dynamics to the economic dispatch problem in networked microgrids, where the quadratic structure of the generation costs and the power balance constraints yield explicit expressions for the sensitivity matrix and enable a physically interpretable, fully distributed implementation.
Theorem 1 (Local Convergence of the Algorithm)
. Under Assumptions 1, 3, and 5, the algorithm (22) converges locally to an optimal solution of problem (14). Proof. Let
be an equilibrium of (
22) with
,
and
. By the implicit function theorem, under Assumptions 1–3, there exists a continuously differentiable mapping
such that
for all
x in a neighborhood of
, and
. Hence, the lower-level solution can be locally represented as
. Then, the reduced cost
satisfies
and
is well defined.
Consider the error variables
and
. Linearizing the
-subsystem of (
22) along the manifold
yields
Let be an orthonormal basis diagonalizing with , where and by strong connectivity and weight balance. In this case, denotes the k-th eigenvalue of the Laplacian matrix .
In the consensus direction (
), the linearized dynamics reduce to
,
, for which asymptotic stability of
follows from
and the fact that
remains bounded and is driven to zero by the integral action on disagreement modes. On the disagreement subspace (eigenvalues
), each mode satisfies
whose characteristic polynomial is
with
. Thus, all roots satisfy
for any
and
, and the real parts are uniformly bounded away from zero by choosing
sufficiently large. Therefore,
is Hurwitz on the disagreement subspace and the consensus mode is stabilized by the strong convexity of
.
To incorporate the
y-dynamics, write the linearization
which is exponentially stable and input-to-state stable (ISS) with respect to
because
. By standard small-gain/ISS interconnection arguments, the full
-system is locally exponentially stable around
. Hence, trajectories of (
22) converge locally to
, which completes the proof. □
Remark 1. The computational cost of the proposed sensitivity-conditioned bilevel dynamics is dominated by the evaluation of local sensitivity matrices, which are computed independently by each microgrid. For quadratic or strongly convex EMS cost functions, the Hessian structure can be precomputed offline, reducing the online sensitivity update to a low-dimensional matrix–vector multiplication. As a result, the per-step computational complexity scales as per microgrid, where denotes the number of local EMS decision variables. In contrast to KKT-based bilevel reformulations, the proposed approach does not require inner optimization loops or centralized matrix factorizations. The overall computational cost scales linearly with the number of microgrids and admits fully parallel implementation, making the method well suited for real-time distributed energy management in large-scale microgrid networks.
On the other hand, although the proposed framework assumes convex lower-level cost functions for analytical tractability, it remains applicable in non-convex EMS scenarios in a local or piecewise-convex sense. In such cases, the method converges to stable locally optimal equilibria, which are often sufficient for real-time microgrid operation. Extending global optimality guarantees to fully non-convex formulations remains an open research direction.
In Algorithm 1, a pseudo-code of the main algorithm is presented to illustrate the states and equations used in the calculations.
| Algorithm 1 Distributed Sensitivity-Conditioned Bilevel Coordination |
Input: Set of microgrids ; cost functions ; upper-level objective ; global constraint . Output: Equilibrium power exchanges and prices .
- 1:
Initialize , for all - 2:
Initialize global dual variable - 3:
Set controller gains and step sizes - 4:
while not converged do - 5:
for each microgrid (in parallel) do - 6:
Lower-level update (EMS optimization): - 7:
Compute local sensitivity: - 8:
Upper-level update (price coordination): - 9:
Global constraint update: - 10:
Integrate dynamics over one time step
|
The theoretical developments presented in this section establish a unified framework for analyzing bilevel optimization in decentralized multi-agent systems. By introducing the sensitivity-based interconnection and proving local convergence under standard convexity and connectivity assumptions, the proposed approach ensures that both the leader and follower dynamics evolve coherently toward a locally optimal equilibrium. These results guarantee that the distributed algorithm achieves coordination without requiring explicit time-scale separation between subsystems. In the following section, we validate the theoretical findings through numerical simulations, demonstrating the performance and convergence behavior of the proposed algorithm in the context of networked microgrids.