Skip to Content
MathematicsMathematics
  • Article
  • Open Access

30 September 2026

32 Pages

Privacy-Preserving Distributed Online Dispatch of Low-Carbon Microgrids with TCN–BiLSTM Probabilistic Photovoltaic Forecasts

and
1
School of Computer Science and Technology, Shandong Xiehe University, Jinan 250109, China
2
School of Automation, Nanjing University of Information Science and Technology, Nanjing 210044, China
*
Author to whom correspondence should be addressed.

Abstract

Low-carbon microgrids are complex energy systems in which photovoltaic (PV) uncertainty, time-varying operating costs, distributed coordination, and information privacy must be addressed simultaneously. This paper couples TCN–BiLSTM probabilistic PV forecasting with privacy-preserving distributed online economic dispatch. Quantile forecasts are converted into a risk-aware net demand by combining the median PV forecast with a lower-side uncertainty reserve. The resulting dispatch problem also includes a step-type carbon-trading cost. To solve the problem when cost gradients are unavailable, we propose a probabilistic PV forecasting-driven differentially private distributed online one-point bandit optimization algorithm (PPF-DP-DOBO). Each generator uses one function-value query per iteration and perturbs its communicated state with Laplace noise. The analysis establishes a per-release differential privacy guarantee, its sequential composition over the dispatch horizon, and an individual dynamic regret bound that explicitly depends on PV forecasting uncertainty. Under bounded weighted path variation and cumulative forecasting uncertainty, the regret is sublinear with order O ( T 3 / 4 ) . Simulations on a modified IEEE 162-bus system show that the method tracks the risk-aware net demand, preserves the expected privacy–performance trade-off, and yields lower average regret and carbon cost in the evaluated probabilistic-PV setting than in the no-PV case.

1. Introduction

Low-carbon microgrids combine renewable generation, conventional distributed generators (DGs), communication networks, and market mechanisms in a tightly coupled complex system. Their operating state is nonlinear and time-varying because electricity demand, renewable generation, and generation costs evolve on different time scales. Photovoltaic (PV) generation further increases this variability, but it also reduces the demand supplied by carbon-emitting units. Effective microgrid operation therefore requires a dispatch method that can use renewable forecasts, coordinate distributed units, and adapt online without exposing sensitive operating information [1,2,3]. The work in Ref. [4] further indicates that distributed economic dispatch research is progressively extending from basic consensus-based allocation toward renewable-rich operation, practical engineering constraints, and privacy-preserving coordination.
A deterministic PV point forecast does not describe the asymmetric risk of PV underproduction. Probabilistic forecasting addresses this limitation by estimating quantiles or prediction intervals rather than a single trajectory [5,6]. In this work, a temporal convolutional network and bidirectional long short-term memory network (TCN–BiLSTM) model short- and long-range dependencies in PV and meteorological time series. Its median forecast represents the expected PV contribution, whereas a lower quantile defines the reserve required to hedge against underproduction. This construction provides a direct interface between machine-learning-based uncertainty estimation and risk-aware dispatch. The work in Ref.  [7] also employs a robust multi-objective optimization method to consider the uncertainties of renewable energy generation and load in uncertainty-aware microgrid research.
The dispatch problem is also challenging when generator costs are time-varying or available only through function-value queries. Distributed gradient methods have been widely studied for economic dispatch [8,9,10], but they require explicit gradients or sufficiently accurate gradient evaluations. Gradient-free online optimization instead constructs an update direction from queried function values [11,12,13]. Two-point estimators offer favorable accuracy but require two evaluations per update [14,15]. One-point bandit feedback reduces this requirement to one query and is therefore attractive for online energy systems with limited computational or communication resources [16,17,18,19].
Distributed communication introduces a second source of risk. Exchanged generator states may reveal operating conditions or cost-related information to an observer. Differential privacy provides a mathematically explicit way to bound the influence of an adjacent input on a released state [20]. Existing privacy-preserving dispatch methods protect state exchange under several network and optimization settings [21,22,23,24]. For microgrid applications, the work in Ref. [25] proposes a distributed differentially private consensus controller that combines nondecaying noise with a time-varying control gain to protect exchanged information while maintaining frequency regulation and active-power sharing. The method proposed in Ref. [26] integrates attenuated differential privacy and fully homomorphic encryption into distributed alternating direction method of multipliers to protect generator information during distributed microgrid energy management. However, privacy noise degrades coordination accuracy, and its cumulative effect must be distinguished from the uncertainty already introduced by renewable forecasting.
Dynamic regret is the appropriate performance measure for this setting because the optimal dispatch changes with load, PV output, and operating costs [27,28,29,30,31]. Existing online bandit and privacy-preserving dispatch studies do not jointly expose how forecast uncertainty, one-point feedback, privacy perturbations, and a time-varying comparator affect the regret bound. Moreover, treating probabilistic forecasting and online dispatch as unrelated modules obscures how prediction intervals alter the mathematical dispatch target.
Accordingly, a research gap remains between uncertainty-aware microgrid scheduling, privacy-preserving distributed optimization, and cyber-resilient reconfiguration: existing studies generally emphasize one or two of these aspects, whereas their combined influence on a distributed online bandit regret bound has received limited attention.
This paper addresses this gap through PPF-DP-DOBO, a probabilistic PV forecasting-driven differentially private distributed online one-point bandit optimization method. The TCN–BiLSTM forecaster first produces PV quantiles. The median and lower quantile are then mapped to a risk-aware net demand, which becomes the time-varying dispatch target. The dispatch stage uses one function-value query to estimate a search direction and adds Laplace noise to communicated generator states. A step-type carbon-trading cost represents the increasing economic penalty of high-emission operation. The comparison with recent works is listed in Table 1.
Table 1. Qualitative comparison of representative distributed economic dispatch and online optimization approaches.
The main contributions are as follows.
(1)
A forecast-to-dispatch coupling is developed for low-carbon microgrids. TCN–BiLSTM quantile forecasts are converted into a median PV contribution and a lower-side uncertainty reserve. These quantities define a risk-aware net demand and connect data-driven temporal uncertainty with the economic dispatch model.
(2)
A privacy-preserving distributed online dispatch method is formulated for unknown time-varying cost functions. PPF-DP-DOBO uses one function-value query per generator and iteration, while Laplace perturbations protect communicated power states. The privacy analysis distinguishes the per-release privacy budget from its sequential composition over the dispatch horizon.
(3)
A conditional dynamic-regret guarantee is derived for the coupled forecasting and dispatch system. The bound contains both the weighted path variation of the dynamic optimum and the cumulative PV forecasting uncertainty. Simulations on a modified IEEE 162-bus system evaluate forecasting behavior, supply–demand tracking, carbon cost, dynamic regret, baseline performance, and the privacy–performance trade-off.
The remainder of this paper is organized as follows. Section 2 introduces the communication and privacy preliminaries, the TCN–BiLSTM probabilistic forecaster, and the low-carbon dispatch problem. Section 3 presents PPF-DP-DOBO and its theoretical guarantees. Section 4 reports the numerical experiments, Section 5 discusses their implications and limitations, and Section 6 concludes the paper.
Notation 1.
The set of real numbers is denoted by R , and R d is the d-dimensional Euclidean space. The feasible decision set is Ω, and Π Ω [ x ] = arg min x ^ ∈ Ω ∥ x ^ − x ∥ 2 denotes the Euclidean projection. The unit sphere in R d is S d . The TCN–BiLSTM quantile forecaster is F θ , where θ contains its trainable parameters. The input window is X t , the q-quantile PV forecast is p ^ p v ( q ) , t , the lower-side reserve is R p v t , the risk-aware net demand is P net t , and cumulative PV forecasting uncertainty is U p v , T .

2. Preliminaries and Problem Formulation

2.1. Graph Theory

A directed graph G : = ( V , E ) is employed to model information sharing among N agents. The graph G comprises a vertex set V = { 1 , 2 , … , N } and edge set E ⊆ V × V . The connectivity of G is ensured by the existence of a path that links any pair of agents. We utilize N i i n = { j ∈ V ( j , i ) ∈ E } to represent the number of edges that point to agent i (in-neighbors), and N i o u t = { j ∈ V ( i , j ) ∈ E } to represent the number of edges departing from agent i (out-neighbors). Let A = [ a i j ] ∈ R N × N be the weighted adjacency matrix of G , where a i j represents the weight assigned to the edge ( i , j ) .

2.2. Differential Privacy

Differential privacy is a pivotal and indispensable concept and technology in the field of privacy protection, serving as a vital tool to address the inherent risks posed by the potential disclosure of sensitive personal information during data analysis and sharing. Traditional approaches to privacy protection, such as de-identification or encryption, commonly prove inadequate in effectively mitigating the re-identification of personal data. In contrast, differential privacy provides a robust and rigorously defined framework that ensures the preservation of individuals’ privacy by incorporating controlled perturbation techniques during the data processing phase. Therefore, this paper will employ differential privacy techniques to effectively protect the privacy of generator power information. Next, we first give the following definitions.
Definition 1.
Given two different datasets Θ and Θ ′ , if Θ can be obtained by adding, deleting, or modifying a single element from Θ ′ , then Θ and Θ ′ are considered adjacent.
Definition 2.
For any adjacent inputs Θ and Θ ′ and any measurable output set ψ, a randomized mechanism M is ε-differentially private if
P [ M ( Θ ) ∈ ψ ] ≤ e ε P [ M ( Θ ′ ) ∈ ψ ] ,
where the probability is taken over the internal randomness of M .
Definition 3.
Let q t ( Θ ) denote the deterministic vector-valued state query to be released at time t. Its global ℓ 1 -sensitivity over adjacent inputs is defined as
Δ t = sup Θ ∼ Θ ′ q t ( Θ ) − q t ( Θ ′ ) 1 ,
where Θ ∼ Θ ′ denotes that Θ and Θ ′ are adjacent according to Definition 1.
In the present dispatch setting, two inputs are adjacent if they differ only in one private local generation or cost record of one generator at one time, while the public load profile, PV forecasts, communication topology, and all other generators’ records remain unchanged. The deterministic query is the stacked generator-state vector q t ( Θ ) = p t ( Θ ) , and the released mechanism is M t ( Θ ) = q t ( Θ ) + η t . Therefore, one release refers to the joint communicated state vector at one dispatch time rather than to each individual scalar component.
The parameter ε t > 0 is the dimensionless privacy-loss budget assigned to the release at time t. Inequality (1) shows that changing one private record can change the probability of any released event by at most the multiplicative factor e ε t . A smaller ε t makes the output distributions under adjacent inputs closer and therefore provides stronger privacy, but it increases the Laplace scale σ t = Δ t / ε t and may reduce dispatch accuracy. Conversely, a larger ε t introduces less noise but provides weaker privacy protection.
Because the mechanism releases one state vector at every dispatch time, ε t is a per-release budget rather than the privacy guarantee for the entire dispatch horizon. Under sequential composition, the pure-differential-privacy budget accumulated from time 0 to time T is
ε tot ( T ) = ∑ t = 0 T ε t .
Thus, using a constant per-release budget ε t = ε results in ε tot ( T ) = ( T + 1 ) ε . For a prescribed finite-horizon budget ε ¯ , the release budgets should instead be selected such that ∑ t = 0 T ε t ≤ ε ¯ ; for example, the uniform allocation ε t = ε ¯ / ( T + 1 ) is privacy-feasible for a fixed horizon. The resulting guarantee protects the released generator states against inference under the stated adjacency relation, but it does not by itself prevent active message falsification, denial-of-service attacks, or physical manipulation of the microgrid.
Remark 1.
The game-theoretic framework in Ref. [32] uses a Stackelberg game to mitigate false-data-injection attacks and dynamically reconfigure interconnected microgrids for secure and cost-efficient operation. In contrast, PPF-DP-DOBO employs sensitivity-calibrated Laplace noise to limit information inference from communicated generator states. Thus, the former addresses active cyberattacks and network resilience, whereas the latter protects information privacy during distributed online dispatch. Their integration is left for future research.

2.3. TCN-BiLSTM-Based Probabilistic Photovoltaic Power Forecasting Model

Accurate photovoltaic (PV) power forecasting is essential for reliable economic dispatch under renewable energy uncertainty. However, PV generation exhibits strong temporal fluctuations and nonlinear characteristics due to the combined influence of historical power outputs and meteorological conditions. To address these challenges, a hybrid temporal convolutional network and bidirectional long short-term memory network (TCN-BiLSTM) is developed in this paper for probabilistic PV power forecasting.
The proposed forecasting model integrates the advantages of temporal convolutional networks (TCNs) and bidirectional long short-term memory networks (BiLSTMs). The TCN module is responsible for extracting multi-scale temporal patterns from historical PV generation sequences through causal dilated convolutions, while the BiLSTM module captures long-range temporal dependencies by considering information from both forward and backward directions.
Given the historical PV power and meteorological feature sequence, the input vector at time t is defined as follows:
X t = [ x t − N + 1 , x t − N + 2 , ⋯ , x t ]
where N represents the length of the historical observation window. Specifically, the feature vector at time t is defined as x t = [ p p v t , G t GHI , G t GTI , G t DHI , T t , H t ] T , where p p v t denotes the historical PV power output, G t GHI , G t GTI , and G t DHI denote the global horizontal irradiance, global tilted irradiance, and diffuse horizontal irradiance, respectively, T t denotes the ambient temperature, and H t denotes the relative humidity. Therefore, X t contains both the recent PV-generation trajectory and the meteorological conditions associated with solar-resource availability and environmental variations.
The TCN module consists of multiple residual temporal convolution blocks. Different from conventional convolutional neural networks, TCN adopts causal convolution to preserve the temporal order of historical information and utilizes dilation factors to enlarge the receptive field.
The feature extraction process of the TCN module can be formulated as follows:
Z t ( l ) = σ ( W l ∗ d Z t ( l − 1 ) + b l )
where Z t ( l ) denotes the output feature representation of the l-th TCN layer, W l and b l represent convolution parameters, ∗ d denotes dilated causal convolution operation with dilation factor d, and σ ( · ) is the nonlinear activation function.
Through the dilation mechanism, the TCN module can effectively extract both short-term fluctuation characteristics and long-term variation trends of PV generation without significantly increasing computational complexity.
After temporal feature extraction, the obtained feature sequence is further processed by the BiLSTM network. The hidden representation generated by the BiLSTM module is expressed as follows:
h t = B i L S T M ( Z t ( L ) )
where h t represents the temporal feature representation learned by the BiLSTM network.
The BiLSTM structure contains two independent LSTM networks, including forward and backward propagation paths. Therefore, the model can simultaneously utilize historical and future information within the input window, improving the ability to capture nonlinear temporal dependencies caused by weather variations.
The final probabilistic PV power prediction output is obtained through a fully connected layer:
p ^ p v t = F θ ( X t ) = W o h t + b o
where F θ denotes the proposed TCN-BiLSTM forecasting network parameterized by θ , and p ^ p v t represents the predicted PV power under different quantile levels.
To characterize forecasting uncertainty, multiple quantile outputs are generated:
p ^ p v t = [ p ^ q 1 t , p ^ q 2 t , ⋯ , p ^ q m t ]
where q m denotes the corresponding quantile level.
The forecasting model is trained by minimizing the quantile regression loss function:
L p v ( θ ) = 1 m T ∑ t = 1 T ∑ r = 1 m ρ q r [ F θ ( X t ) ] r , p p v t
where the pinball loss for quantile q ∈ ( 0 , 1 ) is
ρ q ( p ^ , p ) : = q − 1 { p < p ^ } ( p − p ^ ) ,
and 1 { · } is the indicator function. This objective penalizes under- and over-prediction asymmetrically and yields the quantiles used by the dispatch stage.
To improve the reproducibility of the forecasting model, its main hyperparameters are summarized in Table 2. The input window contains 24 consecutive samples. The TCN component comprises three temporal convolution layers with a kernel size of 3 and dilation factors of 1, 2, and 4. The extracted temporal features are processed by a two-layer BiLSTM network with 64 hidden units. The model is trained for 100 epochs using the Adam optimizer with a learning rate of 0.001 and a batch size of 32.
Table 2. Hyperparameters of the proposed TCN–BiLSTM forecasting model.
Compared with conventional CNN-based forecasting approaches, the proposed TCN-BiLSTM model provides two advantages. First, the TCN structure captures multi-scale temporal characteristics through dilated causal convolutions, which is suitable for PV sequences with periodic fluctuations and sudden changes. Second, the BiLSTM structure enhances the modeling capability of long-term dependencies by utilizing bidirectional temporal information.
The resulting median and lower-quantile forecasts provide, respectively, the expected operating trajectory and the lower-tail uncertainty information required by the subsequent dispatch formulation.

2.4. Problem Formulation

The EDP within a microgrid entails considering every conventional generator as a node and synchronizing all nodes to minimize production costs while adhering to supply–demand equilibrium and generation constraints. Consequently, the EDP within a microgrid fundamentally represents a constrained optimization challenge. Additionally, in the context of low-carbon electricity, energy conservation and emission reduction have become important objectives for microgrid systems. To simultaneously consider the economic and environmental aspects of a microgrid, this paper integrates a carbon-trading mechanism into its economic dispatch model.
Different from deterministic PV-assisted dispatch models, this paper incorporates TCN–BiLSTM probabilistic forecasts into the dispatch formulation. PV output is influenced by solar irradiance, ambient temperature, relative humidity, and recent generation, so a single point forecast cannot express the lower-tail risk relevant to power balance. The forecaster therefore provides multiple PV quantiles:
p ^ p v t = F θ ( X t ) , p ^ p v ( q r ) , t = F θ ( X t ) r , r = 1 , … , m .
Here, X t denotes the sliding-window input sequence constructed from meteorological variables, historical PV generation data, and temporal information. The TCN module enables local temporal feature extraction from PV-related sequences and outputs the quantile forecasts required by the dispatch model.
For a given confidence parameter α ∈ ( 0 , 0.5 ) , the PV prediction interval at time t is defined as
I p v α , t = p ^ p v ( α ) , t , p ^ p v ( 1 − α ) , t ,
where p ^ p v ( α ) , t and p ^ p v ( 1 − α ) , t represent the lower and upper quantile forecasts of PV generation, respectively. In particular, p ^ p v ( 0.5 ) , t denotes the median forecast of PV output.
To hedge against the risk that the actual PV output is lower than the median forecast, a PV uncertainty reserve term is introduced as
R p v t = p ^ p v ( 0.5 ) , t − p ^ p v ( α ) , t + ,
where [ x ] + = max { x , 0 } . A larger R p v t indicates stronger lower-side PV uncertainty and requires more conventional generation or external grid compensation to maintain the supply–demand balance.
Assume that the microgrid contains N conventional generators. At each time t, their outputs and the scheduled external-grid exchange must meet the demand remaining after the probabilistic PV contribution. The low-carbon economic dispatch problem is
min P t F t ( P t ) = ∑ i = 1 N C t i ( p t i ) + C CO 2 t ( P t ) ,
subject to
∑ i = 1 N p t i + p t grid = p D t − p ^ p v ( 0.5 ) , t + R p v t ,
and
p min i ≤ p t i ≤ p max i , i = 1 , … , N .
Constraint (16) imposes the static generation-capacity limits of each generator. Inter-temporal generator ramp-rate limits are not explicitly included in the present formulation. Therefore, the feasible set considered here restricts the instantaneous generator output but does not constrain the change in output between two consecutive dispatch intervals.
For compactness, the right-hand side of the balance constraint is denoted as the risk-aware net demand:
P net t = p D t − p ^ p v ( 0.5 ) , t + R p v t .
where P t = [ p t 1 , … , p t N ] T is the conventional-generation vector and F t ( P t ) is the total low-carbon dispatch cost. The function C t i ( p t i ) is the private local cost of generator i. The term C CO 2 t ( P t ) denotes the carbon-trading cost incurred during the dispatch interval t and is determined by the difference between the actual carbon emissions and the allocated free carbon emission quota. The variable p t grid denotes a scheduled grid import when positive and export when negative. It is treated as known when the distributed generator set point is computed. Only the positive part [ p t grid ] + is treated as purchased electricity when calculating the carbon quota and actual emissions, whereas exported electricity does not generate purchased-electricity emissions within the microgrid. The total load is p D t , and p min i and p max i are the generator limits. In the simulations, the local cost is evaluated through the quadratic model
C t i ( p t i ) = a i ( p t i ) 2 + b i p t i + c i , i = 1 , … , N ,
where a i , b i , and c i are the cost parameters.
The formulation combines carbon cost, a time-varying balance requirement, and TCN–BiLSTM forecast uncertainty. The median forecast p ^ p v ( 0.5 ) , t reduces the net demand assigned to dispatchable resources, whereas R p v t restores a margin against PV underproduction. Thus, the forecasting model influences dispatch only through two auditable quantities: the median forecast and the lower-side reserve.
Remark 2.
The time-varying lower-side reserve R p v t is derived from probabilistic forecasts, not a fixed constant. The parameter α sets the accepted lower-tail probability: with a calibrated forecaster, 1 − α is the nominal probability that actual PV output exceeds p ^ p v ( α ) , t . Smaller α selects a more conservative quantile and thus a larger reserve; larger α reduces it. Hence, α should be set by the microgrid operator to balance protection against PV underproduction and reserve procurement cost. When R p v t = 0 , only the median forecast enters the balance equation. The deterministic PV-assisted dispatch model is therefore a special case of the proposed probabilistic formulation.
Remark 3.
In the reported case study, p t grid is a prescheduled quantity and is set to zero in order to isolate the effects of PV forecasting, carbon trading, differential privacy, and distributed online optimization. If active grid exchange under dynamic electricity prices is considered, p t grid becomes an additional decision variable and is jointly optimized with the local generator outputs. Let λ t buy and λ t sell denote the time-varying purchase and selling prices, respectively. The corresponding grid-transaction cost can be defined as
C t grid p t grid = λ t buy p t grid + − λ t sell − p t grid + , 0 ≤ λ t sell ≤ λ t buy ,
and the extended objective becomes
F ˜ t P t , p t grid = ∑ i = 1 N C t i ( p t i ) + C CO 2 t P t , p t grid + C t grid p t grid ,
subject to the power-balance constraint and the grid-interconnection limits. Low purchase prices would generally encourage grid imports and reduce local generation, whereas high purchase prices or attractive selling prices would favor local generation or electricity exports. However, the resulting carbon cost also depends on the emission coefficient of imported electricity and therefore need not decrease monotonically. Because active grid exchange enlarges the decision space and dynamic prices increase the temporal variation of the optimal dispatch trajectory, the present convergence and regret results are not claimed directly for this extension. Its complete theoretical and numerical investigation is left for future work.
Remark 4.
This paper investigates the EDP under unknown cost functions, which typically involves the challenge of obtaining gradient information. In general, the cost functions and parameters of generators are provided by manufacturers. However, if the cost functions are too complex or the cost of computing gradients in large-scale power systems is prohibitively high, it becomes necessary to study the EDP with unknown gradient information. In the subsequent simulations, this paper provides the specific form and parameters of the cost functions only for obtaining the required function values, rather than for calculating exact gradients.
In a carbon-trading market, the carbon payment or revenue is determined by the difference between the actual carbon emissions and the allocated free carbon emission quota. In the proposed dispatch model, the carbon-accounting quantities are evaluated separately for each dispatch interval so that the corresponding carbon-trading cost can be incorporated consistently into the time-varying objective function F t ( P t ) . PV generation is regarded as a zero-marginal-cost and zero-direct-emission energy source, whereas purchased grid electricity and gas-fired conventional generation contribute to carbon emissions.
The free carbon emission quota allocated during dispatch interval t is defined as
e f t = η 1000 ∑ i = 1 N p t i + p t grid + ,
where e f t denotes the free carbon emission quota at time t, η is the emission allowance allocated per unit of electricity supplied, and the factor 1 / 1000 is used for unit conversion. The external-grid exchange appears only once in (21), and [ x ] + = max { x , 0 } ensures that only grid imports are included in the carbon accounting.
The corresponding actual carbon emissions of the microgrid during dispatch interval t are calculated as
e p t = a 1 1000 p t grid + + a 2 n t gen ,
where e p t denotes the actual carbon emissions at time t, a 1 is the equivalent emission coefficient of purchased electricity, and a 2 is the equivalent emission coefficient of natural gas. The quantity n t gen = ∑ i ∈ G gas n t i ( p t i ) denotes the total natural gas consumption of the gas-fired generator set G gas during interval t, as determined from the corresponding generator outputs and heat-rate characteristics.
For clarity, define the carbon-emission deviation from the allocated quota as Δ e t = e p t − e f t . A negative value of Δ e t indicates that the actual emissions are below the allocated quota and therefore produces carbon-trading revenue, whereas a positive value represents excess emissions and results in a carbon payment. The step-type carbon-trading cost at time t is then formulated as
C CO 2 t ( P t ) = π CO 2 ( 1 + 2 ω ) ( Δ e t + v ) − π CO 2 ( 1 + ω ) v , Δ e t ≤ − v , π CO 2 ( 1 + ω ) Δ e t , − v < Δ e t ≤ 0 , π CO 2 Δ e t , 0 < Δ e t ≤ v , π CO 2 v + π CO 2 ( 1 + μ ) ( Δ e t − v ) , v < Δ e t ≤ 2 v , π CO 2 ( 1 + 2 μ ) ( Δ e t − 2 v ) + π CO 2 ( 2 + μ ) v , 2 v < Δ e t ≤ 3 v , π CO 2 ( 1 + 3 μ ) ( Δ e t − 3 v ) + π CO 2 ( 3 + 3 μ ) v , 3 v < Δ e t ≤ 4 v , π CO 2 ( 1 + 4 μ ) ( Δ e t − 4 v ) + π CO 2 ( 4 + 6 μ ) v , 4 v < Δ e t .
Here, π CO 2 denotes the base carbon-trading price, ω is the incentive coefficient applied when the emissions are below the free quota, μ is the price growth rate for successive excess-emission tiers, and v is the width of each carbon-emission interval. Equation (23) is algebraically continuous at the boundaries of the successive trading intervals. A negative value of C CO 2 t represents carbon-trading revenue, while a positive value represents the payment for excess emissions.
The carbon-trading mechanism is directly integrated into the online dispatch objective through
F t ( P t ) = ∑ i = 1 N C t i ( p t i ) + C CO 2 t ( P t ) .
Therefore, changing P t affects not only the conventional generation cost but also e f t , e p t , Δ e t , and consequently the carbon-trading cost. The carbon charge is thus included in the function value evaluated by the online dispatch procedure rather than being calculated only as a post-processing performance indicator. Over the complete dispatch horizon, the total low-carbon operating cost is
∑ t = 1 T F t ( P t ) = ∑ t = 1 T ∑ i = 1 N C t i ( p t i ) + ∑ t = 1 T C CO 2 t ( P t ) .
The subsequent theoretical results remain conditional on the convexity, boundedness, and Lipschitz-continuity requirements imposed on the resulting queried cost functions in Assumptions 2 and 3. Since the proposed algorithm accesses the composite low-carbon cost only through function-value queries, an analytical derivative of the piecewise carbon-trading function is not required.
In the practical context of a microgrid system, the exact formulation of the cost function may remain uncertain, complicating the computation of its gradient. Therefore, this paper assumes that the generator cost function is unknown and can only be queried through function values. A one-point bandit gradient estimator is then developed to replace exact gradient information. Additionally, the information exchange process among generator units poses a potential risk of privacy breaches, as eavesdroppers may infer private generator states or cost-related information from intercepted data. Therefore, this paper proposes a privacy-preserving distributed online economic dispatch algorithm for microgrids, which jointly handles unknown cost functions, TCN–BiLSTM-based PV forecasting uncertainty, and privacy leakage risks while maintaining online dispatch performance. Prior to presenting the specific algorithm for solving the EDP (14), the necessary assumptions and lemmas are given below.

2.5. Useful Assumptions and Lemmas

Assumption 1.
The communication topology graph G = ( V , E ) is strongly connected. The matrices A r and A c are compatible with G , where A r is row-stochastic and A c is column-stochastic, i.e.,
∑ j = 1 N [ A r ] i j = 1 , ∑ i = 1 N [ A c ] i j = 1 .
Moreover, all positive weights are uniformly lower-bounded, and the augmented matrix
A = A r ϵ I I − A r A c − ϵ I
satisfies the geometric mixing property stated in Lemma 2, where ϵ > 0 is the auxiliary coupling parameter in the dispatch recursion.
Assumption 2.
The feasible set Ω ⊂ R d is nonempty, compact, and convex. There exists a constant ρ > 0 such that ∥ p ∥ ≤ ρ for all p ∈ Ω . For each generator i ∈ V and time t, the local cost function f t i is convex on the extended feasible region containing Ω + δ S d and is L-Lipschitz continuous. Furthermore, there exists a constant M > 0 such that | f t i ( p ) | ≤ M for all p ∈ Ω + δ S d , for all i ∈ V , and for all t.
Assumption 3.
There exists a constant D ^ > 0 such that the subgradients of each local cost function are uniformly bounded, i.e.,
∂ f t i ( p ) ≤ D ^ , ∀ p ∈ Ω , ∀ i ∈ V , ∀ t .
Assumption 4.
For each generator i and time t, the random direction u t i used in the one-point bandit estimator is independently generated from the unit sphere S d and is independent of the filtration F t . For the convergence analysis, the per-release privacy-budget sequence satisfies  
0 < ε ̲ ≤ ε t , ∀ t ≥ 0 ,
where ε ̲ is independent of t. Conditional on F t , the coordinates of the privacy-noise vector η t i are independently drawn from a Laplace distribution with zero mean and scale σ t = Δ t / ε t , and the noise vectors are independent across generators and dispatch times. Here, Δ t is the global ℓ 1 -sensitivity defined in Definition 3, and ε t is the privacy budget assigned to the time-t release. Consequently,
E η t i F t = 0 , E ∥ η t i ∥ F t ≤ 2 d σ t .
Assumption 5.
For a given confidence parameter α ∈ ( 0 , 0.5 ) , the TCN–BiLSTM-based probabilistic PV forecasting model F θ satisfies the following prediction interval coverage condition:
P p p v t ∈ p ^ p v ( α ) , t , p ^ p v ( 1 − α ) , t ≥ 1 − 2 α , t = 1 , … , T .
Furthermore, define the cumulative PV forecasting uncertainty over the time horizon T as
U p v , T = ∑ t = 1 T p p v t − p ^ p v ( 0.5 ) , t + ∑ t = 1 T R p v t .
There exists a constant B p v > 0 such that the cumulative dispatch cost perturbation caused by TCN–BiLSTM-based PV forecasting uncertainty is bounded by
∑ t = 1 T Δ f p v , t ≤ B p v U p v , T ,
where Δ f p v , t denotes the cost variation caused by the mismatch between the actual PV output and the TCN–BiLSTM-based probabilistic PV forecasting-driven dispatch model at time t.
Remark 5.
Assumption 5 concerns the output quality of the trained TCN–BiLSTM forecaster rather than global optimality of its training problem. The forecasting module affects the dispatch analysis only through the median forecast p ^ p v ( 0.5 ) , t and the lower-side reserve R p v t , summarized by U p v , T .
Then, the definition and properties of the one-point bandit gradient estimator are given. Suppose that the objective function of problem (14) is f : R d → R . The one-point bandit gradient estimator is formulated as
∇ ˜ f ( z ) = d δ f ( z + δ u ) u ,
where δ > 0 represents the exploration parameter, and u ∈ S d is an independently generated unit random direction.
Lemma 1
([17]). For δ > 0 , define the smoothed function as
f ^ t i ( p ) = E u t + 1 i ∈ S d f t i ( p + δ u ) .
Under Assumption 2 and Assumption 4, the one-point bandit gradient estimator satisfies the following properties: (a) The estimator is unbiased for the smoothed gradient:
E u t + 1 i ∈ S d ∇ ˜ f t i ( p t i ) = ∇ f ^ t i ( p t i ) .
(b) The estimator is uniformly bounded:
∇ ˜ f t i ( p t i ) ≤ d M δ .
(c) The smoothing error is bounded:
f ^ t i ( p t i ) − f t i ( p t i ) ≤ δ L , ∀ p t i ∈ Ω .
Lemma 2
([16]). If Assumption 1 is satisfied and the constant ϵ satisfies
0 < ϵ < 1 ( 20 + 8 N ) N ( 1 − | λ 3 | ) N ,
then, for all i , j ∈ { 1 , … , 2 N } and t ≥ 1 , there exist constants Γ > 0 and κ ∈ ( 0 , 1 ) such that
A t − 1 N 1 N 1 N T 1 N 1 N 1 N T 0 0 ∞ ≤ Γ κ t ,
where λ 3 is the third largest eigenvalue of the augmented matrix A.
Typically, the efficacy of online optimization algorithms is evaluated using dynamic regret. For any generator j ∈ V , the individual dynamic regret is defined as
Reg T j = ∑ t = 1 T ∑ i = 1 N f t i ( p t j ) − ∑ t = 1 T ∑ i = 1 N f t i ( p t ⋆ ) ,
where p t ⋆ denotes the dynamic optimal solution associated with the risk-aware net demand P net t . If
lim T → ∞ Reg T j T = 0 ,
then the online algorithm achieves sublinear average dynamic regret.
In the following section, the proposed algorithm along with its convergence analysis will be presented.

3. PPF-DP-DOBO and Its Convergence Analysis

3.1. TCN–BiLSTM Probabilistic PV Forecasting-Driven DP-DOBO Algorithm

Algorithm 1 separates uncertainty estimation from online dispatch. At each time step, the trained TCN–BiLSTM model maps the current input window to PV quantiles. The median forecast estimates the available PV contribution, and the lower quantile determines the reserve against underproduction. Their combination produces the time-varying risk-aware net demand in (45). The dispatch analysis does not depend on the internal network architecture once these forecast outputs are available.
The second part of each iteration addresses unknown costs and private communication. Generator i evaluates its cost once at a randomly perturbed point and uses the resulting one-point estimator in (47). Before communication, the generator adds Laplace noise calibrated to the time-t sensitivity and privacy budget. The projected update keeps the released decision within the generator limits. The risk-aware net demand specifies the balance target in (15); the numerical study uses balance-feasible schedules and reports the aggregate tracking error explicitly.
Algorithm 1 PPF-DP-DOBO with TCN–BiLSTM probabilistic PV forecasts
  • Input: Trained quantile forecaster F θ ; quantile set Q = { q 1 , … , q m } ; lower-tail risk parameter α ∈ Q ∩ ( 0 , 0.5 ) ; PV input windows { X t } t = 0 T ; load profile { p D t } t = 0 T ; scheduled grid exchange { p t grid } t = 0 T ; initial states p 0 i = 0 and y 0 i = 0 ; exploration radius δ > 0 ; auxiliary coupling parameter ϵ > 0 ; per-release privacy budgets { ε t } t = 0 T ; and step sizes { γ t } t = 0 T .
  • For  t = 0 to T
  • Forecast and uncertainty mapping:
  • Generate the PV quantiles and lower-side reserve:
    p ^ p v t = F θ ( X t ) ,
    p ^ p v ( q r ) , t = F θ ( X t ) r , r = 1 , … , m ,
    R p v t = p ^ p v ( 0.5 ) , t − p ^ p v ( α ) , t + .
  •     Compute the risk-aware net demand:
    P net t = p D t − p ^ p v ( 0.5 ) , t + R p v t .
  •     Privacy-preserving one-point bandit dispatch:
  • Each generator i draws η t i ∼ Lap ( 0 , Δ t / ε t ) , communicates s t i = p t i + η t i , and queries one function value to form
    s t i = p t i + η t i ,
    ∇ ˜ f t i ( p t i ) = d δ f t i ( p t i + δ u t i ) u t i .
  •     Update the dispatch and auxiliary states:
    p t + 1 i = Π Ω ∑ j = 1 N [ A r ] i j s t j + ϵ y t i − γ t ∇ ˜ f t i ( p t i ) ,
    y t + 1 i = p t i − ∑ j = 1 N [ A r ] i j p t j + ∑ j = 1 N [ A c ] i j y t j − ϵ y t i .
  • End For
  • Output: PV quantile forecasts, lower-side reserves, risk-aware net demands, and feasible generator outputs.
The privacy budget ε t and the auxiliary coupling coefficient ϵ have different roles. The former controls the Laplace scale and the privacy–performance trade-off, whereas the latter controls the augmented consensus recursion. Specifically, decreasing ε t strengthens the per-release privacy guarantee but increases the injected noise through σ t = Δ t / ε t , while increasing ε t has the opposite effect. Moreover, ε t should not be interpreted as the privacy budget of the complete dispatch process, because the horizon-level budget is obtained by accumulating all per-release budgets according to (3). The lower bound ε t ≥ ε ̲ in Assumption 4 is imposed for the convergence analysis so that the privacy-noise moments remain uniformly controlled; the differential privacy guarantee itself remains valid for every individual release with any ε t > 0 . For the inverse-square-root step-size schedules used in the finite-horizon analysis below, write γ t = c γ / t + 1 with a fixed c γ > 0 . These schedules satisfy
∑ t = 0 ∞ γ t = ∞ , ∑ t = 0 T γ t 2 ≤ c γ 2 [ 1 + ln ( T + 1 ) ] , γ t ≤ γ s , ∀ t > s ≥ 0 .
Next, we reformulate the DP-DOBO-based dispatch stage in a compact form for the subsequent convergence analysis. Define the augmented variables h t i and perturbation terms v t i . Then, the dispatch recursion can be written as
h t + 1 i = ∑ j = 1 2 N [ A ] i j h t j + v t i .
For i ∈ { 1 , … , N } , let h t i = p t i and
v t i = p t + 1 i − ∑ j = 1 N [ A r ] i j p t j − ϵ y t i .
For i ∈ { N + 1 , … , 2 N } , let h t i = y t i − N and v t i = 0 . Since p t i is projected onto Ω at each iteration, it holds that ∥ p t i ∥ ≤ ρ for all i ∈ V and t ≥ 0 . In addition, define
h ¯ t = 1 N ∑ i = 1 2 N h t i = 1 N ∑ i = 1 N p t i + 1 N ∑ i = 1 N y t i ,
which represents the mean of p t i + y t i across all agents at time t. Furthermore, the matrix A is constructed by merging the submatrices A r and A c as follows:
A = A r ϵ I I − A r A c − ϵ I .
Remark 6.
It is worth noting that the injection of differential privacy noise may cause the perturbed variable s t i to temporarily violate the power constraint. However, as shown in (48), the decision variable p t + 1 i is projected onto the feasible set in each iteration, ensuring that the constraint p min i ≤ p t i ≤ p max i is always satisfied. Meanwhile, the difference between the actual carbon emissions and the allocated free quota is converted into the interval carbon-trading cost through (23), and this cost is directly included in the time-varying objective through (24). In addition, the TCN–BiLSTM-based probabilistic PV forecasting results are incorporated into the dispatch process through P net t , where the TCN–BiLSTM-generated median forecast reduces the conventional generation burden and the reserve term R p v t enhances robustness against PV underproduction.

3.2. Convergence Analysis

This subsection presents the main theoretical results of the proposed PPF-DP-DOBO algorithm. The TCN–BiLSTM-based probabilistic PV forecasting module provides quantile forecasts before the dispatch recursion is executed. Since the dispatch-stage privacy mechanism, the one-point bandit estimator, and the distributed projection recursion do not depend on the internal architecture of the TCN–BiLSTM forecaster, the main privacy and convergence analysis of the DP-DOBO recursion can be preserved. The TCN–BiLSTM forecasting error and the lower-side reserve are treated as exogenous time-varying perturbations in the dynamic regret analysis. For completeness, recall the cumulative PV forecasting uncertainty defined in Assumption 5:
U p v , T = ∑ t = 1 T p p v t − p ^ p v ( 0.5 ) , t + ∑ t = 1 T R p v t .
The first term in U p v , T measures the cumulative deviation between the actual PV output and the TCN–BiLSTM-generated median forecast, while the second term represents the cumulative lower-side reserve introduced to hedge against PV underproduction.
The network dynamic regret under TCN–BiLSTM-based probabilistic PV forecasting-driven dispatch is defined as
Reg T net = ∑ t = 1 T ∑ i = 1 N f t i ( p t i ) − ∑ t = 1 T ∑ i = 1 N f t i ( p t ⋆ ) ,
where p t ⋆ denotes the dynamic optimal solution associated with the risk-aware net demand P net t .
Before presenting the main theorems, several crucial lemmas are given. Since the TCN–BiLSTM-based PV forecasting module is incorporated into the dispatch problem only through P net t , Lemmas 3–5 are independent of the CNN architecture and remain valid under Assumptions 1–4.
Lemma 3.
Suppose Assumptions 1–4 hold. Let ϵ satisfy
ϵ ≤ min ϵ ¯ , 1 − κ 2 N Γ κ , ϵ ¯ = 1 ( 20 + 8 N ) N ( 1 − | λ 3 | ) N .
Then, there exists a positive constant G such that
∑ j = 1 N E v t j F t ≤ G γ t , ∀ t ≥ 0 .
Lemma 4.
Suppose Assumptions 1–4 hold, and let the sequence { h t i } t ≥ 0 be generated by the dispatch stage of Algorithm 1. Then, for each i ∈ { 1 , … , N } and t ≥ 1 , it holds that
E h t i − h ¯ t F t − 1 ≤ 2 N ρ Γ ^ κ t + G Γ ^ ∑ r = 1 t κ t − r γ r − 1 ,
where Γ ^ = max { Γ , 1 } .
Lemma 5.
Suppose Assumptions 1–4 hold. Then, the global sensitivity of the dispatch mechanism satisfies
Δ t ≤ 2 M γ t d 3 2 δ .
Lemma 6.
Suppose Assumptions 1–5 hold, and let the step size be chosen as
γ t = 1 t + 1 .
Then, for T > 0 , the network dynamic regret Reg T net satisfies
E Reg T net ≤ C 1 + C 2 ( T + 1 ) 1 2 + 2 ( T + 1 ) N δ L + 4 N ρ V γ t − 1 p + B p v E U p v , T ,
where
C 1 = 2 G + 3 d M δ + D ^ N 2 ρ Γ ^ 1 − κ ,
and
C 2 = V T + 4 N ρ 2 + 2 G + 3 d M δ + D ^ N G Γ ^ κ ( 1 − κ ) + 2 N G G + d N M δ .
Here,
V γ t − 1 p = ∑ t = 0 T 1 γ t p t + 1 ⋆ − p t ⋆
denotes the 1 / γ t -weighted path variation of the dynamic optimal trajectory.
Proof. 
The proof follows the network regret analysis of the DP-DOBO dispatch recursion. The key difference is that the supply–demand balance is determined by the risk-aware net demand P net t , which depends on the TCN–BiLSTM-generated median PV forecast p ^ p v ( 0.5 ) , t and the forecast-derived lower-side reserve R p v t . By Assumption 5, the cumulative dispatch cost perturbation caused by the TCN–BiLSTM forecasting error and the reserve requirement over the horizon T is bounded by B p v U p v , T . Therefore, adding this perturbation term to the original network dynamic regret bound yields (62). □
Subsequently, we present the main results concerning ε -differential privacy and individual dynamic regret.
Theorem 1.
Suppose that Assumptions 1–4 hold. At time t, implement the dispatch release with Laplace scale
σ t = Δ t ε t ,
where ε t > 0 . The joint released state vector s t = [ ( s t 1 ) T , … , ( s t N ) T ] T at time t is ε t -differentially private. By sequential composition, the releases up to time T are
∑ t = 0 T ε t - differentially private .
Thus, a constant per-release budget ε t = ε gives a horizon-level budget of ( T + 1 ) ε .
More generally, for a prescribed finite-horizon privacy budget ε ¯ , any allocation satisfying ∑ t = 0 T ε t ≤ ε ¯ guarantees that the complete release sequence is ε ¯ -differentially private.
Proof. 
Condition on an arbitrary realization of the release history up to time t. Let
p t = [ ( p t 1 ) T , … , ( p t N ) T ] T = q t ( Θ )
and
p t ′ = [ ( p ′ t 1 ) T , … , ( p ′ t N ) T ] T = q t ( Θ ′ )
be the state queries induced by two adjacent inputs Θ ∼ Θ ′ . By Definition 3,
∥ p t − p t ′ ∥ 1 ≤ Δ t .
For an arbitrary possible released vector s , let g t ( s ∣ p t ) denote the conditional joint density of the Laplace mechanism. Because the coordinates of the noise vectors are independent and have scale σ t = Δ t / ε t , the likelihood ratio satisfies
g t ( s ∣ p t ) g t ( s ∣ p t ′ ) = exp ε t Δ t ∥ s − p t ′ ∥ 1 − ∥ s − p t ∥ 1 ≤ exp ε t Δ t ∥ p t − p t ′ ∥ 1 ≤ exp ( ε t ) ,
where the first inequality follows from the triangle inequality. Integrating the pointwise density inequality over any measurable output set ψ gives
P M t ( Θ ) ∈ ψ ≤ e ε t P M t ( Θ ′ ) ∈ ψ .
Therefore, the time-t joint release is ε t -differentially private according to Definition 2. Finally, applying sequential composition to the possibly adaptive release sequence s 0 , … , s T yields the horizon-level privacy budget ∑ t = 0 T ε t . This completes the proof. □
Theorem 2.
Suppose that Assumptions 1–5 hold and the step size is chosen as γ t = 1 / t + 1 . Then, for any j ∈ V and T > 0 , the individual dynamic regret Reg T j generated by the dispatch stage of PPF-DP-DOBO satisfies
E Reg T j ≤ C ˜ 1 + C ˜ 2 ( T + 1 ) 1 2 + 2 N δ L ( T + 1 ) + 4 N ρ V γ t − 1 p + B p v E U p v , T ,
where
C ˜ 1 = V 4 , T + 2 G + 5 d M δ + D ^ N 2 ρ Γ ^ 1 − κ ,
and
C ˜ 2 = 4 N ρ 2 + V T + 2 G + 5 d M δ + D ^ N G Γ ^ κ ( 1 − κ ) + 2 N G G + d N M δ .
Proof. 
First, denote
Reg ^ T net = ∑ t = 0 T f ^ t ( h t i ) − ∑ t = 0 T f ^ t ( h t ⋆ ) , Reg ^ T j = ∑ t = 0 T f ^ t ( h t j ) − ∑ t = 0 T f ^ t ( h t ⋆ ) ,
where h t ⋆ corresponds to the dynamic optimal solution associated with P net t . Then, one has
E Reg ^ T j − E Reg ^ T net = E ∑ t = 0 T f ^ t ( h t j ) − f ^ t ( h t i ) ≤ E ∑ t = 0 T ∇ f ^ t ( h t j ) T h t j − h t i ≤ ∑ t = 0 T d N M δ E h t j − h t i + V 4 , T ,
where V 4 , T denotes the accumulated covariance term
V 4 , T = ∑ t = 0 T Cov ∇ f ^ t ( z t j ) , h t j − h t i .
Using Lemma 4, we have
E h t j − h t i ≤ E h t j − h ¯ t + h ¯ t − h t i ≤ 4 N ρ Γ ^ κ t + 2 G Γ ^ ∑ r = 1 t κ t − r γ r − 1 .
Combining (69) and (70), it follows that
E Reg ^ T j − E Reg ^ T net ≤ d N M δ 4 N ρ Γ ^ 1 − κ + 2 G Γ ^ κ ( 1 − κ ) ∑ t = 0 T γ t + V 4 , T .
Finally, using Lemma 6, Lemma 4, and (71), we obtain
E Reg ^ T j ≤ V T + 4 N ρ 2 γ T + G + 5 d M δ + D ^ N G Γ ^ κ ( 1 − κ ) + G G + d N M δ N ∑ t = 0 T γ t + 2 G + 5 d M δ + D ^ N 2 ρ Γ ^ 1 − κ + 4 N ρ V γ t − 1 p + 2 ( T + 1 ) N δ L + V 4 , T + B p v E U p v , T .
Substituting γ t = 1 / t + 1 into (72) gives (66). □
Theorem 2 establishes an individual dynamic regret bound for each generator under TCN–BiLSTM-driven dispatch. Relative to the corresponding bound without PV uncertainty, the term B p v E [ U p v , T ] quantifies the contribution of median-forecast error and the reserve requirement. Thus, the trained forecaster enters the online analysis through an explicit exogenous uncertainty term rather than through its network parameters.
Corollary 1.
Under the same conditions as Theorem 2, let
δ * = ξ ( T + 1 ) 1 2 + ξ 2 ( T + 1 ) + 8 N L ϑ ( T + 1 ) 3 2 4 N L ( T + 1 ) .
If V γ t − 1 p ≤ O ( T 3 / 4 ) and E [ U p v , T ] ≤ O ( T 3 / 4 ) , then
E Reg T j ≤ O ( T 3 4 ) ,
where
ξ = 4 N ρ 2 + V T + 2 ( G + D ^ ) N G Γ ^ κ ( 1 − κ ) + 2 N G 2 ,
and
ϑ = 10 d M N G Γ ^ κ ( 1 − κ ) + 2 d M G N .
Proof. 
Recall that the regret bound is provided in Theorem 2. By balancing the terms related to the exploration parameter δ , we set
ξ + ϑ δ ( T + 1 ) 1 2 = 2 N δ L ( T + 1 ) .
Solving (77) yields
δ * = ξ ( T + 1 ) 1 2 + ξ 2 ( T + 1 ) + 8 N L ϑ ( T + 1 ) 3 2 4 N L ( T + 1 ) .
If V γ t − 1 p ≤ O ( T 3 / 4 ) and E [ U p v , T ] ≤ O ( T 3 / 4 ) , then substituting δ * into (66) yields
E Reg T j ≤ O ( T 3 4 ) .
This completes the proof. □
Remark 7.
Corollary 1 shows that the proposed PPF-DP-DOBO algorithm preserves the sublinear individual dynamic regret rate when both the 1 / γ t -weighted path variation V γ t − 1 p and the cumulative TCN–BiLSTM-based PV forecasting uncertainty E [ U p v , T ] grow no faster than O ( T 3 / 4 ) . This condition is reasonable for TCN–BiLSTM-based probabilistic PV forecasting-driven dispatch, because a highly volatile optimal trajectory or an excessively large PV forecasting error would make accurate online tracking theoretically difficult. Under this condition, the average individual dynamic regret E [ Reg T j ] / T converges to zero as T increases, which implies that each generator can asymptotically track the dynamic optimal dispatch solution while accounting for TCN–BiLSTM-based PV forecasting uncertainty, differential privacy, and low-carbon operation.

4. Numerical Experiments

4.1. Experimental Setup

The numerical study evaluates four links in the proposed framework: PV forecasting, uncertainty-to-demand mapping, distributed online dispatch, and the privacy–performance trade-off. The dispatch case uses a modified IEEE 162-bus system with 10 conventional DGs and seven representative loads. Consistent with the formulation in Section 2.4, the reported numerical case imposes the generation-capacity bounds in Table 3 but does not include inter-temporal generator ramp-rate constraints. The experiments are numerical case studies rather than a field deployment, so the conclusions are limited to the reported topology, profiles, and parameter settings.
Table 3. The cost factors a i , b i , c i and generation limits of DGs [33].
The 10 controllable resources in this case are conventional DGs, not 10 PV units. PV generation is represented by one aggregate exogenous forecast profile and enters the dispatch model through the risk-aware net demand P net t . Because the present study does not solve a bus-level power-flow problem or impose line-flow constraints, the aggregate PV profile is not assigned to a particular bus of the IEEE 162-bus system.
The same aggregate dispatch test system is considered in three scenarios: no PV, deterministic PV using the median forecast, and probabilistic PV using the median forecast plus a lower-side reserve. Section 4.3 reports these comparisons explicitly: Figures 5 and 6 show generator trajectories and supply–demand tracking for the probabilistic-PV scenario. Figure 7 gives the average dynamic regret and Figure 8 gives the carbon-trading cost.
The forecasting and dispatch stages are evaluated sequentially. The trained TCN–BiLSTM produces the PV point and interval forecasts shown below. Its median and lower quantile are then treated as exogenous inputs to PPF-DP-DOBO. This separation isolates the effect of forecast uncertainty on dispatch and is consistent with the theoretical analysis, which depends on forecasting outputs through U p v , T rather than on neural-network parameters.
The cost parameters in Table 3 are used only to return queried function values; the algorithm does not evaluate analytic gradients. The auxiliary coupling coefficient is ϵ = 0.01 , the local decision dimension is d = 1 , and δ = 1 is used as the baseline exploration radius in the main numerical comparisons; its sensitivity with respect to generator scale is examined separately in Section 4.4. The step size is
γ t = c t + 1 , c = 0.1 .
Here, c = 0.1 is the fixed step-size scale used in the numerical study, giving an initial step size of γ 0 = 0.1 . This scale has a precedent in distributed resource optimization: Hao et al. [34] use the diminishing subgradient step size v k = 0.1 / k , which has the same form after reindexing k = t + 1 . For the same one-point gradient estimate, the factor 0.1 makes the gradient-based correction one tenth of that under the unit-scale schedule 1 / t + 1 , thereby moderating the update magnitude. The coefficient controls the scale of the update, whereas the denominator controls its inverse-square-root decay; 0.1 is therefore a numerator coefficient. It is a literature-supported numerical setting, not an optimal or universal constant. The regret bounds in Section 3 are written for unit scale; a fixed positive scale changes the constants associated with ∑ t γ t and 1 / γ T , but not their T order. The scheduled grid exchange is set to p t grid = 0 in the reported case, and the initial conventional outputs sum to the initial risk-aware net demand. Figure 1 shows the strongly connected communication topology. The row-stochastic matrix A r and column-stochastic matrix A c are constructed as
[ A r ] i j = 1 d i , in , j ∈ N i in , 0 , j ∉ N i in ,
and
[ A c ] i j = 1 d j , out , i ∈ N j out , 0 , i ∉ N j out ,
where N i in denotes the in-neighbor set of DG i, N j out denotes the out-neighbor set of DG j, and d i , in and d j , out represent the corresponding in-degree and out-degree, respectively.
Figure 1. Communication topology diagram for 10 DGs.

4.2. Probabilistic PV Forecasting and Risk-Aware Demand

The evaluated forecasting sequence contains the actual PV output p p v t , the median forecast p ^ p v ( 0.5 ) , t , and the lower and upper quantiles p ^ p v ( α ) , t and p ^ p v ( 1 − α ) , t . In the reported case study, the lower-tail risk parameter is fixed at α = 0.05 before the online dispatch process. Accordingly, the fifth-percentile PV forecast is used as the lower forecast bound, corresponding to a nominal lower-side reliability of 1 − α = 0.95 . Together with the upper quantile p ^ p v ( 0.95 ) , t , it forms a central prediction interval with a nominal coverage of 1 − 2 α = 0.90 . This setting is adopted to provide protection against PV underproduction without introducing an excessively conservative reserve requirement.
The interval between two consecutive samples is 5 min; therefore, each complete day contains 288 samples. The time axes in the PV forecasting and prediction-interval figures use this same sampling interval.
The lower-side reserve is computed as
R p v t = p ^ p v ( 0.5 ) , t − p ^ p v ( α ) , t + .
Accordingly, the dispatch target is transformed from the original load demand p D t into the risk-aware net demand
P net t = p D t − p ^ p v ( 0.5 ) , t + R p v t .
The conventional DGs therefore track P net t rather than the original load p D t .
To further characterize the meteorological information contained in X t , a supplementary correlation analysis was performed using the 5 min PV and meteorological measurements from the DKASC Alice Springs DKA-M5 C-Phase data over 2020–2022. To avoid artificially increasing the correlations because of the large number of simultaneous zero-PV and zero-irradiance observations during nighttime, only daylight samples satisfying G t GHI > 20 W / m 2 and p p v t > 0.01 kW were retained. This resulted in 138,048 valid samples. Both Pearson and Spearman correlation coefficients were calculated to quantify, respectively, the linear and monotonic associations between each meteorological variable and the measured PV output. The results are reported in Table 4.
Table 4. Correlation between meteorological input features and measured PV output.
As shown in Table 4, GTI and GHI exhibit the strongest positive associations with the measured PV output, with Pearson coefficients of 0.9811 and 0.9626, respectively. Their corresponding Spearman coefficients are also high, indicating that these relationships remain strong under a rank-based measure. Relative humidity exhibits a negative association with PV generation, whereas DHI and ambient temperature show weaker marginal correlations. These pairwise coefficients are used to characterize the association of individual meteorological variables with PV output rather than to represent model-specific feature importance; weaker pairwise correlation does not imply that a variable is uninformative within the nonlinear multivariate temporal forecasting model.
Figure 2 compares the TCN–BiLSTM prediction with the observed PV profile. On the evaluated sequence, the reported R 2 , RMSE, and MAE are 0.9629 , 55.31 MW, and 29.75 MW, respectively. The enlarged region also shows visible errors during several rapid ramps, so the forecast is informative but not exact.
Figure 2. Prediction performance of the proposed TCN-BiLSTM-based PV forecasting model. The figure compares the actual PV output and predicted values, including the overall forecasting results, the enlarged daytime high-power period, and the actual-predicted correlation analysis.
To further validate the forecasting model, three representative baselines—ARIMA, CNN, and GRU—were implemented using the same dataset division and evaluation metrics. Table 5 reports the resulting RMSE, MAE, and pinball loss.
Table 5. Performance comparison among different forecasting models. Lower values indicate better performance.
In this comparison, TCN–BiLSTM achieved the lowest error for all three metrics: its RMSE, MAE, and pinball loss were 55.31 MW, 29.75 MW, and 0.0328, respectively. These results support the use of the hybrid architecture for the evaluated sequence and are consistent with its design rationale: the TCN extracts multi-scale temporal patterns, whereas the BiLSTM represents bidirectional temporal dependencies. The comparison is limited to the reported dataset split and does not establish universal superiority across sites or operating conditions.
Figure 3 visualizes the associated prediction bands. These bands are used as an uncertainty diagnostic; the dispatch reserve itself is computed from the lower quantile in (13). Figure 4 then shows the resulting forecast-to-dispatch mapping. The median forecast lowers the conventional-generation requirement, whereas R p v t partially offsets that reduction when lower-tail uncertainty is larger. The horizontal axis in Figure 4, Figure 5, Figure 6, Figure 7, Figure 8, Figure 9 and Figure 10 is labeled Iterations(t), where t is the dispatch iteration index.
Figure 3. Probabilistic PV forecasts and prediction intervals generated by the TCN–BiLSTM model. The intervals characterize the uncertainty range used to construct the lower-side dispatch reserve.
Figure 4. Construction of the risk-aware net demand under probabilistic PV forecasting.

4.3. Dispatch Tracking and Low-Carbon Performance

Figure 5 shows the output trajectories of the 10 conventional generators. All reported trajectories remain within the limits in Table 3 and respond to the time-varying dispatch target despite perturbations in the exchanged states.
Figure 5. Output power trajectories of conventional generators under probabilistic PV forecasting-driven dispatch.
Because the load, PV contribution, and queried operating costs vary with time, the optimal dispatch is a time-indexed trajectory rather than a single static vector of generator outputs. Table 3 reports the cost coefficients and admissible output range of each controllable DG, while Figure 5 reports the corresponding optimized power trajectories over the evaluated horizon. PV is not optimized as an additional DG; its forecasted aggregate contribution is incorporated through P net t .
Figure 6 compares aggregate conventional generation with P net t . The two curves are visually coincident in the reported zero-grid-exchange case, showing that the balance-feasible schedule follows the target constructed from the median forecast and reserve. This experiment verifies numerical tracking for the tested profile; it does not by itself establish robustness to unmodeled network or device constraints.
Figure 6. Supply–demand balance between total conventional generation and risk-aware net demand.
Figure 7 and Figure 8 present the with-PV and without-PV comparisons for the modified IEEE 162-bus dispatch case introduced in Section 4.1. The no-PV scenario requires the conventional DGs to meet the original load demand. The deterministic-PV scenario subtracts the median PV forecast from the load, without adding an uncertainty reserve. The probabilistic-PV scenario uses the median forecast together with the lower-side reserve to define the risk-aware net demand. Thus, the no-PV curve is the common baseline for the two PV-assisted scenarios in both figures. These are comparisons within the aggregate economic-dispatch model; they do not constitute a bus-level power-flow validation.
Figure 7. Average dynamic regret for the no-PV, deterministic-PV, and probabilistic-PV scenarios in the modified IEEE 162-bus dispatch case.
Figure 8. Carbon-trading costs for the no-PV, deterministic-PV, and probabilistic-PV scenarios in the modified IEEE 162-bus dispatch case.
Figure 7 compares the reported average-regret values for the three scenarios. Over the displayed 100 iterations, all three curves increase and begin to flatten; therefore, this finite-horizon plot should not be described as evidence of decreasing average regret. The probabilistic-PV curve remains below the no-PV and deterministic-PV curves throughout most of the evaluated horizon. The asymptotic O ( T 3 / 4 ) statement follows from Corollary 1 under its path-variation and forecasting-uncertainty conditions, rather than from this finite simulation alone.
Figure 8 shows that both PV-assisted cases reduce carbon-trading cost relative to the no-PV case over most of the horizon. The deterministic-PV case is usually lowest because it carries no lower-tail reserve. The probabilistic-PV case incurs a moderate reserve-related cost but remains below the no-PV case, illustrating the economic trade-off between renewable utilization and protection against PV underproduction.

4.4. Online Optimization and Privacy–Performance Trade-Off

Figure 9 compares PPF-DP-DOBO with RGF-DPGD [14], D-DPS [35], and D-OOA [36]. D-DPS uses exact gradients, RGF-DPGD uses two function queries, and D-OOA and PPF-DP-DOBO use one. The exact-gradient and two-point methods show faster finite-horizon reduction in regret. The proposed method pays an additional performance cost for Laplace perturbation, but it requires only one function query and protects released states. The comparison therefore illustrates an information–privacy–performance trade-off rather than uniform dominance over all baselines.
Figure 9. Average individual dynamic regret over time for the proposed method, RGF-DPGD, D-DPS, and D-OOA.
Figure 10 varies the constant per-release budget ε t = ε . A larger value reduces the Laplace scale and improves the reported regret, but it weakens privacy. By Theorem 1, the corresponding horizon-level budget is ( T + 1 ) ε unless a nonuniform budget schedule or another composition strategy is used. Privacy parameters should therefore be selected at the horizon level rather than interpreted only per iteration.
Figure 10. Impact of the privacy budget ε on the average dynamic regret over time.
To examine the sensitivity of PPF-DP-DOBO to the exploration radius and to determine whether a fixed radius remains appropriate when the generator capacity changes, an additional scale-sensitivity experiment is conducted. Since the same absolute exploration radius represents different relative perturbation magnitudes for generators with different operating ranges, we define the normalized exploration radius as
δ ¯ = δ p i , max ( s ) − p i , min ( s ) ,
where s denotes the generator-scale factor. Three capacity scales, s ∈ { 0.5 , 1 , 2 } , are considered. For each scale, the generation limits and the corresponding risk-aware net demand are scaled consistently as p i , min ( s ) = s p i min , p i , max ( s ) = s p i max , and P net t , ( s ) = s P net t , while the generation-cost coefficients and the other algorithmic settings are kept unchanged. The normalized radii δ ¯ ∈ { 0.05 , 0.10 , 0.25 , 0.50 , 0.75 , 0.90 , 1.00 } are evaluated, and each configuration is repeated over 20 independent random runs.
For cross-scale comparison, the cumulative dynamic regret is normalized by the cumulative magnitude of the corresponding optimal dispatch cost, i.e.,
NReg ( T ) = 100 Reg ( T ) ∑ t = 1 T ∑ i = 1 N f i t ( p i , t ⋆ ) ,
and is reported in percentage form.
Figure 11 shows that the algorithm is sensitive to an excessively small exploration radius. For δ ¯ = 0.05 , the normalized cumulative dynamic regret is 8.8571 % , 8.2467 % , and 7.6956 % for the 0.5 × , 1 × , and 2 × generator scales, respectively. Increasing δ ¯ to 0.25 substantially reduces these values to 2.5412 % , 2.1838 % , and 2.1553 % , respectively. The performance subsequently enters a broad low-regret region. At δ ¯ = 0.50 , the corresponding values are only 2.0726 % , 1.9190 % , and 1.9924 % , and further increases in the normalized radius produce only small additional changes.
Figure 11. Sensitivity of the normalized cumulative dynamic regret to the normalized exploration radius under different generator scales. Error bars represent the mean ± standard deviation over 20 independent runs.
The results also show that a fixed absolute exploration radius is not scale invariant. In particular, the baseline value δ = 1 corresponds to normalized radii of approximately 0.00690 , 0.00345 , and 0.00172 for the 0.5 × , 1 × , and 2 × cases, respectively, and produces normalized cumulative regrets of 19.3101 % , 30.9603 % , and 46.3544 % . Hence, the same absolute value becomes progressively smaller relative to the generator operating range as the capacity scale increases. In contrast, scaling δ with the generator operating range produces similar low-regret behavior across the three cases.
The lowest observed regrets in the tested range occur at δ ¯ = 1.00 for the 0.5 × case and at δ ¯ = 0.75 for the 1 × and 2 × cases. However, the differences among the results for δ ¯ ∈ [ 0.50 , 1.00 ] are small relative to the run-to-run variability. Therefore, these values should not be interpreted as unique global optima. Rather, the experiment indicates a broad low-regret region once a moderate normalized exploration radius is reached. This observation is also consistent with Corollary 1, where the theoretically balanced exploration radius depends on problem-dependent constants rather than being a universal fixed value.
Overall, the simulations support three bounded conclusions. First, probabilistic forecasts can be mapped to a risk-aware net demand that is tracked by the reported schedule. Second, PV-assisted cases reduce carbon-trading cost relative to the no-PV case, while the reserve adds cost relative to deterministic PV dispatch. Third, privacy perturbation and one-point feedback reduce the information requirement but slow finite-horizon optimization. The simulations do not establish the asymptotic regret rate, which remains a conditional theoretical result.

5. Discussion

The main contribution of PPF-DP-DOBO is the explicit coupling of learned temporal uncertainty with privacy-aware online control. The TCN–BiLSTM forecaster does not enter the dispatch recursion through its internal weights. Instead, its outputs are reduced to a median forecast and a lower-side reserve, which determine the risk-aware net demand. This modular interface makes the dispatch analysis applicable to another probabilistic forecaster if it supplies comparable quantiles and satisfies Assumption 5.
The results also expose three trade-offs. First, the lower-side reserve protects against PV underproduction but increases conventional generation relative to median-only dispatch. Second, one-point feedback reduces each update to one function query but converges more slowly than exact-gradient or two-point alternatives in the finite-horizon comparison. Third, stronger privacy requires more noise and therefore increases regret. These trade-offs are intrinsic to the problem and should not be interpreted as implementation defects.
From a practical perspective, these results provide several design implications for privacy-aware microgrid dispatch. First, the probabilistic PV forecasts are used not only as prediction outputs but also as an operational interface that converts lower-tail forecast uncertainty into a reserve-aware net-demand target. This allows the reserve level to respond to the predicted PV uncertainty instead of relying on a fixed deterministic margin. Second, the one-point bandit architecture requires only function-value feedback and is therefore relevant when analytical cost gradients are unavailable, expensive to evaluate, or undesirable to disclose. Third, the privacy-budget results indicate that privacy protection and dispatch performance should be configured jointly over the dispatch horizon, since stronger privacy increases the perturbation magnitude and may degrade finite-horizon performance. Finally, the exploration-radius sensitivity study shows that a fixed absolute exploration radius is not scale-invariant; therefore, the exploration parameter should be calibrated with respect to the generator operating range when the framework is transferred to systems with different capacity scales.
Several limitations define the scope of the conclusions. Although the updated forecasting evaluation includes ARIMA, CNN, and GRU baselines under the same dataset division, it remains limited to one site and one evaluated sequence and does not include a distribution-shift test or a broader multi-site benchmark. The forecasting experiment reports one evaluated sequence and does not provide a multi-site benchmark, distribution-shift test, or comparison with alternative probabilistic forecasters. Prediction-interval quality should also be assessed with empirical coverage and interval-width metrics in addition to point-forecast error. The dispatch study uses a modified IEEE 162-bus numerical case and omits network losses, line-flow constraints, generator ramp-rate limits, storage dynamics, demand response, and reactive power. Incorporating generator ramp-rate limits would introduce inter-temporal feasibility constraints because the admissible output at time t would depend on the generator output at time t − 1 . This extension would require a time-dependent feasible set and a corresponding extension of the present online regret analysis. Moreover, the reported simulation assumes balance-feasible schedules; a fully decentralized feasibility-restoration mechanism for detailed power-flow constraints is outside the present model. These modeling simplifications also imply several risks for practical deployment. Forecasting errors or poorly calibrated prediction intervals may lead to an inaccurate reserve requirement and consequently increase the risk of PV underproduction being insufficiently compensated. In addition, a dispatch schedule that is feasible under the present aggregate power-balance model may not remain physically feasible when network losses, line-flow limits, generator ramp-rate constraints, and other device dynamics are explicitly considered. The differential-privacy mechanism also introduces a privacy–performance trade-off, since stronger privacy requires larger perturbations and may increase finite-horizon dispatch error or regret. Therefore, the reported numerical results should be interpreted as evidence under the stated model assumptions rather than as a guarantee of identical economic or carbon-reduction performance in field operation.
The theoretical guarantees are conditional on convex bounded costs, a strongly connected directed communication graph, bounded sensitivity, controlled path variation, and controlled cumulative forecast uncertainty. The privacy guarantee is per release, and sequential composition causes the total pure-differential-privacy budget to grow over the horizon. Future work should therefore examine adaptive privacy-budget allocation, tighter composition, changing communication graphs, distributionally robust forecast-to-dispatch coupling, and network-constrained validation with real operational data.

6. Conclusions

This paper connected TCN–BiLSTM probabilistic PV forecasting with privacy-preserving distributed online low-carbon dispatch. The median forecast and lower quantile define a risk-aware net demand, while PPF-DP-DOBO combines one-point bandit feedback with Laplace-perturbed state exchange. The analysis provides a per-release differential privacy guarantee and a conditional O ( T 3 / 4 ) individual dynamic regret bound containing an explicit cumulative forecasting-uncertainty term. In the modified IEEE 162-bus dispatch case, Figure 5 and Figure 6 show that the reported probabilistic-PV schedules track the risk-aware target. The with-PV and without-PV comparisons are reported in Figure 7 and Figure 8: the probabilistic-PV scenario exhibits lower reported average regret over most of the evaluated horizon, and both PV-assisted scenarios exhibit lower carbon-trading costs than the no-PV baseline over most of that horizon. The method and privacy-budget comparisons in Figure 9 and Figure 10 show the expected query and privacy trade-offs.
From a practical perspective, the results indicate that probabilistic PV forecasts can be translated into reserve-aware dispatch targets, one-point feedback can support online optimization when analytical gradients are unavailable, and privacy and exploration parameters should be selected jointly with the acceptable dispatch-performance degradation and generator operating scale.
These findings should be interpreted within the stated convex online model and numerical setting rather than as direct field-deployment guarantees. The observed carbon-cost reductions are case-dependent, and broader practical conclusions require further validation under detailed network constraints, generator dynamics, forecast-distribution shifts, and real operational data.

Author Contributions

Conceptualization, C.Z. and Z.Z.; methodology, C.Z.; software, C.Z.; validation, C.Z. and Z.Z.; investigation, C.Z.; resources, C.Z. and Z.Z.; supervision, Z.Z.; writing—original draft preparation, C.Z. and Z.Z.; writing—review and editing, C.Z. and Z.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by the National Natural Science Foundation of China under Grant U23B2061, the Open Fund under Grant WDZC20245250409, and the Shandong Provincial Natural Science Foundation under Grant ZR2024QF163.

Data Availability Statement

The data presented in this study are available in the Desert Knowledge Australia Solar Centre (DKASC) data repository at https://dkasolarcentre.com.au/download?location=alice-springs (accessed on 30 August 2026). These data were derived from the following resources available in the public domain: Desert Knowledge Australia Solar Centre (DKASC), Alice Springs BP Solar, 5.0 kW, poly-Si, Fixed, 2008 generation and meteorological dataset (accessed on 30 August 2026), https://dkasolarcentre.com.au/download?location=alice-springs.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Cheng, J.; Wang, L.; Pan, T. Optimized Configuration of Distributed Power Generation Based on Multi-Stakeholder and Energy Storage Synergy. IEEE Access 2023, 11, 129773–129787. [Google Scholar] [CrossRef] [Scilit]
  2. El-Rifaie, A.M.; Shaheen, A.M.; Tolba, M.A.; Smaili, I.H.; Moustafa, G.; Ginidi, A.R.; Elshahed, M.A. Modified Gradient-Based Algorithm for Distributed Generation and Capacitors Integration in Radial Distribution Networks. IEEE Access 2023, 11, 120899–120917. [Google Scholar] [CrossRef] [Scilit]
  3. Han, H.; Chen, X.; Liu, Z.; Liu, Y.; Su, M.; Chen, S. A Completely Distributed Economic Dispatching Strategy Considering Capacity Constraints. IEEE J. Emerg. Sel. Top. Circuits Syst. 2021, 11, 210–221. [Google Scholar] [CrossRef] [Scilit]
  4. Sun, L.; An, W.; Chen, Y.; Zhao, P.; Ding, D. An Overview of Distributed Economic Dispatch of Microgrids: Advances and Challenges. Syst. Sci. Control Eng. 2025, 13, 2467077. [Google Scholar] [CrossRef] [Scilit]
  5. Amnuaypongsa, W.; Wangdee, W.; Songsiri, J. Probabilistic solar power forecasting using multi-objective quantile regression. In 2024 18th International Conference on Probabilistic Methods Applied to Power Systems (PMAPS); IEEE: Piscataway, NJ, USA, 2024; pp. 1–6. [Google Scholar]
  6. Wu, Y.-K.; Phan, Q.-T.; Lo, H.-Y.; Zhan, K.-C.; Tan, W.-S. Weighted Quantile Regression-based Probabilistic Forecasting for Solar Photovoltaic Systems. IEEE Trans. Ind. Appl. 2026, 62, 5699–5715. [Google Scholar] [CrossRef] [Scilit]
  7. Boroumandfar, G.; Khajehzadeh, A.; Eslami, M. A Single and Multiobjective Robust Optimization of a Microgrid in Distribution Network Considering Uncertainty Risk. Sci. Rep. 2024, 14, 28195. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Yan, J.; Cao, J.; Cao, Y. Distributed Continuous-Time Algorithm for Economic Dispatch Problem Over Switching Communication Topology. IEEE Trans. Circuits Syst. II Express Briefs 2021, 68, 2002–2006. [Google Scholar] [CrossRef] [Scilit]
  9. Guo, F.; Li, G.; Wen, C.; Wang, L.; Meng, Z. An Accelerated Distributed Gradient-Based Algorithm for Constrained Optimization with Application to Economic Dispatch in a Large-Scale Power System. IEEE Trans. Syst. Man Cybern. Syst. 2021, 51, 2041–2053. [Google Scholar] [CrossRef] [Scilit]
  10. Mi, Y.; Yuan, D.; Ratnam, E.L.; Verbic, G.; Shi, G. Distributed online algorithms for economic power dispatch. In 2025 IEEE 64th Conference on Decision and Control (CDC); IEEE: Piscataway, NJ, USA, 2025; pp. 7220–7225. [Google Scholar]
  11. Wei, M.; Yu, W.; Liu, H.; Chen, D. Byzantine-Resilient Distributed Bandit Online Optimization in Dynamic Environments. IEEE Trans. Ind. Cyber-Phys. Syst. 2024, 2, 154–165. [Google Scholar] [CrossRef] [Scilit]
  12. Yuan, D.; Hong, Y.; Ho, D.W.C.; Xu, S. Distributed Mirror Descent for Online Composite Optimization. IEEE Trans. Autom. Control 2021, 66, 714–729. [Google Scholar] [CrossRef] [Scilit]
  13. Li, J.; Li, C.; Yu, W.; Zhu, X.; Yu, X. Distributed Online Bandit Learning in Dynamic Environments Over Unbalanced Digraphs. IEEE Trans. Netw. Sci. Eng. 2021, 8, 3034–3047. [Google Scholar] [CrossRef] [Scilit]
  14. Pang, Y.; Hu, G. Randomized Gradient-Free Distributed Optimization Methods for a Multiagent System with Unknown Cost Function. IEEE Trans. Autom. Control 2020, 65, 333–340. [Google Scholar] [CrossRef] [Scilit]
  15. Li, J.; Gu, C.; Wu, Z.; Huang, T. Online Learning Algorithm for Distributed Convex Optimization with Time-Varying Coupled Constraints and Bandit Feedback. IEEE Trans. Cybern. 2022, 52, 1009–1020. [Google Scholar] [CrossRef] [PubMed]
  16. Wang, C.; Xu, S.; Yuan, D.; Zhang, B.; Zhang, Z. Push-Sum Distributed Online Optimization with Bandit Feedback. IEEE Trans. Cybern. 2022, 52, 2263–2273. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Flaxman, A.D.; Kalai, A.T.; McMahan, H.B. Online Convex Optimization in the Bandit Setting: Gradient Descent Without a Gradient. arXiv 2004, arXiv:cs/0408007. [Google Scholar]
  18. Yuan, D.; Proutiere, A.; Shi, G. Distributed Online Optimization with Long-Term Constraints. IEEE Trans. Autom. Control 2022, 67, 1089–1104. [Google Scholar] [CrossRef] [Scilit]
  19. Yi, X.; Li, X.; Yang, T.; Xie, L.; Chai, T.; Johansson, K.H. Distributed Bandit Online Convex Optimization with Time-Varying Coupled Inequality Constraints. IEEE Trans. Autom. Control 2021, 66, 4620–4635. [Google Scholar] [CrossRef] [Scilit]
  20. Dwork, C. Differential Privacy. In International Colloquium on Automata, Languages, and Programming; Springer: Berlin/Heidelberg, Germany, 2006; pp. 1–12. [Google Scholar]
  21. Sun, L.; Ding, D.; Dong, H.; Bai, X. Privacy-Preserving Distributed Economic Dispatch for Microgrids Based on State Decomposition with Added Noises. IEEE Trans. Smart Grid 2023, 15, 2424–2433. [Google Scholar] [CrossRef] [Scilit]
  22. Zhao, D.; Liu, D.; Liu, L. Distributed Privacy Preserving Algorithm for Economic Dispatch Over Time-Varying Communication. IEEE Trans. Power Syst. 2024, 39, 643–657. [Google Scholar] [CrossRef] [Scilit]
  23. Lü, Q.; Deng, S.; Li, H.; Huang, T. Privacy Protection Decentralized Economic Dispatch Over Directed Networks with Accurate Convergence. IEEE Trans. Emerg. Top. Comput. Intell. 2023, 7, 1702–1716. [Google Scholar] [CrossRef] [Scilit]
  24. Xiong, Y.; Xu, J.; You, K.; Liu, J.; Wu, L. Privacy-Preserving Distributed Online Optimization Over Unbalanced Digraphs via Subgradient Rescaling. IEEE Trans. Control Netw. Syst. 2020, 7, 1366–1378. [Google Scholar] [CrossRef] [Scilit]
  25. Lian, Z.; Yu, H.; Chen, Y.; Shahidehpour, M.; Zhou, Q. Secure Microgrid Clusters Under Integrated Satellite–Terrestrial Networks: A Differential Privacy-Preserving Control Strategy. IEEE Trans. Ind. Electron. 2026, 73, 2104–2115. [Google Scholar] [CrossRef] [Scilit]
  26. Liu, J.; Ma, E.; Zha, L.; Tian, E. Distributed Energy Management for Microgrids: When Privacy Preservation Meets Multiple Mechanisms. IEEE Trans. Consum. Electron. 2026, 72, 1514–1523. [Google Scholar] [CrossRef] [Scilit]
  27. Yan, F.; Sundaram, S.; Vishwanathan, S.V.N.; Qi, Y. Distributed Autonomous Online Learning: Regrets and Intrinsic Privacy-Preserving Properties. IEEE Trans. Knowl. Data Eng. 2013, 25, 2483–2493. [Google Scholar] [CrossRef] [Scilit]
  28. Li, X.; Yi, X.; Xie, L. Distributed Online Optimization for Multi-Agent Networks with Coupled Inequality Constraints. IEEE Trans. Autom. Control 2020, 66, 3575–3591. [Google Scholar] [CrossRef] [Scilit]
  29. Huang, B.; Zou, Y.; Chen, F.; Meng, Z. Distributed Time-Varying Economic Dispatch via a Prediction-Correction Method. IEEE Trans. Circuits Syst. I Regul. Pap. 2022, 69, 4215–4224. [Google Scholar] [CrossRef] [Scilit]
  30. Lesage-Landry, A.; Callaway, D.S. Dynamic and Distributed Online Convex Optimization for Demand Response of Commercial Buildings. IEEE Control Syst. Lett. 2020, 4, 632–637. [Google Scholar] [CrossRef] [Scilit]
  31. Hosseini, S.; Chapman, A.; Mesbahi, M. Online Distributed Convex Optimization on Dynamic Networks. IEEE Trans. Autom. Control 2016, 61, 3545–3550. [Google Scholar] [CrossRef] [Scilit]
  32. Abdollahi, A.; Amato, G.; Savastio, L.P.; De Tuglie, E.E. A Game-Theoretic Optimization Framework for Secure and Cost-Efficient Dynamic Reconfiguration of Multi-microgrids. In Proceedings of the International Symposium on Intelligent Technology for Power and Energy Systems; Springer: Cham, Switzerland, 2026; pp. 21–35. [Google Scholar]
  33. Chen, G.; Yang, Q. An ADMM-Based Distributed Algorithm for Economic Dispatch in Islanded Microgrids. IEEE Trans. Ind. Inform. 2018, 14, 3892–3903. [Google Scholar] [CrossRef] [Scilit]
  34. Hao, Y.; Ni, Q.; Li, H.; Hou, S. On the Energy and Spectral Efficiency Tradeoff in Massive MIMO Enabled HetNets with Capacity-Constrained Backhaul Links. IEEE Trans. Commun. 2017, 65, 4720–4733. [Google Scholar] [CrossRef] [Scilit]
  35. Xi, C.; Khan, U.A. Distributed Subgradient Projection Algorithm Over Directed Graphs. IEEE Trans. Autom. Control 2017, 62, 3986–3992. [Google Scholar] [CrossRef] [Scilit]
  36. Zhao, Z.; Xia, L.; Jiang, L.; Ge, Q.; Yu, F. Distributed Bandit Online Optimisation for Energy Management in Smart Grids. Int. J. Syst. Sci. 2023, 54, 2957–2974. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Article metric data becomes available approximately 24 hours after publication online.