Next Article in Journal
Global Optimization Research on Parametric Design of Printed Circuit Heat Exchanger Airfoil Plates Based on Integrated Machine Learning and Simulated Annealing
Previous Article in Journal
Comparative Analysis Between Cylindrical-Type Electrostatic Precipitators Used in Household Applications
Previous Article in Special Issue
FTimeDD: A Time–Frequency Collaborative Model for Multi-Energy Load Forecasting
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Decision-Focused Wind Power Forecasting for Economic Dispatch Under Asymmetric Imbalance Costs

1
School of Applied Technology, Nanjing Institute of Technology, Nanjing 211167, China
2
Department of Electrical and Computer Engineering, National University of Singapore, Singapore 117581, Singapore
3
College of Automation Engineering, Nanjing University of Aeronautics and Astronautics, Nanjing 210016, China
*
Authors to whom correspondence should be addressed.
Energies 2026, 19(17), 4053; https://doi.org/10.3390/en19174053
Submission received: 10 July 2026 / Revised: 16 August 2026 / Accepted: 25 August 2026 / Published: 28 August 2026
(This article belongs to the Special Issue Artificial Intelligence for Energy Forecasting)

Abstract

Wind power forecasts are commonly trained to minimize statistical errors, although the thermal schedule based on a forecast ultimately incurs asymmetric recourse costs. To connect forecast training with this dispatch consequence, a dispatch-value-oriented forecasting (DVOF) framework is developed by coupling a Mamba-style forecaster with a solver-based differentiable economic-dispatch layer. The forecast determines the first-stage thermal schedule; after wind realization, load shedding and wind curtailment settle the imbalance, and the resulting recourse-cost gradient is propagated to the forecaster by KKT-based implicit differentiation. On Global Energy Forecasting Competition 2014 (GEFCom2014) data, DVOF reduces the day-ahead realized imbalance penalty from 55.5 to 26.6 cost units per hour relative to an MSE Loss model with the same backbone. The resulting forecast approaches the conservative operating point implied by the asymmetric recourse model. In the simplified single-area, single-period case, this result shows that the dispatch gradient can guide the forecaster toward a cost-relevant adjustment without prescribing or searching over a quantile level.

1. Introduction

With the transition toward low-carbon power systems, wind power has become an increasingly important source of electricity supply. According to the Global Wind Report 2026, global cumulative installed wind power capacity reached 1299 GW in 2025 and is expected to continue increasing in the coming years [1]. Despite its environmental benefits, wind generation remains inherently weather-dependent and nondispatchable. Its variability and uncertainty affect generation scheduling, economic dispatch, reserve scheduling, and balancing operations [2,3]. Therefore, reliable wind power forecasts (WPFs) are essential inputs to the secure and economic operation of wind-rich power systems [4,5].
Over the past decades, a large number of studies have been devoted to improving wind power forecasting models. Early approaches mainly relied on physical models and numerical weather prediction (NWP). In these methods, meteorological variables are converted into wind power estimates using hub-height wind estimation, turbine power curves, and site-specific corrections. Although physical methods are interpretable, their performance can be affected by NWP errors, terrain complexity, and model mismatch. To reduce reliance on detailed physical modeling, statistical methods [6] were introduced to capture temporal dependence directly from historical wind power data. With the increasing availability of supervisory control and data acquisition (SCADA) data, machine learning methods further improved forecasting performance by learning nonlinear relationships between meteorological inputs and power outputs [7]. More recently, deep learning models, such as recurrent neural networks (RNNs) [8,9], convolutional neural networks (CNNs) [10,11], and attention-based models [12,13,14], have been widely used to model temporal and multi-scale wind patterns. These methods have substantially improved forecast accuracy and provided system operators with more reliable inputs for generation scheduling and economic dispatch.
Despite these advances, wind power forecasting remains challenging because wind power exhibits rapid local fluctuations as well as long-range dependencies driven by meteorological evolution. Extending the input context can provide richer temporal information but also increases the computational burden, particularly for self-attention-based models whose complexity grows quadratically with sequence length. Selective state-space models, exemplified by Mamba, provide an efficient alternative by introducing input-dependent state transitions and processing a length-T sequence in O ( T ) time [15,16]. Existing Mamba-based WPF studies have focused primarily on improving statistical forecast accuracy, whereas their integration with downstream economic dispatch for decision-focused or value-oriented training remains insufficiently explored.
In power system operation, a WPF is not merely a standalone statistical output. It serves as an input to generation scheduling before wind power is realized. In a two-stage forecast-based economic dispatch process, the WPF affects the first-stage generation schedule, whereas the deviation revealed upon wind power realization gives rise to the subsequent balancing requirement [17]. Therefore, two forecasts with the same mean absolute error (MAE) or mean squared error (MSE) can induce different realized recourse costs. For example, suppose that the realized wind power is 0.50 p.u. Forecasts of 0.45 and 0.55 p.u. have the same absolute error and are therefore equivalent under MAE. However, from an operational perspective, they may lead to different consequences. A forecast of 0.55 p.u., which overestimates the realized wind power, may result in insufficient scheduled thermal generation, thereby increasing the need for upward balancing and, in extreme cases, load shedding. In contrast, a forecast of 0.45 p.u., which underestimates the realized wind power, increases scheduled thermal generation and may therefore reduce shortage risk at the expense of greater wind curtailment when the realized wind power is high. When overforecasting and underforecasting have asymmetric operational consequences, symmetric statistical losses cannot distinguish between equal-magnitude errors of opposite signs [18]. Therefore, a forecast should be evaluated not only by its statistical accuracy but also by the realized second-stage recourse cost induced by its forecast-based generation schedule. However, most existing wind power forecasting approaches, including advanced deep learning models, are still trained using conventional statistical losses and evaluated primarily using forecast-accuracy metrics [19]. They do not explicitly account for the impact of forecast errors on first-stage dispatch decisions and second-stage recourse costs [20,21]. Therefore, it is necessary to study whether higher forecast accuracy necessarily leads to lower realized recourse cost.
The literature bearing on this accuracy–value mismatch falls into four streams, which the following two paragraphs address in turn. These are cost-oriented and probabilistic renewable forecasting, decision-focused learning and differentiable optimization, forecasting–optimization integration for trading and commitment, and forecasting under imbalance settlement in power-system operation.
The first stream designs forecast products according to downstream operational consequences. For example, Ref. [20] proposed cost-oriented prediction intervals and showed that statistically calibrated intervals do not necessarily minimize downstream operating cost. Moreover, Ref. [21] considered contextual selection of value-oriented prediction intervals, allowing the preferred interval to change with the operating condition. Both build on the asymmetric quantile-regression loss introduced by [22], whose minimizer is the conditional quantile of the target. These approaches exploit the asymmetry between power shortfalls and surpluses by selecting decision-relevant forecast products without explicitly differentiating through the downstream optimization problem. The second stream uses end-to-end decision-focused learning together with differentiable optimization. Ref. [23] proposed the Smart "Predict, then Optimize" framework, in which predictions are evaluated according to the decision loss they induce rather than solely by their deviation from the realized outcome. Earlier, Ref. [24] trained a probabilistic model directly against the cost of the stochastic-optimization task it feeds, establishing that a task-based objective can outperform a likelihood-based one on the decision that matters. However, applying this principle to a constrained economic dispatch problem requires differentiation of the downstream operating objective with respect to the forecast. Ref. [25] addressed this requirement by introducing differentiable convex optimization layers, which enable differentiation through the solution map of a parameterized convex program. This development makes it possible to train a forecasting model end-to-end by differentiating through the downstream economic dispatch problem rather than relying on a manually constructed statistical surrogate.
The third stream integrates forecasting with trading and commitment decisions. In renewable-energy trading, economic value has been used to identify decision-relevant forecasting features, while prescriptive models have been developed to integrate forecasting and trading decisions [26,27]. For network-constrained unit commitment, a closed-loop predict-and-optimize framework was developed in which forecasts jointly affect commitment and dispatch decisions [28]. In addition, Ref. [29] considered decision-focused electricity-price forecasting for commitment decisions. The fourth stream links renewable-energy forecasts directly to the operational cost they incur once the imbalance is settled. Ref. [19] formulated cost-oriented WPF as a decision-focused learning problem. An iterative value-oriented renewable-energy forecasting framework was proposed in [17], in which forecast updates are evaluated through their operational consequences. In addition, Ref. [18] coupled WPFs with system operating costs through an economic dispatch model and Karush–Kuhn–Tucker (KKT)-based sensitivity analysis, and Ref. [30] extended decision-focused learning to settings in which the decision itself alters the uncertainty it faces. Collectively, these studies establish that operational benefits can arise either from forecast products that reflect asymmetric costs or from decision-derived learning signals propagated through a downstream optimization model.
Despite these advances, the relationship between forecast accuracy and recourse-based dispatch value remains insufficiently understood. First, the recourse consequence of a WPF depends on how it changes the first-stage generation schedule and the resulting second-stage recourse actions. It therefore cannot be adequately assessed using forecast-accuracy metrics alone or by evaluating the first-stage dispatch under the assumption that the forecast equals the realization. Second, different training objectives may produce systematically different point forecasts. Symmetric forecasting losses prioritize statistical accuracy, fixed-quantile pinball losses encourage an asymmetric forecast shift, and decision-focused learning incorporates the downstream recourse consequence into training. The relative recourse-cost benefits of these formulations have not been sufficiently compared. Third, it remains unclear whether decision-focused learning provides benefits beyond the conservative forecast shift induced by asymmetric recourse penalties and how any additional benefit depends on forecast uncertainty and the sensitivity of the dispatch solution to forecast perturbations. These questions remain insufficiently addressed and motivate the present study.
Motivated by these research gaps, this paper proposes a dispatch-value-oriented forecasting (DVOF) framework for wind-integrated economic dispatch. The framework couples a Mamba-style wind forecaster with a solver-based differentiable economic dispatch layer. The forecaster is trained using a blended objective that combines an accuracy-oriented forecasting loss with the normalized second-stage recourse cost induced by the forecast-based generation schedule. The main contributions of this paper are summarized as follows:
1.
A Mamba-style forecaster is developed as the forecasting backbone of the proposed DVOF framework. The model combines a selective state-space architecture with a forecast-origin information set comprising historical wind power, horizon-aligned NWP forecasts, and calendar encodings. This architecture has linear sequence-length complexity and retains the meteorological information required for WPF.
2.
A DVOF architecture is proposed by coupling the Mamba-style forecaster with a solver-based differentiable economic dispatch layer. The dispatch layer solves a convex first-stage economic dispatch problem parameterized by the WPF and propagates the downstream recourse-cost gradient to the forecaster through KKT-based implicit differentiation. The training objective combines an accuracy-oriented forecasting loss with the normalized second-stage recourse cost arising from load shedding and wind curtailment. This formulation trains the forecaster to reduce realized recourse cost rather than minimizing pointwise forecast error alone.
3.
A verifiable test of the learning mechanism is introduced. Because the recourse model admits a closed-form optimal quantile, the quantile level a trained forecast attains can be measured directly and compared with it. This turns the question of whether the differentiable layer transmits the intended economic signal into a measurement rather than an assertion, and it applies equally to point and quantile forecasters.
4.
Case studies at the 1-h-ahead and 24-h-ahead forecast horizons evaluate the relationship between forecast accuracy and dispatch value on the Global Energy Forecasting Competition 2014 (GEFCom2014) dataset. Two quantile controls are used, a fixed reference quantile and one selected on validation data, so that the incremental contribution of the differentiable layer is judged against a properly tuned cost-aware forecaster. The results separate what the layer contributes, namely, automatic and condition-stable placement of the operating point, from what it does not, namely, a lower realized cost in a single-period problem.
The remainder of this paper is organized as follows. Section 2 formulates the problem. Section 3 presents the proposed method. Section 4 reports and discusses the experimental results. Section 5 concludes the paper.

2. Problem Formulation

A WPF is commonly evaluated according to its deviation from the realized wind power. However, in power-system operation, the forecast also affects the thermal generation schedule determined before wind power is realized and, consequently, the recourse actions required afterward. Therefore, this paper studies how to generate WPFs that reduce the asymmetric second-stage recourse cost induced by forecast-based economic dispatch through a learning objective that retains an accuracy-oriented term.

2.1. Forecast-Driven Dispatch Process

Let c t denote the causal information available when the forecast for delivery hour t is issued, including historical wind power, NWP features, and calendar encodings. Let w t and w ^ t denote the realized wind power and its deterministic forecast, respectively. A forecaster f θ ( · ) maps the causal context c t to the WPF w ^ t as
w ^ t = f θ ( c t ) ,
where θ denotes the trainable model parameters. In an accuracy-oriented forecasting framework, θ is optimized by minimizing a statistical loss between w ^ t and w t , such as MSE or MAE.
Operationally, w ^ t serves as an input to economic dispatch and determines the thermal generation schedule before w t is known. After wind power is realized, any residual mismatch among scheduled thermal generation, wind power, and load is settled through load shedding or wind curtailment. Thus, the dispatch value considered in this paper is quantified by the asymmetric recourse cost induced by the forecast-based generation schedule rather than by forecast error alone.

2.2. Forecast-Based First-Stage Economic Dispatch

Consider one dispatch interval with G conventional generators, deterministic load demand d, and a generic wind power input w ˜ . The active-power output of generator k is p k , with quadratic fuel cost
C k ( p k ) = a k p k 2 + b k p k ,
where a k > 0 and b k are the quadratic and linear cost coefficients, respectively, and p k min p k p k max are the generation limits.
Let p = ( p 1 , , p G ) and x = ( p , s curt , s shed ) collect the first-stage decision variables. Given w ˜ , the forecast-based economic dispatch layer solves the following convex quadratic program:
x ( w ˜ ) = arg min p , s curt , s shed k = 1 G C k ( p k ) + M plan ( s curt + s shed ) ,
s . t . k = 1 G p k + w ˜ + s shed = d + s curt ,
p k min p k p k max , 0 s curt w ˜ , 0 s shed d .
where s curt and s shed are first-stage feasibility slacks representing forecast-based surplus and shortage, respectively. These variables preserve feasibility under forecast-driven surplus or shortage conditions. The common coefficient M plan penalizes their use symmetrically. In the present implementation, M plan is numerically set equal to M shed . However, it is used only in the first-stage dispatch problem and should not be confused with the asymmetric second-stage settlement coefficients M shed and M curt . Because both first-stage slacks have the same positive penalty, an optimal solution cannot contain positive values of both simultaneously.
The total scheduled thermal generation is
p ( w ˜ ) = k = 1 G p k ( w ˜ ) .
In operation, the layer is solved with w ˜ = w ^ t . Consequently, equal-magnitude forecast errors can induce different generation schedules and asymmetric realized recourse costs.

2.3. Schedule-Then-Recourse Settlement

For a WPF w ^ , the first-stage dispatch problem determines p ( w ^ ) before the realization w is observed. After realization, the net supply is
n ( w ^ , w ) = p ( w ^ ) + w .
The resulting load shedding and wind curtailment are
shed = ( d n ) + , curt = ( n d ) + ,
where ( z ) + = max ( z , 0 ) .
The realized second-stage recourse cost is
C rec ( w ^ , w ) = M shed shed + M curt curt ,
where M shed > M curt represents the greater consequence of an energy shortfall than that of wind curtailment. The relative values are treated as controlled recourse scenarios rather than universal market-settlement prices.
For cost-accounting purposes within the stylized dispatch model, the realized cost associated with a forecast can be decomposed as
C real ( w ^ , w ) = k = 1 G C k p k ( w ^ ) scheduled fuel cost + C rec ( w ^ , w ) second - stage recourse cost .
The proposed DVOF framework incorporates C rec into its learning objective and uses it as the primary value metric. The complete realized cost C real is retained only for cost-accounting purposes and is not included in the DVOF objective or used as the primary evaluation metric. Accordingly, all reported value improvements refer to reductions in realized recourse cost and do not imply reductions in C real .
When generator limits are nonbinding and the first-stage feasibility slacks are inactive,
p ( w ^ ) = d w ^ .
The recourse quantities then reduce to
shed = ( w ^ w ) + , curt = ( w w ^ ) + ,
and the recourse cost becomes
C rec ( w ^ , w ) = M shed ( w ^ w ) + + M curt ( w w ^ ) + .
This expression is a scaled asymmetric pinball loss. Its expected value is minimized by the conditional τ -quantile of w given c t , where
τ = M curt M shed + M curt .
Thus, the recourse-only objective favors a conservative WPF when shortage is penalized more heavily than curtailment.

2.4. Accuracy-Oriented and DVOF Objectives

The conventional accuracy-oriented objective is
min θ E f θ ( c t ) w t 2 .
This objective targets the conditional mean but does not account for the recourse consequences of the forecast-induced generation schedule.
The corresponding recourse-cost-oriented objective is
min θ E C rec f θ ( c t ) , w t .
To retain a statistical-accuracy signal, DVOF adopts the blended objective
min θ E f θ ( c t ) w t 2 + β C rec f θ ( c t ) , w t M shed ,
where division by M shed controls the numerical scale of the recourse term, and β determines its relative contribution to training.
Unlike the MSE objective, the DVOF objective also accounts for how a WPF changes the first-stage generation schedule and the resulting asymmetric second-stage recourse cost. It does not minimize the complete realized cost in Equation (10).

3. Proposed Dispatch-Value-Oriented Forecasting Architecture

Figure 1 illustrates the proposed DVOF framework. At each forecast origin, the forecaster f θ maps the causal information set c t to a deterministic WPF w ^ t . The differentiable dispatch layer then solves the first-stage economic dispatch problem in Equations (3)–(5) using w ^ t , thereby determining the scheduled thermal generation p ( w ^ t ) . Once the actual wind power w t is observed, the residual imbalance is settled through load shedding or wind curtailment, yielding the realized recourse cost C rec ( w ^ t , w t ) . Thus, w t is used to evaluate the forecast-induced schedule but is unavailable to the first-stage dispatch decision.
During training, the recourse-cost gradient is propagated from the settlement stage through the dispatch layer to the forecaster. The required sensitivity of the dispatch solution to the WPF is obtained primarily through KKT-based implicit differentiation, with a finite-difference fallback at numerically degenerate active-set configurations.

3.1. Mamba Forecaster

The wind-power forecaster f θ maps a causal input X R T × d in , which is the matrix form of the context c t , to a wind forecast w ^ . The input has d in = 11 features, and T denotes the number of tokens available at the forecast time. NWP variables are aligned with the target horizon. Each token contains historical wind power, six NWP variables (10 m and 100 m wind components u , v and the derived wind speeds), and four calendar encodings (sin/cos of hour-of-day and day-of-year).
The mapping is implemented as a Mamba-style state-space sequence model (SSM) [15]. To make the architecture explicit, each module is denoted by a separate mapping function. The complete computation is
H ( 0 ) = F Embed ( X ) = F L ( X ) + P pos ,
H ( ) = F M ( ) H ( 1 ) , = 1 , , L ,
H ¯ = F LN H ( L ) ,
h ¯ T = H ¯ T ,
w ^ = F out F MLP ( h ¯ T ) ,
where F L maps each d in -dimensional observation to a d m -dimensional token by a time-shared linear map applied to every row of X, P pos R T × d m denotes the learnable positional embedding, F M ( ) denotes the th Mamba-style selective-SSM block, and F LN denotes layer normalization. The vector h ¯ T is the representation of the last causal token. The forecasting head F MLP comprises a linear layer, a GELU activation, and dropout, while F out maps the resulting representation to the scalar wind-power forecast.

Selective-SSM Block

Let H in R T × d m denote the sequence entering one block. The block is written as a composition of six functions, namely, normalization, linear projection, local convolution, selective scan, gate modulation, and residual output.
A residual copy R = H in is retained, and the input is normalized to V = F LN ( H in ) , which stabilizes the relative scale of the wind-power, NWP and calendar features before they enter the recurrent state update. The normalized sequence is then expanded and split into a candidate dynamics branch U and a gate branch Z,
[ U , Z ] = F in ( V ) = V W in + 1 b in , U , Z R T × d i ,
where W in R 2 d i × d m and b in R 2 d i are the weight and bias of the in-block projection, 1 R T is a vector of ones, and d i is the expanded inner width. The candidate branch carries the temporal content to be propagated, whereas the gate branch controls how much of the resulting dynamic feature reaches the next layer.
The candidate branch is first processed by a causal depthwise convolution,
U ˜ = F SiLU F Conv . ( U ) ,
where F Conv . applies a per-channel one-dimensional convolution of kernel width k with k 1 steps of left padding, so that each output position depends only on the current and preceding tokens, and F SiLU is the sigmoid linear unit activation function.
The selective part of the block generates the SSM coefficients from the current token rather than using fixed coefficients for all hours, which is
Δ t = F Δ ( u ˜ t ) = softplus ( W Δ u ˜ t + ρ Δ ) + ε Δ , η t = F B ( u ˜ t ) , ξ t = F C ( u ˜ t ) ,
where Δ t R d i , η t , ξ t R N are the input-dependent state input and output projections, N is the latent state dimension, ρ Δ is the step bias, and ε Δ = 10 4 is a small positive floor that keeps the step strictly positive. The symbols η t and ξ t replace the B t and C t of the standard selective-SSM notation, so that they are not confused with the generator cost coefficients b k and C k in Equation (2) or the context set c t in Equation (1). With A = exp ( A log ) R d i × N and learned skip coefficient D R d i , the selective-SSM scan Y = F Selective SSM ( U ˜ ) is computed channel-wise as
s t ( j ) = exp Δ t ( j ) A ( j ) s t 1 ( j ) + Δ t ( j ) u ˜ t ( j ) η t , o t ( j ) = ξ t , s t ( j ) + D ( j ) u ˜ t ( j ) ,
initialized by s 0 ( j ) = 0 , where s t ( j ) R N is the latent state of channel j. The input-dependent step Δ t ( j ) sets the memory length of that channel, so a block can hold a persistent wind state or discard it in favour of the current observation according to the token it is reading.
Finally, the scan output is modulated by the gate branch and projected back to the model width
Γ = F SiLU ( Z ) , Y g = Y Γ , H out = R + F Dropout F BlockOut ( Y g ) ,
where Γ is the gating signal, Y is the output of the selective scan, ⊗ is the element-wise product, F BlockOut is the output projection of the block, and R restores the residual retained at its input.
The complete architecture is summarized in Figure 2. Equations (18)–(22) define the overall forward path.

3.2. Differentiable Economic Dispatch Layer

To train f θ through dispatch, Equations (3)–(5) are exposed as a layer x ( w ^ ) . The forward pass solves the optimization problem, and the backward pass returns the sensitivity of the optimal solution to the wind forecast.

3.2.1. Forward Process

Collect the first-stage decision variables as x = ( p 1 , , p G , s curt , s shed ) R G + 2 . For numerical differentiation, the implementation augments the first-stage objective with a small quadratic regularization term ε s [ ( s curt ) 2 + ( s shed ) 2 ] . The resulting regularized QP is
x ( w ^ ) = arg min x 1 2 x P x + q x s . t . A bal x = b ( w ^ ) , x min x x max ( w ^ ) ,
where
P = diag 2 a 1 , , 2 a G , 2 ε s , 2 ε s , q = b 1 , , b G , M plan , M plan , A bal = ( 1 , , 1 , 1 , + 1 ) , b ( w ^ ) = d w ^ .
The box constraints collect the generator limits, 0 s curt w ^ , and 0 s shed d . Because a k > 0 and ε s > 0 , the matrix P is positive-definite, and the regularized QP is strictly convex. The coefficient ε s = 10 6 is used to stabilize differentiation while introducing only a small perturbation to the first-stage objective.
The forward QP is solved through CVXPY using the first available solver in the order ECOS, CLARABEL, OSQP, and SCS. In the computing environment used for the reported experiments, ECOS was unavailable, and CLARABEL was therefore used for all forward dispatch solves. The solver is used only to obtain the optimal first-stage solution, whereas gradients with respect to the WPF are computed separately through the KKT-based implicit-differentiation procedure described below. The scheduled thermal generation of Equation (6) is read from the first G coordinates of the stacked solution:
p ( w ^ ) = k = 1 G x k ( w ^ ) .

3.2.2. Backward Process by Implicit Differentiation

Let A denote the locally active set of box constraints at x ( w ^ ) . Writing the box constraints in the canonical form E x h ( w ^ ) , where E collects the box-constraint rows and is distinct from the generator count G, the active constraints satisfy
E A x = h A ( w ^ ) .
Let ν R and λ A 0 denote the multipliers associated with the balance equation and active box constraints, respectively. The active-set KKT conditions are
P x + q + A bal ν + E A λ A = 0 , A bal x = b ( w ^ ) , E A x = h A ( w ^ ) .
Under a locally unchanged active set and a nonsingular KKT system, differentiation with respect to w ^ gives
P A bal E A A bal 0 0 E A 0 0 K x w ^ ν w ^ λ A w ^ = 0 b w ^ h A w ^ .
Here, b w ^ = 1 .
The vector h A / w ^ is zero except when the forecast-dependent upper bound s curt w ^ is active, in which case its corresponding entry equals one.
Because the training loss is scalar, the full solution Jacobian need not be formed. Let
g x = L x
denote the upstream gradient with respect to the dispatch solution. The system is assembled over the variables that are strictly inside their box constraints at x , because a variable held at a forecast-independent bound has zero local sensitivity. A diagonal term of 10 8 is added to K so that the factorization remains well posed. The adjoint vector z = ( z x , z ν , z λ ) is obtained from
K z = g x 0 0 .
Then, the gradient with respect to the WPF is
L w ^ = z ν b w ^ + z λ h A w ^ = z ν + z λ h A w ^ .
The second term of Equation (36) is written for the general case. Because the balance equation gives s curt = k p k + w ^ + s shed d , the bound s curt w ^ is attained only when the demand equals the minimum total thermal output. Whenever the demand is larger, the gradient reduces to z ν , which is the case at every dispatch instance solved here.

3.2.3. Degeneracy Handling

The implicit-differentiation procedure may become unreliable when the active set changes or when the KKT matrix is singular or ill-conditioned. The layer therefore invokes the finite-difference fallback when the active-set construction is inconsistent, the condition number of K exceeds 10 12 , the factorization of K fails, or the resulting gradient is nonfinite or has a magnitude greater than 10 6 .
If any safeguard is triggered, the layer uses the central finite-difference approximation
x w ^ x ( w ^ + ϵ ) x ( w ^ ϵ ) 2 ϵ , ϵ = 10 3 p . u .
using two additional forward solves. The estimated solution Jacobian is then contracted with the upstream gradient L / x .
Estimating the solution Jacobian rather than differentiating a particular downstream objective keeps the fallback applicable to the realized recourse cost in Equation (9). At active-set transitions, the finite-difference fallback provides a numerical directional approximation for stochastic optimization. The frequency of the fallback was measured over the training split for three of the matched seeds at both horizons. The exact KKT branch was used in 100.00 % of the 11,603 backward calls per seed at H = 24 and in at least 99.98 % of the 11,626 calls per seed at H = 1. The dispatch quadratic program is therefore well conditioned over the operating range visited in this study, and the reported gradients come from exact implicit differentiation.

3.3. Dispatch-Value-Oriented Training Objective

The network produces a raw normalized output w ˜ s , where the subscript s marks a quantity expressed on the wind-power base S w , which equals the wind rated power and is 1.0  p.u. in the case study. Because the normalized wind power cannot exceed unity, the dispatch input is additionally capped at one. The non-negative prediction used in the accuracy term and the bounded dispatch input are defined, respectively, as
w ^ s acc = ReLU ( w ˜ s ) , w ^ s disp = min { 1 , w ^ s acc } , w ^ = S w w ^ s disp , w s = w S w .
The DVOF loss is
L ( θ ) = w ^ s acc w s 2 accuracy term + β C rec ( w ^ , w ) M shed normalized recourse term ,
where C rec is defined in Equation (9). Division by M shed controls the numerical scale of the recourse term, whereas β determines its relative contribution to training.
Applying the chain rule through the differentiable dispatch layer gives
L θ = 2 w ^ s acc w s w ^ s acc θ accuracy term + β M shed C rec p settlement sensitivity p w ^ dispatch sensitivity w ^ θ .
Away from the nondifferentiable point n = d , the settlement sensitivity is
C rec p = M shed 1 { n < d } + M curt 1 { n > d } , n = p ( w ^ ) + w .
In the interior regime, scheduled thermal generation changes approximately one-for-one in the opposite direction to the WPF:
p w ^ 1 .
Consequently, shortage hours generate a stronger conservative learning signal than curtailment hours when M shed > M curt . For the recourse-only component, the equilibrium of these asymmetric signals corresponds to the conditional quantile τ in Equation (14). Under the blended DVOF objective, the MSE term moderates this conservative shift.

3.4. Training Procedure

Algorithm 1 states the two-stage procedure: MSE pretraining, then DVOF fine-tuning from the checkpoint with the lowest validation MAE, with the recourse-cost gradient propagated through the dispatch layer. Two properties of the protocol are worth making explicit. The realized wind power w t never enters the forecaster input or the first-stage dispatch decision. It is the supervised target during training and validation and the reference realization at test time. Checkpoint selection also changes with the objective, from validation MAE in stage 1 to validation recourse cost in stage 2, so each arm is selected by the criterion it is meant to optimize.
Because each recourse evaluation requires a dispatch solve, the recourse term is applied to a uniformly spaced subset of mini-batches while the MSE term is applied to every one. This sparse evaluation approximates the blended objective of Equation (39), so β is an empirically selected coefficient conditional on the adopted update frequency rather than a transferable constant.
Algorithm 1 Training of the proposed DVOF model
Require: Training set D tr , validation set D val , forecaster f θ , differentiable dispatch layer G ,
  wind-power scale S w , and trade-off coefficient β .
Ensure: Best MSE model θ MSE and best DVOF model θ DVOF .
1:
Stage 1: MSE pretraining
2:
Initialize θ
3:
for each pretraining epoch do
4:
   for each mini-batch B D tr  do
5:
     Compute w ˜ s , t = f θ ( c t ) , w ^ s , t acc = ReLU ( w ˜ s , t ) , w ^ s , t disp = min { 1 , w ^ s , t acc } , and w ^ t = S w w ^ s , t disp
6:
     Compute w s , t = w t / S w
7:
     Compute L MSE = 1 | B | t B ( w ^ s , t acc w s , t ) 2
8:
     Update θ using θ L MSE
9:
   end for
10:
 Evaluate validation MAE and save the checkpoint if it improves
11:
end for
12:
Restore the checkpoint with the lowest validation MAE
13:
Set θ MSE θ
14:
Stage 2: DVOF fine-tuning
15:
Initialize θ θ MSE
16:
Select a uniformly spaced subset of mini-batches B dec for recourse-loss evaluation
17:
for each fine-tuning epoch do
18:
   for each mini-batch B D tr  do
19:
     Compute w ˜ s , t = f θ ( c t ) , w ^ s , t acc = ReLU ( w ˜ s , t ) , w ^ s , t disp = min { 1 , w ^ s , t acc } , and w ^ t = S w w ^ s , t disp
20:
     Compute w s , t = w t / S w
21:
     Compute L MSE = 1 | B | t B ( w ^ s , t acc w s , t ) 2
22:
     if  B B dec  then
23:
        Obtain p t = G ( w ^ t , d t ) and compute C t rec ( w ^ t , w t )
24:
        Set L = L MSE + β | B | t B C t rec ( w ^ t , w t ) / M shed
25:
     else
26:
        Set L = L MSE
27:
     end if
28:
     Update θ using θ L
29:
   end for
30:
   Evaluate validation recourse cost and save the checkpoint if it improves
31:
end for
32:
Restore the checkpoint with the lowest validation recourse cost
33:
Set θ DVOF θ
34:
return  θ MSE and θ DVOF

4. Experimental Results and Discussion

4.1. Experimental Setup

4.1.1. Data and Evaluation Metrics

The forecasting data are drawn from Zone 3 of the GEFCom2014 wind track, which contains hourly normalized wind-power observations and numerical weather prediction (NWP) variables [31]. At each forecast origin, the causal input sequence combines observed historical wind power, horizon-aligned NWP information, and sine–cosine calendar encodings for hour of day and day of year. The data are divided chronologically, without shuffling, into training, validation, and test segments in the ratio 0.70 / 0.15 / 0.15 , corresponding to 11,698 , 2506, and 2508 h, respectively. The evaluation considers one-hour-ahead (H = 1) and day-ahead (H = 24) forecasts.
The demand shape is obtained from Zone 1 of the GEFCom2014 load track. Because the public wind and load tracks span different calendar periods, the load series is aligned to the wind calendar from a common January origin and paired hour by hour. The resulting demand profile defines a standardized dispatch scenario rather than an observed joint wind–load trajectory. Consequently, the case study isolates the forecast–dispatch mechanism under a controlled demand shape, and it does not reproduce the temporal co-movement of a particular power system. Validation using contemporaneous wind and load measurements is required before operational generalization. Let e t = w ^ t w t denote the signed forecast error over the N test samples. Forecasting performance is evaluated using the MAE, RMSE and R 2 in their standard definitions, where lower MAE and RMSE and higher R 2 indicate better statistical accuracy. The signed error is reported as a distribution rather than as a single bias value, because its location relative to zero is the quantity that carries operational meaning. A forecast displaced below the realization is conservative, and shortage and surplus incur asymmetric recourse penalties.
Dispatch value is evaluated by settling every forecast through the schedule-then-recourse procedure in Section 2. The primary value metric is the realized imbalance penalty, M shed shed + M curt curt , because it is the recourse cost induced directly by the forecast-based generation schedule.
For each seed, all eligible timestamps in the held-out chronological test segment are settled before the seed-level metrics are calculated. The comparison of loss-defined models uses the six matched seeds { 0 , 11 , 22 , 33 , 44 , 55 } , so that every model sees an identical set of initializations and paired comparisons are available. The penalty-ratio and β studies of Section 4.3.5 require a separate training run per scenario point and use the three matched seeds { 0 , 11 , 22 } , as their captions state. The resulting dispersion characterizes sensitivity to initialization within the present data and scenario and does not establish robustness to distribution shift or alternative market designs.

4.1.2. Compared Models and Experimental Configuration

The benchmarks are organized in two stages, which separates the choice of forecasting backbone from the choice of training objective. First, the Mamba-style forecaster is compared with persistence, LSTM, GRU, Transformer and DLinear to establish the statistical competence of the selected backbone, and all learned models share the causal input of Section 4.1.1, the Adam optimizer, gradient clipping and early stopping on validation MAE. Second, the architecture is held fixed and only the objective is varied across the four loss-defined models below, which together separate the contribution of statistical accuracy, of a cost-aware conservative shift, and of the differentiable dispatch layer itself.
MSE Loss model. This accuracy-oriented model is trained using MSE and selected using validation MAE. As an estimate of the conditional mean, it provides the principal statistical baseline for evaluating whether DVOF training can reduce realized imbalance penalty.
Pinball Loss model, fixed τ = 2/9. This quantile-regression model uses the same backbone and is trained with the pinball loss. Its dispatch quantile is fixed at τ = 2/9, the value adopted in the related value-oriented forecasting formulation [17]. It is retained for comparability with that work and is not presented as the optimal quantile of the present scenario.
Pinball Loss model, validation selected. This control uses the same quantile-regression backbone but chooses its dispatch quantile per seed from the candidate set { 0.05 , 0.0909 , 0.10 , 1 / 6 , 2 / 9 , 0.30 , 0.50 } by minimizing the realized imbalance penalty on the validation split. The test split is never used in the selection. Because it locates the cost-optimal quantile by direct search under the operative recourse scenario, it is the strongest control available to a forecaster that does not differentiate through the dispatch problem, and it is the reference against which the incremental contribution of the dispatch layer is judged.
Proposed DVOF Loss model. This model starts from the MSE-trained forecaster and is fine-tuned using Equation (39), which combines MSE with a differentiable recourse-cost term. It receives no quantile level and performs no search over one.
Table 1 reports the backbone and its selection protocol. A 50-trial Optuna study, of which 47 completed, was run on the one-hour-ahead MSE-trained forecaster and ranked by validation MAE under the chronological training–validation split. The searched variables and their candidate ranges are listed beside every selected value. The resulting architecture was then fixed for all four loss-defined models at both horizons, with DVOF fine-tuning using its objective-specific optimization schedule. Because the selection criterion is validation MAE, dispatch-cost information does not enter the choice of backbone and no arm of the comparison is tuned preferentially.
The dispatch environment is deliberately simplified so that the effect of asymmetric recourse on the forecasting gradient can be identified. It retains power balance, continuous generation-output bounds, and non-negative load-shedding and wind-curtailment recourse, while omitting network power flows, intertemporal ramping, reserve procurement, storage dynamics, multi-area exchanges, and binary commitment decisions. Holding these coupled effects outside the model separates the forecast adjustment induced by the recourse asymmetry from adjustments attributable to a particular network or intertemporal constraint. Accordingly, the supplementary scenario analysis changes the recourse-cost ratio only and holds the remaining operating conditions fixed. The model is therefore used to study the forecast–decision mechanism under controlled conditions, rather than to reproduce the market design of a specific control area. The configuration of the economic-dispatch test system is shown in Table 2.

4.2. Forecasting Competence of the Mamba Backbone

The forecasting backbone is first evaluated independently of the dispatch objective. A value-oriented training objective is informative only when it is applied to a statistically competitive forecaster. Figure 3 and Table 3 therefore compare the Mamba-style backbone with the learned baselines and the persistence reference.
All learned models are trained under the same causal-input and MSE-training protocol and are evaluated across six random initializations. Table 3 reports the resulting accuracy metrics, and Figure 3 provides an illustrative trajectory from the held-out test segment.
The purpose of this comparison is to establish that the selected backbone is a competitive predictor, so that the subsequent comparison of training objectives is not confounded by a weak forecaster. It is not to claim architectural superiority. Table 3 shows that the Mamba-style forecaster attains the lowest mean MAE and RMSE and the highest mean R 2 at both horizons, but the margin over the strongest learned baselines is within the seed-to-seed dispersion. At H = 24, its MAE is 2.0 % below LSTM, a difference of about one pooled standard error over six seeds, and at H = 1, it is 0.6 % below GRU, well inside the dispersion. Neither gap is statistically separated, and the appropriate conclusion is that the Mamba-style backbone performs on par with the strongest recurrent baselines rather than better than them.
The differences that are unambiguous lie elsewhere. All learned models improve substantially on persistence at the day-ahead horizon, where the MAE and RMSE fall by 67.2 % and 64.8 % , and the R 2 rises from 0.813 to 0.776 ; and all outperform DLinear, whose MAE is 14.0 % higher at H = 24 and whose seedwise dispersion at H = 1 is an order of magnitude larger than that of the other learned models. The signed bias of every accuracy-trained model is small. The backbone is therefore adequate for the purpose at hand, and the results of Section 4.3 can be attributed to the training objective rather than to the choice of architecture.
The representative trajectories in Figure 3 are consistent with the aggregate results. The Mamba-style forecast follows the principal ramps in the displayed interval, whereas persistence exhibits the expected horizon-dependent lag and misses several ramp onsets. The figure is illustrative, and the quantitative comparison and its variability across seeds are reported in Table 3.

4.2.1. Computational Cost

Computational efficiency is evaluated under an identical input and hardware configuration for all backbones. Timings are averaged over repeated runs on an NVIDIA GeForce RTX 5070 Ti GPU using PyTorch 2.13 and CUDA 13.2. Latency is reported per sample, calculated as batch time divided by batch size, at T = 72. Table 4 reports the number of trainable parameters and the inference and training latencies. The Mamba-style forecaster contains approximately 330 k trainable parameters, about 2.1 times the Transformer size and 3.8 times the GRU size. Its inference and training latencies are 88.4 and 424.5 μ s /sample, respectively. Thus, at the context length used in the main experiments, the proposed backbone does not offer a latency advantage over the alternative models.
One qualification is essential to interpreting these timings. The selective scan is implemented in plain PyTorch as an explicit recurrence over the T time steps, so that the model runs without the fused CUDA kernels of the original Mamba implementation. The measured latency therefore reflects a reference implementation with a large per-step constant rather than an intrinsic property of the selective-SSM architecture, and it should not be read as a like-for-like comparison against the highly optimized fused kernels behind the recurrent and attention baselines. What the measurements do establish is the scaling behavior of the selective scan, which is linear in T as the architecture predicts, and the absolute cost of the configuration actually used in this study.
The choice of backbone therefore reflects a trade-off among statistical accuracy, realized recourse value and computational demand rather than a preference for model complexity. In the present use case, the trade-off is not binding, because the absolute inference time stays below 0.1 ms/sample at T = 72, which is negligible against the one-hour and day-ahead decision horizons, and training and DVOF fine-tuning are offline and do not enter the online execution time after deployment.
Selective scanning has linear sequence-length complexity, whereas standard self-attention has quadratic complexity, but no latency crossover is observed between T = 24 and T = 768, because the per-step constant of the reference implementation dominates at these context lengths. The measured growth rates are nonetheless the ones the two complexity classes predict once the sequence is long enough that fixed overheads no longer dominate. Over T = 192 to 768, the proposed model slows by a factor of 4.0 for a fourfold increase in T, whereas the Transformer slows by a factor of 7.3 . Below T = 192, the comparison is governed by constant terms rather than by asymptotic order. The proposed backbone is therefore selected for its forecasting and dispatch behavior rather than for execution speed at the studied context length.
Three routes are available to a deployment-oriented study, and they act on different parts of the cost. The first is to replace the reference scan with the fused selective-scan kernel, which removes the per-step Python 3.14 overhead identified above and is expected to give the largest single reduction without changing the model. The second is structured pruning of the width parameters that dominate the parameter count, namely, the state dimension N and the expansion factor, together with post-training quantization of the linear projections, both of which reduce memory traffic in the recurrence. The third is knowledge distillation of the trained forecaster into a smaller recurrent or linear student, which is attractive here because the deployed artefact is a deterministic point forecaster and the DVOF objective is applied only during offline fine-tuning. A distilled student would nevertheless have to be re-checked against the operating-point diagnostics of Section 4.3.1, since compression that preserves MAE does not automatically preserve the attained quantile level that determines dispatch value.

4.2.2. Backbone-Selection Sensitivity

Figure 4 summarizes the completed Optuna study. The functional analysis of variance (fANOVA) diagnostic estimates that weight decay and model dimension account for 55.7 % and 24.0 % , respectively, of the normalized relative importance for validation MAE within the sampled search space. The remaining hyperparameters have lower marginal importance. This ranking is conditional on the trial distribution and parameter interactions and should not be interpreted as a causal decomposition. Its role is to document the stability of backbone selection, not to infer sensitivity of realized recourse cost.

4.3. Evaluation of the Proposed DVOF Framework

The evaluation proceeds in three steps. Section 4.3.1 establishes where each training objective places the forecast by measuring the quantile level the forecast attains and comparing it with the level implied by the recourse model. Section 4.3.3 quantifies what that placement costs against an accuracy-oriented control and a cost-aware quantile control, and Section 4.3.5 examines how both results respond to the recourse scenario and to the training weight. Unless stated otherwise, the configuration is M shed : M curt = 10 : 1 with β = 35 , and all values are test-set results over the six matched seeds { 0 , 11 , 22 , 33 , 44 , 55 } .
Two quantile controls are used throughout and are kept distinct. The fixed control uses τ = 2 / 9 , the value inherited from the related value-oriented forecasting formulation [17] and is retained for comparability with that work rather than as an optimum for the present scenario. The validation-selected control chooses τ separately for each seed from the candidate set { 0.05 , 0.0909 , 0.10 , 1 / 6 , 2 / 9 , 0.30 , 0.50 } using validation imbalance penalty only and is therefore the stronger benchmark.

4.3.1. Attained Operating Point of the Trained Forecasts

Equation (14) identifies the recourse-optimal forecast as the conditional τ -quantile of wind power, with τ = 0.0909 under the nominal penalty ratio. This closed form admits a direct test of whether the differentiable layer transmits the intended economic signal. The quantile level that a trained forecast attains is measurable on the test set as the empirical frequency Pr ( w w ^ ) and can be compared with τ . The measurement requires no quantile supervision and applies equally to point and quantile forecasters.
Table 5 reports the result. The MSE Loss model operates near the conditional median, at 0.530 and 0.539 , as expected for an estimator of the conditional mean. The Proposed DVOF Loss model operates at 0.135 and 0.164 , having moved most of the distance from the median toward τ . It reaches this operating point from the dispatch gradient alone, without being supplied with a quantile level and without any search over τ . The validation-selected Pinball Loss model, which reaches its operating point by an explicit seven-point search on the validation split, settles at the closely similar levels 0.149 and 0.166 . The fixed τ = 2/9 control stops short of both, at 0.245 and 0.240 , which is the expected consequence of importing a quantile calibrated for a different penalty ratio.
The agreement between an end-to-end trained point forecaster and an explicitly tuned quantile forecaster indicates that the KKT-based gradient of Section 3.2 carries the asymmetry of the recourse model faithfully enough to recover the economically correct operating point without supervision. Both learned levels remain above τ for two reasons. The blended objective of Equation (39) retains an MSE term that moderates the conservative shift, and the generator bounds truncate the interior solution on which Equation (14) is derived.

4.3.2. Stability of the Operating Point Across Operating Conditions

Equation (14) further shows that τ depends only on the penalty ratio and not on the weather regime. A value-oriented forecaster should therefore hold a condition-independent operating point, whereas a forecaster that fixes a quantile of the conditional distribution need not, because the same probability level maps to different economic positions as the conditional spread changes. This provides a second and sharper test of the differentiable layer.
The held-out period is stratified by terciles of the horizon-aligned 100 m NWP wind speed. This variable is available at the forecast origin and is independent of the realized wind power, so the stratification cannot induce a selection artefact. Table 6 reports the attained quantile level within each stratum. The Proposed DVOF Loss model varies by 0.048 at H = 1 and 0.035 at H = 24 across the three strata, whereas every control varies more. The fixed- τ control drifts from 0.319 in low wind to 0.190 in high wind at H = 1, a spread of 0.130 , and the validation-selected control spreads by 0.097 and 0.075 .
Because each seed supplies a matched observation of every model, the comparison can be made pairwise. At H = 24, the Proposed DVOF Loss model has the smaller spread in all six seeds against each of the three controls, and the paired differences are statistically separated, with | t | between 3.44 and 5.65 against a critical value of 2.57 . At H = 1, it has the smaller spread in five of six seeds against each control, with | t | between 2.02 and 2.95 , so the ordering is consistent but separated only against the fixed- τ control. The distinction is one of the quantity being held fixed. A quantile head equalizes a probability level, whereas the dispatch gradient equalizes the marginal economic trade-off, and it is the latter that Equation (14) identifies as constant across operating conditions.

4.3.3. Realized Dispatch Cost Against the Accuracy and Quantile Controls

Table 7 first establishes the forecast–decision mismatch that motivates DVOF. From the table, the MSE Loss model attains the lowest MAE but the highest realized imbalance penalty at both horizons. Relative to this accuracy-oriented reference, the Proposed DVOF Loss model reduces the mean penalty by 47.8 % at H = 1 and 52.1 % at H = 24, while accepting a larger MAE and a more negative error location. The validation-selected Pinball Loss model then provides the relevant cost-aware comparison. Its mean penalty is close to that of DVOF at H = 1 ( 18.42 and 18.29 c.u. h 1 , respectively) and lower at H = 24 ( 26.11 vs. 26.59 c.u. h 1 ). The paired six-seed comparison detects no significant difference at either horizon ( | t | = 1.06 and 1.12 , respectively). The two methods nevertheless arrive at these outcomes differently. The Pinball model selects a quantile from a validation bank, whereas DVOF obtains its forecast adjustment from the dispatch gradient without prescribing a quantile level.
Table 8 gives the seedwise distribution behind these means. Figure 5 decomposes the penalty into its physical parts and is discussed in Section 4.3.6.
Figure 6 reports the distribution of the signed forecast error itself, pooled over the matched seeds and the complete held-out set, and identifies what the value-oriented objective changes and what it leaves alone. The standard deviation is nearly common to all four models, at 0.092 to 0.098 at H = 1 and 0.145 to 0.153 at H = 24, so the training objective does not measurably alter the sharpness of the forecast. What moves is the location. The whole distribution is displaced downward, with the median falling from + 0.003 to 0.087 at H = 1 and from + 0.009 to 0.139 at H = 24. This is the signature of a quantile shift rather than of an improvement in predictive skill, and it is the distributional counterpart of the attained levels in Table 5. The left panel of Figure 7 shows the weekly aggregation of the realized penalty over the whole held-out period.

4.3.4. Event-Conditioned Evaluation

An aggregate mean can conceal the hours that matter operationally. The held-out period is therefore partitioned by rules fixed before any test result was examined. High-ramp hours are those whose realized wind change | w t w t 1 | reaches the 90th percentile of the same statistic on the training split, a threshold of 0.1516 p.u., which selects 305 of the 2508 test h. High-imbalance hours are those whose realized recourse under the accuracy-oriented MSE reference reaches its 90th test percentile, which selects 251 h. The high-imbalance partition is derived once from the reference model and applied unchanged to every model, so all models are scored on a common set of hours.
Table 9 examines whether the aggregate result persists when the dispatch is stressed. The MSE Loss model remains the highest-penalty model in every stratum, and the gap between the accuracy-oriented and cost-aware objectives widens markedly in high-imbalance hours. For example, at H = 24, its mean penalty rises to 240.25 c.u. h 1 , compared with 74.75 for DVOF, whereas the corresponding all-hours values are 55.47 and 26.59 c.u. h 1 .
Within the cost-aware comparison, the horizon-dependent ordering in Table 7 is retained. DVOF has the lower mean penalty in all H = 1 strata, whereas the validation-selected Pinball model is lower in every H = 24 stratum. Thus, the aggregate comparison is not driven by the inclusion of undemanding hours. At H = 1, ramp events remain the most difficult stratum for all four models, consistent with the reduced value of persistence information during rapid wind changes.

4.3.5. Sensitivity to the Recourse Scenario and the Training Weight

Two quantities are varied here and must not be confused. The penalty ratio is a property of the operating environment and changes the recourse gradient itself. The trade-off weight β is a property of the training objective in Equation (39) and changes how strongly the forecaster responds to that gradient. Neither is a transferable physical constant.
The penalty-ratio analysis examines whether the forecast–decision mismatch depends on the assumed shortage asymmetry. The ratios 2 : 1 , 5 : 1 , 10 : 1 , and 15 : 1 are stylized recourse-cost scenarios, not calibrated settlement rules for particular jurisdictions. They represent progressively greater relative consequences of energy shortage that may arise under different reliability requirements, adequacy conditions, and imbalance-management arrangements.
For each ratio, the Proposed DVOF Loss model is retrained because the recourse gradient changes with the cost scenario. The MSE Loss model is re-settled under the corresponding ratio, and the Pinball quantile is selected from validation penalty independently for each seed. The right panel of Figure 7 reports the matched H = 24 comparison, with the quantile selected at each ratio shown beneath the corresponding tick. As the ratio increases from 2 : 1 to 15 : 1 , the observed mean penalty reduction of DVOF relative to MSE rises from 4.1 % to 58.8 % . Thus, stronger shortage asymmetry induces a stronger incentive for conservative wind forecasts.
The validation-selected Pinball Loss model nevertheless attains a lower observed mean penalty than the Proposed DVOF Loss model at every examined ratio. The penalty-ratio scenarios therefore support a bounded conclusion. The value of conservative forecasting increases systematically with recourse asymmetry, but the differentiable layer does not universally dominate a well-tuned quantile policy in the simplified dispatch model.
Turning to the training weight, Table 10 reports the matched H = 24 analysis for β { 0.1 , 5 , 20 , 35 , 60 } under the nominal penalty ratio. The observed mean penalty decreases from 47.98 ± 6.42 at β = 0.1 to 26.51 ± 1.48 at β = 60 . Over the same range, the mean MAE increases from 0.1110 ± 0.0016 to 0.1702 ± 0.0214 , bias becomes more negative, load shedding decreases, and wind curtailment increases. These coordinated changes show that β traces the transition from an accuracy-oriented forecast to an increasingly conservative value-oriented one, and that the operating point of Section 4.3.1 is reached by moving along this path rather than by a discrete switch.
The marginal reduction in penalty diminishes in the high- β region, and the confidence intervals for β = 35 and β = 60 overlap. Accordingly, β = 35 is retained as a representative high-value configuration, not as a uniquely optimal or transferable constant. Its interpretation is conditional on the loss normalization and recourse configuration used here.

4.3.6. Decision-Level Interpretation

Figure 5 provides a recourse-level interpretation of the lower DVOF penalty despite a larger MAE. At H = 24, mean load shedding decreases from 0.046 p.u. for the MSE Loss model to 0.009 p.u. for the Proposed DVOF Loss model, while wind curtailment increases from 0.097 to 0.175 p.u. Because load shedding carries the larger penalty, the avoided shortage cost exceeds the additional curtailment cost. The effect is therefore a redistribution of recourse, not a simultaneous reduction in every imbalance quantity.
The differentiable dispatch layer makes this trade-off visible at the decision level. When thermal generation is marginal, a lower wind forecast increases scheduled thermal generation. If realized wind is low, this adjustment can avoid expensive load shedding, whereas if realized wind is high, the same adjustment increases curtailment. The value of a conservative forecast is therefore conditional on the realized operating regime and is not guaranteed at every timestamp.
Figure 8 illustrates both outcomes using the pre-specified H = 24, seed-0 case-selection rules. In the shortage-risk case, the DVOF forecast decreases from 0.277 to 0.119 p.u., scheduled thermal generation increases from 0.667 to 0.825 p.u., and load shedding decreases from 0.150 p.u. to zero. The realized penalty consequently decreases from 149.62 to 0.85 c.u./h. In the high-wind case, the forecast decreases from 0.533 to 0.387 p.u., and the thermal schedule increases from 0.082 to 0.228 p.u. Wind curtailment then increases from 0.268 to 0.415 p.u., and the penalty increases from 26.82 to 41.46 c.u./h. These cases demonstrate why the same conservative adjustment can be beneficial under shortage risk yet costly under a high-wind realization.
This analysis concerns the forecast-to-dispatch mapping rather than the attribution of meteorological inputs. It identifies when a forecast adjustment changes scheduled generation and recourse, but it does not identify which NWP variable or historical observation causes that adjustment. Feature-level attribution would require a separate pre-specified analysis and is not inferred from these decision traces.

4.4. Discussion and Limitations

Overall, the operating-point analysis, the recourse decomposition, and the event-conditioned results explain what the differentiable layer contributes. The layer translates asymmetric recourse costs into a systematic conservative adjustment of the forecast and thereby reaches an operating point close to that obtained by validation-selected quantile regression. The latter reaches this point by an explicit quantile search, whereas DVOF learns it from the dispatch gradient. The matched six-seed comparison does not resolve a realized-penalty difference between the two approaches at either horizon. However, the present evidence supports the gradient-based construction of the operating point.
This interpretation is restricted to the controlled case study. The wind and load tracks form a standardized rather than contemporaneous scenario, and the dispatch model is single-area, single-period, and convex. The penalty-ratio analysis varies the assumed recourse asymmetry but not wind penetration, generator limits, demand levels, network conditions, or intertemporal resources. Independent wind datasets, contemporaneous wind–load records, broader operating-condition studies, and network-constrained multi-period scheduling are therefore needed to assess operational generality.

5. Conclusions

This paper develops a dispatch-value-oriented forecasting framework for wind-integrated economic dispatch. A Mamba-style forecaster is coupled with a differentiable economic-dispatch layer so that the gradient of realized second-stage recourse cost complements the conventional forecasting loss. The case studies show that forecast accuracy and dispatch value need not coincide. Relative to the MSE Loss model, DVOF substantially reduces the realized imbalance penalty while shifting the forecast toward the conservative operating point implied by the asymmetric recourse model.
The operating-point and recourse analyses clarify this adjustment. DVOF approaches the recourse-implied quantile without being supplied with a quantile level or searching over one, and it maintains this operating point more consistently across the examined wind regimes than the quantile controls. The validation-selected Pinball Loss model attains a closely related operating point through direct search. Its realized penalties are close to those of DVOF, with no significant difference detected in the matched six-seed comparison at either horizon.
The conclusions are limited to the standardized wind–load scenario and the single-area, single-period dispatch formulation studied here. Further studies should test the mechanism using independent and contemporaneous data, market-calibrated settlement rules, and broader operating conditions. Extending the framework to multi-period network-constrained scheduling with unit commitment, ramping, reserves, and storage is particularly important because the cost-relevant operating point need not remain stationary in such settings.

Author Contributions

Conceptualization, S.W. and W.Z.; methodology, S.W.; software, S.W. and W.Z.; validation, W.Z.; formal analysis, S.W.; investigation, W.Z.; resources, Y.S. and D.S.; data curation, S.W.; writing—original draft preparation, S.W.; writing—review and editing, Y.S. and D.S.; visualization, S.W.; supervision, Y.S. and D.S.; project administration, S.W.; funding acquisition, S.W. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the Natural Science Foundation of Jiangsu Province, under grant BK20251065, and in part by the Scientific Research Foundation of Nanjing Institute of Technology, under grant YKJ202518.

Data Availability Statement

This study uses the publicly available Global Energy Forecasting Competition 2014 (GEFCom2014) dataset [31].

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Global Wind Energy Council. Global Wind Report 2026; Global Wind Energy Council: Lisbon, Portugal, 2026. [Google Scholar]
  2. Yan, J.; Li, Y.; Wang, H.; Han, S.; Shang, W.; Liu, Y. Wind Power Forecasting Based on Large Time Series Model. Engineering 2025, in press. [Google Scholar] [CrossRef] [Scilit]
  3. Zhao, J.; Li, F.; Zhang, Q. Impacts of renewable energy resources on the weather vulnerability of power systems. Nat. Energy 2024, 9, 1407–1414. [Google Scholar] [CrossRef] [Scilit]
  4. Wang, S.; Sun, Y.; Zhang, W.; Srinivasan, D. Optimization of deterministic and probabilistic forecasting for wind power based on ensemble learning. Energy 2025, 319, 134884. [Google Scholar] [CrossRef] [Scilit]
  5. Wang, Y.; Zou, R.; Liu, F.; Zhang, L.; Liu, Q. A review of wind speed and wind power forecasting with deep neural networks. Appl. Energy 2021, 304, 117766. [Google Scholar] [CrossRef] [Scilit]
  6. Sideratos, G.; Hatziargyriou, N.D. An Advanced Statistical Method for Wind Power Forecasting. IEEE Trans. Power Syst. 2007, 22, 258–265. [Google Scholar] [CrossRef] [Scilit]
  7. Wang, S.; Sun, Y.; Zhou, Y.; Jamil Mahfoud, R.; Hou, D. A New Hybrid Short-Term Interval Forecasting of PV Output Power Based on EEMD-SE-RVM. Energies 2020, 13, 87. [Google Scholar] [CrossRef] [Scilit]
  8. Hewamalage, H.; Bergmeir, C.; Bandara, K. Recurrent neural networks for time series forecasting: Current status and future directions. Int. J. Forecast. 2021, 37, 388–427. [Google Scholar] [CrossRef] [Scilit]
  9. Kisvari, A.; Lin, Z.; Liu, X. Wind power forecasting—A data-driven method along with gated recurrent neural network. Renew. Energy 2021, 163, 1895–1909. [Google Scholar] [CrossRef] [Scilit]
  10. Song, Y.; Tang, D.; Yu, J.; Yu, Z.; Li, X. Short-term forecasting based on graph convolution networks and multiresolution convolution neural networks for wind power. IEEE Trans. Ind. Inf. 2023, 19, 1691–1702. [Google Scholar] [CrossRef] [Scilit]
  11. Liu, J.; Shi, Q.; Han, R.; Yang, J. A Hybrid GA–PSO–CNN Model for Ultra-Short-Term Wind Power Forecasting. Energies 2021, 14, 6500. [Google Scholar] [CrossRef] [Scilit]
  12. Dai, X.; Liu, G.P.; Hu, W. An online-learning-enabled self-attention-based model for ultra-short-term wind power forecasting. Energy 2023, 272, 127173. [Google Scholar] [CrossRef] [Scilit]
  13. Zhu, S.; Zheng, J.; Ma, Q. MR-Transformer: Multiresolution Transformer for Multivariate Time Series Prediction. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 1171–1183. [Google Scholar] [CrossRef] [Scilit]
  14. Zeng, A.; Chen, M.; Zhang, L.; Xu, Q. Are transformers effective for time series forecasting? In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; AAAI Press: Washington, DC, USA, 2023; Volume 37, pp. 11121–11128. [Google Scholar]
  15. Gu, A.; Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar]
  16. Dao, T.; Gu, A. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality. In Proceedings of the 41st International Conference on Machine Learning, Vienna, Austria, 21–27 July 2024. [Google Scholar]
  17. Zhang, Y.; Jia, M.; Wen, H.; Bian, Y.; Shi, Y. Toward Value-Oriented Renewable Energy Forecasting: An Iterative Learning Approach. IEEE Trans. Smart Grid 2025, 16, 1962–1974. [Google Scholar] [CrossRef] [Scilit]
  18. Wahdany, D.; Schmitt, C.; Cremer, J.L. More than accuracy: End-to-end wind power forecasting that optimises the energy system. Electr. Power Syst. Res. 2023, 221, 109384. [Google Scholar] [CrossRef] [Scilit]
  19. Li, G.; Chiang, H.D. Toward Cost-Oriented Forecasting of Wind Power Generation. IEEE Trans. Smart Grid 2018, 9, 2508–2517. [Google Scholar] [CrossRef] [Scilit]
  20. Zhao, C.; Wan, C.; Song, Y. Cost-Oriented Prediction Intervals: On Bridging the Gap Between Forecasting and Decision. IEEE Trans. Power Syst. 2022, 37, 3048–3062. [Google Scholar] [CrossRef] [Scilit]
  21. Zhang, Y.; Wen, H.; Wu, Q. A Contextual Bandit Approach for Value-Oriented Prediction Interval Forecasting. IEEE Trans. Smart Grid 2024, 15, 2271–2281. [Google Scholar] [CrossRef] [Scilit]
  22. Koenker, R.; Bassett, G. Regression Quantiles. Econometrica 1978, 46, 33–50. [Google Scholar] [CrossRef] [Scilit]
  23. Elmachtoub, A.N.; Grigas, P. Smart “Predict, then Optimize”. Manag. Sci. 2022, 68, 9–26. [Google Scholar] [CrossRef] [Scilit]
  24. Donti, P.L.; Amos, B.; Kolter, J.Z. Task-Based End-to-End Model Learning in Stochastic Optimization. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  25. Agrawal, A.; Amos, B.; Barratt, S.; Boyd, S.; Diamond, S.; Kolter, J.Z. Differentiable Convex Optimization Layers. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 8–14 December 2019; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
  26. Muñoz, M.Á.; Morales, J.M.; Pineda, S. Feature-Driven Improvement of Renewable Energy Forecasting and Trading. IEEE Trans. Power Syst. 2020, 35, 3753–3763. [Google Scholar] [CrossRef] [Scilit]
  27. Stratigakos, A.; Camal, S.; Michiorri, A.; Kariniotakis, G. Prescriptive trees for integrated forecasting and optimization applied in trading of renewable energy. IEEE Trans. Power Syst. 2022, 37, 4696–4708. [Google Scholar] [CrossRef] [Scilit]
  28. Chen, X.; Yang, Y.; Liu, Y.; Wu, L. Feature-Driven Economic Improvement for Network-Constrained Unit Commitment: A Closed-Loop Predict-and-Optimize Framework. IEEE Trans. Power Syst. 2022, 37, 3104–3118. [Google Scholar] [CrossRef] [Scilit]
  29. Sang, L.; Xu, Y.; Long, H.; Hu, Q.; Sun, H. Electricity Price Prediction for Energy Storage System Arbitrage: A Decision-Focused Approach. IEEE Trans. Smart Grid 2022, 13, 2822–2832. [Google Scholar] [CrossRef] [Scilit]
  30. Ellinas, P.; Kekatos, V.; Tsaousoglou, G. Decision-Focused Learning under Decision Dependent Uncertainty for Power Systems with Price-Responsive Demand. Electr. Power Syst. Res. 2024, 235, 110665. [Google Scholar] [CrossRef] [Scilit]
  31. Hong, T.; Pinson, P.; Fan, S.; Zareipour, H.; Troccoli, A.; Hyndman, R.J. Probabilistic energy forecasting: Global Energy Forecasting Competition 2014 and beyond. Int. J. Forecast. 2016, 32, 896–913. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Proposed DVOF framework. Solid arrows show the forward path from wind forecasting to scheduled thermal generation, recourse settlement and realized recourse cost. The dashed arrow shows the realized-recourse-cost gradient propagated through the differentiable economic dispatch layer back to the forecaster.
Figure 1. Proposed DVOF framework. Solid arrows show the forward path from wind forecasting to scheduled thermal generation, recourse settlement and realized recourse cost. The dashed arrow shows the realized-recourse-cost gradient propagated through the differentiable economic dispatch layer back to the forecaster.
Energies 19 04053 g001
Figure 2. Mamba-style forecaster used in the proposed DVOF framework. The main path runs from the causal input through the linear embedding with positional embedding, the L selective-SSM blocks, layer normalization and the MLP head to the wind forecast. The panel on the right is an expanded view of one selective-SSM block. Its dashed leader is an explanatory link to that block only and does not represent a data-flow connection to the linear embedding.
Figure 2. Mamba-style forecaster used in the proposed DVOF framework. The main path runs from the causal input through the linear embedding with positional embedding, the L selective-SSM blocks, layer normalization and the MLP head to the wind forecast. The panel on the right is an expanded view of one selective-SSM block. Its dashed leader is an explanatory link to that block only and does not represent a data-flow connection to the linear embedding.
Energies 19 04053 g002
Figure 3. Forecast trajectories of the compared models at the one-step-ahead horizon.
Figure 3. Forecast trajectories of the compared models at the one-step-ahead horizon.
Energies 19 04053 g003
Figure 4. Forecasting-backbone hyperparameter sensitivity for the completed Optuna study. The left panel reports relative importance from functional analysis of variance for validation MAE across 47 completed trials out of 50 total trials. The right panel shows the observed validation MAE against weight decay, with colour indicating model dimension.
Figure 4. Forecasting-backbone hyperparameter sensitivity for the completed Optuna study. The left panel reports relative importance from functional analysis of variance for validation MAE across 47 completed trials out of 50 total trials. The right panel shows the observed validation MAE against weight decay, with colour indicating model dimension.
Energies 19 04053 g004
Figure 5. Decomposition of the realized imbalance penalty into load shedding and wind curtailment for the four loss-defined models at H = 1 (left) and H = 24 (right). Values are the mean recourse quantity per dispatch period in p.u., averaged over the six matched seeds. Moving from the accuracy-oriented model to the cost-aware models trades shedding for curtailment monotonically, and the two cost-aware controls that reach the recourse-optimal operating point are almost indistinguishable in physical terms.
Figure 5. Decomposition of the realized imbalance penalty into load shedding and wind curtailment for the four loss-defined models at H = 1 (left) and H = 24 (right). Values are the mean recourse quantity per dispatch period in p.u., averaged over the six matched seeds. Moving from the accuracy-oriented model to the cost-aware models trades shedding for curtailment monotonically, and the two cost-aware controls that reach the recourse-optimal operating point are almost indistinguishable in physical terms.
Energies 19 04053 g005
Figure 6. Density of the signed forecast error e t = w ^ t w t over the complete held-out set, pooled across the matched seeds, at H = 1 (left) and H = 24 (right). The three cost-aware models are displaced toward negative error relative to the accuracy-oriented model while retaining a comparable width, which is the signature of a quantile shift rather than of a change in predictive sharpness.
Figure 6. Density of the signed forecast error e t = w ^ t w t over the complete held-out set, pooled across the matched seeds, at H = 1 (left) and H = 24 (right). The three cost-aware models are displaced toward negative error relative to the accuracy-oriented model while retaining a comparable width, which is the signature of a quantile shift rather than of a change in predictive sharpness.
Energies 19 04053 g006
Figure 7. (Left) realized imbalance penalty over the complete held-out period at H = 24, aggregated to weekly means so that the full series is legible. (Right) sensitivity of the penalty to the shedding-to-curtailment ratio, with the proposed model retrained at each ratio and the quantile benchmark reselected on validation data; the quantile selected at each ratio is given beneath the corresponding tick. Error bars are two-sided 95 % confidence intervals over the six matched seeds on the left and over the three matched seeds of the retraining study on the right.
Figure 7. (Left) realized imbalance penalty over the complete held-out period at H = 24, aggregated to weekly means so that the full series is legible. (Right) sensitivity of the penalty to the shedding-to-curtailment ratio, with the proposed model retrained at each ratio and the quantile benchmark reselected on validation data; the quantile selected at each ratio is given beneath the corresponding tick. Error bars are two-sided 95 % confidence intervals over the six matched seeds on the left and over the three matched seeds of the retraining study on the right.
Energies 19 04053 g007
Figure 8. Decision-level interpretation of conservative forecast adjustments at H = 24. The upper row is the shortage-risk case, and the lower row is the high-wind case, both drawn from seed 0 and selected by pre-specified rules. In each row, forecast trajectories over the surrounding day are shown on the left, and scheduled thermal generation, realized recourse quantities and the imbalance penalty at the settled hour are shown on the right. A recourse bar of zero height is a result rather than a missing entry, because the net position in a given hour is either short or long, so load shedding and wind curtailment cannot both be positive; such bars are labelled with their value.
Figure 8. Decision-level interpretation of conservative forecast adjustments at H = 24. The upper row is the shortage-risk case, and the lower row is the high-wind case, both drawn from seed 0 and selected by pre-specified rules. In each row, forecast trajectories over the surrounding day are shown on the left, and scheduled thermal generation, realized recourse quantities and the imbalance penalty at the settled hour are shown on the right. A recourse bar of zero height is a result rather than a missing entry, because the net position in a given hour is either short or long, so load shedding and wind curtailment cannot both be positive; such bars are labelled with their value.
Energies 19 04053 g008
Table 1. Main hyperparameters and selection protocol of the MSE-trained Mamba-style forecaster.
Table 1. Main hyperparameters and selection protocol of the MSE-trained Mamba-style forecaster.
GroupHyperparameterValueCandidates
InputHistorical context length72 h { 12 , 24 , 48 , 72 }
Forecast horizon H1/24 h-
Mamba SSMModel dimension96 { 32 , 64 , 96 , 128 }
Number of SSM layers3 { 1 , 2 , 3 }
State dimension16 { 8 , 16 , 32 }
Convolution width7 { 3 , 4 , 5 , 7 }
Expansion factor2 { 1 , 2 , 3 }
TrainingDropout0.2740 [ 0 , 0.3 ]
OptimizerAdam-
Learning rate 9.40 × 10 4 [ 10 4 , 3 × 10 3 ]
Weight decay 4.12 × 10 5 [ 10 7 , 10 3 ]
Batch size256 { 128 , 256 , 512 }
Table 2. Two-unit economic-dispatch test system configuration (per unit on the wind rated power).
Table 2. Two-unit economic-dispatch test system configuration (per unit on the wind rated power).
Thermal Unit p min (p.u.) p max (p.u.)a (c.u./p.u.2h)b (c.u./p.u. h)
T100.801230
T200.501535
Peak load 1.6 p.u.; wind rated power 1.0 p.u.; no PV. Stylized recourse penalties M shed = 1000 , M curt = 100 c.u./p.u. h (ratio 10 : 1 ).
Table 3. Forecasting accuracy of the proposed Mamba forecaster against standard deep-learning baselines and the persistence reference. Values are the mean ± standard deviation over six random seeds. The best MAE, RMSE, and R 2 values in each horizon are shown in bold. Bias is reported as a signed diagnostic.
Table 3. Forecasting accuracy of the proposed Mamba forecaster against standard deep-learning baselines and the persistence reference. Values are the mean ± standard deviation over six random seeds. The best MAE, RMSE, and R 2 values in each horizon are shown in bold. Bias is reported as a signed diagnostic.
HorizonModelMAERMSE R 2 Bias
H = 1Persistence0.07310.10660.8810.0001
LSTM [8]0.0664 ± 0.00040.0927 ± 0.00070.910 ± 0.001 0.0034 ± 0.0027
GRU [9]0.0661 ± 0.00050.0930 ± 0.00050.909 ± 0.001 0.0001 ± 0.0038
Transformer [13]0.0673 ± 0.00090.0941 ± 0.00080.907 ± 0.002 0.0087 ± 0.0060
DLinear [14]0.0990 ± 0.02050.1237 ± 0.02470.834 ± 0.074 0.0105 ± 0.0070
Proposed0.0657 ± 0.00110.0921 ± 0.00110.911 ± 0.002 0.0048 ± 0.0096
H = 24Persistence0.33640.4150 0.813 0.0027
LSTM [8]0.1127 ± 0.00130.1479 ± 0.00120.770 ± 0.004 0.0166 ± 0.0129
GRU [9]0.1143 ± 0.00160.1508 ± 0.00140.761 ± 0.005 0.0250 ± 0.0086
Transformer [13]0.1144 ± 0.00170.1515 ± 0.00300.758 ± 0.010 0.0245 ± 0.0137
DLinear [14]0.1285 ± 0.00140.1593 ± 0.00200.733 ± 0.007 0.0170 ± 0.0036
Proposed0.1105 ± 0.00220.1459 ± 0.00230.776 ± 0.007 0.0132 ± 0.0113
Table 4. Computational cost of the compared backbones under the same input configuration (T = 72).
Table 4. Computational cost of the compared backbones under the same input configuration (T = 72).
ModelParametersInference ( μ s/Sample)Training ( μ s/Sample)
LSTM [8]≈116  k 2.3 6.7
GRU [9]≈87 k 2.0 6.5
Transformer [13]≈158 k 4.8 22.0
DLinear [14]≈1.4  k 0.8 4.5
Proposed≈330  k 88.4 424.5
Table 5. Quantile level attained by each trained forecast on the test split, measured as Pr ( w w ^ ) . The recourse-optimal level under the nominal 10 : 1 ratio is τ = 0.0909 . Values are mean over the six matched seeds { 0 , 11 , 22 , 33 , 44 , 55 } , with the seedwise range in brackets.
Table 5. Quantile level attained by each trained forecast on the test split, measured as Pr ( w w ^ ) . The recourse-optimal level under the nominal 10 : 1 ratio is τ = 0.0909 . Values are mean over the six matched seeds { 0 , 11 , 22 , 33 , 44 , 55 } , with the seedwise range in brackets.
ModelH = 1H = 24
MSE Loss 0.530   [ 0.478 , 0.606 ] 0.539   [ 0.493 , 0.595 ]
Pinball Loss, fixed τ = 2/9 [17] 0.245   [ 0.187 , 0.289 ] 0.240   [ 0.205 , 0.263 ]
Pinball Loss, validation selected 0.149   [ 0.115 , 0.181 ] 0.166   [ 0.130 , 0.197 ]
Proposed DVOF Loss 0.135   [ 0.114 , 0.180 ] 0.164   [ 0.146 , 0.182 ]
Recourse-optimal reference τ = M curt / ( M shed + M curt ) = 0.0909 .
Table 6. Attained quantile level Pr ( w w ^ ) within terciles of the horizon-aligned 100 m NWP wind speed, a conditioning variable known at the forecast origin. Spread is the range across the three strata. Values are means over the six matched seeds { 0 , 11 , 22 , 33 , 44 , 55 } .
Table 6. Attained quantile level Pr ( w w ^ ) within terciles of the horizon-aligned 100 m NWP wind speed, a conditioning variable known at the forecast origin. Spread is the range across the three strata. Values are means over the six matched seeds { 0 , 11 , 22 , 33 , 44 , 55 } .
ModelH = 1H = 24
Low Mid High Spread Low Mid High Spread
MSE Loss 0.551 0.563 0.477 0.086 0.578 0.558 0.483 0.095
Pinball, fixed τ = 2/9 0.319 0.225 0.190 0.130 0.290 0.245 0.187 0.104
Pinball, validation selected 0.204 0.135 0.108 0.097 0.199 0.176 0.123 0.075
Proposed DVOF Loss 0.166 0.123 0.117 0.048 0.178 0.171 0.143 0.035
Table 7. Dispatch-value comparison of the four loss-defined models on the identical Mamba backbone. Both quantile controls are shown: the fixed reference τ = 2/9 and the validation-selected quantile.
Table 7. Dispatch-value comparison of the four loss-defined models on the identical Mamba backbone. Both quantile controls are shown: the fixed reference τ = 2/9 and the validation-selected quantile.
HModelPenaltyMAE Δ MSE Δ Q
H = 1MSE loss 35.02 ± 3.07 0.066
Pinball loss, fixed τ = 2/9 [17] 20.59 ± 1.33 0.084 41.2 %
Pinball loss, validation selected 18.42 ± 0.48 0.105 47.4 %
Proposed DVOF loss 18.29 ± 0.27 0.106 47.8 % 0.7 %
H = 24MSE loss 55.47 ± 4.05 0.111
Pinball loss, fixed τ = 2/9 [17] 29.47 ± 1.25 0.139 46.9 %
Pinball loss, validation selected 26.11 ± 0.84 0.165 52.9 %
Proposed DVOF loss 26.59 ± 1.02 0.167 52.1 % + 1.9 %
Values are mean ± two-sided 95 % confidence interval over the six matched seeds { 0 , 11 , 22 , 33 , 44 , 55 } . Penalty is the realized imbalance penalty under schedule-then-recourse dispatch and is reported in c.u. h 1 . Δ MSE and Δ Q are measured against the MSE and validation-selected Pinball Loss models, respectively. Bold face denotes the numerical minimum.
Table 8. Seedwise distribution of realized imbalance penalty in the controlled test scenario. Values summarize the same three matched runs used in Table 7.
Table 8. Seedwise distribution of realized imbalance penalty in the controlled test scenario. Values summarize the same three matched runs used in Table 7.
HModelMinimumMaximumSD
1MSE Loss 30.74 37.82 2.93
Pinball Loss ( τ = 2/9) 18.85 22.72 1.27
Pinball Loss (validation selected) 18.00 19.29 0.46
Proposed DVOF Loss 18.03 18.72 0.26
24MSE Loss 51.01 61.52 3.86
Pinball Loss ( τ = 2/9) 27.37 30.57 1.19
Pinball Loss (validation selected) 24.82 27.14 0.80
Proposed DVOF Loss 25.61 28.01 0.97
Table 9. Event-conditioned mean imbalance penalty (c.u. h 1 ) over the six matched seeds { 0 , 11 , 22 , 33 , 44 , 55 } . Strata are defined by rules fixed before the test results were examined. Hour counts are given in the header.
Table 9. Event-conditioned mean imbalance penalty (c.u. h 1 ) over the six matched seeds { 0 , 11 , 22 , 33 , 44 , 55 } . Strata are defined by rules fixed before the test results were examined. Hour counts are given in the header.
HModelAll HoursHigh RampHigh Imbalance
(2508 h) (305 h) (251 h)
H = 1MSE Loss 35.02 76.84 147.30
Pinball, fixed τ = 2/9 20.59 53.68 71.48
Pinball, validation selected 18.42 46.00 47.71
Proposed DVOF Loss 18.29 45.92 44.98
H = 24MSE Loss 55.47 77.13 240.25
Pinball, fixed τ = 2/9 29.47 46.18 105.76
Pinball, validation selected 26.11 39.80 70.91
Proposed DVOF Loss 26.59 40.34 74.75
Table 10. Matched H = 24 sensitivity to the DVOF trade-off weight β under the nominal recourse configuration. Penalty and MAE are given as mean ± two-sided 95 % confidence interval over the three matched seeds { 0 , 11 , 22 } ; bias, shedding and curtailment are means. Penalty is in c.u. h 1 , and the recourse quantities are in p.u.
Table 10. Matched H = 24 sensitivity to the DVOF trade-off weight β under the nominal recourse configuration. Penalty and MAE are given as mean ± two-sided 95 % confidence interval over the three matched seeds { 0 , 11 , 22 } ; bias, shedding and curtailment are means. Penalty is in c.u. h 1 , and the recourse quantities are in p.u.
β PenaltyMAEBiasSheddingCurtailment
0.1 47.98 ± 6.42 0.1110 ± 0.0016 0.011 0.0374 0.1056
5 33.26 ± 4.54 0.1283 ± 0.0081 0.073 0.0196 0.1367
20 29.07 ± 3.96 0.1493 ± 0.0335 0.113 0.0133 0.1575
35 27.36 ± 1.84 0.1672 ± 0.0176 0.140 0.0099 0.1745
60 26.51 ± 1.48 0.1702 ± 0.0214 0.148 0.0087 0.1777
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wang, S.; Zhang, W.; Sun, Y.; Srinivasan, D. Decision-Focused Wind Power Forecasting for Economic Dispatch Under Asymmetric Imbalance Costs. Energies 2026, 19, 4053. https://doi.org/10.3390/en19174053

AMA Style

Wang S, Zhang W, Sun Y, Srinivasan D. Decision-Focused Wind Power Forecasting for Economic Dispatch Under Asymmetric Imbalance Costs. Energies. 2026; 19(17):4053. https://doi.org/10.3390/en19174053

Chicago/Turabian Style

Wang, Sen, Wenjie Zhang, Yonghui Sun, and Dipti Srinivasan. 2026. "Decision-Focused Wind Power Forecasting for Economic Dispatch Under Asymmetric Imbalance Costs" Energies 19, no. 17: 4053. https://doi.org/10.3390/en19174053

APA Style

Wang, S., Zhang, W., Sun, Y., & Srinivasan, D. (2026). Decision-Focused Wind Power Forecasting for Economic Dispatch Under Asymmetric Imbalance Costs. Energies, 19(17), 4053. https://doi.org/10.3390/en19174053

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop