Next Article in Journal
Interaction-Specific Diffusion Scheduling Guided by Temporal Priors for Sequential Recommendation
Previous Article in Journal
STAC-ML: A Trust-Aware Security Architecture for Resource-Constrained NDN-IoT Networks with Lightweight ML-Assisted Threat Detection
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Short-Term Electric Power Load Forecasting Based on ISAO-VMD-BiTCN

School of Electrical and Energy Engineering, Shanghai Dianji University, Shanghai 201306, China
*
Author to whom correspondence should be addressed.
Electronics 2026, 15(19), 4393; https://doi.org/10.3390/electronics15194393
Submission received: 2 August 2026 / Revised: 19 September 2026 / Accepted: 21 September 2026 / Published: 24 September 2026

Abstract

Aiming to improve short-term power load forecasting under nonlinear and nonstationary load characteristics, this paper develops an Improved Snow Ablation Optimizer (ISAO) and applies it to the hyperparameter optimization of a Variational Mode Decomposition (VMD)- Bidirectional Temporal Convolutional Network (BiTCN) forecasting framework. First, VMD is used to decompose the original load sequence into multiple intrinsic mode function (IMF) components in order to reduce the non-stationarity of the signal. Second, to address the limitations of the Snow Ablation Optimizer (SAO), four strategies—quantum-chaotic hybrid initialization, elite-pool-guided dual-population updating, the quantum tunneling effect, and dynamic lens opposition-based learning—are introduced to form the ISAO algorithm. Finally, ISAO is used to independently optimize the BiTCN hyperparameters of each IMF component, and the prediction results of all components are superimposed and reconstructed to obtain the final forecast. CEC2022 benchmark function tests and experiments on actual load data show that the ISAO algorithm achieves superior optimization performance; compared with the SAO-VMD-BiTCN model, the RMSE of the proposed model decreases from 2342.23 MW to 2202.14 MW, a reduction of 5.98% and the proposed model outperforms comparison models such as VMD-BiTCN and PSO-VMD-BiTCN across the RMSE, MAE, MAPE, and R2 metrics. Furthermore, Diebold–Mariano (DM) test results provide additional statistical evidence for the favorable forecasting performance of the proposed framework across multiple random initializations.

1. Introduction

With the deepening advancement of smart grids and new-generation power system construction, short-term electric power load forecasting has become a core supporting technology for ensuring the safe dispatch, economic operation, and efficient market trading of electric power systems [1]. However, power load is influenced by multiple factors such as meteorological conditions, holiday schedules, and socio-economic activity, exhibiting pronounced nonlinearity and non-stationarity, which pose considerable challenges to achieving high-precision forecasting.
In terms of prediction-model architecture, research has also extended beyond conventional recurrent networks. Long Short-Term Memory (LSTM) is a widely used sequence-model baseline that alleviates the vanishing-gradient problem of vanilla RNNs, though a unidirectional LSTM only propagates information forward in time [2]. To better exploit both local and global feature correlations, Cui et al. [3] combined a Convolutional Neural Network (CNN)-based feature-extraction module with a self-attention encoder–decoder network for initial forecasting, followed by a residual-refinement module to further optimize the prediction, reporting improved accuracy and stability. Transformer-based encoder–decoder architectures built purely on self-attention have likewise been applied to load forecasting, offering stronger parallelism and long-range dependency modeling than LSTM [4]. These studies illustrate the potential of convolutional feature extraction, self-attention-based dependency modeling, and residual refinement for improving load forecasting performance. To provide a conventional recurrent baseline, LSTM is also included in the comparative experiments.
In terms of load-data preprocessing, signal decomposition methods have been widely introduced to reduce the complexity of load sequences. Variational Mode Decomposition (VMD), proposed by Dragomiretskiy et al., converts signal decomposition into a constrained variational problem and iteratively solves for each modal component in the frequency domain via the Alternating Direction Method of Multipliers (ADMM), effectively alleviating the mode-mixing problem of traditional decomposition methods. In recent years, VMD has been widely applied in short-term load forecasting: Yang Huping et al. [5] combined VMD with a CNN- Bidirectional Gated Recurrent Unit (BiGRU) model, reducing the non-stationarity of the load sequence through decomposition and effectively improving prediction accuracy; Wang Qing et al. [6] constructed a VMD- Temporal Convolutional Network (TCN) combined model, confirming that after VMD preprocessing, the frequency-domain characteristics of each component become clearer and prediction performance is significantly better than modeling the raw load sequence directly.
After VMD, the stationarity of the resulting components is markedly improved, and a compatible time-series prediction model is required to achieve high-precision fitting. Regarding prediction models, the TCN, based on dilated causal convolution and residual connections, enables parallel computation while offering a large receptive field and efficient training [7]. Building on this, the Bidirectional Temporal Convolutional Network (BiTCN) proposed by Sprangers et al. further improves prediction performance by modeling bidirectional contextual features in parallel through forward and backward branches [8]. Huang et al. [9] validated the effectiveness of BiTCN in domestic short-term load forecasting; ablation experiments showed that VMD-BiTCN can further reduce prediction error compared with a unidirectional TCN. Fan et al. [10] likewise confirmed that BiTCN, owing to its advantage of processing local features bidirectionally in parallel, performs better within combined prediction models. However, BiTCN hyperparameters such as the number and size of convolution kernels, the dropout rate, and the learning rate significantly affect prediction results, and manual tuning is inefficient and unlikely to yield a globally optimal solution [11].
To address the hyperparameter-optimization problem, metaheuristic algorithms such as Particle Swarm Optimization (PSO) [12], the Grey Wolf Optimizer (GWO) [13], and the Sparrow Search Algorithm (SSA) [14] have been widely used for neural-network hyperparameter tuning. The Snow Ablation Optimizer (SAO) proposed by Deng and Liu, which simulates the sublimation and melting behavior of snow to construct a dual-population cooperative mechanism, achieves an adaptive balance between global exploration and local exploitation [15]. Compared with classical algorithms such as PSO and GWO, it has a stronger ability to escape local optima and performs well on numerical-optimization benchmark tests, indicating good potential for further improvement and application. However, Xiao et al. [16] pointed out that SAO suffers from four shortcomings—uneven initial population distribution, an imbalance between exploration and exploitation, susceptibility to local optima, and insufficient late-stage convergence accuracy; Li et al. [17] likewise confirmed that the excessively strong exploitation capability of SAO leads to an imbalance between global and local search, making it prone to becoming trapped in local extrema in complex scenarios. These issues limit the effectiveness of SAO for BiTCN hyperparameter optimization, making targeted improvement necessary. To address these issues, this paper develops an Improved Snow Ablation Optimizer (ISAO) and applies it to the hyperparameter optimization of a VMD-BiTCN forecasting framework for short-term power load prediction. The main contributions of this paper are summarized as follows:
(1) To address the four documented shortcomings of SAO (uneven initial population distribution, imbalance between exploration and exploitation, susceptibility to local optima, and insufficient late-stage convergence accuracy), four improvement strategies—quantum-chaotic hybrid initialization, elite-pool-guided dual-population updating, the quantum tunneling effect, and dynamic lens opposition-based learning—are introduced to construct the ISAO.
(2) The optimization performance of ISAO is systematically validated on the CEC2022 benchmark functions against RIME, HBA, ARO, and the original SAO, with statistical significance confirmed via the Wilcoxon rank-sum test.
(3) The proposed ISAO is applied to the hyperparameter optimization of a VMD-BiTCN forecasting framework, in which VMD decomposes the load sequence into IMF components and ISAO independently optimizes the BiTCN hyperparameters for each component.
(4) The proposed ISAO-VMD-BiTCN framework is evaluated on real-world PJM load data against LSTM, VMD-BiTCN, GWO-VMD-BiTCN, PSO-VMD-BiTCN, and SAO-VMD-BiTCN, achieving the best overall performance among the compared forecasting models in terms of RMSE, MAE, MAPE, and R2.
The remainder of this paper is organized as follows. Section 2 introduces VMD and its role in reducing the non-stationarity of the original load sequence. Section 3 presents the BiTCN, which serves as the prediction backbone for each decomposed component. Section 4 reviews the SAO and details the proposed ISAO, which is used to automatically optimize the BiTCN hyperparameters. Section 5 integrates VMD, BiTCN, and ISAO into the proposed combined prediction model. Section 6 validates the proposed model on real-world PJM load data through comparative case studies against four benchmark models. Section 7 concludes the paper.

2. VMD

As the data-preprocessing stage of the proposed ISAO-VMD-BiTCN framework, this section introduces VMD, which is used to decompose the original load sequence into components with reduced non-stationarity. VMD decomposes the original load sequence f into K intrinsic mode function (IMF) components u k ( t ) , each with an independent center frequency and a limited bandwidth. Its constrained variational objective function is expressed as:
m i n { u k } , { w k } ∑ k = 1 K ∂ t δ ( t ) + j π t ∗ u k ( t ) e − j w k t 2 2
s . t . ∑ k = 1 K u k ( t ) = f ( t )
where u k denotes the decomposed modal components; w k denotes the center frequency of each component; δ ( t ) is the Dirac function; and ∗ denotes the convolution operation. Using a quadratic penalty factor α and Lagrange multiplier λ ( t ) , this constrained problem is converted into an unconstrained one and solved iteratively in the frequency domain via the Alternating Direction Method of Multipliers (ADMM).
In practical application, the number of decomposed modes K and the penalty factor α are the key parameters governing decomposition quality. This paper determines K using the center-frequency observation method, progressively increasing K while observing the distribution of the center frequencies of each IMF. Experiments show that when K = 5, the center frequencies of the components are evenly distributed with no overlap; the penalty factor α is set to its default value of 2000. Under these settings, the maximum VMD reconstruction error is only 4.52 × 10 − 2 , indicating that the decomposition accuracy fully satisfies the requirements of subsequent forecasting.

3. BiTCN Network

Section 2 decomposed the original load sequence into IMF components via VMD; this section introduces the BiTCN network, which serves as the prediction backbone for each of these components. BiTCN consists of two parallel branches: a forward TCN and a backward TCN. The forward TCN performs causal convolution on the input sequence in chronological order, while the backward TCN performs causal convolution after reversing the input sequence in time, thereby implicitly capturing the future context of the sequence. The core modules of TCN comprise three key techniques: causal convolution, dilated convolution, and residual connections. For an input sequence x = ( x 0 , x 1 , ⋯ , x t ) , the corresponding output y = ( y 0 , y 1 , ⋯ , y t ) after causal convolution is:
y t = f c ( x 0 , x 1 , x 2 , ⋯ , x t )
where f c denotes the causal convolution operation.
The dilated convolution structure is shown in Figure 1. For an input sequence x = ( x 0 , x 1 , ⋯ , x t ) , the hidden-layer output after dilated convolution is:
F ( l ) ( t ) = ∑ i = 0 k c − 1 w ( l ) ( i ) · x t − d ( l ) · i
where k c is the convolution kernel size, w ( l ) is the i − t h trainable kernel coefficient at layer l , and d ( l ) is the layer-dependent dilation factor, with d ( 1 ) = 1 , d ( 2 ) = 2 , d ( 3 ) = 3 , as shown in Figure 1. This dilated convolution module is shared by the forward and backward branches of BiTCN.
The residual connection adds the input x to the nonlinearly transformed F ( x ) :
y = F ( x ) + x

4. Snow Ablation Optimizer and Its Improvement

Because BiTCN’s prediction accuracy is highly sensitive to hyperparameter choices, this section focuses on their optimization: Section 4.1 reviews the original SAO, and Section 4.2 develops the proposed ISAO, which is later used to automatically tune these hyperparameters for each IMF component.

4.1. Snow Ablation Optimizer (SAO)

SAO is a metaheuristic developed by Deng and Liu that simulates the sublimation and melting of accumulated snow to achieve a trade-off between exploration and exploitation, thereby avoiding premature convergence. SAO mainly consists of an initialization stage, an exploration stage, an exploitation stage, and a dual-population mechanism.

4.1.1. Initialization Stage

In SAO, the iterative process begins with a randomly generated population. As shown below, the entire population is typically modeled as a matrix with N rows and D i m columns, where N represents the population size, and D i m represents the dimensionality of the solution space.
Z = l b + θ × ( u b − l b ) = Z 1,1 Z 1 , 2 ⋯ Z 1 , D i m − 1 Z 1 , D i m Z 2 , 1 Z 2 , 2 ⋯ Z 2 , D i m − 1 Z 2 , D i m ⋮ ⋮ ⋱ ⋮ ⋮ Z N − 1,1 Z N − 1,2 ⋯ Z N − 1 , D i m − 1 Z N − 1 , D i m Z N , 1 Z N , 2 ⋯ Z N , D i m − 1 Z N , D i m N × D i m
where l b and u b denote the lower and upper bounds of the solution space, respectively, and θ denotes a random number in [0, 1].

4.1.2. Exploration Stage

When snow, or the liquid water converted from snow, turns into vapor, the individuals exhibit highly dispersed characteristics owing to irregular motion. In this study, Brownian motion is used to simulate this behavior. As a stochastic process, Brownian motion is widely used to model animal foraging behavior, the ceaseless and irregular motion of particles, stock-price fluctuations, and similar phenomena. For standard Brownian motion, the step length is obtained from the probability density function of a normal distribution with mean 0 and variance 1, expressed as follows:
f B M ( x ; 0 , 1 ) = 1 2 π × e x p ( − x 2 2 )
Brownian motion can explore potential regions of the search space and therefore effectively reflects the diffusion of vapor within it. The position-update formula during the exploration process is:
Z i ( t + 1 ) = E l i t e ( t ) + B M i ( t ) ⨂ θ 1 × G ( t ) − Z i ( t ) + 1 − θ 1 × Z ¯ ( t ) − Z i ( t )
Z ¯ ( t ) = 1 N ∑ i = 1 N Z i ( t )
E l i t e ( t ) ∈ G ( t ) , Z s e c o n d ( t ) , Z t h i r d ( t ) , Z c ( t )
Z c ( t ) = 1 N 1 ∑ i = 1 N 1 Z i ( t )
where Z i ( t + 1 ) is the position of the i − t h individual at time t + 1 ; E l i t e ( t ) is a random variable selected from among G ( t ) , Z s e c o n d ( t ) , Z t h i r d ( t ) and Z c ( t ) ; B M i is a Brownian-motion coefficient governing the random-walk step of the i − t h individual; G ( t ) is the position of the best individual at time t ; Z s e c o n d ( t ) and Z t h i r d ( t ) are the second-best and third-best positions at time t ; and Z c ( t ) is the average position of the individuals at time t .

4.1.3. Exploitation Stage

When snow is converted into liquid water through melting, this process is simulated according to the snow-ablation process as follows:
M = D D F × ( T − T 1 )
where D D F is the snow-ablation coefficient, ranging over [0.35, 0.6]; T is the daily average temperature (°C); and T 1 is the base temperature (°C). D D F varies with time as:
D D F = 0.35 + 0.25 × e t t m a x − 1 e − 1
The position is then updated to simulate the snow-ablation process:
Z i ( t + 1 ) = M × G ( t ) + B M i ( t ) ⨂ θ 2 × G ( t ) − Z i ( t ) + 1 − θ 2 × Z ¯ ( t ) − Z i ( t )
where M is the snow-ablation rate and θ 2 is a random number in the range [−1, 1].

4.1.4. Dual-Population Mechanism

In SAO, balancing the exploration and exploitation stages prevents the algorithm from converging prematurely and enables a more comprehensive search of the solution space in pursuit of the global optimum. Accordingly, in the early stage of iteration, SAO achieves this balance by introducing a dual-population mechanism, whereby the entire population is randomly divided into two subgroups of equal size.

4.2. Improved Snow Ablation Optimizer (ISAO)

4.2.1. Quantum-Chaotic Hybrid Initialization Strategy

Although SAO is simple to implement, it is prone to population clustering in high-dimensional search spaces, resulting in insufficient search coverage and a reduced probability of locating the global optimum. To address this, a quantum-chaotic hybrid initialization strategy is introduced. This strategy uses the quantum logistic map as the chaotic source, with the iterative formula:
q ( n + 1 ) = μ · q ( n ) · ( 1 − q ( n ) ) · ( 1 − 2 · q ( n ) 2 )
where q ( n ) ∈ ( 0 ,   1 ) is the quantum-state amplitude at step n; μ ∈ ( 3.57 , 4 ] is the chaos control parameter, set to μ = 4 in this paper to guarantee full chaotic behavior; and q ( n ) 2 is the squared modulus of the quantum-state probability amplitude, introduced so that the map possesses quantum ergodicity beyond that of the classical logistic map, enabling more uniform probability coverage of the [0, 1] interval. After generating the chaotic-sequence operator Q c h a o s from the above sequence, the initial population is mapped into the search space as follows:
x i = l b + Q c h a o s × u b − l b
where Q c h a o s denotes the operator generated by the quantum-chaotic map. This effectively increases population diversity and lays the foundation for subsequent global exploration.

4.2.2. Elite-Pool-Guided Dual-Population Update Mechanism

To strengthen the guidance of the search process, ISAO constructs an elite pool composed of the best individual Z b e s t , the second-best individual Z s e c o n d , the third-best individual Z t h i r d and the centroid Z c e n t r o i d of the top 50% of individuals. Building on this, a hierarchical dual-population update mechanism is designed, randomly dividing the population into an elite-guided subset N a and a dynamic-scaling subset N b .
The elite-guided subset ( N a ) corresponds to the SAO exploration-stage formula, while retaining its Brownian-motion framework and weighted-displacement structure, the fixed elite term E l i t e ( t ) is extended to a guidance term dynamically sampled from the elite pool E_p, enhancing the diversity of exploration directions.
Z i ( t + 1 ) = E p ( K ) + R B ⨂ r 1 × Z b e s t − Z i ( t ) + ( 1 − r 1 ) × ( Z c e n t r o i d − Z i ( t ) )
where E p ( K ) is the k-th elite individual randomly selected from the elite pool; r 1 ∈ [ − 1 ,   1 ] corresponds to θ 1 in the original formula; Z c e n t r o i d is the population centroid, corresponding to Z ¯ ( t ) and Z b e s t corresponds to G ( t ) . R B is a random perturbation vector following Brownian motion, with each element independently drawn from a standard normal distribution. ⨂ denotes the Hadamard product.
The dynamic-scaling subset ( N b ) corresponds to the SAO exploitation-stage formula, while retaining its snow-ablation coefficient M and weighted-displacement framework, a temperature factor T is introduced to dynamically adjust the update speed, enabling the algorithm to transition smoothly from early-stage global exploration to late-stage local exploitation:
Z i ( t + 1 ) = M × Z b e s t + R B ⨂ r 2 × Z b e s t − Z i ( t ) + 1 − r 2 × Z c e n t r o i d − Z i ( t )
  τ = e x p − t M a x i t e r
where M is the snow-ablation rate; r 2 ∈ [ − 1 ,   1 ] corresponds to θ 2 in the original formula; τ is the temperature factor; t is the current iteration number; and M a x i t e r is the maximum number of iterations. τ decreases nonlinearly from 1 to e − 1 ≈ 0.368 as iterations proceed, allowing the dynamic-scaling subset to maintain a larger update step in the early iterations for thorough exploration, and to gradually contract in later iterations to focus on local exploitation, thereby achieving a dynamic balance between exploration and exploitation.

4.2.3. Quantum Tunneling Effect Exploration Strategy

To break through local extrema in the search space, a quantum tunneling effect mechanism is introduced. This mechanism allows an individual to cross the current search region with a certain probability P t u n n e l and perform a large-scale jump.
P t u n n e l is dynamically adjusted as iterations proceed, being larger in the early stage to maintain search intensity and lower in the later stage to focus on exploitation:
P t u n n e l = 0.25 × 1 − t M a x i t e r 2 + 0.05
When the tunneling effect is triggered, an individual performs the following jump:
Z t u n n e l = X i + S j × u b − l b × r a n d − 0.5
S j is the jump intensity, which decreases dynamically with the iteration number, calculated as:
S j = S m a x · 1 − t M a x i t e r + S m i n
where S m a x = 0.2 is the maximum jump intensity; S m i n = 0.01 is the lower bound of the jump intensity; and t is the current iteration number. S j decreases linearly, providing relatively large jump steps in the early stage to help individuals cross local barriers and explore the search space, while gradually reducing the jump range in the later stage to minimize disturbance to promising solutions.   r a n d ∈ [ 0 ,   1 ] is a uniformly distributed random number. This strategy grants individuals a “tunneling” capability, effectively guiding the population to escape local optima in which it has become stagnant and substantially enhancing the algorithm’s global search capability.

4.2.4. Dynamic Lens Opposition-Based Learning

To further improve the convergence accuracy of the algorithm, dynamic lens opposition-based learning is introduced at the end of the iterative process. By generating an opposition-based solution near the current optimum and deciding, via a greedy selection strategy, whether to replace the current optimum with this solution, a refined search of the neighborhood of the optimal solution is achieved. A dynamic-scaling factor d K is introduced to control the search range of the opposition-based solution, allowing it to cover a larger region in the early stage of iteration and progressively narrow in the later stage.
d K = 10 4 × 1 − t M a x i t e r 2 + 1
Based on the principle of lens imaging, the opposition-based solution is calculated as:
Z b e s t ′ = u b + l b 2 + u b + l b 2 ∗ d K − Z b e s t d K
where d K decreases as the iteration number increases, allowing the opposition-based solution to gradually contract toward the neighborhood of Z b e s t in the later stage and achieve refined exploitation. If the fitness of Z b e s t ′ is better than that of Z b e s t , the current best solution is updated; otherwise, the original best solution is retained. This greedy selection ensures that the best-so-far fitness does not deteriorate during the update.

4.3. Performance Testing

This paper uses the CEC2022 benchmark test function set (F1–F12) to evaluate the performance of the proposed ISAO algorithm. The experimental parameters are set as follows: population size 30, maximum number of iterations 500, problem dimension 10, and 30 independent runs. The comparison algorithms include the Rime Ice Algorithm (RIME) [18], the Honey Badger Algorithm (HBA) [19], Artificial Rabbit Optimization (ARO) [20], and SAO. The statistical indicators include the minimum value (Min), standard deviation (Std), mean value (Avg), median (Median), and worst value (Worst).
Table 1 presents the statistical results of each algorithm on the CEC2022 benchmark functions F1–F12.
Taken together, the results in Table 1 and Figure 2 show that ISAO achieves substantial improvement over the original SAO while maintaining competitive optimization performance across the CEC2022 benchmark suite. A particularly notable improvement is observed on F1, where the average value decreases from 990.9372 for SAO to 301.3804 for ISAO, corresponding to a reduction of approximately 69.58%. ISAO achieves the best average and median results among all compared algorithms on six test functions, namely F2, F3, F4, F9, F10, and F12. On F3, ISAO reaches the theoretical optimum of 600 with a very small standard deviation, indicating accurate and stable optimization performance. The convergence curves in Figure 2 further show that ISAO generally achieves rapid improvement during the early iterations and maintains a stable convergence trend thereafter.
Nevertheless, the performance of ISAO remains dependent on the characteristics of individual benchmark functions. For F6, ARO achieves substantially better performance than ISAO, with average values of 2811.7 and 4232.7, respectively, and the difference is statistically significant. On F7, ARO also achieves a lower average value than ISAO, with a statistically significant difference. For F11, both ARO and HBA obtain lower average values than ISAO, with average values of 2685.8, 2707.8, and 2728.9 for ARO, HBA, and ISAO, respectively. However, the differences between ISAO and ARO and between ISAO and HBA are not statistically significant. These results suggest that different metaheuristic search mechanisms may exhibit different levels of adaptability to individual benchmark landscapes. In particular, the search strategies incorporated into ISAO improve the overall search capability of SAO, but their effectiveness may vary across functions with different modality, ruggedness, and local-optimum distributions. Therefore, the benchmark results demonstrate clear improvements of ISAO over the original SAO and strong competitiveness across a broad range of functions, rather than universal superiority on every test function.

4.4. Wilcoxon Rank-Sum Test

To scientifically assess whether the performance improvement of ISAO is statistically significant, this section adopts the Wilcoxon rank-sum test at a significance level of α = 0.05. The calculated p-values are given in Table 2 (p < 0.05 indicates a statistically significant difference between the two algorithms).
To further interpret these results, the following analysis considers the p-values in Table 2 jointly with the mean and median values in Table 1 and the optimization objective of each benchmark function. A p-value below 0.05 indicates a statistically significant difference between the two distributions, but does not by itself determine the direction of performance superiority. Therefore, the statistical results are interpreted jointly with the mean and median values in Table 1 and the optimization objective of each benchmark function.
Compared with RIME, ISAO exhibits statistically significant differences on 10 of the 12 test functions and achieves better performance on F2, F3, F4, F9, F10, F11, and F12, whereas RIME performs better on F1. Compared with HBA, ISAO achieves better performance on F1, F3, F4, F5, F9, F10, and F12, with statistically significant differences for these comparisons. Compared with ARO, ISAO performs better on F1, F3, F4, F5, F9, F10, and F12, whereas ARO performs better on F6 and F7 with statistically significant differences. For F11, ARO achieves lower mean and median values than ISAO, but the difference is not statistically significant. Compared with the original SAO, ISAO achieves better performance on F1, F2, F3, F9, F11, and F12, while SAO performs better on F8 with a statistically significant difference.
Overall, the Wilcoxon results indicate statistically significant differences between ISAO and the compared algorithms on a considerable number of benchmark functions. When considered together with the corresponding mean and median values, these results support the competitive performance of ISAO on a substantial subset of the benchmark functions. More importantly, the comparison with the original SAO confirms that the proposed improvement strategies can substantially enhance the search performance of SAO on multiple functions, while the results on F6, F7, and F11 demonstrate that the effectiveness of the optimizer remains problem-dependent. Thus, ISAO provides a competitive and improved search strategy rather than a universally dominant optimizer for all benchmark landscapes.

5. ISAO-VMD-BiTCN Combined Prediction Model

Building on the study of the ISAO algorithm and the VMD method, this paper establishes the ISAO-VMD-BiTCN combined prediction model. The model first applies VMD to decompose the original load data into multiple IMF components, then applies ISAO to optimize the BiTCN hyperparameters for each component, and finally aggregates and denormalizes the component-wise predictions to obtain the final forecast. The specific steps are as follows.
(1) Read the actual load data from the PJM electricity market and perform normalization preprocessing. The normalization expression is:
x ~ = x − x m i n x m a x − x m i n
where x ~ is the normalized data; x is the original data; and x m a x and x m i n are the maximum and minimum values of the data, respectively.
(2) Use VMD to decompose the normalized load sequence into K IMF components, reducing the nonlinearity and non-stationarity of the data and improving the predictability of each component.
(3) Construct a separate BiTCN prediction sub-model for each IMF component, and use ISAO to independently optimize the BiTCN network hyperparameters (number of convolution kernels NumFilters, convolution kernel size FilterSize, dropout rate, and learning rate LearnRate) for each component, obtaining the optimal hyperparameter combination for each.
(4) Train the BiTCN prediction model for each component using its optimal hyperparameters to obtain the predicted values for each IMF component. Superimpose the predicted values of all components and perform denormalization to obtain the final power load forecast.
The overall workflow of the proposed ISAO-VMD-BiTCN combined prediction model is shown in Figure 3.

6. Case Study Analysis

To validate the effectiveness of the ISAO-VMD-BiTCN model established in Section 5, this section conducts a case study using hourly power load data from the PJM electricity market from 1 July 2016 to 30 September 2016, comprising 2185 data points at a 1 h resolution. This summer period exhibits pronounced load fluctuations, substantial peak-valley differences, and strong nonlinearity, providing a challenging scenario for evaluating the forecasting accuracy and adaptability of the proposed model. All simulation experiments were carried out in the MATLAB 2024 environment.

6.1. Model Parameter Settings

Using the same VMD configuration adopted above, the normalized load data are decomposed into five IMF components. The resulting VMD decomposition results are shown in Figure 4. It should be noted that both the normalization and the VMD are applied to the entire load series prior to the subsequent train/test partitioning; because VMD jointly decomposes the whole signal, this introduces a mild form of information leakage from the test period into the preprocessing stage. The dataset is strictly partitioned into chronological sets: the first 80% is used for training and the remaining 20% for testing. During the optimization phase, 15% of the training data is used to calculate model fitness, ensuring the test set remains entirely unseen until final evaluation. The time-series sliding window length is set to 24 h with a prediction step of 1 h. The BiTCN hyperparameter search space is the number of convolution kernels NumFilters ∈ [16, 128], convolution kernel size FilterSize ∈ [2, 6], dropout rate ∈ [0, 0.4], and learning rate LearnRate ∈ [1 × 10−4, 5 × 10−3]. To improve the reliability of the comparative evaluation and assess the influence of stochasticity, all prediction models were independently evaluated using six random seeds. The results obtained from these independent runs were summarized using the mean. For all metaheuristic algorithms involved, the population size was set to 25, and the maximum number of iterations was set to 30.

6.2. Prediction Model Hyperparameter Optimization

The ISAO algorithm is used to independently optimize the BiTCN hyperparameters for each IMF component within the search space defined above. The final optimal hyperparameter configuration obtained for each component is shown in Table 3.
Analysis of Table 3 shows that the low-frequency components exhibit strong regularity, and ISAO selects a lower learning rate and a simpler network structure for them; the high-frequency components, in contrast, fluctuate drastically, and ISAO automatically assigns more convolution kernels and a higher dropout rate to suppress noise interference. This component-by-component, customized optimization strategy ensures—at the underlying architectural level—the overall prediction accuracy after subsequent superposition and reconstruction.

6.3. Data Comparison

To evaluate the forecasting performance of the proposed ISAO-VMD-BiTCN model, comparative experiments were conducted using the dataset described above, with LSTM, VMD-BiTCN, GWO-VMD-BiTCN, PSO-VMD-BiTCN, and SAO-VMD-BiTCN as comparison models.
To evaluate the performance of each prediction model, the evaluation is conducted from two complementary perspectives: predictive accuracy and statistical significance. Predictive accuracy is assessed using four quantitative metrics: Root Mean Square Error (RMSE), Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), and the Coefficient of Determination (R2). The experimental results obtained from six independent random seeds are summarized using the mean value to characterize the predictive performance of the model. Statistical significance is further assessed using the Diebold–Mariano (DM) test, which is performed separately for each random seed to examine whether the forecasting loss differences between ISAO-VMD-BiTCN and the competing models are statistically significant.
The specific results are reported in Table 4, while the corresponding DM test results are presented in Table 5. Furthermore, Figure 5 and Figure 6 display the error-metric comparisons and actual-versus-predicted scatter plots, respectively.
As shown in Table 4, ISAO-VMD-BiTCN achieves the best mean predictive performance among the five comparison models. Based on the mean values, compared with the other prediction models, ISAO-VMD-BiTCN reduces RMSE by 17.80%, 12.91%, 6.52%, 9.58%, and 5.98%, respectively; reduces MAE by 21.94%, 14.70%, 6.71%, 10.09%, and 6.55%, respectively; reduces MAPE by 19.94%, 13.19%, 7.87%, 9.75%, and 6.64%, respectively; and improves R2 by 3.81%, 2.58%, 1.13%, 1.78%, and 1.00%, respectively. These results indicate that the complete ISAO-VMD-BiTCN framework provides meaningful predictive improvements under the investigated experimental setting. However, the observed improvements cannot be attributed solely to the global search capability of ISAO, since differences in model structure, hyperparameter configuration, and training procedures may also contribute to the final forecasting performance. Figure 5 further visualizes the comparative performance, showing that ISAO-VMD-BiTCN has the lowest mean error levels and the highest mean R2 among the compared models.
Figure 6 presents the actual-versus-predicted scatter plots for all models. The prediction points of ISAO-VMD-BiTCN are more closely distributed around the reference line than those of the competing models, indicating a stronger agreement between predicted and actual load values.
The Diebold–Mariano (DM) test results further assess whether the forecasting differences between ISAO-VMD-BiTCN and the competing models are statistically significant across different random initializations. Statistically significant favorable differences for ISAO-VMD-BiTCN are observed for 50–83.3% of the evaluated seeds across the competing models. According to the error-difference definition, d t = L e o t h e r , t − L e I S A O , t . A positive DM statistic with p < 0.05 indicates that the forecasting error of ISAO-VMD-BiTCN is significantly lower than that of the corresponding competing model. Although the statistical significance and direction of the differences are not consistent across every random seed, the DM results provide additional statistical evidence of favorable forecasting performance across multiple random initializations.
It should be noted that the present study has three main scope limitations. First, this study evaluates only a single summer season, July–September 2016, from the PJM market. Without coverage of other seasons or multiple years, this case study does not establish cross-seasonal or multi-year generalization. Second, electricity load is also influenced by exogenous factors beyond historical load patterns: temperature and humidity affect heating and cooling demand, while weekday type, weekends, and holidays alter electricity consumption patterns due to differences in human activity and operational schedules; seasonal changes in these factors may therefore contribute to variations in electricity demand and provide additional information for load forecasting [21]. However, the present study adopts a univariate forecasting setting and uses only historical load observations, so these exogenous variables are not explicitly modeled; accordingly, the proposed model should be interpreted as a load-history-based forecasting approach validated within the investigated single-season setting, rather than as a general-purpose, cross-seasonal forecasting model. Third, normalization and VMD are applied to the entire load series prior to the chronological train/test split; because VMD jointly decomposes the whole signal via ADMM, this introduces a mild form of information leakage from the test period into the preprocessing stage. This leakage is confined to an unsupervised, label-free preprocessing step rather than the supervised training loss itself, but a strictly causal or online decomposition scheme applied separately to the training and test partitions would eliminate this concern and is left for future work.

7. Conclusions

To address the nonlinear and nonstationary characteristics of power load data and the difficulty of determining BiTCN network hyperparameters, this study develops the ISAO and applies it to the hyperparameter optimization of a VMD-BiTCN forecasting framework. The ISAO incorporates four coordinated improvement strategies: quantum-chaotic hybrid initialization, elite-pool-guided dual-population updating, the quantum tunneling effect, and dynamic lens opposition-based learning. VMD is used to decompose the original load sequence into multiple IMF components, while ISAO independently optimizes the BiTCN hyperparameters for each component. On the CEC2022 benchmark suite, ISAO achieves the best mean and median performance on six of the 12 test functions and shows competitive performance against RIME, HBA, ARO, and SAO, with statistically significant differences observed on a substantial proportion of functions according to the Wilcoxon rank-sum test. On the PJM hourly load dataset, ISAO-VMD-BiTCN achieves the lowest RMSE, MAE, and MAPE and the highest R2 among the compared models, with a 5.98% reduction in RMSE relative to SAO-VMD-BiTCN. The DM test based on six independent random seeds further indicates that ISAO-VMD-BiTCN exhibits statistically significant forecasting advantages under most random initializations, providing additional statistical support for its favorable forecasting performance, although the statistical significance and direction of the differences vary across random seeds.
These findings are naturally bounded by the single-season, univariate scope of the present evaluation, which motivates the following directions for future work. Future work will extend the evaluation to multiple seasons and years, incorporate relevant meteorological and calendar variables into a multivariate forecasting framework, and further investigate computational efficiency and real-time deployment.

Author Contributions

Conceptualization, H.L.; Methodology, H.L.; Software, H.L.; Validation, H.L.; Formal analysis, H.L.; Investigation, H.L.; Data curation, H.L.; Writing—original draft, H.L.; Writing—review & editing, F.W. and X.L.; Visualization, H.L.; Supervision, F.W. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The hourly load data used in this study are publicly accessible via PJM Interconnection’s Data Miner 2 platform at https://dataminer2.pjm.com (accessed on 13 February 2026)/(feed: Hourly Load: Metered). Access to and use of this third-party data are subject to PJM’s data usage and redistribution policies.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Liang, H.T.; Liu, H.J.; Li, J.; Wang, Y.; Guo, C.N. Survey on short-term load forecasting algorithm based on machine learning. Comput. Syst. Appl. 2022, 31, 25–35. (In Chinese) [Google Scholar] [CrossRef]
  2. Wang, Y.; Chen, Q.; Hong, T.; Zhang, C. Short-Term Residential Load Forecasting Based on LSTM Recurrent Neural Network. IEEE Trans. Smart Grid 2019, 10, 841–851. [Google Scholar] [CrossRef] [Scilit]
  3. Cui, Y.; Zhu, H.; Wang, Y.J.; Zhang, L.; Li, Y. Short-term power load forecasting method based on CNN-SAEDN-Res. Electr. Power Autom. Equip. 2024, 44, 164–170. (In Chinese) [Google Scholar] [CrossRef]
  4. Zhao, Z.; Xia, C.; Chi, L.; Chang, X.; Li, W.; Yang, T.; Zomaya, A.Y. Short-Term Load Forecasting Based on the Transformer Model. Information 2021, 12, 516. [Google Scholar] [CrossRef] [Scilit]
  5. Yang, H.P.; Yu, Y.; Wang, C.; Li, X.J.; Hu, Y.T.; Rao, C.C. Short-term load forecasting of power system based on VMD-CNN-BiGRU. Electr. Power 2022, 55, 71–76. (In Chinese) [Google Scholar] [CrossRef]
  6. Wang, Q.; Chen, Z.R.; Li, G.M.; Jing, Z.; Zhang, Z.; Wang, P.X.; Cui, Q. Research on a short-term load forecasting algorithm for distribution areas based on VMD and TCN. J. Harbin Univ. Sci. Technol. 2024, 29, 121–129. (In Chinese) [Google Scholar] [CrossRef]
  7. Bai, S.; Kolter, J.Z.; Koltun, V. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv 2018, arXiv:1803.01271. [Google Scholar]
  8. Sprangers, O.; Schelter, S.; de Rijke, M. Parameter-efficient deep probabilistic forecasting. Int. J. Forecast. 2023, 39, 332–345. [Google Scholar] [CrossRef] [Scilit]
  9. Huang, Y.; Feng, Q.; Han, F. Short-term power load forecasting in China: A Bi-SATCN neural network model based on VMD-SE. PLoS ONE 2024, 19, e0311194. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Fan, C.; Li, G.; Xiao, L.; Yi, L.; Nie, S. Short-term power load forecasting in city based on ISSA-BiTCN-LSTM. Cogn. Comput. 2025, 17, 39. [Google Scholar] [CrossRef] [Scilit]
  11. Guo, X.; Gao, Y.; Song, W.; Zeng, Y.; Shi, X. Short-term power load forecasting based on CEEMDAN-WT-VMD joint denoising and BiTCN-BiGRU-attention. Electronics 2025, 14, 1871. [Google Scholar] [CrossRef] [Scilit]
  12. Kennedy, J.; Eberhart, R. Particle swarm optimization. In Proceedings of the ICNN’95—International Conference on Neural Networks, Perth, WA, Australia, 27 November–1 December 1995; pp. 1942–1948. [Google Scholar]
  13. Mirjalili, S.; Mirjalili, S.M.; Lewis, A. Grey wolf optimizer. Adv. Eng. Softw. 2014, 69, 46–61. [Google Scholar] [CrossRef] [Scilit]
  14. Sun, X.L.; Li, S.X.; Wang, K.; Liu, Q.Q. Research on short-term load forecasting based on an improved sparrow search algorithm. Comput. Sci. Appl. 2021, 11, 2271–2279. (In Chinese) [Google Scholar] [CrossRef]
  15. Deng, L.; Liu, S. Snow ablation optimizer: A novel metaheuristic technique for numerical optimization and engineering design. Expert Syst. Appl. 2023, 225, 120069. [Google Scholar] [CrossRef] [Scilit]
  16. Xiao, Y.; Cui, H.; Hussien, A.G.; Hashim, F.A. MSAO: A multi-strategy boosted snow ablation optimizer for global optimization and real-world engineering applications. Adv. Eng. Inform. 2024, 61, 102464. [Google Scholar] [CrossRef] [Scilit]
  17. Li, W.; Chen, X.; Okere, H.C. MSAO-EDA: A modified snow ablation optimizer by hybridizing with estimation of distribution algorithm. Biomimetics 2024, 9, 603. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Su, H.; Zhao, D.; Heidari, A.A.; Liu, L.; Zhang, X.; Mafarja, M.; Chen, H. RIME: A physics-based optimization. Neurocomputing 2023, 532, 183–214. [Google Scholar] [CrossRef] [Scilit]
  19. Hashim, F.A.; Houssein, E.H.; Hussain, K.; Mabrouk, M.S.; Al-Atabany, W. Honey badger algorithm: New metaheuristic algorithm for solving optimization problems. Math. Comput. Simul. 2022, 192, 84–110. [Google Scholar] [CrossRef] [Scilit]
  20. Wang, L.; Cao, Q.; Zhang, Z.; Mirjalili, S.; Zhao, W. Artificial rabbits optimization: A new bio-inspired meta-heuristic algorithm for solving engineering optimization problems. Eng. Appl. Artif. Intell. 2022, 114, 105082. [Google Scholar] [CrossRef] [Scilit]
  21. Li, W.; Sui, W.; Cheng, L.; Ji, Y.; Guo, Y.; Zhu, J. Quantifying seasonal demand-side flexibility in residential air conditioning under diverse control strategies. Energy Build. 2026, 352, 116764. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Dilated convolution structure.
Figure 1. Dilated convolution structure.
Electronics 15 04393 g001
Figure 2. Algorithm convergence curves.
Figure 2. Algorithm convergence curves.
Electronics 15 04393 g002
Figure 3. Flowchart of the ISAO-VMD-BiTCN short-term power load forecasting model.
Figure 3. Flowchart of the ISAO-VMD-BiTCN short-term power load forecasting model.
Electronics 15 04393 g003
Figure 4. VMD results.
Figure 4. VMD results.
Electronics 15 04393 g004
Figure 5. Bar charts of each model’s indicators: (a) RMSE; (b) MAE; (c) MAPE; (d) R2.
Figure 5. Bar charts of each model’s indicators: (a) RMSE; (b) MAE; (c) MAPE; (d) R2.
Electronics 15 04393 g005aElectronics 15 04393 g005b
Figure 6. Scatter plots of actual-versus-predicted load values for each model: (a) LSTM; (b) VMD-BiTCN; (c) GWO-VMD-BiTCN; (d) PSO-VMD-BiTCN; (e) SAO-VMD-BiTCN; (f) ISAO-VMD-BiTCN.
Figure 6. Scatter plots of actual-versus-predicted load values for each model: (a) LSTM; (b) VMD-BiTCN; (c) GWO-VMD-BiTCN; (d) PSO-VMD-BiTCN; (e) SAO-VMD-BiTCN; (f) ISAO-VMD-BiTCN.
Electronics 15 04393 g006aElectronics 15 04393 g006b
Table 1. Statistical comparison of CEC2022 benchmark function results.
Table 1. Statistical comparison of CEC2022 benchmark function results.
FuncMetricISAORIMEHBAAROSAO
Min300.0003300.0908300.0002304.7959300.0260
Std3.75480.70920.1252191.63721.5922 × 103
F1Avg301.3804300.9447300.0452492.1757990.9372
Median300.1449300.6781300.0086451.1535567.3950
Worst319.9029303.2662300.6727988.42048.7853 × 103
Min400.0994400.0331400.0010400.0008400.0061
Std1.950715.797322.877717.42183.2856
F2Avg404.9227410.3404415.0236407.6390406.3220
Median404.7841407.5445408.9161400.7275408.9161
Worst407.5227470.2210475.6756471.1503408.9161
Min600.0000600.0589600.0043600.0013600.0000
Std7.7407 × 10−50.28081.64820.56680.0352
F3Avg600.0000600.2744600.9259600.2105600.0090
Median600.0000600.2258600.1995600.0280600.0000
Worst600.0004601.5666607.4770602.8726600.1734
Min802.9966807.9627806.9647803.9799801.9899
Std6.852513.19786.67976.26188.2546
F4Avg811.0183828.0266818.2610815.9582813.6012
Median809.5156823.2814818.4067815.0092812.4370
Worst828.7462861.6931830.8437828.8538844.7360
Min900.0000900.0062900.0000900.0001900.0000
Std1.03816.122838.261121.12730.1660
F5Avg900.4601902.6485926.1769914.8638900.0693
Median900.0000900.7311912.9602904.8294900.0000
Worst904.0415933.52651.0718 × 103985.3316900.6334
Min1.9620 × 1031.8378 × 1031.8491 × 1031.8124 × 1031.8188 × 103
Std2.1894 × 1032.0025 × 1032.1889 × 1031.5152 × 1032.0880 × 103
F6Avg4.2327 × 1034.1825 × 1034.1042 × 1032.8117 × 1033.9603 × 103
Median3.9154 × 1033.8159 × 1033.2090 × 1032.1365 × 1033.4378 × 103
Worst1.2590 × 1038.1163 × 1038.2302 × 1036.7019 × 1037.6322 × 103
Min2.0002 × 1032.0004 × 1032.0020 × 1032.0000 × 1032.0010 × 103
Std9.06418.44905.91138.939223.8749
F7Avg2.0163 × 1032.0167 × 1032.0231 × 1032.0149 × 1032.0234 × 103
Median2.0212 × 1032.0211 × 1032.0226 × 1032.0201 × 1032.0216 × 103
Worst2.0278 × 1032.0243 × 1032.0369 × 1032.0216 × 1032.1443 × 103
Min2.2033 × 1032.2006 × 1032.2010 × 1032.2009 × 1032.2007 × 103
Std7.23936.722243.84776.522523.2900
F8Avg2.2237 × 1032.2187 × 1032.2347 × 1032.2182 × 1032.2228 × 103
Median2.2270 × 1032.2212 × 1032.2226 × 1032.2207 × 1032.2205 × 103
Worst2.2310 × 1032.2222 × 1032.4331 × 1032.2217 × 1032.3428 × 103
Min2.4856 × 1032.5293 × 1032.5293 × 1032.5293 × 1032.5293 × 103
Std2.876926.825626.675826.82390
F9Avg2.4874 × 1032.5342 × 1032.5352 × 1032.5342 × 1032.5293 × 103
Median2.4859 × 1032.5293 × 1032.5298 × 1032.5293 × 1032.5293 × 103
Worst2.4956 × 1032.6762 × 1032.6762 × 1032.6762 × 1032.5293 × 103
Min2.5002 × 1032.4436 × 1032.5003 × 1032.5002 × 1032.5002 × 103
Std46.380565.0293155.892049.227156.3875
F10Avg2.5231 × 1032.5505 × 1032.5988 × 1032.5272 × 1032.5425 × 103
Median2.5004 × 1032.5006 × 1032.6125 × 1032.5005 × 1032.5005 × 103
Worst2.6253 × 1032.6475 × 1033.3686 × 1032.6233 × 1032.6252 × 103
Min2.6000 × 1032.6249 × 1032.6000 × 1032.6000 × 1032.6000 × 103
Std131.279794.6216166.5669129.3503126.5838
F11Avg2.7289 × 1032.9059 × 1032.7078 × 1032.6858 × 1032.8001 × 103
Median2.7504 × 1032.9222 × 1032.6000 × 1032.6006 × 1032.9000 × 103
Worst3.0004 × 1033.1715 × 1033.1835 × 1032.9122 × 1032.9000 × 103
Min2.8476 × 1032.8614 × 1032.8605 × 1032.8643 × 1032.8614 × 103
Std4.05113.253429.81241.93504.1797
F12Avg2.8502 × 1032.8667 × 1032.8885 × 1032.8673 × 1032.8658 × 103
Median2.8487 × 1032.8663 × 1032.8831 × 1032.8671 × 1032.8649 × 103
Worst2.8651 × 103 2.8786 × 103 3.0004 × 1032.8728 × 1032.8845 × 103
Table 2. Wilcoxon rank-sum test p-values.
Table 2. Wilcoxon rank-sum test p-values.
FuncRIMEHBAAROSAO
F10.00147.2208 × 10−66.6955 × 10−111.4294 × 10−8
F20.00350.00580.12600.0114
F32.5158 × 10−112.5158 × 10−112.5158 × 10−110.0032
F45.0922 × 10−81.7836 × 10−40.00460.2339
F55.2943 × 10−62.8725 × 10−98.4372 × 10−81.0000
F60.99410.66274.2175 × 10−50.3632
F70.44640.00300.03780.2581
F88.1465 × 10−50.21707.1988 × 10−54.7138 × 10−4
F93.0199 × 10−112.5562 × 10−113.0199 × 10−111.2118 × 10−12
F100.00121.5846 × 10−40.00350.0850
F112.0035 × 10−80.09620.51050.0132
F127.3891 × 10−116.0290 × 10−114.9752 × 10−111.9523 × 10−10
Table 3. Optimal hyperparameter configuration for each IMF component.
Table 3. Optimal hyperparameter configuration for each IMF component.
ComponentNumFiltersFilterSizeDropout RateInitial Learning RateValidation-Set RMSE
IMF17620.00001.00 × 10−40.0086
IMF24120.01651.85 × 10−40.0241
IMF35230.05131.28 × 10−30.0231
IMF48050.23253.93 × 10−30.0221
IMF512830.40002.88 × 10−30.0106
Table 4. Comparison of prediction indicators across models.
Table 4. Comparison of prediction indicators across models.
ModelRMSE/MWMAE/MWMAPE/%R2
LSTM2679.142245.85 6.2330.8932
VMD-BiTCN2528.682055.165.7480.9039
GWO-VMD-BiTCN2355.841879.025.4160.9168
PSO-VMD-BiTCN2435.591949.815.5290.9110
SAO-VMD-BiTCN2342.231875.945.3450.9180
ISAO-VMD-BiTCN2202.141752.994.9900.9272
Table 5. Diebold–Mariano test statistics and p-values.
Table 5. Diebold–Mariano test statistics and p-values.
ModelStatisticS1S2S3S4S5S6
LSTMDM3.56035.24041.25793.95852.47152.0651
P3.7042 × 10−41.6033 × 10−72.0843 × 10−17.5418 × 10−50.01350.0389
VMD-BiTCNDM6.2112−1.52104.26643.16884.18022.8940
P5.2651 × 10−100.12831.9868 × 10−51.5307 × 10−32.9124 × 10−53.8043 × 10−3
GWO-VMD-BiTCNDM−1.32104.58434.65550.52262.4195−1.4120
P0.18654.5543 × 10−63.2323 × 10−60.60130.01550.1579
PSO-VMD-BiTCNDM3.86874.1273−1.7179−1.38501.15826.3570
P1.0935 × 10−43.6698 × 10−50.85820.16610.24682.0560 × 10−10
SAO-VMD-BiTCNDM6.2156−1.0455−1.21403.32912.2418−0.5695
P5.1128 × 10−100.29580.22488.7138 × 10−40.02490.5690
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Liu, H.; Wang, F.; Li, X. Short-Term Electric Power Load Forecasting Based on ISAO-VMD-BiTCN. Electronics 2026, 15, 4393. https://doi.org/10.3390/electronics15194393

AMA Style

Liu H, Wang F, Li X. Short-Term Electric Power Load Forecasting Based on ISAO-VMD-BiTCN. Electronics. 2026; 15(19):4393. https://doi.org/10.3390/electronics15194393

Chicago/Turabian Style

Liu, Hanchi, Fang Wang, and Xinyi Li. 2026. "Short-Term Electric Power Load Forecasting Based on ISAO-VMD-BiTCN" Electronics 15, no. 19: 4393. https://doi.org/10.3390/electronics15194393

APA Style

Liu, H., Wang, F., & Li, X. (2026). Short-Term Electric Power Load Forecasting Based on ISAO-VMD-BiTCN. Electronics, 15(19), 4393. https://doi.org/10.3390/electronics15194393

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop