1. Introduction
With the global energy transition and the widespread adoption of renewable energy, the International Energy Agency (IEA) has set a target in its World Energy Outlook for global installed photovoltaic (PV) capacity to reach approximately 16,000 gigawatts (GW) by 2050. By the end of 2023, China’s PV power generation capacity had reached 610 million kilowatts, establishing PV power generation as a cornerstone of renewable energy [
1]. However, PV power generation systems are highly susceptible to environmental factors such as solar irradiance, exhibiting significant randomness, intermittency, uncertainty, and volatility [
2,
3]. If integrated into the grid on a large scale directly, these characteristics can disrupt scheduling plans and compromise the safe operation of the power grid [
4]. Therefore, accurate forecasting of PV power output is crucial for grid dispatchers to reasonably adjust generation schedules and maintain the stable operation of the power system [
5].
At present, photovoltaic (PV) power forecasting can be primarily categorized into physical methods, direct forecasting methods, and artificial intelligence (AI) methods [
6,
7]. Physical methods establish models based on photoelectric theory to predict power output, an approach that does not rely on historical PV output data. However, due to the high complexity of factors influencing solar irradiance, current theoretical models are difficult to apply directly to practical scenarios [
8]. Direct forecasting methods establish a mapping relationship between historical and future PV output data by analyzing characteristics such as the periodicity and trends of PV power output curves [
9]. With the rise in AI methods, data-driven PV power forecasting approaches have been widely applied [
10]. Zhang et al. [
11] propose a multi-quantile photovoltaic power forecasting model based on a Bidirectional Temporal Convolutional Network. By using the K-means algorithm for same-day weather clustering and integrating quantile regression, the model unifies point and interval forecasting within a single framework, thereby enabling reliable uncertainty quantification and improving prediction accuracy. Deng et al. [
12] employ a Long Short-Term Memory (LSTM) model to conduct long-term power generation forecasting for high-rise Building-Integrated Photovoltaic (BIPV) systems, emphasizing the importance of the LSTM model in improving the reliability and efficiency of long-term solar forecasting. Literature [
13] adopts a Convolutional Graph Neural Network to mine the spatial correlation features among distributed photovoltaic nodes, highlighting the advantages of this deep learning architecture in enhancing the universality of model training and prediction accuracy.
Since single forecasting models struggle to cope with the variability of weather conditions, leading to a decline in their generalization ability and prediction accuracy, researchers have proposed solutions such as ensemble forecasting models. Nadour et al. [
14] combined a convolutional neural network with a bidirectional long short-term memory network, leveraging the spatial feature extraction capability of CNN and the bidirectional temporal modeling advantage of BiLSTM to achieve high-accuracy photovoltaic power forecasting under the extreme Saharan climate. Zhan et al. [
15] constructed a multi-module collaborative hybrid forecasting framework that organically integrates clustering analysis, similar-day matching, variational mode decomposition, a temporal convolutional network-long short-term memory-attention mechanism, and physical information constraints, while adaptively tuning key parameters using the rime optimization algorithm, demonstrating excellent generalization ability across different climatic regions. Wang et al. [
16] also adopted a hybrid modeling approach, cascading improved complete ensemble empirical mode decomposition with adaptive noise and variational mode decomposition as a preprocessing module, then building the core prediction module with a temporal convolutional network, a bidirectional gated recurrent unit, and an attention mechanism, and finally introducing an improved crested porcupine optimizer for global hyperparameter optimization, forming a complete hybrid chain from data decomposition to prediction to parameter optimization.
When performing PV power prediction, the input features directly affect the prediction results. Furthermore, PV forecasting often involves variables such as meteorological data, so the input features are multidimensional, which not only affects the power output but also increases the model’s complexity. On the one hand, feature decomposition techniques can be used. Sun et al. [
17] utilized the discrete wavelet transform to decompose complex raw time series into multi-band components of varying frequencies. Building upon this, they cascaded continuous sampling and interval down-sampling mechanisms, which allowed them to accurately extract both short-term local fluctuations and long-term evolutionary trends within each frequency band. Meanwhile, Wang et al. [
18] applied variational mode decomposition as their primary data-smoothing technique to process clustered photovoltaic sequences. By adaptively breaking down the non-stationary data into multiple discrete sub-modes, they effectively circumvented the endpoint effects that frequently occur in traditional decomposition methods. Building on this, Liang et al. [
19] proposed a secondary deep feature decomposition architecture. After employing VMD in the initial layer to extract several intrinsic mode sequences alongside a single residual term, they specifically targeted this residual with complete ensemble empirical mode decomposition with adaptive noise to perform a secondary decomposition. This stepped approach maximized mitigation of the mode-mixing drawbacks inherent in single-decomposition methods. These methods improve prediction accuracy through feature decomposition or reconstruction; however, when handling large-scale, high-dimensional data, the decomposed features may still retain noise or unstable components, ultimately affecting the model’s accuracy and generalization. While the decomposition methods in the aforementioned literature effectively mitigate data non-stationarity, they typically input all decomposed sub-sequences directly into the forecasting model. This direct input approach inevitably leads to a surge in computational complexity and the accumulation of prediction errors from each component.
On the other hand, deep learning can be used to fully explore the features of meteorological factors and photovoltaic power generation, and effectively extract the deep information of the features. Han et al. [
20] utilized one-dimensional convolution operations with diverse activation functions to specifically extract various complex dynamic features in electricity load sequences, such as linear monotonic, nonlinear gradual, nonlinear abrupt, and periodic fluctuations. Sheng et al. [
21] proposed a feature filtering and adaptive reconstruction module that performs multi-scale, parallel feature extraction by combining spatial and channel reconstruction convolutions with receptive-field attention convolutions, effectively reducing feature redundancy in both spatial and channel dimensions. Liu et al. [
22] utilized convolution kernels to conduct local spatial scanning on long-term time-series three-dimensional matrices containing multi-dimensional meteorological factors, extracting high-dimensional spatial features at both shallow and deep levels, and employed pooling layers for dimensionality reduction to decrease computational parameters. However, these methods fail to account for information at different data frequencies, and multiple feature-extraction and dimensionality-reduction operations may lead to high computational complexity. Although the aforementioned frameworks are adept at extracting spatial–temporal features, they rely on static, single-scale convolution kernels and fail to dynamically separate macroscopic meteorological trends from transient power fluctuations across different frequency domains, resulting in severely restricted receptive fields.
To address the limited forecasting accuracy stemming from limitations of existing photovoltaic models, such as error accumulation from over-decomposition, restricted receptive fields in convolutional networks, and dimensional confounding among covariates, this paper proposes a robust short-term forecasting framework. First, unlike conventional direct decomposition prediction frameworks, we introduce fuzzy entropy to quantify and reconstruct CEEMDAN-derived IMFs. This fundamentally prevents the computational surge and error accumulation caused by over-decomposition, providing highly analyzable and stationary sub-sequences for subsequent deep learning layers. Second, overcoming the limited receptive fields of standard CNNs, we introduce the WTConv module. By seamlessly embedding discrete wavelet transforms into convolution operations, it dynamically extracts frequency-band features, thereby mathematically isolating low-frequency climate trends from high-frequency chaotic noise at minimal parameter cost. Finally, to resolve the dimensional confounding inherent in standard forecasting networks, we adopt the iTransformer. Its inverted encoding structure treats the entire temporal evolution of each meteorological factor as an independent token, accurately capturing pure, long-term cross-variable physical dependencies without interference. Extensive experiments, including cross-regional generalization validation and statistical significance tests, confirm that the proposed framework achieves highly superior forecasting results in terms of robustness, accuracy, and stability.
3. Case Analysis
3.1. Experimental Data and Platform
The experimental data used in this study were collected from a photovoltaic power station in Ningxia, China, covering January, April, July, and October of 2019, which represent the four seasons of spring, summer, autumn, and winter, respectively, in order to validate the effectiveness and generalization capability of the proposed model. The dataset exhibits typical temporal characteristics of photovoltaic power generation, showing a clear diurnal cycle. Specifically, the power output is nearly zero during nighttime, while significant fluctuations occur during daytime. The maximum output power is approximately 128 MW, and the standard deviation is 34.59, indicating strong variability and pronounced non-stationarity in the data. Given the intermittency of photovoltaic power generation, data from 8:00 to 22:00 were selected with a 15 min sampling interval. The influencing factors include 16 variables, such as air temperature, air pressure, humidity, dew point temperature, and solar radiation.
To preserve the chronological integrity of the time series, the data for each selected month are sequentially partitioned into training, validation, and test sets in a 7:1:2 ratio. Furthermore, data preprocessing procedures, including normalization and feature selection, are applied exclusively to the training set to prevent potential data leakage. During model training, the Mean Squared Error (MSE) is used as the loss function to quantify the deviation between the forecast and actual photovoltaic power values, and the Adam optimizer is used to adaptively update the network weights. The input sequence length is set to 56-time steps, enabling the model to capture a complete diurnal fluctuation cycle. Concurrently, the forecasting horizon is set to 1 time step, enabling 15 min ahead short-term predictions.
The experiments were conducted on an experimental platform with an Intel Core i9-11900K processor, NVIDIA GeForce RTX 3080 GPU, and 32 GB DDR4 memory.
After constructing the dataset and analyzing its fundamental characteristics, it is necessary to appropriately configure the key hyperparameters of the proposed model to fully exploit its advantages in feature extraction and multivariate modeling. The selection of hyperparameters directly affects the model’s feature representation capability, convergence speed, and prediction accuracy. Consequently, to identify the optimal configuration for the proposed model, the final parameter settings were determined via a systematic grid search method, accounting for the data characteristics and model complexity. The main hyperparameter configurations used in this study are listed in
Table 1.
3.2. Data Preprocessing and Evaluation Metrics
Outliers can occur during the collection of PV power data, and such data can affect the predictability of the model and reduce the prediction accuracy. Outliers are detected using the 3Sigma rule, which identifies those data points that exceed the range of mean ±3 standard deviations by calculating the mean and standard deviation of the data. Detected outliers are replaced with null values, and then cubic spline interpolation is used to fill these null values.
In order to measure the prediction performance of the model, Root Mean Squared Error (RMSE), Mean Absolute Error (MAE), Mean Absolute Percentage Error (MAPE), and goodness-of-fit were used to analyze the prediction results. The four evaluation metrics were analyzed as follows. The formulas for the four evaluation indexes are as follows:
where
are the number of predicted samples;
is the
actual value of the first sample;
is the
predicted value of the first sample;
is the average value of the actual sample.
3.3. Sequence Decomposition and Reconstruction
Considering the strong periodicity and uncertainty of the PV power data, CEEMDAN is initially used to decompose the raw data into multiple IMF components. However, establishing independent prediction modules for all generated sub-sequences directly would inevitably lead to a surge in computational complexity and severe error accumulation due to over-decomposition. Therefore, Fuzzy Entropy is introduced to quantify the dynamic complexity and randomness of each individual IMF.
The results of the fuzzy entropy value of PV sequences in January, April, July, and October 2019 are shown in
Figure 5. Scientifically, the FE values exhibit a distinctly hierarchical distribution that aligns with the physical characteristics of the PV signals. The initial high-frequency IMFs naturally exhibit the highest FE values, indicating intense stochasticity driven by sudden changes in cloud cover or sensor noise. The intermediate IMFs exhibit moderate FE values, consistent with the macroscopic diurnal cycles. Finally, the low-frequency residual components exhibit FE values approaching zero, reflecting smooth, long-term evolutionary trends.
Rather than merely describing the data, this entropy-based complexity assessment provides a rigorous mathematical criterion for feature compression. By superimposing IMFs with similar entropy magnitudes, subsequences with isomorphic dynamical characteristics are adaptively merged into distinct reconstructed components. This reconstruction strategy effectively mitigates the severe non-stationarity of the original data and significantly alleviates the “curse of dimensionality,” yielding highly analyzable, stationary sub-sequences without discarding any critical frequency-domain information for the subsequent modeling phase.
3.4. Correlation Analysis
Input variables directly impact the model’s predictive ability and are key elements of a PV power prediction model. The input variables should cover sufficient features, and too many meteorological factors may lead to redundant information and reduce model accuracy. Therefore, Pearson correlation analysis was used to determine the linear relationship between PV power and meteorological factors; however, many non-linear, indirect, and deeper relationships exist between the two, so the maximum mutual information coefficient (MIC) was then used for correlation analysis. The threshold value for the correlation coefficient is set to 0.3, and the correlation results for each reconstructed component with meteorological features in January are shown in
Table 2.
The table further reveals significant differences in correlations between reconstructed components and meteorological features. Specifically, for the component IMF1, strong positive correlations are observed with radiation-related variables, while a significant negative correlation is found with the solar zenith angle. This indicates that IMF1 primarily reflects rapid fluctuations in photovoltaic power driven by variations in solar irradiance. For the component IMF2, the overall correlation pattern is similar to that of IMF1, although the correlation coefficients are slightly lower. This suggests that IMF2 retains the dominant influence of radiation factors while also being affected to some extent by other meteorological variables, reflecting relatively smoother variation trends. In addition, the MIC values are generally higher than the corresponding Pearson coefficients, indicating that IMF2 may contain more nonlinear relationships. In contrast, the component IMF3 exhibits a distinctly different correlation pattern. It shows relatively strong correlations with opacity and snow depth, while its linear correlations with radiation variables are significantly weaker. However, the MIC values remain relatively high, suggesting that this component has more implicit nonlinear and indirect influences.
Overall, the differences in correlations among the reconstructed components validate the need to decompose and reconstruct the original sequence. On the one hand, this process transforms complex non-stationary series into sub-sequences with distinct physical meanings; on the other hand, the combined use of Pearson and MIC analysis enables the capture of both linear and nonlinear relationships, thereby providing a more reliable basis for feature selection and multi-scale modeling.
In order to eliminate the differences in physical dimensions, it is necessary to normalize the input variables with the following formula:
where
are raw data;
is the minimum value in the data set;
is the maximum value in the data set;
is normalized value.
3.5. Ablation Experiment
In order to verify the model’s validity, the proposed model is compared with the CEEMDAN-iTransformer model, with the elimination of the feature extraction module to verify the validity of WTConv, and with the WTConv-iTransformer model without CEEMDAN to verify the validity of modal decomposition, and then with the CEEMDAN-WTConv-Transformer model to verify the performance of iTransformer compared to Transformer.
Table 3 compares evaluation metrics for each model across the ablation experiments conducted in January, April, July, and October. The tables show that the proposed model achieves the best performance across all four seasons for MAE, MAPE, RMSE, and R
2. Taking January as an example for detailed analysis, the MAE, MAPE, and RMSE of the proposed model are reduced by 9.03%, 8.35%, and 26.56%, respectively, while the R
2 is improved by 0.31% compared to the CEEMDAN-WTConv-Transformer model. This is because standard Transformers tokenize all variables within a single time step into a unified vector, which blurs the physical independence of different meteorological factors. In contrast, the iTransformer module embeds each variable as an independent token, making it better suited to handling multivariate time-series data. Compared to the WTConv-iTransformer model, the proposed model’s MAE, MAPE, and RMSE are reduced by 16.08%, 33.46%, and 30.5%, respectively, with an R
2 improvement of 1.65%. This demonstrates that the entropy-based CEEMDAN decomposition, combined with the data reconstruction module, effectively mitigates the highly non-stationary nature of the raw photovoltaic power sequences, thereby substantially enhancing the fitting capability. For the CEEMDAN-iTransformer model, MAE, MAPE, and RMSE are reduced by 9.80%, 21.92%, and 5.31%, respectively, and R
2 improves by 0.82%. This demonstrates that the WTConv module can expand the receptive field of feature decomposition and acts as a learnable secondary decomposition mechanism. It mathematically isolates meteorological trends in the low-frequency space while sensitively capturing transient power drops in the high-frequency space, thereby continuously providing highly refined multi-scale feature maps for the prediction phase.
A comprehensive analysis of the tables shows that while the improved Transformer contributes to performance gains, its full potential may not be realized within a limited monthly dataset. Conversely, the combination of CEEMDAN and WTConv produces a quadratic decomposition effect that significantly suppresses noise interference in non-stationary sequences, leading to exceptionally high R2 values of 0.995 in July and 0.996 in October, thereby substantially boosting the model’s prediction accuracy and fitting stability.
In order to clearly observe the prediction performance of the proposed model and the comparison model, the prediction curves of a certain period in January, April, July, and October are selected, respectively. As shown in
Figure 6, the proposed model’s prediction curve is closest to the real value among all models, with the smallest fluctuation range, indicating good prediction accuracy and fitting ability.
3.6. Performance Analysis of the Loss Function
To evaluate the learning performance of the proposed model on the reconstructed components, this paper plots the loss function convergence curves of the reconstructed components during the training process. Taking January as an example, a brief analysis is conducted on the convergence curves of the three reconstructed components: The IMF1 component corresponds to the reconstructed sub-sequence with the lowest complexity; its loss converges at the fastest rate and achieves the lowest final stabilized loss value. This indicates that the model can highly effectively capture and fit the stationary variation patterns inherent in this component. The loss sequences of the IMF2 and IMF3 components also maintain a highly stable downward trend. Due to their greater fitting difficulty, their final stabilized loss values are slightly higher than those of the IMF1 component, which objectively reflects the model’s strong robustness and feature extraction capabilities.
As shown in the overall loss convergence analysis results in
Figure 7, the loss values of all components decrease rapidly as the number of training epochs increases, and they generally stabilize after approximately 20 epochs. This phenomenon clearly demonstrates the excellent convergence of the model and fully proves that the proposed model possesses strong, stable learning capability.
3.7. Significance Test
In summary, BiLSTM and CNN-LSTM are constrained by their inherent network structures and fall short in adequately extracting multi-scale time–frequency features from PV power data. Although Transformer and Informer have made progress in long-sequence modeling, they still have limitations in multivariate independent modeling and the effective utilization of frequency information. The proposed model addresses these shortcomings by applying CEEMDAN decomposition to reduce sequence non-stationarity, employing WTConv to extract multi-scale time–frequency features, and leveraging iTransformer for independent multivariate modeling, thereby achieving optimal prediction performance.
To verify the statistical significance of the performance improvement achieved by the proposed model (denoted as model 1) over its ablated variants, the Diebold–Mariano (DM) test is employed to compare the prediction error sequences of different models. The comparison models include CEEMDAN-WTConv-Transformer (denoted as model 2), CEEMDAN-iTransformer (denoted as model 3), and WTConv-iTransformer (denoted as model 4). Pairwise DM tests are conducted on the absolute prediction error sequences for four months: April, July, October, and January.
Table 4 shows that for all seasons, the absolute DM statistics between model1 and model 2, model1 and model 3, and model1 and model 4 all exceed the critical values, with corresponding
p-values less than 0.001, thereby rejecting the null hypothesis of equal prediction accuracy. This indicates that the contributions of CEEMDAN decomposition and reconstruction, WTConv multi-scale feature extraction, and iTransformer independent variable modeling are all statistically significant and not due to random fluctuations. Therefore, the proposed model outperforms all ablated variants in a statistically significant manner, confirming the effectiveness and robustness of its architectural design.
3.8. Comparative Experiments
To thoroughly validate the effectiveness of the CEEMDAN-WTConv-iTransformer model proposed in this paper for short-term photovoltaic power forecasting, this section conducts comparative experiments using four classic and widely adopted models in the field of time series forecasting. The experimental results are presented in
Table 5, while
Figure 8 visually demonstrates the superiority of the proposed method.
A detailed analysis is provided using the January data as an example: The CNN-LSTM model records an R2 of 0.941, with MAE, MAPE, and RMSE values of 2.512 MW, 4.512%, and 3.956 MW, respectively. Relative to the proposed model, its MAE and RMSE are 37.12% and 45.92% higher, respectively. Although CNN-LSTM extracts local features through convolutional layers, the use of a single-scale convolution kernel restricts its ability to simultaneously capture fluctuation patterns across different time scales. Moreover, the absence of a dynamic weighting mechanism for input features reduces its adaptability when correlations with meteorological variables change over time. The CNN-Transformer model achieves an R2 of 0.969, with MAE, MAPE, and RMSE values of 2.111 MW, 3.908%, and 3.257 MW, respectively. Compared to the proposed model, its MAE and RMSE are 15.23% and 20.14% higher, respectively. While introducing the Transformer architecture improves the modeling of temporal dependencies compared to CNN-LSTM, it still utilizes standard convolution for feature extraction, which limits the receptive field and struggles with complex multi-scale frequency components. Furthermore, it retains the conventional tokenization strategy that blends multivariate variables. The Informer model achieves an R2 of 0.975, with MAE, MAPE, and RMSE values of 2.015 MW, 3.612%, and 3.045 MW, respectively. In comparison with the proposed model, its MAE and RMSE are 9.99% and 12.32% higher, respectively. While Informer improves the efficiency of long-sequence forecasting on the basis of Transformer, it still relies on the same tokenization strategy, which imposes constraints when handling complex interactions among multiple variables. It also lacks a dedicated module designed to extract multi-scale frequency features from PV power sequences. The standalone iTransformer model yields an R2 of 0.980, with MAE, MAPE, and RMSE values of 1.977 MW, 3.416%, and 2.902 MW, respectively. Compared to the proposed model, its MAE and RMSE are 7.91% and 7.05% higher, respectively. Although the iTransformer successfully addresses the issue of multivariate confounding by treating the entire temporal evolution of each variable as an independent token, its direct application to raw, highly volatile PV sequences without prior noise reduction and multi-scale frequency decomposition limits its ability to reach optimal forecasting accuracy.
In summary, traditional deep learning models like CNN-LSTM are constrained by single-scale feature extraction, while advanced variants like CNN-Transformer and Informer face bottlenecks in multivariate independent modeling and frequency information utilization. Even the advanced iTransformer falls short of its full potential due to the lack of a dedicated data decomposition module. The proposed model successfully addresses these shortcomings by applying CEEMDAN decomposition to reduce sequence non-stationarity, employing WTConv to extract multi-scale time–frequency features, and leveraging iTransformer for independent multivariate modeling, thereby achieving optimal prediction performance.
3.9. Generalization Verification
To evaluate the generalization capability of the proposed model beyond the specific climatic conditions and plant scale of the primary dataset, additional experiments were conducted on the Desert Knowledge Australia Solar Center (DKASC) photovoltaic dataset. This dataset is collected from a PV plant located in a different region, featuring distinct geographical characteristics and meteorological patterns. The variations in installed capacity and localized weather dynamics provide a rigorous test for the spatial robustness of the model.
Following the exact same data preprocessing, feature selection, and hyperparameter configuration protocols described above, the proposed framework and its ablation models were evaluated on this Australian PV dataset. The performance comparison results are summarized in
Table 6.
A comparison of the ablation experiment results reveals that the proposed model maintains a leading advantage on the DKASC dataset, achieving the highest R2 of 0.975, along with the lowest MAE, MAPE, and RMSE among all models at 3.143, 4.316, and 3.232, respectively. This once again proves that removing the data decomposition module, omitting the wavelet convolution, or employing the traditional Transformer architecture leads to a significant degradation in forecasting performance.
A horizontal comparison reveals that in this photovoltaic plant with completely different meteorological characteristics, the magnitude of the forecasting error remains approximately consistent with the results from the Ningxia plant. This strongly demonstrates that the proposed model performs exceptionally well across multiple datasets. Furthermore, because the climate dynamics in the Australian desert region are more drastic and high-frequency fluctuations occur more frequently, the model’s goodness of fit on this dataset experiences a slight decrease, while the relative errors exhibit an increase within a reasonable range. Such subtle metric variations perfectly align with the objective physical laws governing cross-regional climate characteristic transfers. Overall, the proposed model maintains high forecasting accuracy and demonstrates outstanding robustness and generalization capabilities.