4.3. Scenario Generation Performance
After determining reliable data sources, the next step is to verify the effectiveness of the two-stage dimensionality reduction and correlation modeling method. This section will present the performance of scenario generation, including the results of dimensionality reduction, correlation capture, and the accuracy of scenario reconstruction.
4.3.1. Dimension Reduction Effect Verification
To verify the effectiveness of the first-stage dimensionality reduction strategy (
Section 3.1), this study processed NASA’s historical wind and solar energy data, reducing the time dimension of each virtual station from 5475 days (365 × 15) to 225 days. It selected representative solar terms based on climatic characteristics: The representative solar terms for wind power (monsoon transition and the peak period of strong winds) are Fresh Green, First Frost, and Great Cold. The representative solar terms for solar power (with clear skies and high radiation before and after the rainy season) are Spring Equinox, Lesser Fullness, and White Dew. The retained typical daily scenarios are crucial for a complementarity analysis, as shown in
Figure 10 and
Figure 11. For detailed clustering results, please refer to
Appendix B (
Figure A3 and
Figure A4).
CSD and labels (in the upper right corner) represent the clustering tightness. The closer the value is to 0, the more concentrated the intra-class curve is and the higher the centroid representativeness. The analysis of the clustering results is as follows:
Wind power clustering results: It can be seen from
Figure 10 that Fresh Green, First Frost, and Great Cold have the typical features of high and large fluctuations in wind power during spring and winter, which is highly consistent with the monsoon climate characteristics of this basin. The clustering result C1 of wind power generation is significantly higher because the wind speed is too low or even close to 0 on some days during the solar term. These patterns are real but atypical operating states; in the subsequent Kernel PCA dimensionality reduction process, their corresponding features will mainly be distributed in the principal components with lower interpretation variances and are usually not selected for scenario generation. Except for C1, the sum of the CSD within each cluster is close to 0, indicating a good clustering effect.
Solar power clustering results:
Figure 11 shows the clustering results of the solar output. The peak data of the three typical solar terms, namely, the Spring Equinox, Lesser Fullness, and White Dew, are relatively high and concentrated, which conforms to the natural laws before and after the rainy season. The total CSD within each cluster is basically maintained between 0 and 3. The closer the value is to 0, the more effectively the CSD can aggregate curves with similar shapes.
After dividing and clustering the original data by solar terms, the time dimension was significantly compressed, and 50 to 80 typical daily datasets were obtained for each solar term. This method verifies the effectiveness of the CSD-improved clustering algorithm in capturing fluctuation patterns and reduces the complexity of the original high-dimensional time series data. It provides high-correlation and low-dimensional input for the subsequent Kernel PCA. This process also divides some extreme events into independent clusters and converts them into higher-order principal components with a lower variance interpretation rate in the subsequent Kernel PCA analysis. When generating scenarios based on principal components (
Section 4.3.3), some extreme events will be covered.
4.3.2. Principal Component Loading and Correlation Verification
To analyze the nonlinear coupling characteristics proposed in this study (
Section 3.2.2), this section takes the three representative solar terms of the Spring Equinox, Lesser Heat, and White Dew to deeply analyze the physical significance of the principal component load and the collaborative mechanism of the hydro–wind–solar correlation model. The principal component load is shown in
Figure 12.
The subgraphs at the lower right corner of each chart contain the following: the vertical axis on the left represents the variance interpretation rate of each principal component (Blue bar chart), while the horizontal axis on the right shows the cumulative variance rate (Red dot line graph). Select the top 15 principal components with a cumulative variance exceeding 85% to ensure a balance between the research dimension and the correlation. The vertical axis of the main graph represents the contribution of each energy type to its principal component, which can be extended to understand the correlation among different energy types.
To analyze the principal component load, it is necessary to analyze the contributions of each principal component. In terms of contribution, the first three main components are of vital importance. ① The contribution value of PC1 in
Figure 12a shows that wind power dominates (with a contribution value of approximately 0.75), while the contribution value of solar power is much lower than that of runoff and wind power. Runoff, a sharp decrease in wind power and solar power, has become a positive correlation, representing a windless and cloudy weather pattern. ② The significant increase in PC2 runoff and the negative correlation with the decrease in wind and solar energy clearly indicate that rainy days are approaching and other weather patterns are captured. ③ PC3 reflects a situation where solar energy is severely suppressed, wind energy almost disappears, and runoff decreases, that is, the characteristics of cloudy weather.
An analysis of typical weather patterns is equally important. Based on the law in
Figure 12a, combined with the principal component loads in
Figure 12b,c, the complex nonlinear correlation characteristics presented by the hydro–wind–solar system can be summarized. The main pattern is as follows: ① The runoff volume is positively correlated with significantly weakened wind and solar energy. ② The increase in runoff volume is negatively correlated with the attenuation of wind and solar energy. ③ The sudden increase in wind force forms a confrontation with the suppression of hydropower and solar energy. ④ Solar energy has close to 0 growth, wind power is gradually disappearing, and the runoff volume continues to increase.
This section provides a physical interpretation basis for the Vine Copula model in
Section 3.3 by quantifying the nonlinear correlation of the hydro–wind–solar system. At the same time, the effectiveness and innovation of the method described in
Section 1.2 were verified, and it was proved that the Kernel PCA method does not have a black box problem, and its steps and mechanisms have clear physical interpretability.
4.3.3. Scenario Generation Result Analysis
In order to accurately describe the nonlinear dependency structure among principal components after dimensionality reduction and generate scenarios, this study adopts the Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC) criteria to screen the optimal Copula function and implements the simplified Vine Copula method for scenario generation. The average AIC/BIC value calculation results of each solar term are shown in
Table 3.
As shown in
Table 3, the systematic assessment based on the AIC/BIC criteria indicates that t-Copula demonstrates significant advantages in characterizing the combined distribution of hydropower, wind power, and solar output. Its average AIC value of −28.221 and average BIC value of −23.797 are both much lower than those of other candidate Copula.
This study generates scenarios based on the division of solar terms. During the Winter Solstice and Great Cold, solar output and runoff is at a trough level, while wind power reaches its annual peak. The Summer Solstice and Great Heat correspond to peak solar output, with increased variability in wind power coinciding with high hydropower demand. At the Spring Equinox and Autumn Equinox, the correlation between wind and solar energy reverses. To balance the typicality and conciseness of the research results, this study selects four typical solar terms: the Spring Equinox, the Summer Solstice, the Autumnal Equinox, and the Great Cold. The solar terms selected in this plan cover the entire hydrological cycle and key energy states throughout the year.
Based on the Kernel PCA dimensionality reduction results and Vine Copula reconstruction,
Figure 13 presents the hydro–wind–solar joint correlation scenarios of four typical solar terms. The red line represents the average value of historical scenarios, and the black line represents the average value of generated scenarios. The colored lines represent the different scenarios generated.
① During the rainy season solar terms, the generated wind–solar output scenarios can reflect typical temporal characteristics. For instance, during the Summer Solstice, the wind power output curve shows a distinct peak from 18:00 to 24:00 at night. This is basically consistent with the fluctuation trend of the average curve in the original scenario. The solar energy curve reaches a single-peak between 12:00 and 13:00 in the afternoon, with a very small deviation from the average value of the original scenario. The above results fully demonstrate that the generated scenarios have high accuracy in depicting the collaborative characteristics of the rainy season’s wind and solar energy.
② During the Great Cold period, the generated runoff–wind–solar scenarios are generally consistent with the historical scenarios. It effectively reproduced the operational characteristics of the dry season with hydropower as the base load. Specifically, the average runoff volume is between 500 and 800 cubic meters per second. The fluctuation of wind power output in the morning from 04:00 to 08:00 is consistent with the historical variation characteristics. The solar energy curve was approximately 5% lower than the historical value during the morning rush hour from 11:00 to 13:00. This model accurately reflects the typical characteristics of operation based on hydro power during the dry season.
The deviations in the simulation of solar energy scenarios under drought conditions cannot be ignored. To assess the impact of this deviation in actual operation, this article conducted a sensitivity analysis based on the indicators in
Section 3.4. Detailed results are in
Appendix B (
Table A5 and
Table A6). During this period, if the solar power generation decreases by 5% and 10%, respectively, from 11:00 to 13:00 in the dry season (such as the Winter Solstice and the Great Cold), the volatility error (
) of the hybrid power system will increase by approximately 3.74% to 4.31%, while the other indicators remain stable. This increase in volatility means that during drought periods, the system requires more regulation capacity and reserve capacity to ensure its reliability. Therefore, when formulating response plans for drought periods, it is necessary to fully consider such potential systemic deviations to ensure the stable operation of the system.
③ During the monsoon transition period (the Spring Equinox and the Autumnal equinox), solar power shows a high degree of stability, while wind power exhibits significant volatility. Taking the scenarios of the Spring Equinox and the Autumnal Equinox as examples, the overlap degree of solar energy output during the morning rush hour (5:00–18:00) with the historical average curve is as high as 96%. The trough value of wind energy deviates significantly from 8:00 to 13:00. This comparison result indicates that during the period of seasonal transition, the output of solar power can remain stable, while the output of wind power is prone to significant fluctuations.
In this study, the normalized root mean square error (NRMSE) was obtained by using normalization calculation for Equation (9). The hydro–wind–solar combined scenario successfully captured the unique climatic characteristics of each period. The NRMSEs of all the scenarios were less than 3%, and for key features such as peak and valley values and inflection points, this error was controlled within 5%. This result demonstrates the following: Firstly, the unique climatic characteristics embedded in the division of solar terms restore the complex statistical distribution of the original hydro–wind–solar energy scenario. Secondly, this study confirmed the feasibility and reliability of the Kernel PCA + Vine Copula method in generating comprehensive scenarios with both physical consistency and statistical representativeness in data-scarce river basins.
4.3.4. Generate Scenario Application Analysis
The hydro–wind–solar combined power generation scenario generation method proposed in this study has high accuracy and resolution. Its core advantages lie in the following: ① It can accurately extract the key features and change patterns in historical data. ② It can provide quantitative data support for the calculation of core performance indicators of complementary systems. The effectiveness and practicality of the generated scenario are mainly reflected in the following two aspects.
The verification of physical consistency and statistical reliability: ① The generated scenarios strictly follow the climate cycle laws of the Yalong River Basin. The scenario reproduces the amplitude of output fluctuations, morphological trends, and complementary temporal relationships in the actual system. It has been proved that this method is reliable in characterizing the operational characteristics of the actual system. ② The generated scenarios provide a verified input data basis for quantifying indicators such as the distribution of peak shaving pressure and the fluctuation of wind and solar curtailment rates under different hydrological years. This helps to enhance the accuracy of subsequent planning and scheduling research.
Its supporting role in system analysis and planning design. ① The nonlinear coupling mode of hydro–wind–solar output at different solar terms is clearly presented in the scenario (
Section 4.3.2), providing a quantitative basis for the optimization of the wind and solar capacity ratio based on the stable base load of hydropower. ② The generated scenario captures the periods of sharp fluctuations in wind and solar output, which can be used to analyze the power gap of the system and deduce the power or capacity of the energy storage configuration. ③ The monsoon transition characteristics contained in the scenario can provide a testing basis for designing dynamic scheduling rules that adapt to climate change.
4.4. Comparison of Clustering Algorithms and Scenario Generation Methods
- (1)
Comparison of clustering algorithms
To verify the advantages of the improved clustering algorithm proposed in this study, a comparative analysis was conducted between the proposed method and the traditional K-means algorithm. The comparison results are presented in
Figure 14. In this evaluation, three indicators were employed: The Silhouette Coefficient is used to measure clustering quality, and the higher the value, the better the performance. The Davies–Bouldin (DB) Index is also used, wherein lower values reflect higher within-cluster similarity and greater inter-cluster separation, and as is a CSD indicator. To facilitate consistent interpretation, both the DB Index and the CSD were positively transformed during standardization—thus, for all indicators, larger values correspond to better algorithm performance.
As observed from the line chart, the improved method demonstrated significantly superior performance in all three indicators. In some solar terms, the indicators are particularly prominent (for example, the seventh solar term, Beginning of Summer). The average performance of the indicators of the improved algorithm further confirmed the effectiveness of the improvement: the CSD indicator decreased by 4.6%, the contour coefficient increased by 2.3% and the DB index decreased by 4.5%. These enhancements can be attributed to the algorithm’s integrated consideration of multiple features including power amplitude, pattern trend, and fluctuation position characteristics. By comprehensively capturing these aspects, the improved algorithm more accurately describes the distribution characteristics of complex renewable energy curves, thereby assigning each curve to the most appropriate cluster. By integrating the improvement effects of the three indicators, the traditional clustering accuracy was ultimately increased by 3.8%.
- (2)
Comparison of scenario generation methods
To verify the comprehensive performance of the scenario generation method proposed in this study, two typical comparison schemes are set up in this section for systematic comparison. The scenario quality assessment system constructed based on
Section 3.4 analyzed the three schemes. The comparison scheme is as follows.
Scheme 1 is an improved K-means clustering method + PCA + Copula. This scheme inherits the PCA–Copula model [
23] and extends the two-dimensional inputs of wind and solar energy to a three-dimensional coupled system of hydro–wind–solar energy. Scheme 1 serves as the comparison benchmark for the proposed schemes in this study. Scheme 2 is an improved K-means clustering method + Kernel PCA + Vine Copula, which is the scenario generation method proposed in this study. Scheme 3 is a traditional K-means clustering method + Kernel PCA + Vine Copula. The comparison and analysis results of the scenarios generated by different schemes for the six typical solar terms with the original scenarios are shown in
Table 4.
Based on the analysis of the metrics in
Table 4, the following conclusions can be drawn:
① Analysis of the CSD results. The average CSD values for Scheme 1 (Benchmark), Scheme 2, and Scheme 3 are 0.849, 0.593, and 0.613, respectively. Compared to the benchmark, Scheme 2 shows a 30.2% improvement, while Scheme 3 shows a 27.8% improvement.
② Analysis of the ACF results. The average ACF values are 0.105, 0.031, and 0.093, respectively. Scheme 2 demonstrates a 70.5% improvement over Scheme 1, and Scheme 3 shows an 11.4% improvement.
③ Analysis of the e results. The e values are 0.235, 0.229, and 0.182, respectively. This corresponds to a 2.6% improvement for Scheme 2 and a 22.6% improvement for Scheme 3 compared to the benchmark.
④ Analysis of the results. The values of in the three schemes are 2.651, 0.696, and 0.529, respectively. The volatility errors of both Scheme 2 and Scheme 3 are smaller than those of Scheme 1, improving by 73.7% and 80%, respectively, compared to Scheme 1.
After comparing all the schemes, it can be found that Scheme 2 performs best in terms of comprehensive similarity distance (CSD) and temporal correlation (ACF). The scenario generated by Scheme 2 is the closest to the historical reality in terms of overall form and change pattern, while inheriting the inherent time-dependent structure of the historical sequence. In terms of spatial correlation (e) and volatility error (), although Scheme 2 has improved compared to the benchmark scheme, it is inferior to Scheme 3 in terms of the degree of advantage. This phenomenon precisely indicates that the traditional K-means algorithm is based on Euclidean distance and is extremely sensitive to numerical outliers (extreme output days). It will cluster the highly volatile extreme days separately and retain them as typical patterns. Therefore, Scheme 3 is more accurate in replicating the extreme spatial distribution and fluctuation characteristics in historical data, which directly leads to better values of its e and indicators. This precise replication of extreme values is a double-edged sword. Although it can capture extreme scenarios and spatial distributions, it comes at the cost of sacrificing the morphological fidelity and temporal structure restoration of mainstream and typical patterns.
Taking into account all the indicators and their physical significance comprehensively, Scheme 2 pays more attention to capturing the overall morphological change trend and temporal structure rather than deliberately replicating all the details of spatial correlation. Therefore, it remains the best and most balanced solution for generating high-fidelity combined scenarios of hydro–wind–solar power.
- (3)
Analysis of algorithm scalability
The codes of all the schemes in this article were run on the MATLAB R2024a platform. The computing times of the three schemes were 1620 s, 1487 s, and 1380 s, respectively. The running time of a PCA is longer than that of a Kernel PCA, while the traditional clustering method is faster than the improved clustering method. Taking Scheme 2 as an example, when ten more stations were added, the operation time of this method was 186 s, an increase of 374 s (25.2%). When 20 more stations were added, the operation time reached 2493 s, an increase of 632 s, which was 1006 s (67.7%) longer than the initial operation time of Scheme 2. This result indicates that as the number of stations increases, the computational complexity of the algorithm in this study also rises. The increase in complexity is not linear because the scenario is generated based on principal components, and each time new stations are added, the number of principal components required will increase.
Specifically, in the original method, the first 15 principal components are required to reflect the complex coupling relationship. After the number of stations increases, the first 20 or even the first 30 principal components are needed. This will directly lead to a linear increase in the time for extracting principal components, and the computational complexity of Copula shows a combinatorial increase, which in turn affects the computational complexity of the entire algorithm. Based on this, it can be inferred that when the number of stations increases by 50 or more, more principal components will be required, and the computational complexity will rapidly reach an unbearable level. This scale might be the maximum that the algorithm can handle.
4.5. Sensitivity Analysis of the Solar Term Division Generation Scenario
This section designs the misalignment sensitivity experiment of the system. The core purpose is to quantitatively assess the impact on the fidelity of the scenarios generated by this framework when the preset astronomical/climatic segments (the 24 solar terms) deviate from the real meteorological processes on the time axis.
In this study, the actual climate was offset by 7 days and 15 days compared to the solar term division standard. Its actual meaning is to assume that the actual climate model is half a solar term to one solar term earlier or later than the solar term division model. Based on this, a sensitivity analysis was conducted. All data were generated through scenario generation in Scheme 2, and the generated scenarios were compared with the original scenario to quantify various indicators. As the misalignment involves the influence of adjacent solar terms, this section uses the average values of all the indicators for the entire solar term for analysis. The results of the sensitivity analysis are shown in
Table 5.
The sensitivity analysis indicates that the mismatch between the division of solar terms and the actual climate model generally leads to the deterioration of the morphological similarity (CSD) and temporal correlation (ACF) indicators of the generated scenarios. It is worth noting that in specific offset scenarios (such as 15 days in advance), the CSD value may be slightly lower than the standard scenario. This might be due to the misalignment accidentally grouping similar weather patterns into the same analysis window, but the overall trend indicates that the misalignment introduces additional uncertainty. The spatial correlation (e) and volatility error () fluctuated up and down in all cases, and there was no obvious tendency towards optimization or deterioration. The numerical changes of these two indicators are not significant, indicating that they have certain robustness in terms of spatial dependence and system fluctuation characteristics during misalignment.
The variation patterns of the above indicators can be explained by the data-driven modeling principle: Firstly, the general increase in ACF is due to the misalignment that disrupted the temporal coherence of weather processes within the original climate stage, making it difficult for the model to capture the precise phase-dependent temporal autoregressive structure. Secondly, the deterioration of the CSD is due to the misalignment that blurs the morphological boundaries of different climatic sub-models, causing the generated typical daily curves to deviate from the true best representatives in terms of amplitude, trend, and position. Relatively speaking, the stability of and e suggests that the fluctuation intensity and spatial correlation of the system’s net output are less sensitive to the precise temporal alignment of the climate stage. As long as the distribution of weather types within the misaligned window is similar, its statistical fluctuation characteristics and spatial correlation can be roughly retained.
This case study shows that a time window of approximately 15 days defined by astronomical solar terms exhibits certain robustness against climate phase shifts within a half-month scale. However, this does not prove that the boundaries of solar terms remain valid under long-term climate change. If climate change causes continuous phase shifts in dominant weather systems such as monsoons, the long-term cumulative offset will eventually exceed the duration of a single solar term (15 days). At that time, continuing to use a fixed astronomical boundary will lead to a serious mixture of information from different climate mechanisms in the training data, fundamentally weakening the physical consistency and predictive ability of the model. Therefore, the results of this sensitivity analysis further prove the necessity of developing adaptive climate stage identification algorithms in future directions. A dynamic segmentation method capable of automatically identifying statistical characteristics from data is a fundamental approach to addressing the uncertainties of climate change.