Next Article in Journal
Impacts of Distributed Renewable Energy Source Integration on the Reliability of Distribution Networks: A Bibliometric Review
Previous Article in Journal
Joint Planning of Battery Swapping Stations and Distribution Networks to Enhance Photovoltaic Utilization
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Climate-Informed Scenario Generation Method for Stochastic Planning of Hybrid Hydro–Wind–Solar Power Systems in Data-Scarce Regions

1
College of Electrical Engineering and New Energy, China Three Gorges University, Yichang 443002, China
2
College of Hydraulic and Environmental Engineering, China Three Gorges University, Yichang 443002, China
*
Author to whom correspondence should be addressed.
Energies 2026, 19(1), 74; https://doi.org/10.3390/en19010074
Submission received: 12 November 2025 / Revised: 19 December 2025 / Accepted: 20 December 2025 / Published: 23 December 2025
(This article belongs to the Section A: Sustainable Energy)

Abstract

The high penetration rate of renewable energy poses significant challenges to the planning and operation of power systems in regions with scarce data. In these regions, it is impossible to accurately simulate the complex nonlinear dependencies among hydro–wind–solar energy resources, which leads to huge operational risks and investment uncertainties. To bridge this gap, this study proposes a new data-driven framework that embeds the natural climate cycle (24 solar terms) into a physically consistent scenario generation process, surpassing the traditional linear approach. This framework introduces the Comprehensive Similarity Distance (CSD) indicator to quantify the curve similarity of power amplitude, pattern trend, and fluctuation position, thereby improving the K-means clustering. Compared with the K-means algorithm based on the standard Euclidean distance, the accuracy of the improved clustering pattern extraction is increased by 3.8%. By embedding the natural climate cycle and employing a two-stage dimensionality reduction architecture: time compression via improved clustering and feature fusion via Kernel PCA, the framework effectively captures cross-source dependencies and preserves climatic periodicity. Finally, combined with the simplified Vine Copula model, high-fidelity joint scenarios with a normalized root mean square error (NRMSE) of less than 3% can be generated. This study provides a reliable and computationally feasible tool for stochastic optimization and reliability analysis in the planning and operation of future power systems with high renewable energy grid integration.

1. Introduction

Hybrid systems that combine hydro–wind–solar energy are widely regarded as a key solution for achieving a high penetration of renewable energy applications while maintaining the stability of the power grid [1]. The effective planning and reliable operation of such integrated systems fundamentally rely on reliable long-term joint energy prediction data. However, in many remote areas or newly developed river basins, there exists a serious obstacle: long-term, high-resolution meteorological and hydrological observation data are extremely scarce [2]. This data shortage forces planners and researchers to rely on other often imperfect data sources, which directly weakens the foundation for subsequent stochastic optimization analysis or investment decisions.
Therefore, researchers have had to turn to multi-source public datasets, such as reanalysis data, as an alternative [3]. This dependency brings two interrelated challenges to generating reliable scenario analysis results. Firstly, public datasets usually contain systematic biases. When matched with local climatic conditions, these data also show inconsistencies in morphological characteristics, especially in areas with complex terrain [4,5]. More crucially, traditional point-to-point evaluation metrics are not sufficient to assess the overall similarity of data curves in terms of distribution patterns and fluctuation patterns [6,7]. If unverified data are directly used, these uncertainties are very likely to continue to spread throughout all subsequent modeling stages. The second issue is that even if relevant data are obtained, due to their origin from interannual time series, the high-dimensional characteristics of multiple spatial observation points, and the interrelationships among different energy types still pose significant difficulties for the modeling work [8]. The existing dimensionality reduction methods have shortcomings when dealing with such problems: the clustering algorithm based on Euclidean distance cannot accurately capture the changing trend of time series [9]. PCA is essentially a linear analysis method, and thus cannot reflect the nonlinear correlation among different energy sources [10]. However, if each type of energy is reduced in dimension separately, the cross-source spatiotemporal correlation mechanism will be disrupted [11]. At present, there is still a lack of a method that can effectively compress the data volume while retaining the regularity of the climate cycle and the nonlinear relationship between energy. This problem remains a key bottleneck, restricting the modeling work.
To address the above challenges, this study constructs a coherent framework from data screening to scenario generation: (1) To ensure the reliability of the input data, this study introduces a comprehensive similarity distance indicator. This indicator quantifies the similarity between data curves from three perspectives: power amplitude, pattern trend, and fluctuation position, thereby providing an effective tool for screening reliable public data. (2) To deal with high-dimensional complexity, a two-stage dimensionality reduction strategy was designed. Firstly, time compression is achieved through improved cluster analysis based on the 24 solar terms. Secondly, the Kernel PCA technique is jointly applied to the hydro–wind–solar datasets for feature compression. This can effectively extract and retain the nonlinear and cross-source dependencies among these data. (3) To achieve efficient and high-fidelity scenario generation, this study adopted a simplified rattan Copula model based on the low-dimensional data integrated with various features obtained in the previous step. This method not only avoids the problem of excessively high dimensions but also generates a joint scenario that is statistically representative and physically conforms to the actual laws. This framework is designed to provide a reliable basis for the stochastic planning and operation of hybrid hydro–wind–solar systems in regions where measurement data are relatively scarce.

1.1. Literature Review

1.1.1. Multi-Source Data Quality Assessment

In the research of renewable energy, ensuring the reliability of multi-source public data is often not given sufficient attention. The current widely adopted methods rely on traditional statistical error indicators and correlation coefficients to evaluate data quality [6,7]. These indicators are used to measure the accuracy of numerical values and the similarity between different data sources [12]. The entropy weight method is used to further integrate various indicators and conduct a comprehensive evaluation of the dataset [13]. Although these indicators are necessary for assessing the accuracy of the data, they are not sufficient to evaluate the overall representativeness of the data curve. More crucially, these indicators cannot quantify whether different data sources have similar distribution characteristics, dynamic change trends, and fluctuation patterns under the same climatic conditions. This oversight is extremely serious: if data are selected merely based on numerical similarity, it may lead to the obtained data failing to accurately reflect the local hydro-meteorological conditions, thereby causing errors in subsequent complementarity analysis and system optimization work.
In addition to statistical indicators, pattern-based sequence comparison methods provide more detailed analytical means for evaluating curve similarity. Dynamic Time Warping algorithms and shape-based clustering techniques have been developed to capture time alignment relationships and morphological features [14,15]. However, when these methods are applied to the specific scenario of screening multi-source hydro–wind–solar data, they often lack a comprehensive and easily understandable evaluation framework. These methods may only focus on one aspect of similarity (such as shape alignment), while ignoring other key factors, such as the amplitude of total energy fluctuations or the specific temporal characteristics of the fluctuations. Therefore, there is still an urgent need for a specialized similarity assessment framework that can simultaneously quantify multiple key features of multi-energy curves. Moreover, its assessment results not only have physical significance but can also be directly applied to data screening work, especially in regions where data is scarce.

1.1.2. Dimensionality Reduction and Uncertainty Expression in Hydro–Wind–Solar Systems

Cluster analysis and probabilistic modeling are the main research directions for reducing the complexity of multi-energy systems and characterizing uncertainties. Widely adopted methods such as K-means clustering and Monte Carlo simulation provide important tools for identifying representative data patterns [16,17]. These studies provide practical tools for simplifying high-dimensional data and characterizing uncertainties. However, the effectiveness of these methods is essentially limited by the evaluation indicators and assumptions they employ. Specifically, the standard K-means clustering method, due to its reliance on Euclidean distance, is unable to capture the temporal dynamic changes of the energy curve. It can only compare the data values at the same time point, but cannot reflect the characteristics based on data form such as the amplitude change trend or the timing of fluctuation occurrence [18]. Therefore, the typical scenarios derived from this may wrongly reflect the real spatiotemporal interaction mechanism in a hydro–wind–solar system, thereby bringing potential deviations to the subsequent planning models.
Another parallel study employed the Copula function and dimensionality reduction techniques to analyze the interdependence among various energy sources. Relevant studies have adopted the joint confidence constraint mechanism and Copula fusion technology to quantify various correlations within hydraulic, wind, and solar energy systems [19,20]. In terms of feature compression and relationship preservation, Principal Component Analysis (PCA) and Kernel Principal Component analysis (Kernel PCA) are commonly used techniques [21,22]. PCA captures the linear correlations in data during the dimensionality reduction process [23], while the integration of Kernel PCA and variational mode decomposition enhances the nonlinear feature extraction and dimensionality reduction capabilities [24]. These methods highlight the prospects of statistical dependency modeling; they have key gaps in the context of integrated multi-energy systems. PCA is essentially linear and unable to capture the complex nonlinear couplings existing in the output of hydro–wind–solar energy. More notably, even more powerful Kernel PCA is usually applied in isolation to a single energy source, ignoring the fundamental task of cross-source feature fusion. This oversight neglects the key spatiotemporal dependencies that determine the complementary effects among various energy sources, thereby resulting in the generated simulation results lacking physical authenticity.

1.1.3. Method for Generating Combined Scenarios of Hydro–Wind–Solar Energy

In the final stage of generating joint scenarios, various methods mainly rely on the combination of the Copula model and Monte Carlo simulation. Advanced tools such as Vine Copula provide structured approaches for modeling multivariable dependencies [25], while combining them with techniques such as TimeGAN helps capture temporal dynamics [26]. These techniques have laid a solid foundation for the construction of stochastic complementarity and high-dimensional uncertainty models. However, the linear Copula model has difficulty handling nonlinear and asymmetric dependencies, while the Vine Copula model faces the challenge that parameter estimation and computing time increase exponentially with the increase of dimensions [27]. In addition, traditional scenario generation methods often fail to capture short-term climate changes that are crucial for renewable energy generation in terms of time division. For instance, the onset or fading of a monsoon within a month may lead to sudden wind–solar complementary changes, while the seasonal average smoothed out the fluctuations within the cycle. This inconsistency between the administrative calendar (for example, months or seasons) and the natural climate rhythm limits the fidelity of the scenario. These issues collectively create a comprehensive problem among statistical accuracy, computational feasibility, and physical authenticity, which remains unsolved at present.
To address the computational burden, researchers have designed a hybrid framework that integrates the knowledge of dimensionality reduction, decomposition, and operations. Through multi-stage feature extraction (Kernel PCA) and model optimization strategies, its computational efficiency and prediction accuracy have been improved [28]. The idea of multi-level segmentation is an effective strategy for dealing with complex nonlinear problems. Reference [19] divides the entire cycle into multiple consecutive stages according to the time scale or the operational characteristics of the system. At each stage, the originally highly nonlinear complex relationships such as power grid constraints are simplified. Each stage is reconstructed and linearized into a mixed integer linear programming model, thereby reducing complexity and improving computational efficiency. These research achievements represent significant progress in building scenarios that better meet real needs and have practical application value. However, when the research objective shifts from simple prediction to generating physically reasonable joint scenarios for planning, these methods reveal obvious limitations. Firstly, these methods have shortcomings in integrating the physical mechanisms of the climate; due to the adoption of relatively rough or operational-demand-defined time division methods, they are unable to effectively analyze the changing processes that determine the generation patterns of renewable energy. Secondly, even if advanced technologies such as Kernel PCA are adopted, the current common practice of separately treating various energy sources still neglects the nonlinear mutual influence relationship among the outputs of hydro–wind–solar energy. Therefore, although such frameworks have advantages in computational efficiency, the scenarios they generate often have problems in terms of physical rationality, which seriously limits their practical application value in high-risk and long-term planning decisions.

1.2. Unresolved Problems and Contributions

Although existing research has made significant progress in data screening, dimensionality reduction analysis and scenario generation of hydro–wind–solar energy systems, there are still several key deficiencies at present.
(1)
In river basins with scarce data, the existing methods have failed to effectively address the fundamental issue of the reliability of input data [29]. Traditional evaluation indicators are indispensable, but they have limitations. Existing research has demonstrated that merely verifying reliability through numerical values is not comprehensive; instead, it is necessary to start from the overall characteristics [30]. Although they can reflect the accuracy and correlation of some values, they struggle to capture the characteristic similarities among curves such as the amplitude, shape, and fluctuation position of multi-source data. More importantly, the existing studies usually assume that the selected data sources can represent the local actual situation, but lack a systematic assessment of the consistency among multi-source data and their fit with the natural climate cycle [31]. This issue is more prominent in river basins where measured data are limited. Although reanalysis data can be obtained, its unverified system bias and resolution error will affect the accuracy of subsequent scheduling processes and capacity optimization. Therefore, establishing a reliable data evaluation framework that combines distribution feature analysis with traditional error detection methods is a prerequisite for achieving reliable scenario generation.
(2)
When conducting dimensionality reduction processing on ultra-high-dimensional interannual data, trade-offs need to be made between fidelity and the mutual feedback relationship between the data. Although the K-means algorithm can compress the time dimension, it relies on the Euclidean distance calculation method, which requires comparisons at strictly aligned time points. Therefore, it cannot capture the essential morphological similarities between curves that may experience time shifts or phase misalignments [9]. Although PCA can reduce the number of features, it will ignore nonlinear features [10]. Independent dimensionality reduction of energy dimensions may ignore the cross-source spatiotemporal coupling relationship [11]. If there is a lack of a joint compression mechanism that can both reduce the data dimension and maintain the correlation and morphological characteristics, researchers will face two choices: either accept the distortion of the synergy effect of hydro–wind–solar energy, or bear the high computational cost caused by the data dimension. Therefore, there is an urgent need for an architecture capable of performing hierarchical spatiotemporal dimension compression, which is explicitly designed to preserve the periodic and nonlinear cross-energy characteristics of the climate.
(3)
Most of the existing methods decompose the three variables into binary relationships such as wind–solar and hydro–wind for modeling [32]. This simplification tends to overlook the triple coupling mechanism of runoff, wind power, and solar power [33]. Although a limited number of studies have attempted to quantify the correlations among hydropower, wind energy, and solar energy, they are still insufficient in quantifying the complex nonlinear correlation among the three [34]. The application effects of the Vine Copula model and the Monte Carlo method outside binary scenarios are not very satisfactory: the parameter growth and estimation error rise sharply, and the complementary models (such as the runoff buffer period coinciding exactly with the low period of wind and solar energy) may not hold. Purely data-driven processes also underestimate climatic cycle factors, leading to impaired physical interpretability and causing deviations in the assessment of planners flexibility.
To address these interconnected challenges, this study makes three core contributions:
(1)
To address the issues of incomplete quality assessment of multi-source data, this study proposes a new CSD indicator. The overall distribution similarity among multi-source data curves is quantified from three dimensions: similarity of power amplitude ( Δ E ), similarity of morphological trend (SBD), and similarity of fluctuation position ( δ ), thereby providing key pattern insights that are lacking in traditional error indicators. This study combined this indicator with the traditional error indicator to construct a comprehensive evaluation system.
(2)
To address the inherent limitations of traditional dimension reduction methods, this study proposes an innovative two-stage dimension reduction framework. This framework systematically compresses data while maintaining the relevant structure of the data core. In the first stage, the traditional K-means algorithm is enhanced by integrating the CSD to better capture the curve distribution characteristics that Euclidean distance fails to represent. The second stage addresses the feature compression problem by innovatively applying Kernel PCA to the entire hydro–wind–solar dataset. This step clearly captures the complex, nonlinear interactions and spatiotemporal dependencies among all three energy sources, overcoming the limitations of isolated processing.
(3)
To break the deadlock between high-dimensional dependency modeling and computational feasibility, this study pioneered a generative paradigm that collaboratively couples Kernel PCA with Vine Copula. This method converts high-dimensional hydro–wind–solar energy data into low-dimensional latent space, making complex dependencies easier to handle. Such a preprocessing step effectively avoids the curse of dimensionality that plagues the direct application of Vine Copula. Then, a simplified Vine Copula was implemented on this refined latent representation, generating high-fidelity scenarios efficiently for the first time. These scenarios essentially reflect the complex nonlinearity of the combined system and the key climate cycles. It has been verified that the normalized root mean square error (NRMSE) between the scenario generated by the method in this study and the original scenario is less than 3%.
To fill the aforementioned gap, this study proposes a comprehensive framework that integrates climate information into the scenario generation process. The core contributions of this framework are reflected in three aspects: Firstly, a CSD indicator has been constructed to address the issue of the insufficient reliability of input data, thereby enabling the conduction of morphological feature-based evaluations on multi-source data. Secondly, a two-stage dimensionality reduction method was developed; firstly, the time series data were compressed based on the 24 solar terms to retain the climate change patterns within it. Then, the Kernel PCA technique was applied to the joint dataset including hydro–wind–solar energy data, thereby fusing various features and capturing the nonlinear correlations among the data sources. Finally, combining this dimensionality reduction processing result with the simplified Vine Copula model makes the generation process of high-fidelity joint scenarios computationally feasible. Overall, this framework can provide planning work with a basis that conforms to physical laws and is statistically reliable in an environment where data are scarce.
Section 2 elaborates in detail on the construction of a multi-source data quality assessment and screening framework. Section 3 introduces the scenario generation method based on two-stage dimensionality reduction. Section 4 presents case studies. Section 5 summarizes the research results. Section 6 explores the limitations and future research directions.

2. Multi-Source Data Quality Assessment and Screening Framework

To build a reliable foundation for scenario generation, it is first necessary to establish a multi-source data quality assessment and screening framework. Without this foundation, subsequent modeling steps may introduce systematic biases or ignore the key long-term patterns contained in heterogeneous data. Based on this research motivation, this section proposes a comprehensive method for the identification, screening, and verification of multi-source public data that employs the following: (1) A data distribution feature assessment based on an improved clustering algorithm. (2) Use error indicators to verify the accuracy of the data point by point. (3) A comprehensive assessment that combines the above two methods. Such a mechanism can ensure that only high-quality data can enter the subsequent processing stage.

2.1. Data Distribution Feature Evaluation Based on Improved Clustering

When building a reliable scenario generation model, the input datum not only needs to have high numerical accuracy but also should pay attention to its distribution characteristics. Among various clustering methods, K-means is widely adopted due to its simplicity and efficiency based on Euclidean distance, and it can effectively capture the global structure of data [35]. Although it can identify typical cluster centroids, it falls short in characterizing amplitude, morphological trend, and the fluctuation position characteristics in time series data. To address this issue, this article introduces a comprehensive similarity measurement method—this algorithm can provide reliable support for evaluating the distribution characteristics of data.

2.1.1. Comprehensive Similarity Measurement of Curves

To enhance the accuracy and generalization ability of the algorithm, this study decomposes the overall similarity of any two sequences into three interpretable dimensions: power amplitude similarity, morphological trend similarity, and fluctuation position similarity.
Taking the solar output sequence as an example. The comprehensive similarity between sequence x and sequence y is specifically evaluated through the following three aspects:
(1)
The similarity of power amplitude, denoted as Δ E ( x , y ) , reflects the relative difference between the two curves and the total energy enclosed by the time axis. The idea of this article is to compare the total power generation or calculate the general practice of energy error through the integrated power curve [36]. However, in this study, this fundamental concept is constructed as a normalized distance metric with a range of [1, 2], and the smaller the value, the more similar the amplitude.
Δ E ( x , y ) = 1 + t = 1 T x t t = 1 T y t M a x ( t = 1 T x t , t = 1 T y t , )
where x t and y t represent the values of the power output sequences x and y at the t-th time period. The range of values for Δ E ( x , y ) is [1, 2]. The smaller the value of Δ E ( x , y ) , the greater the similarity between the power areas of sequences x and y.
(2)
The similarity of morphological trend, denoted as S B D ( x , y ) , reflects the alignment coefficient when sequence y is shifted to best match the trend of sequence x. This study adopts the normalized correlation number as the basis and converts this coefficient into the form of distance. The normalized cross-relation number is a classic metric in the shape matching of time series [37].
N C C ( x , y ) = i = 1 n ( x i x ¯ ) ( y i y ¯ ) i = 1 n ( x i x ¯ ) 2 i = 1 n ( y i y ¯ ) 2
where x ¯ and y ¯ are the mean values of the time series x and y. For the convenience of calculation, the similarity distance between the morphological trends of sequences x and y is defined as S B D ( x , y ) :
S B D ( x , y ) = 1 max ( N C C ( x , y ) )
The range of S B D ( x , y ) is [0, 2], where 0 indicates that the morphological trends of sequence x and y are completely similar, and 2 indicates that they are completely opposite.
(3)
The similarity of fluctuation position, denoted as δ ( x , y ) , captures the displacement required to align the fluctuation features (peaks and troughs) of y with x. This metric aims to address the phase misalignment problem between sequences, which has been deeply explored in classic time series alignment algorithms such as Dynamic Time Warping [18]. Unlike the complex curved paths of DTW, this method simplifies it to finding an optimal global translation to measure the minimum displacement cost required to align the wave features.
δ ( x , y ) = 1 + s ( x ) b e s t T
where s ( x ) b e s t represents the displacement when the output sequence y is translated to the most similar morphological trend to sequence x. The smaller the value of δ ( x , y ) is, the more similar the fluctuation positions of the two curves are. The value range is [1, 2].
Because the unit types of these three indicators are different, the traditional addition and subtraction calculation methods are not applicable. Moreover, in the context of scarce data, there is a lack of reliable basis for allocating reasonable and universal weights for the three. At the same time, linear-weighted summation may mask the defects of specific dimensions due to the compensation effect among different dimensions. Multiplicative aggregation equally focuses on all three physical dimensions in a parameterless manner, simplifying the process and enhancing the interpretability and reproducibility of the method. This processing has an amplifying effect on any significant difference in a single dimension, thus enabling the acquisition of a larger CSD value, which can effectively distinguish the curves in subsequent screening or clustering. Therefore, in this study, the three similarity indicators are multiplied to construct a comprehensive similarity distance C S D ( x , y ) . The smaller the C S D ( x , y ) value is, the higher the similarity between the two sequences is. The indicator equation is as follows:
C S D ( x , y ) = Δ E ( x , y ) S B D ( x , y ) δ ( x , y )
As shown in Figure 1, the area difference corresponds to the similarity of power amplitude, representing the relative difference between the two curves and the regions enclosed by the time axis. The optimal coefficient corresponds to the normalized cross-relation number that aligns two curves in the similarity of morphological trends. The optimal displacement corresponds to the translation amount required for feature alignment in the similarity of the fluctuating position.
In mathematics, a strict distance metric is typically required to satisfy four conditions: non-negativity, symmetry, identity, and triangle inequality. The CSD proposed in this study naturally satisfies the first three conditions. ① Its non-negativity stems from the non-negative property of the product of the values of the three constituent components. ② Symmetry is guaranteed by the computational symmetry of each component. ③ Identity is ensured because if and only if the power amplitude, morphological trend, and fluctuation position of the two curves are exactly the same, all components are simultaneously 0, resulting in C S D ( x , y ) = 0. ④ Triangle inequality, however, is not enforced in this study. This is a deliberate design choice: the goal of a CSD is not to construct a rigorous metric space but to provide a more faithful pairwise comparison of time-series curves. The CSD is only used to assign data points to the nearest center point and update the center point, thus deliberately relaxing the restrictions of triangular inequality. These comparisons do not require the calculation of distance through a third point, so triangle inequality is independent of the correctness of the algorithm. Furthermore, subsequent modeling processes (such as Kernel PCA or Vine Copula models) also rely solely on the clustered data points themselves and are independent of the CSD distance matrix. Therefore, this relaxation of triangular inequality does not affect the subsequent modeling results, which is also verified by the high fidelity of the final output (Section 4.3.4 and Section 4.4).

2.1.2. Adaptive Method for Determining the K Method Based on the CSD

Because CSD violates the Euclidean space axiom [38], traditional K-selection methods become unreliable or theoretically invalid. To address this fundamental incompatibility issue, this study proposes a novel adaptive K-selection framework specifically designed for distance measurement methods such as the CSD. This framework has made three key adjustments to the standard process:
(1)
It redefines the intra-cluster dispersion: This method replaces the traditional Euclidean distance sum of squares equation, namely Equation (6), with the total dispersion calculated based on the CSD.
W C S D ( K ) = k = 1 K x i C k t C S D ( x i , μ k t )
where K represents the number of candidate clusters; W C S D ( K ) represents the total CSD from all sample curves to the representative curves of their respective clusters. C k is the k-th cluster, x i is the intra-cluster sample curve, and t represents the number of iterations. C S D ( x i , μ k t ) represents the comprehensive similarity distance, which is calculated by Equation (5). μ k t is the representative curve within the cluster, that is, the actual sample curve in the selected cluster that minimizes the sum of the CSD from all curves to this curve.
(2)
It maintains physical rationality: The representative value of a cluster (the center point of the cluster) is defined as the actual sample curve existing in the dataset (Equation (7)), rather than the arithmetic mean of each point. This ensures that these representative values can truly reflect the changing trend in the daily power generation process.
μ k t + 1 = arg   min x C k t C S D ( x , x i )
(3)
The optimal K value is determined by the inflection point criterion: The optimal K value is determined by identifying the most significant inflection point in the relationship curve between the functions W C S D ( K ) and K in the chart, as shown in Figure 2.
This study assesses the reliability of data from the perspective of data distribution characteristics. The data were transformed into a low-dimensional typical diurnal variation pattern using the improved CSD-K-means algorithm. By comparing the CSD values of typical days in multi-source data, the data source that is most consistent with the distribution characteristics of the measured data is selected as the reliable data source. The distribution feature extraction process based on the improved K-means clustering algorithm is as follows:
Step 1: Data cleaning and preprocessing into matrix form. Clean up the original data and handle missing values, delete duplicate values, and correct outliers. Organize the preprocessed data into a matrix form. The matrix format is as follows:
R e = r 1 , 1 r 1 , 2 r 1 , N r 2 , 1 r 2 , 2 r 2 , N r T , 1 r T , 2 r T , N R T × N
where R e represents the data matrix of energy type e, T represents the time period of the dataset, and N is the number of selected data samples. r represents the sample data within a certain period of time. The values of T and N are determined by the time scale of the original data and the number of data samples taken within the time period. Based on this generalized data input structure, for this study, the input matrix (T × N) represents the number of days T for each solar term and the number of hours N for each day. Therefore, the feature space of clustering in this article is the space composed of T × 24 data curves.
Step 2: The samples are divided into clusters. For each sample curve x i , calculate its CSD from all the clustering centers C S D ( x i , μ k t ) and assign it to the cluster corresponding to the minimum CSD.
Step 3: Update the clustering centers. For each cluster C k t , select the actual sample that minimizes the total CSD within the cluster as the new representative point, as shown in Equation (7).
Step 4: Iterate in a loop until convergence. Repeat Step 2 and Step 3 until satisfied: the representative point change rate is less than or equal to 1% or the algorithm has reached the maximum number of iterations.
When selecting the optimal K value based on the above steps, for each candidate K, calculate the total CSD W C S D ( K ) from all sample curves to the representative curve of the cluster to which it belongs. When K is less than the true number of clusters, the newly added clustering centers significantly reduce the intra-class differences, and W C S D drops rapidly. When K exceeds the true number of clusters, the decline of W C S D slows down. The optimal K value can be determined by locating the inflection point of the W C S D ~ K curve. This adaptive K-value selection method is an indispensable part of improving the K-means clustering algorithm to ensure an accurate number of clusters.

2.2. Point-by-Point Accuracy Verification Based on Error Indicators

The distribution characteristics of data can help gain insights into the overall pattern, but short-term biases may still affect the credibility of multi-source datasets in practical applications. To capture these deviations, this section introduces several commonly used error indicators.
The root mean square error is used to measure the deviation between the reference value and the original data value.
R M S E = 1 N i = 1 N ( y i y ^ i ) 2
where y i   represents the reference value, and y ^ i represents the original data value.
The mean bias error reflects the systematic bias direction of the original data.
M B E = 1 N i = 1 N ( y ^ i y i )
where N represents the number of samples.
The index of agreement is used to evaluate the matching degree between the original data and the reference value.
I O A = 1 i = 1 N ( y i y ^ i ) 2 i = 1 N ( y ^ i y ¯ + y i y ¯ ) 2
where y ¯ represents the mean of the reference value.
The Pearson correlation coefficient is used to measure the linear correlation between the original data and the reference value.
R = i = 1 N ( y i y ¯ ) ( y ^ i y ^ ¯ ) i = 1 N ( y i y ¯ ) 2 i = 1 N ( y ^ i y ^ ¯ ) 2
The Spearman rank correlation coefficient is used to conduct non-parametric correlation tests based on ranks. The error indication equation is as follows:
ρ = 1 6 i = 1 N d i 2 N ( N 2 1 )
where d i represents the rank difference between the reference value and the original data.

2.3. A Comprehensive Assessment Combining Characteristics and Error Indicators

To more accurately assess and screen renewable energy data, in this study, the evaluation results of distribution characteristics and error indicators are combined through the entropy weight method to construct a comprehensive evaluation indicator. The specific processing procedure is as follows:
Step 1: Positive processing of indicators. Take the reciprocal of the reverse indicators RMSE and CSD, and take the absolute value of the MBE indicator before taking the reciprocal. Ensure that the larger the values of all the indicators are, the better the data performance or reliability will be.
Step 2: Standardized processing of indicators. Using the minimum–maximum normalization method [39], the data of each indicator that has undergone forward normalization processing is linearly transformed into the interval of [0, 1].
Step 3: Determine the comprehensive weight and indicators. The entropy weight method is applied to calculate the weights of each standardized indicator [40]. Multiply the values of each standardized indicator by their corresponding weights and sum them up to obtain the final comprehensive evaluation indicator S. Its core equation is as follows:
S = w 1 S 1 / R M S E + w 2 S 1 / M B E + w 3 S I O A + w 4 S R + w 5 S 1 / C S D + w 6 S ρ
where S ( · )   is the weighted comprehensive evaluation indicator. The comprehensive evaluation indicator is a positive indicator, with a value range of [0, 1]. As the indicator value increases, the overall reliability of the data improves.

3. Scenario Generation Based on Two-Stage Dimensionality Reduction

Reliable multi-source hydro–wind–solar energy data were screened out based on the comprehensive quality assessment and screening framework established in Section 2. These data have a relatively high dimension due to the large volume of temporal observation data, the wide distribution of spatial sites, and the diversity of energy types. Thus, complex spatiotemporal correlations arise among seasons, day–night cycles, and geographical locations. This chapter proposes a two-stage dimensionality reduction strategy for scenario generation. The first stage compresses the time dimension via clustering based on solar term division, preserving seasonal and climatic cycles. The second stage employs Kernel PCA to extract nonlinear cross-source dependencies, overcoming the limitations of conventional linear approaches. Together, these steps generate structured and high-quality input data for subsequent correlation modeling and scenario generation. Unlike traditional methods that rely on linear PCA and simple Copula models, the proposed approach incorporates nonlinear feature extraction and embeds seasonal physical constraints, significantly improving both the accuracy and robustness of dependency modeling.

3.1. First Stage: Time Dimension Compression Based on Improved Clustering

Section 1.2 (2) identifies the challenge of high-dimensional data caused by long-term interannual observations, multiple spatial sites, and cross-energy couplings. To address this, the first compression stage employs two complementary strategies while retaining key climate cycle features: (1) The 15-year data is divided into 24 climate periods according to the solar terms. (2) The extraction of typical daily scenarios within each solar term using the CSD-improved clustering method. The solar terms are astronomical definition stages based on the ecliptic longitude of the sun (with an increment of 15 degrees), which directly determine the intensity of solar radiation. Crucially, they coincide with the key monsoon transitions in East Asia. This fine-grained 15-day scale addresses monthly institutional changes that cannot be reached by monthly/seasonal time frames. The resulting data matrix for solar term division is structured as follows:
R s t = [ R s t 1 , R s t 2 R s t 24 ]
where s t ( · ) represents the name of the solar term. Taking the above matrix as the input, the improved clustering algorithm (Section 2.1.2) is applied to cluster the runoff–wind–solar data of all sub-matrices. Cluster each data source in the sub-matrix respectively on the same time scale to obtain the typical daily scenarios of all the historical data of the epoch.
As shown in Figure 3, the solid lines in the figure represent the cluster center curves, and the dashed lines represent the similar sample curves within each cluster. This process reduces the time dimension from 5475 (15 years × 365 days) to 24 solar terms. Each solar term incorporates clustering outcomes from all stations to produce n typical days, preserving the diurnal fluctuation patterns that are critical for complementarity analysis.
Organize the 24 solar terms into 24 worksheets, and integrate the typical daily scenarios of runoff–wind–solar energy at all the stations in each solar term worksheet into the form of Equation (16). Equation (17) is a simplified expression of the data format.
R u n = r 1 r 2 r n R n × 72 ,   r i = h 1 , h 2 , , h 24 Runoff , w 1 , w 2 , , w 24 W i n d p o w e r , s 1 , s 2 , s 24 S o l a r p o w e r R n × 72
R u n = [ R h , R w , R s ] R n × ( N h + N w + N s )
where R u n is the combined n × 72-column matrix, h i , w i and s i represent the hourly data of typical intraday runoff–wind–solar energy. Since the number of typical daily scenarios obtained by clustering different energy sources separately in each solar term sub-matrix is different, to meet the requirements of the matrix format and the feasibility of subsequent research calculations, the n of the R u n matrix is set to the minimum number of scenarios among the three types of runoff–wind–solar energy data after clustering. r i represents the combined data vector of runoff–wind–solar energy on a certain day within n days.
This method provides low-dimensional and highly correlated input for the secondary dimension reduction of the subsequent Kernel PCA, laying the foundation for the generation of hydro–wind–solar correlation scenarios.

3.2. Second Stage: Feature Dimension Compression Based on Kernel PCA

After completing the first stage of time dimension compression based on the division of solar terms, a low-dimensional scenario matrix representing the typical daily operation mode of the hydro–wind–solar hybrid system was obtained. However, this matrix still retains multi-dimensional spatiotemporal features (n × 72 dimensions), and the implicit nonlinear correlations among runoff–wind–solar energy across energy sources have not yet been explicitly modeled. As described in Section 1.2, the traditional linear PCA method has the inherent limitations of insufficient characterization when dealing with the complex nonlinear correlation among the power outputs of hydro–wind–solar energy. In order to retain the nonlinear correlation to the greatest extent while reducing the dimension, this section innovatively adopts the principal component analysis method to apply to the hydro–wind–solar combined dataset. This method reveals nonlinear correlations that cannot be seen by traditional methods and retains the most influential interaction patterns among energy sources. The feature dimension has been reduced while retaining the key nonlinear correlation, laying a foundation for the subsequent precise reconstruction of the joint dependency relationship among multiple energy sources.

3.2.1. Nonlinear Feature Extraction via Kernel PCA

This section specifically elaborates on the application mechanism of Kernel PCA in a typical daily scenario matrix. Given that traditional PCA has difficulty capturing nonlinear interactions when dealing with hydro–wind–solar hybrid systems, independent dimensionality reduction by energy sources will lose spatial correlation features. Based on the structured scenario output in Section 3.1, this study innovatively uses Kernel PCA to transform the complex interaction relationship among runoff–wind–solar energy into a more manageable and recognizable low-dimensional pattern. The data input format is a standardized matrix obtained from Equations (16) and (17). The standardized matrix format is as follows:
R ^ = r ^ 1 r ^ 2 r ^ n R ^ n × 72 ,   h ^ i = h i μ h σ h , w ^ i = w i μ w σ w , s ^ i = s i μ s σ s
where R ^ is the standardized matrix obtained, μ h , μ s , μ w represent the mean of each energy historical datum, and σ h   , σ w ,   σ s represent the standard deviation of each energy historical datum.
The Kernel PCA is carried out on the standardized matrix of each solar term. The Gaussian kernel function can effectively capture the complex nonlinear correlations among energy sources. Therefore, this study selects the Gaussian kernel function for analysis. The specific steps are as follows:
Step 1: Capture the nonlinearity among energy sources and eliminate data bias. Calculate the sample kernel values pair by pair [41], based on the input matrix, to construct the symmetric kernel matrix. To eliminate the influence of data offset, it is also necessary to centralize the kernel matrix. The kernel matrix and the centering equation are as follows:
K = k ( r ^ 1 , r ^ 1 ) k ( r ^ 1 , r ^ 2 ) k ( r ^ 1 , r ^ n ) k ( r ^ 2 , r ^ 1 ) k ( r ^ 2 , r ^ 2 ) k ( r ^ 2 , r ^ n ) k ( r ^ n , r ^ 1 ) k ( r ^ n , r ^ 2 ) k ( r ^ n , r ^ n )
K c = K 1 N 1 N K 1 N K 1 N + 1 N 2 1 N K 1 N
where K represents the symmetric kernel matrix, K c represents the centralized kernel matrix, 1 N is the full 1 matrix, and the centering operation is equivalent to subtracting the mean from the data in the high-dimensional feature space.
Step 2: Feature decomposition and low-dimensional representation. Decompose the eigenvalues of the characteristic equation and extract the principal components [42]. The selection conditions of the characteristic equation and principal component are as follows:
K c α = λ α λ 1 λ 2 λ d
where λ represents the eigenvalue obtained from the matrix calculation. Take the first, largest d eigenvalues and the corresponding eigenvectors α 1 , α 2 α d , then normalize the eigenvectors [43]. Construct the load matrix V:
V = [ α 1 , α 2 , , α d ] d i a g ( 1 / λ 1 , , 1 / λ d )
Step 3: Visualize the complex interaction relationships among the energy sources. Map the original data to a high-dimensional space, and then project the data onto a low-dimensional subspace composed of the first d principal components. The principal component matrix is obtained:
Z = K c V T R n × d

3.2.2. Principal Component Selection and Interpretation

After obtaining the projection matrix Z through Kernel PCA, as shown in Equation (23), it is necessary to establish a criterion to balance information retention and dimension compression. Due to the possibility of noise interference (with redundant fluctuations in secondary components) when mapping data to a high-dimensional feature space, this study employs a feature selection criterion based on the cumulative variance contribution rate threshold. This threshold represents a predefined value that determines the number of principal components to retain by specifying the minimum proportion of total variance that must be collectively explained.
For example, if the threshold is set to 85%, the smallest set of components that capture at least 85% of the total variance would be retained, preserving the majority of the original variability while maintaining manageable dimensionality. Conversely, adopting a higher threshold—such as 95%—would retain more variance but require more components, substantially increasing the computational burden in subsequent modeling steps.
Principal component physical interpretation and cumulative variance analysis remain essential to ensure the retention of key features after dimensionality reduction. The physical meaning of the components reflects intrinsic characteristics and correlations within the data, while the cumulative contribution rate quantifies how much of the original information is preserved. Figure 4 illustrates the variance interpretation rates under the chosen criterion.
This graph reflects the changing trends of the variance contribution rates and the cumulative contribution rates of the top n principal components. The left vertical axis represents the independent variance contribution of each principal component. The contributions of the first few principal components significantly retained the most typical data characteristics, which might be the complementarity among runoff–wind–solar energy (Figure 5). The rapid decline in subsequent component contributions may represent some special weather conditions. The cumulative curve on the right vertical axis shows the number of principal components that reach the required cumulative variance threshold. This result provides an intuitive basis for selecting the quantity of principal components.
The figure shows the contribution of energy data to the principal components. The horizontal axis represents the contribution of the eigenvectors of the runoff–wind–solar energy data in the principal component to this principal component (expressed as a correlation), reflecting the coupling strength of different energy sources in the principal component. This step not only explicitly describes the complex nonlinear correlation between runoff–wind–solar energy, but also is a key step in explaining the physical meaning of principal components. This study has made a detailed analysis in Section 4.3.2.

3.3. Vine Copula-Based Scenario Generation

The principal component matrix generated by Kernel PCA provides a low-dimensional representation of the combined characteristics of hydro–wind–solar energy, in which complex nonlinear interactions are encoded and the climatic structure imposed by the solar term division is inherently retained. Traditional scenario generation methods, such as directly sampling the raw data or using linear models [44,45], still face challenges in capturing these high-dimensional correlations. To overcome this challenge and reduce computational difficulty, this study proposes a simplified D-vine Copula framework. Unlike the standard Vine Copula method, this method uses a uniform Copula class throughout the hierarchical tree instead of independently selecting binary functions based on variable pairs. This simplification significantly reduces the complexity of parameter estimation, while retaining the ability to capture tail correlations through thick-tail distributions. The optimal Copula type is selected through the Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC) criteria [46]; Section 4.3.3 analyzes the selection results. To ensure statistical optimality under high-dimensional constraints, the input of this method is the principal component matrix Z reduced by the Kernel PCA, with a dimension of m × d ' (where m is the number of samples and d ' is the number of retained principal components), representing the nonlinear core characteristics of the typical output scenario of hydro–wind–solar energy within each solar term.
The specific process of scenario generation is as follows:
Step 1: Fit the marginal distribution of the principal components. For the principal component matrix Z obtained from the Kernel PCA, the marginal distribution is constructed using non-parametric kernel density estimation [47]:
f ^ j ( z ) = 1 n h i = 1 n K z ( z Z i j h )
where K z is the Gaussian kernel function and h is the bandwidth parameter; j represents the j-th principal component; i represents the i-th historical scenario sample; and z represents the principal component. Obtain the cumulative distribution function through numerical integration:
U j = F j ( Z j ) = Z j f j ( t ) d t
Step 2: Screen the optimal Copula and conduct hierarchical modeling. Construct an adaptive tree structure based on the Kendall Tau correlation coefficient matrix:
(1)
Calculate the correlation coefficient matrix τ i j among the principal components.
(2)
Give priority to fitting the binary Copula with the largest τ i j variable pair:
C ( u i , u j ; θ ) = exp [ ( ( ln u i ) θ + ( ln u j ) θ ) 1 θ ]
The AIC and BIC were used to screen the optimal Copula. Capture the complex dependence relationship between the principal components through a hierarchical Copula.
Step 3: Conditional sampling and inverse transformation driven by dependence structure. Generate samples that follow the dependency structure through the inverse transformation method. Firstly, generate independent and identic-distributed random variables, then solve the conditional distribution based on the hierarchical relationship of the tree structure, and finally obtain the principal component samples through the inverse transformation of the edge distribution. The core equation is as follows:
w ~ u [ ε , 1 ε ]
u j c o n d = F 1 ( w u i ; θ )
Z n e w = F j 1 ( U n e w )
where ε is the boundary tolerance parameter ( ε = 0 .005 in this study), u j c o n d is the conditional distribution solved by tree structure hierarchy, and Z n e w is the principal component sample.
Step 4: Reconstruct and generate new scenarios. Reconstruct the original output scenario through the inverse transformation of the Kernel PCA:
X n e w = ( Z n e w V T ) d i a g ( σ ) + M
where X n e w is the reconstructed scenario matrix, V is the kernel PCA load matrix, d i a g ( σ ) is the standard deviation diagonal matrix, and σ j is the standard deviation of the j-th characteristic. M is the mean matrix, and each row has the mean μ . This reconstruction method employs the mean and standard deviation calculated in the original standardization step (Equation (18)). This step aims to ensure the stability of numerical calculations and does not assume that the data have Gaussian distribution characteristics. The non-Gaussian features in the original data are retained in the joint dependency structure—a structure described by the Vine Copula model. Meanwhile, the marginal distribution of the principal component variables is also modeled using non-parametric methods (Equation (24)). Therefore, regardless of the distribution pattern of the original data, this equation can accurately reproduce the actual scale of the data and its variation characteristics.
The output generated through the above steps is the reconstructed typical daily scenario matrix of hydro–wind–solar energy combined. Each row of the matrix represents a generated typical daily scenario, including the hourly runoff scenario, the wind power output scenario, and the solar output scenario within that day. By implementing the above scenario generation process for all 24 solar terms respectively, a complete typical daily scenario library of hydro–wind–solar correlation covering all climate stages throughout the year was ultimately constructed. The generated typical daily scenario library provides key input for the long-term planning and operation research of hydro–wind–solar hybrid systems.

3.4. Scenario Quality Evaluation System

After generating the hydro–wind–solar energy scenarios using Kernel PCA and Vine Copula, a multi-dimensional evaluation system is needed to verify its physical consistency and statistical reliability. The existing indicators usually focus on spatiotemporal correlations or cumulative and probability distributions between scenarios, but rarely capture curve shapes and feature positions, possibly ignoring the structural differences of the original data. To address this issue, this study proposes an evaluation system that combines a comprehensive similarity indicator with a spatiotemporal correlation indicator to evaluate the quality of scenarios more comprehensively.
(1)
Comprehensive similarity distance (CSD)
The CSDs defined in Section 2.1.1 were adopted to quantify the differences between the generated scenarios and the historical scenarios from three dimensions: similarity of power amplitude, similarity of morphological trend, and similarity of fluctuation position.
(2)
Temporal correlation indicator
An autocorrelation coefficient (ACF) is used as an indicator of time correlation to evaluate the simulation performance of the generated hydro–wind–solar power scenario set for time series characteristics [48].
A C F ( k ) = t = 1 T k ( X t X ¯ ) ( X t + k l a g X ¯ ) t = 1 T ( X t X ¯ ) 2
where ACF(k) represents the autocorrelation coefficient of the time series at lag k. k is the lag time, that is, the time difference corresponding to the autocorrelation coefficient calculated. X t is the value of the time series at time t, X ¯ is the mean of the time series, and T is the total number of time periods.
(3)
Spatial correlation indicator
In this study, the absolute error of the average Kendall correlation coefficient is used as the evaluation indicator of spatial correlation. The average absolute error of the spatial dependence structure between the generated scenarios and the historical data scenarios is calculated to systematically evaluate the quality of the spatial correlation characteristics of multi-energy joint output scenarios.
e = 1 W w = 1 W τ w τ ¯
where W is the number of scenarios, τ w is the Kendall correlation coefficient between the generated group of scenarios, and τ ~ is the Kendall correlation coefficient between the historical time series data.
(4)
Net power volatility indicator
To evaluate the quality of the generated scenarios from the perspective of system operation, this article proposes the net power volatility indicator of the system. This indicator quantifies the intensity of the time series variation of the total output of the hydro–wind–solar hybrid system, directly reflecting the net power fluctuation characteristics presented by the system externally. It is a key feature that has a direct impact on the dispatching of the power grid and the demand for reserve capacity.
P total ( t ) = P h ( t ) + P w ( t ) + P s ( t )
R hour l y = 1 T 1 t = 2 T P t o t a l ( t ) P t o t a l ( t 1 )
R range = max t ( P t o t a l ( t ) ) min t ( P t o t a l ( t ) )
P h , P w , and P s represent the sunrise power scenarios of hydro–wind–solar energy, respectively. P t o t a l is defined as the total system power at time t. R h o u r l y represents the hourly volatility, specifically the average value of the absolute value of the total power change in adjacent hours. R r a n g e represents the average daily fluctuation range, defined as the difference between the maximum total power and the minimum total power within a day. The volatility error can be calculated based on the following equation:
ε h = R h , g e n R h , o r i g / R h , o r i g
ε r = R r , g e n R r , o r i g / R r , o r i g
ε c = ( ε h + ε r ) / 2
where ε h represents the relative error of hourly fluctuations, ε r represents the relative error of the daily fluctuation range, and ε c represents the volatility error. gen and orig represent the generated scenario and the original scenario, respectively.
The flowchart of the scenario generation method for the interrelationship among hydro-wind-solar energy is shown in Figure 6:

4. Case Study

4.1. Study Area and Data Description

4.1.1. Basin Characteristics and Station Layout

This study takes the Yalong River Basin in Sichuan Province as the research object. This basin is rich in wind and solar energy resources, which complement hydropower generation. As most of the wind and solar energy facilities in this basin are still in the planning stage, the scarcity of historical meteorological data poses a significant challenge to the modeling of integrated hydro–wind–solar energy. Therefore, in this study, a multi-source data-driven scenario generation method was developed using publicly available datasets to support the optimization of integrated hydro–wind–solar systems.
To address data limitations, a virtual location selection method based on grid resource sampling and multi-source spatial consistency verification was adopted. Sites were chosen within a 60 km range of hydropower stations in the middle and lower reaches of the basin to enhance spatial correlation. A total of seven hydropower stations were selected, including two with annual regulation capacity, one with seasonal regulation capacity, and four with daily regulation capacity. Around these, 12 virtual wind farms and 13 virtual solar power stations were established (Figure 7). The installed capacity was allocated based on a widely adopted hydropower-to-renewables ratio of 1:1, which simplifies system design while ensuring operational coordination. Specifically, the total hydropower capacity is 16,000 MW, with 8000 MW each allocated to virtual wind and solar power stations. The total wind and solar capacity of 16,000 MW enables coordinated operation and maximizes renewable utilization.

4.1.2. Multi-Source Data Acquisition and Preprocessing

This study integrated wind and solar energy data from multiple platforms and sources to construct a long sequence dataset covering a 15-year hourly resolution from 2008 to 2022. The data sources and main parameter specifications are as follows:
Solar power data:
Data source: NASA POWER [49], the European Centre for Medium-Range Weather Forecasts (ECMWF) [50], the Photovoltaic Geographic Information System (PVGIS) [51], Meteonorm [52], and the National Meteorological Information Center (NMIC) [53].
Data type: Hourly downward shortwave radiation intensity, 2 m air temperature and solar elevation angle from 2008 to 2022. Among them, only the hourly measured shortwave radiation intensity data of the NMIC in 2022 was obtained.
Wind power data:
Data source: NASA POWER, the Greenwich Platform [54], the ECMWF, and the NMIC.
Data type: Hourly wind speed and altitude from 2008 to 2022. Among them, only the hourly wind speed data of the NMIC in 2022 was obtained.
In view of the differences of the multi-source data, this study conducted preprocessing on the data. To do so, firstly, uniformly convert to the UTC + 8 time zone and eliminate leap days to ensure strict alignment of the time series. Secondly, based on the longitude and latitude coordinates of the planned stations, the nearest neighbor grid data is extracted to ensure spatial consistency. Finally, convert all different file formats to Excel format uniformly. All the NetCDF files were exported in CSV format through the official Panoply Data Viewer (4.12.11.0) provided by NASA and then integrated into an Excel workbook. The runoff datum adopts the measured data of the study basin. This study converts wind speed data and solar data into the output power of renewable energy based on the standardized power generation equation. All programs were developed using MATLAB R2024a.

4.2. Reliability Assessment Results of Multi-Source Data

This section adopts a two-stage verification method to evaluate the reliability of the multi-source reanalysis data. Firstly, a consistency test was conducted on the 15-year multi-source dataset to verify the long-term distribution pattern and determine the most reliable data source. Then, cross-validation was conducted using the one-year measured data of the Yalong River Basin and multi-source datasets to screen out data sources that conformed to the short-term fluctuation pattern. Finally, the most reliable data source was comprehensively evaluated based on the results of the two-stage verification.

4.2.1. The Principle of Time Window Division

The purpose of reasonable time division is to embed the inherent climate periodicity into the data-driven scenario generation process. To achieve this goal, a segmented scheme that aligns with the key physical drivers of renewable energy in terms of time is needed, rather than using a fixed administrative calendar (such as months and quarters). The latter often introduces bias due to the inability to capture the intra-seasonal mutations that determine the output mode. The 24 solar terms are a traditional time division system in China, reflecting the changes in the position of the sun and the monsoon climate. Taking the 24 solar terms as the analysis range, the influence of seasonal and climatic factors can be incorporated and applied to the output characteristics of wind and solar energy. This effectively reflects the output patterns and variation characteristics of the landscape at different times. This study fully utilized the inherent climatic regularity contained in the solar terms. Each solar term (approximately 15 days) corresponds to a period with similar astronomical characteristics, including key parameters such as the solar altitude angle and sunshine duration. Thus, in a short-term study, a deep coupling of astronomical driving laws and climatic response mechanisms was achieved.
The limitations of month, season, and ten-day divisions are as follows. ① The traditional division of months and quarters (30–90 days) cannot capture the sudden changes in the mid-month monsoon, and the generated scenarios may deviate from the actual monsoon patterns. ② The ten-day division lacks an astronomical driving basis and may overlook key climate events. The division of solar terms can better solve these problems. The outstanding performance of typical solar terms in capturing relevance and generating scenarios in Section 4.3 once again confirms the advantages of the solar term division mechanism: it not only ensures the consistency of time segments with the natural climate rhythm, but also significantly enhances the physical authenticity of the renewable energy output scenarios.
The principle of time division based on climatic physical processes is universal. Although the type of solar terms selected in this article is customized for the characteristics of the East Asian monsoon climate, its core principle: the periodic division based on the astronomy–climate coupling mechanism has universal applicability. Researchers in other regions can adapt to local conditions and adopt climate cycle division methods that conform to local seasonal characteristics.① For instance, in the tropical monsoon region of India, the monsoon outbreak day is used instead of the solar term to capture the pattern of double peak precipitation. ② The construction of the solar thermal–snow cover coupling cycle (snowmelt period/polar day period/dark period) in the Nordic frigid zone is adapted to the solar fluctuation characteristics at high latitudes. The value of this scheme lies in providing a standard process for embedding such physical prior information. Subsequent case studies will demonstrate how to generate high-quality scenarios guided by this principle and using the division method of the 24 solar terms.

4.2.2. Consistency Assessment of Multi-Source Data

This section conducts a long-term consistency assessment of the distribution characteristics of multi-source data. To do so, extract the typical daily curves of each solar term from the different data sources and calculate the average of the CSDs of the typical daily curves among the different data sources. The input data includes 15-year historical power generation data from two of the best virtual stations (W3/S7) of wind and solar resources, multiple reanalysis data sources, covering 24 solar terms. Each solar term contains 225 daily power generation curves (each solar term has 15 days, that is, 15 years × 15 days = 225). Each “typical sunrise power curve” is composed of continuous power values over 24 h and thus contains 24 data points. This is consistent with the original data at hourly resolution adopted in this study. The total number of data points in the original power time series is of the order of 5475 × 24 = 131,400. The results of the consistency assessment of the long-term distribution characteristics are shown in Table 1.
The stability score in the table is obtained by adding up the CSDs of each website and then taking the reciprocal.
Wind data sources: As shown in the CSD comparison results of the wind power data source in Table 1, NASA performs relatively well among the wind data sources. Its CSD from Greenwich is 0.062 (the lowest value), while its CSD from the ECMWF is 0.097, which is lower than the 0.099 between Greenwich and the ECMWF. NASA surpassed the other two data sources with a stability score of 12.578.
Solar data sources: It can be seen from the CSD comparison results of solar power data sources in Table 1 that NASA has significant advantages. The CSD values of NASA and all other data sources do not exceed 0.008, and the minimum average CSD is only 0.006. Its stability score is as high as 142.857, which is 42.6% higher than that of the PVGIS (103.039), which ranks second.
Overall, whether it is wind data sources or solar data sources, NASA performed the best in the long-term multi-source consistency assessment. This method effectively solves the problem of the long-term reliability assessment of data sources in data-scarce river basins and provides a reliable data screening idea for the situation where measured data is seriously insufficient. The consistency assessment results provide a preliminary basis for determining reliable scenario input data.

4.2.3. Cross-Validation and Weighted Evaluation Based on Measured Data

The consistency assessment results in Section 4.2.2 prove that NASA data have the highest credibility among numerous open-source data, but this assessment result still needs to be verified by actual measurement data. This section takes the measured data of 2022 and multi-source data of the same year as input. For the optimal virtual stations (W3/S7) of wind power and solar energy resources, two different strategies are adopted to cross-verify the short-term fluctuation patterns and numerical accuracy of the data sources: Strategy 1 is to extract the cluster centers of each multi-source datum and the measured data, and then compare and analyze the distribution characteristics of their cluster center curves and CSD values. Strategy 2 is to conduct point-to-point error calculations on each multi-source datum and the measured data, and then compare and analyze their error indicator values (the error indicators are shown in Section 2.2). Ultimately, the multi-source data were comprehensively evaluated based on all indicators and the comprehensive score was calculated.
(1)
Consistency verification based on measured data
To overcome the challenge of lacking measured data for long-term consistency assessment, this section employs one year of measured data and validates the results using a single-clustering-based global consistency evaluation strategy [55]. Based on the solar terms, the number of clusters (K) is set to one to extract cluster centers from both the multi-source and measured data. The distribution characteristics of the cluster center curves are then compared across data sources and against the measured data. The CSD and its annual average between the clustering results of each data source and the measured data are also computed. Figure 8 shows the clustering results of the wind and solar data sources during the Beginning of Spring in 2022.
Figure 8 describes the differences in the shapes of the daily variation curves of different wind–solar data sources and the corresponding measured data, where the bolded curves represent the centers of clusters. The red NMIC represents the benchmark data. The assessment of wind power and solar power data quality are analyzed respectively below.
Wind power data: ① Analysis of distribution characteristics. As shown in Figure 8a, the clustering results of different wind data websites are presented. The clustering center curves of NASA and Greenwich align closely with the measured data in terms of fluctuation patterns, with particularly high consistency in the position and amplitude of nighttime peaks. In contrast, the ECMWF exhibits a significant deviation from the NMIC during the morning hours. Overall, NASA demonstrates superior performance in distribution characteristics. ② Comprehensive similarity distance analysis. The CSD values between each data source and the measured clustering centers were calculated quantitatively in Figure 8a. NASA achieved the lowest CSD values across all 16 solar terms, with an average of only 0.397, indicating the highest similarity to the reference data. Although the ECMWF performed well in a limited number of solar terms, its average CSD was higher than that of NASA and showed relatively unstable results. Greenwich only matches the measured data in a few solar terms, and the CSD value is significantly higher and fluctuates sharply in most periods, indicating that its reliability is obviously insufficient.
Solar power data: ① Analysis of distribution characteristics. As shown in Figure 8b, it is the clustering result of the solar data website. The clustering centers of NASA, the ECMWF, and the PVGIS show high consistency with the measured data during peak sunlight hours, such as noon. Among these, NASA exhibits fluctuation characteristics most similar to the NMIC. In contrast, Meteonorm displays significant deviation during dusk periods. Overall, NASA demonstrates the closest alignment with the actual measured values in terms of distribution characteristics. ② Comprehensive similarity distance analysis. The CSD values between each data source and the measured cluster centers were computed in Figure 8b to quantitatively evaluate similarity. NASA achieves the lowest CSD values over 11 solar terms, with an annual average of only 0.082, confirming its highest agreement with the NMIC. Both the ECMWF and PVGIS show considerably higher annual average CSD values and poorer stability. Notably, Meteonorm performs exceptionally well under extreme weather conditions—for example, during the Great Heat period, its CSD drops to 0.018. However, such low values occur in only six solar terms, indicating limited consistency across the entire year.
Overall, whether it is wind energy or solar energy, NASA has performed the best. The low value and stability of the CSD fully demonstrate that NASA is highly consistent with the measured data in terms of the similarity of distribution characteristics. This method verified the accuracy of the evaluation results in Section 4.2.2 and the effectiveness of the method through limited measured data. The detailed calculation results of the CSDs for each solar term in 2022 are shown in Appendix A (Table A1).
(2)
Analysis of Error Indicators
In this section, by calculating the error indicator values between the measured data of 2022 and the multi-source data, and combining the CSD results, a comprehensive evaluation indicator S was constructed (Section 2.3). The results are shown in Table 2, and the distribution of each indicator is detailed in Figure 9. The detailed calculation results of each solar term indicator in 2022 are shown in Appendix A (Table A2, Table A3 and Table A4).
Wind data sources: ① An analysis of the results of each indicator of wind was performed, as shown in Table 2, and the wind data sources received the following scores: Greenwich (0.350), ECMWF (0.397), and NASA (0.413). NASA achieved the highest score, indicating its overall advantage in wind data performance. ② To further analyze the strengths of each data source, Figure 9a displays the radar chart of performance across various metrics. It shows that NASA excels particularly in having the smallest root mean square error and the strongest linear correlation with measured data. While the ECMWF also performs well in mean bias error and nonlinear correlation, it does not surpass NASA significantly in these aspects.
Solar data sources: ① It can be known from the results of each indicator of solar energy in Table 2 that NASA and Meteonorm obtained the top two comprehensive scores, with NASA at 0.551 and Meteonorm at 0.546—a difference of only 0.005. ② A detailed comparison via the radar chart in Figure 9b reveals that NASA is slightly weaker in short-term indicators such as root mean square error and mean bias error. However, it outperforms Meteonorm across all other indicators, demonstrating broader consistency and reliability.
The reliability, accuracy, and bias of research data are fundamental considerations in energy applications. Based on comprehensive indicator performance, NASA demonstrates the closest agreement with the measured values, making it the most suitable data source for both wind and solar energy applications. The ECMWF and Meteonorm represent optimal alternative options for wind and solar data, respectively. This section confirms the effectiveness of the multi-source data evaluation framework introduced in Section 2, which supplies statistically reliable and physically consistent input data for subsequent high-dimensional scenario generation. Detailed results of the short-term error indicators are provided in Appendix A (Figure A1 and Figure A2).

4.3. Scenario Generation Performance

After determining reliable data sources, the next step is to verify the effectiveness of the two-stage dimensionality reduction and correlation modeling method. This section will present the performance of scenario generation, including the results of dimensionality reduction, correlation capture, and the accuracy of scenario reconstruction.

4.3.1. Dimension Reduction Effect Verification

To verify the effectiveness of the first-stage dimensionality reduction strategy (Section 3.1), this study processed NASA’s historical wind and solar energy data, reducing the time dimension of each virtual station from 5475 days (365 × 15) to 225 days. It selected representative solar terms based on climatic characteristics: The representative solar terms for wind power (monsoon transition and the peak period of strong winds) are Fresh Green, First Frost, and Great Cold. The representative solar terms for solar power (with clear skies and high radiation before and after the rainy season) are Spring Equinox, Lesser Fullness, and White Dew. The retained typical daily scenarios are crucial for a complementarity analysis, as shown in Figure 10 and Figure 11. For detailed clustering results, please refer to Appendix B (Figure A3 and Figure A4).
CSD and labels (in the upper right corner) represent the clustering tightness. The closer the value is to 0, the more concentrated the intra-class curve is and the higher the centroid representativeness. The analysis of the clustering results is as follows:
Wind power clustering results: It can be seen from Figure 10 that Fresh Green, First Frost, and Great Cold have the typical features of high and large fluctuations in wind power during spring and winter, which is highly consistent with the monsoon climate characteristics of this basin. The clustering result C1 of wind power generation is significantly higher because the wind speed is too low or even close to 0 on some days during the solar term. These patterns are real but atypical operating states; in the subsequent Kernel PCA dimensionality reduction process, their corresponding features will mainly be distributed in the principal components with lower interpretation variances and are usually not selected for scenario generation. Except for C1, the sum of the CSD within each cluster is close to 0, indicating a good clustering effect.
Solar power clustering results: Figure 11 shows the clustering results of the solar output. The peak data of the three typical solar terms, namely, the Spring Equinox, Lesser Fullness, and White Dew, are relatively high and concentrated, which conforms to the natural laws before and after the rainy season. The total CSD within each cluster is basically maintained between 0 and 3. The closer the value is to 0, the more effectively the CSD can aggregate curves with similar shapes.
After dividing and clustering the original data by solar terms, the time dimension was significantly compressed, and 50 to 80 typical daily datasets were obtained for each solar term. This method verifies the effectiveness of the CSD-improved clustering algorithm in capturing fluctuation patterns and reduces the complexity of the original high-dimensional time series data. It provides high-correlation and low-dimensional input for the subsequent Kernel PCA. This process also divides some extreme events into independent clusters and converts them into higher-order principal components with a lower variance interpretation rate in the subsequent Kernel PCA analysis. When generating scenarios based on principal components (Section 4.3.3), some extreme events will be covered.

4.3.2. Principal Component Loading and Correlation Verification

To analyze the nonlinear coupling characteristics proposed in this study (Section 3.2.2), this section takes the three representative solar terms of the Spring Equinox, Lesser Heat, and White Dew to deeply analyze the physical significance of the principal component load and the collaborative mechanism of the hydro–wind–solar correlation model. The principal component load is shown in Figure 12.
The subgraphs at the lower right corner of each chart contain the following: the vertical axis on the left represents the variance interpretation rate of each principal component (Blue bar chart), while the horizontal axis on the right shows the cumulative variance rate (Red dot line graph). Select the top 15 principal components with a cumulative variance exceeding 85% to ensure a balance between the research dimension and the correlation. The vertical axis of the main graph represents the contribution of each energy type to its principal component, which can be extended to understand the correlation among different energy types.
To analyze the principal component load, it is necessary to analyze the contributions of each principal component. In terms of contribution, the first three main components are of vital importance. ① The contribution value of PC1 in Figure 12a shows that wind power dominates (with a contribution value of approximately 0.75), while the contribution value of solar power is much lower than that of runoff and wind power. Runoff, a sharp decrease in wind power and solar power, has become a positive correlation, representing a windless and cloudy weather pattern. ② The significant increase in PC2 runoff and the negative correlation with the decrease in wind and solar energy clearly indicate that rainy days are approaching and other weather patterns are captured. ③ PC3 reflects a situation where solar energy is severely suppressed, wind energy almost disappears, and runoff decreases, that is, the characteristics of cloudy weather.
An analysis of typical weather patterns is equally important. Based on the law in Figure 12a, combined with the principal component loads in Figure 12b,c, the complex nonlinear correlation characteristics presented by the hydro–wind–solar system can be summarized. The main pattern is as follows: ① The runoff volume is positively correlated with significantly weakened wind and solar energy. ② The increase in runoff volume is negatively correlated with the attenuation of wind and solar energy. ③ The sudden increase in wind force forms a confrontation with the suppression of hydropower and solar energy. ④ Solar energy has close to 0 growth, wind power is gradually disappearing, and the runoff volume continues to increase.
This section provides a physical interpretation basis for the Vine Copula model in Section 3.3 by quantifying the nonlinear correlation of the hydro–wind–solar system. At the same time, the effectiveness and innovation of the method described in Section 1.2 were verified, and it was proved that the Kernel PCA method does not have a black box problem, and its steps and mechanisms have clear physical interpretability.

4.3.3. Scenario Generation Result Analysis

In order to accurately describe the nonlinear dependency structure among principal components after dimensionality reduction and generate scenarios, this study adopts the Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC) criteria to screen the optimal Copula function and implements the simplified Vine Copula method for scenario generation. The average AIC/BIC value calculation results of each solar term are shown in Table 3.
As shown in Table 3, the systematic assessment based on the AIC/BIC criteria indicates that t-Copula demonstrates significant advantages in characterizing the combined distribution of hydropower, wind power, and solar output. Its average AIC value of −28.221 and average BIC value of −23.797 are both much lower than those of other candidate Copula.
This study generates scenarios based on the division of solar terms. During the Winter Solstice and Great Cold, solar output and runoff is at a trough level, while wind power reaches its annual peak. The Summer Solstice and Great Heat correspond to peak solar output, with increased variability in wind power coinciding with high hydropower demand. At the Spring Equinox and Autumn Equinox, the correlation between wind and solar energy reverses. To balance the typicality and conciseness of the research results, this study selects four typical solar terms: the Spring Equinox, the Summer Solstice, the Autumnal Equinox, and the Great Cold. The solar terms selected in this plan cover the entire hydrological cycle and key energy states throughout the year.
Based on the Kernel PCA dimensionality reduction results and Vine Copula reconstruction, Figure 13 presents the hydro–wind–solar joint correlation scenarios of four typical solar terms. The red line represents the average value of historical scenarios, and the black line represents the average value of generated scenarios. The colored lines represent the different scenarios generated.
① During the rainy season solar terms, the generated wind–solar output scenarios can reflect typical temporal characteristics. For instance, during the Summer Solstice, the wind power output curve shows a distinct peak from 18:00 to 24:00 at night. This is basically consistent with the fluctuation trend of the average curve in the original scenario. The solar energy curve reaches a single-peak between 12:00 and 13:00 in the afternoon, with a very small deviation from the average value of the original scenario. The above results fully demonstrate that the generated scenarios have high accuracy in depicting the collaborative characteristics of the rainy season’s wind and solar energy.
② During the Great Cold period, the generated runoff–wind–solar scenarios are generally consistent with the historical scenarios. It effectively reproduced the operational characteristics of the dry season with hydropower as the base load. Specifically, the average runoff volume is between 500 and 800 cubic meters per second. The fluctuation of wind power output in the morning from 04:00 to 08:00 is consistent with the historical variation characteristics. The solar energy curve was approximately 5% lower than the historical value during the morning rush hour from 11:00 to 13:00. This model accurately reflects the typical characteristics of operation based on hydro power during the dry season.
The deviations in the simulation of solar energy scenarios under drought conditions cannot be ignored. To assess the impact of this deviation in actual operation, this article conducted a sensitivity analysis based on the indicators in Section 3.4. Detailed results are in Appendix B (Table A5 and Table A6). During this period, if the solar power generation decreases by 5% and 10%, respectively, from 11:00 to 13:00 in the dry season (such as the Winter Solstice and the Great Cold), the volatility error ( ε c ) of the hybrid power system will increase by approximately 3.74% to 4.31%, while the other indicators remain stable. This increase in volatility means that during drought periods, the system requires more regulation capacity and reserve capacity to ensure its reliability. Therefore, when formulating response plans for drought periods, it is necessary to fully consider such potential systemic deviations to ensure the stable operation of the system.
③ During the monsoon transition period (the Spring Equinox and the Autumnal equinox), solar power shows a high degree of stability, while wind power exhibits significant volatility. Taking the scenarios of the Spring Equinox and the Autumnal Equinox as examples, the overlap degree of solar energy output during the morning rush hour (5:00–18:00) with the historical average curve is as high as 96%. The trough value of wind energy deviates significantly from 8:00 to 13:00. This comparison result indicates that during the period of seasonal transition, the output of solar power can remain stable, while the output of wind power is prone to significant fluctuations.
In this study, the normalized root mean square error (NRMSE) was obtained by using normalization calculation for Equation (9). The hydro–wind–solar combined scenario successfully captured the unique climatic characteristics of each period. The NRMSEs of all the scenarios were less than 3%, and for key features such as peak and valley values and inflection points, this error was controlled within 5%. This result demonstrates the following: Firstly, the unique climatic characteristics embedded in the division of solar terms restore the complex statistical distribution of the original hydro–wind–solar energy scenario. Secondly, this study confirmed the feasibility and reliability of the Kernel PCA + Vine Copula method in generating comprehensive scenarios with both physical consistency and statistical representativeness in data-scarce river basins.

4.3.4. Generate Scenario Application Analysis

The hydro–wind–solar combined power generation scenario generation method proposed in this study has high accuracy and resolution. Its core advantages lie in the following: ① It can accurately extract the key features and change patterns in historical data. ② It can provide quantitative data support for the calculation of core performance indicators of complementary systems. The effectiveness and practicality of the generated scenario are mainly reflected in the following two aspects.
The verification of physical consistency and statistical reliability: ① The generated scenarios strictly follow the climate cycle laws of the Yalong River Basin. The scenario reproduces the amplitude of output fluctuations, morphological trends, and complementary temporal relationships in the actual system. It has been proved that this method is reliable in characterizing the operational characteristics of the actual system. ② The generated scenarios provide a verified input data basis for quantifying indicators such as the distribution of peak shaving pressure and the fluctuation of wind and solar curtailment rates under different hydrological years. This helps to enhance the accuracy of subsequent planning and scheduling research.
Its supporting role in system analysis and planning design. ① The nonlinear coupling mode of hydro–wind–solar output at different solar terms is clearly presented in the scenario (Section 4.3.2), providing a quantitative basis for the optimization of the wind and solar capacity ratio based on the stable base load of hydropower. ② The generated scenario captures the periods of sharp fluctuations in wind and solar output, which can be used to analyze the power gap of the system and deduce the power or capacity of the energy storage configuration. ③ The monsoon transition characteristics contained in the scenario can provide a testing basis for designing dynamic scheduling rules that adapt to climate change.

4.4. Comparison of Clustering Algorithms and Scenario Generation Methods

(1)
Comparison of clustering algorithms
To verify the advantages of the improved clustering algorithm proposed in this study, a comparative analysis was conducted between the proposed method and the traditional K-means algorithm. The comparison results are presented in Figure 14. In this evaluation, three indicators were employed: The Silhouette Coefficient is used to measure clustering quality, and the higher the value, the better the performance. The Davies–Bouldin (DB) Index is also used, wherein lower values reflect higher within-cluster similarity and greater inter-cluster separation, and as is a CSD indicator. To facilitate consistent interpretation, both the DB Index and the CSD were positively transformed during standardization—thus, for all indicators, larger values correspond to better algorithm performance.
As observed from the line chart, the improved method demonstrated significantly superior performance in all three indicators. In some solar terms, the indicators are particularly prominent (for example, the seventh solar term, Beginning of Summer). The average performance of the indicators of the improved algorithm further confirmed the effectiveness of the improvement: the CSD indicator decreased by 4.6%, the contour coefficient increased by 2.3% and the DB index decreased by 4.5%. These enhancements can be attributed to the algorithm’s integrated consideration of multiple features including power amplitude, pattern trend, and fluctuation position characteristics. By comprehensively capturing these aspects, the improved algorithm more accurately describes the distribution characteristics of complex renewable energy curves, thereby assigning each curve to the most appropriate cluster. By integrating the improvement effects of the three indicators, the traditional clustering accuracy was ultimately increased by 3.8%.
(2)
Comparison of scenario generation methods
To verify the comprehensive performance of the scenario generation method proposed in this study, two typical comparison schemes are set up in this section for systematic comparison. The scenario quality assessment system constructed based on Section 3.4 analyzed the three schemes. The comparison scheme is as follows.
Scheme 1 is an improved K-means clustering method + PCA + Copula. This scheme inherits the PCA–Copula model [23] and extends the two-dimensional inputs of wind and solar energy to a three-dimensional coupled system of hydro–wind–solar energy. Scheme 1 serves as the comparison benchmark for the proposed schemes in this study. Scheme 2 is an improved K-means clustering method + Kernel PCA + Vine Copula, which is the scenario generation method proposed in this study. Scheme 3 is a traditional K-means clustering method + Kernel PCA + Vine Copula. The comparison and analysis results of the scenarios generated by different schemes for the six typical solar terms with the original scenarios are shown in Table 4.
Based on the analysis of the metrics in Table 4, the following conclusions can be drawn:
① Analysis of the CSD results. The average CSD values for Scheme 1 (Benchmark), Scheme 2, and Scheme 3 are 0.849, 0.593, and 0.613, respectively. Compared to the benchmark, Scheme 2 shows a 30.2% improvement, while Scheme 3 shows a 27.8% improvement.
② Analysis of the ACF results. The average ACF values are 0.105, 0.031, and 0.093, respectively. Scheme 2 demonstrates a 70.5% improvement over Scheme 1, and Scheme 3 shows an 11.4% improvement.
③ Analysis of the e results. The e values are 0.235, 0.229, and 0.182, respectively. This corresponds to a 2.6% improvement for Scheme 2 and a 22.6% improvement for Scheme 3 compared to the benchmark.
④ Analysis of the ε c results. The values of ε c in the three schemes are 2.651, 0.696, and 0.529, respectively. The volatility errors of both Scheme 2 and Scheme 3 are smaller than those of Scheme 1, improving by 73.7% and 80%, respectively, compared to Scheme 1.
After comparing all the schemes, it can be found that Scheme 2 performs best in terms of comprehensive similarity distance (CSD) and temporal correlation (ACF). The scenario generated by Scheme 2 is the closest to the historical reality in terms of overall form and change pattern, while inheriting the inherent time-dependent structure of the historical sequence. In terms of spatial correlation (e) and volatility error ( ε c ), although Scheme 2 has improved compared to the benchmark scheme, it is inferior to Scheme 3 in terms of the degree of advantage. This phenomenon precisely indicates that the traditional K-means algorithm is based on Euclidean distance and is extremely sensitive to numerical outliers (extreme output days). It will cluster the highly volatile extreme days separately and retain them as typical patterns. Therefore, Scheme 3 is more accurate in replicating the extreme spatial distribution and fluctuation characteristics in historical data, which directly leads to better values of its e and ε c indicators. This precise replication of extreme values is a double-edged sword. Although it can capture extreme scenarios and spatial distributions, it comes at the cost of sacrificing the morphological fidelity and temporal structure restoration of mainstream and typical patterns.
Taking into account all the indicators and their physical significance comprehensively, Scheme 2 pays more attention to capturing the overall morphological change trend and temporal structure rather than deliberately replicating all the details of spatial correlation. Therefore, it remains the best and most balanced solution for generating high-fidelity combined scenarios of hydro–wind–solar power.
(3)
Analysis of algorithm scalability
The codes of all the schemes in this article were run on the MATLAB R2024a platform. The computing times of the three schemes were 1620 s, 1487 s, and 1380 s, respectively. The running time of a PCA is longer than that of a Kernel PCA, while the traditional clustering method is faster than the improved clustering method. Taking Scheme 2 as an example, when ten more stations were added, the operation time of this method was 186 s, an increase of 374 s (25.2%). When 20 more stations were added, the operation time reached 2493 s, an increase of 632 s, which was 1006 s (67.7%) longer than the initial operation time of Scheme 2. This result indicates that as the number of stations increases, the computational complexity of the algorithm in this study also rises. The increase in complexity is not linear because the scenario is generated based on principal components, and each time new stations are added, the number of principal components required will increase.
Specifically, in the original method, the first 15 principal components are required to reflect the complex coupling relationship. After the number of stations increases, the first 20 or even the first 30 principal components are needed. This will directly lead to a linear increase in the time for extracting principal components, and the computational complexity of Copula shows a combinatorial increase, which in turn affects the computational complexity of the entire algorithm. Based on this, it can be inferred that when the number of stations increases by 50 or more, more principal components will be required, and the computational complexity will rapidly reach an unbearable level. This scale might be the maximum that the algorithm can handle.

4.5. Sensitivity Analysis of the Solar Term Division Generation Scenario

This section designs the misalignment sensitivity experiment of the system. The core purpose is to quantitatively assess the impact on the fidelity of the scenarios generated by this framework when the preset astronomical/climatic segments (the 24 solar terms) deviate from the real meteorological processes on the time axis.
In this study, the actual climate was offset by 7 days and 15 days compared to the solar term division standard. Its actual meaning is to assume that the actual climate model is half a solar term to one solar term earlier or later than the solar term division model. Based on this, a sensitivity analysis was conducted. All data were generated through scenario generation in Scheme 2, and the generated scenarios were compared with the original scenario to quantify various indicators. As the misalignment involves the influence of adjacent solar terms, this section uses the average values of all the indicators for the entire solar term for analysis. The results of the sensitivity analysis are shown in Table 5.
The sensitivity analysis indicates that the mismatch between the division of solar terms and the actual climate model generally leads to the deterioration of the morphological similarity (CSD) and temporal correlation (ACF) indicators of the generated scenarios. It is worth noting that in specific offset scenarios (such as 15 days in advance), the CSD value may be slightly lower than the standard scenario. This might be due to the misalignment accidentally grouping similar weather patterns into the same analysis window, but the overall trend indicates that the misalignment introduces additional uncertainty. The spatial correlation (e) and volatility error ( ε c ) fluctuated up and down in all cases, and there was no obvious tendency towards optimization or deterioration. The numerical changes of these two indicators are not significant, indicating that they have certain robustness in terms of spatial dependence and system fluctuation characteristics during misalignment.
The variation patterns of the above indicators can be explained by the data-driven modeling principle: Firstly, the general increase in ACF is due to the misalignment that disrupted the temporal coherence of weather processes within the original climate stage, making it difficult for the model to capture the precise phase-dependent temporal autoregressive structure. Secondly, the deterioration of the CSD is due to the misalignment that blurs the morphological boundaries of different climatic sub-models, causing the generated typical daily curves to deviate from the true best representatives in terms of amplitude, trend, and position. Relatively speaking, the stability of ε c and e suggests that the fluctuation intensity and spatial correlation of the system’s net output are less sensitive to the precise temporal alignment of the climate stage. As long as the distribution of weather types within the misaligned window is similar, its statistical fluctuation characteristics and spatial correlation can be roughly retained.
This case study shows that a time window of approximately 15 days defined by astronomical solar terms exhibits certain robustness against climate phase shifts within a half-month scale. However, this does not prove that the boundaries of solar terms remain valid under long-term climate change. If climate change causes continuous phase shifts in dominant weather systems such as monsoons, the long-term cumulative offset will eventually exceed the duration of a single solar term (15 days). At that time, continuing to use a fixed astronomical boundary will lead to a serious mixture of information from different climate mechanisms in the training data, fundamentally weakening the physical consistency and predictive ability of the model. Therefore, the results of this sensitivity analysis further prove the necessity of developing adaptive climate stage identification algorithms in future directions. A dynamic segmentation method capable of automatically identifying statistical characteristics from data is a fundamental approach to addressing the uncertainties of climate change.

5. Conclusions

This study developed an integrated data-driven framework to generate high-fidelity hydro–wind–solar complementary scenarios for basins with limited measured data. The key contributions include the following:
(1)
A multi-source data evaluation framework that strictly assesses short-term accuracy and long-term pattern consistency has established a new standard for reliable input selection in the context of data scarcity.
(2)
A comprehensive similarity distance (CSD) metric that integrates power amplitude, morphological trend, and fluctuation position, improving clustering accuracy by 3.8% and enhancing typical pattern extraction.
(3)
A novel two-stage dimensionality reduction method that seamlessly combines time series compression based on the solar period with feature fusion based on kernel PCA, successfully preserving the climate cycle while capturing key nonlinear dependencies across energy sources.
(4)
A Kernel principal component analysis (Kernel PCA) + Vine Copula scenario generation method that can generate physically consistent and statistically representative joint scenarios, with an NRMSE of less than 3%, as shown in the case study of the Yalong River Basin.
The framework provides a practical and scalable tool for supporting the planning, optimization, and operation of integrated renewable energy systems in regions lacking historical data. By deeply embedding climate characteristics into advanced data-driven workflows, this study offers a paradigm shift from traditional, typically physically decoupled approaches to a new generation of trusted and climate-conscious integrated research on renewable energy.

6. Limitations and Future Directions

The high-dimensional correlation scenario generation method proposed in this study has been verified for its effectiveness in the hydro–wind–solar complementary system of the Yalong River Basin, but it still has the following limitations.
Inherent reliance on the quality of reanalyzed data: NASA data performs best in 16 solar terms, but there are still deviations in radiation data during the dry season. This reflects that systematic errors in reanalysis data are still difficult to eliminate under extreme weather conditions, and measured data remain the first choice. This once again confirms that high-quality local measurements are indispensable for achieving the highest fidelity and points out that more powerful deviation correction techniques are needed in future work.
The scalability of high-dimensional modeling architecture: Kernel PCA already requires 15 principal components when preserving 85% variance. Further expansion of the system (such as incorporating more energy sources or more power generation stations) will intensify the dimensional curse and may lead to a combinatorial and rapid increase in the complexity of the algorithm. This article makes reasonable inferences about the maximum scale of the algorithm, but does not conduct more detailed tests. At the same time, due to the limited data obtained, the impact of the increase in energy types on the computational complexity of the algorithm could not be determined. In addition, the higher-order principal components with a lower variance interpretation rate cover extreme events, but this article does not explicitly quantify them.
The trade-off between the Copula simplification method and standard modeling accuracy: Although the t-Copula structure of the uniform type has achieved the optimal statistical sufficiency and is applicable to the vast majority of variable pairs, the simplified Vine Copula method adopted in this study still cannot rule out the possibility that a few extreme cases may affect the accuracy of modeling. Vine Copula with fully adaptive multivariate pairs was not achieved in this study due to its complexity.
Future recommendations are as follows:
Build an adaptive universal model: The unsupervised climate adaptive model can automatically identify the operation mode cycle from multi-dimensional climate-energy measured data, thereby achieving a fully adaptive climate division method. This will free the framework from reliance on any preset calendar, making it a truly globally applicable and highly climate-adaptable tool. Based on this model, the framework can be further expanded to integrate climate model predictions (such as CMIP6), using future climate sequences as the input to replace historical data, and capturing changing climate patterns based on adaptive climate models. Meanwhile, in clustering and dependency modeling, it can enhance the ability to identify and represent extreme samples. Ultimately, the framework will generate a library of future joint output scenarios corresponding to different scenarios, providing key inputs for the long-term climate adaptability planning of energy systems.
Intelligent algorithms integrating physical mechanisms: Exploring how generative AI (such as diffusion models and Gans) can further break through the bottleneck of high-dimensional nonlinear modeling by embedding physical constraints and fine climate features. This can ensure that the algorithm strictly adheres to the laws of physics and can cover more stations and energy types under reasonable computational complexity, thereby enhancing its broader applicability and credibility. In addition, in the future, if an automated, standard-based variable pair Copula function selection mechanism can be integrated within the standardized vine framework, it will be possible to achieve a more flexible balance between modeling accuracy and computational efficiency.
Realize the closed loop from scenario generation to application: Focus on the application value of scenarios, and deeply integrate scenarios into robust planning (resisting extreme weather) and real-time scheduling (model predictive control). Evolve this framework from a scenario provider to an end-to-end decision support system, quantifying how climate-driven uncertainties directly affect operational and investment strategies. In practical applications, to ensure the long-term validity of the model, it is recommended to use a rolling time window for annual updates. To avoid overfitting, at least three years of data are required when updating Kernel PCA, while at least five years of data are needed to stably estimate the dependency structure when updating the Vine Copula model. In this process, verifying the model’s extrapolation ability in adjacent periods is a valuable future test that can serve as an important benchmark for evaluating the robustness of its statistical learning. The generalization performance of the updated model can be evaluated through time cross-validation, thereby ensuring its ability to adapt to climate change and changes in the energy structure.

Author Contributions

P.G.: Software, Writing—Original Draft, Conceptualization. X.C.: Methodology, Writing—Review and Editing, Funding acquisition. W.M.: Supervision, Investigation, Data Curation. X.Z.: Validation, Supervision, Formal analysis. J.S.: Supervision, Data Curation. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original achievements presented in this study have been included in the main text of the article. For further information, please contact the corresponding author.

Conflicts of Interest

The authors declare no conflict of interest.

Appendix A. Detailed Indicator Result Tables and Graphs

Table A1 calculates the CSD values and their annual averages between the 2022 wind–solar multi-source data and the clustering results of the measured data. Table A2, Table A3 and Table A4 show the significant statistical indicators as RMSE, IOA, and R, respectively.
Table A1. The CSD values between the clustering results of the multi-source data and the measured data.
Table A1. The CSD values between the clustering results of the multi-source data and the measured data.
Solar TermsWind-CSDSolar-CSD
NASAECMWFGreenwichNASAECMWFPVGISMeteonorm
Beginning of Spring0.8910.7280.7370.0090.0280.0420.051
Rain Water0.5550.5830.5010.0120.0380.0530.041
Insects Awakening0.4260.5570.5150.0130.0420.0360.046
Spring Equinox0.3810.5280.4740.0190.0510.0370.052
Fresh Green0.4550.5610.5680.0110.0370.0360.054
Grain Rain0.4160.4910.8310.0560.0990.0990.092
Beginning of Summer0.6550.5960.7260.3310.3850.3470.186
Lesser Fullness0.5160.6960.4260.2130.2410.1920.110
Grain in Ear0.1990.1430.1080.0680.1060.0990.030
Summer Solstice0.1160.1530.2940.2510.2920.2830.222
Lesser Heat0.0720.1350.2250.1340.2110.2110.063
Great Heat0.0840.1930.2200.0410.0800.0520.018
Beginning of Autumn0.1210.1940.3920.0580.0330.0290.101
End of Heat0.1300.2150.2720.0310.0320.0190.070
White Dew0.1310.2170.1490.0380.0740.0730.025
Autumnal Equinox0.0520.1420.0850.0450.0870.0770.031
Cold Dew0.0930.2400.2030.3120.3190.2340.245
First Frost0.2120.2190.1390.0730.1210.1260.059
Beginning of Winter0.4850.4600.2880.1710.1350.1860.244
Light Snow0.5220.2190.1580.0080.0330.0650.050
Heavy Snow0.3600.4910.6150.0140.0470.0760.044
Winter Solstice1.0370.9021.4240.0130.0320.0420.061
Lesser Cold0.8990.9060.9120.0260.0160.0150.077
Great Cold0.7100.7820.6350.0250.0230.0260.075
Average Value0.3970.4310.4540.0820.1070.1020.085
Table A2. The RMSE values between the clustering results of the multi-source data and the measured data.
Table A2. The RMSE values between the clustering results of the multi-source data and the measured data.
Solar TermsWind-RMSESolar-RMSE
NASAECMWFGreenwichNASAECMWFPVGISMeteonorm
Beginning of Spring1.451.691.93165135257224
Rain Water1.462.292.83202182212190
Insects Awakening1.582.641.97196174252238
Spring Equinox1.422.182.48308306322228
Fresh Green1.151.791.5337308311311
Grain Rain1.131.591.35247278288336
Beginning of Summer0.991.061.26264264270261
Lesser Fullness1.532.011.78397412419469
Grain in Ear0.90.711.02385401401359
Summer Solstice0.770.851.49316299311294
Lesser Heat0.850.740.84250271283344
Great Heat0.660.550.64276278280237
Beginning of Autumn0.710.660.72193232229214
End of Heat0.660.51.36248243251275
White Dew0.620.730.86310312313359
Autumnal Equinox0.870.621.07325345329391
Cold Dew0.80.650.87381391395394
First Frost0.680.440.42266287279310
Beginning of Winter0.70.530.86291308396345
Light Snow0.720.430.27151156251186
Heavy Snow1.041.252.09172171305231
Winter Solstice0.870.50.8170198238219
Lesser Cold1.472.472.6149141223192
Great Cold1.282.031.96113131201195
Average Value1.011.211.37255259292283
Table A3. The IOA values between the clustering results of the multi-source data and the measured data.
Table A3. The IOA values between the clustering results of the multi-source data and the measured data.
Solar TermsWind-IOASolar-IOA
NASAECMWFGreenwichNASAECMWFPVGISMeteonorm
Beginning of Spring0.280.330.160.940.960.890.87
Rain Water0.420.370.230.870.890.850.89
Insects Awakening0.460.320.30.940.960.930.92
Spring Equinox0.470.280.260.860.850.840.91
Fresh Green0.460.360.220.820.850.860.84
Grain Rain0.330.330.220.860.840.840.77
Beginning of Summer0.230.330.170.40.380.370.41
Lesser Fullness0.430.420.370.680.660.640.61
Grain in Ear0.350.380.20.630.610.610.78
Summer Solstice0.340.290.220.520.490.480.58
Lesser Heat0.270.280.130.840.820.820.78
Great Heat0.240.290.150.810.810.820.88
Beginning of Autumn0.250.20.130.850.840.840.76
End of Heat0.220.190.130.830.850.840.79
White Dew0.310.170.160.770.770.780.74
Autumnal Equinox0.270.260.150.770.750.770.66
Cold Dew0.20.240.110.510.490.480.53
First Frost0.150.310.090.80.760.760.77
Beginning of Winter0.130.180.060.770.760.650.71
Light Snow0.10.160.060.930.930.850.89
Heavy Snow0.190.190.10.940.940.790.9
Winter Solstice0.090.140.040.860.820.750.81
Lesser Cold0.490.380.250.910.920.840.82
Great Cold0.320.360.170.920.920.820.79
Average Value0.290.280.170.790.780.760.77
Table A4. The R values between the clustering results of the multi-source data and the measured data.
Table A4. The R values between the clustering results of the multi-source data and the measured data.
Solar TermsWind-RSolar-R
NASAECMWFGreenwichNASAECMWFPVGISMeteonorm
Beginning of Spring0.15−0.140.250.910.940.880.81
Rain Water0.3−0.250.430.80.830.820.89
Insects Awakening0.26−0.230.270.930.950.950.88
Spring Equinox0.29−0.210.460.760.770.730.83
Fresh Green0.21−0.170.250.730.780.770.76
Grain Rain0.05−0.190.180.750.710.710.6
Beginning of Summer−0.120.010.090.210.180.160.21
Lesser Fullness0.110.050.10.470.450.420.44
Grain in Ear0.130.080.080.520.50.50.63
Summer Solstice0.1−0.050.190.280.260.240.34
Lesser Heat−0.010.03−0.040.720.680.670.63
Great Heat0.070.130.30.670.670.680.79
Beginning of Autumn0−0.15−0.040.790.770.770.66
End of Heat−0.04−0.090.170.720.740.740.65
White Dew0.15−0.230.090.680.640.630.56
Autumnal Equinox0.12−0.070.180.680.640.640.55
Cold Dew−0.03−0.020.140.20.160.150.24
First Frost00.210.280.650.630.60.64
Beginning of Winter0.030.030.060.60.580.410.52
Light Snow0.020.030.110.880.870.770.82
Heavy Snow−0.1−0.25−0.210.890.90.660.81
Winter Solstice−0.21−0.01−0.20.770.770.690.66
Lesser Cold0.28−0.170.080.860.860.730.74
Great Cold0.08−0.110.120.90.870.720.71
Average Value0.08−0.070.140.680.670.630.64
Figure A1 and Figure A2 show the error indicator values calculated based on multi-source data under different solar conditions. The label on the left side of the figure represents the website that obtained the wind and solar data, and the name of the error indicator used is marked on the far right. The numbers 1 to 24 on the horizontal axis represent the 24 solar terms. For RMSE and MBE, the lighter the color and the smaller the value, the smaller the data error. For IOA and R, the darker the color, the larger the value, and the higher the data consistency and correlation.
Figure A1. Short-term error indicator value of wind speed in solar terms scale.
Figure A1. Short-term error indicator value of wind speed in solar terms scale.
Energies 19 00074 g0a1
Figure A2. Short-term error indicator value of solar power in solar terms scale.
Figure A2. Short-term error indicator value of solar power in solar terms scale.
Energies 19 00074 g0a2

Appendix B. Detailed Clustering Results and Sensitivity Analysis During Drought Periods

Figure A3. The clustering result graph of solar output by solar term.
Figure A3. The clustering result graph of solar output by solar term.
Energies 19 00074 g0a3
Figure A4. The clustering result graph of wind output by solar term.
Figure A4. The clustering result graph of wind output by solar term.
Energies 19 00074 g0a4
The results of the drought sensitivity analysis conducted using the scenario generation method in this article are as follows:
Table A5. Analysis results of a 5% reduction in solar power generation (11:00–13:00).
Table A5. Analysis results of a 5% reduction in solar power generation (11:00–13:00).
Solar TermsCSDACFe ε c
Spring Equinox0.524 0.021 0.290 0.791
Summer Solstice0.551 0.019 0.335 0.794
Great Heat0.951 0.079 0.143 0.883
Autumnal Equinox0.702 0.012 0.208 0.843
Winter Solstice0.331 0.019 0.150 0.484
Great Cold0.508 0.026 0.259 0.539
Average Value0.5940.0290.230 0.722
Table A6. Analysis results of a 10% reduction in solar power generation (11:00–13:00).
Table A6. Analysis results of a 10% reduction in solar power generation (11:00–13:00).
Solar TermsCSDACFe ε c
Spring Equinox0.536 0.018 0.290 0.797
Summer Solstice0.551 0.019 0.325 0.794
Great Heat0.951 0.079 0.143 0.853
Autumnal Equinox0.712 0.025 0.208 0.844
Winter Solstice0.329 0.019 0.170 0.479
Great Cold0.508 0.018 0.259 0.592
Average Value0.5980.030 0.233 0.726

References

  1. Wang, Z.; Wen, X.; Tan, Q.; Fang, G.; Lei, X.; Wang, H.; Yan, J. Potential assessment of large-scale hydro-photovoltaic-wind hybrid systems on a global scale. Renew. Sustain. Energy Rev. 2021, 146, 111154. [Google Scholar] [CrossRef] [Scilit]
  2. Minola, L.; Zhang, G.; Ou, T.; Kukulies, J.; Curio, J.; Guijarro, J.A.; Deng, K.; Azorin-Molina, C.; Shen, C.; Pezzoli, A.; et al. Climatology of near-surface wind speed from observational, reanalysis and high-resolution regional climate model data over the Tibetan Plateau. Clim. Dynam. 2024, 62, 933–953. [Google Scholar] [CrossRef] [Scilit]
  3. Ren, G.; Wan, J.; Liu, J.; Yu, D. Spatial and temporal assessments of complementarity for renewable energy resources in China. Energy 2019, 177, 262–275. [Google Scholar] [CrossRef] [Scilit]
  4. Rafati, A.; Joorabian, M.; Mashhour, E.; Shaker, H.R. High dimensional very short-term solar power forecasting based on a data-driven heuristic method. Energy 2021, 219, 119647. [Google Scholar] [CrossRef] [Scilit]
  5. Feng, Y.; Qi, Y.; Chen, D.; Li, D.; Li, Z.; Xu, X. Multi-scale analysis of satellite, reanalysis and muti-source precipitation estimates over the Tibetan Plateau. Atmos. Res. 2024, 309, 107484. [Google Scholar] [CrossRef] [Scilit]
  6. Yousef, L.A.; Temimi, M.; Molini, A.; Weston, M.; Wehbe, Y.; Mandous, A.A. Cloud Cover over the Arabian Peninsula from Global Remote Sensing and Reanalysis Products. Atmos. Res. 2020, 238, 104866. [Google Scholar] [CrossRef] [Scilit]
  7. Yao, B.; Teng, S.; Lai, R.; Xu, X.; Yin, Y.; Shi, C.; Liu, C. Can atmospheric reanalyses (CRA and ERA5) represent cloud spatiotemporal characteristics? Atmos. Res. 2020, 244, 105091. [Google Scholar] [CrossRef] [Scilit]
  8. Ray, P.; Reddy, S.S.; Banerjee, T. Various dimension reduction techniques for high dimensional data analysis: A review. Artif. Intell. Rev. 2021, 54, 3473–3515. [Google Scholar] [CrossRef] [Scilit]
  9. Ikotun, A.M.; Ezugwu, A.E.; Abualigah, L.; Abuhaija, B.; Heming, J. K-means clustering algorithms: A comprehensive review, variants analysis, and advances in the era of big data. Inf. Sci. 2023, 622, 178–210. [Google Scholar] [CrossRef] [Scilit]
  10. Sihare, S.R. Dimensionality Reduction for Data Analysis with Quantum Feature Learning. Wires Data Min. Knowl. 2024, 15, e1568. [Google Scholar] [CrossRef] [Scilit]
  11. Wang, C.; Wang, Y.; Ding, Z.; Zheng, T.; Hu, J.; Zhang, K. A Transformer-Based Method of Multienergy Load Forecasting in Integrated Energy System. IEEE Trans. Smart Grid 2022, 13, 2703–2714. [Google Scholar] [CrossRef] [Scilit]
  12. Free, M.; Sun, B.; Yoo, H.L. Comparison between Total Cloud Cover in Four Reanalysis Products and Cloud Measured by Visual Observations at U.S. Weather Stations. J. Clim. 2016, 29, 2015–2021. [Google Scholar] [CrossRef] [Scilit]
  13. Li, Z.; Luo, Z.; Wang, Y.; Fan, G.; Zhang, J. Suitability evaluation system for the shallow geothermal energy implementation in region by Entropy Weight Method and TOPSIS method. Renew. Energy 2022, 184, 564–576. [Google Scholar] [CrossRef] [Scilit]
  14. Zhao, J.; Itti, L. shapeDTW: Shape Dynamic Time Warping. Pattern. Recogn. 2018, 74, 171–184. [Google Scholar] [CrossRef] [Scilit]
  15. Ji, C.; Hu, Y.; Liu, S.; Pan, L.; Li, B.; Zheng, X. Fully convolutional networks with shapelet features for time series classification. Inf. Sci. 2022, 612, 835–847. [Google Scholar] [CrossRef] [Scilit]
  16. Liu, H.; Chen, J.; Dy, J.; Fu, Y. Transforming Complex Problems into K-means Solutions. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 9149–9168. [Google Scholar] [CrossRef] [Scilit]
  17. Wang, H.; Liao, S.; Liu, B.; Zhao, H.; Ma, X.; Zhou, B. Long-term complementary scheduling model of hydro-wind-solar under extreme drought weather conditions using an improved time-varying hedging rule. Energy 2024, 305, 132285. [Google Scholar] [CrossRef] [Scilit]
  18. Berndt, D.J.; Clifford, J. Using dynamic time warping to find patterns in time series. In Proceedings of the 3rd International Conference on Knowledge Discovery and Data Mining, Phuket, Thailand, 9–10 January 2010; pp. 359–370. [Google Scholar]
  19. Jin, X.; Liu, B.; Liao, S.; Cheng, C.; Li, G.; Liu, L. Impacts of different wind and solar power penetrations on cascade hydroplants operation. Renew. Energy 2022, 182, 227–244. [Google Scholar] [CrossRef] [Scilit]
  20. Camal, S.; Teng, F.; Michiorri, A.; Kariniotakis, G.; Badesa, L. Scenario generation of aggregated Wind, Photovoltaics and small Hydro production for power systems applications. Appl. Energy 2019, 242, 1396–1406. [Google Scholar] [CrossRef] [Scilit]
  21. Feng, Z.; Huang, Q.; Niu, W.; Su, H.; Li, S.; Wu, H.; Wang, J. Peak operation optimization of cascade hydropower reservoirs and solar power plants considering output forecasting uncertainty. Appl. Energy 2024, 358, 122533. [Google Scholar] [CrossRef] [Scilit]
  22. Wang, R.; Ma, R.; Zeng, L.; Yan, Q.; Johnston, A.J. Improved bidirectional long short-term memory network-based short-term forecasting of photovoltaic power for different seasonal types and weather factors. Comput. Electr. Eng. 2025, 123, 110219. [Google Scholar] [CrossRef] [Scilit]
  23. Densing, M.; Wan, Y. Low-dimensional scenario generation method of solar and wind availability for representative days in energy modeling. Appl. Energy 2022, 306, 118075. [Google Scholar] [CrossRef] [Scilit]
  24. Hou, G.; Wang, J.; Fan, Y. Multistep short-term wind power forecasting model based on secondary decomposition, the kernel principal component analysis, an enhanced arithmetic optimization algorithm, and error correction. Energy 2024, 286, 129640. [Google Scholar] [CrossRef] [Scilit]
  25. Krishna, A.B.; Abhyankar, A.R. An Efficient Data-Driven Conditional Joint Wind Power Scenario Generation for Day-Ahead Power System Operations Planning. IEEE Trans. Power Syst. 2024, 39, 3105–3117. [Google Scholar] [CrossRef] [Scilit]
  26. Fan, Y.; Liu, W.; Zhu, F.; Wang, S.; Yue, H.; Zeng, Y.; Xu, B.; Zhong, P. Short-term stochastic multi-objective optimization scheduling of wind-solar-hydro hybrid system considering source-load uncertainties. Appl. Energy 2024, 372, 123781. [Google Scholar] [CrossRef] [Scilit]
  27. Ávila, R.L.; Mine, M.R.M.; Kaviski, E.; Detzel, D.H.M.; Fill, H.D.; Bessa, M.R.; Pereira, G.A.A. Complementarity modeling of monthly streamflow and wind speed regimes based on a copula-entropy approach: A Brazilian case study. Appl. Energy 2020, 259, 114127. [Google Scholar] [CrossRef] [Scilit]
  28. Peng, S.; Zhu, J.; Wu, T.; Yuan, C.; Cang, J.; Zhang, K.; Pecht, M. Prediction of wind and PV power by fusing the multi-stage feature extraction and a PSO-BiLSTM model. Energy 2024, 298, 131345. [Google Scholar] [CrossRef] [Scilit]
  29. Fileni, F.; Fowler, H.J.; Lewis, E.; McLay, F.; Yang, L. A quality-control framework for sub-daily flow and level data for hydrological modelling in Great Britain. Hydrol. Res. 2023, 54, 1357–1367. [Google Scholar] [CrossRef] [Scilit]
  30. Chen, S.; Huang, J.; Huang, J. Improving daily streamflow simulations for data-scarce watersheds using the coupled SWAT-LSTM approach. J. Hydrol. 2023, 622, 129734. [Google Scholar] [CrossRef] [Scilit]
  31. Abelló, A.; Cheney, J. Eris: Efficiently measuring discord in multidimensional sources. VLDB J. 2024, 33, 399–423. [Google Scholar] [CrossRef] [Scilit]
  32. Ming, B.; Liu, P.; Guo, S.; Zhang, X.; Feng, M.; Wang, X. Optimizing utility-scale photovoltaic power generation for integration into a hydropower reservoir by incorporating long- and short-term operational decisions. Appl. Energy 2017, 204, 432–445. [Google Scholar] [CrossRef] [Scilit]
  33. Fan, J.; Huang, X.; Shi, J.; Li, K.; Cai, J.; Zhang, X. Complementary potential of wind-solar-hydro power in Chinese provinces: Based on a high temporal resolution multi-objective optimization model. Renew. Sustain. Energy Rev. 2023, 184, 113566. [Google Scholar] [CrossRef] [Scilit]
  34. Ding, Z.; Wen, X.; Tan, Q.; Yang, T.; Fang, G.; Lei, X.; Zhang, Y.; Wang, H. A forecast-driven decision-making model for long-term operation of a hydro-wind-photovoltaic hybrid system. Appl. Energy 2021, 291, 116820. [Google Scholar] [CrossRef] [Scilit]
  35. Zhao, G.; Yu, C.; Huang, H.; Yu, Y.; Zou, L.; Mo, L. Optimization Scheduling of Hydro–Wind–Solar Multi-Energy Complementary Systems Based on an Improved Enterprise Development Algorithm. Sustainability 2025, 17, 2691. [Google Scholar] [CrossRef] [Scilit]
  36. Wang, Y.; Hu, Q.; Li, L.; Foley, A.M.; Srinivasan, D. Approaches to wind power curve modeling: A review and discussion. Renew. Sust. Energy Rev. 2019, 116, 109422. [Google Scholar] [CrossRef] [Scilit]
  37. Paparrizos, J.; Gravano, L. k-shape: Efficient and accurate clustering of time series. In Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data, Melbourne, Victoria, Australia, 31 May–4 June 2015; pp. 1855–1870. [Google Scholar] [CrossRef] [Scilit]
  38. Liu, Q.; Ma, J.; Zhao, X.; Zhang, K.; Xiangli, K.; Meng, D. A novel method for fault diagnosis and type identification of cell voltage inconsistency in electric vehicles using weighted Euclidean distance evaluation and statistical analysis. Energy 2024, 293, 130575. [Google Scholar] [CrossRef] [Scilit]
  39. Cardo-Miota, J.; Beltran, H.; Pérez, E.; Khadem, S.; Bahloul, M. Deep reinforcement learning-based strategy for maximizing returns from renewable energy and energy storage systems in multi-electricity markets. Appl. Energy 2025, 388, 125561. [Google Scholar] [CrossRef] [Scilit]
  40. Leal, A.B.; de Souza Filho, H.N.; Gabriel Zanon, L.; Sigahi, F.A.C.T.; Simon Rampasso, I.; Anholon, R. Selection of photovoltaic panels for floating systems: An analysis based on Entropy, CRITIC and TOPSIS. Int. J. Sustain. Eng. 2024, 17, 1173–1185. [Google Scholar] [CrossRef] [Scilit]
  41. Zhu, A.; Zhao, Q.; Yang, T.; Zhou, L.; Zeng, B. Condition monitoring of wind turbine based on deep learning networks and kernel principal component analysis. Comput. Electr. Eng. 2023, 105, 108538. [Google Scholar] [CrossRef] [Scilit]
  42. Gou, H.; Ning, Y. Forecasting Model of Photovoltaic Power Based on KPCA-MCS-DCNN. Comput. Model. Eng. Sci. 2021, 128, 803–822. [Google Scholar] [CrossRef] [Scilit]
  43. Lin, H.; Gao, L.; Cui, M.; Liu, H.; Li, C.; Yu, M. Short-term distributed photovoltaic power prediction based on temporal self-attention mechanism and advanced signal decomposition techniques with feature fusion. Energy 2025, 315, 134395. [Google Scholar] [CrossRef] [Scilit]
  44. Zhou, S.; Han, Y.; Zalhaf, A.S.; Chen, S.; Zhou, T.; Yang, P.; Elboshy, B. A novel multi-objective scheduling model for grid-connected hydro-wind-PV-battery complementary system under extreme weather: A case study of Sichuan, China. Renew. Energy 2023, 212, 818–833. [Google Scholar] [CrossRef] [Scilit]
  45. Woodruff, D.L.; Deride, J.; Staid, A.; Watson, J.; Slevogt, G.; Silva-Monroy, C. Constructing probabilistic scenarios for wide-area solar power generation. Sol. Energy 2018, 160, 153–167. [Google Scholar] [CrossRef] [Scilit]
  46. Cohen, N.; Berchenko, Y. Normalized Information Criteria and Model Selection in the Presence of Missing Data. Mathematics 2021, 9, 2474. [Google Scholar] [CrossRef] [Scilit]
  47. Wang, Y.; Ding, Y.; Shahrampour, S. TAKDE: Temporal Adaptive Kernel Density Estimator for Real-Time Dynamic Density Estimation. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 13831–13843. [Google Scholar] [CrossRef] [Scilit]
  48. Lv, M.; Gou, K.; Chen, H.; Lei, J.; Zhang, G.; Liu, T. Optimal Design of Wind-Solar complementary power generation systems considering the maximum capacity of renewable energy. Energy 2024, 312, 133650. [Google Scholar] [CrossRef] [Scilit]
  49. NASA POWER Data Access Viewer (DAV). Available online: https://power.larc.nasa.gov/data-access-viewer/ (accessed on 1 September 2025).
  50. Climate Data Store. Available online: https://cds.climate.copernicus.eu/cdsapp#!/dataset/reanalysis-era5-complete?tab=form (accessed on 1 September 2025).
  51. JRC Photovoltaic Geographical Information System (PVGIS)—European Commission. Available online: https://re.jrc.ec.europa.eu/pvg_tools/en/ (accessed on 1 September 2025).
  52. Meteonorm. Available online: https://meteonorm.com/ (accessed on 1 September 2025).
  53. National Meteorological Information Center—China Meteorological Data Network. Available online: https://data.cma.cn/ (accessed on 1 September 2025).
  54. Greenwich. Available online: https://greenwich.envisioncn.com/web/portal/login (accessed on 1 September 2025).
  55. Kim, H.; Park, I.; Park, J.; Kim, J.K.; Seo, M.; Kim, J.K. scICE: Enhancing clustering reliability and efficiency of scRNA-seq data with multi-cluster label consistency evaluation. Nat. Commun. 2025, 16, 6031. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Schematic diagram of comprehensive similarity distance.
Figure 1. Schematic diagram of comprehensive similarity distance.
Energies 19 00074 g001
Figure 2. Determine the optimal K schematic diagram.
Figure 2. Determine the optimal K schematic diagram.
Energies 19 00074 g002
Figure 3. Schematic diagram of dimensionality reduction for solar and wind clustering.
Figure 3. Schematic diagram of dimensionality reduction for solar and wind clustering.
Energies 19 00074 g003
Figure 4. Principal component variance interpretation rate.
Figure 4. Principal component variance interpretation rate.
Energies 19 00074 g004
Figure 5. Example of the main component load of the solar term.
Figure 5. Example of the main component load of the solar term.
Energies 19 00074 g005
Figure 6. The overall process of scenario generation.
Figure 6. The overall process of scenario generation.
Energies 19 00074 g006
Figure 7. Schematic diagram of the Yalong River Basin and wind–solar planned power stations.
Figure 7. Schematic diagram of the Yalong River Basin and wind–solar planned power stations.
Energies 19 00074 g007
Figure 8. The clustering comparison results of wind and solar data for the Beginning of Spring.
Figure 8. The clustering comparison results of wind and solar data for the Beginning of Spring.
Energies 19 00074 g008
Figure 9. Evaluation chart of performance indicators for wind and solar websites.
Figure 9. Evaluation chart of performance indicators for wind and solar websites.
Energies 19 00074 g009
Figure 10. The clustering result graph of wind output in typical solar terms.
Figure 10. The clustering result graph of wind output in typical solar terms.
Energies 19 00074 g010
Figure 11. The clustering result graph of solar output in typical solar terms.
Figure 11. The clustering result graph of solar output in typical solar terms.
Energies 19 00074 g011
Figure 12. Principal component loading and correlation capture results.
Figure 12. Principal component loading and correlation capture results.
Energies 19 00074 g012
Figure 13. The generation scenarios of runoff, wind, and solar power output in typical solar terms.
Figure 13. The generation scenarios of runoff, wind, and solar power output in typical solar terms.
Energies 19 00074 g013
Figure 14. Comparison results of clustering methods.
Figure 14. Comparison results of clustering methods.
Energies 19 00074 g014
Table 1. The CSD mean value derived from the cross-calculation of wind and solar data sources.
Table 1. The CSD mean value derived from the cross-calculation of wind and solar data sources.
Wind PowerNASAGreenwichECMWFStability score
NASA0.0620.09712.578
Greenwich0.0620.09912.422
ECMWF0.0970.09910.204
Solar PowerNASAPVGISMeteonormECMWFStability score
NASA0.0070.0080.006142.857
PVGIS0.0070.0130.009103.093
Meteonorm0.0080.0130.01583.333
ECMWF0.0060.0090.015100
Table 2. Calculation results of wind and solar data source indicators.
Table 2. Calculation results of wind and solar data source indicators.
Error Indicator1/RMSE1/|MBE|IOAR1/CSDρS
WindGreenwich0.1930.1040.3840.5240.2670.4970.350
ECMWF0.4150.2430.5060.3830.3700.6070.397
NASA0.4580.0550.5090.5640.2300.4650.413
SolarPVGIS0.4530.0880.6930.5950.2600.6760.522
NASA0.2950.0550.7240.6610.3020.6970.551
Meteonorm0.5070.1780.7030.6310.2750.6590.546
ECMWF0.3670.1210.7030.6510.2820.670 0.541
Table 3. AIC/BIC calculation results of different Copula.
Table 3. AIC/BIC calculation results of different Copula.
Copula TypeAverage AIC ValueAverage AIC Value
Gaussian−5.986−3.773
Clayton−3.019−0.807
Frank−16.956−14.743
Gumbel−5.553−3.341
t−28.221−23.797
Table 4. Calculation results of different scheme indicators.
Table 4. Calculation results of different scheme indicators.
Solar TermsScheme 1Scheme 2Scheme 3
CSDACFe ε c CSDACFe ε c CSDACFe ε c
Spring Equinox0.878 0.035 0.295 1.026 0.519 0.022 0.290 0.786 0.525 0.007 0.149 0.618
Summer Solstice0.612 0.028 0.320 2.005 0.551 0.019 0.325 0.794 0.435 0.017 0.264 0.637
Great Heat0.823 0.136 0.092 2.595 0.951 0.079 0.133 0.853 0.458 0.038 0.092 0.752
Autumnal Equinox0.731 0.008 0.252 1.204 0.698 0.012 0.208 0.841 0.942 0.053 0.237 0.805
Winter Solstice1.130 0.206 0.171 4.142 0.333 0.023 0.160 0.386 0.590 0.222 0.211 0.343
Great Cold0.920 0.215 0.278 4.925 0.509 0.025 0.260 0.516 0.727 0.223 0.131 0.019
Average value0.8490.1050.2352.6510.5930.0310.2290.6960.6130.0930.1820.529
Table 5. Scenario quality indicator results under advance and delay conditions.
Table 5. Scenario quality indicator results under advance and delay conditions.
Average ValueCSDACFe ε c
7 days in advance0.601 0.051 0.211 0.709
7 days behind0.593 0.052 0.208 0.704
15 days in advance0.557 0.049 0.198 0.686
15 days behind0.588 0.056 0.204 0.687
Scheme 2 (Standard)0.562 0.021 0.207 0.698
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Guo, P.; Cheng, X.; Min, W.; Zeng, X.; Sun, J. A Climate-Informed Scenario Generation Method for Stochastic Planning of Hybrid Hydro–Wind–Solar Power Systems in Data-Scarce Regions. Energies 2026, 19, 74. https://doi.org/10.3390/en19010074

AMA Style

Guo P, Cheng X, Min W, Zeng X, Sun J. A Climate-Informed Scenario Generation Method for Stochastic Planning of Hybrid Hydro–Wind–Solar Power Systems in Data-Scarce Regions. Energies. 2026; 19(1):74. https://doi.org/10.3390/en19010074

Chicago/Turabian Style

Guo, Pu, Xiong Cheng, Wei Min, Xiaotao Zeng, and Jingwen Sun. 2026. "A Climate-Informed Scenario Generation Method for Stochastic Planning of Hybrid Hydro–Wind–Solar Power Systems in Data-Scarce Regions" Energies 19, no. 1: 74. https://doi.org/10.3390/en19010074

APA Style

Guo, P., Cheng, X., Min, W., Zeng, X., & Sun, J. (2026). A Climate-Informed Scenario Generation Method for Stochastic Planning of Hybrid Hydro–Wind–Solar Power Systems in Data-Scarce Regions. Energies, 19(1), 74. https://doi.org/10.3390/en19010074

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop