1. Introduction
Global hydrological models (GHMs) play a crucial role in evaluating climate change impacts on hydrological processes [
1,
2]. The societal relevance of reliable hydrological assessment is increasing as climate change is projected to expose future generations to unprecedented river-flood conditions and to amplify the economic consequences of climate extremes [
3,
4]. The Inter-Sectoral Impact Model Intercomparison Project (ISIMIP) has pioneered an innovative framework for GHM simulations by integrating multidisciplinary models with standardized multi-scenario datasets [
5,
6]. The recent ISIMIP3a water-sector ensemble provides standardized GHM simulations under a common forcing and output framework [
5,
7], creating a consistent basis for model intercomparison. However, the standard configurations retain differences in inherited calibration, routing schemes, and human–water representations [
5,
8], making regional benchmark evaluations necessary to identify model-specific strengths and limitations.
International GHM intercomparison studies have shown that model performance varies substantially across models, regions, temporal scales, and flow regimes [
8,
9]. A recent evaluation of eight global water models further showed that discharge performance varies across geographic and climatic regions and can deteriorate as human impacts intensify [
8]. Using multi-model validation, Veldkamp et al. [
9] showed that human-impact parameterizations can improve estimates of monthly discharge and hydrological extremes, while other studies emphasized the importance of regionalization and parameter estimation when transferring GHMs to basin-scale applications [
10,
11]. Nevertheless, systematic discharge biases remain at the global scale [
12], and flood simulations continue to show substantial uncertainty [
13,
14]. These findings indicate that GHM performance should be assessed using multiple metrics, temporal scales, and flow conditions rather than a single aggregate measure.
The Yangtze River Basin is a strategic hydrological region with major socioeconomic importance in China and is strongly affected by climate variability and intensive human activities [
15,
16,
17,
18]. Previous Yangtze-focused studies have examined runoff responses to extreme precipitation [
19], reservoir-induced changes in streamflow [
16], and climate change impacts on discharge [
18]. More recently, Zhao et al. [
20] demonstrated that basin-specific calibration can improve regional flood- and drought-related simulations of a global hydrological model. Such work highlights the value of regional calibration for basin-scale applications but does not show how multiple standard GHM configurations, used without additional Yangtze-specific recalibration, compare under a common evaluation framework. Gauge-based multi-model benchmarking of standard ISIMIP3a configurations at major Yangtze mainstem stations therefore remains limited [
8,
12,
20].
Here, we evaluate eight standard ISIMIP3a GHM configurations at four principal Yangtze mainstem gauges under common forcing and a consistent output framework. Using complementary metrics, the analysis addresses three diagnostic questions: (1) whether model skill and relative ranking remain consistent across daily discharge, monthly discharge, the seasonal cycle (i.e., the mean annual cycle of monthly discharge), and annual peak discharge; (2) whether models that reproduce overall discharge dynamics also reproduce low-flow percentiles and annual peak discharge; and (3) how inherited calibration and model-configuration differences constrain the interpretation of apparent model skill. By integrating these diagnostics, the study identifies cross-scale and flow-regime-specific strengths and weaknesses that would be obscured by a single aggregate ranking.
2. Materials and Methods
2.1. The Yangtze River
The Yangtze River Basin (24°30′–35°45′ N, 90°33′–122°25′ E), China’s largest river system, spans approximately 1.8 million km
2 and delivers a mean annual runoff volume of approximately 960 billion m
3 [
21,
22]. The basin experiences a subtropical monsoon climate marked by pronounced spatiotemporal variability in precipitation, driving diverse hydrological regimes across its upper, middle, and lower reaches [
21,
23]. The locations of the Yangtze River Basin and the four gauging stations are shown in
Figure 1.
2.2. Datasets
Eight GHM simulations from the historical ISIMIP3a water-sector simulations were used as the model discharge data [
5]. All simulations were driven by the same daily GSWP3-W5E5 meteorological forcing [
5,
7], comprising precipitation, 2 m air temperature, shortwave and longwave radiation, specific humidity, surface pressure, and near-surface wind speed on a 0.5° × 0.5° latitude–longitude grid. The evaluated output was daily river discharge (dis), expressed in m
3 s
−1 [
5]. Examination of the original NetCDF metadata showed that all eight discharge outputs were stored as dis (time, lat, lon) on a regular global grid comprising 360 latitude and 720 longitude cells, with coordinates representing grid-cell centers, and used the proleptic Gregorian calendar.
Observed daily discharge records at four principal Yangtze River gauging stations were used as the reference data for model evaluation. The four stations are Cuntan, Yichang, Hankou, and Datong, representing the upper reach, the upper-middle transition, the middle reach, and the lower reach of the Yangtze River, respectively. These stations were selected as major mainstem gauges to evaluate integrated discharge responses along the upper-to-lower Yangtze River. The evaluation therefore focuses on integrated mainstem discharge responses rather than separate tributary or sub-basin-scale responses. The reference discharge records were derived from in situ measurements at the four gauging stations; thus, model performance was evaluated against observations rather than against another simulation product. Detailed metadata on the administering agency/database and station-level quality-control procedures were unavailable in the records accessible for this study; therefore, no unverified provenance or quality-control information was inferred. The observed discharge records represent the observed river-flow conditions during the respective evaluation periods, including any effects of reservoir operations present in the measured flows.
The available evaluation periods were determined by the overlap between observed discharge records and GHM simulations: Cuntan, 1975–2008; Yichang, 1950–2002; Hankou, 1976–2001; and Datong, 1976–2001. The unequal record lengths reflect differences in the availability of continuous in situ observations. The records from the three downstream gauges end in 2002, whereas the longer Cuntan record is from a station upstream of the Three Gorges Reservoir.
2.3. Global Hydrological Models
WEB-DHM-SG is a previously developed hydrological–hydrodynamic coupled model and was not newly developed in this study [
24,
25]. The model has been described in previous work, and its routing representation builds on the CaMa-Flood framework [
26]. The WaterGAP2-2e model integrates natural hydrological cycles with anthropogenic water use—including irrigation, industrial consumption, and reservoir operations—to simulate global water resource dynamics [
27]. The H08 model, an integrated water resource management framework, incorporates irrigation demand, reservoir operations, and groundwater abstraction [
28].
The CWatM model employs a modular architecture that integrates surface and groundwater hydrological processes with human water use, reservoir regulation, water demand, and return flows [
29,
30]. The HydroPy model, a global distributed hydrological framework, simulates large-scale basin hydrological cycles by integrating physical mechanisms with spatially heterogeneous drivers such as topography, vegetation, and soil properties [
31]. The JULES-W2-DDM30 model is based on the Joint UK Land Environment Simulator (JULES), which couples vegetation and hydrological processes, with DDM30 used as the river-routing network [
32,
33,
34]. MIROC-INTEG-LAND is an integrated land model that couples water resources, crop production, ecosystem processes, and land-use change [
35]. ORCHIDEE-MICT extends the ORCHIDEE land-surface framework with high-latitude process representations that describe interactions among soil carbon, soil temperature, and hydrology and their feedbacks on water and CO
2 fluxes [
36]; the evaluated ISIMIP3a water-sector configuration is documented by ISIMIP [
37].
Taken together, the eight GHMs differ substantially in their routing schemes and representations of human water use and reservoirs [
25,
30,
33,
37,
38,
39,
40,
41]. Among the evaluated configurations, only WEB-DHM-SG incorporates the modified CaMa-Flood hydrodynamic routing module based on DDM30 [
25], whereas the other models use different routing schemes or river-network formulations [
30,
33,
37,
38,
39,
40,
41]. Reservoir operations and human–water processes were not harmonized separately in this study but were retained as specified in each model’s standard ISIMIP3a configuration [
5,
25,
30,
33,
37,
38,
39,
40,
41]. These structural and configuration differences provide important context for interpreting inter-model performance. Detailed routing schemes, human-water-use sectors, and reservoir representations are provided in
Table S1 [
42].
All eight GHM simulations were evaluated as provided in the ISIMIP3a ensemble [
5], and no model was recalibrated against the four Yangtze River gauges in this study, because the objective was to benchmark the standard ISIMIP3a configurations under a common forcing framework. WaterGAP2-2e was globally calibrated by its development team against observed river discharge, with the Global Runoff Data Centre (GRDC) serving as the main source of streamflow gauging-station data [
27,
43]. Based on the coordinates of the WaterGAP2-2e global calibration gauges, the calibration-station dataset was spatially screened using the Yangtze River Basin boundary, identifying eight calibration gauges within the basin. The identities and locations of these gauges were then compared with the four evaluation gauges, showing that Yichang, Hankou, and Datong overlap with the WaterGAP2-2e calibration dataset, whereas Cuntan does not. H08 retained the climate-zone-specific parameter optimization following Yoshida et al. (2022) [
44]. The remaining models were evaluated using their standard ISIMIP3a configurations without additional Yangtze-specific calibration or optimization [
5]. Accordingly, no Yangtze-specific parameter ranges or optimal parameter values were estimated in this study.
Table 1, therefore, summarizes the key ISIMIP3a characteristics and calibration or parameter-setting status of the eight GHMs rather than basin-specific optimization results.
2.4. Data Processing and Evaluation Design
To make full use of the available observational information, each gauge was evaluated over its complete available record rather than restricting all four gauges to their common overlapping period. At each gauge, all eight GHMs were evaluated over the same station-specific period, ensuring temporally consistent within-gauge comparisons. Consequently, the gauge-level evaluations correspond to different historical periods across the four stations. To assess the sensitivity of the cross-gauge summary metrics to the unequal historical evaluation windows, a supplementary common-period analysis was conducted for 1976–2001, which is shared by all four gauges. Daily, monthly, seasonal-cycle, and annual peak-discharge metrics were recalculated using the same data-processing and metric calculation procedures as in the primary analysis. Cross-gauge averages were calculated using the same unweighted averaging procedure, with absolute gauge-level RB and RBpeak values averaged for the bias-based summaries. The common-period results were compared with the primary station-specific full-record results to assess the robustness of the cross-model patterns.
ISIMIP3 provides the DDM30 river-routing network as a standardized global drainage-direction dataset for water-sector simulations [
45,
46]. DDM30 represents surface-water drainage directions at 30′ spatial resolution and has been extensively validated and applied in large-scale global hydrological studies [
45,
47]. All four gauges are located along the Yangtze River mainstem. For each gauge, the corresponding 0.5° × 0.5° model grid cell was identified using the gauge coordinates together with the mapped alignment of the Yangtze mainstem. No automated river-network snapping algorithm was applied. Because all eight GHM discharge outputs share the same spatial grid, the same station-to-grid assignments were applied consistently across all models. Simulated discharge was extracted from the selected grid cells, whose spatial extents are shown in
Figure 1. The observed gauge coordinates and corresponding model grid-cell center coordinates are provided in
Table S23. The observed and simulated daily discharge series were aligned by calendar date without lag optimization or temporal shifting. In the data-processing procedure, discharge values coded as −999 and dates without available observations were represented as missing values (NaN). Model performance metrics were calculated only from paired observed and simulated values after excluding dates with missing values in either series. The upstream drainage areas represented by the selected model grid cells were not quantitatively matched to the reported drainage areas of the corresponding gauges.
Monthly discharge series were derived from the aligned daily series by calculating the monthly mean discharge. The seasonal cycle was subsequently obtained by averaging the monthly discharge values for each calendar month across all years within the corresponding station-specific evaluation period. Thus, both the monthly and seasonal-cycle evaluations were derived from the original daily discharge series rather than from independent monthly model simulations. Annual peak discharge was defined as the maximum daily discharge within each calendar year and was used for the subsequent peak-flow magnitude and trend analyses.
2.5. Evaluation Metrics
This study used a multi-metric framework to evaluate the performance of the eight GHMs across daily discharge, monthly discharge, the seasonal cycle, different parts of the discharge distribution, and annual peak discharge. Daily discharge was evaluated using Nash–Sutcliffe Efficiency (NSE), Relative Bias (RB), Pearson’s correlation coefficient (R), and mean absolute error (MAE), whereas monthly discharge and the seasonal cycle were evaluated using NSE, R, and MAE. Flow-percentile performance was assessed using percentile-specific RB, and annual peak-discharge magnitude was assessed using the Relative Bias of peak flow (RBpeak). For each metric, the station-averaged value was calculated as the unweighted arithmetic mean of the corresponding gauge-level values across the four gauges, with each gauge-level value derived from its station-specific evaluation period. For bias-based cross-gauge summaries, the absolute gauge-level RB or RBpeak values were averaged to avoid cancellation between positive and negative biases. These averages provide only a descriptive cross-gauge summary of model performance and should not be interpreted as rankings based on a common historical period or as area-weighted basin-wide hydrological averages.
NSE quantifies the goodness of fit between simulated and observed discharge and is widely used for hydrological model evaluation [
48]. It is defined as follows:
where
denotes the observed values,
represents the simulated values, and
is the mean of the observed values. RB measures volumetric errors and systematic deviations in long-term discharge-volume bias:
R measures the temporal correspondence between simulated and observed discharge series [
49]:
where
is the mean of the simulated values. MAE was used to quantify the mean absolute magnitude of simulation errors, with the same unit as discharge [
50]:
where
and
denote the observed and simulated discharge values at time step
, respectively, and
is the number of valid paired observations.
Taylor diagrams were also used as a complementary diagnostic to jointly summarize the Pearson correlation coefficient, R, the normalized standard deviation, and the centered root-mean-square difference between simulated and observed discharge [
51]. The Taylor diagrams were constructed from the same temporally aligned or aggregated observed and simulated discharge series used for the corresponding scale-specific evaluations. Kling–Gupta efficiency (KGE) was not additionally calculated because its principal diagnostic dimensions—correlation, bias, and variability—were examined separately using R, RB, and the Taylor-diagram diagnostics [
52].
To evaluate model performance across different parts of the discharge distribution, flow-percentile diagnostics were calculated from daily discharge series. The 5th, 10th, 25th, 50th, and 95th percentiles were used to represent low-flow to high-flow conditions. The Relative Bias for each percentile was calculated as follows:
where
and
are the
th percentiles of observed and simulated daily discharge, respectively. In particular,
and
were used to assess low-flow performance, whereas
was used as a high-flow percentile diagnostic.
Trends in annual peak discharge were estimated using ordinary least-squares linear regression between annual peak discharge and year, with slopes reported in m
3 s
−1 year
−1. Trend significance was assessed using the Mann–Kendall test at α = 0.05. A
value below 0.05 indicates a statistically significant monotonic trend, whereas
> 0.05 indicates that the trend is not statistically significant [
53]. Accordingly, the flood-focused evaluation in this study was limited to Q95 bias, annual maximum daily discharge magnitude, RB
peak, and the OLS-fitted slopes and Mann–Kendall significance of the annual maxima. Peak timing, threshold-exceedance frequency, flood volume, and flood duration were not evaluated.
4. Discussion
This study reveals marked performance differences among the eight ISIMIP3a GHMs. Across the descriptive station-averaged diagnostics shown in
Figure 6, WaterGAP2-2e showed the most consistently strong performance among the evaluated standard ISIMIP3a configurations. Because the gauge-level metrics underlying this cross-gauge synthesis were calculated over different historical evaluation periods,
Figure 6 should not be interpreted as a controlled common-period ranking. Nevertheless, the supplementary 1976–2001 common-period analysis reproduced the principal cross-model patterns represented by the cross-gauge summary metrics (
Tables S24–S27), with only minor metric-specific changes in relative ordering. This indicates that the principal interpretation of
Figure 6 is not solely an artifact of the unequal historical evaluation windows. This result must be interpreted in the context of its inherited global discharge calibration and other configuration features, including model structure, routing, and human–water representation, rather than as evidence of intrinsic structural superiority. Yichang, Hankou, and Datong coincide with gauges in the WaterGAP2-2e calibration-station dataset, whereas Cuntan does not; however, this spatial correspondence does not establish that the calibration and evaluation used identical observation periods or discharge records. The relatively strong performance at Cuntan provides evidence of performance beyond the identified calibration-gauge locations, but the present design cannot isolate calibration effects from other configuration differences. Accordingly, the following comparisons concern complete model configurations rather than isolated calibration, routing, or structural effects.
The higher metric-based performance at monthly and seasonal-cycle scales partly reflects the smoothing effect of temporal aggregation. Monthly discharge was calculated by averaging the daily series, and the seasonal cycle was further obtained by averaging the same calendar months across years. These successive averaging steps reduce the influence of short-term fluctuations and event-scale timing or magnitude errors, allowing the evaluation to emphasize the more persistent seasonal discharge pattern. In contrast, annual peak discharge was defined as the maximum daily discharge in each year and therefore retains event-scale errors rather than averaging them out. Peak-flow performance is consequently more sensitive to errors in the magnitude and timing of individual flood events and to uncertainties in meteorological forcing, particularly precipitation, as well as limitations in runoff-generation and routing processes.
Among the remaining configurations, WEB-DHM-SG showed the strongest overall performance in terms of NSE and bias-related diagnostics and had the smallest station-averaged absolute peak-flow bias (|RBpeak| = 15.52%). Its relatively strong performance is consistent with its explicit coupling of hydrological and hydrodynamic simulations [
24,
25,
26], although the contribution of this routing representation cannot be isolated from other configuration differences in the present evaluation. However, this characterization applies primarily to general discharge dynamics and peak-flow magnitude rather than to all flow regimes. WEB-DHM-SG substantially underestimated Q05 at all four gauges, with biases reaching −95.69% at Yichang and −96.78% at Cuntan, indicating limited reliability for low-flow simulation. Although Q05 provides a low-flow diagnostic, dedicated drought and ecological-flow indicators were not evaluated in this study; therefore, such applications would require additional independent validation rather than being inferred from the overall NSE, R, or RB results.
In contrast, the H08 model systematically underestimated daily discharge at all four gauges (RB = −36.78% to −45.06%;
Tables S2–S5) and showed weak daily scale performance, with negative NSE values at Datong, Hankou, and Yichang. These results indicate limited transferability of the standard H08 configuration to the Yangtze River Basin under the present evaluation framework. Given that H08 retained climate-zone-specific parameter optimization rather than Yangtze-specific discharge calibration [
44], its weak performance may partly reflect limitations in transferring this parameterization to the basin; however, the present evaluation cannot isolate parameterization effects from other configuration differences.
The ORCHIDEE-MICT configuration showed strongly station-dependent performance, with an extreme discharge-magnitude underestimation at Hankou across daily, monthly, seasonal-cycle, flow-percentile, and annual peak-discharge diagnostics. The diagnostic check further showed that this anomaly was not attributable to missing or invalid records (
Table S22). Because the daily biases at Datong, Yichang, and Cuntan were substantially smaller, the near-total discharge-magnitude underestimation appears specific to Hankou rather than reflecting a uniform basin-wide scaling bias. However, the available diagnostics confirm the magnitude anomaly but do not identify its cause. Possible runoff-generation, routing, grid-cell matching, or other configuration issues therefore require further investigation, and the Hankou result should be interpreted cautiously. At the monthly scale, CWatM performed better at Cuntan (NSE = 0.75) than at the three downstream gauges (NSE = 0.56–0.59;
Tables S7–S10). This spatial contrast may be associated with station-dependent differences in hydrological regulation, routing, contributing-area representation, water use, or other configuration factors; however, their individual contributions cannot be isolated in the present evaluation.
These contrasting behaviors indicate that model suitability is application-dependent rather than represented by a single overall ranking. Models that reproduce general discharge dynamics well may still perform poorly for low flows or flood peaks. Model selection should therefore consider the target temporal scale and flow regime, particularly whether the application emphasizes overall discharge dynamics, low-flow conditions, or flood-peak simulation.
These findings are broadly consistent with, but also extend, previous GHM evaluations. The strong performance of the evaluated WaterGAP2-2e configuration is consistent with broader evidence that parameter estimation and calibration can improve GHM performance [
11,
20]. In the present study, however, models representing human–water processes did not perform uniformly better than those without such representations, indicating that human–water process representation is only one component of regional model skill. This finding complements recent global evidence that discharge performance can deteriorate as human impacts intensify and that challenges remain in representing human activities in heavily managed regions [
8].
In contrast to the general tendency toward global discharge overestimation reported by Heinicke et al. [
12], the Yangtze results include both positive and negative biases, indicating substantial regional and model dependence. The pronounced station-dependent performance of several models further supports the need for basin-scale evaluation and cautious assessment of regional model transferability [
10]. The difficulty in reproducing annual peak discharge is also consistent with previous GHM flood-evaluation and multi-model flood-assessment studies [
13,
14]. The flood-related findings should therefore be interpreted as diagnostics of high-flow magnitude and annual-peak behavior rather than as a comprehensive evaluation of flood-event dynamics. Because peak timing, exceedance frequency, flood volume, and flood duration were outside the evaluation scope, the results should not be interpreted as a complete assessment of flood hazards. Future evaluations could also incorporate complementary temporal diagnostics beyond monotonic trends, as recent ISIMIP-based studies have demonstrated the value of examining changes in the temporal regularity and dominant periods of climate-impact extremes under global warming [
54].
This study has several limitations. First, the evaluation was based on four major mainstem gauges and therefore primarily constrains integrated mainstem discharge responses rather than tributary, headwater, high-elevation, or reservoir-regulated sub-basin processes. The selected model grid cells were identified using gauge coordinates and mainstem alignment, but their upstream contributing areas were not quantitatively matched to the reported gauge drainage areas, and no lag optimization applied. Residual grid-to-gauge, contributing-area, and travel-time mismatches may therefore affect the station-based metrics. Second, the primary station-specific evaluation periods differ because longer continuous in situ records were unavailable for all gauges. The primary station-averaged metrics therefore combine different historical windows and should still be interpreted as descriptive cross-gauge summaries rather than strictly controlled common-period rankings. However, the supplementary 1976–2001 analysis produced broadly consistent cross-model summary patterns, indicating that the principal
Figure 6 conclusions are relatively robust to this difference in evaluation windows, despite minor metric-specific changes in relative ordering.
Third, the WaterGAP2-2e evaluation was not fully independent of calibration because three of the four evaluation gauges overlapped with its calibration dataset. Reservoir regulation also remains an unresolved source of inter-model uncertainty because reservoir processes differ among the standard ISIMIP3a configurations and the available downstream observations do not support a consistent before-and-after assessment of Three Gorges regulation. In addition, uncertainty in the GSWP3-W5E5 forcing was not independently evaluated, and none of the models were recalibrated specifically for the Yangtze River Basin. Because no process-isolation experiments were performed, the individual effects of calibration, parameterization, routing, reservoir representation, and other human–water processes cannot be quantified separately. Accordingly, the reported rankings should be interpreted as comparisons of complete standard model configurations rather than as controlled comparisons of intrinsic model structures or the performance achievable after basin-specific calibration.
Despite these limitations, this study provides more than an overall ranking of the eight GHMs. Relative to previous Yangtze studies that have mainly focused on the regional calibration and application of individual models [
20], the present analysis establishes a mainstem gauge-based benchmark of standard ISIMIP3a configurations across temporal scales and flow regimes. By integrating daily, monthly, seasonal-cycle, flow-percentile, and annual peak-discharge diagnostics, the study shows that model suitability is scale- and flow-regime dependent and that apparent superiority must be interpreted in the context of inherited calibration and configuration. Identifying these performance trade-offs and regime-specific weaknesses constitutes the main contribution of the study.
5. Conclusions
The evaluation of eight standard ISIMIP3a GHM configurations at four principal Yangtze mainstem gauges revealed strong dependence of model performance on temporal scale and flow regime. Metric-based performance generally increased from daily to monthly and seasonal-cycle scales, consistent with the smoothing of short-term timing and magnitude errors through temporal aggregation, whereas annual peak discharge remained more difficult to reproduce. WaterGAP2-2e showed the most consistently strong performance across these temporal-scale evaluations, although this result should be interpreted in light of its inherited global discharge calibration and other configuration differences.
For the other evaluated configurations, WEB-DHM-SG showed the strongest overall performance in terms of NSE and bias-related diagnostics and had the smallest station-averaged absolute peak-flow bias (|RBpeak| = 15.52%), although MIROC-INTEG-LAND showed slightly higher station-averaged correlations across the daily, monthly, and seasonal-cycle evaluations. In contrast, H08 and ORCHIDEE-MICT showed the weakest overall performance, with negative NSE values and large peak-flow biases.
Overall, no single configuration performed uniformly well across temporal scales and flow regimes. The supplementary 1976–2001 common-period analysis yielded broadly consistent cross-gauge summary patterns, indicating that the principal
Figure 6 results were robust to the unequal gauge-specific evaluation windows. The results therefore provide a cross-diagnostic benchmark for application-specific model selection and targeted model improvement in regional discharge simulation and peak-flow-related applications. Several limitations should be noted explicitly: the evaluation was restricted to four Yangtze mainstem gauges and therefore provides limited coverage of tributary, headwater, high-elevation, and reservoir-regulated sub-basin processes; the primary cross-gauge summaries combine gauge-specific records from different historical periods; the common GSWP3-W5E5 forcing dataset was not independently evaluated; and no model was recalibrated specifically for the Yangtze River Basin. Accordingly, the apparent model rankings should be interpreted as diagnostics of the evaluated standard ISIMIP3a configurations rather than as a controlled comparison of intrinsic model structures or as the upper bound of performance achievable after basin-specific calibration.