Enhancing Daily Runoff Prediction via Uniform Design and Meta-Learning Integrated Hyperparameter Optimization Embedded in Transformer
Abstract
1. Introduction
2. Materials and Methodology
2.1. Study Area
2.2. Data Sources and Preprocessing
2.3. Constructing Runoff Prediction Model
2.3.1. UD-ML Optimization Algorithm
2.3.2. Modified Transformer Model
2.3.3. UD-ML-Transformer Model
- (1)
- Based on the 10 selected hyperparameters of the modified Transformer, an initial UD matrix was constructed via the good lattice point method. Taking the minimum centered discrepancy as the optimization objective, the optimal generator vector was identified through column-wise replacement and iterative search, ensuring uniform sampling distribution in the high-dimensional parameter space. Accordingly, two UD tables with 10 factors and 100/2000 levels were established. The discrete level indices were linearly mapped to the actual value ranges of each hyperparameter, finally yielding 100 and 2000 valid hyperparameter configuration sets for model training.
- (2)
- The modified Transformer model was trained and tested for 100 epochs using the 100 hyperparameter trial configurations, and the corresponding model training results were collected. These hyperparameter trial configurations and their associated performance results were combined to form the training set for the meta-learner.
- (3)
- Four hyperparameters with significant impacts on the MLP meta-learner performance were selected as optimization targets, namely the number of neurons in the hidden layer, the number of hidden layers, learning rate, and optimizer. A UD table with 4 factors and 100 levels was generated using the UD method, and 300 hyperparameter trial configurations were derived by setting three groups of different value ranges for the selected hyperparameters.
- (4)
- The meta-learner training set was input into the MLP meta-learner, which was then trained for 500 epochs to obtain a well-trained model. Subsequently, the 2000 hyperparameter trial configurations were fed into this trained meta-learner to generate simulated results. Only the best simulated result and its corresponding hyperparameter configuration were retained. By training the meta-learner with 300 hyperparameter trial configurations, one optimal simulated prediction result and its corresponding configuration were obtained per trial, resulting in a total of 300 optimal configurations.
- (5)
- Duplicate configurations were removed from the 300 optimal outcomes, and the remaining configurations were used as the parameter tuning schemes for the Transformer model. The Transformer model was trained and tested for 100 epochs using each of these configurations, and the hyperparameter scheme yielding the best test set result was identified as the optimal hyperparameter configuration for the Transformer model.
| Algorithm 1. Meta-learner based on uniform design | |
| Input: Training dataset Dtrain: 100 hyperparameter trial configurations for Transformer and their corresponding training results (NSE) Hyperparameter candidate table H: 300 hyperparameter trial configurations for meta-learners Candidate scheme set to be predicted Dpred: 2000 hyperparameter trial configurations for Transformer | |
| Output: Optimal hyperparameter set Dpred′ and its predicted accuracy a∗ Best model parameters θ∗ for each evaluated scheme | |
| 1: | Split Dtrain into training (80%) and validation (20%) sets |
| 2: | for each hyperparameter scheme hi in H do |
| 3: | Create MLP model with He initialization |
| 4: | Set loss function to MAE (L1Loss) and L2 regularization (λ = 0.001) |
| 5: | Initialize best_loss ← ∞, best_params ← None |
| 6: | Initialize empty lists for training and validation losses |
| 7: | Initialize early stopping flag stop ← False |
| 8: | for epoch e = 1 to 500 do |
| 9: | Perform one training step (forward, backward, update) |
| 10: | Compute validation loss lval |
| 11: | if lval < best_loss then |
| 12: | best_loss ← lval |
| 13: | Save current model parameters as best_params |
| 14: | end if |
| 15: | if e mod 10 = 0 then |
| 16: | Find minimum validation loss among last 10 epochs, denote as lmin at epoch ek |
| 17: | Store (lmin, ek) in list M |
| 18: | if ∣M∣ ≥ 3 then |
| 19: | Let (A, eA), (B, eB), (C, eC) be the last three entries |
| 20: | if B < A and B < C and eA < eB <eC then |
| 21: | Stop ← True, stop_epoch ← eB |
| 22: | break |
| 23: | end if |
| 24: | end if |
| 25: | end if |
| 26: | end for |
| 27: | Load best_params into model |
| 28: | Predict accuracies a∗ on Dpred using the trained model |
| 29: | Identify index j∗ = argmax a∗ |
| 30: | Return Dpred′, a∗ |
| 31: | end for |
2.4. Shapely Additive Explanations (SHAP)
2.5. Evaluation Metrics
3. Results and Discussion
3.1. Advantages of UD-ML-Transformer Model
3.2. Robustness of Peak Flow Prediction
3.3. Hyperparameter Optimization for Transformer-Based Models
3.4. Explainability of UD-ML-Transformer Model
4. Model Applicability and Limitation
4.1. Applicability of UD-ML-Transformer Model
4.2. Limitations and Outlooks
5. Conclusions
- (1)
- The proposed UD-ML-Transformer model outperformed five benchmark models in providing reliable and accurate runoff predictions. The hyperparameter optimization strategy integrating UD and ML consistently enhanced the predictive performance of both Transformer and RWKV architectures. Comprehensive performance comparison further verified the stability and superiority of the proposed UD-ML optimization strategy.
- (2)
- The proposed UD-ML-Transformer model also exhibited robust and reliable performance in peak runoff prediction. Benefiting from the synergistic optimization of optimization strategies, the model can effectively capture the complex nonlinear characteristics of peak runoff, thereby substantially improving its practical applicability for flood simulation and early warning scenarios.
- (3)
- The proposed UD-ML-Transformer model demonstrated strong regional generalization capability, as verified by cross-watershed validation in the Ford River Watershed. The model maintained stable performance superiority across watersheds with heterogeneous hydrological and topographic characteristics, demonstrating promising extensibility for practical applications in diverse basins.
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Sivakumar, B. Global climate change and its impacts on water resources planning and management: Assessment and challenges. Stoch. Environ. Res. Risk Assess. 2011, 25, 583–600. [Google Scholar] [CrossRef] [Scilit]
- Juutinen, A.; Virk, Z.; Huuki, H.; Ruokamo, E.; Kopsakangas-Savolainen, M.; Torabi Haghighi, A.; Marttila, H. Impact of environmental flow policy on power system balancing costs and river ecosystem service benefits. Water Resour. Econ. 2025, 52, 100269. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Ji, G.; Hu, Y. Effect of vegetation growth, agricultural irrigation and climatic variability on streamflow in Wujiang, China. Forests 2024, 15, 1928. [Google Scholar] [CrossRef] [Scilit]
- Qi, J.; Yan, F.; Tian, Q.; Yang, C.; Tian, Y.; Li, X.; Guo, L.; Ma, Q.; Ma, Y. Analysis of high–low runoff encounters between the water source and receiving areas in the Xinyang urban water supply project. Water 2025, 17, 2618. [Google Scholar] [CrossRef] [Scilit]
- Tian, W.; Wang, W.; Wang, Y.; Shi, C.; Ma, Q. Accurate runoff prediction in nonlinear and nonstationary environments using a novel hybrid model. J. Hydrol. 2025, 662, 133949. [Google Scholar] [CrossRef] [Scilit]
- Fatichi, S.; Vivoni, E.R.; Ogden, F.L.; Ivanov, V.Y.; Mirus, B.; Gochis, D.; Downer, C.W.; Camporese, M.; Davison, J.H.; Ebel, B.; et al. An overview of current applications, challenges, and future trends in distributed process-based models in hydrology. J. Hydrol. 2016, 537, 45–60. [Google Scholar] [CrossRef] [Scilit]
- Singh, V.P. Hydrologic modeling: Progress and future directions. Geosci. Lett. 2018, 5, 15. [Google Scholar] [CrossRef] [Scilit]
- Ng, K.W.; Huang, Y.F.; Koo, C.H.; Chong, K.L.; El-Shafie, A.; Najah Ahmed, A. A review of hybrid deep learning applications for streamflow forecasting. J. Hydrol. 2023, 625, 130141. [Google Scholar] [CrossRef] [Scilit]
- Crawford, N.H.; Linsley, R.K. Digital Simulation in Hydrology: Stanford Watershed Model IV; Technical Report No. 39; Stanford University: Stanford, CA, USA, 1966. [Google Scholar]
- McCuen, R.H. A Guide to Hydrologic Analysis Using SCS Methods; Prentice Hall: Englewood Cliffs, NJ, USA, 1982. [Google Scholar]
- Zhao, R.J. Brief description of rainfall-runoff watershed models. People’s Yellow River 1983, 2, 40–43. (In Chinese) [Google Scholar]
- Arnold, J.G.; Srinivasan, R.; Muttiah, R.S.; Williams, J.R. Large area hydrologic modeling and assessment part I: Model development. J. Am. Water Resour. Assoc. 1998, 34, 73–89. [Google Scholar] [CrossRef] [Scilit]
- Gochis, D.J.; Barlage, M.; Dugger, A.; Cabell, R.; Casali, M.; FitzGerald, K.; McAllister, M.; McCreight, J.; RafieeiNasab, A.; Read, L.; et al. The WRF-Hydro® Modeling System Technical Description (Version 5.1.1); National Center for Atmospheric Research: Boulder, CO, USA, 2021. [Google Scholar] [CrossRef]
- Kuffour, B.N.O.; Engdahl, N.B.; Woodward, C.S.; Maxwell, R.M.; Kollet, S.J. Simulating coupled surface–subsurface flows with ParFlow v3.5.0: Capabilities, applications, and ongoing development of an open-source, massively parallel, integrated hydrologic model. Geosci. Model Dev. 2020, 13, 1373–1397. [Google Scholar] [CrossRef] [Scilit]
- Wagener, T.; Sivapalan, M.; Troch, P.A.; McGlynn, B.L.; Harman, C.J.; Gupta, H.V.; Kumar, P.; Rao, P.S.C.; Basu, N.B.; Wilson, J.S. The future of hydrology: An evolving science for a changing world. Water Resour. Res. 2010, 46, W05301. [Google Scholar] [CrossRef] [Scilit]
- Su, Y.; Ding, Z.; Zhang, R.; Tang, W.; Huang, W.; Wang, Z.; Zhao, K.; Wang, X.; Liu, S.; Li, Y. High-efficiency organic solar cells processed from a halogen-free solvent system. Sci. China Chem. 2023, 66, 2380–2388. [Google Scholar] [CrossRef] [Scilit]
- Valipour, M.; Banihabib, M.E.; Behbahani, S.M.R. Comparison of the ARMA, ARIMA, and the autoregressive artificial neural network models in forecasting the monthly inflow of Dez dam reservoir. J. Hydrol. 2013, 476, 433–441. [Google Scholar] [CrossRef] [Scilit]
- Zhang, G.P. Time series forecasting using a hybrid ARIMA and neural network model. Neurocomputing 2003, 50, 159–175. [Google Scholar] [CrossRef] [Scilit]
- Sabzipour, B.; Arsenault, R.; Troin, M.; Martel, J.-L.; Brissette, F. Sensitivity analysis of the hyperparameters of an ensemble Kalman filter application on a semi-distributed hydrological model for streamflow forecasting. J. Hydrol. 2023, 626, 130251. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Tian, W.; Xu, D.; Chau, K.; Ma, Q.; Liu, C. Muskingum models’ development and their parameter estimation: A state-of-the-art review. Water Resour. Manag. 2023, 37, 3129–3150. [Google Scholar] [CrossRef] [Scilit]
- Jahangir, M.S.; You, J.; Quilty, J. A quantile-based encoder-decoder framework for multi-step ahead runoff forecasting. J. Hydrol. 2023, 619, 129269. [Google Scholar] [CrossRef] [Scilit]
- Nourani, V. An emotional ANN (EANN) approach to modeling rainfall-runoff process. J. Hydrol. 2017, 544, 267–277. [Google Scholar] [CrossRef] [Scilit]
- Ren, M.; Sun, W.; Chen, S.; Zeng, D.; Xie, Y. Inconsistent monthly runoff prediction models using mutation tests and machine learning. Water Resour. Manag. 2024, 38, 5235–5254. [Google Scholar] [CrossRef] [Scilit]
- Van, S.P.; Le, H.M.; Thanh, D.V.; Dang, T.D.; Loc, H.H.; Anh, D.T. Deep learning convolutional neural network in rainfall–runoff modelling. J. Hydroinf. 2020, 22, 541–561. [Google Scholar] [CrossRef] [Scilit]
- Gao, S.; Huang, Y.; Zhang, S.; Han, J.; Wang, G.; Zhang, M.; Lin, Q. Short-term runoff prediction with GRU and LSTM networks without requiring time step optimization during sample generation. J. Hydrol. 2020, 589, 125188. [Google Scholar] [CrossRef] [Scilit]
- Man, Y.; Yang, Q.; Shao, J.; Wang, G.; Bai, L.; Xue, Y. Enhanced LSTM model for daily runoff prediction in the upper Huai River basin, China. Engineering 2023, 24, 229–238. [Google Scholar] [CrossRef] [Scilit]
- Le, X.-H.; Nguyen, D.-H.; Jung, S.; Yeon, M.; Lee, G. Comparison of deep learning techniques for river streamflow forecasting. IEEE Access 2021, 9, 71805–71820. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Jiang, J.; Xu, H.; Xu, X.; Cui, Y.; Wu, J. Transformer-based fused attention combined with CNNs for image classification. Neural Process. Lett. 2023, 55, 11905–11919. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Shao, S.; Bai, Y.; Deng, J.; Lin, Y. Multiscale wavelet graph AutoEncoder for multivariate time-series anomaly detection. IEEE Trans. Instrum. Meas. 2023, 72, 1–11. [Google Scholar] [CrossRef] [Scilit]
- Sun, W.; Chang, L.-C.; Chang, F.-J. Deep dive into predictive excellence: Transformer’s impact on groundwater level prediction. J. Hydrol. 2024, 636, 131250. [Google Scholar] [CrossRef] [Scilit]
- Wei, X.; Wang, G.; Schmalz, B.; Hagan, D.F.T.; Duan, Z. Evaluation of transformer model and self-attention mechanism in the Yangtze River basin runoff prediction. J. Hydrol. Reg. Stud. 2023, 47, 101438. [Google Scholar] [CrossRef] [Scilit]
- Yin, H.; Guo, Z.; Zhang, X.; Chen, J.; Zhang, Y. RR-former: Rainfall-runoff modeling based on transformer. J. Hydrol. 2022, 609, 127781. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Xu, J.; Wang, J.; Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. arXiv 2021, arXiv:2106.13008. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Yan, J. Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting. In Proceedings of the 11th International Conference on Learning Representations (ICLR 2023), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Zhang, X.; Liu, F.; Yin, Q.; Qi, Y.; Sun, S. A runoff prediction method based on hyperparameter optimisation of a kernel extreme learning machine with multi-step decomposition. Sci. Rep. 2023, 13, 19341. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rong, G.; Li, K.; Su, Y.; Tong, Z.; Liu, X.; Zhang, J.; Zhang, Y.; Li, T. Comparison of tree-structured parzen estimator optimization in three typical neural network models for landslide susceptibility assessment. Remote Sens. 2021, 13, 4694. [Google Scholar] [CrossRef] [Scilit]
- Mirjalili, S.; Mirjalili, S.M.; Lewis, A. Grey wolf optimizer. Adv. Eng. Softw. 2014, 69, 46–61. [Google Scholar] [CrossRef] [Scilit]
- Kennedy, J.; Eberhart, R. Particle swarm optimization. In Proceedings of the ICNN’95—International Conference on Neural Networks, Perth, Australia, 27 November–1 December 1995; pp. 1942–1948. [Google Scholar] [CrossRef] [Scilit]
- Xue, J.; Shen, B. A novel swarm intelligence optimization approach: Sparrow search algorithm. Syst. Sci. Control Eng. 2020, 8, 22–34. [Google Scholar] [CrossRef] [Scilit]
- Zhang, L.; Zeng, T.; Wang, L.; Li, L. Advancing seismic landslide susceptibility modeling: A comparative evaluation of deep learning models through particle swarm optimization. Earth Sci. Inf. 2024, 17, 3547–3566. [Google Scholar] [CrossRef] [Scilit]
- Mabdeh, A.N.; Ajin, R.S.; Razavi-Termeh, S.V.; Ahmadlou, M.; Al-Fugara, A. Enhancing the performance of machine learning and deep learning-based flood susceptibility models by integrating grey wolf optimizer (GWO) algorithm. Remote Sens. 2024, 16, 2595. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Wang, X.; Li, H.; Sun, S.; Liu, F. Monthly runoff prediction based on a coupled VMD-SSA-BiLSTM model. Sci. Rep. 2023, 13, 13149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, B.-J.; Sun, G.-L.; Li, Y.-P.; Zhang, X.-L.; Huang, X.-D. A hybrid variational mode decomposition and sparrow search algorithm-based least square support vector machine model for monthly runoff forecasting. Water Supply 2022, 22, 5698–5715. [Google Scholar] [CrossRef] [Scilit]
- Samantaray, S.; Das, S.S.; Sahoo, A.; Satapathy, D.P. Monthly runoff prediction at Baitarani River basin by support vector machine based on salp swarm algorithm. Ain Shams Eng. J. 2022, 13, 101732. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Fang, K.T. On uniform distribution and experimental design (number-theoretic method). Chin. Sci. Bull. 1981, 26, 65–70. (In Chinese) [Google Scholar]
- Elsawah, A.M. A novel non-heuristic search technique for constructing uniform designs with a mixture of two- and four-level factors: A simple industrial applicable approach. J. Korean Stat. Soc. 2022, 51, 716–757. [Google Scholar] [CrossRef] [Scilit]
- Lai, J.; Fang, K.-T.; Peng, X.; Lin, Y. Construction of uniform designs over continuous domain in computer experiments. Commun. Stat. Simul. Comput. 2024, 53, 130–146. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Xu, B.; Sun, G.; Yang, S. A two-phase differential evolution for uniform designs in constrained experimental domains. IEEE Trans. Evol. Comput. 2017, 21, 665–680. [Google Scholar] [CrossRef] [Scilit]
- Bai, D.; Ma, S.; Yang, X.; Ma, D.; Ma, X.; Ma, H. A recommendation model for optimizing transfer learning hyper-parameter settings in building heat load prediction with limited data samples. Energy Build. 2024, 325, 115021. [Google Scholar] [CrossRef] [Scilit]
- Finn, C.; Abbeel, P.; Levine, S. Model-agnostic meta-learning for fast adaptation of deep networks. In Proceedings of the 34th International Conference on Machine Learning (ICML 2017), Sydney, Australia, 6–11 August 2017; pp. 1126–1135. [Google Scholar]
- Ye, L.; Wang, W.; Sun, H.; Ye, W.; Hou, Y.; Zhang, Y.; Zhang, Y.; Ren, G.; Gao, Z.; Qu, X. A meta-learning approach for multicenter and small-data single-cell image analysis. Anal. Chem. 2025, 97, 16812–16821. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yoon, J.; Kim, T.; Dia, O.; Kim, S.; Bengio, Y.; Ahn, S. Bayesian model-agnostic meta-learning. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, Montréal, QC, Canada, 3–8 December 2018. [Google Scholar]
- Shahriari, B.; Swersky, K.; Wang, Z.; Adams, R.P.; de Freitas, N. Taking the human out of the loop: A review of Bayesian optimization. Proc. IEEE 2016, 104, 148–175. [Google Scholar] [CrossRef] [Scilit]
- Peng, B.; Goldstein, D.; Anthony, Q.; Albalak, A.; Alcaide, E.; Biderman, S.; Cheah, E.; Du, X.; Ferdinan, T.; Hou, H.; et al. Eagle and finch: RWKV with matrix-valued states and dynamic recurrence. arXiv 2024, arXiv:2404.05892. [Google Scholar] [CrossRef] [Scilit]
- Addor, N.; Newman, A.J.; Mizukami, N.; Clark, M.P. CAMELS: Catchment Attributes and Meteorology for Large-Sample Studies [Dataset]; UCAR/NCAR: Boulder, CO, USA, 2017. [Google Scholar] [CrossRef]
- Plaut, D.C.; Hinton, G.E. Learning sets of filters using back-propagation. Comput. Speech Lang. 1987, 2, 35–61. [Google Scholar] [CrossRef] [Scilit]
- Granata, F.; Zhu, S.; Di Nunno, F. Advanced streamflow forecasting for central European rivers: The cutting-edge Kolmogorov-Arnold networks compared to transformers. J. Hydrol. 2024, 645, 132175. [Google Scholar] [CrossRef] [Scilit]
- Moosavi, V.; Mostafaei, S.; Berndtsson, R. Temporal cluster-based local deep learning or signal processing-temporal convolutional transformer for daily runoff prediction? Appl. Soft Comput. 2024, 155, 111425. [Google Scholar] [CrossRef] [Scilit]
- Granata, F.; Di Nunno, F. Forecasting short- and medium-term streamflow using stacked ensemble models and different meta-learners. Stoch. Environ. Res. Risk Assess. 2024, 38, 3481–3499. [Google Scholar] [CrossRef] [Scilit]
- Abdoulhalik, A.; Ahmed, A.A. A comparative analysis of advanced machine learning techniques for river streamflow time-series forecasting. Sustainability 2024, 16, 4005. [Google Scholar] [CrossRef] [Scilit]
- Lundberg, S.M.; Lee, S.I. A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. 2017, 30, 4765–4774. [Google Scholar]
- Moriasi, D.N.; Arnold, J.G.; Van Liew, M.W.; Bingner, R.L.; Harmel, R.D.; Veith, T.L. Model evaluation guidelines for systematic quantification of accuracy in watershed simulations. Trans. ASABE 2007, 50, 885–900. [Google Scholar] [CrossRef] [Scilit]
- Botterill, T.E.; McMillan, H.K. Using machine learning to identify hydrologic signatures with an encoder–decoder framework. Water Resour. Res. 2023, 59, e2022WR033091. [Google Scholar] [CrossRef] [Scilit]
- Ji, Y.; Zhang, H.; Zhang, Z.; Liu, M. CNN-based encoder-decoder networks for salient object detection: A comprehensive review and recent advances. Inf. Sci. 2021, 546, 835–857. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Y.; Wu, X.; Zhang, W.; Lan, P.; Qin, G.; Li, X.; Li, H. A deep learning-based probabilistic approach to flash flood warnings in mountainous catchments. J. Hydrol. 2025, 652, 132677. [Google Scholar] [CrossRef] [Scilit]
- Han, J.; Liu, Z.; Woods, R.; McVicar, T.R.; Yang, D.; Wang, T.; Hou, Y.; Guo, Y.; Li, C.; Yang, Y. Streamflow seasonality in a snow-dwindling world. Nature 2024, 629, 1075–1081. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, L.; Wang, J.; Lu, T.; He, W.; Zou, X.; Xia, H.; Albano, R.; Ozga-Zielinski, B.; Adamowski, J.; Feng, Q. Impact of the soil freeze-thaw process on runoff generation and water balance in an alpine region of the northeast Qinghai-Tibet Plateau. Agric. Water Manag. 2026, 325, 110191. [Google Scholar] [CrossRef] [Scilit]
- Xie, S.; Xie, Y.; Zhang, Y.; Li, J.; Wang, G.; Zeng, C. Connecting effects of precipitation, soil hydrological processes, and groundwater dynamics in a continuous permafrost catchment on runoff of northeastern Qinghai-Tibet Plateau. Glob. Planet. Change 2026, 260, 105396. [Google Scholar] [CrossRef] [Scilit]








| Models | NSE | MSE | RMSE | MAE |
|---|---|---|---|---|
| PSO-Transformer | 0.867 | 0.006 | 0.075 | 0.039 |
| UD-Transformer | 0.903 | 0.004 | 0.064 | 0.035 |
| UD-ML-Transformer | 0.906 | 0.004 | 0.062 | 0.034 |
| PSO-RWKV | 0.824 | 0.007 | 0.085 | 0.043 |
| UD-RWKV | 0.883 | 0.005 | 0.068 | 0.038 |
| UD-ML-RWKV | 0.887 | 0.004 | 0.067 | 0.037 |
| Models | MSE | RMSE | MAE |
|---|---|---|---|
| PSO-Transformer | 0.025 | 0.159 | 0.095 |
| UD-Transformer | 0.014 | 0.118 | 0.074 |
| UD-ML-Transformer | 0.013 | 0.117 | 0.063 |
| PSO-RWKV | 0.016 | 0.128 | 0.078 |
| UD-REKV | 0.015 | 0.120 | 0.073 |
| UD-ML-REKV | 0.014 | 0.118 | 0.066 |
| Hyperparameters | PSO-Transformer | UD-Transformer | UD-ML-Transformer |
|---|---|---|---|
| Data preprocessing method | Normalization | Normalization | Standardization |
| Embedding dimension | 80 | 64 | 80 |
| Number of multi-head attention heads | 4 | 4 | 2 |
| Encoder–decoder layer count | 3 | 4 | 2 |
| Fully connected layer neuron count | 97 | 189 | 197 |
| Batch size | 16 | 256 | 64 |
| Input sequence length | 90 | 60 | 30 |
| Learning rate | 0.0001 | 0.001 | 0.001 |
| Optimizer type | Adam | Adam | Adam |
| Dropout rate | 0.2 | 0.1 | 0.1 |
| Models | NSE | MSE | RMSE | MAE |
|---|---|---|---|---|
| PSO-Transformer | 0.783 | 0.179 | 0.422 | 0.187 |
| UD-Transformer | 0.854 | 0.116 | 0.340 | 0.154 |
| UD-ML-Transformer | 0.890 | 0.088 | 0.296 | 0.141 |
| Hyperparameters | PSO-Transformer | UD-Transformer | UD-ML-Transformer |
|---|---|---|---|
| Data preprocessing method | Normalization | Standardization | Standardization |
| Embedding dimension | 64 | 96 | 80 |
| Number of multi-head attention heads | 4 | 8 | 2 |
| Encoder–decoder layer count | 4 | 3 | 3 |
| Fully connected layer neuron count | 95 | 89 | 208 |
| Batch size | 32 | 64 | 64 |
| Input sequence length | 60 | 60 | 30 |
| Learning rate | 0.0001 | 0.001 | 0.001 |
| Optimizer type | Adam | Adam | Adam |
| Dropout rate | 0.3 | 0.3 | 0.1 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, W.; Li, L.; Su, D.; Zhang, X.; Tong, H.; Shao, T.; Fan, J. Enhancing Daily Runoff Prediction via Uniform Design and Meta-Learning Integrated Hyperparameter Optimization Embedded in Transformer. Hydrology 2026, 13, 201. https://doi.org/10.3390/hydrology13080201
Wang W, Li L, Su D, Zhang X, Tong H, Shao T, Fan J. Enhancing Daily Runoff Prediction via Uniform Design and Meta-Learning Integrated Hyperparameter Optimization Embedded in Transformer. Hydrology. 2026; 13(8):201. https://doi.org/10.3390/hydrology13080201
Chicago/Turabian StyleWang, Wenxue, Liuyang Li, Donghui Su, Xin Zhang, Haibin Tong, Tiantian Shao, and Jiaxin Fan. 2026. "Enhancing Daily Runoff Prediction via Uniform Design and Meta-Learning Integrated Hyperparameter Optimization Embedded in Transformer" Hydrology 13, no. 8: 201. https://doi.org/10.3390/hydrology13080201
APA StyleWang, W., Li, L., Su, D., Zhang, X., Tong, H., Shao, T., & Fan, J. (2026). Enhancing Daily Runoff Prediction via Uniform Design and Meta-Learning Integrated Hyperparameter Optimization Embedded in Transformer. Hydrology, 13(8), 201. https://doi.org/10.3390/hydrology13080201

