Benchmarking Deep Learning Against Statistical Baselines and a Physical Climate-Model Comparator for Station-Scale Meteorological Forecasting: A 100-Station Study from the Western Balkans
Abstract
1. Introduction
1.1. Deep Learning for Time Series Forecasting: State of the Art
1.2. The Benchmarking Gap: Statistical Rigour and Physical-Model Baselines
1.3. Meteorological Station Networks as an Evaluation Environment
1.4. Aims
1.5. Main Contributions
- (1)
- We present, to our knowledge, one of the first statistically comprehensive evaluation frameworks to simultaneously compare five state-of-the-art deep learning forecasting architectures, classical statistical baselines, and a bias-corrected physically based CMIP6 climate model ensemble on a large, heterogeneous, 100-station meteorological panel.
- (2)
- We integrate four complementary statistical validation procedures—the Friedman test, Nemenyi post hoc analysis, pairwise Diebold-Mariano testing, and bootstrap confidence intervals—providing stronger, station-matched evidence of significance than the single-metric comparisons typical of the broader forecasting literature.
- (3)
- We show that deep learning architectures achieve lower temperature errors than the classical statistical and machine learning baselines, with four of the five architectures significantly outperforming climatology, whereas a simple climatological-mean baseline remains the strongest approach for precipitation forecasting, establishing that the AI forecasting advantage is variable-specific rather than universal.
- (4)
- We evaluate robustness through a rolling-origin temporal backtest across five non-overlapping two-year evaluation windows spanning 2011–2020 and through repetition of the physical-model comparison across all five individual members of an independent multi-model GCM ensemble.
- (5)
- We offer this comprehensive, temporally robust benchmarking protocol as a reusable evaluation framework for future environmental AI studies comparing data-driven and physically based forecasting approaches.
2. Materials and Methods
2.1. Dataset
2.2. Benchmarked Models
2.3. Training Configuration and Computational Environment
2.4. Statistical Significance Framework
Robustness Analysis Protocols
2.5. Multi-GCM Robustness Protocol
2.6. Seasonal Decomposition
2.7. Use of Generative AI Tools
3. Results
3.1. Headline Forecast Accuracy Comparison
3.2. Statistical Significance: Friedman Test and Critical-Difference Diagrams
3.3. Pairwise Diebold-Mariano Testing
3.4. Bootstrap Confidence Intervals
3.5. Robustness to Choice of Physical-Model Comparator
3.6. Seasonal Decomposition of Forecast Skill
3.7. Per-Station Heterogeneity
3.8. Rolling-Origin Temporal Robustness
3.9. Additional Robustness Analyses
3.10. Sensitivity to Random Initialisation (Multi-Seed Training)
4. Discussion
4.1. Inductive Bias and Architectural Performance
4.2. Why the Climatological Baseline Remains Competitive for Precipitation
4.3. Positioning Within the Broader Environmental AI Benchmarking Literature
4.4. Limitations
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| CMIP6 | Coupled Model Intercomparison Project, Phase 6 |
| DM | Diebold-Mariano |
| GCM | General Circulation Model |
| HAC | Heteroscedasticity- and Autocorrelation-Consistent |
| MAE | Mean Absolute Error |
| RMSE | Root Mean Square Error |
| SSP | Shared Socioeconomic Pathway |
| TFT | Temporal Fusion Transformer |
| N-HiTS | Neural Hierarchical Interpolation for Time Series |
| CD | Critical Difference |
| CI | Confidence Interval |
References
- Lim, B.; Arik, S.O.; Loeff, N.; Pfister, T. Temporal Fusion Transformers for Interpretable Multi-Horizon Time Series Forecasting. Int. J. Forecast. 2021, 37, 1748–1764. [Google Scholar] [CrossRef]
- Challu, C.; Olivares, K.G.; Oreshkin, B.N.; Garza, F.; Mergenthaler-Canseco, M.; Dubrawski, A. N-HiTS: Neural Hierarchical Interpolation for Time Series Forecasting. Proc. AAAI Conf. Artif. Intell. 2023, 37, 6989–6997. [Google Scholar] [CrossRef]
- Oreshkin, B.N.; Carpov, D.; Chapados, N.; Bengio, Y. N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting. In Proceedings of the International Conference on Learning Representations (ICLR), Addis Ababa, Ethiopia, 26–30 April 2020. [Google Scholar]
- Nie, Y.; Nguyen, N.H.; Sinthong, P.; Kalagnanam, J. A Time Series Is Worth 64 Words: Long-Term Forecasting with Transformers. In Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Das, A.; Kong, W.; Leach, A.; Mathur, S.; Sen, R.; Yu, R. Long-Term Forecasting with TiDE: Time-Series Dense Encoder. arXiv 2023, arXiv:2304.08424. [Google Scholar]
- Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; Zhang, W. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. Proc. AAAI Conf. Artif. Intell. 2021, 35, 11106–11115. [Google Scholar] [CrossRef]
- Wu, H.; Xu, J.; Wang, J.; Long, M. Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting. Adv. Neural Inf. Process. Syst. 2021, 34, 22419–22430. [Google Scholar]
- Zhou, T.; Ma, Z.; Wen, Q.; Wang, X.; Sun, L.; Jin, R. FEDformer: Frequency Enhanced Decomposed Transformer for Long-Term Series Forecasting. In Proceedings of the ICML 2022, Baltimore, MD, USA, 17–23 July 2022; pp. 27268–27286. [Google Scholar]
- Zeng, A.; Chen, M.; Zhang, L.; Xu, Q. Are Transformers Effective for Time Series Forecasting? Proc. AAAI Conf. Artif. Intell. 2023, 37, 11121–11128. [Google Scholar] [CrossRef]
- Wu, H.; Hu, T.; Liu, Y.; Zhou, H.; Wang, J.; Long, M. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. arXiv 2023, arXiv:2210.02186. [Google Scholar] [CrossRef]
- Wang, S.; Wu, H.; Shi, X.; Hu, T.; Luo, H.; Ma, L.; Zhang, J.Y.; Zhou, J. TimeMixer: Decomposable Multiscale Mixing for Time Series Forecasting. arXiv 2024, arXiv:2405.14616. [Google Scholar]
- Alharthi, M.; Mahmood, A. xLSTMTime: Long-Term Time Series Forecasting with xLSTM. AI 2024, 5, 1482–1495. [Google Scholar] [CrossRef]
- Demšar, J. Statistical Comparisons of Classifiers over Multiple Data Sets. J. Mach. Learn. Res. 2006, 7, 1–30. [Google Scholar]
- Nearing, G.; Cohen, D.; Dube, V.; Gauch, M.; Gilon, O.; Harrigan, S.; Hassidim, A.; Klotz, D.; Kratzert, F.; Metzger, A.; et al. Global Prediction of Extreme Floods in Ungauged Watersheds. Nature 2024, 627, 559–563. [Google Scholar] [CrossRef] [PubMed]
- Djalović, I.; Stojanović, D.B.; Stojsavljević, R.; Jovanović, M.; Nikolić, D. Projected Aridity Dynamics Across the Western Balkans Using a Multi-Model CMIP6 Ensemble and Short-Term AI Benchmarking. Atmosphere 2026, 17, 712. [Google Scholar] [CrossRef]
- Diebold, F.X.; Mariano, R.S. Comparing Predictive Accuracy. J. Bus. Econ. Stat. 1995, 13, 253–263. [Google Scholar] [CrossRef]
- Döscher, R.; Acosta, M.; Alessandri, A.; Anthoni, P.; Arsouze, T.; Bergman, T.; Bernardello, R.; Boussetta, S.; Caron, L.-P.; Carver, G.; et al. The EC-Earth3 Earth System Model for the Coupled Model Intercomparison Project 6. Geosci. Model Dev. 2022, 15, 2973–3020. [Google Scholar] [CrossRef]
- Eyring, V.; Bony, S.; Meehl, G.A.; Senior, C.A.; Stevens, B.; Stouffer, R.J.; Taylor, K.E. Overview of the Coupled Model Intercomparison Project Phase 6 (CMIP6) Experimental Design and Organization. Geosci. Model Dev. 2016, 9, 1937–1958. [Google Scholar] [CrossRef]
- Box, G.E.P.; Jenkins, G.M. Time Series Analysis: Forecasting and Control; Holden-Day: San Francisco, CA, USA, 1970. [Google Scholar]
- Breiman, L. Random Forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
- Hawkins, E.; Sutton, R. The Potential to Narrow Uncertainty in Regional Climate Predictions. Bull. Am. Meteorol. Soc. 2009, 90, 1095–1107. [Google Scholar] [CrossRef]
- Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar] [CrossRef]
- Lundberg, S.M.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. Adv. Neural Inf. Process. Syst. 2017, 30, 4765–4774. [Google Scholar]
- Das, A.; Kong, W.; Sen, R.; Zhou, Y. A Decoder-Only Foundation Model for Time-Series Forecasting. arXiv 2024, arXiv:2310.10688v4. [Google Scholar]
- Ansari, A.F.; Stella, L.; Turkmen, C.; Zhang, X.; Mercado, P.; Shen, H.; Shchur, O.; Rangapuram, S.S.; Arango, S.P.; Kapoor, S.; et al. Chronos: Learning the Language of Time Series. arXiv 2024, arXiv:2403.07815. [Google Scholar]
- Garza, A.; Mergenthaler-Canseco, M. TimeGPT-1. arXiv 2023, arXiv:2310.03589. [Google Scholar]
- Rasul, K.; Ashok, A.; Williams, A.R.; Ghonia, H.; Bhagwatkar, R.; Khorasani, A.; Darvishi Bayazi, M.J.; Adamopoulos, G.; Riachi, R.; Hassen, N.; et al. Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting. arXiv 2023, arXiv:2310.08278. [Google Scholar]
- Woo, G.; Liu, C.; Kumar, A.; Xiong, C.; Savarese, S.; Sahoo, D. Unified Training of Universal Time Series Forecasting Transformers. arXiv 2024, arXiv:2402.02592. [Google Scholar]
- Liang, Y.; Wen, H.; Nie, Y.; Jiang, Y.; Jin, M.; Song, D.; Pan, S.; Wen, Q. Foundation Models for Time Series Analysis: A Tutorial and Survey. arXiv 2024, arXiv:2403.14735. [Google Scholar]
- Jin, M.; Wang, S.; Ma, L.; Chu, Z.; Zhang, J.Y.; Shi, X.; Chen, P.-Y.; Liang, Y.; Li, Y.-F.; Pan, S. Time-LLM: Time Series Forecasting by Reprogramming Large Language Models. arXiv 2024, arXiv:2310.01728. [Google Scholar]




| Model | Input Information Used |
|---|---|
| Climatology | Station- and calendar-month-specific mean of training-period observations only. |
| SARIMA | Own historical series only, with built-in seasonal (period-12) autoregressive and moving-average terms; fitted independently per station. |
| Random Forest | Station identity (categorical), calendar month, calendar year, lag-1, lag-12, trailing 12-month rolling mean; trained globally across all stations. |
| TFT, N-HiTS, PatchTST, TiDE, xLSTM | Own historical series only (input window = 36 months, horizon = 24 months); no explicit calendar, station-identity, or engineered lag covariates; trained as a single global model per architecture across all 100 stations (unique_id, ds, y format). |
| Architecture | Inductive Bias | Key Structural Hyperparameters | Parameters |
|---|---|---|---|
| TFT | Attention + recurrent | hidden = 128; heads = 4; LSTM layers = 1; GRN gating | 870, 158 |
| N-HiTS | Hierarchical MLP | 3 stacks; mlp_units = [128, 128]/stack | 177, 558 |
| PatchTST | Patch-based attention | hidden = 128; heads = 16; layers = 3; patch = 16; stride = 8 | 412, 443 |
| TiDE | Dense encoder–decoder | hidden = 128; enc/dec layers = 1; decoder_dim = 32 | 247, 708 |
| xLSTM | Extended recurrent (mLSTM) | hidden = 128; encoder blocks = 2; decoder layers = 2 | 239, 625 |
| Model | Temp. MAE (°C) | Temp. RMSE (deg °C) | Precip. MAE (mm) | Precip. RMSE (mm) |
|---|---|---|---|---|
| xLSTM | 1.39 | 1.98 | 44.43 | 68.80 |
| TiDE | 1.45 | 1.94 | 42.78 | 67.21 |
| TFT | 1.47 | 1.84 | 42.73 | 66.10 |
| N-HiTS | 1.52 | 1.90 | 42.96 | 67.36 |
| PatchTST | 1.62 | 1.98 | 42.81 | 66.45 |
| Random Forest | 1.79 | 2.35 | 49.04 | 72.12 |
| Climatology | 1.81 | 2.24 | 40.65 | 63.37 |
| SARIMA | 2.28 | 2.88 | 60.57 | 91.28 |
| Variable | Best Model | Benchmark | Mean DM | % Favouring Best | % Favouring Benchmark |
|---|---|---|---|---|---|
| Temperature | xLSTM | Climatology | −2.49 | 67% | 2% |
| Temperature | xLSTM | SARIMA | −3.14 | 87% | 3% |
| Temperature | xLSTM | Random Forest | −1.83 | 41% | 2% |
| Temperature | xLSTM | TFT | −0.87 | 13% | 2% |
| Temperature | xLSTM | N-HiTS | −1.30 | 25% | 2% |
| Temperature | xLSTM | PatchTST | −1.40 | 21% | 2% |
| Temperature | xLSTM | TiDE | −0.68 | 8% | 3% |
| Precipitation | Climatology | SARIMA | −2.63 | 64% | 0% |
| Precipitation | Climatology | Random Forest | −1.78 | 44% | 0% |
| Precipitation | Climatology | TFT | −1.05 | 16% | 1% |
| Precipitation | Climatology | N-HiTS | −0.85 | 20% | 0% |
| Precipitation | Climatology | PatchTST | −0.89 | 12% | 0% |
| Precipitation | Climatology | TiDE | −0.82 | 9% | 1% |
| Precipitation | Climatology | xLSTM | −1.19 | 17% | 0% |
| Model | Temp. MAE [95% CI] | Precip. MAE [95% CI] |
|---|---|---|
| xLSTM | 1.39 [1.27, 1.61] | 44.43 [39.38, 49.48] |
| TiDE | 1.45 [1.37, 1.59] | 42.81 [38.15, 47.57] |
| TFT | 1.47 [1.44, 1.52] | 42.76 [38.78, 47.22] |
| N-HiTS | 1.52 [1.49, 1.56] | 43.00 [38.38, 47.86] |
| PatchTST | 1.62 [1.55, 1.73] | 42.84 [38.45, 47.32] |
| Random Forest | 1.80 [1.73, 1.87] | 49.04 [44.16, 54.13] |
| Climatology | 1.81 [1.72, 1.95] | 40.67 [36.47, 45.43] |
| SARIMA | 2.29 [2.16, 2.44] | 60.65 [56.01, 66.54] |
| Model | Temp. MAE | Temp. RMSE | Precip. MAE | Precip. RMSE |
|---|---|---|---|---|
| xLSTM (best single-seed AI, temperature) | 1.39 | 1.98 | 44.43 | 68.80 |
| Climatology (best, precipitation) | 1.81 | 2.24 | 40.65 | 63.37 |
| CNRM-CM6-1 (best GCM, temp.) | 1.61 | 2.16 | 50.63 | 81.91 |
| IPSL-CM6A-LR (best GCM, precip.) | 1.96 | 2.55 | 47.48 | 78.24 |
| MPI-ESM1-2-HR (GCM) | 2.37 | 2.86 | 55.24 | 87.05 |
| EC-Earth3 (GCM) | 2.39 | 2.98 | 47.82 | 76.15 |
| MRI-ESM2-0 (GCM) | 2.50 | 3.18 | 51.60 | 82.12 |
| Variable | Season | xLSTM | TiDE | TFT | N-HiTS | PatchTST | Climatology | RF | SARIMA |
|---|---|---|---|---|---|---|---|---|---|
| Temperature | Summer | 1.20 | 1.17 | 1.18 | 1.16 | 1.73 | 1.66 | 1.26 | 1.51 |
| Temperature | Rest of year | 1.45 | 1.54 | 1.57 | 1.64 | 1.59 | 1.86 | 1.97 | 2.54 |
| Precipitation | Summer | 36.60 | 37.13 | 37.37 | 36.36 | 37.34 | 34.78 | 37.75 | 60.84 |
| Precipitation | Rest of year | 47.01 | 44.67 | 44.52 | 45.17 | 44.64 | 42.61 | 52.75 | 60.49 |
| Model | Variable | 2011–2012 | 2013–2014 | 2015–2016 | 2017–2018 | 2019–2020 | Mean +/− SD |
|---|---|---|---|---|---|---|---|
| Climatology | Temp. (°C) | 1.96 | 1.54 | 1.45 | 1.95 | 1.81 | 1.74 +/− 0.24 |
| TFT | Temp. (°C) | 1.94 | 1.76 | 1.18 | 1.61 | 1.47 | 1.60 +/− 0.29 |
| TiDE | Temp. (°C) | 1.88 | 1.77 | 1.32 | 1.66 | 1.45 | 1.62 +/− 0.23 |
| xLSTM | Temp. (°C) | 1.88 | 1.80 | 1.29 | 1.86 | 1.39 | 1.64 +/− 0.28 |
| N-HiTS | Temp. (°C) | 1.87 | 1.82 | 1.35 | 1.72 | 1.52 | 1.66 +/− 0.22 |
| PatchTST | Temp. (°C) | 1.89 | 1.94 | 1.25 | 1.88 | 1.62 | 1.72 +/− 0.29 |
| Random Forest | Temp. (°C) | 2.10 | 2.59 | 1.72 | 2.17 | 1.79 | 2.07 +/− 0.35 |
| SARIMA | Temp. (°C) | 2.40 | 2.32 | 2.01 | 2.52 | 2.28 | 2.31 +/− 0.19 |
| Climatology | Precip. (mm) | 40.09 | 46.00 | 37.32 | 34.86 | 40.65 | 39.78 +/− 4.18 |
| TFT | Precip. (mm) | 41.30 | 47.42 | 37.79 | 38.71 | 42.73 | 41.59 +/− 3.81 |
| TiDE | Precip. (mm) | 41.86 | 48.66 | 37.96 | 38.00 | 42.78 | 41.85 +/− 4.39 |
| N-HiTS | Precip. (mm) | 40.85 | 47.43 | 38.27 | 41.75 | 42.96 | 42.25 +/− 3.37 |
| xLSTM | Precip. (mm) | 42.10 | 51.17 | 41.89 | 40.05 | 44.40 | 43.92 +/− 4.34 |
| PatchTST | Precip. (mm) | 45.55 | 51.69 | 40.45 | 39.52 | 42.81 | 44.01 +/− 4.89 |
| Random Forest | Precip. (mm) | 64.07 | 54.10 | 72.09 | 45.00 | 49.01 | 56.85 +/− 11.11 |
| SARIMA | Precip. (mm) | 47.05 | 73.56 | 57.75 | 57.49 | 60.57 | 59.29 +/− 9.49 |
| Horizon | TFT | N-HiTS | PatchTST | TiDE | xLSTM |
|---|---|---|---|---|---|
| h = 1 | p < 0.0001 * (clim better) | p < 0.0001 * (clim better) | p < 0.0001 * (clim better) | p = 0.0009 * (clim better) | p < 0.0001 * (clim better) |
| h = 6 | p < 0.0001 * (model better) | p = 0.0002 * (model better) | p = 0.246 (model better) | p = 0.0004 * (model better) | p < 0.0001 * (model better) |
| h = 12 | p < 0.0001 * (model better) | p < 0.0001 * (model better) | p = 0.0003 * (model better) | p < 0.0001 * (model better) | p < 0.0001 * (model better) |
| h = 24 | p < 0.0001 * (clim better) | p < 0.0001 * (clim better) | p < 0.0001 * (clim better) | p < 0.0001 * (clim better) | p < 0.0001 * (clim better) |
| Model | Temperature MAE (degC) | Precipitation MAE (mm) |
|---|---|---|
| N-HiTS | 1.373 +/− 0.031 | 43.03 +/− 0.50 |
| TFT | 1.409 +/− 0.081 | 42.89 +/− 0.39 |
| TiDE | 1.480 +/− 0.037 | 43.07 +/− 0.17 |
| PatchTST | 1.517 +/− 0.133 | 43.40 +/− 0.48 |
| xLSTM | 1.548 +/− 0.060 | 43.61 +/− 0.38 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Nikolić, D.; Djalović, I.; Vitezović, I.; Stojanović, D.B.; Pavkov, S.; Stojsavljević, R.; Jovanović, M. Benchmarking Deep Learning Against Statistical Baselines and a Physical Climate-Model Comparator for Station-Scale Meteorological Forecasting: A 100-Station Study from the Western Balkans. AI 2026, 7, 329. https://doi.org/10.3390/ai7090329
Nikolić D, Djalović I, Vitezović I, Stojanović DB, Pavkov S, Stojsavljević R, Jovanović M. Benchmarking Deep Learning Against Statistical Baselines and a Physical Climate-Model Comparator for Station-Scale Meteorological Forecasting: A 100-Station Study from the Western Balkans. AI. 2026; 7(9):329. https://doi.org/10.3390/ai7090329
Chicago/Turabian StyleNikolić, Dalibor, Ivica Djalović, Ivan Vitezović, Dejan B. Stojanović, Sara Pavkov, Rastislav Stojsavljević, and Mlađen Jovanović. 2026. "Benchmarking Deep Learning Against Statistical Baselines and a Physical Climate-Model Comparator for Station-Scale Meteorological Forecasting: A 100-Station Study from the Western Balkans" AI 7, no. 9: 329. https://doi.org/10.3390/ai7090329
APA StyleNikolić, D., Djalović, I., Vitezović, I., Stojanović, D. B., Pavkov, S., Stojsavljević, R., & Jovanović, M. (2026). Benchmarking Deep Learning Against Statistical Baselines and a Physical Climate-Model Comparator for Station-Scale Meteorological Forecasting: A 100-Station Study from the Western Balkans. AI, 7(9), 329. https://doi.org/10.3390/ai7090329

