Next Article in Journal
Advancing Clear-Air Turbulence Detection with Hybrid Predictive Models for a Regional Aviation Corridor in Southeast Brazil
Previous Article in Journal
Numerical Weather Prediction of Hurricane Florence (2018) and Potential Climate Impacts Through Thermodynamic and Moisture Modification
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Enhancing the Temperature Forecast Accuracy of the ZJOCF Model Using AI-Based Station-Level Bias Correction

1
Longyou County Meteorological Bureau, Quzhou 324400, China
2
Plateau Atmosphere and Environment Key Laboratory of Sichuan Province, Sichuan Provincial Engineering Research Center for Meteorological Disaster Prediction and Early Warning, Chengdu Plain Urban Meteorology and Environment Observation and Research Station of Sichuan Province, School of Atmospheric Sciences, Chengdu University of Information Technology, Chengdu 610225, China
*
Author to whom correspondence should be addressed.
Atmosphere 2026, 17(5), 439; https://doi.org/10.3390/atmos17050439
Submission received: 9 March 2026 / Revised: 7 April 2026 / Accepted: 13 April 2026 / Published: 26 April 2026
(This article belongs to the Section Atmospheric Techniques, Instruments, and Modeling)

Abstract

Liuchun Lake area, located in the high-elevation and topographically complex western region of Zhejiang Province, exhibits temperature variability strongly influenced by terrain-induced dynamics and local microclimates. The Zhejiang Operational Consensus Forecasts (ZJOCF) model shows pronounced systematic biases in this area, making it difficult to meet the demand for short-term, fine-scale forecasts in cultural-tourism applications. Using observational data from four stations at different elevations, this study analyzes how ZJOCF temperature forecast errors vary with altitude, develops a station-level machine-learning temperature bias-correction model, and evaluates its performance in terms of accuracy, mean absolute error (MAE), error distribution, and control of extreme errors. Results show that the accuracy of the raw forecasts decreases significantly with increasing elevation, with high-altitude sites exhibiting distinct warm biases and strong fluctuations. After correction, the 72 h forecast accuracy at the four stations increases to 69–71% (up to 40.8% at the mountaintop station), MAE is reduced by more than 60% on average, extreme-error cases decrease by 40–60%, and the error distribution shifts from a scattered multi-peak pattern to a concentrated single-peak structure. These findings demonstrate that station-level machine-learning correction can effectively mitigate structural errors in ZJOCF temperature forecasts over complex terrain, providing a reliable technical pathway for refined meteorological services in mountainous regions.

1. Introduction

In recent years, the growing demand for fine-scale meteorological information across multiple sectors has made local-scale temperature forecasting an essential foundation for agriculture, tourism management, energy dispatching, and public safety. Agricultural production, for example, relies on field-level temperature forecasts to guide sowing schedules, frost prevention, and moisture regulation, as microclimatic conditions often deviate substantially from regional averages [1]. In emergency management, extreme temperature events such as cold waves and heatwaves can severely threaten human safety and disrupt social operations, thereby requiring short-term, high-accuracy temperature predictions [2]. The need for ultra-local refined temperature forecasts is particularly prominent in high-elevation scenic areas with complex terrain, where strong topographic relief and heterogeneous surface conditions induce significant horizontal temperature gradients [3].
Although numerical weather prediction models have made notable advances in spatial resolution and data assimilation, providing a solid basis for refined forecasting [4], substantial gaps remain in accurately representing near-surface temperature variability over complex mountainous regions. Influenced by terrain shielding, radiative contrasts, wind-field structures, and heterogeneous land-surface conditions, many sub-kilometer-scale physical processes cannot be explicitly resolved by the models, resulting in cumulative systematic biases [5]. In the mountainous areas of southern China, forecasts produced by models with horizontal resolutions of 3 km still commonly exhibit warm or cold biases and struggle to capture detailed features such as valley cold-air pooling and slope-dependent temperature gradients [3]. Consequently, raw numerical weather prediction outputs are insufficient for meeting the requirements of ultra-local forecasting services, making targeted bias-correction strategies indispensable.
Traditional statistical post-processing techniques—such as Model Output Statistics [6], Kalman filtering [7], and ensemble averaging approaches [8]—have been widely applied in operational forecasting to improve model outputs through historical error statistics. However, because these methods generally rely on linear assumptions, they struggle to characterize the nonlinear structure of error fields in regions with complex terrain [9] and tend to perform less effectively under extreme temperature conditions [10]. In recent years, artificial intelligence (AI) techniques, including random forests, support vector machines, and deep learning, have shown remarkable potential in temperature bias correction. These methods can automatically learn sophisticated relationships between forecast errors and factors such as topography, radiation, and air-mass characteristics, achieving higher accuracy in predicting extreme temperatures [11,12]. Nevertheless, in mountainous areas where observations are sparse and elevation varies sharply, challenges such as limited training samples, insufficient representativeness of features, and constrained model generalization remain substantial [13].
To more effectively utilize numerical model element-based forecast products, the Zhejiang Meteorological Observatory has conducted comprehensive analyses of model forecast errors and developed corresponding correction strategies using observational data from township-level weather stations across Zhejiang Province. Based on these efforts, a dynamic multi-model integration workflow has been established. This workflow accounts for the consistency among different forecast variables and incorporates high-resolution data from the European Center for Medium-Range Weather Forecasts, the Japan Meteorological Agency, and the National Centers for Environmental Prediction. By applying error statistics across forecast lead times, the system dynamically adjusts the weights of individual models within the integrated product and performs fine-scale corrections on forecast variables using recent error characteristics. The resulting enhanced objective forecast product—Zhejiang Operational Consensus Forecasts (ZJOCF)—features a spatial resolution of 2.5 km and provides hourly predictions of key meteorological variables. It has now become a fundamental data source for regional operational forecasting in Zhejiang Province.
Liuchun Lake area, the study area considered in this paper, is situated in the high-elevation mountainous region of western Zhejiang, where steep terrain and highly variable microclimates lead to pronounced vertical temperature gradients driven by thermal circulations, localized subsidence, and valley cold-air pooling. With the rapid development of tourism in this area, the need for short-term, fine-scale forecasts of key meteorological variables—particularly temperature—has become increasingly urgent. However, under such complex topographic influences, the current ZJOCF model still exhibits marked systematic biases in the Liuchun Lake area, especially during nighttime radiative cooling and early-morning cold-air accumulation, which limits its effectiveness in real-time operational applications.
From a post-processing perspective, grid-based correction methods can preserve the spatial structure of forecast fields but depend heavily on dense observations, resulting in higher implementation costs and reduced timeliness. In contrast, station-based correction treats individual observation stations as the basic units, offering a simpler structure and higher operational efficiency, making it more suitable for mountainous regions where rapid response is essential [14]. Therefore, applying station-level corrections to ZJOCF temperature forecasts is of considerable practical significance and operational value for the Liuchun Lake area and other mountainous areas with similar complexity.
This study utilizes observational data from four weather stations distributed at different elevations in the Liuchun Lake area—Bajiaodian, Octagonal Palace, Liuchun Lake Mid-slope, and the Mountaintop station. We first conduct a systematic analysis of ZJOCF temperature forecast biases under complex terrain conditions and examine how elevation and local climate factors influence these errors. On this basis, we develop a machine-learning-based station-level correction model to improve the accuracy and stability of ZJOCF temperature forecasts in mountainous environments. Although the verification is conducted for the Liuchun Lake area, the main error characteristics identified in this study are closely related to general forecasting difficulties over complex terrain, including strong elevation dependence, local thermal contrasts, and terrain-induced microclimatic effects. By comparing forecast performance before and after correction, we assess the practical effectiveness of AI-based correction methods in complex mountainous regions and provide a feasible technical pathway and practical reference for refined meteorological services in Liuchun Lake and similar scenic areas.

2. Materials and Methods

2.1. Materials

This study employs temperature forecasts from the ZJOCF model together with in situ observations from surface weather stations as the fundamental data sources for constructing the temperature bias-correction model. The ZJOCF model, developed by the Zhejiang Meteorological Observatory, is operated within the provincial operational forecasting network of Zhejiang and is used to support routine forecasting at weather stations across the province. The model is issued twice daily at 08:00 and 20:00, providing forecasts with a lead time of up to 144 h at 3 h intervals. Its spatial resolution is 0.05° × 0.05°, and the model outputs include multiple meteorological variables such as 2 m air temperature, 10 m mean wind speed, and 3 h accumulated precipitation. The observational dataset is derived from four surface weather stations—Octagonal Palace, Bajiaodian, Liuchun Lake mid-slope, and Liuchun Lake mountaintop—each with an hourly temporal resolution. Recorded variables include air temperature, 24 h maximum temperature, precipitation, wind speed, and wind direction. The variables listed above describe the full set of available model outputs and station observations. However, for the purpose of this study, only the 2 m air temperature forecast and the corresponding observed air temperature were used as the correction target and verification variable.
To develop the training and testing datasets, the ZJOCF temperature forecasts were rigorously matched with station observations in both time and space. In space, each weather station was paired with the nearest ZJOCF grid point according to its geographic coordinates, and no spatial extrapolation was performed. In time, the model forecasts were matched to the corresponding observational records at the same valid time. This was followed by quality-control procedures to remove missing or anomalous records. Ultimately, continuous three-year matched samples from 1 January 2021 to 31 December 2023 were obtained for all four stations, yielding 2190 valid samples per station for model training and evaluation. The geographic characteristics of the four stations are as follows: Bajiaodian (119.13° E, 28.79° N; elevation ~273 m), Octagonal Palace (119.08° E, 28.79° N; elevation ~608 m), Liuchun Lake mid-slope (119.08° E, 28.78° N; elevation ~903 m), and Liuchun Lake mountaintop (119.07° E, 28.75° N; elevation ~1327 m). Their spatial distribution within the Liuchun Lake area and the surrounding complex terrain is illustrated in Figure 1. All station datasets have passed consistency checks and quality-control procedures, ensuring suitability for temperature bias correction and verification analyses. In addition, no prolonged snow-cover conditions were observed at the stations during the study period, so snow-cover effects were not analyzed separately.

2.2. Method

2.2.1. XGBoost

XGBoost (Extreme Gradient Boosting) is an ensemble learning framework built upon the principle of gradient boosting. It enhances overall predictive performance by sequentially adding multiple weak learners—typically regression trees. In each iteration, the model uses the residuals from the current prediction as the new learning target, fitting these residuals to progressively approach the true values and thereby improving predictive accuracy.
y ^ i = x i = i = 1 k f k x ,   f k     F
XGBoost represents the final prediction as the sum of multiple decision trees, where K denotes the number of trees and fk is the function corresponding to the k-th tree.
To achieve efficient optimization during training, XGBoost applies a second-order Taylor expansion to the objective function using both first- and second-order gradient information, while incorporating a regularization term to penalize model complexity. The resulting optimization objective balances model fit and structural simplicity.
Obj x = 1 2 j = 1 T G j 2 H j + ε + γ T
When splitting tree nodes, XGBoost evaluates the gain in the objective function before and after the split to determine whether the division is beneficial.
Gain = 1 2 G L 2 H L + ε + G R 2 H R + ε - G 2 H + ε
In this process, Gj and Hj represent the cumulative first- and second-order gradients, γ is the regularization term controlling model structure, and ε is the smoothing coefficient. By leveraging second-order approximation and explicit regularization, XGBoost effectively suppresses overfitting and enhances both generalization capability and training efficiency.

2.2.2. Bias-Correction Model

In this study, the Extreme Gradient Boosting (XGBoost) algorithm is employed to construct bias-correction models for temperature. All correction procedures were implemented in a Python 3.9 computing environment. The model development process begins with thorough data preprocessing, during which the model outputs are temporally and spatially matched with observational records from weather stations. This produces a unified dataset containing both simulated and observed values, forming the basis for subsequent learning and model formulation. During feature construction, multiple meteorological variables are examined to assess their correlation with the target variables. Observational factors exhibiting significant linear or nonlinear relationships with the correction targets are selected as feature inputs, while the original model outputs are retained to strengthen the model’s fitting capability. The correction target is the simultaneously observed air temperature.
The dataset is divided chronologically into training and testing subsets, with 80% of samples used for model training and hyperparameter tuning, and the remaining 20% reserved exclusively for accuracy validation. The XGBoost models are implemented using the xgboost library, with root-mean-square error (RMSE) specified as the optimization objective. A Bayesian optimization strategy is adopted to explore and select key hyperparameters—including max_depth, n_estimators, learning_rate, subsample, and colsample_bytree—to identify the configuration yielding the best overall performance. The optimized model is then applied to predict the test-set features, producing normalized correction results that are subsequently transformed back into physical units through inverse normalization. The reliability and improvement of the correction performance are evaluated using multiple metrics, including error (E), RMSE, mean absolute error (MAE), and accuracy, to ensure comprehensive verification of the model’s effectiveness.

2.2.3. Model Forecast Error Analysis and Evaluation Methods

To quantitatively assess the performance of the ZJOCF model in forecasting 2 m air temperature, temperature predictions at the grid points corresponding to the four stations—Bajiaodian, Octagonal Palace, Liuchun Lake mid-slope, and Liuchun Lake mountaintop—were extracted from the model’s gridded products and compared against in situ observations. The evaluation framework employs four statistical metrics: the forecast E, MAE, RMSE, and forecast accuracy. Together, these indicators provide a comprehensive depiction of bias characteristics and the overall reliability of the model’s temperature forecasts.
The forecast error Ei is defined as the difference between the model-predicted value and the observed value for each sample:
E i   =   F i   -   f i
In the formulas, Fi denotes the predicted air temperature, and fi represents the observed temperature at the station. For different forecast lead times, boxplots are constructed with forecast time on the horizontal axis and error E on the vertical axis, enabling a clear visualization of the statistical distribution and dispersion characteristics of temperature prediction errors. The MAE is used to quantify the overall magnitude of deviation and is defined as follows:
MAE   =   1 n i = 1 n E i
The forecast accuracy is evaluated according to the criteria specified in the Methods for Quality Verification of Medium-Range Weather Forecasting issued by the China Meteorological Administration [15]. A forecast is considered correct when the temperature error satisfies ∣Ei∣ < 2K. The proportion of all time steps that meet this condition is computed as the forecast accuracy. A line chart with forecast time on the horizontal axis and accuracy on the vertical axis is then plotted to illustrate the evolution of forecasting performance with lead time before temperature correction.

3. Results

3.1. Accuracy Characteristics and Correction Performance of ZJOCF Temperature Forecasts Under Different Elevation Backgrounds

To systematically assess the effectiveness of the correction method, the forecasts before and after correction were evaluated from four aspects: forecast accuracy, MAE, error-distribution characteristics, and extreme-error control. This section first examines the accuracy dimension by comparing the 72 h hourly temperature-forecast performance at four stations: Bajiaodian (273 m), Octagonal Palace (608 m), Mountainside (903 m), and Mountaintop (1327 m) (Figure 2). The results indicate that the raw forecast accuracy decreases markedly with increasing elevation, accompanied by substantially enhanced temporal fluctuations. The low-elevation stations, Bajiaodian and Octagonal Palace, exhibit relatively stable performance, with pre-correction accuracies of 63% and 60.5%, respectively. However, noticeable instability remains in the mid-to-late forecast periods—for example, the accuracy at Bajiaodian decreases to 55.1% at the 39th hour, while Octagonal Palace reaches a minimum of only 28.4% at the 51st hour. After correction, the accuracies at both stations increase significantly to approximately 69–70%, with markedly smoother temporal curves and prolonged periods of high accuracy.
In contrast, improvement at the high-elevation stations is even more pronounced. Before correction, the forecast accuracies at the Mountainside and summit stations are only about 18% and 17.8%, respectively, with the strongest fluctuations occurring during the 24–48 h period, where values remain below 20% for most hours. After correction, accuracy at the mid-slope station increases to 71.3%, and that at the summit station rises to 40.8%. Notably, the mid-slope station exhibits an improvement exceeding 50 percentage points, highlighting the strong capability of the correction method in mitigating systematic biases in high-altitude regions. Additionally, all four stations show substantial improvements in temporal consistency after correction—particularly Bajiaodian and Octagonal Palace, which maintain accuracies above 60% almost throughout the full 72 h period—demonstrating a marked enhancement in forecast stability and practical usability.
Further analysis of the MAE (Figure 3) shows that the original temperature errors increase markedly with elevation and exhibit pronounced discontinuities. Before correction, the average MAE at Bajiaodian and Octagonal Palace is approximately 1.9 °C and 2.3 °C, respectively, with certain periods reaching 2.8–3.7 °C. In contrast, errors at the mid-slope and summit stations are substantially larger, with peak values of 5.69 °C and 7.55 °C. After correction, errors at all four stations converge significantly: the MAE at Bajiaodian and Octagonal Palace decreases to ≤1.7 °C with notably reduced fluctuations; the mid-slope station maintains values below 2 °C for most hours, and the summit station’s peaks are reduced to within 2.57 °C. The average reduction in error exceeds 60% in high-elevation areas, demonstrating the strong capability of the correction method in mitigating structural biases induced by complex terrain.
Overall, under complex topographic conditions, the original ZJOCF temperature forecasts display a typical pattern characterized by higher accuracy at low elevations, substantial warm biases at high elevations, and pronounced temporal fluctuations. The correction method significantly enhances overall forecast accuracy, temporal stability, and error control at high altitudes, thereby providing an effective technical approach for mountain microclimate operational forecasting.

3.2. Structural Characteristics of Temperature Forecast Errors and Evaluation of Extreme-Error Control Capability

To further examine the error-structure characteristics of the ZJOCF model under complex terrain conditions, this study evaluates changes before and after correction from three perspectives—error-interval distribution, worst-case forecast scenarios, and the number of extreme-error samples (Figure 4, Figure 5 and Figure 6). Together with the accuracy and MAE analyses, these diagnostics form an integrated evaluation framework for assessing the effectiveness of the correction method in improving forecast reliability, error concentration, and robustness under complex terrain conditions.
From the perspective of error-interval distribution (Figure 4), the comparison focuses on whether the correction method can reduce systematic bias and increase the concentration of forecast errors within the core interval. This diagnostic is particularly useful for identifying structural shifts in the error distribution that cannot be fully captured by scalar metrics such as MAE alone. Errors at the Bajiaodian and Octagonal Palace stations are mainly concentrated within the −2 °C to 2 °C range, although cold-biased accumulation is still observed; for instance, 15% of Bajiaodian samples fall within the −4 °C to −2 °C range. The high-elevation stations display even stronger warm deviations: approximately 35% and 40% of the mid-slope and summit samples, respectively, fall within the 2 °C to 4 °C interval, with some extending into the 4 °C to 6 °C range. This pattern highlights the systematic warm bias inherent in the raw model over regions with significant vertical terrain gradients. After correction, error distributions across all four stations converge substantially, with more than 95% of samples clustered within the −2 °C to 2 °C core interval. Notably, the summit station achieves complete convergence, indicating a remarkable improvement in structural bias within high-altitude areas.
The analysis of the worst forecast scenarios (Figure 5) is intended to assess the robustness of the correction method under unfavorable conditions, which is particularly important for operational forecasting in complex mountainous terrain. Under the raw model, the lowest-accuracy periods correspond to highly dispersed error distributions. This is particularly evident at the Mountaintop station, where the proportion of large errors exceeding 6 °C reaches 67.5%, indicating severely limited forecast reliability under extreme conditions. After correction, extreme errors shrink substantially: the proportion of >6 °C samples at the summit station decreases to 6%, and at the mid-slope station, the proportion of samples within the 4 °C to 6 °C interval decreases from 36.2% to 6.3%. Meanwhile, the concentration of errors within the −2 °C to 2 °C interval increases markedly, and all four stations exhibit a clear shift toward a unimodal distribution.
The number of extreme-error samples (|error| > 2 °C) was further examined to quantify the correction effect on operationally high-risk forecasts and to determine whether the reduction in large errors was achieved without introducing new systematic deviations (Figure 6). The raw forecasts exhibit a typical warm-bias pattern, with frequent peaks of large errors occurring during the 3–30 h period at both the mid-slope and summit stations. After correction, the number of >2 °C samples decreases by 40–60% across all stations. For example, at the summit station, the count at the 51st hour decreases from 383 to 165, and no notable cold bias is introduced. This indicates that the correction method effectively reduces systematic warm errors without imposing new systematic deviations.
In summary, the correction method substantially improves the error-structure characteristics of the original ZJOCF, transforming the error distribution from a dispersed, multi-peaked pattern to a more concentrated, unimodal form. It also significantly suppresses the frequency of extreme errors, with particularly strong performance in high-elevation, complex-terrain regions. These results demonstrate that the correction strategy can markedly enhance model stability and operational applicability, providing reliable support for meteorological forecasting of temperature in mountainous areas.

3.3. Comparison of Temperature-Error Distribution Shapes Before and After Correction

To more intuitively demonstrate the structural changes in the ZJOCF model’s error distribution before and after correction, this study uses the 24 h forecast lead time as an example and compares the error boxplot structures at the four stations (Figure 7 and Figure 8). This analysis serves as a complementary diagnostic to the preceding accuracy-, MAE-, and frequency-based evaluations by highlighting changes in central tendency, dispersion, asymmetry, and outlier behavior.
The results show that before correction, substantial spatial differences and systematic biases exist across stations at different elevations. At Bajiaodian, the median errors are consistently negative (−1.20 °C to −0.30 °C), indicating a stable cold bias. Some time periods also exhibit pronounced dispersion, with interquartile ranges exceeding 10 °C, suggesting strong variability in the errors. At the Octagonal Palace, the errors are generally close to zero, although occasional large deviations still occur, implying that the raw model retains instability under certain local conditions. In contrast, the high-altitude stations show considerably poorer performance. Both the mid-slope and summit stations exhibit significantly positive median errors (2.60 °C to 3.73 °C), with wide box ranges and upper quartiles reaching 7–11 °C. The number of outliers is markedly higher, revealing a persistent warm bias in complex terrain settings.
After correction, the error-distribution shapes improve substantially, with systematic biases effectively suppressed and error concentration greatly enhanced. The median errors at all four stations converge to the range of −0.4 °C to 0.5 °C, indicating a strong debiasing effect. Box sizes shrink markedly, with upper and lower quartiles constrained within the [−2 °C, 5 °C] interval, and the number of extreme outliers decreases significantly—by more than 50% at the mid-slope and summit stations. Importantly, error dispersion at the high-elevation sites is greatly reduced, and the systematic warm bias is eliminated, demonstrating the ability of the correction method to adapt effectively to terrain-sensitive error characteristics.
Overall, the combined use of accuracy, MAE, error-distribution diagnostics, extreme-error statistics, and boxplot analysis provides a coherent and systematic evaluation of the correction effect. Across these complementary perspectives, the corrected forecasts consistently show improved accuracy, reduced dispersion, fewer extreme errors, and enhanced stability, particularly at the high-elevation stations.

4. Conclusions and Discussion

This study investigated the elevation-dependent error characteristics of ZJOCF temperature forecasts in the Liuchun Lake area and evaluated a station-level machine-learning correction method under complex terrain conditions. The main conclusions are as follows.
(1)
Raw ZJOCF temperature forecasts show clear terrain-dependent deficiencies. Forecast skill decreases with elevation, while high-altitude stations exhibit systematic warm bias, stronger temporal variability, and more frequent extreme errors, indicating substantial limitations of the model in complex mountainous environments.
(2)
The station-level machine-learning correction significantly improves forecast performance. After correction, forecast accuracy increases, MAE decreases, error distributions become more concentrated, and extreme-error samples are greatly reduced, without introducing new systematic bias. The corrected forecasts are also more stable across lead times and show clear operational value for short-term mountainous temperature forecasting.
(3)
Overall, station-level correction provides a practical and efficient approach for refining temperature forecasts in complex terrain. Although this study is based on the Liuchun Lake area, the results are relevant to other mountainous regions affected by strong elevation gradients and terrain-induced microclimates.
The results of this study are consistent with the current understanding of temperature forecasting over complex terrain. The deterioration of raw ZJOCF forecast skill with increasing elevation, together with the warm bias and stronger temporal variability at high-altitude stations, agrees with mountain-meteorology theory and with previous studies showing that temperature forecasts tend to perform poorly under strong cooling conditions and in topographically complex environments [16]. In mountainous areas, cold-air pooling, slope-valley circulations, and stable boundary layers can generate strong sub-grid thermal contrasts that are difficult for gridded operational forecasts to resolve. In this sense, the Liuchun Lake area should be regarded not as an isolated case, but as a representative example of temperature-forecast challenges in complex terrain.
The post-correction results are also consistent with the broader literature on statistical post-processing. The increase in forecast accuracy, reduction in MAE, and suppression of extreme errors indicate that the station-level machine-learning model successfully learned part of the systematic mismatch between gridded model output and local observations, which is a central objective of forecast post-processing [17]. Similar station-based machine-learning correction studies have reported clear improvements in temperature forecasts in both regional applications and recent complex-terrain settings [18]. At the same time, this study is based on only four stations within one mountainous area, so the broader transferability of the correction relationships still requires validation in other regions. Even so, the present results demonstrate that station-level correction is a practical and operationally efficient pathway for refined temperature forecasting in mountainous areas with similar topographic complexity.

Author Contributions

Conceptualization, Y.W. and S.Y.; methodology, Y.W. and Y.S.; software, Y.S. and S.M.; validation, Y.S.; investigation, Y.W.; resources, T.Q.; data curation, Y.S., T.Q., X.L. and S.M.; writing—original draft preparation, Y.S. and Z.Z.; writing—review and editing, X.L.; visualization, Z.Z.; supervision, K.X. and S.Y.; project administration, K.X.; funding acquisition, Y.W. and S.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the 2024 Sichuan Science and Technology Program, grant number 2024YFTX0016; the Longyou County Science and Technology Plan Project, grant number JHXM2022170; and the Scientific and Technological Innovation Capacity Improvement Project of Chengdu University of Information Technology, grant number KYQN202205.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Acknowledgments

The authors would like to thank the anonymous reviewers and the editors for their valuable comments and suggestions, which helped improve the quality of this manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Seitter, K.L. Weather as fuel—The wicked problem of renewable energy. Bull. Am. Meteorol. Soc. 2024, 105, 2033–2036. [Google Scholar] [CrossRef] [Scilit]
  2. Cui, L.; Wang, J.; Tabas, S.S.; Carley, J.R. A Machine Learning-Based Bias Correction Method for Global Forecast System Products; NOAA NCEP Office Note 520; NOAA: Washington, DC, USA, 2025; 23p. [Google Scholar]
  3. Zhang, Q.; Shi, Y.; Wang, Y.; Mou, S.; Zhu, Z.; Qian, T.; Mao, Z.; Yuan, S.; Han, L.; Lao, X. Assessment of the ZJWARMS Forecast Model’s Adaptability and AI-Based Bias Correction over Complex Terrain. Atmosphere 2025, 16, 1151. [Google Scholar] [CrossRef] [Scilit]
  4. Bauer, P.; Thorpe, A.; Brunet, G. The quiet revolution of numerical weather prediction. Nature 2015, 525, 47–55. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Xue, W.; Yu, H.; Tang, S.; Huang, W. Relationships between terrain features and forecasting errors of surface wind speeds in a mesoscale numerical weather prediction model. Adv. Atmos. Sci. 2024, 41, 1161–1170. [Google Scholar] [CrossRef] [Scilit]
  6. Glahn, H.R.; Lowry, D.A. The use of Model Output Statistics (MOS) in objective weather forecasting. J. Appl. Meteorol. 1972, 11, 1203–1211. [Google Scholar] [CrossRef] [Scilit]
  7. Delle Monache, L.; Hacker, J.P.; Zhou, Y.; Deng, X.; Stull, R.B. Kalman filter and analog schemes to postprocess numerical weather prediction ensemble forecasts. Mon. Weather Rev. 2011, 139, 3554–3570. [Google Scholar] [CrossRef] [Scilit]
  8. Raftery, A.E.; Gneiting, T.; Balabdaoui, F.; Polakowski, M. Using Bayesian model averaging to calibrate forecast ensembles. Mon. Weather Rev. 2005, 133, 1155–1174. [Google Scholar] [CrossRef] [Scilit]
  9. Dueben, P.D.; Bauer, P. Challenges and design choices for global weather and climate models based on machine learning. Geosci. Model Dev. 2018, 11, 3999–4009. [Google Scholar] [CrossRef] [Scilit]
  10. Vannitsem, S.; Bremnes, J.B.; Demaeyer, J.; Evans, G.R.; Flowerdew, J.; Hemri, S.; Lerch, S.; Roberts, N.; Theis, S.; Atencia, A.; et al. Statistical postprocessing for weather forecasts: Review, challenges, and avenues in a big data world. Bull. Am. Meteorol. Soc. 2021, 102, E681–E699. [Google Scholar] [CrossRef] [Scilit]
  11. Cho, D.; Yoo, C.; Im, J.; Cha, D.H. Comparative assessment of machine learning-based bias correction methods for NWP model forecasts of extreme air temperatures in urban areas. Earth Space Sci. 2020, 7, e2019EA000740. [Google Scholar] [CrossRef] [Scilit]
  12. Sha, Y.; Gagne, D.J., II; West, G.; Stull, R.B. Deep-learning-based gridded downscaling of surface meteorological variables in complex terrain. Part I: Daily maximum and minimum 2-m temperature. J. Appl. Meteorol. Climatol. 2020, 59, 2057–2073. [Google Scholar] [CrossRef] [Scilit]
  13. Chen, L.; Han, B.; Wang, X.; Zhao, J.; Yang, W.; Yang, Z. Machine learning methods in weather and climate applications: A survey. Appl. Sci. 2023, 13, 12019. [Google Scholar] [CrossRef] [Scilit]
  14. Xia, J.; Li, H.; Kang, Y.; Yu, C.; Ji, L.; Wu, L.; Lou, X.; Zhu, G.; Wang, Z.; Yan, Z.; et al. Machine learning-based weather support for the 2022 Winter Olympics. Adv. Atmos. Sci. 2020, 37, 927–932. [Google Scholar] [CrossRef] [Scilit]
  15. China Meteorological Administration. Methods for Quality Verification of Short- and Medium-Range Weather Forecasting (Trial); China Meteorological Administration: Beijing, China, 2005. [Google Scholar]
  16. Whiteman, C.D. Mountain Meteorology: Fundamentals and Applications; Oxford University Press: Oxford, UK, 2000. [Google Scholar]
  17. Gneiting, T.; Raftery, A.E.; Westveld, A.H., III; Goldman, T. Calibrated probabilistic forecasting using ensemble model output statistics and minimum CRPS estimation. Mon. Weather Rev. 2005, 133, 1098–1118. [Google Scholar] [CrossRef] [Scilit]
  18. Chen, Y.; Ning, Y.; Tang, R.; Xie, X. Tropical temperature correction for numerical forecast in Hainan based on spatiotemporal independence random forest model. Nat. Sci. Hainan Univ. 2025, 38, 356–364. [Google Scholar]
Figure 1. Topographic map of the study area (Liuchun Lake area) and distribution of weather stations.
Figure 1. Topographic map of the study area (Liuchun Lake area) and distribution of weather stations.
Atmosphere 17 00439 g001
Figure 2. Seventy-two-hour temperature-forecast accuracy of the raw model and the corrected predictions (a) Bajiaodian; (b) Octagonal Palace; (c) Mountainside; (d) Mountaintop.
Figure 2. Seventy-two-hour temperature-forecast accuracy of the raw model and the corrected predictions (a) Bajiaodian; (b) Octagonal Palace; (c) Mountainside; (d) Mountaintop.
Atmosphere 17 00439 g002
Figure 3. Seventy-two-hour MAE of the raw and corrected temperature forecasts (a) Bajiaodian; (b) Octagonal Palace; (c) Mountainside; (d) Mountaintop.
Figure 3. Seventy-two-hour MAE of the raw and corrected temperature forecasts (a) Bajiaodian; (b) Octagonal Palace; (c) Mountainside; (d) Mountaintop.
Atmosphere 17 00439 g003
Figure 4. Frequency distribution of intervals corresponding to maximum forecast accuracy (a) ZJOCF model; (b) corrected ZJOCF model.
Figure 4. Frequency distribution of intervals corresponding to maximum forecast accuracy (a) ZJOCF model; (b) corrected ZJOCF model.
Atmosphere 17 00439 g004
Figure 5. Frequency distribution of intervals corresponding to minimum forecast accuracy (a) ZJOCF model; (b) corrected ZJOCF model.
Figure 5. Frequency distribution of intervals corresponding to minimum forecast accuracy (a) ZJOCF model; (b) corrected ZJOCF model.
Atmosphere 17 00439 g005
Figure 6. Number of raw and corrected temperature-error samples exceeding 2 °C or below −2 °C (a) Bajiaodian; (b) Octagonal Palace; (c) Mountainside; (d) Mountaintop.
Figure 6. Number of raw and corrected temperature-error samples exceeding 2 °C or below −2 °C (a) Bajiaodian; (b) Octagonal Palace; (c) Mountainside; (d) Mountaintop.
Atmosphere 17 00439 g006
Figure 7. Boxplots of 24 h temperature-forecast errors of the raw model (a) Bajiaodian; (b) Octagonal Palace; (c) Mountainside; (d) Mountaintop.
Figure 7. Boxplots of 24 h temperature-forecast errors of the raw model (a) Bajiaodian; (b) Octagonal Palace; (c) Mountainside; (d) Mountaintop.
Atmosphere 17 00439 g007
Figure 8. Boxplots of 24 h corrected temperature-forecast errors (a) Bajiaodian; (b) Octagonal Palace; (c) Mountainside; (d) Mountaintop.
Figure 8. Boxplots of 24 h corrected temperature-forecast errors (a) Bajiaodian; (b) Octagonal Palace; (c) Mountainside; (d) Mountaintop.
Atmosphere 17 00439 g008
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wang, Y.; Shi, Y.; Qian, T.; Zhu, Z.; Lao, X.; Xiang, K.; Mou, S.; Yuan, S. Enhancing the Temperature Forecast Accuracy of the ZJOCF Model Using AI-Based Station-Level Bias Correction. Atmosphere 2026, 17, 439. https://doi.org/10.3390/atmos17050439

AMA Style

Wang Y, Shi Y, Qian T, Zhu Z, Lao X, Xiang K, Mou S, Yuan S. Enhancing the Temperature Forecast Accuracy of the ZJOCF Model Using AI-Based Station-Level Bias Correction. Atmosphere. 2026; 17(5):439. https://doi.org/10.3390/atmos17050439

Chicago/Turabian Style

Wang, Yifan, Yiwen Shi, Tu Qian, Zhidan Zhu, Xiaocan Lao, Keyi Xiang, Shiyun Mou, and Shujie Yuan. 2026. "Enhancing the Temperature Forecast Accuracy of the ZJOCF Model Using AI-Based Station-Level Bias Correction" Atmosphere 17, no. 5: 439. https://doi.org/10.3390/atmos17050439

APA Style

Wang, Y., Shi, Y., Qian, T., Zhu, Z., Lao, X., Xiang, K., Mou, S., & Yuan, S. (2026). Enhancing the Temperature Forecast Accuracy of the ZJOCF Model Using AI-Based Station-Level Bias Correction. Atmosphere, 17(5), 439. https://doi.org/10.3390/atmos17050439

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop