3.1. Comparison of Prediction Strategies
Considering that the prediction strategy directly affects how the model utilizes input information and influences the subsequent evaluation of model structures, this study first compares the forecast performance of different prediction strategies. Existing deep learning-based subseasonal prediction methods can generally be divided into two categories: iterative prediction and direct prediction (
Figure 4).
Iterative prediction starts from the current atmospheric state and progressively generates future sequences through rolling recursion. Its advantage is that it more closely resembles the continuous evolution of atmospheric states. However, its limitation is that errors are propagated and gradually accumulated during multi-step recursion, which becomes particularly evident at longer lead times. In contrast, direct prediction does not rely on the prediction result from the previous step, but directly learns the mapping relationship between the current state and the target lead time. This strategy can, to some extent, avoid recursive error accumulation and also has relatively high computational efficiency. Based on this, the different prediction strategies were compared using their mean prediction skill over Weeks 1–6 to determine a temporal modeling strategy more suitable for subseasonal 2-m temperature prediction over East Asia.
Figure 5 presents the mean ACC results of different prediction strategies for subseasonal 2-m temperature prediction over East Asia during the evaluation period from 2019 to 2024. To ensure a fair comparison, all experiments used the same model structure, input variables, and training strategy, with only the prediction strategy being changed. Overall, the prediction skill of all four strategies decreases markedly with increasing lead time, indicating that the predictable signals of 2-m temperature on the subseasonal time scale decay rapidly over time. However, different temporal-evolution modeling strategies show clear differences in their ability to maintain prediction skill.
The iterative and direct prediction strategies exhibit distinct lead-time-dependent characteristics. Iterative prediction progressively generates future states through rolling recursion, which is closer to the continuous evolution process of atmospheric states, and therefore performs relatively well at short lead times. Among them, the temporally non-overlapping iterative prediction strategy performs best during Weeks 1–2, with ACC values of approximately 0.73 and 0.38, respectively, which are clearly higher than those of the other methods. This indicates that this strategy can more effectively utilize neighboring state information and better capture short-term persistence characteristics. However, as the lead time increases, its prediction skill declines rapidly, reaching only about 0.06 by Week 6. This suggests that error accumulation caused by multi-step recursion clearly limits its medium- and extended-range prediction skill. In contrast, the temporally overlapping iterative prediction strategy shows the weakest overall performance, especially during Weeks 5–6, when the ACC decreases to approximately 0.04 and 0.02, respectively. This indicates that high-frequency rolling recursion is more likely to amplify local errors and is unfavorable for the stable extraction of subseasonal low-frequency signals.
Compared with iterative prediction, direct prediction shows better stability in the medium- and extended-range stages. Lead-time prompted prediction has slightly lower skill than the optimal method during Weeks 1–4, but its rate of skill decline is relatively slow, and it still maintains positive skill of approximately 0.09 by Week 6. Lead-time simultaneous prediction achieves the best overall performance, especially during Weeks 3–6, when it achieves higher ACC than the other prediction strategies. This suggests that direct prediction can, to some extent, avoid recursive error propagation and may better support the learning of slowly varying background signals associated with different target weeks. Overall, iterative prediction is more suitable for short-lead prediction tasks, whereas direct prediction has greater advantages for medium- and extended-range forecasts during Weeks 3–6. Since this study focuses on improving subseasonal prediction skill during Weeks 3–6, the subsequent experiments adopt the lead-time simultaneous prediction strategy with the best overall performance as the basic prediction framework; that is, forecasts for Weeks 1–6 are generated simultaneously through a single forward pass.
Based on the selected prediction strategy, the influence of different historical input lengths on model performance is further analyzed to clarify the model’s ability to utilize recent atmospheric states and low-frequency background information.
Figure 6 shows the sensitivity experiment results for different historical input lengths under the lead-time simultaneous prediction framework, where h1–h4 represent historical information from the previous 1–4 weeks, respectively. The results indicate that a longer historical input does not necessarily lead to better performance; instead, the optimal historical input length varies with the target lead time. For Week 1 prediction, h1 performs best, with an ACC of approximately 0.77. As the historical window increases to h4, the skill decreases to approximately 0.69. This suggests that short-lead prediction mainly depends on recent atmospheric state information, while an excessively long history may introduce additional background differences and weaken the model’s response to the current dominant signals. During Weeks 3–4, however, longer historical windows provide clearer benefits. Among them, h4 achieves the highest skill, with ACC values of approximately 0.16 and 0.14 for Weeks 3 and 4, respectively, which are higher than the corresponding h1 values of approximately 0.12 and 0.08. This indicates that medium-range prediction relies more on the cumulative representation of low-frequency and slowly varying processes, and that appropriately extending the historical input length helps enhance the model’s ability to extract subseasonal predictable signals.
However, during Weeks 5–6, the benefit of increasing the historical input length diminishes and may even reverse. In Week 5, h2 performs slightly better, while in Week 6, h1 and h2 are relatively more stable, whereas h3 and h4 show decreased skill. This suggests that at longer lead times, although a longer historical window can provide more background information, it may also introduce irrelevant noise and redundant signals, thereby weakening the model’s generalization performance. Overall, the optimal historical input length varies clearly among different lead times: shorter historical inputs are more suitable for short-range forecasts, whereas longer historical inputs are more beneficial for medium-range prediction during Weeks 3–4. Although the 4-week historical input achieves better ACC during Weeks 3–4, it does not provide a consistent advantage across all lead weeks. In particular, its skill decreases during Weeks 5–6, suggesting that longer historical windows may introduce redundant or less relevant information at longer lead times. Therefore, the one-week historical input was retained as the default setting as a conservative and unified configuration for the following model-comparison experiments. This choice also reduces computational cost and memory consumption, which is important under the current experimental conditions. This setting is not intended to be optimal for every lead week; rather, it provides a consistent baseline for comparing input variables, model components, and CFSv2 under the same input-length condition.
3.2. Feature Ablation Experiments
After the prediction strategy was selected, feature-ablation experiments were conducted to examine the effects of different input-variable combinations on model performance. SwinUNet-AR was used as the common backbone in all experiments. The experiments differed only in their input-variable configurations. Specifically, the experiments began with a univariate setting using only T2m as input to evaluate the basic predictability of the near-surface temperature field itself. Vertical temperature information was then gradually introduced to examine the contribution of thermal features to subseasonal prediction. On this basis, dynamical and circulation-related factors, including wind fields, geopotential height, relative humidity, and surface pressure, were further added to construct multiple variable-combination schemes (
Table 3), and their prediction performance at different lead weeks was compared. Through this design, the incremental contributions of different predictor groups to the model’s medium- and extended-range prediction skill can be systematically evaluated, providing a basis for determining the default input configuration in subsequent experiments.
Table 4 presents the mean ACC results for subseasonal 2-m temperature prediction over East Asia during Weeks 3–6 under different meteorological-variable input configurations. To ensure a fair comparison, all experiments use the same model architecture and training strategy, with only the input-variable combination being changed, so as to quantify the contribution of different predictors to medium- and extended-range forecast skill. Overall, as the input variables are gradually expanded from a single near-surface temperature field to multivariable combinations containing thermal, dynamical, and circulation information, the prediction skill of the model during Weeks 3–6 generally increases. However, the gains associated with adding different variables are not strictly monotonic. The full-variable input scheme (ALL) achieves the best performance, with a cumulative ACC of 0.621 during Weeks 3–6, which is 0.067 higher than that obtained using only T2m. This indicates that multisource meteorological information can effectively enhance the model’s ability to extract predictable signals in the medium- and extended-range stages.
In terms of the role of different variable combinations, the ACC values at all lead weeks are lowest when only T2m is used, indicating that a single near-surface state is insufficient to fully represent the large-scale low-frequency background on which subseasonal temperature-anomaly evolution depends. After introducing vertical temperature information, the model skill improves markedly. The cumulative ACC of the T scheme during Weeks 3–6 increases to 0.597, suggesting that thermal-structure information can provide a more stable background constraint for prediction after Week 3. When wind-field variables are further added, the model performance shows a modest further improvement. For example, the cumulative ACC of the TUV scheme reaches 0.603, although the increase is relatively limited. This indicates that dynamical information contributes mainly through its synergistic representation with thermal structures rather than independently determining the forecast performance.
After further adding geopotential height and relative humidity, the model performance continues to improve. The cumulative ACC values of the TUVZ and TUVZR schemes reach 0.604 and 0.612, respectively, indicating that large-scale circulation patterns and humidity-related processes provide additional constraints for medium- and extended-range prediction. In particular, humidity information contributes more clearly to the improvement of Week 6 skill. Finally, the ALL scheme achieves the best performance, with the largest gains occurring during Weeks 4–6. This suggests that the complete variable combination not only improves the overall prediction skill, but also more effectively suppresses skill degradation at longer lead times. Overall, multivariable input can significantly improve subseasonal 2-m temperature prediction over East Asia. Among the input variables, vertical thermal-structure information provides the largest incremental contribution under the present sequential-addition design, while wind fields, geopotential height, and humidity variables offer further complementary constraints. Based on forecast skill and stability, the ALL scheme is adopted as the default input configuration in subsequent experiments. When input dimensionality and computational efficiency are also considered, TUVZR can be regarded as an alternative scheme with performance close to the optimum.
From a physical perspective, the feature-ablation results provide a preliminary indication of the incremental contributions of different predictor groups under the specified addition order. The improvement from T2m to T suggests that vertical thermal-structure information provides a larger incremental gain than near-surface temperature alone under the present addition order for Weeks 3–6 prediction. The further improvement after adding wind fields and geopotential height indicates that large-scale dynamical circulation and temperature advection provide additional predictive information. The increase after adding relative humidity and surface pressure suggests that moisture-related processes and near-surface pressure conditions also contribute to the maintenance of subseasonal temperature anomalies. Therefore, the model skill is not only derived from local temperature persistence, but also from the combined representation of thermal, dynamical, circulation, and moisture-related background conditions.
3.3. Module Ablation Results
Based on the selected prediction strategy and input-variable combination, module-ablation experiments were then designed to evaluate the effects of the proposed architectural modifications. The baseline SwinUNet was used as the control model. Several comparative configurations were constructed by adding or removing AFNO and Resize-Conv from the baseline SwinUNet. All experiments adopted the same dataset split, input-variable configuration, training epochs, and optimization parameters. Differences in forecast skill among the model configurations for Weeks 1–6 were first compared using the RMSE curves shown in
Figure 7. To further summarize the medium- and extended-range performance quantitatively,
Table 5 presents the mean RMSE values for Weeks 3–4 and Weeks 5–6, together with the cumulative ACC for Weeks 3–6.
Overall, the RMSE values of the five models increase markedly with lead time, indicating that forecast errors in 2-m temperature prediction generally increase as the forecast range extends from Week 1 to Week 6. The RMSE values of all models are relatively low in Week 1, suggesting that short-lead prediction is mainly constrained by initial-condition information and short-term atmospheric persistence, and performance differences among the model architectures remain small. From Week 3 onward, structural differences among the models become more apparent as forecast errors increase. At this stage, differences among network structures in representing low-frequency background signals, improving spatial reconstruction quality, and suppressing error accumulation become more evident.
The week-by-week RMSE curves show that UNet achieves the lowest RMSE in Week 1, with a value of 0.40 °C, indicating that a simple convolutional structure can effectively capture short-term local temperature persistence at short lead times. However, as the lead time increases, the RMSE of UNet increases rapidly, reaching 1.30 °C in Week 6. This indicates its limited ability to maintain slowly varying signals and large-scale anomalous structures in the medium- and extended-range stages. The baseline SwinUNet shows more stable performance than UNet during Weeks 3–6, with the Week 6 RMSE decreasing to 1.24 °C. This suggests that the introduction of a hierarchical window-attention structure enhances the model’s ability to represent multiscale spatial features.
Further comparison of different modules shows that, when only the AFNO module is added, the RMSE values in Weeks 1 and 2 are 0.71 °C and 1.09 °C, respectively, which are higher than those of the baseline SwinUNet. In addition, no stable advantage is observed during Weeks 3–6. This indicates that, although the frequency-domain module extracts large-scale spectral features, the additional processing does not translate into lower overall RMSE under the present configuration without sufficient recovery of spatial detail in the decoder.
The Resize-Conv-only configuration generally yields slightly lower RMSE than the AFNO-only configuration during Weeks 3–6, with the clearest difference occurring during Weeks 5–6. When AFNO and Resize-Conv are introduced simultaneously, the model achieves among the lowest RMSE values in the medium- and extended-range stages. The RMSE values of SwinUNet + AFNO + ResizeConv from Weeks 2 to 6 are 1.02 °C, 1.10 °C, 1.18 °C, 1.22 °C, and 1.21 °C, respectively, which are generally lower than those of the baseline SwinUNet and the single-module schemes. In particular, during Weeks 3–6, the RMSE curve of this model remains at a relatively low level, indicating a certain complementarity between the large-scale low-frequency signal modeling capability of AFNO and the spatial reconstruction capability of Resize-Conv. The combined configuration yields lower RMSE and higher cumulative ACC than either single-module configuration, suggesting that AFNO and Resize-Conv may provide complementary benefits under the present architecture.
Table 5 further presents the mean RMSE values of different module combinations during Weeks 3–4 and Weeks 5–6, as well as the cumulative ACC during Weeks 3–6. The baseline SwinUNet has RMSE values of 1.16 °C and 1.24 °C during Weeks 3–4 and Weeks 5–6, respectively, indicating that error accumulation at longer lead times remains an important factor limiting the model’s medium- and extended-range prediction skill. After adding only AFNO, the RMSE values in the two time windows are 1.17 °C and 1.25 °C, respectively, which are not lower than those of the baseline SwinUNet. This suggests that AFNO alone does not reduce RMSE relative to the baseline under the present configuration. After adding only Resize-Conv, the RMSE values during Weeks 3–4 and Weeks 5–6 are 1.17 °C and 1.23 °C, respectively. The RMSE during Weeks 5–6 is slightly lower than that of the baseline SwinUNet, indicating that this module can partly alleviate spatial reconstruction errors at longer lead times. By contrast, when AFNO and Resize-Conv are added simultaneously, the RMSE values during Weeks 3–4 and Weeks 5–6 decrease to 1.14 °C and 1.21 °C, respectively, both of which are the lowest among all ablation schemes. Meanwhile, the cumulative ACC during Weeks 3–6 increases to 0.62, indicating that the dual-module combination not only reduces the magnitude of prediction errors but also improves the consistency of anomalous spatial patterns.
Overall, AFNO does not consistently reduce RMSE when used alone, whereas the combined AFNO-Resize-Conv configuration yields lower RMSE and higher cumulative ACC than either single-module configuration. Resize-Conv helps improve the spatial reconstruction quality in the decoding stage and reduces error propagation that may be caused by the upsampling process. The resulting SwinUNet-AR, which combines both modules, shows lower RMSE and higher cumulative ACC during the medium- and extended-range period of Weeks 3–6. These results indicate that the proposed module combination improves the stability and accuracy of medium- and extended-range subseasonal 2-m temperature prediction under the present experimental setting.
Figure 8 further shows the spatial distribution of RMSE differences between SwinUNet-AR and the baseline SwinUNet at different lead times. The shading represents the RMSE difference between the two models, namely ΔRMSE = RMSE(SwinUNet-AR) − RMSE(SwinUNet), while the contours indicate the spatial distribution of RMSE for SwinUNet-AR itself. The figure further illustrates the spatial characteristics of the error differences between the improved and baseline models and their evolution with lead time.
Overall, the shaded fields at different lead times are dominated by small positive and negative differences close to zero, but negative values prevail over most regions. This indicates that SwinUNet-AR can reduce prediction errors to some extent in most areas of East Asia. The regional mean ΔRMSE values shown in the subplots are −0.04 °C, −0.03 °C, −0.01 °C, −0.02 °C, −0.00 °C, and −0.01 °C, respectively. These results suggest that, compared with the baseline SwinUNet, SwinUNet-AR achieves slight but relatively stable RMSE improvements over Weeks 1–6, with more evident average improvements in Weeks 1–2 and Week 4.
From the spatial distributions at different lead times, the differences between the two models are generally small during Weeks 1–2. The shaded fields mainly consist of scattered weak negative values and localized weak positive values, indicating that short-lead prediction is still mainly constrained by the initial atmospheric state, and the error reduction brought by structural improvements has not yet fully emerged. Nevertheless, SwinUNet-AR already exhibits a certain range of negative ΔRMSE values over parts of the mid- and low-latitude regions of East Asia and surrounding areas, suggesting that the improved structure has begun to suppress errors even at short lead times.
As the lead time extends to Weeks 3–6, the differences between the two models become clearer. Although a few localized regions still show positive ΔRMSE, negative ΔRMSE values are more widely distributed over key mid- and low-latitude regions of East Asia, South Asia, and adjacent oceanic areas. This indicates that SwinUNet-AR can maintain lower forecast errors over more regions during the medium- and extended-range stages. Meanwhile, the contours show that the RMSE of SwinUNet-AR generally increases with lead time, and high-error regions gradually expand toward low-latitude areas and regions with complex circulation activity. This is consistent with the general pattern of error accumulation with increasing lead time in subseasonal prediction. Although the overall RMSE increases in the medium- and extended-range stages, SwinUNet-AR still generally maintains slightly lower RMSE than the baseline SwinUNet, indicating that the improved structure can partly buffer and suppress error accumulation.
Taken together,
Figure 7 and
Figure 8 show that SwinUNet-AR generally yields slightly lower RMSE than the baseline SwinUNet, although the improvements are modest and spatially heterogeneous. Its advantage is reflected not only in the reduction of mean RMSE during the medium- and extended-range stages, but also in the broader distribution of negative ΔRMSE values across most regional error fields. The AFNO module enhances the direct modeling of large-scale low-frequency background signals at the bottleneck layer, thereby improving the model’s ability to represent cross-regional coherent anomalous structures. Resize-Conv improves the quality of spatial reconstruction in the decoding stage and reduces local artifacts and upsampling errors. These two modules jointly improve model performance from the key aspects of feature extraction and output recovery, providing modest improvements in error control and spatial continuity under the present experimental setting.
3.4. Comparative Evaluation of SwinUNet-AR and CFSv2
The final SwinUNet-AR model was further compared with CFSv2 to evaluate its subseasonal forecast performance over East Asia. CFSv2 is a widely used coupled dynamical prediction system. It can simulate atmospheric circulation evolution based on physical processes and provide subseasonal forecast products. Comparing SwinUNet-AR with CFSv2 helps assess the advantages and limitations of the regional deep learning model relative to the dynamical prediction system in subseasonal 2-m temperature prediction over East Asia. Before the comparison, the CFSv2 forecasts were processed to match the verification framework of SwinUNet-AR as closely as possible. The original 6-hourly CFSv2 forecasts were averaged into daily means, and the 32 ensemble members were averaged to obtain the ensemble-mean forecast. The daily CFSv2 forecasts were then aggregated into the same lead-week windows as the SwinUNet-AR targets. Both systems were evaluated over the same East Asian domain, on the same 1.5° × 1.5° grid, and against ERA5 as the verifying reference. RMSE was calculated using weekly mean 2-m temperature fields, while ACC was calculated using weekly mean anomaly fields after removing the corresponding lead-dependent climatological mean. This consistent preprocessing and verification framework was used to reduce differences caused by temporal aggregation, spatial resolution, and anomaly definition. However, the CFSv2 forecasts were used as raw model outputs without additional bias correction. Therefore, the RMSE comparison should be interpreted as a comparison between SwinUNet-AR and raw CFSv2 forecasts, and part of the RMSE difference may be associated with the systematic climatological bias of CFSv2.
Figure 9 shows the comparison of RMSE and ACC between SwinUNet-AR and CFSv2 for Lead Weeks 1–6.
In terms of RMSE, both systems show increasing errors with lead time. Under the present verification framework, SwinUNet-AR shows lower RMSE than raw CFSv2 at all lead weeks. Specifically, the RMSE of CFSv2 increases from 1.437 °C in Week 1 to 2.04 °C in Week 6, with a mean RMSE of approximately 1.87 °C over Weeks 1–6. In comparison, the RMSE of SwinUNet-AR increases from 0.45 °C in Week 1 to 1.21 °C in Week 6, with a mean RMSE of approximately 1.03 °C. However, this RMSE difference should be interpreted cautiously because CFSv2 was evaluated using raw ensemble-mean forecasts without additional bias correction, whereas SwinUNet-AR was trained directly against ERA5. Therefore, the lower RMSE of SwinUNet-AR reflects both lower forecast errors and the model’s ability to fit the ERA5-based temperature distribution and should not be regarded as a fully bias-corrected operational comparison.
In terms of ACC, the comparison is based on climatological anomalies and is therefore less affected by mean-state bias. CFSv2 shows higher ACC during Weeks 1–3, indicating that the dynamical model retains advantages in representing physically consistent anomaly evolution at short to medium lead times. During Weeks 4–6, SwinUNet-AR becomes more competitive and shows slightly higher ACC than CFSv2 at some lead times. However, the ACC difference is relatively small and should be interpreted cautiously. Overall, the CFSv2 comparison should be viewed as a reference benchmark rather than as direct evidence that the data-driven model fully outperforms the dynamical system. These results suggest that SwinUNet-AR may provide a useful regional complement to dynamical forecasts, particularly in achieving lower RMSE than raw CFSv2 forecasts, while CFSv2 still retains advantages in short-lead anomaly-pattern prediction.
To further analyze the differences in the spatial error distributions between SwinUNet-AR and the dynamical model CFSv2,
Figure 10 shows the spatial distribution of RMSE differences between the two models at different lead times. The regional mean ΔRMSE values for Weeks 1–6 are all negative, at approximately −0.99, −0.72, −0.85, −0.84, −0.82, and −0.83 °C, respectively. Because ΔRMSE is defined as RMSE(SwinUNet-AR) − RMSE(CFSv2), these negative values indicate that SwinUNet-AR has lower RMSE than raw CFSv2 at all lead times under the present verification framework. The largest difference occurs in Week 1, when the RMSE of SwinUNet-AR is approximately 0.99 °C lower than that of CFSv2. During Weeks 2–6, the regional mean differences remain between approximately −0.72 and −0.85 °C. However, because CFSv2 was evaluated without additional bias correction, these differences reflect both forecast errors and mean-state biases and should not be interpreted as evidence that SwinUNet-AR comprehensively outperforms the dynamical forecasting system.
From a spatial perspective, most regions in
Figure 10 are dominated by blue shading, indicating that the RMSE of SwinUNet-AR is lower than that of CFSv2 at most grid points over East Asia. During Weeks 1–2, the negative-value regions are widely distributed, suggesting that SwinUNet-AR has lower RMSE than raw CFSv2 over most grid points in the verification domain at short lead times. By Weeks 3–6, the blue-shaded regions still persist over a large area, indicating that even in the medium- and extended-range subseasonal stage, SwinUNet-AR continues to exhibit lower RMSE than raw CFSv2 across most regions. The localized red regions indicate that, in some areas and at certain lead times, the RMSE of CFSv2 may be lower than that of SwinUNet-AR. These localized positive differences indicate regional and lead-dependent variations in relative model performance; their physical causes require further investigation.
The RMSE contours for SwinUNet-AR show that its error level generally increases with lead time. The contour values are relatively low in Week 1, indicating that SwinUNet-AR shows relatively low RMSE against ERA5 at short lead times. As the lead time extends to Weeks 3–6, the area covered by higher RMSE contours gradually expands, reflecting the typical increase in forecast error with lead time in subseasonal prediction. Nevertheless, the shaded differences are still dominated by negative values, indicating that although the RMSE of SwinUNet-AR increases with lead time, it still maintains a lower error level than CFSv2.
Taken together,
Figure 9 and
Figure 10 show that SwinUNet-AR has consistently lower RMSE than raw CFSv2 under the present verification framework. However, part of this RMSE difference may reflect the mean-state bias of the uncorrected CFSv2 forecasts. The results therefore suggest that SwinUNet-AR may serve as a useful regional complement to dynamical subseasonal forecasts.