1. Introduction
Ultraviolet (UV) radiation is the leading preventable risk factor for skin cancer, both melanoma and non-melanoma type. According to the latest update from the World Health Organization [
1], most of these cancers are directly related to cumulative exposure to UV radiation, whether solar or artificial. In 2020 alone, this exposure was responsible for approximately 1.2 million new cases of non-melanoma skin cancer and 325,000 cases of melanoma, with more than 120,000 associated premature deaths. In light-skinned populations, between 70% and 95% of cutaneous melanomas and up to 95% of non-melanoma skin cancers are attributable to UV radiation [
1,
2].
In this context, UV radiation intensity, defined as the incident UV radiant energy per unit area, represents a key variable for understanding and modeling the level of sun exposure in areas with high solar exposure, such as various regions of Peru. Although the Ultraviolet Radiation Index (UVI) is widely used for population warning purposes, quantitative analysis of UV intensity allows a more detailed characterization of the physical phenomenon and is essential for predicting extreme risk scenarios. Skin cancer represents an increasing public health concern due to cumulative exposure to ultraviolet radiation. In Peru, elevated solar UV radiation levels represent a persistent environmental and public health challenge, underlining the need for robust predictive UV intensity models as input for preventive health policies.
UV radiation is measured using specialized equipment that allows for assessing its impact on human health. Therefore, there are two main indicators for monitoring it:
UV Radiation Intensity: Refers to the amount of UV energy reaching a surface, expressed in milliwatts per square centimeter (). This parameter quantifies the actual energy of UV radiation at a specific time and place.
UV Index (UVI): This is a standardized scale that measures the risk of exposure to human health. Its value ranges from 0 (no risk) to over 11 (extreme risk), serving as a warning tool for the population.
For these reasons, the present study focuses on UV radiation intensity as the target variable for short-term forecasting.
Various studies have shown that climate change can intensify ultraviolet (UV) radiation intensity levels as a result of decreased cloud cover, variability in atmospheric aerosol concentrations, and the residual effects of ozone layer depletion. These factors modify the radiative balance of the atmosphere, increasing the transmission of direct UV radiation to the Earth’s surface, especially in vulnerable regions such as the Peruvian coast. In this scenario, accurate prediction of UV intensity becomes essential for assessing the actual risk of exposure and designing adaptation strategies to the effects of climate change.
These challenges have motivated the development of forecasting methodologies capable of anticipating short-term variations in UV radiation under changing atmospheric conditions. Traditional statistical forecasting models have been widely used in environmental time-series analysis due to their interpretability and computational efficiency. However, their ability to represent highly nonlinear patterns and abrupt temporal variations may be limited under complex atmospheric conditions. In recent years, machine learning and deep learning approaches have demonstrated promising capabilities for modeling nonlinear relationships and temporal dependencies in environmental variables, providing alternative tools for improving forecasting accuracy.
Despite these advances, comparative studies evaluating statistical and machine learning approaches under consistent forecasting conditions remain limited. Furthermore, aspects such as forecasting horizons, temporal validation strategies, and model assessment procedures are not always addressed within a unified evaluation framework. These limitations motivate the development of forecasting methodologies that allow a more rigorous assessment of UV radiation intensity prediction models.
Following the growing concern about UV radiation exposure and its health effects, our study is based on data obtained from the Meteorological Station located at the National University of Engineering (UNI) in the Rímac district of Lima, Peru. Traditional statistical models (ARIMA, SARIMA, and Prophet), recurrent deep learning architectures, and hybrid neural models were applied to develop UV radiation intensity forecasting models. These models were used to identify temporal patterns and trends that can support prevention strategies for skin cancer and other diseases related to sun exposure.
In addition, this study adopts a unified forecasting framework based on consistent temporal aggregation, daytime filtering, and a fixed short-term forecasting horizon. Statistical and machine learning models are evaluated under the same chronological conditions, while complementary validation and diagnostic procedures are considered to provide a more reliable assessment of forecasting performance. The results contribute to understanding the suitability of different forecasting approaches for short-term UV radiation prediction in coastal urban environments and may support future environmental monitoring and public health initiatives.
This paper is structured as follows:
Section 2 presents a literature review;
Section 3 describes the materials and methods used;
Section 4 details the results obtained;
Section 5 discusses the results; and finally
Section 6 presents the conclusions of the study.
2. Literature Review
Ultraviolet (UV) radiation is recognized as an important environmental and public health risk factor due to its effects on skin and ocular health [
1,
2]. The literature includes several related works on solar radiation prediction, UV radiation intensity forecasting, and ultraviolet index (UVI) estimation. Although the UVI has been widely used as a categorical indicator of public health risk, it does not directly represent the continuous physical magnitude of UV radiation. In contrast, UV radiation intensity, usually expressed in
, provides a continuous measurement that allows for a more detailed characterization of temporal variability and supports the development of robust predictive models for environmental monitoring and public health applications.
Among classical forecasting approaches, ARIMA, SARIMA, and Prophet models have been widely used in solar and environmental time-series prediction. Since ref. [
3] highlighted the lack of UV radiation intensity data as a modeling challenge, several studies have contributed to the development of statistical and machine learning-based predictive approaches. In [
4], the ARIMA model was investigated for spatial solar radiation forecasting. Meanwhile, ref. [
5] used an optimized Prophet algorithm for long-term solar irradiance prediction. Similarly, ref. [
6] analyzed ARIMA models for solar radiation forecasting across different geographic locations. In [
7], ARIMA, SARIMA, and Prophet were compared as time-series forecasting models, showing their usefulness for trends and seasonal patterns under different forecasting conditions. In addition, ref. [
8] proposed a statistical–mathematical model for solar radiation prediction, reinforcing the relevance of classical modeling approaches in radiation-related forecasting problems.
In [
9], the SARIMA model was applied to weather forecasting, demonstrating its usefulness for climatic variables due to its ability to capture repetitive seasonal behavior. Although these classical models remain valuable due to their interpretability and computational efficiency, their performance may be limited when dealing with highly nonlinear UV radiation dynamics, abrupt peak events, and rapidly changing atmospheric conditions.
Machine learning and deep learning methods have increasingly been applied to UV radiation-related prediction, estimation, and monitoring tasks. In [
10], a mathematical model was developed to quantify the effect of urban trees on reducing human exposure to UV radiation in Seoul, Korea, contributing to sustainable urban design. Later, ref. [
11] proposed an Artificial Neural Network (ANN) model to predict global UVI based on astronomical parameters. In [
12], Deep Belief Networks and Backpropagation were used to predict safe sun exposure time for individuals with phototypes I and II. Furthermore, ref. [
13] proposed a direct UV radiation measurement system using smartphone cameras and fog computing, improving the accessibility of UV monitoring technologies.
Deep learning architectures have also shown strong potential for UV-related forecasting problems. In [
14], a deep learning model was designed to predict extreme ultraviolet radiation (EUV) emissions from He I 1083 nm solar images, showing the applicability of neural models to UV-related radiation prediction. In [
15], hybrid deep learning methods and optimization algorithms were used to develop a UVI forecasting model. Along the same line, ref. [
16] developed a cloud-affected solar UV prediction system based on a three-phase wavelet hybrid convolutional long short-term memory network for multi-step forecasting. These studies support the use of hybrid architectures that combine convolutional feature extraction and recurrent sequence modeling for UV radiation forecasting.
In [
17], machine learning models, particularly deep neural networks, were shown to be effective for mapping clear-sky surface solar ultraviolet radiation with high spatial resolution. Independently, ref. [
18] introduced an explainable artificial intelligence model for one-hour-ahead solar UVI prediction, incorporating LIME, SHAP, and permutation feature importance techniques to improve interpretability in UV-related forecasting tasks. In [
19], UVBoost was proposed as a Gradient Boosting-based model for erythemally weighted UV radiation estimation, achieving competitive predictive performance and highlighting the relevance of continuous UV radiation variables for environmental forecasting.
Remote sensing, physical modeling, and statistical approaches have also contributed to UV radiation analysis. In [
20], satellite-derived UVI values were evaluated against ground-based measurements in Egypt, identifying differences in accuracy across climatic conditions. In [
21], random matrix theory was used to study and predict solar ultraviolet radiation behavior, contributing to the physical and statistical understanding of UV radiation variability. Later, ref. [
22] developed a machine learning approach to estimate total UV, UVA, and UVB irradiance from solar radiation data and atmospheric parameters.
Similarly, ref. [
23] developed ANN and regression-based models to estimate horizontal ultraviolet irradiance under different sky conditions. The study considered variables such as global horizontal irradiance, clearness index, solar altitude angle, temperature, cloud cover, and diffuse irradiance, concluding that ANN-based models achieved higher accuracy than regression approaches. Subsequently, ref. [
24] implemented machine learning models to generate a daily ultraviolet radiation prediction dataset in China at 10 km spatial resolution, improving the applicability of UV prediction for public health and environmental exposure studies.
Recent hybrid and advanced machine learning approaches further support the relevance of nonlinear models for radiation-related forecasting. In [
25], MTM-LSTM and MLP models were combined for solar radiation time-series forecasting, illustrating the usefulness of hybrid neural structures for radiation-related environmental prediction. Finally, ref. [
26] integrated statistical evaluation and recurrent neural network-based hybrid forecasting for surface solar irradiation, demonstrating the potential of hybrid statistical–deep learning strategies in radiation forecasting problems.
Beyond application-oriented studies, several authors have highlighted the growing importance of deep learning architectures for time-series forecasting problems. A comprehensive review presented by [
27] demonstrated the effectiveness of recurrent and hybrid neural architectures in modeling complex temporal dependencies. The theoretical foundations of modern deep learning systems are extensively discussed in [
28,
29], while the recurrent architectures employed in this study are supported by the seminal Long Short-Term Memory (LSTM) model proposed by [
30] and the Gated Recurrent Unit (GRU) architecture introduced by [
31]. Furthermore, hybrid deep learning strategies combining convolutional and recurrent components have shown strong performance in solar irradiance forecasting applications [
32,
33].
Recent advances have also emphasized the importance of rigorous model evaluation and reproducibility in deep learning research. The selection of loss functions and performance metrics has been extensively reviewed in [
34]. Moreover, recent studies specifically focused on ultraviolet radiation forecasting have demonstrated the advantages of explainable hybrid deep learning frameworks for UV-B prediction [
35], while comparative analyses between satellite-derived and ground-based ultraviolet irradiance measurements continue improving environmental monitoring capabilities [
36]. These developments reinforce the need for robust experimental protocols incorporating temporal validation, statistical significance testing, and residual diagnostics when evaluating forecasting models.
Overall, the reviewed literature shows that UV radiation prediction has evolved from classical statistical models toward machine learning, deep learning, and hybrid neural architectures. However, several methodological gaps remain, including the need for standardized forecasting horizons, chronological train–test partitioning, leakage-free normalization, temporal validation, statistical significance testing, and residual diagnostics. These gaps motivate the comparative framework proposed in this study for short-term UV radiation intensity forecasting in Lima, Peru.
3. Materials and Methods
Figure 1 summarizes the methodological workflow followed in this study, including data preprocessing, model implementation, validation, and final model selection.
3.1. Study Area and Dataset
The study was conducted using ultraviolet (UV) radiation measurements collected in Lima, Peru. The meteorological monitoring station was located at the Universidad Nacional de Ingenieria (UNI), in the Rímac district of Lima, Peru, at approximately S and W. Measurements were collected using an MKIII-RTI-LR meteorological station (RAINWISE Inc., Trenton, ME, USA). The original dataset contained 827,559 observations, each including date, time, and UV radiation intensity measurements.
Table 1 presents a representative sample of the original UV radiation intensity records.
Because UV radiation values are negligible or equal to zero during nighttime periods, observations were restricted to the interval between 06:00 and 18:00 local time. Including nighttime observations would introduce long sequences of near-zero values that are not representative of actual UV exposure conditions and could artificially inflate forecasting performance. Therefore, the analysis focused exclusively on daylight periods, which are both physically meaningful and operationally relevant for UV monitoring and early-warning applications.
To reduce high-frequency irregularity and improve temporal consistency, the original observations were aggregated into 5 min intervals. After removing invalid timestamps, duplicated observations, and nighttime records, the final modeled dataset contained 89,571 valid observations covering the period from April 2023 to November 2024.
Figure 2 shows the UV radiation time series after preprocessing, daytime filtering, and temporal aggregation. The series exhibits strong daily variability, seasonal fluctuations, and abrupt radiation peaks, highlighting the nonlinear and highly dynamic nature of the forecasting problem.
3.2. Data Preprocessing
Several preprocessing steps were performed before model development. First, invalid timestamps and duplicated observations were removed. Then, the UV radiation series was sorted chronologically and aggregated using a fixed 5 min temporal resolution.
To ensure numerical stability and reproducibility during neural network training, UV radiation values were normalized using Min–Max scaling. To avoid temporal information leakage, the scaler was fitted exclusively on the chronological training subset and subsequently applied to the validation and testing subsets. This procedure ensured that information from the future testing period was not used during model training or preprocessing.
The forecasting problem was formulated as a supervised learning task using a sliding-window strategy, where previous observations were used to predict UV radiation intensity at a fixed future horizon.
3.3. Forecasting Horizon
A fixed-horizon forecasting strategy was adopted to ensure that all models were compared under the same operational forecasting condition.
The input window size was defined as , and the forecasting horizon was defined as . Since the temporal aggregation interval was 5 min, each model used the previous 60 min of UV radiation observations to forecast the value 60 min ahead.
This configuration was selected because one-hour-ahead forecasting is operationally relevant for short-term UV warning systems, outdoor exposure prevention, and public health decision-making.
3.4. Chronological Train–Test Partition
Environmental time-series forecasting requires preserving temporal order to avoid information leakage from future observations. Therefore, the dataset was divided chronologically into training and testing subsets without random shuffling.
The adopted partition was
The training subset contained 80,613 observations, whereas the testing subset contained 8958 observations. This chronological holdout configuration was used as the primary out-of-sample evaluation protocol.
Table 2 summarizes the experimental configuration adopted throughout the study.
3.5. Forecasting Models
Eight forecasting models were evaluated under the same preprocessing pipeline, forecasting horizon, and chronological testing protocol.
3.5.1. Statistical Models
The statistical forecasting baseline included the following models:
These models represent widely used approaches in environmental and meteorological time-series forecasting and were included as classical benchmarks.
3.5.2. Deep Learning Models
Two recurrent neural network architectures were evaluated:
Both architectures were selected because they are designed to capture nonlinear temporal dependencies and sequential patterns in time-series data.
3.5.3. Hybrid Neural Architectures
Two hybrid neural architectures were evaluated to combine local temporal feature extraction with recurrent sequence modeling:
Hybrid Model 1 was designed to extract short-term local patterns using a causal one-dimensional convolutional layer, followed by batch normalization, spatial dropout, and a GRU layer. This structure allows the model to capture abrupt local variations in UV radiation while preserving temporal dependence.
Hybrid Model 2 extends this idea through a multiscale convolutional structure with two parallel causal Conv1D branches using kernel sizes 3 and 7. These branches extract temporal patterns at different local scales. Their outputs are concatenated, normalized, regularized through spatial dropout, and then processed by a bidirectional GRU layer. A global average pooling layer summarizes the sequential representation before the final dense prediction layers. This architecture was selected to improve temporal representation while maintaining computational efficiency.
Architectural details of the recurrent and hybrid neural models are summarized in
Table 3.
3.6. Training Configuration
All neural models were trained using the Adam optimizer with a learning rate of 0.0003. The recurrent models (LSTM and GRU) employed the Mean Squared Error (MSE) loss function:
The hybrid architectures were trained using the Huber loss function to improve robustness against abrupt UV radiation peaks and extreme deviations:
where
denotes the prediction error.
To reduce overfitting, dropout and L2 regularization were incorporated into the recurrent and hybrid architectures. The L2 penalty is defined as
where
is the regularization coefficient and
represents the trainable model parameters.
Early stopping and adaptive learning-rate reduction were additionally employed to improve convergence and prevent unnecessary training epochs. To ensure reproducibility, a fixed random seed of 2026 was adopted throughout all experiments.
3.7. Temporal Cross-Validation
In addition to the primary chronological holdout evaluation, a five-fold temporal cross-validation procedure was performed to assess model stability across different chronological partitions.
Unlike conventional random cross-validation, temporal cross-validation preserves the temporal order of observations. This is essential in forecasting problems because random shuffling may introduce future information into the training process and produce overly optimistic results.
Let the ordered time series be denoted by
For each fold
k, the training subset contains past observations and the validation subset contains future observations:
This expanding-window strategy ensures that each validation fold simulates a realistic forecasting scenario, where only past information is available for model training.
The temporal cross-validation analysis was applied to the following deep learning and hybrid neural models:
LSTM.
GRU.
Hybrid Model 1.
Hybrid Model 2.
The statistical models were evaluated on the chronological holdout partition. This strategy allowed us to compare all models under the same final test scenario while using temporal cross-validation to assess the stability of the neural forecasting architectures.
3.8. Evaluation Metrics
Model performance was assessed using three standard forecasting metrics: Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), and the coefficient of determination (
).
and
Lower MAE and RMSE values indicate better predictive accuracy, whereas larger values indicate greater explanatory capability.
3.9. Statistical Significance Testing
To determine whether differences between competing forecasting models were statistically meaningful, pairwise comparisons were performed against the best-performing model. Let
and
denote the forecasting errors of models
A and
B, respectively.
Three complementary statistical procedures were applied: the Wilcoxon signed-rank test, the paired t-test on absolute forecasting errors, and the Diebold–Mariano test.
For the Wilcoxon signed-rank test, the paired differences in absolute errors were defined as
After removing zero differences, ranks were assigned to
, and the Wilcoxon statistic was computed as
where
This nonparametric test evaluates whether the median difference in forecasting errors is significantly different from zero.
The paired
t-test was also applied to the paired absolute error differences:
The test statistic is given by
where
is the mean of the paired differences,
is their standard deviation, and
n is the number of paired observations.
The Diebold–Mariano test was considered the primary forecasting comparison criterion because it directly evaluates whether two competing models have significantly different predictive losses. The loss differential was defined as
where
is a forecasting loss function. In this study, squared error loss was used:
The Diebold–Mariano statistic is computed as
where
is the average loss differential and
is the estimated long-run variance of
, adjusted for autocorrelation up to the forecasting horizon. A statistically significant result indicates that the compared models differ in predictive accuracy.
3.10. Residual Diagnostic Methodology
Residual diagnostics were performed to evaluate model adequacy and identify remaining systematic forecasting errors. For each model, the residuals were defined as
The residual mean was computed as
and the residual standard deviation was computed as
Residual skewness was calculated as
whereas residual kurtosis was computed as
To evaluate remaining autocorrelation in the residuals, the Ljung–Box test was applied. Its statistic is defined as
where
n is the number of residuals,
h is the number of lags, and
is the sample autocorrelation of the residuals at lag
k. A small
p-value indicates remaining autocorrelation, whereas a larger
p-value suggests that residuals behave more like white noise.
Residual plots were additionally inspected to evaluate systematic forecasting bias, remaining temporal dependence, and the behavior of prediction errors during extreme UV radiation peaks.
4. Results
4.1. Comparative Forecasting Performance
All forecasting approaches were evaluated using the same forecasting horizon and temporal partition to ensure methodological consistency.
Table 4 summarizes the predictive performance obtained by statistical, deep learning, and hybrid forecasting models on the independent test set.
Hybrid Model 2 achieved the best overall predictive performance in terms of RMSE and explanatory capability, obtaining the lowest RMSE (0.3849) and the highest coefficient of determination (). Hybrid Model 1 ranked second, followed by the GRU and LSTM architectures. Although the Naive model achieved the lowest MAE (0.1509), its RMSE (0.4302) and explanatory power () were substantially inferior to those of Hybrid Model 2, indicating reduced capability to accurately predict large UV radiation fluctuations. Therefore, considering all evaluation metrics jointly, Hybrid Model 2 provided the most balanced and reliable forecasting performance. These results indicate that hybrid deep learning approaches are better suited for modeling the nonlinear dynamics of UV radiation intensity.
Although the Naive model achieved the lowest MAE, model selection was primarily based on RMSE and statistical significance analyses because RMSE penalizes large forecasting errors more heavily and is particularly relevant for extreme UV radiation events.
Traditional statistical models exhibited substantially lower predictive capability. ARIMA and SARIMA produced negative values, indicating predictive performance below that of a simple baseline predictor. Prophet also showed limited forecasting capability, yielding a negative coefficient of determination and failing to capture the variability present in the UV radiation series.
Overall, the superior performance of the deep learning and hybrid architectures suggests that nonlinear temporal dependencies play a dominant role in UV radiation dynamics and cannot be adequately represented using conventional statistical approaches.
Figure 3 compares the forecasts generated by all evaluated approaches.
Deep learning and hybrid architectures more closely follow the temporal evolution of the observed UV radiation series, whereas statistical models tend to oversmooth the signal and fail to reproduce high-intensity radiation peaks.
4.2. Detailed Analysis of the Best Model
Given its superior predictive performance, Hybrid Model 2 was selected for a more detailed analysis.
Figure 4 presents a detailed comparison between observed values and predictions generated by Hybrid Model 2 during the final testing period.
The model successfully reproduces the overall temporal structure of the UV radiation series while maintaining close agreement with the observed measurements. Most radiation peaks are correctly identified, and the timing of major fluctuations is accurately represented.
Although the highest peaks remain partially underestimated, Hybrid Model 2 captures the underlying temporal dynamics more effectively than all competing methods. This behavior demonstrates the ability of the proposed architecture to balance forecasting accuracy and stability in the presence of highly variable environmental conditions.
4.3. Temporal Cross-Validation Analysis
To evaluate model robustness and stability, a five-fold temporal cross-validation procedure was conducted. The analysis was applied to the deep learning and hybrid architectures, which represent the primary methodological contribution of this study.
Table 5 summarizes the average performance obtained across all folds.
Figure 5 presents the distribution of RMSE values obtained across the five temporal folds.
The temporal cross-validation results confirm that the evaluated deep learning and hybrid architectures exhibit stable predictive behavior across chronological partitions. While GRU achieved the best average RMSE during cross-validation, Hybrid Model 2 obtained the best performance on the independent holdout test set, indicating superior generalization capability under real forecasting conditions.
4.4. Statistical Significance Analysis
To determine whether the observed forecasting improvements were statistically meaningful, Hybrid Model 2 was selected as the reference model and compared against all competing approaches using Wilcoxon signed-rank tests, paired t-tests, and Diebold–Mariano (DM) tests.
Table 6 summarizes the statistical comparison between Hybrid Model 2 and the remaining forecasting approaches. The RMSE improvement column quantifies the relative reduction in forecasting error achieved by Hybrid Model 2 with respect to each competing model. Wilcoxon signed-rank tests and paired
t-tests were included as complementary pairwise analyses of forecasting errors, whereas the Diebold–Mariano test was adopted as the primary criterion for comparing predictive accuracy between forecasting models.
The Wilcoxon signed-rank and paired t-test results indicated statistically significant differences between Hybrid Model 2 and all competing models. However, the Diebold–Mariano test, which directly evaluates predictive accuracy differences, only confirmed significant improvements over the traditional statistical models. These findings support the conclusion that the proposed hybrid architecture achieved statistically significant forecasting improvements over ARIMA, SARIMA, and Prophet.
The Diebold–Mariano test further confirmed that Hybrid Model 2 significantly outperformed ARIMA, SARIMA, and Prophet (). In contrast, no statistically significant differences were observed between Hybrid Model 2 and the remaining neural architectures (LSTM, GRU, and Hybrid Model 1), despite Hybrid Model 2 achieving the lowest RMSE on the independent holdout test set. These results suggest that the deep learning and hybrid approaches belong to the same high-performance forecasting group, while collectively outperforming the classical statistical models.
The largest forecasting improvement was observed against SARIMA, where Hybrid Model 2 reduced RMSE by approximately 67.67%, followed by ARIMA (41.45%) and Prophet (20.85%). Overall, the statistical analyses provide robust evidence that modern deep learning and hybrid architectures are more suitable than conventional statistical approaches for short-term ultraviolet radiation intensity forecasting.
4.5. Comparative RMSE Analysis
Figure 6 provides a visual comparison of the RMSE values obtained by all evaluated forecasting models.
The figure clearly illustrates the superiority of deep learning and hybrid approaches over traditional statistical methods. Hybrid Model 2 achieved the lowest RMSE and the highest coefficient of determination among all evaluated forecasting models, followed closely by Hybrid Model 1, GRU, and LSTM. Statistical models, particularly SARIMA and ARIMA, exhibited substantially higher RMSE values and poorer overall predictive performance. These results confirm that nonlinear deep learning architectures are more effective than conventional statistical methods for short-term UV radiation forecasting.
4.6. Residual Analysis
Residual diagnostics were conducted for the best-performing forecasting model to assess systematic prediction errors and identify potential sources of model bias.
Figure 7 presents the residual series obtained from Hybrid Model 2.
Residuals fluctuate around zero during most of the testing period, indicating the absence of major systematic bias. This behavior is also supported by the residual mean reported in
Table 7. However, the positive skewness (4.6013) and high kurtosis (38.3022) indicate the presence of occasional large forecasting errors, mainly associated with abrupt and extreme UV radiation peaks. These results indicate that extreme UV radiation events remain the principal source of forecasting uncertainty and prediction error variability.
The Ljung–Box test yielded a p-value below 0.05, indicating that some temporal dependence remains in the residual series. This result suggests that although Hybrid Model 2 captures the main temporal dynamics of UV radiation intensity, part of the residual structure may still be associated with unobserved atmospheric drivers, such as cloud cover, ozone concentration, humidity, aerosols, and other meteorological covariates not included in the present study.
Overall, the residual diagnostics support the practical robustness of Hybrid Model 2 for short-term UV radiation forecasting. The combination of the lowest RMSE, the highest coefficient of determination (), favorable residual behavior, and statistically significant improvements over conventional forecasting methods positions Hybrid Model 2 as the most effective approach evaluated in this study. Nevertheless, forecasting uncertainty remains concentrated around extreme UV radiation events, highlighting the potential value of incorporating additional atmospheric predictors in future research.
5. Discussion of Results
The results confirm that ultraviolet (UV) radiation intensity forecasting is a highly nonlinear and temporally complex problem. The analyzed series exhibited strong daily variability, abrupt peaks, and nonstationary behavior, characteristics that are consistent with the observed residual skewness, kurtosis, and remaining temporal dependence identified in the residual diagnostics. These factors partially explain the limited performance of classical statistical models. Under the unified evaluation protocol adopted in this study, all models were assessed using the same 5 min temporal resolution, a chronological 90%/10% train–test partition, and a fixed forecasting horizon of 60 min.
Hybrid Model 2 achieved the best predictive performance on the independent holdout set, obtaining MAE = 0.1618, RMSE = 0.3849, and . This result suggests that the hybrid architecture was more effective in capturing nonlinear temporal dependencies than both traditional statistical approaches and standard recurrent neural networks. Hybrid Model 1 ranked second, followed by GRU and LSTM, confirming the advantage of hybrid architectures for short-term UV radiation forecasting.
The superior performance of Hybrid Model 2 can be attributed to the complementary integration of convolutional and recurrent components. The convolutional layers capture short-term local patterns and abrupt fluctuations, whereas the recurrent layers model temporal dependencies across the input sequence. This combination is particularly suitable for environmental time series characterized by rapid variability and extreme events.
The observed behavior is also consistent with the atmospheric characteristics of Lima, a coastal city influenced by marine humidity, seasonal cloud variability, and rapidly changing radiative conditions. These factors may contribute to the nonlinear temporal dynamics observed in the UV radiation series, favoring forecasting approaches capable of capturing complex short-term fluctuations.
In contrast, ARIMA and SARIMA produced negative () values, indicating that linear autoregressive structures were insufficient to represent the complex dynamics of the UV radiation series. Prophet also showed limited predictive capability, likely because its underlying assumptions are better suited to smoother seasonal patterns than to highly volatile environmental signals.
The statistical significance analysis reinforced these findings. Diebold–Mariano tests indicated that Hybrid Model 2 significantly outperformed ARIMA, SARIMA, and Prophet. However, differences among the deep learning models were not statistically significant, suggesting that neural and hybrid architectures constitute a higher-performing group than classical statistical methods under the evaluated forecasting scenario.
Temporal cross-validation provided additional evidence regarding model robustness. Although GRU achieved the lowest average RMSE across temporal folds, Hybrid Model 2 consistently delivered the best performance on the final holdout test set, indicating stronger out-of-sample generalization capability.
Residual diagnostics further supported the adequacy of Hybrid Model 2. The residual mean (0.1034) and standard deviation (0.3708) indicated that prediction errors remained relatively centered around zero with moderate dispersion. However, the positive skewness (4.6013) and high kurtosis (38.3022) revealed the presence of occasional large errors associated with abrupt UV radiation peaks. Additionally, the Ljung–Box test indicated residual temporal dependence, suggesting that further improvements could be achieved through the incorporation of additional meteorological predictors.
Comparison with previous studies should be interpreted cautiously because reported performance metrics depend strongly on the UV variable considered, temporal resolution, geographical conditions, forecasting horizon, and experimental protocol. Previous works have focused on UVI forecasting [
15,
18], cloud-affected UV prediction [
16], clear-sky UV mapping [
17], UV estimation [
19,
22,
23], and satellite-based UV monitoring [
20]. Consequently, direct numerical comparisons of RMSE and MAE values across studies may be misleading.
The main contribution of the present work is therefore methodological rather than purely predictive. Unlike many previous studies, the proposed framework incorporates a fixed forecasting horizon, chronological holdout evaluation, temporal cross-validation, statistical significance testing, and quantitative residual diagnostics within a fully reproducible experimental design.
From an applied perspective, a reliable 60 min ahead UV radiation forecast may support short-term monitoring and early warning systems, providing useful information for public health communication, occupational safety, and outdoor exposure management. Nevertheless, because the study was conducted using data from a single meteorological station in Lima, the findings should be interpreted as locally valid and not directly generalizable to broader geographical regions.
A limitation of this study is that only UV radiation intensity measurements were used as forecasting inputs. Consequently, potentially informative atmospheric drivers such as cloud cover, humidity, ozone concentration, aerosols, and temperature were not explicitly modeled.
Broader Implications for Climate Action
The proposed forecasting framework contributes to climate adaptation and public health preparedness by enabling short-term anticipation of high-radiation events. Timely UV forecasts can support preventive communication strategies, risk management, and environmental awareness in urban areas exposed to intense solar radiation.
However, operational implementation requires further validation using multiple monitoring stations and additional meteorological covariates. Future research should evaluate spatial generalization, incorporate atmospheric predictors, and assess performance under diverse climatic conditions. In this context, the proposed model represents a methodological step toward climate-resilient environmental monitoring systems and data-driven decision support tools.
6. Conclusions
This study evaluated statistical, recurrent deep learning, and hybrid deep learning models for short-term ultraviolet (UV) radiation intensity forecasting in Lima, Peru. Using a unified experimental protocol based on 5 min temporal aggregation, daytime filtering, a fixed 60 min forecasting horizon, chronological train–test partitioning, temporal cross-validation, statistical significance testing, quantitative residual diagnostics, and temporal dependence analysis, the results demonstrated that deep learning and hybrid neural architectures substantially outperformed traditional statistical forecasting models.
Among all evaluated approaches, Hybrid Model 2 achieved the best performance on the independent holdout test set, obtaining MAE = 0.1618, RMSE = 0.3849, and . This model outperformed Hybrid Model 1, GRU, LSTM, Prophet, ARIMA, SARIMA, and the naive baseline, indicating that the combination of convolutional feature extraction and recurrent sequence modeling is effective for capturing the nonlinear dynamics of UV radiation intensity.
The Diebold–Mariano test indicated statistically significant forecasting improvements of Hybrid Model 2 over the classical statistical models, whereas differences among neural forecasting models were not statistically significant. Temporal cross-validation showed that GRU achieved slightly greater stability across folds, whereas Hybrid Model 2 consistently obtained the best predictive accuracy on the independent test set. These results suggest a trade-off between forecasting stability and maximum predictive performance.
Quantitative residual diagnostics revealed a residual mean of 0.1034 and moderate residual dispersion, indicating limited systematic forecasting bias. However, the observed positive skewness, high kurtosis, and statistically significant Ljung–Box results indicate that extreme UV radiation peaks remain the principal source of forecasting uncertainty and that residual temporal dependence persists. These findings suggest that additional meteorological predictors, including cloud cover, humidity, ozone concentration, temperature, and atmospheric aerosols, may further improve forecasting accuracy.
From an applied perspective, the proposed framework provides a reproducible methodology for short-term UV monitoring and early warning systems. A 60 min ahead forecasting capability may support preventive public health actions, occupational exposure management, and climate-resilient environmental monitoring. Nevertheless, the findings should be interpreted as locally valid for the analyzed meteorological station. Future research should evaluate multi-station datasets, incorporate additional atmospheric covariates, and investigate longer forecasting horizons and spatial generalization across different climatic regions.
The results demonstrate that hybrid deep learning architectures can effectively capture the nonlinear temporal dynamics of UV radiation intensity, making them promising tools for operational short-term forecasting applications.
Overall, this study contributes to the application of artificial intelligence in environmental risk forecasting by providing a statistically validated and reproducible framework for UV radiation prediction. The proposed methodology supports future intelligent UV warning systems aligned with Sustainable Development Goal 13 (Climate Action) and Sustainable Development Goal 3 (Good Health and Well-being).