1. Introduction
The issue of energy consumption has increased the likelihood of a global energy crisis having environmental consequences. Sustainable energy is becoming a necessity, as significantly declining fossil fuel sources have raised severe economic, political, and societal challenges [
1]. Continuous measurement and forecasting are quite necessary for optimizing power generation through wind, water, and solar energy sources according to energy demand. Solar energy is among the most popular renewable energy sources because it does not harm the environment, is cost-effective, and is readily available. The sun, on average, impinges Earth with about 1.5 × 10
8 kWh of energy every year, and that is much more than what can suffice for the entire world’s energy needs [
2]. The great market growth of photovoltaics (PV) has indeed been due to the decrease in the cost of PV modules, the continued support of national policies, and particularly renewable energy targets that have unlocked this potential. Accordingly, global PV capacity increased substantially, from 6.6 GW in 2006 to more than 1 TW by 2024. However, the increased benefits associated with solar energy have some disadvantages, one of them being dependence on sunlight. Other conditions that influence solar energy generation at different times include the time of day, the season, the weather, and the location on Earth. Demand forecasting effectively avoids costly errors and disruption of a day-ahead market energy network. This need arises from the global increase in energy demand and the growing environmental concerns. Quite complicated forecasting systems have been developed, particularly for renewable sources such as solar and wind power, which on their own contribute hugely to balancing grid power and optimizing resource use [
3].
High forecasting accuracy is achieved by analyzing temperature, humidity, rainfall, solar radiation, cloud cover, wind speed, and wind direction. Many recent artificial intelligence (AI)-based methods have been developed to capture the non-linear aspects of renewable energy data [
4]. Whereas the traditional physical model could not adequately replicate the periodicities and spectral energy distributions of renewable energy data, AI models, especially soft computing and deep learning, were quite superior in this regard. The current research shows a trend towards using sophisticated renewable energy forecasting methods, according to [
5]. Energy systems now depend on AI and machine learning (ML) technologies to achieve greater operational flexibility, which enables better utilization of renewable energy resources because of their fluctuation. The system upgrades of the machine learning system improve its predictions while providing current energy distribution information. Hybrid models currently use AI together with traditional forecasting methods to create more flexible forecasting systems. The deep learning-based forecasting models demonstrate excellent accuracy and efficiency, which allows them to deliver consistent results for solar energy management and solar energy forecasting. Various machine learning models, including Random Forest (RF), Linear Regression (LR), Support Vector Regression (SVR), and Artificial Neural Network (ANN), were employed for this purpose [
6]. The ANN technique primarily serves short-term solar energy forecasting because it can accurately forecast solar radiation production through its ability to model complicated nonlinear and dynamic solar radiation patterns. Researchers have found that deep learning methods perform best when they need to analyze intricate datasets that include nonlinear patterns [
5]. The advanced models improve their performance, but they face challenges when dealing with multi-level data interactions, which lead to unpredictable model behavior in real-life situations that need exact energy management and grid stability forecasting [
6].
The solar radiation levels exhibit non-stationary behavior because they depend primarily on three factors, which include cloud cover, atmospheric disturbances, and seasonal changes. The time-domain features fail to provide sufficient information because they obstruct seasonal and cyclical patterns that forecasters need for their predictions.
Thus, temporal models may struggle to generalize across weather conditions, reducing forecast accuracy and robustness. This shortcoming highlights the need for further representations that can better separate periodic behaviors from instantaneous fluctuations, particularly in fast-changing weather conditions.
Frequency-domain analysis is currently receiving attention in solar radiation forecasting, but most research uses Fourier transforms for signal denoising or as incorporated components of deep learning frameworks. Modern FFT-based approaches use implicit feature learning to incorporate spectrum changes directly into convolutional or lightweight neural network models without explicitly providing frequency-domain characteristics. Thus, seasonal frequencies, spectral energy distributions, and entropy-related patterns are rarely explored in a way that typical machine learning and deep learning models can use. Solar radiation forecasting has not been widely researched using explicit spectral feature engineering and rigorous hyperparameter optimization. In this study, an FFT-based spectral feature engineering approach extracts interpretable frequency-domain characteristics from meteorological time-series data and integrates them with Bayesian Optimization to increase prediction accuracy and robustness under varied weather situations. Using FFT-based features [
7], models can generalize their results under a variety of conditions by transitioning from the time to the frequency. Besides feature engineering, we also employ Bayesian Optimization (BO) for optimal adjustment of hyperparameters in machine learning models. BO effectively explores the parameter space so that forecasting errors are minimal, thereby increasing predictive performance. By integrating BO, the following main advantages contribute to solar radiation forecasting.
We introduced an FFT spectral feature engineering procedure so that solar radiation model forecasts could become more accurate.
Bayesian Optimization (BO) of the hyperparameters of machine learning models ensured that their performance and predictive accuracy were vastly improved.
Even while the time-related structure of solar radiation data is preserved, assessments of models are strictly carried out by the Time-Series Cross-Validation (TSCV) methodology.
Based on SHAP, distribution plots, and performance curves, the model’s performance is compared under seasonal variations and complex data patterns.
Performance evaluation of machine-learning models is performed under multiple weather conditions, namely, Clear Sky, Cloudy, High Humidity, and Mixed Conditions.
The remaining sections of this paper are structured as follows. Following a review of prior work and the research methods, we present an overview of the dataset, the preprocessing stages, and the FFT-based framework. This paper analyzes environmental elements that affect solar irradiance and discusses how the framework captures complicated relationships. We explain the TSCV approach, model applications, and performance evaluation, followed by a discussion analyzing the experimental data and situating the suggested forecasting paradigm within existing research.
2. Related Work
The global transition to sustainable energy has raised interest in renewable energy, particularly solar power. Future energy scenarios will rely heavily on solar energy due to significant expansion and technological advances [
5]. The process of managing electricity distribution becomes more difficult because solar power generation depends on unpredictable weather conditions. As a result, addressing these issues requires accurate forecasting of solar radiation [
8].
Li et al. [
9] utilized numerical weather prediction (NWP) models to forecast solar irradiance on the basis of atmospheric conditions. Although NWP models are useful, results indicate that they can be further improved. Of utmost importance is the quality and accuracy of the meteorological data that serve as the inputs. This is a major drawback in any place where there is not sufficient weather data. Mittal [
10] examined time-series methods, including ARIMA and exponential smoothing, for the forecasting of solar energy. These methods have been effective in using historical data for short-term forecasting, assuming that future trends will become similar to those of historical data, which is not always the case, especially with the continuous changes in weather conditions.
Regression models using linear or non-linear analysis of solar radiation data were studied by Antoñanzas [
11]. The effectiveness of such models is highly dependent on the complexity of the selected model. Simple models may lead to underfitting, while complex models may lead to overfitting. Ramirez-Vergara [
12] investigated the probabilistic forecasting of solar radiation using Markov Chain and Bayesian probabilistic approaches, where prediction uncertainty is quantified, which is an essential factor in the risk management of power systems. However, in effecting its application, relatively very large datasets and processing resources not commonly available have to be used. Voyant [
12,
13,
14] incorporated supervised machine learning techniques in their research, such as support vector regression. These models exhibit quite believable assumptions but do not promise very high accuracy in predictions. This is because efficacy in historical data can be maximized only when appropriate input factors and models are employed. They involve a very significant amount of computing resources, especially when using a large number of databases. The prediction accuracy of the above forecasts is increased due to their incorporation of nonlinear data interactions.
Ahmad and Chenn [
15] assert that ensemble approaches like bagging and boosting are effective. With many models combined, increased forecasting precision will be achieved. Such processes are very complex and tend to be resource-intensive despite inherent advantages. Their research pointed out that ensemble approaches like bagging, boosting, and random forests are effective in enhancing accuracy in prediction through the combination of many models. The strategies do enhance predictive accuracy, but their implementation becomes difficult because they need powerful computing systems, which demand extensive computer resources.
Kumari and Toshniwa [
14] used deep learning to forecast solar irradiance via CNNs. More importantly, CNNs are great at recognizing spatial patterns, which makes them really effective in the analysis of satellite pictures. Still, they might be constrained by huge datasets and resource-poor processing usage. Rajagukguk [
16] examined sequential data LSTM networks for long-term dependency. Due to their substantial processing requirements and extensive datasets, Long Short-Term Memory networks are highly effective in forecasting solar radiation and power output. While an LSTM model may have yielded superior predictive accuracy through its extended memory capacity, the model did not exhibit learning behaviors.
Alameen et al. [
17] produced accurate solar irradiance data by Generative Adversarial Networks (GANs) to enhance the quality of training data, which in turn increases prediction accuracy. GANs are beneficial in many aspects; however, for their utility, they require high computation and model skill. The framework of flexible hybrid ensemble FHE in Song 18 establishes a system that selects base models according to their performance in prediction errors to achieve better forecasting results through hybrid model feature combination. Integrated models provide superior accuracy compared to their standalone counterparts, but they demand additional processing power to operate.
Significant advancements have been made in the field of solar forecasting; however, numerous challenges persist due to the complexities and variability inherent in weather patterns and solar radiation. The development of models that deliver both precise results and fast processing times presents a major obstacle [
18]. The methodology we present here provides a solution to the existing problem. We applied FFT-based spectral feature engineering to identify essential frequency patterns that additional analysis methods fail to detect. Traditional time-based modeling methods fail to detect yearly cycles and recurring trends that remain hidden in their time-based modeling approaches. The model improves its ability to recognize fundamental patterns in solar and meteorological data changes, which leads to better and more precise predictions. The results demonstrate that further feature engineering work is needed to improve solar radiation prediction results through spectral analysis, which will enhance accuracy, stability, and operational value.
4. Results
In this section, we present and explain the results of our experiments to validate the efficiency of our proposed forecasting models. All machine learning (ML) models were developed using Python with the Scikit-learn library. The performance of the model is assessed in terms of three standard metrics: MAE, RMSE, and R
2. These measures examine predictive accuracy, magnitude of error, and model fit, respectively, providing a robust framework for judging the effectiveness of the DTIFS framework. The improvement seen across the various models due to Bayesian optimization has been greatest with the MLP model among others.
Table 4 indicates the comparison between the basic model setups and three hyperparameter tuning methods, which include random search, grid search, and Bayesian Optimization, for all evaluated models. The tuning strategy shows a consistent improvement pattern that starts from baseline settings and reaches its highest point with Bayesian Optimization. The Bayesian Optimization method produces the best MAE and RMSE results and the highest R
2 values for all models, resulting in improved prediction accuracy compared to both baseline configurations and simpler tuning methods. The absolute magnitude of improvement introduced by Bayesian Optimization shows clear progress, and its impact becomes evident when compared with random and grid search methods. The Random Forest model experiences an MAE reduction from 2.52 to 2.47, while R
2 increases from 0.88 to 0.90 after applying Bayesian Optimization. The Decision Tree, Linear Regression, and KNN models demonstrate the same pattern of continuous improvement, as Bayesian Optimization achieves better results than random and grid search for all evaluation metrics, although the models’ limited hyperparameter spaces result in smaller performance gains. Neural network models demonstrate the most pronounced effects from Bayesian Optimization. The MLP model shows the largest improvement, with MAE decreasing from 1.83 to 1.78, RMSE decreasing from 2.82 to 2.75, and R
2 increasing from 0.90 to 0.92. The LSTM and GRU models exhibit similar trends, with consistent reductions in MAE and corresponding increases in R
2. Overall, the results indicate that Bayesian Optimization performs best for models operating in high-dimensional hyperparameter spaces that require complex searches, as structured optimization enables more effective exploration than random or exhaustive tuning approaches.
Once the implemented spectral feature engineering and optimization were applied, the second test proved quite significant, which can be seen from
Table 5. The largest improvement in this respect was observed for the MLP, wherein MAE reduced from 1.82 to 1.68, RMSE decreased from 2.80 to 2.55, and R
2 rose from 0.89 to 0.91, which translates into better accuracy and stability. However, RF also improved by decreasing MAE from 2.51 to 2.36 and increasing R
2 from 0.87 to 0.90. There was a modest improvement for LR and KNN, while DT showed a mixed bag of benefits, showing MAE improvement from 2.72 to 2.61 and R
2 improvement from 0.81 to 0.83. The LSTM and GRU were also found to benefit from feature engineering, as measured by MAE and R
2. The feature engineering enhancement differentially improved all models, but particularly MLP and RF, to produce increased accuracy and reliability in the prediction of solar radiation.
MLP showed superior prediction over all other models across different weather scenarios: Clear Sky (CS), Cloudy (CL), High Humidity (HH), and Mixed Conditions (MC). Under Clear Sky, MLP achieved its best performance with an R
2 of 0.92. On the other hand, RF was fairly accurate and stable across scenarios. KNN and LR, however, showed poor performance under conditions of High Humidity and Mixed Conditions, indicating their lack of robustness. In contrast, although LSTM and GRU gave fairly uniformly good results, their accuracy could not stand up to that of MLP and RF. This proved that weather variability has an impact on forecasting accuracy, whereas the more sophisticated models, such as MLP, tend to be more robust against environmental fluctuations.
Table 6 summarizes model performances over different weather scenarios.
Forecasting models use time-series cross-validation to maintain temporal ordering and prevent data leakage [
20]. Model stability evaluation in the temporal domain was conducted using Time-Series Cross-Validation (TSCV), which involves segmenting the dataset into four chronological folds, the results of which are presented in
Table 7. The method maintains temporal integrity for subdivided data while conducting model tests throughout seasonal changes. MLP delivered the best performance, which showed continuous MAE and RMSE reduction together with R
2 values between 0.92 and 0.93. RF achieved second place, with an R
2 range between 0.88 and 0.90. KNN ranked as the weakest performer among the three because it experienced significant drops in performance during early folds, which demonstrated its high sensitivity to time-based changes in solar data. LSTM and GRU maintained stable performance across all folds, but they achieved lower accuracy rates than MLP and RF. We applied a four-fold rolling-window TSCV to maintain temporal continuity while ensuring that all training windows contained sufficient seasonal information.
The experimental design enables observation of model performance differences resulting from different model architectures and feature selection methods, rather than from variations in datasets or evaluation procedures. The proposed FFT-based framework demonstrates superior performance compared to CNN and LSTM methods because it allows identification of specific frequency-domain features of the system. The proposed method enables direct encoding of dominant periodicities, seasonal cycles, and spectral energy patterns, whereas CNN and LSTM models rely on automatic feature extraction from time-domain sequences. The explicit spectral representation of data reduces learning difficulty, improves model stability, and enhances generalization ability under severe weather variability.
The scatter plot shows the relationship between predicted and actual values for solar radiation output, with the data points clustering around the ideal diagonal line, indicating strong predictive accuracy.
Figure 5 shows the performance of various machine learning models, comparing the performance of different machine learning models. With the lowest MAE of 1.78 and RMSE of 2.75, and the highest R
2 (0.92), MLP is the best model and performs better than all others. This shows how reliable and precise MLP is at predicting solar radiation. The next closest is RF, which is not far behind but has a lower R
2 value. Conversely, models with lower R
2 and higher error, like KNN and LR, show poorer predictive abilities. All findings support that MLP is the most accurate model for forecasting solar energy generation in this study.
As shown in
Figure 6, the smoothed cumulative error distribution (CED) plot clearly indicates that the MLP model achieves the most reliable prediction, with more than 90% of the predictions being less than a 3 MJ/m
2 error threshold. Random Forest and LSTM models roughly follow the same trends, showing a great consistency with respect to their predictability. KNN and LR fall far behind in slow cumulative growth in predictive accuracy. Again, the steepness of the MLP curve proves how well this model generalizes on varying conditions. Such outcome results confirm the success of deep learning, further improved by engineered features and BO. The same pattern results have also authenticated the robustness of the simulation model.
A thorough study of the 3D surface that represents day length and temperature against solar irradiance clearly exhibits evidence of nonlinear interactions among these variables, as illustrated in
Figure 7. With an increase in day length, solar irradiance initially increases, following a seasonal pattern aligned with that of the solar cycle. Temperature positively affects irradiance, but beyond about 30 °C, the effect appears to plateau, probably due to thermal saturation or atmospheric attenuation. The joint effect maximizes irradiance under conditions of long day length and moderate-to-high temperatures. These findings endorse the rationale behind including temporal variables in the forecasting model and thus justify the use of spectral feature engineering and MLP-based architectures for adequately modeling seasonal and nonlinear behaviors in solar radiation prediction.
Figure 8 depicts the residual distributions of seven models, which provide insight into how each of those models performed both in prediction accuracy and consistency. The distributions of LSTM and GRU are compact, symmetric, and zero-centered, which means that their error variance is low, along with bias, both being characteristics of an equally structured and generalized model. The performance of MLP is more or less comparable but relatively breached at the end. In turn, there are wider and skewed distributions in Linear Regression and KNN with obvious bias, thus representing under- or over-predictions systematically. Random Forest sits between the two extremes, having medium dispersion and a few extreme residuals. These residual characteristics increase the credibility for the understanding that deep learning models (LSTM and GRU) indeed outperform other, more straightforward or “shallow” models for capturing nonlinearities and temporal dependencies that are embedded in solar irradiance forecasting.
6. Conclusions
In this study, we propose a novel algorithm combining FFT with BO for solar radiation forecasting. While the algorithm is developed to find meaningful patterns in solar data by designing specific features, BO is used to tune the model parameters. The proposed approach was applied to MLP and successfully achieved an R2 of 0.91, MAE of 1.68, and RMSE of 2.55, outperforming traditional models. The results show that domain-specific knowledge, feature engineering, and hyperparameter optimization can help improve prediction accuracy. Time-series cross-validation was used to test the model, which showed that the proposed approach outperforms and is more stable than other recent approaches across different time horizons. In summary, the combination of the FFT framework and BO provides more reliable solar energy predictions, which are required for the integration of renewable energy into the electrical grid and for improved power system management. From a theoretical perspective, this study demonstrates the value of integrating explicit spectral feature engineering with machine learning and deep learning models for non-stationary solar radiation forecasting. By transforming meteorological time-series data into the frequency domain, the proposed framework enables more effective learning of dominant periodicities and seasonal patterns that are often obscured in purely time-domain representations, thereby improving model generalization under variable weather conditions. From a practical standpoint, the proposed FFT-based framework enhances forecasting accuracy and robustness, supporting more reliable grid operation, photovoltaic generation scheduling, and energy management. The combination of spectral feature engineering with Bayesian Optimization also improves adaptability across different data characteristics and computational settings, making the approach suitable for operational use. Future research may explore multi-resolution spectral techniques, such as wavelet or hybrid FFT–wavelet methods, as well as probabilistic forecasting to quantify uncertainty. Further studies could also examine the generalization of the proposed framework across different geographic regions, climatic conditions, and higher temporal resolutions.