Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (115)

Search Parameters:
Keywords = Diebold–Mariano test

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
35 pages, 2715 KB  
Article
A Flexible Kernel Regression Decomposition-Based Hybrid Framework for Forecasting Foreign Exchange Rates: Uncovering Symmetric and Asymmetric Market Dynamics
by Hasnain Iftikhar, Said Farooq Shah, Fatimah E. Almuhayfith and Paulo Canas Rodrigues
Symmetry 2026, 18(8), 1372; https://doi.org/10.3390/sym18081372 - 14 Aug 2026
Abstract
Foreign exchange rate forecasting is difficult, as its time series have complex nonlinear and nonstationary dynamics. To this end, the paper proposes a hybrid forecasting framework based on kernel regression decomposition that combines nonparametric trend extraction with linear and nonlinear forecasting models. In [...] Read more.
Foreign exchange rate forecasting is difficult, as its time series have complex nonlinear and nonstationary dynamics. To this end, the paper proposes a hybrid forecasting framework based on kernel regression decomposition that combines nonparametric trend extraction with linear and nonlinear forecasting models. In the framework we propose, we decompose each series of exchange rates into a long-run trend component and a short-run fluctuation component, which are modeled separately and then combined to generate forecasts. An empirical study based on five major foreign exchange markets shows that the proposed framework consistently outperforms conventional hybrid and standalone forecasting models in multiple forecasting accuracy measures. The statistical validation of the forecasting improvements is confirmed by the Diebold–Mariano test. Moreover, the decomposition offers an interpretable representation of exchange-rate dynamics by distilling persistent long-run movements from localized short-run nonlinear fluctuations. The results indicate that the proposed framework improves forecasting accuracy and offers insight into the structural behavior of foreign exchange markets. Full article
(This article belongs to the Section B: Mathematics)
Show Figures

Figure 1

20 pages, 6359 KB  
Article
Ramp Event Directional Forecasting for Wind Power Integration: A Regime-Stratified Ensemble Framework with Direction-Focused Training
by Konstantinos Stergiou and Theodoros E. Karakasidis
Energies 2026, 19(16), 3794; https://doi.org/10.3390/en19163794 - 12 Aug 2026
Viewed by 146
Abstract
Wind power ramp events (abrupt swings in output driven by frontal passages, sea-breeze transitions, and turbulence) are among the hardest problems for operators integrating renewables. Forecasting models are usually judged by aggregate error metrics (MAE, RMSE, overall directional accuracy) that average stable and [...] Read more.
Wind power ramp events (abrupt swings in output driven by frontal passages, sea-breeze transitions, and turbulence) are among the hardest problems for operators integrating renewables. Forecasting models are usually judged by aggregate error metrics (MAE, RMSE, overall directional accuracy) that average stable and ramp periods together, masking how a model behaves during the ramps that actually stress the grid. We address this on two fronts. First, we propose a regime-stratified evaluation that reports ramp event directional accuracy (ramp-DA) separately from stable-period accuracy and argue that ramp-DA should be a primary metric for grid-integration forecasting. Second, we build an ensemble of five regime-specialised sub-models trained with a direction-focused loss that penalises sign errors in the forecast power change, using only on-site SCADA wind speed and power. On 33,411 held-out samples from three onshore Greek farms, the ensemble reaches 76.5% ramp-DA, against 70.1% for a two-layer LSTM (+6.4 pp) and 74.3% and 74.1% for the PatchTST and iTransformer baselines. A strict leave-one-farm-out test retains 77.4% ramp-DA on a fully unseen farm. Overall directional accuracy rises 3.4 points, evidence that aggregate metrics understate the ramp-focused gain, while mean absolute error falls 19% compared to the LSTM (Diebold–Mariano p < 0.001). Full article
(This article belongs to the Special Issue Application of Machine Learning in Modern Power Systems)
Show Figures

Figure 1

28 pages, 3224 KB  
Article
Forecasting Multivariate Time Series: A Comparison of Machine Learning, Statistical and Deep Learning Models
by Dler Hussein Kadir, Diyar Muadh Khalil and Azhin Muhammed Khudhur
Forecasting 2026, 8(4), 73; https://doi.org/10.3390/forecast8040073 - 12 Aug 2026
Viewed by 297
Abstract
This study develops a rigorous, leakage-free forecasting framework for monthly Robusta coffee prices using historical observations from January 1975 to December 2025. A comprehensive set of explanatory variables is constructed from lagged coffee prices, moving averages, logarithmic returns, rolling volatility, and exogenous variables [...] Read more.
This study develops a rigorous, leakage-free forecasting framework for monthly Robusta coffee prices using historical observations from January 1975 to December 2025. A comprehensive set of explanatory variables is constructed from lagged coffee prices, moving averages, logarithmic returns, rolling volatility, and exogenous variables such as the Oceanic Niño Index (ONI), the U.S. Dollar Index, and Brent crude oil prices. To ensure methodological fairness, all predictors are generated exclusively from information available at the forecast origin, and all competing models are evaluated under a unified expanding-window walk-forward validation framework. Seven forecasting models are compared: Naïve, Exponential Smoothing (ETS), ARIMA, ARIMAX, Extreme Gradient Boosting (XGBoost), Long Short-Term Memory (LSTM), and Gated Recurrent Unit (GRU). Forecasting performance is evaluated using R2, RMSE, MAE, and MAPE, while Taylor diagrams and the Diebold–Mariano test are employed to assess model agreement and differences in predictive accuracy. The results show that XGBoost achieves the highest forecasting accuracy (R2 = 0.956, RMSE = 0.264), followed closely by the Naïve (R2 = 0.954, RMSE = 0.271) and ARIMA (R2 = 0.954, RMSE = 0.270) benchmarks, whereas ARIMAX and ETS provide comparable performance and the deep learning models (LSTM and GRU) produce substantially larger prediction errors. Feature importance analysis further indicates that the first lag of coffee price is the dominant predictor, accounting for approximately 94% of the predictive gain in XGBoost. Overall, the findings demonstrate that rigorous leakage-free validation is essential for reliable forecasting research and that, for monthly Robusta coffee prices, increased model complexity does not necessarily yield superior predictive performance. Full article
Show Figures

Figure 1

33 pages, 8015 KB  
Article
A Physics-Constrained Super-Resolution Framework for Wind Resource Mapping: Application to Brazil
by José Péricles Freire, Lihki Rubio, Jin Yang and Carlos E. Velasquez
Energies 2026, 19(16), 3722; https://doi.org/10.3390/en19163722 - 7 Aug 2026
Viewed by 313
Abstract
High-resolution wind resource maps are essential for wind farm siting and renewable-energy planning, but national-scale assessment is often limited by sparse meteorological stations and the coarse resolution of reanalysis products. This study introduces a self-supervised Physics-Constrained Super-Resolution CNN (PC-SRCNN) that enhances ERA5 wind [...] Read more.
High-resolution wind resource maps are essential for wind farm siting and renewable-energy planning, but national-scale assessment is often limited by sparse meteorological stations and the coarse resolution of reanalysis products. This study introduces a self-supervised Physics-Constrained Super-Resolution CNN (PC-SRCNN) that enhances ERA5 wind fields from 0.25° to 0.025°, for settings lacking a high-resolution, hourly-resolved reference wind field, by embedding kinematic and vertical-profile consistency as soft regularization penalties, without solving the full Navier-Stokes momentum balance. The framework was evaluated against 190 independent INMET stations, alongside interpolation, data-driven, and climatological baselines, using a station-wise bootstrap and the Diebold-Mariano test. The full model reduced the bootstrap MAE from 1.22 to 1.04 m/s relative to native ERA5, a 14.85% improvement confirmed by both tests, and achieved a vertical-profile R2 of 0.90 against baselines below 0.70. Ablation shows physical regularization is not automatically beneficial: a variant constrained only by horizontal kinematics performed significantly worse than ERA5, with gains obtained only when vertical logarithmic-profile consistency was included. All other baselines evaluated were statistically indistinguishable from native ERA5 against station observations. These results indicate that sharper spatial detail alone is insufficient to improve wind-resource estimates unless supported by physically meaningful, vertically consistent constraints. Full article
(This article belongs to the Section F5: Artificial Intelligence and Smart Energy)
Show Figures

Figure 1

32 pages, 5243 KB  
Article
A Comparative Study of Multi-Scale Hybrid Deep Learning Frameworks for Estimation of Domestic Load Demand of Pakistan’s Central Region
by Muhammad Yousouf Bashir, Mustafa Shakir, Ali Raza, Manzoor Ellahi and Mohsin Jamil
Sensors 2026, 26(15), 4991; https://doi.org/10.3390/s26154991 - 6 Aug 2026
Viewed by 304
Abstract
In this technological era, electrical energy is the bloodstream for the economic and social development of any country. It is the need of the time that developing countries like Pakistan have strategic planning for efficient generation and utilisation of electricity and have as [...] Read more.
In this technological era, electrical energy is the bloodstream for the economic and social development of any country. It is the need of the time that developing countries like Pakistan have strategic planning for efficient generation and utilisation of electricity and have as much cheap electricity as possible at their disposal while respecting environmental constraints. The overloaded and ageing infrastructure of an electrical power network can impact system reliability and the sustainability of power generation, transmission and distribution mechanisms. The initiation of the planning process depends upon accurate load estimation to optimally fulfil consumers’ power needs. This paper compares statistical, hybrid and deep learning (DL) mechanisms, including SARIMAX, SARIMA with gradient boosting (SARIMA-GB), long short-term memory (LSTM) network, STL decomposition with LSTM, and CWT-LeNet-5-LSTM, for the prediction of residential electricity demand in the LESCO region of central Pakistan. The study uses 7670 daily feeder observations recorded between 1 January 2002 and 31 December 2022. The series is modelled at its native daily resolution and partitioned chronologically into a fitting span of 5216 days, a validation span of 920 days and a test span of 1534 days beginning 20 October 2018. All models receive the same block of 14 exogenous calendar variables, four annual Fourier harmonic pairs, day-of-week and month sine and cosine terms, a weekend indicator and a linear trend, and all neural models are trained with a validation split, early stopping and restoration of the best weights rather than for a fixed number of epochs. Accuracy is assessed with MAE, RMSE, MAPE and peak normalized RMSE under one recursive protocol at forecast leads of 1, 7, 14 and 30 days, because the ranking of the frameworks depends on the lead. Averaged over three random initialisations, the proposed CWT-LeNet-5-LSTM attains the lowest error at every multi-step lead, reaching an MAPE of 2.35 ± 0.28% at lead 30 against 3.32% for the multivariate LSTM, 3.89% for SARIMA-GB and 4.47% for SARIMAX. At lead 1, SARIMA-GB is the more accurate model (0.76% against 1.20 ± 0.17%) because the previous day’s observed load dominates one-step prediction for a series whose lag-one autocorrelation is 0.978. An architecture ablation isolates the contribution of the wavelet stage, the convolutional stage, the anisotropic pooling and the calendar fusion. A stratified analysis across seasons, weekdays, weekends and high-, medium- and low-load days shows where the advantage is concentrated. Additionally, Diebold–Mariano tests identify the statistical significance of the differences. Full article
(This article belongs to the Section Intelligent Sensors)
Show Figures

Figure 1

28 pages, 880 KB  
Article
Fractional Long-Memory Dynamics and Residual Machine Learning for Medium-Horizon Agricultural Commodity Price Forecasting
by Sergio Orozco Cirilo, Juan Manuel Vargas-Canales, Dora María Sangerman Jarquín, Sergio Ernesto Medina Cuéllar, Juan Antonio Bautista, Alberto Valdes Cobos, Benito Rodríguez Haros and Belén Hernández Hernández
Fractal Fract. 2026, 10(8), 536; https://doi.org/10.3390/fractalfract10080536 - 6 Aug 2026
Viewed by 159
Abstract
This paper develops a hybrid fractional-order framework for medium-horizon forecasting of wheat, corn, and soybean prices, combining a Caputo fractional differential equation with exogenous macro-climatic drivers (weather, crude oil, exchange rate, inflation) and a machine learning residual-correction layer. Existence, uniqueness, and Ulam–Hyers stability [...] Read more.
This paper develops a hybrid fractional-order framework for medium-horizon forecasting of wheat, corn, and soybean prices, combining a Caputo fractional differential equation with exogenous macro-climatic drivers (weather, crude oil, exchange rate, inflation) and a machine learning residual-correction layer. Existence, uniqueness, and Ulam–Hyers stability are established for both Caputo and Atangana–Baleanu formulations via Banach fixed-point theory, with numerical illustration through a fractional Adams–Bashforth–Moulton predictor–corrector scheme. Formal unit-root tests (ADF, KPSS, Phillips–Perron) confirm I(1) behaviour in log-price levels and stationarity in first differences. Residuals are corrected using XGBoost and two feedforward neural networks (MLP-A, MLP-B), producing a family of hybrid forecasting models. Using 25 years of monthly FRED data (January 2000–December 2024), multi-method long-memory diagnostics (Hurst exponent, Lo’s modified R/S test, DFA, local Whittle estimation) confirm near unit-root fractional integration in log-price levels (H[0.985,1.044]), with estimated fractional orders α^{0.737,0.884,0.451} for wheat, corn, and soybean. Under a strict rolling-origin protocol (17 origins, horizons h{1,5,6,12} months), conventional benchmarks remain competitive at h=1, but their MAPE degrades to 16.6–31.9% at h=12, while Frac+XGBoost error stays flat at 3.9–14.5%. This horizon-robust advantage holds across three market regimes (COVID-19 pandemic shock, 2021–2022 super-cycle, 2023–2024 normalisation) and is confirmed by Holm–Bonferroni-corrected Diebold–Mariano tests (p<0.001 at h=12). The model’s structural advantage emerges at h5 months, supporting procurement planning, food-security buffer stocks, and import budgeting. Full article
Show Figures

Figure 1

36 pages, 3231 KB  
Article
Predicting Commodity ETF Returns with Deep Learning: Overnight Versus Daytime Predictability Across Forecast Horizons
by Triparna Kundu, Sarthak Pattnaik and Eugene Pinsky
Commodities 2026, 5(3), 16; https://doi.org/10.3390/commodities5030016 - 1 Aug 2026
Viewed by 228
Abstract
Commodity prices are notoriously hard to forecast, and whether the returns of commodity exchange-traded funds (ETFs) can be predicted remains an open question. We compare three deep learning models, Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), and Transformer, for forecasting the returns [...] Read more.
Commodity prices are notoriously hard to forecast, and whether the returns of commodity exchange-traded funds (ETFs) can be predicted remains an open question. We compare three deep learning models, Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), and Transformer, for forecasting the returns of six Deutsche Bank commodity ETFs covering agriculture (DBA), base metals (DBB), broad commodities (DBC), energy (DBE), oil (DBO), and precious metals (DBP). Using daily price data from January 2007 to December 2025, we predict daytime returns (open to close) and overnight returns (previous close to open) separately, over five horizons of 1, 5, 30, 60, and 180 trading days. Each model sees a 20-day window of price-based features, returns, rolling averages and volatilities, momentum, and recent lags, built from all six ETFs. All models are trained on a strict chronological split and judged by two simple, decision-oriented measures: how often they call the direction correctly, and the risk-adjusted return (annualized Sharpe ratio) of a stylized long–short strategy that ignores transaction costs. Formal significance tests with HAC corrections for overlapping targets, bootstrap confidence intervals, and comparisons with ARIMA, random forest, and simpler benchmarks corroborate strong predictability in overnight DBP and daytime DBB at medium horizons. Predictability turns out to be highly specific to the asset, the trading session, and the horizon. Overnight returns of the precious metals ETF (DBP) are by far the most predictable: the correct direction is called 71.6% of the time at 60 days and 76.7% at 180 days, with Sharpe ratios reaching about 15. Base metals (DBB) daytime returns are predictable at 30 days and oil (DBO) daytime returns at 180 days, whereas one-day-ahead forecasts and agricultural returns (DBA) stay essentially unpredictable. The Transformer has a slight edge at longer horizons and the GRU at shorter ones. Key directional accuracy and Sharpe ratio results are confirmed by Newey–West HAC significance tests and Diebold–Mariano forecast comparison tests with the Harvey–Leybourne–Newbold small-sample correction; HAC standard errors at the 180-day horizon exceed naïve OLS errors by a factor of approximately 7.4, and we explicitly flag results that do not survive this correction. A three-fold expanding walk-forward validation scheme corroborates the main findings, with DBP overnight and DBO daytime predictability persisting across all evaluation windows. Deep learning architectures statistically and economically outperform logistic regression, ridge regression, and momentum baselines on the most predictable configurations. An anomalous failure of all models on DBA daytime returns at the 180-day horizon is diagnosed as a regime-driven artefact associated with post-2021 commodity inflation, not a general feature of agricultural return dynamics. The broader lesson is that splitting returns into daytime and overnight components exposes predictable structure that conventional close-to-close returns hide. Full article
Show Figures

Figure 1

42 pages, 4364 KB  
Article
Week-Ahead Electricity Price Forecasting for Battery Arbitrage: Benchmarking ML/DL Models and Interpreting Feature Importance Through Merit-Order Pricing in Spain
by Amgad Khamis, Francesco Crespi and David Sánchez
Forecasting 2026, 8(4), 61; https://doi.org/10.3390/forecast8040061 - 21 Jul 2026
Viewed by 604
Abstract
Accurate electricity price forecasting is essential for market participants seeking to optimise bidding and arbitrage strategies. This paper presents a week-ahead (168 h) hourly electricity price forecasting study for the Spanish day-ahead market. Nine competing models—two naïve baselines (a Seasonal Naïve and a [...] Read more.
Accurate electricity price forecasting is essential for market participants seeking to optimise bidding and arbitrage strategies. This paper presents a week-ahead (168 h) hourly electricity price forecasting study for the Spanish day-ahead market. Nine competing models—two naïve baselines (a Seasonal Naïve and a Day-of-Week persistence), a Lasso-estimated auto-regressive (LEAR) statistical benchmark, and six machine- and deep-learning models (CatBoost, Random Forest, LSTM, GRU, CNN, and a hybrid CNN–LSTM)—are benchmarked; the two leading models, CNN–LSTM and CatBoost, are then compared under exogenous-feature configurations. The analysis is complemented by an ex-post Add-One-In and Leave-One-Out feature-importance analysis, a controlled comparison of weather-input scenarios, and a rolling battery-arbitrage backtest that translates forecast quality into economic value. Under an endogenous benchmark of weekly rolling origins across 2024 (with a rotating start weekday) and Diebold–Mariano testing, a recursive CatBoost and the hybrid CNN–LSTM are statistically indistinguishable and both significantly outperform a direct multi-horizon CatBoost; once an operational (forecasted) weather input is added, recursive CatBoost becomes significantly the most accurate while remaining simpler and more stable to train, a ranking confirmed on a fully out-of-sample 2025 year. Operational weather forecasts are found to be the best weather input, recovering about 84% of the perfect-foresight weather improvement over a no-weather baseline, with the advantage concentrated at longer lead times. Natural-gas-fired generation emerged as the dominant explanatory feature, consistent with the marginal-pricing mechanism governing the Spanish market. In a rolling battery-arbitrage backtest on the out-of-sample 2025 year, a deployable forecast-driven 4-h grid-scale unit (200 MW/800 MWh) captured about 89% of perfect-foresight value at a 168 h optimisation horizon and about 87% at 24 h; extending the horizon from 24 h to 168 h added about 2.4% of profit, an optimisation-horizon (look-ahead) effect bounded at +4.5% under perfect foresight. Full article
(This article belongs to the Collection Energy Forecasting)
Show Figures

Figure 1

28 pages, 8314 KB  
Article
Predictive Model Based on Machine Learning to Determine Gold Price Fluctuation and Improve Trading Decisions
by Alexander Vladimir Velez Flores, Arturo Rafael Chayña Rodriguez, Wildor Jazmany Jara Vilca, Carlos Paul Hancco Ramos, Esteban Marín Paucara, Lucio Quea-Gutierrez, Juan Carlos Chayña-Contreras, Julian Apaza-Chino, Mario Serafín Cuentas Alvarado, Yesenia Fátima Llanque Añacata and Anibal Sucari León
J. Risk Financ. Manag. 2026, 19(7), 533; https://doi.org/10.3390/jrfm19070533 - 17 Jul 2026
Viewed by 731
Abstract
Gold’s price reflects currency, opportunity-cost, and safe-haven channels whose strength shifts across regimes, motivating an empirical, data-driven forecasting approach. This study develops a monthly gold price forecasting system for ASM sales-timing decisions in Peru (January 2020–June 2026) using macro-financial predictors including a geopolitical [...] Read more.
Gold’s price reflects currency, opportunity-cost, and safe-haven channels whose strength shifts across regimes, motivating an empirical, data-driven forecasting approach. This study develops a monthly gold price forecasting system for ASM sales-timing decisions in Peru (January 2020–June 2026) using macro-financial predictors including a geopolitical risk index and three U.S. monetary indicators, none of which were Granger-causal and were therefore excluded from the production set. After confirming non-stationarity and Johansen cointegration (four vectors), thirty-two model-feature-set combinations, including Elastic Net, Bayesian Ridge, and a PCA factor, were compared under strict temporal validation with bounded hyperparameter search. The selected model, Ridge regression on the CONTROL feature set, achieved a cross-validation MAPE of 2.29% and test MAPE of 3.62% (official)/3.15% (extended sensitivity window). It was benchmarked against random walk, historical mean, and exponential smoothing and evaluated via the Diebold–Mariano, Clark–West, encompassing, and Model Confidence Set tests (low-power caveats given the small sample). A dual-horizon Monte Carlo simulation, robust to heavy-tailed shocks, projected USD 4482/oz (December 2026) and USD 5106/oz (December 2027). A sales-timing backtest showed a statistically significant result (−0.67%) versus a passive strategy, indicating calibrated price information alone does not yet yield a reliable trading edge, supporting the model’s role as decision support rather than an autonomous trading signal. Full article
(This article belongs to the Section Financial Technology and Innovation)
Show Figures

Figure 1

46 pages, 2809 KB  
Article
Do Supply-Chain Stress and Geopolitical Risk Predict Strategic Commodity and Clean Energy Market Returns? Evidence from Explainable Machine Learning
by Nader Naifar
Forecasting 2026, 8(4), 59; https://doi.org/10.3390/forecast8040059 - 15 Jul 2026
Viewed by 465
Abstract
This study examines whether daily supply-chain stress and geopolitical risk improve the forecasting of strategic commodity and clean energy market returns. Using daily data on aluminum, copper, nickel, and clean energy from 10 February 2015 to 27 February 2026, the analysis compares a [...] Read more.
This study examines whether daily supply-chain stress and geopolitical risk improve the forecasting of strategic commodity and clean energy market returns. Using daily data on aluminum, copper, nickel, and clean energy from 10 February 2015 to 27 February 2026, the analysis compares a baseline forecasting model based on conventional market controls with augmented specifications that incorporate supply-chain stress, geopolitical risk, and their joint effects. The empirical framework combines multiple machine-learning algorithms with SHAP-based explainability to evaluate both forecast performance and the relative importance of predictors. Formal Diebold-Mariano tests are also used to assess whether the forecasting gains from augmented specifications are statistically significant. A Model Confidence Set analysis is further used to identify statistically superior model groups across the full set of algorithm-specification combinations. The results show that disruption-related predictors contain asset-specific forecasting information, while the comparison across algorithms indicates that no single model uniformly dominates across all assets and loss functions. The forecasting gains from disruption-related predictors, however, are strongly asset-specific and statistically uneven. For aluminum returns, augmented specifications that include supply-chain stress and/or geopolitical risk significantly improve forecast accuracy relative to the baseline. For copper returns, the evidence is weaker and mainly associated with geopolitical risk. For nickel returns, the joint inclusion of supply-chain stress and geopolitical risk provides the greatest improvement. By contrast, clean energy returns remain more closely tied to conventional macro-financial conditions, with no statistically significant incremental gains from disruption-related variables. SHAP evidence further indicates that predictor importance is asset-specific rather than dominated by a single market factor across all assets. The findings highlight the importance of combining flexible forecasting methods with economically interpretable tools when evaluating disruption-sensitive commodity and clean energy markets. Full article
Show Figures

Figure 1

19 pages, 793 KB  
Article
Enhancing Agricultural Decision-Making: Banana Yield Forecasting in Colombia Using Tuned Ensemble Machine Learning Models
by María I. Arrieta-Escobar, Carlos D. Paternina-Arboleda, Jorge I. Vélez and Guisselle A. García-Llinás
AgriEngineering 2026, 8(7), 289; https://doi.org/10.3390/agriengineering8070289 - 14 Jul 2026
Viewed by 339
Abstract
Accurate short-term forecasts of banana productivity can improve harvest scheduling, packing-capacity allocation, and export logistics, yet commercial forecasts often deviate substantially from realized yields. Objective: We evaluated whether tuned ensemble machine learning could reduce that error for seven-week-ahead forecasting of weekly productivity [...] Read more.
Accurate short-term forecasts of banana productivity can improve harvest scheduling, packing-capacity allocation, and export logistics, yet commercial forecasts often deviate substantially from realized yields. Objective: We evaluated whether tuned ensemble machine learning could reduce that error for seven-week-ahead forecasting of weekly productivity (boxes/Ha). Methods: Using weekly records from nine commercial farms in northern Colombia (2007–2018), we benchmarked six tuned ensembles—Random Forest, Extra Trees, Gradient Boosting, LightGBM, XGBoost, and CatBoost—against a company forecast and three statistical baselines (Seasonal Naive, Historical Mean, and a per-farm ARIMA). Hyperparameters were tuned by time-series cross-validation on 2007–2017, and models assessed on a 2018 holdout (405 observations). Results: CatBoost obtained the lowest errors (MAE = 3.95, R2 = 0.61), a 57.0% MAE and 81.2% MSE reduction versus the company baseline (MAE = 9.18, R2 = −1.06). Bootstrap 95% CIs and the Diebold–Mariano test showed CatBoost, XGBoost, LightGBM, Gradient Boosting, Random Forest, and the Historical Mean to be statistically equivalent (CatBoost’s lead was not significant), all outperforming the ARIMA (MAE = 5.64), Seasonal Naive (MAE = 7.31), and company baselines; an ablation confirmed no data leakage. Conclusions: Tuned ensembles can substantially improve short-term harvest planning in commercial banana production, with model choice guided by operational and computational constraints. Full article
Show Figures

Graphical abstract

25 pages, 762 KB  
Article
Mode-Dependent Performance of ARIMA, SARIMA, Holt-Winters, and Prophet in Electricity Load Forecasting: A Held-Out Comparison in Jordan
by Ahmed BaniMustafa, Zakaria Al-Omari, Nour Khlaifat and Mike Haddad
Energies 2026, 19(14), 3257; https://doi.org/10.3390/en19143257 - 10 Jul 2026
Viewed by 426
Abstract
Accurate forecasting of electricity demand is important for generation and distribution planning, especially in Jordan, where distinct seasons, variable weather, a fast-growing economy, a rising population, and shifting consumption habits affect demand. This paper applies four univariate time-series methods (ARIMA, SARIMA, Holt-Winters triple [...] Read more.
Accurate forecasting of electricity demand is important for generation and distribution planning, especially in Jordan, where distinct seasons, variable weather, a fast-growing economy, a rising population, and shifting consumption habits affect demand. This paper applies four univariate time-series methods (ARIMA, SARIMA, Holt-Winters triple exponential smoothing, and Prophet) to one year of aggregated national daily electricity consumption data collected by the Jordanian Electric Power Company (JEPCO). The forecasting models are trained and then tested on two horizons that span held-out periods of three months and six months. Fitting is performed both with dynamic rolling multi-step forecasts across the entire held-out period and with one-step-ahead out-of-sample forecasts whose parameters are frozen at their training-window values. All models are compared with a weekly seasonal-naive forecast, and pairwise accuracy differences are assessed with Diebold–Mariano tests. The main finding in this work shows that model performance is dependent on evaluation mode. In one-step-ahead mode, Holt-Winters is the most accurate model on both horizons (MAPE of 1.90% and 2.27%; R20.91); weekly SARIMA is statistically indistinguishable from it at three months, whereas ARIMA and Prophet are significantly less accurate. In dynamic multi-step mode, accuracy degrades sharply across all models (MAPE of 7.7–16.8%); no model attains a positive held-out R2; and none significantly outperforms the seasonal-naive benchmark. Exponential smoothing with weekly seasonality is therefore the strongest univariate choice for short-lead operational prediction, whereas reliable medium-term multi-step forecasting requires exogenous information. Full article
(This article belongs to the Special Issue New Progress in Electricity Demand Forecasting—2nd Edition)
Show Figures

Figure 1

40 pages, 12219 KB  
Article
Integrating Explainability into an Adaptive Transfer Learning with Uncertainty Quantification for PM2.5 Prediction in the Data-Scarce Region of South Africa
by Israel Edem Agbehadji and Ibidun Christiana Obagbuwa
Forecasting 2026, 8(4), 57; https://doi.org/10.3390/forecast8040057 - 4 Jul 2026
Viewed by 527
Abstract
South Africa faces significant challenges in monitoring air pollution from different provinces due to the sparse nature of the sensor network and heterogeneous pollutant sources. Notably, some provinces continue to record a limited amount of data on air pollution, thus making monitoring in [...] Read more.
South Africa faces significant challenges in monitoring air pollution from different provinces due to the sparse nature of the sensor network and heterogeneous pollutant sources. Notably, some provinces continue to record a limited amount of data on air pollution, thus making monitoring in those locations problematic. Fortunately, the capabilities of deep learning models to facilitate effective monitoring in data-scarce locations have been highlighted by researchers; however, these models within the context of transfer learning still lack transparency and uncertainty quantification. Using air pollutants and meteorological factors, this study proposes a transfer learning model for particulate matter (PM2.5) prediction in a data-scarce region. This transfer learning (TL) model leverages an adaptive Bi-directional Gated Recurrent Unit (adaBiGRU) with explainable artificial intelligence (xAI) and uncertainty quantification (UQ) to provide a novel uncertainty-aware adaptation transfer learning (UATL_adaBiGRU) model for a data-scarce location. Variant models based on the adaBiGRU technique, such as the temporal convolution network adaBiGRU (TCN-adaBiGRU) and domain-adversarial neural network adaBiGRU (DANNadaBiGRU), are presented as comparative models. The performance evaluation metrics are root mean squared, R2 score and mean squared error. The R2 score of pre-trained models in source domain is adaBiGRU (0.888), DANN_adaBiGRU (0.7788) and TCN_adaBiGRU (0.876). Furthermore, other comparative TL models include GRU (0.898), MLP (0.802) and adaptive LSTM (0.886). Afterwards, the pre-trained baseline model (adaBiGRU) was fine-tuned in the target domain dataset and the unpromising result contributed to the proposition of the UATL_adaBiGRU model for a data-scarce location, with R2 score of 0.9618. Uncertainty assessment metrics results were also presented for the proposed model. Ablation assessment demonstrates that each component of the UATL_adaBiGRU contributes to enhancing the predictive performance. Again, the Diebold–Mariano (DM) test statistic demonstrates a statistically significant difference between baseline model and UATL_adaBiGRU model. Finally, the local interpretable model-agnostic explanation highlights multi-scaled features as contributing towards the prediction of PM2.5 in the target domain. In view of this result, model fine-tuning is strongly recommended to enhance the robustness of the proposed uncertainty-aware adaption model in data-limited regions in South Africa. Full article
Show Figures

Figure 1

22 pages, 1941 KB  
Article
Lightweight Graph Embedding Augmentation for Airport Traffic Forecasting
by Ahmed Alharbi
Electronics 2026, 15(13), 2923; https://doi.org/10.3390/electronics15132923 - 3 Jul 2026
Cited by 1 | Viewed by 218
Abstract
Short-term airport traffic forecasting faces a structural gap: temporal ensemble models ignore route-network dependencies that shape hub operations, while deep graph neural networks require synchronised multi-airport operational data streams unavailable to single-airport operators, who have access only to their own operational records and [...] Read more.
Short-term airport traffic forecasting faces a structural gap: temporal ensemble models ignore route-network dependencies that shape hub operations, while deep graph neural networks require synchronised multi-airport operational data streams unavailable to single-airport operators, who have access only to their own operational records and publicly available route topology. To our knowledge, this study provides the first systematic evaluation of three graph representation classes—centrality measures, DeepWalk, and Node2Vec—as structural augmentations to Random Forest (RF), XGBoost, and LightGBM for hourly aircraft movement prediction at King Khalid International Airport (RUH), using a two-hop aviation graph combining RUH operational data with the OpenFlights database. Across all three ensemble families, random-walk graph augmentations consistently reduce MAE by approximately 9–17% relative to temporal-only baselines, whereas handcrafted centrality measures provide smaller and less consistent gains. Diebold–Mariano tests confirm that both RF+DeepWalk and RF+Node2Vec significantly outperform all nine baseline models (p<0.05), while no statistically significant difference is observed between the two embedding methods within any ensemble family, indicating that the benefit arises from the class of random-walk representations rather than a specific algorithm. RF+DeepWalk achieves the lowest observed MAE of 1.810 (RMSE = 2.481, sMAPE = 6.17%). SHAP analysis indicates that graph embedding dimensions rank among the top predictors, suggesting that they capture structural signal absent from temporal features. Full article
(This article belongs to the Special Issue AI Innovations in Smart Transportation)
Show Figures

Figure 1

15 pages, 750 KB  
Proceeding Paper
Enhancing Bitcoin Price Forecasting Through Integrated Sentiment Analysis and XGBoost Models
by Vasileios Dellopoulos, Ioannis Antoniadis, Evanggelos Saprikis and George Fragulis
Eng. Proc. 2026, 143(1), 31; https://doi.org/10.3390/engproc2026143031 - 2 Jul 2026
Viewed by 633
Abstract
This study investigates Bitcoin price forecasting using integrated sentiment analysis and gradient boosting within digital financial ecosystems. Two XGBoost models were developed using sentiment scores derived from Bitcoin news (2021–2024) and technical indicators, including GARCH-estimated volatility, Bollinger Bands, MACD, and RSI. The analysis [...] Read more.
This study investigates Bitcoin price forecasting using integrated sentiment analysis and gradient boosting within digital financial ecosystems. Two XGBoost models were developed using sentiment scores derived from Bitcoin news (2021–2024) and technical indicators, including GARCH-estimated volatility, Bollinger Bands, MACD, and RSI. The analysis uses 1042 daily Bitcoin observations and 10,025 sentiment records. Two model configurations were evaluated: one using only technical indicators and another incorporating daily aggregated sentiment scores. Model performance was assessed using Diebold–Mariano tests with Newey–West HAC variance estimation and walk-forward validation across 40 rolling windows. Contrary to expectations, sentiment features provided no statistically significant improvement over the technical-only model (p = 0.4888). Both models achieved identical test performance (R2 = −0.16%). Walk-forward validation revealed substantial temporal instability (Mean R2 = −126.30%, Std = 233.05%), highlighting the challenges of forecasting daily Bitcoin returns. Nevertheless, both XGBoost models significantly outperformed the random walk benchmark (DM statistic = −8.58, p < 0.0001), indicating that technical indicators capture exploitable market structure despite limited predictive accuracy for practical trading. These findings support the efficient market hypothesis and have implications for digital financial ecosystems integrating multimodal information. Full article
Show Figures

Figure 1

Back to TopTop