Next Article in Journal
Knowledge-Injected Transformer (KIT): A Modular Encoder–Decoder Architecture for Efficient Knowledge Integration and Reliable Question Answering
Previous Article in Journal
A Dual-Branch Ensemble Learning Method for Industrial Anomaly Detection: Fusion and Optimization of Scattering and PCA Features
Previous Article in Special Issue
A Distributed Instance Selection Algorithm Based on Cognitive Reasoning for Regression Tasks
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Attention-Based Deep Learning Hybrid Model for Cash Crop Price Forecasting: Evidence from Global Futures Markets with Implications for West Africa

1
School of Economics and Management, University of Electronic Science and Technology of China, No. 2006, Xiyuan Ave, West Hi-Tech Zone, Chengdu 611731, China
2
Center for West African Studies, University of Electronic Science and Technology of China, No. 2006, Xiyuan Ave, West Hi-Tech Zone, Chengdu 611731, China
3
School of Logistics, Sichuan University of Culture and Arts, No. 32, Yandong Road, Economic and Technological Development Zone, Mianyang 621000, China
4
School of Public Administration, University of Electronic Science and Technology of China, No. 2006, Xiyuan Ave, West Hi-Tech Zone, Chengdu 611731, China
*
Authors to whom correspondence should be addressed.
Appl. Sci. 2026, 16(3), 1600; https://doi.org/10.3390/app16031600
Submission received: 26 December 2025 / Revised: 30 January 2026 / Accepted: 2 February 2026 / Published: 5 February 2026
(This article belongs to the Special Issue Big Data Driven Machine Learning and Deep Learning)

Abstract

Accurate forecasting of agricultural commodity prices is essential for managing market volatility, improving supply chain coordination, and supporting food security-related decision-making. Recent advances in deep learning have demonstrated strong potential for capturing nonlinear and temporal dependencies in commodity price dynamics. In this study, we propose a hybrid long short-term memory–multi-head attention (LSTM–MHA) framework for agricultural commodity price forecasting using global futures market data. The model is trained and evaluated on multivariate global commodity futures prices, reflecting internationally traded benchmark markets rather than region-specific domestic prices. While the empirical analysis is based on global data, the study is motivated by the relevance of international price movements for import-dependent regions, particularly West Africa, where global price transmission plays a critical role in domestic market dynamics. The experimental results demonstrate that the proposed model effectively captures short-term temporal dependencies and provides interpretable attention-based insights into lag relevance. An ablation study further highlights the trade-offs between forecasting accuracy and interpretability across different model configurations. The hybrid architecture combines the time-based pattern identification and weighting capabilities of multi-head attention with the sequential learning capabilities of LSTM. Mean absolute error (MAE), root mean squared error (RMSE), and mean squared error (MSE) were used to evaluate the model’s performance. With an MSE of 0.0124, an RMSE of 0.1114, and an MAE of 0.1097, the model outperformed conventional models like ARIMA and standalone LSTM by three to four times in error reduction. The findings suggest that attention-enhanced deep learning models can serve as valuable analytical tools for understanding global price dynamics and informing policy analysis and risk management in West African agricultural markets.

1. Introduction

Agricultural commodity price volatility poses a significant challenge to global food security, trade stability, and agricultural supply chain resilience. For farmers, inaccurate price forecasts can result in suboptimal planting and marketing decisions, directly affecting incomes and livelihood stability. For traders and exporters, forecast errors translate into contract risks and inventory mismanagement, while, for policymakers, unreliable projections may lead to poorly timed interventions and ineffective stabilization measures. These risks are particularly acute in regions such as West Africa, where national economies and household incomes are highly dependent on a narrow set of globally traded cash crops, including cocoa, coffee, and cotton [1]. Consumers of agricultural commodities are increasingly aware of the quality and freshness of their food products to enhance their quality of life and health [2,3,4].
Consequently, producers are endeavoring to optimize their revenues through appropriate pricing of their products, so commodity price forecasting is crucial in the agricultural supply chain. In make-to-order manufacturing workshops, accurately predicting production progress (PP) is a crucial reference index for dynamically optimizing the production process and ensuring on-time delivery of production orders [5]. It aids in decision-making processes related to production planning, inventory management, risk mitigation, and policy formulation [6]. Moreover, regarding increasing global food security concerns and climate change impacts, reliable price forecasts become even more critical [7]. Predicting the prices of agricultural commodities has become more important for keeping the agricultural supply chain strong. This is especially true as traditional forecasting methods increasingly fail to account for farm markets’ complexity and nonlinear dynamics.
This study proposes a hybrid LSTM–multi-head attention (LSTM–MHA) architecture for multivariate multi-commodity price forecasting. Unlike existing hybrid models that combine LSTM with convolutional or statistical components, the proposed approach integrates multi-head attention directly into the temporal learning pipeline, enabling the model to capture both long-range dependencies and dynamic relevance patterns across time. This design enhances both predictive accuracy and interpretability.
Empirical analysis is conducted using global agricultural commodity futures prices, which serve as benchmark prices for internationally traded cash crops. West Africa is treated as an application context, reflecting the region’s strong exposure to global price movements, rather than as a domestic price market. This positioning allows the study to draw policy-relevant insights while maintaining methodological consistency. The forecasting of agricultural commodity prices constitutes a significant aspect of global economics, carrying extensive implications for trade, policy formulation, and investment strategies [8]. The inherent volatility and complexity of commodity markets, particularly in futures trading, pose significant challenges to traditional forecasting methods [9].
Recent advances in deep learning offer promising alternatives. Long short-term memory (LSTM) networks are well suited to modeling sequential data and capturing long-term temporal dependencies, while attention mechanisms enhance interpretability and allow models to selectively emphasize informative historical periods. However, many existing studies apply these methods in isolation or within narrowly scoped hybrid frameworks, often focusing on single commodities, limited feature sets, or short forecasting horizons.
In contrast, multi-head attention mechanisms are adept at identifying relevant patterns across diverse time scales. By integrating these methodologies, our hybrid model aims to leverage the strengths of both frameworks to enhance the accuracy of forecasting and reliability. The crucial importance of the agricultural sector in ensuring global food security and maintaining economic stability underscores the need for precise and dependable price forecasts [10]. These commodities satisfy direct consumption requirements and act as raw materials for diverse businesses. Price variations can produce cascading impacts, influencing stakeholders ranging from small-scale farmers to multinational enterprises and national economies [11].
Developing more accurate forecasting models is crucial for all the stakeholders in the agricultural value chain [6]. It is assumed that the resilience of the agricultural supply chain largely depends on stakeholders’ ability to anticipate and respond to price fluctuations. Accurate pricing forecasts enhance inventory management, optimize resource allocation, and strengthen risk management techniques. However, contemporary forecasting methods frequently inadequately account for the intricate interrelations among diverse market components, including seasonal patterns, climatic variables, global trade influences, and unexpected market disruptions.
This limitation has created a critical need for more robust forecasting approaches that enhance agricultural supply chain resilience. Advanced forecasting models could improve prediction accuracy and agricultural supply chain performance, particularly those incorporating artificial intelligence and machine learning techniques. This study seeks to leverage deep learning and attention mechanisms to identify intricate temporal correlations and pertinent aspects in commodity price data.
Forecasting agricultural commodity prices has progressed from basic statistical methods to advanced machine learning techniques [12]. Traditional statistical models like ARIMA and GARCH have long been used for time-series forecasting. However, they often fail to capture the nonlinear, long-term, and dynamic dependency characteristic of agricultural price movements [13]. In recent years, deep learning models, specifically recurrent neural networks (RNNs) and long short-term memory (LSTM) networks, have demonstrated efficacy in time-series forecasting by effectively capturing long-term dependencies in sequential data [14,15]. While LSTMs can be useful, they may not fully capture all the complexities involved in commodity price dynamics [16]. This is where the multi-head attention mechanism is utilized. Initially, it was developed for natural language processing. Attention mechanisms have demonstrated efficacy in numerous time-series forecasting challenges by allowing the model to focus on specific parts of the input sequence during prediction [17].
Supply chain resilience (SCR) focuses on mitigating risks like climate variability, pest outbreaks, and trade disruptions through tools such as blockchain for enhancing transparency, the Internet of Things (IoT) for real-time monitoring, and public–private collaborations to improve infrastructure and systems. During events such as the COVID-19 pandemic, localized agricultural supply chains proved essential for continuity, underscoring the need for adaptability and climate-smart agricultural practices [18,19]. Price forecasting complements supply chain management by reducing market uncertainty and enabling better planning. Advanced analytics, such as machine learning, utilize historical and real-time data to improve prediction accuracy, aiding inventory and production optimization. For example, forecasting can guide decisions on planting schedules or dynamic pricing, ensuring alignment between supply and market demand [20].
The study examines a diverse range of agricultural commodities, similar to those discussed in some past studies. For instance, Cocoa and coffee are tropical crops that are sensitive to climate conditions and global demand patterns [21]. Cotton is a significant non-food agricultural commodity that is affected by both agricultural practices and industrial factors [22]. Lumber, reflecting the forestry sector, is characterized by distinctive long-term growth cycles [23]. Orange juice is a commodity that exhibits significant seasonal patterns and is vulnerable to weather conditions. Sugar is a widely produced commodity that is influenced by the food and biofuel industries [24].
Integrating LSTM and multi-head attention in a hybrid model offers a new approach to forecasting agricultural commodities [25]. This integration aims to leverage LSTM’s ability to capture long-term dependencies and the attention mechanism’s capacity to concentrate on the important sections of the input sequence [26]. This research is especially pertinent due to the rising volatility in global agricultural markets caused by climate change, evolving consumer tastes, and geopolitical tensions. Enhanced forecasting models can serve as essential instruments for risk management, policy development, and strategic planning within the agriculture industry [1]. Furthermore, this study’s findings may have wider applicability in financial forecasting and time-series analysis across other sectors.
Despite advancements, the existing models for agricultural commodity forecasting often suffer from limitations, such as oversimplified feature sets, a single-commodity focus, or inadequate handling of volatility and seasonality. Many deep learning approaches remain shallow, lack proper regularization, or fail to integrate both long-term trends and short-term fluctuations effectively. Moreover, agricultural markets exhibit unique characteristics, such as high sensitivity to weather, seasonality (orange juice and sugar), and supply chain disruptions that demand more sophisticated, generalizable, and robust models. This study is motivated by the urgent need for a unified, accurate, and scalable forecasting framework that can handle the complexity and diversity of agricultural commodity prices. By combining the sequential modeling strength of LSTM with the dynamic pattern-weighting capability of multi-head attention, we aim to develop a hybrid model that improves predictive accuracy and supports decision-making across the agricultural supply chain.
This research makes significant contributions to the field of agricultural commodity price forecasting through several key advancements. First, it proposes an innovative hybrid model architecture that seamlessly integrates the sequential learning capabilities of long short-term memory (LSTM) networks with the pattern recognition strengths of multi-head attention mechanisms, enabling the model to effectively capture both long-term dependencies and salient short-term fluctuations in price data. Second, the model demonstrates strong generalization performance across a diverse set of agricultural commodities, including cocoa, coffee, cotton, lumber, orange juice, and sugar, indicating its adaptability and applicability to various agricultural markets, thereby providing valuable insights into the effectiveness of deep learning for multi-commodity forecasting. Third, the integration of attention mechanisms allows the model to dynamically identify and weigh critical temporal patterns, such as seasonality and volatility shifts, significantly improving prediction accuracy under complex and volatile market conditions. Finally, the study establishes a practical framework for integrating advanced forecasting models into agricultural supply chain resilience strategies, offering stakeholders actionable tools for optimizing production planning, enhancing resource allocation, and strengthening risk management [27]. Improved price forecasts could lead to better decision-making, more efficient resource allocation, and enhanced risk management strategies [28,29].
The rest of the paper is structured as follows: Section 2 illustrates the related works. Section 3 presents the proposed LSTM–multi-head attention hybrid model. Section 4 shows the results of the proposed model in terms of performance and a comparative analysis with the baseline models. Section 5 concludes the paper.

2. Related Works

This section offers a thorough examination of the current literature concerning agricultural commodity price forecasting, with a particular emphasis on the utilization of machine and deep learning methodologies. Forecasting agricultural food prices is crucial for the academic community and policymakers as it contributes to food security. A substantial amount of research has been conducted in this domain, resulting in the development of numerous models aimed at enhancing prediction accuracy. Vector autoregression (VAR) is a conventional technique employed for time-series forecasting challenges. The VAR model must be evaluated against the rate of an alternative or baseline product, which is why much of the pertinent work in food price forecasting has investigated the correlation between global crude oil prices and food prices. Researchers have recently identified a nonlinear causal link between food prices and crude oil prices [30]. Another study determined that there is a relationship between biofuel prices and food prices, both in the short- and long term. Historically, the forecasting of agricultural commodity prices has depended on several statistical and economic models. Time-series analysis methods, including autoregressive integrated moving averages (ARIMA) and variations, are extensively employed in this field. For instance, Kohzadi et al. [31] compared ARIMA models with artificial neural networks for forecasting commodity prices, finding that neural networks outperformed ARIMA in most cases. Another popular approach is the generalized autoregressive conditional heteroscedasticity (GARCH) model, particularly developed for capturing volatility clustering in commodity prices. The GARCH model has been applied to various metal commodities, demonstrating its effectiveness in modeling price volatility [32].
The emergence of machine learning has introduced novel opportunities in agricultural price forecasting. Support vector machines (SVMs), artificial neural networks (ANNs), and random forests have been utilized with differing levels of efficacy. Xiong et al. [33] used SVMs to forecast agricultural commodity prices, showing improved accuracy compared to traditional methods. Their research on maize price predictions illustrated the capability of machine learning to identify nonlinear correlations in pricing data. Ensemble approaches, especially random forests, have demonstrated potential in this domain. Sujjaviriyasup and Pitiruek [34] applied random forests to predict agricultural commodity prices in Thailand, attaining superior accuracy compared to individual decision tree models.
Deep learning models, especially recurrent neural networks (RNNs) and their variants, have gained significant attention in time-series forecasting due to their ability to capture long-term dependencies. Hochreiter and Schmidhuber [35] introduced LSTM networks; these networks have gained popularity in financial and commodity price forecasting. They are designed to address the vanishing gradient problem seen in traditional RNNs, which makes them better suited for capturing long-term dependencies in time-series data. Attention mechanisms were first introduced in the field of natural language processing by Bahdanau et al. [36] and have recently been applied to time-series forecasting tasks. The self-attention mechanism, particularly as implemented in the transformer architecture by Vaswani et al. [17], has shown remarkable results in various sequence modeling tasks. In financial forecasting, Zhou et al. [26] proposed the informer model, which uses a ProbSparse self-attention mechanism for long-sequence time-series forecasting. Their approach showed efficiency and effectiveness in managing long-range dependencies in time-series data.
Recent studies have concentrated on hybrid models that integrate various techniques to maximize their strengths. For instance, Kristjanpoller and Minutolo [37] proposed a hybrid model that combines artificial neural networks and GARCH for forecasting the volatility of gold prices, demonstrating improved performance compared to individual models in the context of agricultural commodities. Also, Kantasa-ard et al. [38] developed a hybrid LSTM with a genetic algorithm and scatter search to automate the hyperparameter tuning of the hybrid LSTM for demand forecasting in a physical internet supply chain network. In the context of commodity markets, Cen and Wang [39] applied an LSTM network to forecast crude oil prices, showing enhanced performance over traditional time-series models and standard neural networks. Similarly, Livieris et al. [40] utilized LSTM networks in forecasting agricultural commodity prices, demonstrating enhanced accuracy compared to traditional machine learning techniques. Another study by Liu et al. [5] designed a long short-term memory (LSTM) model with transfer learning (TL) to accommodate the nonlinear relationships of the features supplied by the CNN–TL model for predicting production progress.
Olofintuyi et al. [41] applied many machine learning methodologies, including LSTM networks, to predict cocoa prices, illustrating the efficacy of deep learning in encapsulating the intricate dynamics of the cocoa market. For anticipating coffee prices, Mekala et al. introduced a novel approach to predicting coffee prices using BiLSTM (bidirectional long short-term memory) and CNN (convolutional neural network) models. Their method showed how effectively combining CNN and BiLSTM models predicts coffee prices [42]. Dave et al. developed an integrated ARIMA and LSTM machine learning model in the cotton market to accurately predict Indonesia’s future exports. Additionally, they compared various forecasting methods for predicting cotton prices. Their study emphasized the benefits of hybrid approaches [43]. For lumber price forecasting, Lamichhane et al. [44] used machine learning approaches, illustrating their efficacy in encapsulating the distinct attributes of the lumber market, and their research established a foundation for forecasting wood prices in the southern timber market with ANN models. Luo et al. [45] applied deep learning models to predict orange juice futures prices, demonstrating better performance than traditional time-series models. In the sugar market, Ribeiro and Oliveira [46] applied neural networks and neuro-fuzzy systems for price forecasting, proposing a hybrid model for predicting agricultural commodity prices, specifically sugar, illustrating the efficiency of these techniques in capturing the intricate dynamics of sugar prices.
The literature review indicates a distinct trend in the utilization of sophisticated machine learning and deep learning methodologies for forecasting agricultural commodity prices. Although conventional methods remain relevant, there is increasing evidence of the enhanced efficacy of deep learning models, especially LSTM networks and attention-based techniques. The effectiveness of hybrid models that integrate various methodologies is also apparent. However, there is a gap in the reviewed literature on the application and effective integration of LSTM–multi-head attention hybrid models for the diverse agricultural commodities analyzed in this study, highlighting the novelty and potential significance of the proposed research.
Our design tackles the limitations highlighted in several related works, as outlined in Table 1. We compared our proposed LSTM–multi-head attention model with other notable deep learning models to offer a solution for accurate agriculture commodity price predictions.
Table 1 highlights that most existing studies focus on single-commodity forecasting and rely on univariate input structures without explicit attention mechanisms. In addition, several studies report limited evaluation protocols, often omitting validation sets or comprehensive error metrics. In contrast, the proposed LSTM–MHA framework incorporates multivariate inputs, explicitly models temporal relevance through multi-head attention, and evaluates performance using multiple error measures across training, validation, and test sets.

3. Methodology

The methodological framework consists of four stages: data acquisition, preprocessing, model architecture design, training and validation, and performance evaluation. Figure 1 illustrates the overall modeling pipeline.

3.1. Data Collection and Preprocessing

The study employs daily futures price data for six globally traded agricultural commodities—cocoa, coffee, cotton, lumber, orange juice, and sugar—covering the period January 2000 to December 2023. These futures prices represent internationally recognized benchmarks that directly influence export revenues and income stability in West African economies.
The information comprises open, high, low, and closed prices, as well as trading volume. The research employs agricultural commodities futures data [56] from Kaggle, which offers a comprehensive analytical dataset that serves as the basis for implementing the forecasting models examined in this study.
Missing observations arising from non-trading days were forward-filled to preserve temporal continuity. While forward filling may dampen short-term volatility, it ensures sequence completeness for recurrent models. All price variables were normalized using min–max scaling to enhance numerical stability during training; the implications of this choice are discussed in the Section 4.6.
The dataset was split chronologically into training (70%), validation (20%), and test (10%) sets to prevent information leakage and to reflect real-world forecasting conditions.

3.2. Model Architecture

The proposed LSTM–MHA architecture consists of stacked LSTM layers followed by a multi-head attention module, layer normalization, and dense layers. The number of LSTM units and attention heads was selected based on empirical performance stability and computational efficiency. Multi-head attention allows the model to learn diverse temporal relevance patterns, improving robustness across commodities with differing seasonal and volatility characteristics. The proposed architecture of the model is illustrated in Figure 1.
The model was trained using the Adam optimizer with mean squared error as the loss function. Hyperparameters were tuned through controlled experiments, with the final configuration selected based on validation performance stability rather than marginal metric improvements. Training was conducted on a GPU-enabled environment, and consistent random seeds were used to ensure reproducibility.
In the LSTM depicted in Figure 1 information traverses a mechanism regulated by many gates that determine whether to retain or eliminate specific information at each time step. Figure 2 depicts the single-layer structure of the LSTM architecture within the sequence modeling.
The model architecture and procedures for the experiment are illustrated in Algorithm A1.
Algorithm A1: Hybrid LSTM–Multi-Head Attention (LSTM–MHA) Forecasting Framework
Input:  Multivariate commodity futures series { x t } t = 1 T ; lookback window L; epochs E; batch size B; learning rate η ; number of attention heads H.
Output:  Forecasted price y ^ t + 1 ; trained parameters Θ ; attention weights A .
Step 1. Data Preparation:
Normalize inputs using min–max scaling and construct supervised sequences X t = [ x t L , , x t 1 ] . Split data chronologically into training, validation, and test sets.
Step 2. Temporal Feature Extraction:
Encode X t using an LSTM layer to obtain hidden representations H t .
Step 3. Multi-Head Attention:
Apply multi-head self-attention to H t to compute attention-weighted features A t .
Step 4. Feature Refinement:
Combine H t and A t via residual connection and layer normalization to produce Z t .
Step 5. Forecast Generation:
Refine temporal features using a second LSTM layer and generate prediction y ^ t + 1 through a fully connected output layer.
Step 6. Model Training:
Optimize parameters Θ using the Adam optimizer with mean squared error loss and validation-based early stopping.
Step 7. Evaluation and Interpretability:
Evaluate forecasting performance using MSE, RMSE, and MAE. Conduct ablation analysis (LSTM-only and attention-only). Extract attention weights A to compute lag importance.
Step 8: Ablation and Attention Analysis:
Train baseline models (LSTM-only and attention-only).
Extract attention weights A and compute lag importance:
α l = E b , h , t A b , h , t , l
return  y ^ t + 1 , Θ , A
The model architecture comprises the following components:
Input Layer: The input layer acquires sequential data. It is structured as a sequence of lengths and features. The sequence length represents the number of time steps, while features represent the number of variables at each time step.
LSTM Layer 1 (64 units): This represents the first long short-term memory (LSTM) layer. LSTM is a type of recurrent neural network (RNN) that can learn long-term dependencies. It is advantageous for time-series data. This layer consists of 64 units, determining its capacity to identify patterns in the data. The LSTM unit comprises three gates: the input gate, the forget gate, and the output gate. LSTM equations for the various gates include:
Input Gate: The value i t of the corresponding input gate is aggregated to determine the level of updated information, which outputs a value between (0 and 1).
i t = σ ( W i · [ h t 1 , x t ] + b i )
Forget Gate: This variable f t represents the output of the associated forgetting gate, which determines whether to eliminate or retain specific information from the storage unit. h t 1 represents the output value from the preceding moment. x t represents the present input value.
f t = σ ( W f · [ h t 1 , x t ] + b f )
Output Gate: The output gate value o t is calculated to ascertain the extent of memory utilized for output.
o t = σ ( W o · [ h t 1 , x t ] + b o )
Cell State: The cell state is modified by multiplying the previous cell state C t 1 by the output value f t from the associated forgetting gate. The updated output value is subsequently multiplied by the input gate i t , incorporating the second state into the new cell state C t .
C ˜ t = tanh ( W C · [ h t 1 , x t ] + b C )
C t = f t * C t 1 + i t * C ˜ t
Hidden State: The next step is to employ the hyperbolic tangent function to adjust the C t value to fall within the range of 1 and −1. The output gate value o t is ultimately multiplied by i t , yielding the final output value h t at time t.
h t = o t * tanh ( C t )
Here, σ denotes the sigmoid or activation function, * represents element-wise multiplication, W signifies weight matrices, and b indicates bias vectors.
Dropout Layer 1 (20%): Dropout is a regularization method employed to mitigate overfitting. It randomly assigns a proportion of input units to zero during each training update. In this instance, 20% of the inputs are omitted, facilitating the model’s acquisition of more resilient features. The dropout layer is computed as follows:
y = d * x
where d is a vector of independent Bernoulli random variables, each with a probability of 0.8 of being 1.
Multi-Head Attention Layer (8 heads): The attention mechanism, which allows the model to focus on different segments of the input sequence while producing predictions, is incorporated into this layer. The phrase “8 heads” describes the model’s capacity to simultaneously learn eight distinct attention patterns, which enables it to identify a range of dependencies in the data.
The aggregate of the multi-head attention layer is
M u l t i H e a d ( Q , K , V ) = C o n c a t ( { h e a d } 1 , . . . , { h e a d } h ) W O
Each head is calculated as follows:
h e a d i = A t t e n t i o n ( Q W i Q , K W i K , V W i V )
The function of attention is as follows:
A t t e n t i o n ( Q , K , V ) = s o f t m a x ( Q K T d k ) V
Layer Normalization: Layer normalization stabilizes the learning process by normalizing inputs across features instead of the batch dimension, hence accelerating training and enhancing the model’s overall performance.
y = x E [ x ] V a r [ x ] + ϵ * γ + β
where ϵ is a tiny constant for numerical stability and γ and β are learnable parameters.
LSTM Layer 2 (32 units): This is the second LSTM layer with 32 units. Multiple LSTM layers enable the model to learn more complex temporal patterns. This layer processes the output from the attention mechanism, refining the temporal features.
Dropout Layer 2 (20%): Another dropout layer for additional regularization. It helps to prevent overfitting on the patterns learned by the second LSTM layer 2 using y = d * x .
Dense Layer (16 units, ReLU activation): This layer is fully connected with 16 units and uses ReLU (Rectified Linear Unit) activation. It facilitates learning nonlinear combinations of the high-level information retrieved by prior layers. The ReLU activation function provides nonlinearity to the model. The dense layer is expressed as
y = { a c t i v a t i o n } ( W x + b )
where W is the weight matrix, x is the input vector, b is the bias vector, and activation is the activation function (ReLU in this case, except for the final layer, which has no activation).
Output Dense Layer (1 unit, linear activation): This is the final layer that generates the forecast. It consists of one unit with a linear activation function, which is suitable for regression tasks such as forecasting. The linear activation enables the model to output any numerical value, making it appropriate for prediction tasks.
y = W x + b

3.3. Training Process

We employed a 70-20-10 division in the training, validation and testing of datasets. Our proposed model underwent a training process for 50, 100, and 150 epochs, utilizing batch sizes of 64, 128, and 256, alongside learning rates of 0.001, 0.01, and 0.1, respectively, employing Adam optimizer and mean squared error as the loss function. Table 2 below presents the various model parameters adjusted in our studies.
Based on the comparative evaluation of the training configurations summarized in Table 2 and the performance results reported in Table 3, the final experimental results presented in this study are obtained using Configuration 2, corresponding to a training duration of 100 epochs. This configuration achieved the most favorable balance between forecasting accuracy and training stability, yielding consistently lower validation errors across all evaluation metrics without evidence of overfitting.
Although the 150-epoch configuration provided marginally longer training, it did not lead to further improvement in predictive performance and exhibited diminishing returns in validation accuracy. In contrast, the 50-epoch configuration converged more rapidly but resulted in comparatively slightly higher errors, especially during training. Consequently, the 100-epoch configuration was selected as the final training setting and was used uniformly for all reported forecasting results in this work.

4. Experimental Results and Analysis

4.1. Model Performance

The proposed model assessment was done using multiple metrics, including mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE), with the findings displayed alongside the hyperparameter tuning in Table 3. However, the training and validation losses in this experiment exhibit near-overlapping trajectories across epochs, as shown in Table 3, the hyperparameter analysis. While such behavior may appear unusual in some learning scenarios, it is expected in this setting due to three factors: (i) chronological data splitting, which preserves similar statistical regimes across training and validation sets; (ii) regularization effects induced by attention-based temporal averaging and dropout; and (iii) early stopping based on validation stability rather than aggressive loss minimization. This behavior indicates conservative learning and strong generalization rather than data leakage or overfitting.
Figure 3 illustrates the summary of the LSTM–MAH model architecture, which exemplifies an advanced hybrid methodology for forecasting agricultural commodity prices. The model initiates with an input layer that accommodates 30 time steps and seven features per step, encapsulating a month’s historical data. The design next processes this input via a dual LSTM–attention mechanism framework. The initial LSTM layer, with 64 units and 18,432 parameters, identifies temporal patterns within the input stream. The subsequent component is a vital multi-head attention method (265,280 parameters) that enables the model to dynamically concentrate on pertinent historical patterns, augmented by residual connections and layer normalization to ensure steady training. Another LSTM layer with 32 units (12,416 parameters) further analyzes this attention-weighted data. The model culminates in two dense layers that systematically decrease the dimension from 16 units to one unit, yielding the final price forecast. The model has 296,801 trainable parameters (about 1.13 MB), achieving an effective equilibrium between complexity and processing demands. This design integrates the LSTM’s proficiency in capturing long-term dependencies with the attention mechanism’s capacity to emphasize pertinent past patterns. It is especially adept at the intricate process of projecting agricultural commodity prices.
The performance study of the LSTM–MHA model at Epoch 100 reveals strong and stable learning properties, as indicated by the graphical representations in Figure 4 and the numerical metrics in Table 3. The model attained nearly equivalent training and validation losses of 0.03472800, alongside a test loss of 0.03464708, demonstrating exceptional generalization skills without evidence of overfitting. The learning curves exhibit a consistent and persistent convergence pattern, with training and validation measures closely aligning throughout the training phase. The convergence is additionally corroborated by the model’s error measures, which include an MSE of 0.01241407, an RMSE of 0.11141844, and an MAE of 0.10966918. The narrow disparity between training and validation performance, along with the consistent stabilization of loss curves near the 0.034 thresholds, indicates that the model effectively discerned the fundamental patterns in agricultural commodity price fluctuations while preserving robust generalization skills. The results confirm the efficacy of the hybrid LSTM–attention architecture in delivering dependable and consistent price forecasting performance.
Figure 5 presents the comparison between actual and predicted prices on the training dataset, comprising approximately 20,000 time steps. The predicted series closely follows the underlying trend of the true values while exhibiting reduced sensitivity to short-term fluctuations. This behavior indicates that the model effectively learns the dominant temporal structure of the data without overfitting to transient noise.
Although occasional deviations are visible during abrupt price movements, the overall alignment between the predicted and observed values is consistent with the low training error reported in Table 3. The relatively smooth prediction trajectory reflects an intentional regularization effect induced by the LSTM–attention architecture, which prioritizes stable pattern learning over memorization of isolated spikes. The significant difference between the actual and predicted values from Figure 5, Figure 6, Figure 7 and Figure 8 might be influenced by sudden shifts in global supply and demand, unforeseen changes in production levels, or market interventions; for example, the government can set price controls and trade restrictions on various commodities that were not anticipated by our model.
Figure 6 illustrates the model performance on the validation dataset and provides evidence of strong generalization ability. The predicted values maintain a stable range relative to the actual series, with the prediction line exhibiting lower volatility than the observed prices. Importantly, this smoothing behavior is accompanied by nearly identical training and validation losses (approximately 0.0347), indicating that the model does not suffer from overfitting.
From a quantitative perspective, the error magnitudes during validation remain consistent with the training errors, suggesting that the learned temporal representations are transferable to unseen data. This stability confirms that the hybrid LSTM–MHA architecture captures persistent market dynamics rather than dataset-specific noise.
Figure 7 shows the model’s performance on the test dataset (approximately 8000 time steps). The predictions continue to track the overall price trajectory while systematically underestimating sharp price spikes. This effect is reflected in the slightly higher RMSE relative to MAE, indicating that large deviations during high-volatility episodes contribute disproportionately to squared error.
Rather than signaling poor performance, this behavior highlights a conservative forecasting bias, where the model emphasizes robustness and stability over extreme value prediction. Such behavior is common in deep learning models trained with mean squared error loss and reflects a bias–variance trade-off that favors smoother forecasts.
Figure 8 provides a focused examination of the final 1000 observations, a period characterized by elevated market volatility. During this regime, the divergence between actual and predicted values becomes more pronounced, particularly at sharp upward or downward price movements.
Quantitatively, the error contributions in this segment are dominated by a small number of extreme observations rather than persistent misalignment. The model continues to capture the direction and medium-term trend of prices but attenuates the amplitude of sudden shocks. This confirms that the smoothing observed in Figure 5, Figure 6 and Figure 7 is not an artifact of aggregation but a structural characteristic of the model.
While this limits the model’s ability to predict rare price spikes, it enhances forecast stability and reduces sensitivity to noise, an important consideration for medium-term planning and risk management.
Figure 9 illustrates the LSTM–MHA model’s performance over the whole dataset, with the model (red line) constantly aligning with the overall trend of actual commodity prices (blue line). The predicted series consistently aligns with the long-run trend of actual prices, demonstrating strong baseline tracking capability. However, extreme price movements, particularly in later periods characterized by substantial volatility, are visibly dampened.
This global view reinforces the interpretation that the LSTM–MHA model is optimized for trend- and regime learning rather than precise tail-risk estimation. The attenuation of extreme values explains why the model achieves low average error metrics while still underestimating rare but economically significant price spikes.
Figure 10 provides significant insights into its prediction abilities. The majority of data points aggregate in the lower range (0.0–0.4) of both true and forecasted values, with a notable concentration between 0.0 and 0.2, signifying that most commodity price fluctuations transpire within this normalized range. The red dashed line denotes the optimal prediction line (y = x), and the model exhibits strong performance at lower values, with points closely aligned to this line. Nonetheless, it exhibits escalating divergence at elevated levels, frequently underestimating them, especially for actual values beyond 0.4. The distribution exhibits right skewness, indicating that high price spikes are infrequent. The model demonstrates accuracy within usual price ranges while adopting a cautious approach to predicting outlier events, hence ensuring reliable forecasts for standard trading scenarios.

4.2. Hyperparameter Analysis

In optimizing the hyperparameters of our proposed LSTM–multi-head attention (LSTM–MHA) model, we examined several configurations concerning epochs, learning rates, and batch sizes. The analysis focuses on evaluating the influence of these parameters on the model’s predictive accuracy using metrics such as mean squared error (MSE), root mean squared error (RMSE), and root mean absolute error (RMAE). The dataset utilized for this experiment consists of agricultural commodity futures data [56] from Kaggle.
We examine the distinct epochs outlined in Table 3, noting that Epoch 50 produced training and validation losses of roughly 0.0340, accompanied by a test loss nearing 0.034. The MSE, RMSE, and RMAE figures demonstrate reduced error rates, indicating that this epoch setting achieves a balance between underfitting and overfitting. At Epoch 100, the stability of training and validation losses is enhanced, with each approximately 0.0347. The MSE and RMSE values exhibited negligible fluctuation relative to Epoch 50, indicating the model’s stability over training rounds. Epoch 150 did not markedly enhance the loss values compared to Epoch 100, exhibiting only minor increases in both test and training loss parameters. The stability observed with increased epochs indicates that performance improvements peak at 100, suggesting declining returns on prediction accuracy with additional epoch increments.
Similarly, we examine the learning rates as outlined below: A learning rate of 0.001 ensured steady training and validation loss values across the epochs, allowing the model to converge well without substantial changes. A learning rate of 0.01 exhibited comparable stability patterns, although there were minor increases in validation loss values, suggesting the potential for modest overfitting as the learning rate escalated. The learning rate of 0.1 was less effective as the validation loss exhibited increased volatility, indicating that elevated learning rates may cause instability and impede optimal model convergence.
The batch sizes are examined: A batch size of 64 yielded the optimal balance between accuracy and computational efficiency, demonstrating consistent validation and training losses over the epochs. Batch size 128 yielded results comparable to batch size 64, exhibiting low variation in training and test loss metrics. Concurrently, a batch size of 256 exhibited a minor decline in performance as the batch size escalated, resulting in slightly elevated losses. Larger batch sizes may diminish model generalizability, potentially due to inadequate changes in the model weights.

4.3. Comparison with Baseline Models

The proposed LSTM–MHA model exhibits higher performance across all the assessment measures, indicating that the attention mechanism significantly increases the fundamental capabilities of the LSTM architecture in time-series forecasting. The proposed LSTM–MHA model exhibits the lowest errors across all the measures, with errors around 3–4 times lower than those of complicated models and markedly superior to hybrid techniques, as demonstrated in Table 4.
To ensure fair and transparent comparison with existing studies, the same set of evaluation metrics, mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE), were consistently employed throughout the comparative analysis. These metrics are the most widely reported performance measures in the agricultural commodity price forecasting literature and are explicitly provided in the majority of the benchmark studies included in Table 4. Using a common set of metrics allows direct numerical comparison without introducing additional sources of bias arising from metric-specific sensitivities.
Moreover, MSE and RMSE emphasize large deviations and are therefore sensitive to extreme forecasting errors, which is particularly relevant in volatile commodity markets, while MAE provides a complementary measure of average absolute deviation that is more robust to outliers. The combined use of these three metrics enables a balanced assessment of both average forecasting accuracy and error dispersion.
It is important to note that some benchmark studies report only a subset of these metrics, which is reflected by the missing values in Table 4. In such cases, we refrain from imputing or recomputing unreported metrics to avoid methodological inconsistency. Consequently, the comparison is restricted to the metrics explicitly reported in the sources, and missing entries are clearly indicated. This approach prioritizes transparency and reproducibility over exhaustive but potentially misleading metric harmonization.
However, the reported performance differences should be interpreted in terms of relative magnitude rather than absolute equivalence across datasets.

4.4. Ablation Study and Model Comparison

An ablation study was conducted to evaluate the individual and combined contributions of the LSTM and attention components. Three model variants were considered: a standalone LSTM model, an attention-only model, and the proposed hybrid LSTM–MHA architecture. Performance was evaluated using mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE).
Table 5 summarizes the forecasting errors across the three architectures. The standalone LSTM and attention-only models achieve nearly identical MSE values (0.0046), indicating comparable predictive accuracy under the current experimental setup. In contrast, the hybrid LSTM–MHA model exhibits a slightly higher MSE of 0.0054, along with marginally higher RMSE and MAE values.
These results suggest that the incorporation of multi-head attention does not automatically translate into lower average forecasting error, particularly in a low-dimensional price-driven setting. Instead, the hybrid architecture appears to trade a small degree of numerical accuracy for increased model expressiveness and interpretability. The higher RMSE observed for the hybrid model reflects a more conservative forecasting behavior, characterized by smoother predictions and reduced sensitivity to extreme price fluctuations.
This behavior aligns with the attention analysis and highlights an important bias–variance trade-off. While smoothing may lead to underestimation of sharp price spikes, it can also enhance forecast stability under noisy market conditions. Consequently, the hybrid model’s value lies not solely in minimizing point forecast error but in providing structural insights into temporal relevance, which can be particularly useful for policy analysis and medium-term planning.
The MSE in Figure 11 visually confirms the numerical results. The LSTM and attention-only models achieve nearly identical MSE values, while the hybrid LSTM–MHA model exhibits a modest increase in error.
Rather than indicating model inferiority, this result should be interpreted as evidence that attention mechanisms do not automatically improve forecasting accuracy when used as an architectural add-on, especially in relatively low-dimensional price-driven settings. Their primary contribution lies in interpretability and structural flexibility, which become more valuable when additional exogenous features (for instance, weather, exchange rates, or policy indicators) are incorporated.

4.5. Attention-Based Lag Importance Analysis

To enhance the interpretability of the proposed hybrid LSTM–multi-head attention (LSTM–MHA) model, we analyze the distribution of attention weights assigned to historical lags. Attention weights are aggregated across batches, attention heads, and target time steps, providing a global view of the temporal relevance learned by the model.
Figure 12 illustrates the absolute distribution of attention weights across historical lags, while Figure 13 highlights relative deviations from the mean attention weight.
Figure 12 illustrates the absolute lag importance profile. Although the numerical variation in attention weights is small due to normalization and aggregation across heads and time steps, a clear monotonic decay is observed. The first three to five lags receive systematically higher attention weights, indicating that the model prioritizes recent historical information when forming forecasts. Beyond approximately seven to ten lags, the curve flattens and converges toward a stable baseline, suggesting a diminishing marginal contribution of older observations within the 30-day lookback window.
This behavior is consistent with stylized facts of agricultural commodity markets, where short-term dynamics and recent shocks typically exert a stronger influence on price formation than distant historical information. The smooth decay pattern further indicates that the attention mechanism learns a structured temporal relevance profile rather than assigning weights arbitrarily.
To improve visual interpretability, Figure 13 presents the lag importance after centering the attention weights around their mean. This relative representation highlights deviations from the average attention level, making the contribution of individual lags more discernible. The centered profile confirms that early lags contribute positively above the mean attention weight, while later lags contribute negatively with small magnitudes. Meaningful deviations from the mean are concentrated within the first few lags, reinforcing the conclusion that the model’s effective memory is largely short-term.
Taken together, the two panels demonstrate that the multi-head attention mechanism primarily serves as an interpretability tool, revealing how the model selectively emphasizes recent information. While the attention distribution is smooth and gradual, its structure is systematic and economically meaningful. This temporal weighting explains the smoothing of extreme price movements observed in the forecasting results and reflects a bias–variance trade-off in which forecast stability and generalization are favored over sensitivity to isolated extreme shocks.
This pattern indicates that the model primarily relies on short-term temporal dependencies, which is consistent with the stylized facts of agricultural commodity markets, where recent price movements and short-run volatility often dominate price formation. Beyond roughly ten lags, the attention weights converge to near-constant low values, suggesting diminishing marginal relevance of longer historical windows within the chosen 30-day lookback period.
Figure 14 further illustrates this behavior through an attention heatmap averaged across heads. High-intensity regions are concentrated along the leftmost columns, corresponding to recent source time steps. The relative uniformity of the attention structure across target time steps suggests that the model learns a stable temporal relevance pattern, consistently prioritizing recent information regardless of the forecast position. This structural property explains the smoothness observed in the hybrid model’s predictions and highlights how attention mechanisms contribute primarily to interpretability rather than aggressive error minimization.

4.6. Discussion

The forecasting plots reveal that the proposed model produces relatively smooth predictions and tends to underestimate sharp price spikes during periods of high volatility. While this behavior contributes to stable trend estimation and noise reduction, it also represents a limitation of the model. In particular, the use of MSE-based loss functions and attention-driven temporal averaging introduces a bias–variance trade-off that favors robustness over sensitivity to extreme events. Consequently, the model is more suitable for medium-term trend forecasting than for applications requiring accurate tail-risk or volatility prediction.
An attention-weight analysis reveals that the model assigns greater importance to periods characterized by sustained market shifts rather than isolated price spikes. This behavior explains the observed smoothing of extreme price movements in the prediction plots. While such smoothing improves forecast stability and reduces noise sensitivity, it also leads to systematic underestimation of rare but economically significant price spikes. This reflects an inherent bias–variance trade-off and should be interpreted as a limitation rather than an unconditional advantage.
Comparisons with baseline models must be interpreted cautiously as some benchmark results are drawn from heterogeneous datasets and experimental settings. However, the consistent performance gains observed in re-trained baselines support the effectiveness of the proposed architecture.
From a policy and supply chain perspective, the model’s strength lies in its ability to provide reliable trend forecasts rather than precise spike prediction. Such forecasts are valuable for inventory planning, export strategy design, and medium-term policy analysis in commodity-dependent regions such as West Africa. However, for applications requiring accurate tail-risk estimation, such as hedging against extreme shocks, future extensions incorporating external variables (for example, weather indices and geopolitical indicators) and volatility-weighted loss functions are necessary. Consequently, the findings of this study are best interpreted as methodological and analytical contributions rather than prescriptive policy tools.

4.7. Study Ramifications

For agricultural commodity markets and their players, the LSTM–MHA model has a number of significant implications. These implications span a wide range of fields and show how the model may affect several aspects of agricultural commodity trading and management. Our findings are influential regarding forecasting technology, economic and policy issues, industry-specific applications, and market operations and decision-making.
Market activities are immediately impacted by the LSTM–MHA model’s implementation. Through improved price forecasts, the model helps stakeholders to make well-informed decisions about trading tactics, position management, inventory optimization, storage choices, production scheduling, resource allocation, and risk management.
Significant technological implications are shown by the successful integration of LSTM networks with multi-head attention mechanisms. This demonstrates how hybrid deep learning techniques may be used to estimate agricultural prices, provides a foundation for developing more complex forecasting models, and indicates how many neural network architectures can be integrated to address complex market dynamics.
The results of this study should not be interpreted as direct evidence of improved market efficiency or policy effectiveness. Instead, the proposed model contributes indirectly by providing more reliable and interpretable trend-level forecasts of global commodity price dynamics. Such forecasts may support decision-making processes related to planning, risk assessment, and market monitoring, particularly in regions exposed to international price transmission. However, causal policy conclusions would require the integration of domestic price data, institutional factors, and exogenous variables, which remain beyond the scope of this study.
Cocoa, coffee, cotton, timber, orange juice, and sugar are the six different commodities that are examined in this study. The results have particular significance for supply chain optimization, seasonal planning, production scheduling, investment strategy design, cross-commodity market analysis, and correlation studies.
The ablation and attention analyses demonstrate that attention mechanisms contribute more to interpretability and structural understanding than to immediate gains in point forecasting accuracy. The results also suggest that the full potential of the hybrid LSTM–MHA architecture is likely to emerge when richer feature sets, such as weather indicators, exchange rates, or policy variables, are incorporated. In such settings, attention mechanisms can selectively weight heterogeneous information sources, potentially improving both accuracy and robustness under volatile market regimes.

5. Conclusions

This study proposed an attention-enhanced deep learning framework for agricultural commodity price forecasting, integrating long short-term memory (LSTM) networks with a multi-head attention (MHA) mechanism. Using global agricultural commodity futures prices as benchmark data, the model was designed to capture complex temporal dependencies and nonlinear price dynamics that are inadequately addressed by traditional statistical and standalone deep learning approaches. While the empirical analysis relies on globally traded futures markets, the study is motivated by the strong exposure of West African economies to international commodity price movements, positioning the region as an important application context rather than a distinct domestic market.
The empirical results demonstrate that the proposed LSTM–MHA model achieves substantially lower forecasting errors than conventional benchmarks such as ARIMA and standalone LSTM models, reducing mean squared error by approximately three to four times in comparable settings. The hybrid architecture effectively balances sequential learning and dynamic temporal weighting, enabling robust trend prediction across multiple commodities, including cocoa, coffee, cotton, lumber, orange juice, and sugar. The learning curves and validation results further confirm stable convergence and strong generalization performance, indicating that the model captures underlying price structures rather than overfitting short-term noise.
Beyond predictive accuracy, a key contribution of this study lies in the interpretability enabled by the attention mechanism. Attention-weight and lag-importance analyses reveal that the model consistently prioritizes recent historical observations, typically within a short temporal window, while assigning diminishing importance to distant lags. This behavior aligns with stylized facts of commodity markets, where short-term dynamics and recent shocks play a dominant role in price formation. At the same time, the attention structure explains the observed smoothing of extreme price movements in the forecasting plots, highlighting an inherent bias–variance trade-off. While this conservative behavior enhances forecast stability under noisy market conditions, it also leads to systematic underestimation of rare but sharp price spikes, which is acknowledged as a limitation.
The ablation study further clarifies the role of the architectural components. The results indicate that the inclusion of multi-head attention does not automatically yield lower point forecast errors in a low-dimensional price-driven setting. Instead, attention primarily contributes to structural flexibility and interpretability, trading a small degree of numerical accuracy for improved transparency and robustness. This finding suggests that the full benefits of attention mechanisms are likely to emerge in richer modeling environments that incorporate additional exogenous variables, such as weather indicators, exchange rates, or policy signals.
From a practical perspective, the proposed framework offers valuable insights for market participants and policymakers. Reliable trend forecasts can support inventory planning, export strategy formulation, and medium-term policy analysis in commodity-dependent regions such as West Africa, where global price transmission plays a critical role. However, for applications requiring precise tail-risk estimation or hedging against extreme shocks, future extensions are necessary.
Several avenues for future research emerge from this study. First, incorporating external drivers of price volatility, including climatic, macroeconomic, and geopolitical variables, may enhance the model’s ability to capture extreme events. Second, alternative loss functions or volatility-sensitive training objectives could be explored to better balance stability and tail-risk responsiveness. Finally, extending the framework to real-time forecasting and regime-switching environments would further improve its applicability in dynamic agricultural markets.
In summary, this study demonstrates that attention-enhanced deep learning models provide a powerful and interpretable tool for agricultural commodity price forecasting, offering both methodological contributions and practical relevance for globally integrated markets with regional economic implications.

Author Contributions

Conceptualization, M.G.T.; methodology, M.G.T.; software, M.G.T.; validation, M.G.T.; formal analysis, M.G.T.; investigation, M.G.T.; resources, S.Z.; data curation, M.G.T.; writing—original draft preparation, M.G.T.; writing—review and editing, S.Z., J.Z. and Q.X.; visualization, M.G.T.; supervision, S.Z.; project administration, S.Z.; funding acquisition, S.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This study was funded by the Sichuan Social Science Planned Key Research Project (SCJ23ND61), the 2022 Central Chinese University Fundamental Research Program for Humanities and Social Science Cultivation Key Project (No. ZYGX2022FRJH004), and Regional funded Studies of the Ministry of Education of China (No. 2024-N01).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions and raw data presented in this study are included and referenced in this article. However, further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Gouel, C. Trade policy coordination and food price volatility. Am. J. Agric. Econ. 2016, 98, 1018–1037. [Google Scholar] [CrossRef]
  2. Mamoudan, M.M.; Mohammadnazari, Z.; Ostadi, A.; Esfahbodi, A. Food products pricing theory with application of machine learning and game theory approach. Int. J. Prod. Res. 2022, 62, 5489–5509. [Google Scholar] [CrossRef]
  3. Jabir, E.; Panicker, V.V.; Sridharan, R. Environmental friendly route design for a milk collection problem: The case of an Indian dairy. Int. J. Prod. Res. 2020, 60, 912–941. [Google Scholar] [CrossRef]
  4. Gadafi, T.M.; Decui, L.; Darko, A.P. Two stages method-based on Africa smart irrigation system assessment for willingness to pay: A case of Ghana Northern Region. Socio-Econ. Plan. Sci. 2025, 102, 102318. [Google Scholar] [CrossRef]
  5. Liu, C.; Zhu, H.; Tang, D.; Nie, Q.; Li, S.; Zhang, Y.; Liu, X. A transfer learning CNN-LSTM network-based production progress prediction approach in IIoT-enabled manufacturing. Int. J. Prod. Res. 2022, 61, 4045–4068. [Google Scholar] [CrossRef]
  6. Lagi, M.; Bertrand, K.Z.; Bar-Yam, Y. The Food Crises and Political Instability in North Africa and the Middle East. arXiv 2011, arXiv:1108.2455. [Google Scholar] [CrossRef]
  7. Wheeler, T.; Von Braun, J. Climate change impacts on global food security. Science 2013, 341, 508–513. [Google Scholar] [CrossRef] [PubMed]
  8. Prakash, A. Organisation des Nations Unies pour l’alimentation et l’agriculture. In Safeguarding Food Security in Volatile Global Markets; Food and Agriculture Organization of the United Nations: Rome, Italy, 2011. [Google Scholar]
  9. Tomek, W.G.; Kaiser, H.M. Agricultural Product Prices; Cornell University Press: Ithaca, NY, USA, 2014. [Google Scholar] [CrossRef]
  10. Godfray, H.C.J.; Beddington, J.R.; Crute, I.R.; Haddad, L.; Lawrence, D.; Muir, J.F.; Pretty, J.; Robinson, S.; Thomas, S.M.; Toulmin, C. Food security: The challenge of feeding 9 billion people. Science 2010, 327, 812–818. [Google Scholar] [CrossRef]
  11. Bellemare, M.F. Rising food prices, food price volatility, and social unrest. Am. J. Agric. Econ. 2015, 97, 1–21. [Google Scholar] [CrossRef]
  12. Jha, G.K.; Sinha, K. Agricultural price forecasting using neural network model: An innovative information delivery system. Agric. Econ. Res. Rev. 2013, 26, 229–240. [Google Scholar] [CrossRef]
  13. Box, G.E.P.; Jenkins, G.M.; Reinsel, G.C.; Ljung, G.M. Time Series Analysis: Forecasting and Control; John Wiley & Sons: Hoboken, NJ, USA, 2016. [Google Scholar] [CrossRef]
  14. Graves, A. Generating Sequences with Recurrent Neural Networks. arXiv 2013, arXiv:1308.0850. [Google Scholar] [CrossRef]
  15. Heaton, J. Ian Goodfellow, Yoshua Bengio, and Aaron Courville: Deep learning—The MIT Press, 2016, 800 pp, ISBN: 0262035618. Genet. Program. Evolvable Mach. 2018, 19, 305–307. [Google Scholar] [CrossRef]
  16. Pascanu, R.; Mikolov, T.; Bengio, Y. On the difficulty of training recurrent neural networks. In Proceedings of the International Conference on Machine Learning, Edinburgh, UK, 26 June–1 July 2012. [Google Scholar]
  17. Vaswani, A.; Shazeer, N.M.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is All You Need. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
  18. Hobbs, J.E. Food supply chains during the COVID-19 pandemic. Can. J. Agric. Econ. Can. D’Agroecon. 2020, 68, 171–176. [Google Scholar] [CrossRef]
  19. Zhao, S.; Tamimu, M.G.; Luo, A.; Sun, T.; Yang, Y. Hesitant Fuzzy-BWM Risk Evaluation Framework for E-Business Supply Chain Cooperation for China–West Africa Digital Trade. J. Theor. Appl. Electron. Commer. Res. 2025, 20, 233. [Google Scholar] [CrossRef]
  20. Wright, B.D. The economics of grain price volatility. Appl. Econ. Perspect. Policy 2011, 33, 32–58. [Google Scholar] [CrossRef]
  21. Baffes, J.; Gardner, B. The transmission of world commodity prices to domestic markets under policy reforms in developing countries. Policy Reform 2003, 6, 159–180. [Google Scholar] [CrossRef]
  22. MacDonald, S.; Meyer, L.A. Long Run Trends and Fluctuations In Cotton Prices. In Proceedings of a Conference. 2018. Available online: https://mpra.ub.uni-muenchen.de/84484/ (accessed on 10 September 2025).
  23. Prestemon, J.P.; Holmes, T.P. Timber price dynamics following a natural catastrophe. Am. J. Agric. Econ. 2000, 82, 145–160. [Google Scholar] [CrossRef]
  24. Schmitz, C.; Biewald, A.; Lotze-Campen, H.; Popp, A.; Dietrich, J.P.; Bodirsky, B.; Krause, M.; Weindl, I. Trading more food: Implications for land use, greenhouse gas emissions, and the food system. Glob. Environ. Change 2012, 22, 189–209. [Google Scholar] [CrossRef]
  25. Qin, Y.; Song, D.; Chen, H.; Cheng, W.; Jiang, G.; Cottrell, G. A dual-stage attention-based recurrent neural network for time series prediction. arXiv 2017, arXiv:1704.02971. [Google Scholar] [CrossRef]
  26. Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence; Association for the Advancement of Artificial Intelligence: Washington, DC, USA, 2020; Volume 35, pp. 11106–11115. [Google Scholar] [CrossRef]
  27. Tamimu, M.G.; Liang, D. Sustainable mining development in Ghana: An integrated AHP-TOPSIS-Manhattan distance approach. Resour. Policy 2025, 110, 105757. [Google Scholar] [CrossRef]
  28. Tamimu, M.G.; Zhao, S.; Wang, Y.; Yang, Y.; Sun, T.; Luo, A. Fuzzy logic-BWM-based risk assessment for E-business cooperation in China-West Africa digital trade: A path to sustainable business and policies. Glob. Econ. Res. 2025, 1, 100014. [Google Scholar] [CrossRef]
  29. Tamimu, M.G.; Zhao, S.; Xu, Q.; Zhang, J.; Yin, X. Pythagorean Fuzzy-AHP (PF-AHP) Approach for Emerging New Risk Evaluation in China-West Africa Digital Trade Collaboration. J. Theor. Appl. Electron. Commer. Res. 2025, 20, 327. [Google Scholar] [CrossRef]
  30. Chatziantoniou, I.; Degiannakis, S.; Delis, P.; Filis, G. Forecasting oil price volatility using spillover effects from uncertainty indices. Financ. Res. Lett. 2020, 42, 101885. [Google Scholar] [CrossRef]
  31. Kohzadi, N.; Boyd, M.S.; Kermanshahi, B.; Kaastra, I. A comparison of artificial neural network and time series models for forecasting commodity prices. Neurocomputing 1996, 10, 169–181. [Google Scholar] [CrossRef]
  32. Figuerola-Ferretti, I.; Gilbert, C.L. Commonality in the LME aluminum and copper volatility processes through a FIGARCH lens. J. Futur. Mark. Futur. Options Other Deriv. Prod. 2008, 28, 935–962. [Google Scholar] [CrossRef]
  33. Xiong, T.; Li, C.; Bao, Y.; Hu, Z.; Zhang, L. A combination method for interval forecasting of agricultural commodity futures prices. Knowl.-Based Syst. 2015, 77, 92–102. [Google Scholar] [CrossRef]
  34. Sujjaviriyasup, T.; Pitiruek, K. Agricultural product forecasting using machine learning approach. Int. J. Math. Anal. 2013, 7, 1869–1875. [Google Scholar] [CrossRef]
  35. Hochreiter, S. Long Short-term Memory. In Neural Computation; MIT-Press: Cambridge, MA, USA, 1997. [Google Scholar] [CrossRef]
  36. Bahdanau, D.; Cho, K.; Bengio, Y. Neural Machine Translation by Jointly Learning to Align and Translate. arXiv 2014, arXiv:1409.0473. [Google Scholar] [CrossRef]
  37. Kristjanpoller, W.; Minutolo, M.C. Gold price volatility: A forecasting approach using the Artificial Neural Network–GARCH model. Expert Syst. Appl. 2015, 42, 7245–7251. [Google Scholar] [CrossRef]
  38. Kantasa-ard, A.; Nouiri, M.; Bekrar, A.; El Cadi, A.A.; Sallez, Y. Machine learning for demand forecasting in the physical internet: A case study of agricultural products in Thailand. Int. J. Prod. Res. 2020, 59, 7491–7515. [Google Scholar] [CrossRef]
  39. Cen, Z.; Wang, J. Crude oil price prediction model with long short term memory deep learning based on prior knowledge data transfer. Energy 2019, 169, 160–171. [Google Scholar] [CrossRef]
  40. Livieris, I.E.; Pintelas, E.; Pintelas, P. A CNN–LSTM model for gold price time-series forecasting. Neural Comput. Appl. 2020, 32, 17351–17360. [Google Scholar] [CrossRef]
  41. Olofintuyi, S.S.; Olajubu, E.A.; Olanike, D. An ensemble deep learning approach for predicting cocoa yield. Heliyon 2023, 9, e15245. [Google Scholar] [CrossRef]
  42. Mekala, K.; Laxmi, V.; Jagruthi, H.; Dhondiyal, S.A.; Sridevi, R.; Dabral, A.P. Coffee Price Prediction: An Application of CNN-BLSTM Neural Networks. In Proceedings of the 2023 International Conference on Advances in Computing, Communication and Applied Informatics (ACCAI), Chennai, India, 25–26 May 2023; pp. 1–7. [Google Scholar] [CrossRef]
  43. Dave, E.; Leonardo, A.; Jeanice, M.; Hanafiah, N. Forecasting Indonesia exports using a hybrid model ARIMA-LSTM. Procedia Comput. Sci. 2021, 179, 480–487. [Google Scholar] [CrossRef]
  44. Lamichhane, S.; Mei, B.; Siry, J. Forecasting pine sawtimber stumpage prices: A comparison between a time series hybrid model and an artificial neural network. For. Policy Econ. 2023, 154, 103028. [Google Scholar] [CrossRef]
  45. Luo, J.; Klein, T.; Ji, Q.; Hou, C. Forecasting realized volatility of agricultural commodity futures with infinite Hidden Markov HAR models. Int. J. Forecast. 2022, 38, 51–73. [Google Scholar] [CrossRef]
  46. Ribeiro, C.O.; Oliveira, S.M. A hybrid commodity price-forecasting model applied to the sugar–alcohol sector. Aust. J. Agric. Resour. Econ. 2011, 55, 180–198. [Google Scholar] [CrossRef]
  47. Deepa, P.B.; Daisy, J. Deep Learning Based Prediction of Commodity Prices Using LSTM. In Proceedings of the 2023 4th International Conference on Smart Electronics and Communication (ICOSEC), Trichy, India, 20–22 September 2023; pp. 1708–1715. [Google Scholar] [CrossRef]
  48. Ren, B.; Xu, X.; Yu, H. Research of LSTM-RNN Model and Its Application Evaluation on Agricultural Products Circulation. In Proceedings of the 2021 IEEE 3rd Eurasia Conference on IOT, Communication and Engineering (ECICE), Yunlin, Taiwan, 29–31 October 2021; pp. 467–471. [Google Scholar] [CrossRef]
  49. Tami, M.; Owda, A.Y. Efficient commodity price forecasting using long short-term memory model. IAES Int. J. Artif. Intell. (IJ-AI) 2024, 13. [Google Scholar] [CrossRef]
  50. Sari, M.; Duran, S.; Kutlu, H.; Guloglu, B.; Atik, Z. Various optimized machine learning techniques to predict agricultural commodity prices. Neural Comput. Appl. 2024, 36, 11439–11459. [Google Scholar] [CrossRef]
  51. Zhang, Q.; Yang, W.; Zhao, A.; Wang, X.; Wang, Z.; Zhang, L. Short-term forecasting of vegetable prices based on lstm model—Evidence from Beijing’s vegetable data. PLoS ONE 2024, 19, e0304881. [Google Scholar] [CrossRef]
  52. Zhang, T.; Tang, Z. Agricultural commodity futures prices prediction based on a new hybrid forecasting model combining quadratic decomposition technology and LSTM model. Front. Sustain. Food Syst. 2024, 8, 1334098. [Google Scholar] [CrossRef]
  53. Terrada, L.; El Khaili, M.; Ouajji, H. Demand forecasting model using deep learning methods for supply chain management 4.0. Int. J. Adv. Comput. Sci. Appl. 2022, 13. [Google Scholar] [CrossRef]
  54. Ray, S.; Lama, A.; Mishra, P.; Biswas, T.; Das, S.S.; Gurung, B. An ARIMA-LSTM model for predicting volatile agricultural price series with random forest technique. Appl. Soft Comput. 2023, 149, 110939. [Google Scholar] [CrossRef]
  55. Ma, X.Y.; Tong, J.; Jiang, F.; Xu, M.; Sun, L.M.; Chen, Q.Y. Application of Deep Learning to Production Forecasting in Intelligent Agricultural Product Supply Chain. Comput. Mater. Contin. 2023, 74, 6145–6159. [Google Scholar] [CrossRef]
  56. Servera, G. Downloading Agricultural Products Futures [Dataset]. Kaggle. 2021. Available online: https://www.kaggle.com/code/guillemservera (accessed on 2 October 2024).
  57. Kırelli, Y. Comparative Analysis of LSTM and ARIMA Models in Stock Price Prediction: A Technology Company Example. Black Sea J. Eng. Sci. 2024, 7, 15–16. [Google Scholar] [CrossRef]
  58. Murugesan, R.; Mishra, E.; Krishnan, A.H. Deep Learning Based Models: Basic LSTM, Bi LSTM, Stacked LSTM, CNN LSTM and Conv LSTM to Forecast Agricultural Commodities Prices. Res. Sq. 2021; preprint. [Google Scholar] [CrossRef]
  59. Gadafi, T.M.; Mohammed, A.M.A.; Ma, J.; Muhammad, S.R.; Darko, A.P.; Liang, D. BiLSTM-Based Climate and Agricultural Supply Chain Resilience Modeling Using Time Series Forecasting in Ghana. In Proceedings of the 2nd International Conference on Image Processing, Machine Learning, and Pattern Recognition, Kunming, China, 12–13 July 2025; ACM: New York, NY, USA, 2025; pp. 225–230. [Google Scholar] [CrossRef]
  60. Yun, B.; Lai, J.; Ma, Y.; Zheng, Y. Research on Grain Futures Price Prediction Based on a Bi-DSConvLSTM-Attention Model. Systems 2024, 12, 204. [Google Scholar] [CrossRef]
  61. Gadafi, T.M.; Toufic, S.; Sagoe, A.A.; Decui, L.; Anwar, H. Enhancing Climate Resilience through Sustainable Agricultural Intensification: An Attention-Based Neural Network for Climate-Smart Agriculture Modeling in Ghana. In Proceedings of the 2025 International Conference on Smart Agriculture and Artificial Intelligence, Xi’an, China, 13–15 June 2025; ACM: New York, NY, USA, 2025. [Google Scholar] [CrossRef]
Figure 1. LSTM–multi-head attention hybrid model architecture.
Figure 1. LSTM–multi-head attention hybrid model architecture.
Applsci 16 01600 g001
Figure 2. Single-layer LSTM structure.
Figure 2. Single-layer LSTM structure.
Applsci 16 01600 g002
Figure 3. Our proposed model summary.
Figure 3. Our proposed model summary.
Applsci 16 01600 g003
Figure 4. The proposed model loss.
Figure 4. The proposed model loss.
Applsci 16 01600 g004
Figure 5. The true values versus predicted values for train data.
Figure 5. The true values versus predicted values for train data.
Applsci 16 01600 g005
Figure 6. The true values versus predicted values for validation data.
Figure 6. The true values versus predicted values for validation data.
Applsci 16 01600 g006
Figure 7. The true values versus predicted values for test data.
Figure 7. The true values versus predicted values for test data.
Applsci 16 01600 g007
Figure 8. The true values versus predicted values for the last 1000 points.
Figure 8. The true values versus predicted values for the last 1000 points.
Applsci 16 01600 g008
Figure 9. The true values versus predicted values for all data.
Figure 9. The true values versus predicted values for all data.
Applsci 16 01600 g009
Figure 10. The scatterplot of true values versus predicted values for all data.
Figure 10. The scatterplot of true values versus predicted values for all data.
Applsci 16 01600 g010
Figure 11. Ablation study comparing the forecasting performance of LSTM, attention-only, and hybrid LSTM–MHA models in terms of mean squared error (MSE).
Figure 11. Ablation study comparing the forecasting performance of LSTM, attention-only, and hybrid LSTM–MHA models in terms of mean squared error (MSE).
Applsci 16 01600 g011
Figure 12. Lag importance profile based on averaged attention weights. The figure shows a monotonic decay in attention as the lag increases, indicating that recent historical observations contribute more strongly to the forecasting process.
Figure 12. Lag importance profile based on averaged attention weights. The figure shows a monotonic decay in attention as the lag increases, indicating that recent historical observations contribute more strongly to the forecasting process.
Applsci 16 01600 g012
Figure 13. Relative lag importance after centering around the mean attention weight. The centered representation highlights deviations from the average attention level, emphasizing the dominance of early lags and the diminishing contribution of older historical information.
Figure 13. Relative lag importance after centering around the mean attention weight. The centered representation highlights deviations from the average attention level, emphasizing the dominance of early lags and the diminishing contribution of older historical information.
Applsci 16 01600 g013
Figure 14. Attention heatmap averaged across heads, illustrating the temporal focus of the hybrid LSTM–MHA model. Higher attention weights are concentrated on recent source time steps.
Figure 14. Attention heatmap averaged across heads, illustrating the temporal focus of the hybrid LSTM–MHA model. Higher attention weights are concentrated on recent source time steps.
Applsci 16 01600 g014
Table 1. Descriptive comparison of existing deep learning approaches for agricultural commodity price forecasting.
Table 1. Descriptive comparison of existing deep learning approaches for agricultural commodity price forecasting.
ReferenceModel TypeMultivariate InputAttention MechanismEvaluation MetricsTrain/Val/Test SplitForecasting Scope
[47]LSTMNoNoMSENoSingle commodity
[48]LSTMNoNoRMSENoSingle commodity
[49]LSTMNoNoMSEYesSingle commodity
[50]LSTM/CNN–LSTMYesNoMSE, RMSEYesSingle commodity
[51]LSTMNoNoRMSEYesSingle commodity
[52]Hybrid DLYesNoMSE, RMSE, MAEYesSingle commodity
[53]LSTMNoNoRMSENoSingle commodity
[54]ARIMA–LSTMNoNoRMSEYesSingle commodity
[55]LSTMNoNoMSENoSingle commodity
This studyLSTM–MHAYesYesMSE, RMSE, MAEYesMultiple commodities
Table 2. Parameters of our prediction model.
Table 2. Parameters of our prediction model.
ParameterEpochLearning RateBatch SizeOptimizerNumber of Heads
Value 1500.00164Adam8
Value 21000.01128Adam16
Value 31500.1256Adam32
Table 3. Effects of the experiment hyperparameters and result analysis of the LSTM–MHA model using agricultural commodity futures data.
Table 3. Effects of the experiment hyperparameters and result analysis of the LSTM–MHA model using agricultural commodity futures data.
ParameterValue
Epochs50, 100, 150
Learning rate (LR)0.001, 0.01, 0.1
Batch size64, 128, 256
Multi-head attention8, 16, 32
Results
Epoch 50
Training loss0.0349
Validation loss0.0339
Test loss0.03464767336845398
Mean squared error (MSE)0.012414070816788778
Root mean squared error (RMSE)0.11141844917601743
Root mean absolute error (MAE)0.10966918723679892
Epoch 100
Training loss0.0347
Validation loss0.0347
Test loss0.03464708849787712
Mean squared error (MSE)0.012414070816788778
Root mean squared error (RMSE)0.11141844917601743
Root mean absolute error (MAE)0.10966918723679892
Epoch 150
Training loss0.0349
Validation loss0.0339
Test loss0.03464767336845398
Mean squared error (MSE)0.012414070816788778
Root mean squared error (RMSE)0.11141844917601743
Root mean absolute error (MAE)0.10966918723679892
Table 4. Comparative analysis of LSTM–MHA hybrid models.
Table 4. Comparative analysis of LSTM–MHA hybrid models.
Model ArchitectureMSERMSEMAE
Our Proposed LSTM–MHA Model0.01240.11140.1097
LSTM [57]0.03740.19360.1700
ARIMA [57]0.04880.22110.1968
ARIMA–LSTM (random forest) for Gram Series [54]-5.2311.4522
ARIMA–LSTM (random forest) for Moong Series [54]-6.0991.9008
ARIMA–LSTM (random forest) for Urad Series [54]-5.9561.7261
BiLSTM [58]0.1039990.3224900.259999
BiLSTM-based time series [59]-0.18190.1687
Bi-DSConvLSTM–attention [60]-5.613.63
Multi-head attention (8 heads) [60]-5.763.57
Attention-based neural network [61]0.0310.230-
CNN–LSTM [58]0.0929990.3049590.289999
Conv–LSTM [58]0.4699990.2167940.216794
Note: The best metric results are highlighted in bold. Missing values (-) indicate metrics not reported in the original paper.
Table 5. Ablation study: forecasting error comparison across model variants.
Table 5. Ablation study: forecasting error comparison across model variants.
ModelMSERMSEMAE
LSTM0.00460.06780.0350
Attention-only0.00460.06770.0451
Hybrid LSTM–MHA0.00540.07330.0393
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Tamimu, M.G.; Zhao, S.; Xu, Q.; Zhang, J. Attention-Based Deep Learning Hybrid Model for Cash Crop Price Forecasting: Evidence from Global Futures Markets with Implications for West Africa. Appl. Sci. 2026, 16, 1600. https://doi.org/10.3390/app16031600

AMA Style

Tamimu MG, Zhao S, Xu Q, Zhang J. Attention-Based Deep Learning Hybrid Model for Cash Crop Price Forecasting: Evidence from Global Futures Markets with Implications for West Africa. Applied Sciences. 2026; 16(3):1600. https://doi.org/10.3390/app16031600

Chicago/Turabian Style

Tamimu, Mohammed Gadafi, Shurong Zhao, Qianwen Xu, and Jie Zhang. 2026. "Attention-Based Deep Learning Hybrid Model for Cash Crop Price Forecasting: Evidence from Global Futures Markets with Implications for West Africa" Applied Sciences 16, no. 3: 1600. https://doi.org/10.3390/app16031600

APA Style

Tamimu, M. G., Zhao, S., Xu, Q., & Zhang, J. (2026). Attention-Based Deep Learning Hybrid Model for Cash Crop Price Forecasting: Evidence from Global Futures Markets with Implications for West Africa. Applied Sciences, 16(3), 1600. https://doi.org/10.3390/app16031600

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop