A Mathematical Framework for E-Commerce Sales Prediction Using Attention-Enhanced BiLSTM and Bayesian Optimization
Abstract
1. Introduction
- Architectural Innovation: The design of an enhanced BiLSTM network aims to capture bidirectional temporal dependencies more effectively. It integrates an improved attention mechanism to selectively highlight critical sequential features.
- Attention Mechanism Enhancement: The introduction of a refined attention mechanism that adaptively weighs temporal information, enhancing the focus on influential patterns in complex cross-border e-commerce sales data.
- Hyperparameter Optimization: The application of Bayesian optimization to automatically and efficiently tune key model hyperparameters for optimal performance.
- Empirical Validation of Algorithmic Design: Demonstrate through extensive experiments on a large-scale cross-border e-commerce dataset, that covers multiple product categories, that the proposed structural improvements are effective.
2. Related Work
2.1. Neural Networks for Cross-Border E-Commerce
2.2. Attention-Based Sales Prediction
3. Fundamentals
3.1. E-Commerce Sales Forecast
3.2. Bidirectional LSTM
3.3. Hyperparameter Optimization in Deep Learning Models
3.4. Attention Mechanism Enhancement for Sequential Feature Highlighting
4. Methodology
4.1. Enhanced BiLSTM Network with Attention and Residual Connections
| Algorithm 1 Enhanced BiLSTM with Dual-Activation Attention and Residual Connections |
Input: Time series Output: Predicted values
|
4.2. Adaptive Attention Mechanism
4.3. Hyperparameter Optimization
4.4. Overview of the Proposed Method
- BiLSTM Input Processing: The input time series is processed by the BiLSTM layer, which captures both past trends and future events, such as upcoming promotions. Its bidirectional modeling is especially effective for short, volatile sequences where long-range attention may be less reliable.
- Dual-Activation Module: A dual-activation mechanism applies both ReLU and Sigmoid to the BiLSTM hidden states. ReLU amplifies strong, discriminative signals (e.g., promotional spikes), while Sigmoid attenuates weak or noisy signals (e.g., regular months without promotions). This complementary design enables selective feature amplification and gating, improving the model’s ability to capture both short-term fluctuations and long-term trends.
- Adaptive Attention Mechanism: The attention module dynamically adjusts the weight of each time step. During peak seasons, like Black Friday, it assigns higher importance to promotional periods while reducing the influence of irrelevant data. Combined with dual-activation preprocessing, this mechanism ensures that the model focuses on both strong events and contextually relevant patterns.
- Residual Connections: Residual connections are added between the BiLSTM and attention layers. These connections stabilize training by preserving key features, preventing gradient vanishing, and enabling more complex interactions between original and transformed hidden states.
- Bayesian Optimization: Bayesian optimization tunes hyperparameters such as learning rate, hidden layer size, and dropout rate automatically. This adaptive tuning improves model generalization and stability, particularly important for small datasets with high volatility, where manual tuning may fail to achieve consistent performance.
5. Experimental Setup
5.1. Data Source
5.2. Data Preprocessing
- Data Cleaning: Duplicate records and missing entries are removed to ensure data quality. Outliers, such as extremely high or low sales values caused by data entry errors or exceptional events, are detected and handled appropriately. This step ensures that the data fed into the model is reliable and consistent.
- Data Encoding: Time-related features are extracted from the original timestamps, which are at a daily granularity. Numeric fields for year and month are created, e.g., year values as 2022, 2023, and month values as 1–12, to enable temporal modeling at monthly granularity.
- Data Aggregation: Sales are aggregated at the monthly level, which is an important granularity for capturing seasonal patterns and promotional effects. Analysis of historical sales shows that demand varies significantly across months, with peaks during promotional seasons such as Prime Day or Black Friday.
- Feature Construction: To capture seasonal effects, a numeric seasonal feature is added representing the quarter (1 to 4). Promotional activity is encoded as a binary feature, with 1 indicating the presence of a sale or discount campaign during the month and 0 otherwise. Additionally, other relevant product attributes, such as category, price tier, and rating, are included as features to enhance model performance.
5.3. Experimental Environment and Parameter Settings
5.3.1. Experimental Environment Configuration
5.3.2. Model-Specific Parameter Settings
5.4. Comparison Algorithms
5.5. Evaluation Metrics
6. Results
6.1. Comparison Results
6.2. Model Component Evaluation
6.3. Hyperparameter Influence Evaluation
6.4. Training Time and Computational Efficiency Evaluation
6.5. Effect of Sequence Length/Time Window
7. Conclusions
- The model outperforms six baseline models, achieving a RMSE of 13.2, MAE of 10.2, MAPE of 8.7%, and a coefficient of R2 of 0.92, demonstrating its superior accuracy in cross-border e-commerce sales forecasting.
- Ablation studies show the contribution of each component: the adaptive attention mechanism improves prediction accuracy, residual connections reduce errors, and Bayesian optimization enhances overall model performance.
- The model strikes an effective balance between performance and computational efficiency. With fast training and inference times, it demonstrates practicality for real-world deployment, making it suitable for large-scale, real-time forecasting tasks.
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Qi, K.; Wu, W.; Ni, Y. SymOpt-CNSVR: A Novel Prediction Model Based on Symmetric Optimization for Delivery Duration Forecasting. Symmetry 2025, 17, 1608. [Google Scholar]
- Zhu, L.; Liu, W.; Tan, H.; Hu, T. A novel Bayesian optimization prediction framework for four-axis industrial robot joint motion state. Complex Intell. Syst. 2024, 10, 4867–4881. [Google Scholar] [CrossRef]
- Gong, Z. Optimization of cross-border E-commerce (CBEC) supply chain management based on fuzzy logic and auction theory. Sci. Rep. 2024, 14, 14088. [Google Scholar]
- Li, J.; Li, J.; Li, J.; Zhang, G. Bayesian-Optimized GCN-BiLSTM-Adaboost Model for Power-Load Forecasting. Electronics 2025, 14, 3332. [Google Scholar]
- Guo, X.; Mo, Y.; Yan, K. Short-Term Photovoltaic Power Forecasting Based on Historical Information and Deep Learning Methods. Sensors 2022, 22, 9630. [Google Scholar] [CrossRef]
- Kwon, H.; Do, T.N.; Won, W.; Kim, J. An optimization model for the market-responsive operation of naphtha cracking process with price prediction. Chem. Eng. Res. Des. 2022, 188, 681–693. [Google Scholar] [CrossRef]
- Lou, G.; Lin, W.; Huang, G.; Xiang, W. A two-stage online remaining useful life prediction framework for supercapacitors based on the fusion of deep learning network and state estimation algorithm. Eng. Appl. Artif. Intell. 2023, 123, 106399. [Google Scholar] [CrossRef]
- Gena, X.; Zhang, Y. Sales Forecasting using ARIMA Models: A Case Study in the Retail Industry. J. Retail Anal. 2015, 12, 45–56. [Google Scholar]
- Rostan, P.; Rostan, A.; Nurunnabi, M. Options Trading Strategy Based on ARIMA Forecasting. PSU Res. Rev. 2020, 4, 111–127. [Google Scholar] [CrossRef]
- Wang, P.; Xu, Z. A Novel Consumer Purchase Behavior Recognition Method Using Ensemble Learning Algorithm. Math. Probl. Eng. 2020, 2020, 6673535. [Google Scholar] [CrossRef]
- Cho, M.; Kim, C.; Jung, K.; Jung, H. Water Level Prediction Model Applying a Long Short-Term Memory (LSTM)—Gated Recurrent Unit (GRU) Method for Flood Prediction. Water 2022, 14, 2221. [Google Scholar] [CrossRef]
- Tsai, P.-F.; Gao, C.-H.; Yuan, S.-M. Stock Selection Using Machine Learning Based on Financial Ratios. Mathematics 2023, 11, 4758. [Google Scholar] [CrossRef]
- Khandelwal, V.; Chaturvedi, A.K.; Gupta, C.P. Amazon EC2 Spot Price Prediction Using Regression Random Forests. IEEE Trans. Cloud Comput. 2020, 8, 59–72. [Google Scholar] [CrossRef]
- Qin, C.; Chen, L.; Cai, Z.; Liu, M.; Jin, L. Long short-term memory with activation on gradient. Neural Netw. 2023, 164, 135–145. [Google Scholar] [CrossRef]
- Su, Y.; Kuo, C.-C.J. On extended long short-term memory and dependent bidirectional recurrent neural network. Neurocomputing 2019, 356, 151–161. [Google Scholar] [CrossRef]
- Bahdanau, D.; Cho, K.; Bengio, Y. Neural machine translation by jointly learning to align and translate. In Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
- Snoek, J.; Larochelle, H.; Adams, R.P. Practical Bayesian optimization of machine learning algorithms. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Lake Tahoe, NV, USA, 3–8 December 2012; Volume 25, pp. 2951–2959. [Google Scholar]
- Li, G.; Li, N. Customs classification for cross-border e-commerce based on text-image adaptive convolutional neural network. Electron. Commer. Res. 2019, 19, 779–800. [Google Scholar]
- Jiang, C. Research on sales forecasting and consumption recommendation system of e-commerce agricultural products based on LSTM model. GeoJournal 2025, 90, 93. [Google Scholar] [CrossRef]
- Li, J.; Pan, Y.; Yang, Y.; Tse, C.H. Digital platform attention and international sales: An attention-based view. J. Int. Bus. Stud. 2022, 53, 1817–1835. [Google Scholar] [CrossRef]
- Singh, K.; Booma, P.M.; Eaganathan, U. E-commerce system for sale prediction using machine learning technique. J. Phys. Conf. Ser. 2020, 1712, 012042. [Google Scholar] [CrossRef]
- He, M.; Zhuang, L.; Yang, S.; Xu, Z.; Li, W.; Lu, J. An energy-efficient VNE algorithm based on bidirectional long short-term memory. J. Netw. Syst. Manag. 2022, 30, 45. [Google Scholar] [CrossRef]
- Jürgensmeier, L.; Bischoff, J.; Skiera, B. Opportunities for self-preferencing in international online marketplaces. Int. Mark. Rev. 2024, 41, 1118–1132. [Google Scholar] [CrossRef]
- Gers, F.A.; Schmidhuber, J.; Cummins, F. Learning to forget: Continual prediction with LSTM. Neural Comput. 2000, 12, 2451–2471. [Google Scholar] [CrossRef] [PubMed]
- Hochreiter, S.; Schmidhuber, J. Long Short-Term Memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [PubMed]
- Bai, S.; Zhuang, F.; He, X. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. In Proceedings of the International Conference on Learning Representations (ICLR 2019), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Lim, B.; Arık, S.Ö.; Loeff, N.; Pfister, T. Temporal Fusion Transformers for interpretable multi-horizon time series forecasting. Int. J. Forecast. 2021, 37, 1748–1764. [Google Scholar] [CrossRef]




| Month | Year | Product Category | Units Sold | Promotion | Season (Quarter) |
|---|---|---|---|---|---|
| 1 | 2022 | Electronics | 1200 | 0 | 1 |
| 2 | 2022 | Electronics | 1350 | 0 | 1 |
| 3 | 2022 | Electronics | 1420 | 1 | 1 |
| 4 | 2022 | Electronics | 1300 | 0 | 2 |
| 5 | 2022 | Home Appliances | 950 | 0 | 2 |
| 6 | 2022 | Home Appliances | 870 | 1 | 2 |
| 7 | 2022 | Clothing | 2100 | 0 | 3 |
| 8 | 2022 | Clothing | 2350 | 1 | 3 |
| 9 | 2022 | Personal Care | 1500 | 0 | 3 |
| 10 | 2022 | Personal Care | 1650 | 1 | 4 |
| 11 | 2022 | Electronics | 1400 | 1 | 4 |
| 12 | 2022 | Electronics | 1550 | 1 | 4 |
| 13 | 2023 | Electronics | 1600 | 0 | 1 |
| 14 | 2023 | Home Appliances | 980 | 0 | 1 |
| 15 | 2023 | Clothing | 2400 | 1 | 1 |
| Component | Specification |
|---|---|
| Python Version | 3.10 |
| Platform | Anaconda 3 |
| Development IDE | PyCharm 2025.3 |
| Hardware | Intel Core i7, 16 GB RAM, NVIDIA RTX 3060 |
| Dataset Split | Training: Testing = 7:3 (time-based split) |
| First 16 months for training, last 8 months for testing | |
| General Training Parameters | Batch size = 64, Epochs = 300 |
| Hyperparameter Search Ranges | Dropout rate: 0.1–0.5 |
| Learning rate: 0.0001–0.01 |
| Model | Learning Rate | Hidden Units | Attention Dim. | Dropout Rate | Other Key Parameters |
|---|---|---|---|---|---|
| ARIMA | – | – | – | – | Seasonal period = 12 |
| Standard LSTM | 0.001 | 128 | – | 0.2 | Layers = 2 |
| BiLSTM | 0.001 | 128 | – | 0.2 | Layers = 2 |
| BiLSTM + Attention | 0.001 | 128 | 64 | 0.2 | Layers = 2, Attention heads = 1 |
| TCN | 0.001 | 64 | – | 0.2 | Layers = 4, Kernel size = 3, Dilation = 2 |
| TFT | 0.001 | 128 | 64 | 0.2 | Encoder layers = 2, Attention heads = 2 |
| Proposed Enhanced BiLSTM | 0.001 | 128 | 64 | 0.2 | Layers = 2, Residual = 0.1, Dual-attention |
| Model | RMSE (Mean ± Std.) | MAE (Mean ± Std.) | MAPE (%) (Mean ± Std.) | (Mean ± Std.) |
|---|---|---|---|---|
| ARIMA | 21.5 ± 1.2 | 16.8 ± 0.9 | 14.9 ± 0.7 | 0.72 ± 0.03 |
| Standard LSTM | 16.8 ± 0.8 | 13.4 ± 0.6 | 11.7 ± 0.5 | 0.81 ± 0.02 |
| BiLSTM | 16.5 ± 0.7 | 12.6 ± 0.5 | 10.9 ± 0.4 | 0.84 ± 0.02 |
| BiLSTM + Attention | 15.2 ± 0.6 | 11.9 ± 0.4 | 10.1 ± 0.3 | 0.87 ± 0.01 |
| TCN | 15.0 ± 0.5 | 11.5 ± 0.4 | 9.8 ± 0.3 | 0.88 ± 0.01 |
| TFT | 14.8 ± 0.5 | 11.3 ± 0.3 | 9.5 ± 0.2 | 0.89 ± 0.01 |
| Ours | 13.2 ± 0.4 | 10.2 ± 0.3 | 8.7 ± 0.2 | 0.92 ± 0.01 |
| Model Variant | RMSE | MAE | MAPE (%) | |
|---|---|---|---|---|
| Base BiLSTM | 16.5 | 12.6 | 10.9 | 0.84 |
| BiLSTM + Residual | 15.8 | 12.0 | 10.3 | 0.86 |
| BiLSTM + Attention | 15.2 | 11.9 | 10.1 | 0.87 |
| Enhanced BiLSTM w/o BO | 14.0 | 10.9 | 9.2 | 0.90 |
| Proposed Enhanced BiLSTM | 13.2 | 10.2 | 8.7 | 0.92 |
| Hyperparameter | Learning Rate | Hidden Layer Size | Attention Layer Dimension | Dropout Rate | RMSE | MAE |
|---|---|---|---|---|---|---|
| Default Setting | 0.001 | 128 | 64 | 0.2 | 13.2 | 10.2 |
| Higher Learning Rate | 0.01 | 128 | 64 | 0.2 | 14.5 | 11.0 |
| Lower Learning Rate | 0.0001 | 128 | 64 | 0.2 | 15.0 | 11.3 |
| Larger Hidden Layer | 0.001 | 256 | 64 | 0.2 | 13.5 | 10.5 |
| Smaller Hidden Layer | 0.001 | 64 | 64 | 0.2 | 14.0 | 11.0 |
| Larger Attention Layer | 0.001 | 128 | 128 | 0.2 | 13.0 | 10.0 |
| Smaller Attention Layer | 0.001 | 128 | 32 | 0.2 | 14.0 | 11.2 |
| Higher Dropout Rate | 0.001 | 128 | 64 | 0.4 | 13.5 | 10.6 |
| Lower Dropout Rate | 0.001 | 128 | 64 | 0.1 | 14.0 | 11.0 |
| Model | Train Time/Epoch (s) | Inference Time/Sample (ms) | Parameters (M) |
|---|---|---|---|
| ARIMA | 0.5 | 0.01 | 0.001 |
| LSTM | 12 | 0.15 | 1.2 |
| BiLSTM | 14 | 0.18 | 2.3 |
| BiLSTM + Att | 18 | 0.22 | 2.5 |
| TCN | 20 | 0.25 | 2.8 |
| TFT | 22 | 0.30 | 3.1 |
| Enhanced BiLSTM | 19 | 0.23 | 2.6 |
| Sequence Length | RMSE | MAE | MAPE (%) |
|---|---|---|---|
| 3 months | 16.2 | 12.8 | 11.2 |
| 6 months | 13.2 | 10.2 | 8.7 |
| 12 months | 13.0 | 10.1 | 8.5 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hu, H.; Cai, J.; Xu, C. A Mathematical Framework for E-Commerce Sales Prediction Using Attention-Enhanced BiLSTM and Bayesian Optimization. Math. Comput. Appl. 2026, 31, 17. https://doi.org/10.3390/mca31010017
Hu H, Cai J, Xu C. A Mathematical Framework for E-Commerce Sales Prediction Using Attention-Enhanced BiLSTM and Bayesian Optimization. Mathematical and Computational Applications. 2026; 31(1):17. https://doi.org/10.3390/mca31010017
Chicago/Turabian StyleHu, Hao, Jinshun Cai, and Chenke Xu. 2026. "A Mathematical Framework for E-Commerce Sales Prediction Using Attention-Enhanced BiLSTM and Bayesian Optimization" Mathematical and Computational Applications 31, no. 1: 17. https://doi.org/10.3390/mca31010017
APA StyleHu, H., Cai, J., & Xu, C. (2026). A Mathematical Framework for E-Commerce Sales Prediction Using Attention-Enhanced BiLSTM and Bayesian Optimization. Mathematical and Computational Applications, 31(1), 17. https://doi.org/10.3390/mca31010017

