Enhancing Long-Term Forecasting Stability in Smart Grids: A Hybrid Mamba-LSTM-Attention Framework
Abstract
1. Introduction
2. Literature Review
2.1. Evolution of Long-Term Temporal Modeling
2.2. Multivariate Representation in LTSF
2.3. Non-Stationary Time Series Forecasting and Smart Grid Applications
3. Methodology
3.1. Overall Architecture
3.2. Reversible Instance Normalization (RevIN) for Distribution Shift
3.3. Mamba State Space Model
3.4. LSTM Unit
3.5. Multi-Head Attention Mechanism
3.6. Architecture Integration and Optimization Strategy
3.7. Synergistic Mechanism and Theoretical Justification
4. Experimental Results and Analysis
4.1. Benchmark Datasets and Data Preprocessing
4.2. Experimental Setup and Hyperparameters
4.3. Experimental Results and Discussion
4.4. Ablation Study
4.5. Qualitative Evaluation and Visualization
4.6. Computational Efficiency Analysis
4.7. Hyperparameter Sensitivity Analysis
5. Conclusions and Future Work
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| ARIMA | Autoregressive Integrated Moving Average |
| CI | Channel-Independence |
| CNN | Convolutional Neural Network |
| Coefficient of Variation across Horizons | |
| ETT | Electricity Transformer Temperature |
| FC | Fully Connected |
| GPU | Graphics Processing Unit |
| GRU | Gated Recurrent Unit |
| LSTM | Long Short-Term Memory |
| LTSF | Long-Term Time Series Forecasting |
| MAE | Mean Absolute Error |
| MLA | Mamba-LSTM-Attention |
| MLP | Multi-Layer Perceptron |
| MSE | Mean Squared Error |
| ODE | Ordinary Differential Equation |
| RevIN | Reversible Instance Normalization |
| RNN | Recurrent Neural Network |
| SARIMAX | Seasonal Autoregressive Integrated Moving Average with eXogenous variables |
| SSM | Structured State Space Model |
| TPE | Tree-structured Parzen Estimator |
| ZOH | Zero-Order Hold |
References
- Massaoudi, M.; Abu-Rub, H.; Refaat, S.S.; Chihi, I.; Oueslati, F.S. Deep learning in smart grid technology: A review of recent advancements and future prospects. IEEE Access 2021, 9, 54558–54578. [Google Scholar] [CrossRef]
- Liu, Y.; Tian, Y.; Zhao, Y.; Yu, H.; Xie, L.; Wang, Y.; Ye, Q.; Jiao, J.; Liu, Y. Vmamba: Visual state space model. Adv. Neural Inf. Process. Syst. 2024, 37, 103031–103063. [Google Scholar]
- Das, A.; Kong, W.; Leach, A.; Mathur, S.; Sen, R.; Yu, R. Long-term forecasting with Tide: Time-series dense encoder. arXiv 2023, arXiv:2304.08424. [Google Scholar]
- Kim, T.; Kim, J.; Tae, Y.; Park, C.; Choi, J.-H.; Choo, J. Reversible instance normalization for accurate time-series forecasting against distribution shift. In Proceedings of the International Conference on Learning Representations, Virtual Event, 3–7 May 2021. [Google Scholar]
- Bergstra, J.; Bardenet, R.; Bengio, Y.; Kégl, B. Algorithms for hyper-parameter optimization. Adv. Neural Inf. Process. Syst. 2011, 24, 2546–2554. [Google Scholar]
- Lim, B.; Arık, S.Ö.; Loeff, N.; Pfister, T. Temporal fusion transformers for interpretable multi-horizon time series forecasting. Int. J. Forecast. 2021, 37, 1748–1764. [Google Scholar] [CrossRef]
- Zhang, D.; Han, X.; Deng, C. Review on the research and practice of deep learning and reinforcement learning in smart grids. CSEE J. Power Energy Syst. 2018, 4, 362–370. [Google Scholar] [CrossRef]
- Ibrahim, R.A.; Hebala, A. A Feature-Enhanced Approach to Dissolved Gas Analysis for Power Transformer Health Prediction Through Interpretable Ensemble Learning and Multi-Model Evaluation. Technologies 2025, 14, 6. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. arXiv 2017, arXiv:1706.03762. [Google Scholar]
- Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; Zhang, W. Informer: Beyond efficient transformer for long sequence time-series forecasting. In Proceedings of the AAAI Conference on Artificial Intelligence, Palo Alto, CA, USA, 2–9 February 2021; pp. 11106–11115. [Google Scholar]
- Lai, G.; Chang, W.-C.; Yang, Y.; Liu, H. Modeling long-and short-term temporal patterns with deep neural networks. In Proceedings of the 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, Ann Arbor, MI, USA, 12 July 2018; pp. 95–104. [Google Scholar]
- Lim, B.; Zohren, S. Time-series forecasting with deep learning: A survey. Philos. Trans. R. Soc. A 2021, 379, 20200209. [Google Scholar] [CrossRef]
- Salinas, D.; Flunkert, V.; Gasthaus, J.; Januschowski, T. DeepAR: Probabilistic forecasting with autoregressive recurrent networks. Int. J. Forecast. 2020, 36, 1181–1191. [Google Scholar] [CrossRef]
- Wu, H.; Xu, J.; Wang, J.; Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Adv. Neural Inf. Process. Syst. 2021, 34, 22419–22430. [Google Scholar]
- Gu, A.; Goel, K.; Ré, C. Efficiently modeling long sequences with structured state spaces. arXiv 2021, arXiv:2111.00396. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar] [CrossRef]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision mamba: Efficient visual representation learning with bidirectional state space model. arXiv 2024, arXiv:2401.09417. [Google Scholar] [CrossRef]
- Yang, Y.J.; Xing, Z.H.; Yu, L.Q.; Fu, H.Z.; Huang, C.W.; Zhu, L. Vivim: A Video Vision Mamba for Ultrasound Video Segmentation. IEEE Trans. Circ. Syst. Vid. 2025, 35, 10293–10304. [Google Scholar] [CrossRef]
- Zhang, T.; Yuan, H.; Qi, L.; Zhang, J.; Zhou, Q.; Ji, S.; Yan, S.; Li, X. Point cloud mamba: Point cloud learning via state space model. In Proceedings of the AAAI Conference on Artificial Intelligence, Palo Alto, CA, USA, 25 February–4 March 2025; pp. 10121–10130. [Google Scholar]
- Gu, A.; Goel, K.; Gupta, A.; Ré, C. On the parameterization and initialization of diagonal state space models. Adv. Neural Inf. Process. Syst. 2022, 35, 35971–35983. [Google Scholar]
- Wu, L.; Pei, W.; Jiao, J.; Zhang, Q. Umambatsf: A u-shaped multi-scale long-term time series forecasting method using mamba. arXiv 2024, arXiv:2410.11278. [Google Scholar]
- Wardhana, A.K.; Riwanto, Y.; Rauf, B.W. Performance Comparison Analysis on Weather Prediction using LSTM and TKAN. Int. Things Artif. Intell. J. 2024, 4, 518–525. [Google Scholar] [CrossRef]
- Liang, A.; Jiang, X.; Sun, Y.; Shi, X.; Li, K. Bi-Mamba+: Bidirectional mamba for time series forecasting. arXiv 2024, arXiv:2404.15772. [Google Scholar]
- Ahamed, M.A.; Cheng, Q. Timemachine: A time series is worth 4 mambas for long-term forecasting. In Proceedings of the ECAI 2024: 27th European Conference on Artificial Intelligence, Santiago de Compostela, Spain, 19–24 October 2024; p. 1688. [Google Scholar]
- Smith, J.T.; Warrington, A.; Linderman, S.W. Simplified state space layers for sequence modeling. arXiv 2022, arXiv:2208.04933. [Google Scholar]
- Box, G.E.; Jenkins, G.M.; Reinsel, G.C.; Ljung, G.M. Time Series Analysis: Forecasting and Control; Holden-Day: San Francisco, CA, USA, 1970. [Google Scholar]
- Lei, C.F.; Chen, F.S.; Chu, C.W. Optimizing SARIMAX Model with Big Data To Predict Gaming Tourism Destination Demand. Mathematics 2025, 13, 3276. [Google Scholar] [CrossRef]
- Bahdanau, D.; Cho, K.; Bengio, Y. Neural machine translation by jointly learning to align and translate. arXiv 2014, arXiv:1409.0473. [Google Scholar]
- Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neur. Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef]
- Wang, J.; Bello, G.; Cheng, Y. Time-Mamba: Towards better multivariate time series forecasting with Mamba. arXiv 2024, arXiv:2403.11142. [Google Scholar]
- Zeng, A.; Chen, M.; Zhang, L.; Xu, Q. Are transformers effective for time series forecasting? In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; pp. 11121–11128. [Google Scholar]
- Nie, Y.; Nguyen, N.H.; Sinthong, P.; Kalagnanam, J. A time series is worth 64 words: Long-term forecasting with transformers. arXiv 2022, arXiv:2211.14730. [Google Scholar]
- Zhang, Y.; Yan, J. Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting. In Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Liu, Y.; Hu, T.; Zhang, H.; Wu, H.; Wang, S.; Ma, L.; Long, M. itransformer: Inverted transformers are effective for time series forecasting. arXiv 2023, arXiv:2310.06625. [Google Scholar]
- Wu, Z.; Pan, S.; Long, G.; Jiang, J.; Chang, X.; Zhang, C. Connecting the dots: Multivariate time series forecasting with graph neural networks. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, San Diego, CA, USA, 23–27 August 2020; pp. 753–763. [Google Scholar]
- Chen, S.-A.; Li, C.-L.; Yoder, N.; Arik, S.O.; Pfister, T. Tsmixer: An all-mlp architecture for time series forecasting. arXiv 2023, arXiv:2303.06053. [Google Scholar] [CrossRef]
- Wu, H.; Hu, T.; Liu, Y.; Zhou, H.; Wang, J.; Long, M. Timesnet: Temporal 2d-variation modeling for general time series analysis. arXiv 2022, arXiv:2210.02186. [Google Scholar]
- Liu, M.; Zeng, A.; Chen, M.; Xu, Z.; Lai, Q.; Ma, L.; Xu, Q. Scinet: Time series modeling and forecasting with sample convolution and interaction. Adv. Neural Inf. Process. Syst. 2022, 35, 5816–5828. [Google Scholar]
- Chen, K.; Chen, K.; Wang, Q.; He, Z.; Hu, J.; He, J. Short-term load forecasting with deep residual networks. IEEE Trans. Smart Grid 2018, 10, 3943–3952. [Google Scholar] [CrossRef]
- Liu, Y.; Wu, H.; Wang, J.; Long, M. Non-stationary transformers: Exploring the stationarity in time series forecasting. Adv. Neural Inf. Process. Syst. 2022, 35, 9881–9893. [Google Scholar]
- Liang, Z.; Chung, C.; Yang, H.; Liang, J.; Zhang, W.; Dong, H.; Zhu, J. A Heterogeneous Multiple-Experts Approach to Low-Frequency Nonintrusive Load Monitoring. IEEE Trans. Smart Grid 2025, 17, 746–765. [Google Scholar] [CrossRef]
- Ansari, A.F.; Stella, L.; Turkmen, C.; Zhang, X.; Mercado, P.; Shen, H.; Shchur, O.; Rangapuram, S.S.; Arango, S.P.; Kapoor, S. Chronos: Learning the language of time series. arXiv 2024, arXiv:2403.07815. [Google Scholar] [CrossRef]
- Garza, A.; Challu, C.; Mergenthaler-Canseco, M. TimeGPT-1. arXiv 2023, arXiv:2310.03589. [Google Scholar] [CrossRef]
- Huber, P.J. Robust Estimation of a Location Parameter; Springer: Berlin/Heidelberg, Germany, 1992; pp. 492–518. [Google Scholar]
- Susa, D.; Lehtonen, M.; Nordman, H. Dynamic thermal modelling of power transformers. IEEE Trans. Power Deliver. 2005, 20, 197–204. [Google Scholar] [CrossRef]
- Wang, Z.; Kong, F.; Feng, S.; Wang, M.; Yang, X.; Zhao, H.; Wang, D.; Zhang, Y. Is mamba effective for time series forecasting? Neurocomputing 2025, 619, 129178. [Google Scholar] [CrossRef]



| Model | Computational Complexity | Cross-Channel Interaction | Non-Stationary Handling | High-Frequency Refinement |
|---|---|---|---|---|
| Informer | O(L2) | Yes | No | No |
| PatchTST | O((L/P)2) | No (Channel-Independent) | Yes (RevIN) | No |
| S-Mamba | O(L) | Yes | No | No |
| MLA (Proposed) | O(L) | Yes | Yes (RevIN) | Yes (LSTM) |
| Dataset | Domain | Time Granularity | Dimension (D) | Target Variable | Sequence Characteristics Analysis |
|---|---|---|---|---|---|
| ETTh1 | Electricity | 1 h | 7 | Oil Temperature | Moderate periodicity; serves as the baseline reference point. |
| ETTh2 | Electricity | 1 h | 7 | Oil Temperature | Contains significant structural mutations. |
| ETTm1 | Electricity | 15 min | 7 | Oil Temperature | High-frequency sampling; complex short-term fluctuations. |
| ETTm2 | Electricity | 15 min | 7 | Oil Temperature | High-frequency sampling; complex short-term non-linear fluctuations. |
| Electricity | Electricity | 1 h | 321 | Electricity Consumption | True consumer-driven non-stationarity; directly represents real-world smart grid load dynamics. |
| No. | Configuration | Value | Description (Corresponding Code Elements) |
|---|---|---|---|
| 1 | Historical look-back window | 96 | Length of historical observation window (seq_len) and forecasting horizon (pred_len). |
| 2 | Prediction horizon | {96, 192, 336, 720} | Length of historical observation window (seq_len) and forecasting horizon (pred_len). |
| 3 | Internal Feature Embedding Dim | 128 | Dimension of internal feature embedding. |
| 4 | Mamba State Dim | 64 | State space dimension of the SSM. |
| 5 | LSTM Layers | 1 | Number of stacked LSTM layers. |
| 6 | Dropout rate | 0.3 | Unified dropout rate applied across LSTM, Attention, and FC layers. |
| 7 | Learning Rate | 2.84 × 10−4 | Learning rate and L2 weight decay for the AdamW optimizer. |
| 8 | Weight Decay | 3.74 × 10−6 | Learning rate and L2 weight decay for the AdamW optimizer. |
| 9 | Attention Heads | 8 | Number of parallel attention heads for cross-window interaction. |
| 10 | Optimizer | AdamW | Optimization algorithm for network weight updates. |
| 11 | Loss function | Huber Loss (δ = 1.0) | Objective function for robust error optimization. |
| 12 | Batch Size | 64 | Number of sequence samples processed in one training iteration. |
| 13 | Epoch Budget | 15 | Maximum limit of complete passes through the training dataset. |
| Dataset | Model | T = 96 MSE/MAE/RMSE | T = 192 MSE/MAE/RMSE | T = 336 MSE/MAE/RMSE | T = 720 MSE/MAE/RMSE | |
|---|---|---|---|---|---|---|
| ETTh1 | Autoformer | 0.449/0.459/0.670 | 0.500/0.482/0.707 | 0.521/0.496/0.722 | 0.542/0.524/0.736 | 7.93% |
| PatchTST | 0.419/0.424/0.647 | 0.469/0.454/0.685 | 0.501/0.466/0.708 | 0.500/0.488/0.707 | 8.15% | |
| S-Mamba | 0.386/0.405/0.621 | 0.443/0.437/0.666 | 0.499/0.468/0.706 | 0.502/0.499/0.709 | 11.99% | |
| iTransformer | 0.386/0.405/0.621 | 0.441/0.436/0.664 | 0.487/0.458/0.698 | 0.503/0.491/0.709 | 11.57% | |
| MLA | 0.566/0.538/0.752 | 0.601/0.564/0.775 | 0.640/0.589/0.800 | 0.764/0.660/0.874 | 13.43% | |
| ETTh2 | Autoformer | 0.346/0.388/0.588 | 0.456/0.452/0.675 | 0.482/0.486/0.694 | 0.515/0.539/0.718 | 16.29% |
| PatchTST | 0.342/0.384/0.585 | 0.388/0.400/0.623 | 0.426/0.433/0.653 | 0.431/0.446/0.657 | 10.40% | |
| S-Mamba | 0.296/0.348/0.544 | 0.376/0.396/0.613 | 0.424/0.431/0.651 | 0.426/0.444/0.653 | 16.00% | |
| iTransformer | 0.297/0.349/0.545 | 0.380/0.400/0.616 | 0.428/0.432/0.654 | 0.427/0.445/0.653 | 16.07% | |
| MLA | 0.210/0.317/0.458 | 0.237/0.343/0.487 | 0.253/0.353/0.503 | 0.304/0.387/0.551 | 15.75% | |
| ETTm1 | Autoformer | 0.505/0.475/0.711 | 0.553/0.496/0.744 | 0.621/0.537/0.788 | 0.671/0.561/0.819 | 12.47% |
| PatchTST | 0.329/0.367/0.574 | 0.367/0.385/0.606 | 0.399/0.410/0.632 | 0.454/0.439/0.674 | 13.66% | |
| S-Mamba | 0.333/0.368/0.577 | 0.376/0.390/0.613 | 0.408/0.413/0.639 | 0.475/0.448/0.689 | 15.03% | |
| iTransformer | 0.334/0.368/0.578 | 0.377/0.391/0.614 | 0.426/0.420/0.653 | 0.491/0.459/0.701 | 16.57% | |
| MLA | 0.474/0.471/0.688 | 0.533/0.498/0.730 | 0.591/0.537/0.769 | 0.637/0.563/0.798 | 12.66% | |
| ETTm2 | Autoformer | 0.255/0.339/0.505 | 0.348/0.403/0.590 | 0.509/0.472/0.713 | 0.433/0.432/0.658 | 28.34% |
| PatchTST | 0.175/0.259/0.418 | 0.241/0.302/0.491 | 0.305/0.343/0.552 | 0.402/0.400/0.634 | 34.16% | |
| S-Mamba | 0.179/0.263/0.423 | 0.250/0.309/0.500 | 0.312/0.349/0.558 | 0.411/0.406/0.642 | 34.44% | |
| iTransformer | 0.180/0.264 | 0.250/0.309 | 0.311/0.348 | 0.412/0.407/0.642 | 34.12% | |
| MLA | 0.128/0.242/0.358 | 0.160/0.272/0.400 | 0.204/0.305/0.452 | 0.260/0.349/0.510 | 30.44% | |
| Electricity | Autoformer | 0.201/0.317/0.448 | 0.222/0.334/0.471 | 0.231/0.338/0.481 | 0.254/0.361/0.504 | 9.67% |
| PatchTST | 0.181/0.270/0.425 | 0.188/0.274/0.434 | 0.204/0.293/0.452 | 0.246/0.324/0.496 | 14.23% | |
| S-Mamba | 0.139/0.235/0.373 | 0.159/0.255/0.399 | 0.176/0.272/0.420 | 0.204/0.298/0.452 | 16.24% | |
| iTransformer | 0.148/0.240/0.385 | 0.162/0.253/0.402 | 0.178/0.269/0.422 | 0.225/0.317/0.474 | 18.79% | |
| MLA | 0.187/0.291/0.432 | 0.203/0.304/0.451 | 0.212/0.313/0.460 | 0.230/0.326/0.480 | 8.63% |
| Model Variant | Component Removed | ETTh2 | ETTh1 | ||||
|---|---|---|---|---|---|---|---|
| MSE | MAE | ΔMSE (vs. Base) | MSE | MAE | ΔMSE (vs. Base) | ||
| MLA(Proposed) | None | 0.210 | 0.317 | — | 0.566 | 0.538 | — |
| MLA | RevIN | 0.431 | 0.495 | +105.24% | 0.716 | 0.616 | +26.50% |
| LSTM-Attention | Mamba | 0.218 | 0.321 | +3.81% | 0.571 | 0.533 | +0.88% |
| Mamba-LSTM | Attention | 0.223 | 0.329 | +6.19% | 0.633 | 0.552 | +11.84% |
| Mamba-Attention | LSTM | 0.227 | 0.325 | +8.10% | 0.582 | 0.546 | +2.83% |
| Model | Architecture Type | Training Time (s/Epoch) | Peak Memory (MB) |
|---|---|---|---|
| Autoformer | Transformer-based | 4.31 | 233.28 |
| PatchTST | Transformer-based | 3.49 | 1602.87 |
| S-Mamba | State-Space Model | 2.69 | 147.11 |
| iTransformer | Transformer-based | 6.16 | 161.91 |
| MLA (Ours) | Hybrid (SSM + LSTM) | 1.69 | 175.99 |
| Hyperparameter | Range Tested | Optimal Value | Sensitivity | Impact on MSE |
|---|---|---|---|---|
| Hidden Layers (LSTM) | {1, 2, 3} | 1 | High | Sharp degradation from 8.59 (1 layer) to 9.68 (3 layers) |
| Learning Rate | {1.1 × 10−5, 8.5 × 10−4} | 2.84 × 10−4 | High | Suboptimal convergence at boundaries (e.g., MSE 9.13 at 1.1 × 10−5) |
| Hidden Dimension | {128, 256, 512} | 128 | Low | High stability across scales (MSE 8.59 at 128 vs. 8.66 at 512) |
| Dropout Rate | {0.1, 0.2, 0.3, 0.4} | 0.3 | Low | Marginal MSE variance (8.59 to 8.75) across the full range |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Chen, F.; Lei, C.F.; Guo, T.; Chu, C. Enhancing Long-Term Forecasting Stability in Smart Grids: A Hybrid Mamba-LSTM-Attention Framework. Energies 2026, 19, 1855. https://doi.org/10.3390/en19081855
Chen F, Lei CF, Guo T, Chu C. Enhancing Long-Term Forecasting Stability in Smart Grids: A Hybrid Mamba-LSTM-Attention Framework. Energies. 2026; 19(8):1855. https://doi.org/10.3390/en19081855
Chicago/Turabian StyleChen, Fusheng, Chong Fo Lei, Te Guo, and Chiawei Chu. 2026. "Enhancing Long-Term Forecasting Stability in Smart Grids: A Hybrid Mamba-LSTM-Attention Framework" Energies 19, no. 8: 1855. https://doi.org/10.3390/en19081855
APA StyleChen, F., Lei, C. F., Guo, T., & Chu, C. (2026). Enhancing Long-Term Forecasting Stability in Smart Grids: A Hybrid Mamba-LSTM-Attention Framework. Energies, 19(8), 1855. https://doi.org/10.3390/en19081855

