A Multi-Scale Convolutional Neural Network with Residual Blocks and LSTM for Multi-Step Forecasting of Electricity Load
Abstract
1. Introduction
2. Related Work
3. Methodology
3.1. Problem Definition
3.2. Data Preparation and Sample Construction
3.3. Tensor Alignment and Consistency Verification
3.4. Proposed MSCNN-ResLSTM Model
3.4.1. Overall Architecture
3.4.2. Multi-Scale Convolution Module
3.4.3. Residual Convolution Block
3.4.4. LSTM Prediction Layer
3.5. Recursive Multi-Step Forecasting Strategy
3.6. Training Settings and Evaluation Metrics
4. Experimental Results and Analysis
4.1. Experimental Configuration and Baseline Models
4.2. Overall Performance Comparison
4.3. Horizon-Wise Analysis of Recursive Forecasting
4.4. Performance Breakdown Across Seasons and Day Types
4.5. Hyperparameter Sensitivity Analysis
4.6. Interpretability and Feature Importance Analysis
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Dong, Q.; Huang, R.; Cui, C.; Towey, D.; Zhou, L.; Tian, J.; Wang, J. Short-Term Electricity-Load Forecasting by deep learning: A comprehensive survey. Eng. Appl. Artif. Intell. 2025, 154, 110980. [Google Scholar] [CrossRef] [Scilit]
- Eren, Y.; Küçükdemiral, İ. A comprehensive review on deep learning approaches for short-term load forecasting. Renew. Sustain. Energy Rev. 2024, 189, 114031. [Google Scholar] [CrossRef] [Scilit]
- Biswal, B.; Deb, S.; Datta, S.; Ustun, T.S.; Cali, U. Review on smart grid load forecasting for smart energy management using machine learning and deep learning techniques. Energy Rep. 2024, 12, 3654–3670. [Google Scholar] [CrossRef] [Scilit]
- Ferreira, A.B.A.; Leite, J.B.; Salvadeo, D.H.P. Power substation load forecasting using interpretable transformer-based temporal fusion neural networks. Electr. Power Syst. Res. 2025, 238, 111169. [Google Scholar] [CrossRef] [Scilit]
- Fekri, M.N.; Patel, H.; Grolinger, K.; Sharma, V. Deep learning for load forecasting with smart meter data: Online Adaptive Recurrent Neural Network. Appl. Energy 2021, 282, 116177. [Google Scholar] [CrossRef] [Scilit]
- Haque, A.; Rahman, S. Short-term electrical load forecasting through heuristic configuration of regularized deep neural network. Appl. Soft Comput. 2022, 122, 108877. [Google Scholar] [CrossRef] [Scilit]
- Nabavi, S.A.; Mohammadi, S.; Motlagh, N.H.; Tarkoma, S.; Geyer, P. Deep learning modeling in electricity load forecasting: Improved accuracy by combining DWT and LSTM. Energy Rep. 2024, 12, 2873–2900. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Li, X.; Shi, Y.; Jiang, W.; Song, Q.; Li, X. Load forecasting method based on CNN and extended LSTM. Energy Rep. 2024, 12, 2452–2461. [Google Scholar] [CrossRef] [Scilit]
- Zhao, X.; Peng, H.; Zhang, L.; Ma, H. Research on a Short-Term Power Load Forecasting Method Based on a Three-Channel LSTM-CNN. Electronics 2025, 14, 2262. [Google Scholar] [CrossRef] [Scilit]
- Guo, W.; Liu, S.; Weng, L.; Liang, X. Power Grid Load Forecasting Using a CNN-LSTM Network Based on a Multi-Modal Attention Mechanism. Appl. Sci. 2025, 15, 2435. [Google Scholar] [CrossRef] [Scilit]
- Sheng, Z.; An, Z.; Wang, H.; Chen, G.; Tian, K. Residual LSTM based short-term load forecasting. Appl. Soft Comput. 2023, 144, 110461. [Google Scholar] [CrossRef] [Scilit]
- Li, C.; Shi, J. A novel CNN-LSTM-based forecasting model for household electricity load by merging mode decomposition, self-attention and autoencoder. Energy 2025, 330, 136883. [Google Scholar]
- Hua, Q.; Fan, Z.; Mu, W.; Cui, J.; Xing, R.; Liu, H.; Gao, J. A Short-Term Power Load Forecasting Method Using CNN-GRU with an Attention Mechanism. Energies 2025, 18, 5124. [Google Scholar] [CrossRef] [Scilit]
- Ahmad, A.; Xiao, X.; Mo, H.; Dong, D. TFTformer: A novel transformer based model for short-term load forecasting. Int. J. Electr. Power Energy Syst. 2025, 166, 110549. [Google Scholar] [CrossRef] [Scilit]
- Zhu, L.; Gao, J.; Zhu, C.; Deng, F. Short-term power load forecasting based on spatial-temporal dynamic graph and multi-scale Transformer. J. Comput. Des. Eng. 2025, 12, 92–111. [Google Scholar] [CrossRef] [Scilit]
- Rafi, S.H.; Mahdi, M.M. A short-term load forecasting technique using extreme gradient boosting algorithm. In Proceedings of the 2021 IEEE PES Innovative Smart Grid Technologies-Asia (ISGT Asia), Brisbane, Australia, 5–8 December 2021; pp. 1–5. [Google Scholar]
- Zeng, S.; Liu, C.; Zhang, H.; Zhang, B.; Zhao, Y. Short-term load forecasting in power systems based on the Prophet–BO–XGBoost model. Energies 2025, 18, 227. [Google Scholar] [CrossRef] [Scilit]
- Lara-Benítez, P.; Carranza-García, M.; Luna-Romera, J.M.; Riquelme, J.C. Temporal convolutional networks applied to energy-related time series forecasting. Appl. Sci. 2020, 10, 2322. [Google Scholar] [CrossRef] [Scilit]
- Feng, Y.; Zhu, J.; Qiu, P.; Zhang, X.; Shuai, C. Short-term power load forecasting based on TCN-BiLSTM-attention and multi-feature fusion. Arab. J. Sci. Eng. 2025, 50, 5475–5486. [Google Scholar]
- Chen, Y.; Céspedes, N.; Barnaghi, P. A closer look at transformers for time series forecasting: Understanding why they work and where they struggle. In Proceedings of the Forty-Second International Conference on Machine Learning, Vancouver, BC, Canada, 13–19 July 2025. [Google Scholar]
- Debnath, S.; Mia, M.U.; Abubakkar, M.; Islam, M.R.; Mridul, M.S.I.; Biswas, A.K. Hybrid Multi-Scale Deep Learning Enhanced Electricity Load Forecasting Using Attention-Based Convolutional Neural Network and LSTM Model. IEEE Access 2026, 14, 13423–13444. [Google Scholar] [CrossRef] [Scilit]
- Noorizadegan, A. Partition-of-Unity Gaussian Kolmogorov-Arnold Networks. arXiv 2026, arXiv:2604.23599. [Google Scholar]





| Feature Group | Specific Variables Included | Dimensionality |
|---|---|---|
| Base Load | Historical hourly electricity consumption at timestamp t | 1 |
| Meteorological | Dry-bulb temperature, dew-point temperature, wind speed, barometric pressure | 4 |
| Precipitation Dynamics | Rain missingness indicator, 6 h and 24 h rolling cumulative rainfall sums | 3 |
| Temporal Profiles | Calendar indicators (month, day of week, hour), weekend and holiday flags, cyclic components (sine/cosine of hour and day), and grid peak-period binary flags (morning, evening, night) | 12 |
| Thermal Nonlinearity | Absolute deviation from base temperature (18 °C), cooling and heating degree metrics, relative humidity, apparent temperature, and apparent thermal indices | 7 |
| Historical Analytics | Multi-interval chronological lag terms (24 h, 48 h, 168 h) and rolling statistical aggregates (24 h mean/standard deviation, 168 h mean) | 6 |
| Total Dimension (d) | 33 |
| Baseline Model | Detailed Architectural and Hyperparameter Specifications |
|---|---|
| XGBoost | Maximum tree depth = 2, learning rate = 0.1, number of estimators = 50, min child weight = 5, gamma = 0.5, subsample ratio = 0.4, colsample bytree = 0.4, reg alpha = 1.0, reg lambda = 2.0, early stopping rounds = 5. |
| Transformer | 2 Encoder layers, 4 attention heads, feature projection dimension () = 128, attention key/query/value dimension = 32, feed-forward network (FFN) dimension = 256, encoder dropout rate = 0.15, decoding intermediate dense layer = 64 units with a decoding dropout rate of 0.15. |
| LSTM | Single LSTM layer with 64 hidden units and a dropout rate of 0.2, intermediate dense layer = 32 units with ReLU activation, final linear output layer. |
| Direct LSTM (Seq2Seq) | Encoder LSTM layer (64 hidden units, dropout rate = 0.2), Decoder LSTM layer (64 hidden units, dropout rate = 0.2), fully connected dense projection layer mapped directly to a 24-step multi-horizon output vector. |
| CNN-LSTM | 1D Convolution layer (32 filters, kernel size = 3, padding = “same”, ReLU activation) followed by an LSTM layer (64 hidden units, dropout rate = 0.2), an intermediate dense layer (32 units, ReLU), and a linear output layer. |
| MSCNN-LSTM | Three parallel 1D Convolution branches (32 filters each, kernel sizes = 3, 5, 7, padding = “same”, ReLU activation) concatenated and aggregated via a fusion 1D Convolution layer (64 filters, kernel size = 3, padding = “same”, ReLU activation), followed by an LSTM layer (64 hidden units, dropout rate = 0.2), an intermediate dense layer (32 units, ReLU), and a linear output layer. |
| TCN | 4 Stacked residual blocks with dilated causal convolutions, filter count = 64, dilation factors = [1, 2, 4, 8], kernel size = 3, convolution dropout rate = 0.2, GlobalAveragePooling1D layer, followed by a dense layer of 32 units with ReLU activation. |
| Model Architecture | MAE (MW) | MSE () | RMSE (MW) | MAPE (%) |
|---|---|---|---|---|
| XGBoost * | 251.37 | 117,594.05 | 342.93 | 4.30 |
| Transformer | 212.05 ± 3.42 | 120,504.43 ± 1520.50 | 347.14 ± 2.19 | 3.38 ± 0.08 |
| LSTM | 164.13 ± 1.95 | 55,314.75 ± 590.30 | 235.19 ± 1.25 | 2.65 ± 0.05 |
| Direct LSTM (Seq2Seq) # | 153.45 ± 1.35 | 45,890.12 ± 410.50 | 214.22 ± 1.02 | 2.54 ± 0.04 |
| CNN-LSTM | 152.31 ± 1.62 | 44,902.28 ± 480.15 | 211.92 ± 1.13 | 2.50 ± 0.04 |
| MSCNN-LSTM | 151.03 ± 1.45 | 41,111.13 ± 425.60 | 202.76 ± 1.05 | 2.52 ± 0.04 |
| TCN | 148.72 ± 1.12 | 38,762.94 ± 310.45 | 196.88 ± 0.79 | 2.65 ± 0.03 |
| MSCNN-ResLSTM (Proposed) | 122.71 ± 0.65 | 26,130.48 ± 185.20 | 161.65 ± 0.57 | 2.16 ± 0.02 |
| Evaluation Subset | MAE (MW) | RMSE (MW) | MAPE (%) |
|---|---|---|---|
| Workday | 125.35 | 165.23 | 2.16 |
| Weekend | 118.81 | 155.42 | 2.20 |
| Holiday | 166.56 | 206.47 | 3.02 |
| Summer (June–August) | 162.98 | 210.29 | 2.40 |
| Winter (December–February) | 111.94 | 142.41 | 2.02 |
| Input Window () | Kernel Combination (k) | MAE (MW) | RMSE (MW) | MAPE (%) |
|---|---|---|---|---|
| 120 | 125.84 | 165.20 | 2.21 | |
| 168 (Proposed) | (Proposed) | 122.71 | 161.65 | 2.16 |
| 168 | 123.90 | 163.12 | 2.18 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhang, Y.; Zhao, Y.; Meng, Y.; Li, J.; Zhang, T.; Zhang, Y. A Multi-Scale Convolutional Neural Network with Residual Blocks and LSTM for Multi-Step Forecasting of Electricity Load. Computers 2026, 15, 457. https://doi.org/10.3390/computers15070457
Zhang Y, Zhao Y, Meng Y, Li J, Zhang T, Zhang Y. A Multi-Scale Convolutional Neural Network with Residual Blocks and LSTM for Multi-Step Forecasting of Electricity Load. Computers. 2026; 15(7):457. https://doi.org/10.3390/computers15070457
Chicago/Turabian StyleZhang, Yuhang, Yiting Zhao, Yujing Meng, Jingqi Li, Tianze Zhang, and Ying Zhang. 2026. "A Multi-Scale Convolutional Neural Network with Residual Blocks and LSTM for Multi-Step Forecasting of Electricity Load" Computers 15, no. 7: 457. https://doi.org/10.3390/computers15070457
APA StyleZhang, Y., Zhao, Y., Meng, Y., Li, J., Zhang, T., & Zhang, Y. (2026). A Multi-Scale Convolutional Neural Network with Residual Blocks and LSTM for Multi-Step Forecasting of Electricity Load. Computers, 15(7), 457. https://doi.org/10.3390/computers15070457

