Single-Attention Large Language Model for Efficient Multi-Regional Electricity Demand and Generation Forecasting
Abstract
1. Background
1.1. Electricity Forecasting
1.2. LLMs for Forecasting
2. Contribution
- 1.
- Fine-tuning difficulty: Since a lot of optimization is done for pre-trained LLMs, naive fine-tuning often results in negligible improvements or even leads to overfitting issues. In addition, because LLMs have highly complex architectures, it is difficult to determine which components need to be fine-tuned [6].
- 2.
- Prompt design limitations: Prompt engineering has been used as a means of improving LLM performance for natural language processing. However, prompt engineering for time-series forecasting is challenging due to the complex temporal dependencies and the large number of features.
- 3.
- Regional variability: It is observed that there are differences between countries with respect to geography when it comes to climate, electricity demand, and renewable energy generation. These differences make it difficult to identify universal models, and only a limited number of studies are conducted.
Key Contributions
- We introduce the SA-LLM, a new time-series forecasting framework that employs a single-attention mechanism to learn from target variables, yet requires manual prompt tuning.
- We provide a Graph Neural Network (GNN) module to capture spatial and inter-regional relationships, improving forecasting accuracy.
- We evaluate the SA-LLM across major U.S. electricity markets, showing it outperforms LSTM and previous LLM models, reducing the MAE by 22.5%, memory usage by 52.1%, and training time by 38.4%.
- We evaluate the SA-LLM in zero-shot scenarios and find that it reduces the MAE by 18.2% on unseen regions, demonstrating strong generalization across different locations and times.
- We offer a unified approach to spatio-temporal electricity forecasting that combines LLM-based feature extraction with optional graph-based spatial modeling.
3. Methodology
3.1. Mainstream Strategies for Time-Series Forecasting
- Model-Based Optimization (Attention-Enhanced LLM) [39]
- 2.
- Model-Free Reinforcement Learning (TimeLLM) [41]
3.2. Motivation for a Unified Attention Mechanism
3.3. Proposed Single-Attention LLM
3.3.1. Word Projection Layer
3.3.2. Single Unified Attention Module
3.3.3. Frozen Pre-Trained LLM Backbone
3.3.4. Auxiliary Feature Extractor
3.3.5. Frozen LLM Backbone: Benefits and Limitations
3.3.6. Attention-Based Feature Fusion and Feedforward Refinement:
3.3.7. Output Projection Layer
3.4. Attention Mechanism and Feature Interactions
3.5. Clarification of Model Innovation and Relation to Existing LLMs
3.6. Spatio-Temporal Graph Neural Network Module
3.6.1. Graph Construction and Laplacian Operator
3.6.2. Spectral Graph Convolution
3.6.3. Temporal Dynamics Integration
3.6.4. Joint Spatio-Temporal Representation
3.6.5. Training Objective with Regularization
3.6.6. Optional Module Mechanism
4. Proposed SA-LLM Architecture
4.1. Model Setup for SA-LLM
LSTM Baseline
4.2. Sensitivity Analysis of SA-LLM and MultiAttLLM
4.3. Hyperparameter Engineering
5. Implementation and Setup
5.1. Dataset Description
Data Splitting Methodology
5.2. Hardware, Software, and LLM Parameters
5.3. Impact of Prompts on Overall Performance
5.4. Analysis of Model Behavior Under Varying Operating Conditions
5.5. Hyperparameter Optimization and Training
Comprehensive Ablation Study on Spectral and Spatio-Temporal Components
| Model Configuration | MAE ↓ | RMSE ↓ | MAE (%) | Performance Shift |
|---|---|---|---|---|
| Full SA-LLM + Spectral GCN (All Components) | 4.68 | 6.12 | – | Baseline |
| w/o Spectral Graph Propagation () | 6.10 ↑ | 8.45 ↑ | +30.3% | Severe Degradation |
| w/o Temporal Gating Mechanism | 5.72 ↑ | 7.64 ↑ | +22.2% | Major Degradation ↑ |
| w/o Laplacian Normalization | 5.21 ↑ | 6.89 ↑ | +11.3% | Stability Loss ↑ |
| w/o Chebyshev Polynomial Filtering () | 5.34 ↑ | 7.02 ↑ | +14.1% | Reduced Spatial Expressiveness |
| Fully Connected Graph (No Gaussian Kernel) | 5.56 ↑ | 7.31 ↑ | +18.8% | Structural Over-Smoothing ↑ |
| w/o Laplacian Regularization () | 5.11 ↑ | 6.89 ↑ | +9.2% | Moderate Structural Drift |
| Model | Full Benchmark Evaluation | Short-Horizon & Resource Evaluation | |||||
|---|---|---|---|---|---|---|---|
| MAE | RMSE | Params (M) | MAE (MWh) | Training Time (h) | Memory (GB) | ||
| LSTM | 236 | 325 | 0.70 | 61.2 | 7.80 | 0.12 | 2.0 |
| DLinear | 215 | 305 | 0.74 | 61.5 | 8.20 | 0.10 | 1.8 |
| Informer | 210 | 298 | 0.75 | 62.1 | – | – | – |
| Autoformer | 208 | 295 | 0.76 | 62.3 | – | – | – |
| iTransformer | 207 | 293 | 0.76 | 62.0 | – | – | – |
| TimesNet | 206 | 290 | 0.77 | 62.4 | – | – | – |
| MultiAttLLM | 205 | 288 | 0.76 | 62.4 | 5.50 | 0.32 | 4.0 |
| SA-LLM | 197 | 274 | 0.79 | 63.0 | 4.68 | 0.18 | 2.5 |
5.6. Comparative Results
5.7. Electricity Demand Forecasting Performance Analysis
5.8. Performance Across Different Forecasting Horizons
5.9. Zero-Shot Learning Analysis for Multi-Regional Forecasting
5.10. Confusion Matrix-Based Diagnostic Evaluation
6. Performance of SA-LLM (Zero-Shot Learning)
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Appendix A. Multi-Attention Large Language Model (MultiAttLLM)
Appendix A.1. Word Projection
Appendix A.2. Cross-Attention
Appendix A.3. Frozen LLM
Appendix A.4. Feature Extraction and Self-Attention
Appendix A.5. Output Projection
Appendix A.6. Discussion
- Sequential processing introduces high computational costs.
- The frozen LLM may underutilize numerical time-series information.
- Complexity grows with the number of auxiliary features, potentially reducing scalability.
- Some redundant computations over non-target variables may persist.
Appendix B. Prompt Engineering
Summary Table of Prompts
| Prompt | Task Description | Dataset Info | Input Statistics |
|---|---|---|---|
| Prompt1 | ✓ | – | – |
| Prompt2 | ✓ | ✓ | – |
| Prompt3 | ✓ | ✓ | ✓ |
References
- Oswald, Y.; Owen, A.; Steinberger, J.K. Large Inequality in International and Intranational Energy Footprints between Income Groups and across Consumption Categories. Nat. Energy 2020, 5, 231–239. [Google Scholar] [CrossRef] [Scilit]
- Cang, D.; Chen, C.; Chen, Q.; Sui, L.; Cui, C. Does New Energy Consumption Conduce to Controlling Fossil Energy Consumption and Carbon Emissions? Evidence from China. Resour. Policy 2021, 74, 102427. [Google Scholar] [CrossRef] [Scilit]
- Khalili, R.; Khaledi, A.; Marzband, M.; Nematollahi, A.F.; Vahidi, B.; Siano, P. Robust Multi-Objective Optimization for the Iranian Electricity Market Considering Green Hydrogen and Analyzing the Performance of Different Demand Response Programs. Appl. Energy 2023, 334, 120737. [Google Scholar] [CrossRef] [Scilit]
- Danish, M.S.S.; Senjyu, T.; Funabashi, T.; Ahmadi, M.; Ibrahimi, A.M.; Ohta, R.; Howlader, H.O.R.; Zaheb, H.; Sabory, N.R.; Sediqi, M.M. A Sustainable Microgrid: A Sustainability and Management-Oriented Approach. Energy Procedia 2019, 159, 160–167. [Google Scholar] [CrossRef] [Scilit]
- Hsiao, C.-T.; Liu, C.-S.; Chang, D.-S.; Chen, C.-C. Dynamic Modeling of the Policy Effect and Development of Electric Power Systems: A Case in Taiwan. Energy Policy 2018, 122, 377–387. [Google Scholar] [CrossRef] [Scilit]
- Ghiasi, M.; Niknam, T.; Wang, Z.; Mehrandezh, M.; Dehghani, M.; Ghadimi, N. A Comprehensive Review of Cyber-Attacks and Defense Mechanisms for Improving Security in Smart Grid Energy Systems: Past, Present and Future. Electr. Power Syst. Res. 2023, 215, 108975. [Google Scholar] [CrossRef] [Scilit]
- Hong, T.; Pinson, P.; Wang, Y.; Weron, R.; Yang, D.; Zareipour, H. Energy Forecasting: A Review and Outlook. IEEE Open Access J. Power Energy 2020, 7, 376–388. [Google Scholar] [CrossRef] [Scilit]
- Akbary, P.; Ghiasi, M.; Pourkheranjani, M.R.R.; Alipour, H.; Ghadimi, N. Extracting Appropriate Nodal Marginal Prices for All Types of Committed Reserve. Comput. Econ. 2019, 53, 1–26. [Google Scholar] [CrossRef] [Scilit]
- Sharma, M.; Mittal, N.; Mishra, A.; Gupta, A. Survey of Electricity Demand Forecasting and Demand Side Management Techniques in Different Sectors to Identify Scope for Improvement. Smart Grids Sustain. Energy 2023, 8, 9. [Google Scholar] [CrossRef] [Scilit]
- Ghiasi, M.; Wang, Z.; Mehrandezh, M.; Jalilian, S.; Ghadimi, N. Evolution of Smart Grids towards the Internet of Energy: Concept and Essential Components for Deep Decarbonisation. IET Smart Grid 2023, 6, 86–102. [Google Scholar] [CrossRef] [Scilit]
- Aderibigbe, A.O.; Ani, E.C.; Ohenhen, P.E.; Ohalete, N.C.; Daraojimba, D.O. Enhancing Energy Efficiency with AI: A Review of Machine Learning Models in Electricity Demand Forecasting. Eng. Sci. Technol. J. 2023, 4, 341–356. [Google Scholar] [CrossRef] [Scilit]
- Román-Portabales, A.; López-Nores, M.; Pazos-Arias, J.J. Systematic Review of Electricity Demand Forecast Using ANN-Based Machine Learning Algorithms. Sensors 2021, 21, 4544. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sultana, N.; Hossain, S.Z.; Almuhaini, S.H.; Düştegör, D. Bayesian Optimization Algorithm-Based Statistical and Machine Learning Approaches for Forecasting Short-Term Electricity Demand. Energies 2022, 15, 3425. [Google Scholar] [CrossRef] [Scilit]
- Velasquez, C.E.; Zocatelli, M.; Estanislau, F.B.; Castro, V.F. Analysis of Time Series Models for Brazilian Electricity Demand Forecasting. Energy 2022, 247, 123483. [Google Scholar] [CrossRef] [Scilit]
- Jiang, W.; Wang, X.; Huang, H.; Zhang, D.; Ghadimi, N. Optimal Economic Scheduling of Microgrids Considering Renewable Energy Sources Based on Energy Hub Model Using Demand Response and Improved Water Wave Optimization Algorithm. J. Energy Storage 2022, 55, 105311. [Google Scholar] [CrossRef] [Scilit]
- Torres, J.F.; Martínez-Álvarez, F.; Troncoso, A. A Deep LSTM Network for the Spanish Electricity Consumption Forecasting. Neural Comput. Appl. 2022, 34, 10533–10545. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Hu, T.; Liu, Y.; Zhou, H.; Wang, J.; Long, M. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. arXiv 2022, arXiv:2210.02186. [Google Scholar]
- Zhou, H.; Zhang, S.; Peng, J.; Zhang, S.; Li, J.; Xiong, H.; Zhang, W. Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting. Proc. AAAI Conf. Artif. Intell. 2021, 35, 11106–11115. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Xu, J.; Wang, J.; Long, M. Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting. Adv. Neural Inf. Process. Syst. 2021, 34, 22419–22430. [Google Scholar]
- Iftikhar, H.; Turpo-Chaparro, J.E.; Rodrigues, P.C.; López-Gonzales, J.L. Day-Ahead Electricity Demand Forecasting Using a Novel Decomposition Combination Method. Energies 2023, 16, 6675. [Google Scholar] [CrossRef] [Scilit]
- Pallonetto, F.; Jin, C.; Mangina, E. Forecast Electricity Demand in Commercial Building with Machine Learning Models to Enable Demand Response Programs. Energy AI 2022, 7, 100121. [Google Scholar] [CrossRef] [Scilit]
- Grandón, T.G.; Schwenzer, J.; Steens, T.; Breuing, J. Electricity Demand Forecasting with Hybrid Classical Statistical and Machine Learning Algorithms: Case Study of Ukraine. Appl. Energy 2024, 355, 122249. [Google Scholar] [CrossRef] [Scilit]
- Cebekhulu, E.; Onumanyi, A.J.; Isaac, S.J. Performance Analysis of Machine Learning Algorithms for Energy Demand–Supply Prediction in Smart Grids. Sustainability 2022, 14, 2546. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Chen, Z.; Yang, Y.; Liu, C.; Li, X.; Wu, J. A Hybrid Autoformer Framework for Electricity Demand Forecasting. Energy Rep. 2023, 9, 3800–3812. [Google Scholar] [CrossRef] [Scilit]
- Wu, C.; Li, J.; Liu, W.; He, Y.; Nourmohammadi, S. Short-Term Electricity Demand Forecasting Using a Hybrid ANFIS–ELM Network Optimised by an Improved Parasitism–Predation Algorithm. Appl. Energy 2023, 345, 121316. [Google Scholar] [CrossRef] [Scilit]
- Sekhar, C.; Dahiya, R. Robust Framework Based on Hybrid Deep Learning Approach for Short Term Load Forecasting of Building Electricity Demand. Energy 2023, 268, 126660. [Google Scholar] [CrossRef] [Scilit]
- May, E.C.; Bassam, A.; Ricalde, L.J.; Soberanis, M.E.; Oubram, O.; Tzuc, O.M.; Alanis, A.Y.; Livas-García, A. Global Sensitivity Analysis for a Real-Time Electricity Market Forecast by a Machine Learning Approach: A Case Study of Mexico. Int. J. Electr. Power Energy Syst. 2022, 135, 107505. [Google Scholar] [CrossRef] [Scilit]
- Brown, T.B. Language Models Are Few-Shot Learners. arXiv 2020, arXiv:2005.14165. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A. Attention Is All You Need. Adv. Neural Inf. Process. Syst. 2017, 31, 6000–6010. [Google Scholar]
- Hadi, M.U.; Tashi, Q.A.; Qureshi, R.; Shah, A.; Muneer, A.; Irfan, M.; Zafar, A.; Shaikh, M.B.; Akhtar, N.; Hassan, S.Z.; et al. A Survey on Large Language Models: Applications, Challenges, Limitations, and Practical Usage. TechRxiv 2023. [Google Scholar] [CrossRef]
- Nazir, A.; Shaikh, A.K.; Shah, A.S.; Khalil, A. Forecasting Energy Consumption Demand of Customers in Smart Grid Using Temporal Fusion Transformer (TFT). Results Eng. 2023, 17, 100888. [Google Scholar] [CrossRef] [Scilit]
- Jin, M.; Wang, S.; Ma, L.; Chu, Z.; Zhang, J.Y.; Shi, X.; Chen, P.Y.; Liang, Y.; Li, Y.F.; Pan, S.; et al. Time-LLM: Time Series Forecasting by Reprogramming Large Language Models. arXiv 2023, arXiv:2310.01728. [Google Scholar]
- Su, J.; Jiang, C.; Jin, X.; Qiao, Y.; Xiao, T.; Ma, H.; Wei, R.; Jing, Z.; Xu, J.; Lin, J. Large Language Models for Forecasting and Anomaly Detection: A Systematic Literature Review. arXiv 2024, arXiv:2402.10350. [Google Scholar] [CrossRef] [Scilit]
- Xue, H.; Salim, F.D. Promptcast: A New Prompt-Based Learning Paradigm for Time Series Forecasting. IEEE Trans. Knowl. Data Eng. 2023, 36, 6851–6864. [Google Scholar] [CrossRef] [Scilit]
- Cao, D.; Jia, F.; Arik, S.O.; Pfister, T.; Zheng, Y.; Ye, W.; Liu, Y. Tempo: Prompt-Based Generative Pre-Trained Transformer for Time Series Forecasting. arXiv 2023, arXiv:2310.04948. [Google Scholar]
- Tan, M.; Merrill, M.; Gupta, V.; Althoff, T.; Hartvigsen, T. Are Language Models Actually Useful for Time Series Forecasting? Adv. Neural Inf. Process. Syst. 2024, 37, 60162–60191. [Google Scholar]
- Sun, C.; Li, H.; Li, Y.; Hong, S. TEST: Text Prototype Aligned Embedding to Activate LLM’s Ability for Time Series. arXiv 2023, arXiv:2308.08241. [Google Scholar]
- Guo, T.; Hauptmann, E. Fine-tuning large language models for stock return prediction using newsflow. arXiv 2024, arXiv:2407.18103. [Google Scholar] [CrossRef] [Scilit]
- Hu, Z.; Gao, Y.; Sun, L.; Mae, M. A novel attention-enhanced LLM approach for accurate power demand and generation forecasting. Renew. Energy 2025, 252, 123465. [Google Scholar] [CrossRef] [Scilit]
- Liu, G.; Bai, Y.; Wen, K.; Wang, X.; Liu, Y.; Liang, G.; Zhao, J.; Dong, Z.Y. LfLLM: A large language model for load forecasting. TechRxiv 2024. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Karimi, H.A. Exploring large language models for climate forecasting. arXiv 2024, arXiv:2411.13724. [Google Scholar] [CrossRef] [Scilit]
- Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. The LLaMA 3 Herd of Models. arXiv 2024, arXiv:2407.21783. [Google Scholar] [CrossRef] [Scilit]
- Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv 2018, arXiv:1810.04805. [Google Scholar]
- Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I. Language Models are Unsupervised Multitask Learners. OpenAI Blog 2019, 1, 9. [Google Scholar]














| Model | Without Prompt | Prompt1 | Prompt2 | Prompt3 | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MAE | MAPE | RMSE | MAE | MAPE | RMSE | MAE | MAPE | RMSE | MAE | MAPE | RMSE | |||||
| TimeLLM | 236 | 21.02 | 325 | 0.70 | 235 | 21.79 | 325 | 0.70 | 228 | 19.50 | 312 | 0.72 | 217 | 16.24 | 307 | 0.73 |
| MultiAttLLM | 202 | 15.40 | 283 | 0.77 | 209 | 16.46 | 292 | 0.75 | 204 | 15.90 | 285 | 0.77 | 205 | 16.20 | 288 | 0.76 |
| Proposed SA-LLM | 198 | 14.85 | 276 | 0.78 | 199 | 14.92 | 277 | 0.78 | 196 | 14.55 | 273 | 0.79 | 197 | 14.70 | 274 | 0.79 |
| Feature/Parameter | TimeLLM | MultiAttLLM | Proposed SA-LLM |
|---|---|---|---|
| Backbone | Frozen Pre-trained LLM | Frozen Pre-trained LLM | Frozen Pre-trained LLM |
| Attention Mechanism | Single Temporal Attention | Dual/Multi-Attention (Temporal + Cross-feature) | Unified Single Attention (Temporal + Auxiliary Features) |
| Input Dimension | |||
| Output Dimension | Forecast | Forecast | Forecast |
| Hidden Layer Size | 512 | 512 (Temporal) + 256 (Spatial) | 512 (Attention-refined) + 128 (Auxiliary) |
| Number of Layers | 4 Transformer Layers | 4 Transformer Layers + 2 Attention Fusion Layers | 4 Transformer Layers (Frozen) + 1 Unified Attention Layer |
| Activation Function | GELU | GELU | GELU + |
| Dropout Rate | 0.1 | 0.1 | 0.1 |
| Optimizer | AdamW | AdamW | AdamW |
| Learning Rate | |||
| Batch Size | 64 | 64 | 64 |
| Training Epochs | 200 | 200 | 200 |
| Regularization | Weight Decay | Weight Decay | Weight Decay + Laplacian Smoothness |
| Loss Function | MAPE + RMSE | MAPE + RMSE | MAPE + RMSE + + Graph Laplacian Smoothness |
| Special Features | Temporal Embeddings Only | Temporal + Spatial Embeddings | Temporal + Auxiliary + Optional GNN Embeddings |
| Training Region | Similarity to (S) | SA-LLM MAE (MWh) |
|---|---|---|
| 0.92 | 4.8 | |
| 0.87 | 5.1 | |
| 0.81 | 5.5 | |
| 0.75 | 5.9 | |
| 0.70 | 6.2 | |
| 0.65 | 6.5 |
| Model | Actual/Predicted | Accuracy | |||
|---|---|---|---|---|---|
| TimeLLM | 412 (75%) | 58 (11%) | 30 (14%) | 500 (82%) | |
| 76 (12%) | 395 (63%) | 64 (25%) | 635 (62%) | ||
| 42 (8%) | 71 (14%) | 387 (78%) | 500 (78%) | ||
| MultiAttLLM | 438 (85%) | 45 (9%) | 17 (6%) | 500 (85%) | |
| 62 (10%) | 418 (66%) | 55 (24%) | 635 (66%) | ||
| 29 (6%) | 63 (13%) | 408 (81%) | 500 (81%) | ||
| Proposed SA-LLM | 462 (92%) | 28 (6%) | 10 (2%) | 500 (92%) | |
| 41 (6%) | 452 (71%) | 42 (23%) | 635 (71%) | ||
| 18 (4%) | 39 (8%) | 443 (88%) | 500 (88%) |
| Model | Houston | Chicago | Philadelphia | NYC | Boston | Kansas City |
|---|---|---|---|---|---|---|
| LSTM MAE | 258 | 1105 | 885 | 830 | 880 | 910 |
| LSTM RMSE | 345 | 1432 | 1218 | 1175 | 1210 | 1250 |
| LSTM RAE | 0.49 | 0.56 | 0.51 | 0.50 | 0.52 | 0.53 |
| LSTM R | 0.72 | 0.64 | 0.65 | 0.66 | 0.64 | 0.63 |
| DLinear MAE | 242 | 935 | 820 | 810 | 825 | 840 |
| DLinear RMSE | 320 | 1250 | 1165 | 1145 | 1170 | 1185 |
| DLinear RAE | 0.46 | 0.49 | 0.48 | 0.47 | 0.48 | 0.49 |
| DLinear R | 0.74 | 0.72 | 0.69 | 0.70 | 0.68 | 0.67 |
| SA-LLM MAE | 235 | 890 | 800 | 795 | 810 | 825 |
| SA-LLM RMSE | 312 | 1220 | 1130 | 1125 | 1140 | 1155 |
| SA-LLM RAE | 0.44 | 0.46 | 0.47 | 0.46 | 0.47 | 0.48 |
| SA-LLM R | 0.77 | 0.74 | 0.71 | 0.71 | 0.70 | 0.69 |
| Model | LA | Chicago | Philadelphia | NYC | Boston | Kansas City |
|---|---|---|---|---|---|---|
| LSTM MAE | 265 | 1100 | 870 | 835 | 880 | 900 |
| LSTM RMSE | 362 | 1425 | 1200 | 1180 | 1215 | 1240 |
| LSTM RAE | 0.50 | 0.55 | 0.50 | 0.51 | 0.52 | 0.53 |
| LSTM R | 0.71 | 0.63 | 0.65 | 0.66 | 0.64 | 0.63 |
| DLinear MAE | 235 | 930 | 815 | 805 | 820 | 835 |
| DLinear RMSE | 322 | 1245 | 1160 | 1140 | 1165 | 1180 |
| DLinear RAE | 0.46 | 0.48 | 0.48 | 0.47 | 0.48 | 0.49 |
| DLinear R | 0.73 | 0.71 | 0.69 | 0.70 | 0.68 | 0.67 |
| SA-LLM MAE | 237 | 885 | 795 | 790 | 805 | 820 |
| SA-LLM RMSE | 310 | 1215 | 1120 | 1115 | 1130 | 1145 |
| SA-LLM RAE | 0.44 | 0.46 | 0.46 | 0.46 | 0.47 | 0.48 |
| SA-LLM R | 0.77 | 0.74 | 0.71 | 0.71 | 0.70 | 0.69 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zulfiqar, M.; Gamage, K.A.A.; Rasheed, M.B. Single-Attention Large Language Model for Efficient Multi-Regional Electricity Demand and Generation Forecasting. Energies 2026, 19, 1522. https://doi.org/10.3390/en19061522
Zulfiqar M, Gamage KAA, Rasheed MB. Single-Attention Large Language Model for Efficient Multi-Regional Electricity Demand and Generation Forecasting. Energies. 2026; 19(6):1522. https://doi.org/10.3390/en19061522
Chicago/Turabian StyleZulfiqar, Muhammad, Kelum A. A. Gamage, and M. B. Rasheed. 2026. "Single-Attention Large Language Model for Efficient Multi-Regional Electricity Demand and Generation Forecasting" Energies 19, no. 6: 1522. https://doi.org/10.3390/en19061522
APA StyleZulfiqar, M., Gamage, K. A. A., & Rasheed, M. B. (2026). Single-Attention Large Language Model for Efficient Multi-Regional Electricity Demand and Generation Forecasting. Energies, 19(6), 1522. https://doi.org/10.3390/en19061522

