Skip to Content
EnergiesEnergies
  • Editorial
  • Open Access

11 August 2026

6 Pages

Large Language Models in Power Systems: From Grid Operations to Home Energy Management

,
,
,
,
,
and
Department of Engineering, University of Palermo, Viale Delle Scienze, 90128 Palermo, Italy
*
Author to whom correspondence should be addressed.

1. Historical Evolution: From Numerical Intelligence to Language-Based Intelligence

Artificial intelligence has been applied to power system analysis in various ways in recent years. For instance, expert systems for operator support have been created; artificial neural networks have been used as fast surrogates for iterative solvers; and now, deep learning architectures are being applied to optimal power flow, forecasting, and stability assessment. Despite all the changes over time, one assumption has not changed significantly: AI models are task-specific tools trained on structured numerical data to solve a single, well-defined problem. Recently, large language models (LLMs) have appeared and thus changed this view. LLMs have been pre-trained on a large number of diverse datasets to add new functions to the power engineering toolbox, such as understanding unstructured data, general-purpose reasoning, code generation, and natural language interaction with operators and end-users [1,2].
Recent surveys document an extremely rapid growth of this research line, and bibliometric data make the trend concrete: a Scopus title–abstract–keyword search combining large language model terms with power system terms returns a single document published in 2022, 13 in 2023, 128 in 2024, and 464 in 2025—a more than thirty-fold increase in just two years—with a further 348 documents already indexed for the first half of 2026 (query performed on 11 July 2026; Figure 1). Systematic analyses of the literature identify applications spanning fault diagnosis, load forecasting, cybersecurity, control and optimization, system planning, simulation, and knowledge management [1], and map the enabling techniques through which general-purpose models are specialized for the energy domain, including retrieval-augmented generation, parameter-efficient fine-tuning, in-context learning, multimodal models, and multi-agent systems [2]. Other authors read the transition in cognitive terms: while conventional machine learning remains narrow and data-dependent, LLMs extend AI from task-specific functions to broader cognitive support for the operation of increasingly complex, converter-dominated and demand-participated power systems [2]. This editorial retraces this evolution from a specific perspective: the progressive migration of language-based intelligence from the transmission level down to the distribution level and, finally, into the home, where Home Energy Management Systems (HEMSs) represent both a promising application field and a demanding testbed for the reliability of LLM-based automation.
Figure 1. Documents per year indexed in Scopus.

2. LLMs for Grid-Level Analysis and Operations

A first consolidated line of research employs LLMs as an interface and orchestration layer on top of deterministic engineering tools. In this paradigm, the language model does not replace the numerical solver but coordinates it. The GridMind multi-agent system, for instance, couples LLM agents with AC optimal power flow and N–1 contingency analysis routines, allowing power system studies to be requested and interpreted through natural language while preserving numerical precision via function calls. Experiments on IEEE test cases show that the agentic framework consistently reaches correct solutions, with smaller models achieving comparable accuracy at lower latency [3]. Analogous verifiable agentic architectures have been proposed for distribution grid analyses [4] and for interactive grid analysis assistants [5], confirming a clear design trend: LLMs as conversational front-ends and workflow coordinators, and deterministic solvers as the computational back-end.
A second strand addresses fault diagnosis and knowledge management. LLM-based multi-agent pipelines have been used to translate unstructured regulatory documents and tacit expert knowledge into verified, executable fault-diagnosis workflows, with human-in-the-loop verification guaranteeing semantic fidelity and transparency [6]. Agentic approaches to fault analysis in power grids [7] and hybrid schemes coupling knowledge graphs with LLM-based contextual reasoning—applied, for example, to VSC-HVDC systems [8]—show how language models can improve the interpretability of diagnoses, a historical weak point of purely data-driven classifiers. The reach of these techniques extends down to the component level, where hybrid time-series LLM pipelines have been optimized to predict leakage-current progression and anticipate faults in distribution grid insulators [9]. More broadly, the surveyed literature highlights the role of LLMs as translators between natural-language operational requirements and machine-executable artifacts—code, simulation scripts, and optimization formulations—easing a knowledge-transfer bottleneck that purely numerical pipelines cannot address [1,2]. An early and influential example is the integration of GPT-based agents with deep reinforcement learning to embed linguistic stipulations directly into real-time optimal power flow [10]. The paradigm has since reached the operational loops of distribution grids, with in-context model-predictive frameworks for adaptive voltage regulation under frequent topology changes [11], graph-embedding schemes that couple power grid representations with LLM-driven optimization [12], language models supervising the safety of reinforcement learning agents for active distribution network energy management [13], and feedback-driven multi-agent frameworks that enhance the reliability of LLM-generated power system simulations [14].
Forecasting completes the grid-level picture. Following the demonstration that pre-trained language models can act as zero-shot time-series forecasters [15] and that time series can be “reprogrammed” into the language domain [16], a growing body of work has adapted LLMs to electric load and renewable generation prediction, including fine-tuned models that reduce the data requirements of load profile analysis [17], pre-trained architectures for wind power forecasting [18], and time-series LLMs for ultra-short-term distributed photovoltaic prediction [19]. Systematic evaluations on very short-term load forecasting suggest that the general temporal understanding acquired during pre-training gives LLMs remarkable few-shot adaptation capabilities, a valuable property in data-scarce contexts such as newly built areas or individual households [20].

3. LLMs at the Demand Side: Toward Conversational Home Energy Management

The demand side is arguably where the distinctive capabilities of LLMs—rather than their raw predictive power—become decisive. Decades of research have produced HEMS formulations based on mathematical programming, heuristics, and reinforcement learning, yet residential adoption remains limited, and one of the main barriers is not algorithmic but human: conventional HEMSs require users to translate everyday preferences, schedules, and comfort requirements into technical parameters [21]. LLMs attack precisely this barrier, and the recent literature outlines three progressive levels of integration.
At the first level, the LLM acts as an interface. Conversational agents have been designed to elicit user preferences through targeted questions, extract loosely formatted information—such as a textual description of a daily routine—and convert it into the well-structured inputs required by an underlying HEMS optimizer, which then computes the optimal appliance schedule [22]. Here, the division of labor is clean: language understanding is delegated to the LLM, while optimality is guaranteed by the mathematical model.
At the second level, the LLM becomes an autonomous coordinator. Agentic HEMS architectures have been demonstrated in which a hierarchical structure—an orchestrator agent delegating to specialist appliance agents through iterative reasoning-and-acting patterns—manages the complete workflow, from a natural-language request to multi-appliance scheduling and device control, integrating external contexts such as calendars and day-ahead electricity prices. Notably, evaluations against mixed-integer linear programming benchmarks report optimal schedules obtained with fully open-source models, with substantial performance differences across model scales, and coordination of multi-appliance scenarios within seconds [21].
At the third level, the LLM supports sustained human–AI collaboration. Multi-agent home energy assistants combine LLM reasoning with dozens of domain-specific tools and specialized agents for consumption analysis, education, and device control, explicitly reframing occupants from passive recipients of automation into engaged decision makers [23]. This trajectory—interface, coordinator, collaborator—suggests that the long-standing gap between the technical sophistication of HEMS research and its modest real-world adoption may finally be narrowed by conversational interaction, with direct implications for demand-response participation, local renewable integration, and, prospectively, vehicle-to-grid coordination.
The surrounding demand-side ecosystem is evolving in the same direction. For electric mobility, retrieval-augmented LLM pipelines automate problem formulation, code generation, and optimization customization for demand-side management with the Internet of Electric Vehicles [24], while LLM-based agent frameworks reproduce realistic charging behavior by integrating user preferences and psychological factors [25]. On the market side, LLM agents have been employed to assist the bidding of battery energy storage systems in frequency regulation markets [26] and to couple bidding-behavior and market-sentiment agents for electricity price prediction [27], prefiguring a scenario in which household flexibility, aggregators, and markets negotiate through language-based intermediaries. Systematic reviews of transformer- and LLM-based applications across the energy sector condense this trajectory into the concept of the agentic digital twin: an autonomous, LLM-augmented replica of energy assets—including buildings and homes—capable of proactive management [28].
There is, finally, an elegant and somewhat paradoxical closure of the loop. LLMs are not only a management tool for power systems: they are themselves a new and challenging electrical load. Data centers hosting large models exhibit almost zero inertia, extremely fast ramps, and abrupt power peaks and drops that can stress distribution networks [29]. Language-based intelligence, in other words, simultaneously burdens the grid and offers instruments to govern it—a duality that future demand-side research, including HEMS-oriented work, can hardly ignore.

4. Open Challenges and Future Perspectives

Despite the impressive pace of the field, several open issues stand between current prototypes and dependable deployment. The first is reliability in safety-critical contexts. LLMs are probabilistic systems prone to hallucination; when their output parametrizes an optimizer, this risk is contained, but when they directly coordinate physical devices—as in agentic HEMSs—formal validation layers, deterministic fallbacks, and human-in-the-loop mechanisms become indispensable, as recognized both in grid-level agentic frameworks [3,6] and in demand-side implementations [21]. Closely related is the strong dependence of performance on model scale and the scarcity of domain-specific training data, which current surveys identify among the main obstacles to practical adoption [1,2].
A second front is privacy and edge deployment. Household energy data reveal occupancy patterns and lifestyle habits; routing them through cloud-hosted models raises concerns that motivate research on compact, locally executable models, whose capabilities are, however, markedly inferior to those of frontier systems—a tension clearly exposed by the performance gaps observed across open-source model scales in HEMS orchestration [21]. Inference cost and latency add an economic dimension to the same trade-off.
Third, the community lacks shared benchmarks and realistic validation. As already noted for AI in power flow studies, fragmented case studies based on synthetic datasets limit transferability to real systems; for LLMs, the problem is amplified by the rapid obsolescence of the models themselves. Common test scenarios, standardized user simulation methodologies for conversational agents [23], and reproducible evaluation pipelines against optimization benchmarks [21] are promising steps that deserve systematization.
Finally, interoperability and cybersecurity will condition industrial acceptance. Agentic systems that read calendars, query market prices, and actuate appliances enlarge the attack surface of digitalized energy infrastructures, while emerging tool-integration standards are only beginning to mature. The concern is not hypothetical: threat-modeling studies have validated that bad-data injection and domain-knowledge extraction attacks are practical against LLMs deployed in smart grid applications [30], and early analyses of the security threats of applying LLMs to power systems reach similar conclusions [31]. At the same time, language models are proving useful on the defensive side, for example, in detecting cyberattacks on smart inverters through the textual analysis of Volt/VAR command streams [32], confirming that LLMs are simultaneously an attack surface and a protection tool. Taken together with the grid-side impact of AI loads themselves [29], these challenges indicate that the maturity of language-based intelligence in power systems—from the control room to the living room—will require a joint effort between academic research, network operators, device manufacturers, and the AI industry.

Author Contributions

Conceptualization, Z.N., R.M., E.R.S. and G.S.; methodology, Z.N.; investigation, Z.N.; data curation, Z.N.; writing—original draft preparation, Z.N.; writing—review and editing, G.C., S.F., R.M., E.R.S., G.S. and G.Z.; visualization, Z.N.; supervision, E.R.S. and R.M and G.S.; project administration, E.R.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Sarwar, M.; Rizwan, M.; Aziz, M.; Sudais, A.R. Large Language Models for Power System Applications: A Comprehensive Literature Survey. arXiv 2025, arXiv:2512.13004. [Google Scholar]
  2. Fan, X.; Li, Y.; Zhang, C.; Wan, Y.; Du, E.; Liu, G.; Liu, J.; Kang, C. Large Language Models in Power Systems toward Cognitive Intelligence: A Literature Review. Appl. Energy 2026, 421, 128205. [Google Scholar] [CrossRef] [Scilit]
  3. Jin, H.; Kim, K.; Kwon, J. GridMind: LLMs-Powered Agents for Power System Analysis and Operations. In Proceedings of the SC’25 Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis, St. Louis, MO, USA, 16–21 November 2025; pp. 560–568. [Google Scholar]
  4. Badmus, E.O.; Sang, P.; Stamoulis, D.; Pandey, A. PowerChain: A Verifiable Agentic AI System for Automating Distribution Grid Analyses. arXiv 2025, arXiv:2508.17094. [Google Scholar]
  5. Wen, Y.; Chen, X. X-GridAgent: An LLM-Powered Agentic AI System for Assisting Power Grid Analysis. arXiv 2025, arXiv:2512.20789. [Google Scholar]
  6. Wang, Y.; Tian, Y.; Shen, X.; Zhang, G.; Sun, J.; Zhang, H.; Xu, R.; Zhao, F. Fault2Flow: An AlphaEvolve-Optimized Human-in-the-Loop Multi-Agent System for Fault-to-Workflow Automation. arXiv 2025, arXiv:2511.12916. [Google Scholar]
  7. Saha, B.K.; Aarthi, V.; Naidu, O. DrAgent: An Agentic Approach to Fault Analysis in Power Grids Using Large Language Models. In Proceedings of the 2025 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), Fukuoka, Japan, 18–21 February 2025; pp. 938–945. [Google Scholar]
  8. Lan, Y.; Zhang, M.; Su, M.; Zhou, F. Knowledge-Graph-Enhanced and LLM-Guided Fault Diagnosis for VSC-HVDC Systems. AIP Adv. 2025, 15, 115330. [Google Scholar] [CrossRef] [Scilit]
  9. Matos-Carvalho, J.P.; Stefenon, S.F.; Leithardt, V.R.Q.; Yow, K.-C. Time Series Forecasting Based on Optimized LLM for Fault Prediction in Distribution Power Grid Insulators. arXiv 2025, arXiv:2502.17341. [Google Scholar]
  10. Yan, Z.; Xu, Y. Real-Time Optimal Power Flow with Linguistic Stipulations: Integrating GPT-Agent and Deep Reinforcement Learning. IEEE Trans. Power Syst. 2024, 39, 4747–4750. [Google Scholar] [CrossRef] [Scilit]
  11. Jena, A.; Ding, F.; Wang, J.; Yao, Y.; Xie, L. LLM-Based Adaptive Distribution Voltage Regulation Under Frequent Topology Changes: An In-Context MPC Framework. IEEE Trans. Smart Grid 2025, 16, 4297–4300. [Google Scholar] [CrossRef] [Scilit]
  12. Bernier, F.; Cao, J.; Cordy, M.; Ghamizi, S. PowerGraph-LLM: Novel Power Grid Graph Embedding and Optimization with Large Language Models. IEEE Trans. Power Syst. 2025, 40, 5483–5486. [Google Scholar] [CrossRef] [Scilit]
  13. Yang, X.; Lin, C.; Liu, H.; Wu, W. RL2: Reinforce Large Language Model to Assist Safe Reinforcement Learning for Energy Management of Active Distribution Networks. IEEE Trans. Smart Grid 2025, 16, 3419–3431. [Google Scholar] [CrossRef] [Scilit]
  14. Jia, M.; Cui, Z.; Hug, G. Enhancing LLMs for Power System Simulations: A Feedback-Driven Multi-Agent Framework. arXiv 2024, arXiv:2411.16707. [Google Scholar]
  15. Gruver, N.; Finzi, M.; Qiu, S.; Wilson, A.G. Large Language Models Are Zero-Shot Time Series Forecasters. In Proceedings of the Advances in Neural Information Processing Systems 36 (NeurIPS 2023), New Orleans, LA, USA, 10–16 December 2023. [Google Scholar]
  16. Jin, M.; Wang, S.; Ma, L.; Chu, Z.; Zhang, J.; Shi, X.; Chen, P.Y.; Liang, Y.; Li, Y.F.; Pan, S.; et al. Time-LLM: Time Series Forecasting by Reprogramming Large Language Models. In Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 7–11 May 2024. [Google Scholar]
  17. Hu, Y.; Kim, H.; Ye, K.; Lu, N. Applying Fine-Tuned LLMs for Reducing Data Needs in Load Profile Analysis. Appl. Energy 2025, 377, 124666. [Google Scholar] [CrossRef] [Scilit]
  18. Lai, Z.; Wu, T.; Fei, X.; Ling, Q. BERT4ST: Fine-Tuning Pre-Trained Large Language Model for Wind Power Forecasting. Energy Convers. Manag. 2024, 307, 118331. [Google Scholar] [CrossRef] [Scilit]
  19. Lv, C.; Fan, H.; Zhang, Z.; Fan, M.; Run, W.; Yang, L.; Yang, Y.; Liu, D. Ultra-Short-Term Power Prediction for Distributed Photovoltaics Based on Time-Series LLMs. Electronics 2025, 14, 4519. [Google Scholar] [CrossRef] [Scilit]
  20. Chen, Y.; Chen, M.; Hu, C.; Wu, J.; Qin, R. Ele-LLM: A Systematic Evaluation and Adaptation of Large Language Models for Very Short-Term Power Load Forecasting. Energies 2026, 19, 631. [Google Scholar] [CrossRef] [Scilit]
  21. El Makroum, R.; Zwickl-Bernhard, S.; Kranzl, L. Agentic AI Home Energy Management System: A Large Language Model Framework for Residential Load Scheduling. arXiv 2025, arXiv:2510.26603. [Google Scholar]
  22. Michelon, F.; Zhou, Y.; Morstyn, T. Large Language Model Interface for Home Energy Management Systems. In Proceedings of the 16th ACM International Conference on Future and Sustainable Energy Systems (e-Energy), Rotterdam, The Netherlands, 17–20 June 2025; pp. 590–602. [Google Scholar]
  23. Jung, W. Multi-Agent Home Energy Management Assistant. arXiv 2026, arXiv:2602.15219. [Google Scholar]
  24. Zhang, H.; Zhang, R.; Zhang, W.; Niyato, D.; Wen, Y.; Miao, C. Advancing Generative Artificial Intelligence and Large Language Models for Demand Side Management with Internet of Electric Vehicles. arXiv 2025, arXiv:2501.15544. [Google Scholar]
  25. Feng, J.; Cui, C.; Zhang, C.; Fan, Z. Large Language Model Based Agent Framework for Electric Vehicle Charging Behavior Simulation. arXiv 2024, arXiv:2408.05233. [Google Scholar]
  26. Zhang, B.; Li, C.; Chen, G.; Dong, Z.Y. Large Language Model Assisted Optimal Bidding of BESS in FCAS Market: An AI-Agent Based Approach. arXiv 2024, arXiv:2406.00974. [Google Scholar]
  27. Lu, X.; Qiu, J.; Yang, Y.; Zhang, C.; Lin, J.; An, S. Large Language Model-Based Bidding Behavior Agent and Market Sentiment Agent-Assisted Electricity Price Prediction. IEEE Trans. Energy Mark. Policy Regul. 2024, 3, 223–235. [Google Scholar] [CrossRef] [Scilit]
  28. Antonesi, G.; Cioara, T.; Anghel, I.; Michalakopoulos, V.; Sarmas, E.; Toderean, L. From Transformers to Large Language Models: A Systematic Review of AI Applications in the Energy Sector towards Agentic Digital Twins. arXiv 2025, arXiv:2506.06359. [Google Scholar]
  29. Li, Y.; Mughees, M.; Chen, Y.; Li, Y.R. The Unseen AI Disruptions for Power Grids: LLM-Induced Transients. arXiv 2024, arXiv:2409.11416. [Google Scholar]
  30. Li, J.; Yang, Y.; Sun, J. Risks of Practicing Large Language Models in Smart Grid: Threat Modeling and Validation. arXiv 2024, arXiv:2405.06237. [Google Scholar]
  31. Ruan, J.; Liang, G.; Zhao, H.; Liu, G.; Sun, X.; Qiu, J.; Xu, Z.; Wen, F.; Dong, Z.Y. Applying Large Language Models to Power Systems: Potential Security Threats. IEEE Trans. Smart Grid 2024, 15, 3333–3336. [Google Scholar] [CrossRef] [Scilit]
  32. Selim, A.; Zhao, J.; Yang, B. Large Language Model for Smart Inverter Cyber-Attack Detection via Textual Analysis of Volt/VAR Commands. IEEE Trans. Smart Grid 2024, 15, 6179–6182. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.