Next Article in Journal
Research on the Cutting Efficiency of TBM Cutters in Jointed Rock Mass Based on a Multivariate Nonlinear Regression Model
Previous Article in Journal
Spatial Patterns of Bridge Deterioration and Municipal Maintenance Potential for Municipality-Managed Bridges in the Chubu Region of Japan
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Reinforcement Learning for Optimizing Renewable Energy Utilization in Smart Grids: Recent Advances in Power Grids, Microgrids, and Building Energy Systems

by
Panagiotis Michailidis
1,2,*,
Federico Minelli
3,
Hasan Huseyin Coban
4,
Iakovos Michailidis
1,2 and
Elias Kosmatopoulos
1,2
1
Information Technologies Institute, Centre for Research and Technology Hellas (CERTH), Thermi, 57001 Thessaloniki, Greece
2
Department of Electrical and Computer Engineering, Democritus University of Thrace (DUTH), 67100 Xanthi, Greece
3
Department of Industrial Engineering, University of Naples “Federico II”, 80125 Naples, Italy
4
Department of Electrical and Electronics Engineering, Bartin University, 74110 Bartın, Türkiye
*
Author to whom correspondence should be addressed.
Infrastructures 2026, 11(7), 240; https://doi.org/10.3390/infrastructures11070240
Submission received: 11 June 2026 / Revised: 10 July 2026 / Accepted: 13 July 2026 / Published: 15 July 2026

Abstract

The extensive deployment of renewable energy sources (RES) across modern energy infrastructure has introduced significant operational complexity, necessitating the development of advanced data-driven control strategies to ensure reliable and efficient system operation. Among these approaches, reinforcement learning (RL) has emerged as a promising paradigm for managing renewable generation and coordinating interconnected energy subsystems under uncertainty and dynamic operating conditions. The current paper presents a comprehensive review of RL-based control applications across RES-integrated energy domains, including power grids, microgrids, and building energy systems. The paper begins by outlining the fundamental characteristics of these smart grid energy environments along with the mathematical foundations of RL and its principal algorithmic families. A structured analysis of recent peer-reviewed studies is then conducted, with the literature systematically categorized according to the corresponding energy domain. A high number of impactful selected studies are further examined across multiple key dimensions, including RL methodologies, agent architectures, reward design, baseline control strategies, RES-integrated technologies, and control objectives. Based on this multi-dimensional evaluation, the review identifies emerging trends and highlights dominant design patterns across power grid, microgrid, and building-level applications. Finally, the observations are critically discussed and future research directions are outlined towards the development of scalable, practical, and reliable RL-based energy management solutions for next-generation smart grid systems.

1. Introduction

1.1. General

The global transition toward sustainable energy systems has significantly accelerated the deployment of renewable energy sources, including solar photovoltaic (PV) systems, wind turbines (WT), and other distributed energy technologies [1,2,3]. This transition is primarily driven by the urgent need to mitigate climate change, reduce greenhouse gas emissions, and enhance energy security in response to continuously increasing global energy demand [4,5]. As RES technologies continue to mature and become more economically viable, their integration has expanded across multiple layers of modern energy infrastructure, ranging from large-scale power grids to localized microgrids and building energy systems [6,7,8]. Such developments are further supported by international sustainability initiatives and policy frameworks such as the European Green Deal [9] and related renewable energy directives [10], which promote the widespread adoption of RES and the evolution of advanced energy management practices [11].
However, despite their well-established environmental and economic benefits, renewable energy technologies introduce substantial operational challenges [12,13,14] in smart grid frameworks. Unlike conventional generation systems, most RES exhibit inherently intermittent and stochastic behavior, as their output is directly influenced by variable environmental conditions such as solar irradiance and wind availability [15]. This variability increases the complexity of maintaining reliable system operation and necessitates the development of advanced control and coordination strategies capable of balancing generation, storage, and consumption in real time [16,17,18]. Moreover, the current challenges are further intensified in heterogeneous real-world energy environments [19,20]. At the power grid level, high penetration of renewable energy affects system stability, voltage and frequency regulation, and power flow management [21,22], while in microgrid environments the coordinated operation of renewable generators, energy storage systems, and controllable loads requires intelligent scheduling mechanisms to ensure both reliability and economic performance [23]. Similarly, the integration of rooftop PV and other renewable technologies into building energy frameworks necessitates adaptive energy management strategies that may effectively balance occupant comfort, energy costs, and local energy production [24,25,26,27].
Conventional control approaches, including rule-based control (RBC) methods and model-based optimization techniques, often exhibit limited effectiveness under the dynamic and uncertain conditions characteristic of RES-integrated energy systems [28]. Such methods typically rely on predefined rules or simplified system representations, which restrict their ability to respond adequately to rapidly changing operating conditions [29]. In order to cope with such a complicated optimization problem, research attention has shifted toward intelligent data-driven approaches capable of learning effective control policies through direct interaction with the system [27,30,31,32,33,34,35]. Alongside conventional RBC and model-based control practices, more recent hybrid frameworks have emerged for RES-rich energy systems by combining data-driven forecasting, uncertainty estimation, and mathematical optimization with explicit physical models and operational constraints [36,37,38]. A key advantage of such approaches concerns their ability to produce decisions that are both transparent and aware of system limits, especially when the underlying model is accurate enough to reflect real operation. At the same time, their performance might weaken when the system model becomes inaccurate, operating conditions change significantly, or the control problem becomes too complex for real-time coordination [39,40]. More broadly, RES-rich energy systems have also been addressed through stochastic optimization, robust optimization, MPC, and hybrid data-driven optimization frameworks [36,41,42]. Such approaches generally provide stronger feasibility guarantees, more explicit constraint handling, and a clearer interpretation of optimality under uncertainty, particularly when the physical model and uncertainty sets are well-defined [36,41,42]. However, their effectiveness often depends on model accuracy, forecast quality, and computational tractability, which may become limiting in highly dynamic, high-dimensional, or real-time operating environments [39,42].
Among these approaches, RL has emerged as a particularly promising solution for complex energy management problems [43]. RL practice enables autonomous agents to learn decision-making policies based on environmental interactions and observed system responses [44,45], while its model-free nature allows for handling nonlinear dynamics and uncertainty without requiring explicit system modeling. Such characteristics make RL especially suitable for grid environments with high RES penetration, where variability and uncertainty represent fundamental challenges [46,47,48]. However, such flexibility also comes with important challenges, including unstable training, high data requirements, lower interpretability, and lack of guaranteed feasibility unless additional constrained, safe, or physics-informed mechanisms are included [49,50,51]. For this reason, RL should not be seen as a replacement for the aforementioned optimization frameworks but rather as an alternative control paradigm that becomes especially valuable when model accuracy, scalability, and online adaptability begin to limit the effectiveness of more traditional approaches [47,52]. Table 1 provides a brief comparison of RL with classical and hybrid optimization frameworks.
Over the past decade, RL has been increasingly investigated for renewable energy integration across multiple smart grid domains. In power grids, RL-based controllers have been widely applied to problems such as voltage regulation [53], frequency control [54,55], and economic dispatch [56,57] in environments with high renewable penetration [58,59]. In microgrid frameworks, RL has been utilized to coordinate renewable generation, energy storage systems, and flexible loads in order to enhance system efficiency and resilience [60,61,62]. Furthermore, in building energy systems, RL-based approaches have been proposed to optimize renewable self-consumption, facilitate participation in demand response programs, and enable the coordinated operation of HVAC systems, storage units, and electric vehicles [62,63,64].
Given the rapid growth of RL in energy research, a comprehensive synthesis of applications across these interconnected domains is increasingly needed. Although earlier studies have typically focused on individual system layers such as power systems or building energy management, a unified review spanning RES-integrated smart grids considering power grids, microgrids and buildings is still needed. To address this gap, the present review systematically examines recent RL-based frameworks for RES integration across these three well-evaluated energy environments. Thus, the primary scope of this work is to summarize and analyze high-impact RL research in recent years to determine established trends and reveal the research trajectory of RL in RES-integrated smart grids, specifically power grids, microgrids and building energy frameworks. By analyzing key methodological features, system configurations, and application goals, the current work provides a structured overview of how RL is shaping next-generation renewable energy management and highlights future research directions toward scalable and practical smart grid frameworks.

1.2. Previous Works

A growing body of review literature has examined the use of RL in modern energy systems, with particular emphasis on RES integration and intelligent energy management. One of the earlier comprehensive reviews was presented by Cao et al. [65], exploring the role of RL in power and energy systems. This particular study introduced the main concepts of the RL methodology and organized the literature around classical RL, deep reinforcement learning (DRL), and multi-agent RL (MARL) practices applied to system operation, optimization, and energy management. The work also covered key application areas including demand response, grid control, and electricity markets while outlining major challenges and future research directions for RL in the energy sector. Erick et al. [66] focused more specifically on power management in grid-connected microgrids with renewable energy sources and energy storage systems. Their paper examined RL formulations for microgrid scheduling problems, such as battery scheduling and unit commitment, and discussed the design of state–action spaces and reward functions for these environments. Their study also compared classical and deep RL methods for microgrid energy management and highlighted their potential advantages over conventional optimization approaches. More recently, Li et al. [67] reviewed the use of RL and broader machine learning techniques for energy management in smart grid environments within smart cities. Their work emphasized the role of DRL in supporting predictive and adaptive renewable energy management under growing urban energy demand and increasing system complexity. They also discussed the integration of enabling technologies such as blockchain to improve security, transparency, and reliability in intelligent energy management frameworks. Finally, Smart et al. [68] examined the wider contribution of artificial intelligence to RES forecasting and optimization. They considered multiple AI examples for improving both forecasting accuracy and operational performance in solar, wind, hydro, and biomass energy systems, including machine learning, deep learning, and reinforcement learning. In addition to highlighting the advantages of AI-based methods over traditional forecasting approaches, they pointed to persistent challenges related to data quality, computational burden, and model explainability.
Several previous review contributions by the authors are also worth mentioning, since they address related topics within the broader fields of energy systems, smart grids, building energy management, and RL. In 2023, Vamvakas and Michailidis et al. [69] reviewed RL frameworks across smart-grid applications, including RES plants, EV charging stations, and building-related use cases. The same year, Michailidis et al. [70] also reviewed the overall framework of model-free algorithms for HVAC control in building energy systems, integrating the RL application in such frameworks as well as neural networks, FLC, metaheuristics hybrids, and other model-independent control approaches. The application of multi-agent control into building energy management systems was also reviewed in 2024 by [71], elaborating on decentralized and distributed control schemes for integrated building energy management systems (IBEMS). The following year, this work was expanded to cover reinforcement learning for electric vehicle charging systems [43] and the application of RL in RES-integrated building frameworks [45]. Taking advantage of this previously gathered knowledge, the present manuscript focuses on RES-integrated smart grid frameworks and provides a cross-domain synthesis applications across power grids, microgrids, and building energy systems. As such, its contribution is positioned as a unified comparison of algorithmic families, agent architectures, RES integration levels, controlled energy assets, testbeds, baselines, reward formulations, and primary control objectives across the three application layers: power grids, microgrids, and buildings. No text, tables, figures, or analytical frameworks from the previous works are reproduced in the present manuscript.

1.3. Novelty and Contribution

The current review distinguishes itself from the existing literature on RL applications in RES-integrated smart grid frameworks through several important contributions. First, unlike most previous reviews that focus on a single energy domain such as power systems, microgrids, or building energy management, the current study provides a unified and systematic examination of RL-based control frameworks across these three layers of the energy ecosystem. By analyzing RL applications across these interconnected levels, this review offers a broader perspective on how data-driven control may support the integration, coordination, and optimal operation of RES within modern energy infrastructures. Second, this study presents a comprehensive analysis of RL-based optimization and control strategies for RES management. By reviewing a broad set of influential works published over the last decade, we cover RL-driven solutions for voltage and frequency regulation in power grids, energy scheduling and resource coordination in microgrids, demand-side management in building energy systems, and more. Through this synthesis, the present review shows how RL approaches such as value-based, policy-based, and actor–critic methods have evolved to address the growing complexity of renewable energy integration. In addition, this work introduces a structured taxonomy for classifying RL applications based on the energy domain (grid, microgrid, or building), control objective, RES technologies involved, and interacting subsystems such as energy storage, thermal storage, electric vehicles, and demand-side resources. This classification helps to clarify how RL methods are deployed across different operational settings and makes it easier to identify recurring design patterns in RL-based energy management frameworks. Finally, this paper provides comparative analyses of representative studies through detailed summary tables covering key methodological features such as the RL algorithm used, the state–action–reward formulation, the renewable energy configuration, the interacting assets, and the evaluation metrics. By bringing all of these elements together in the evaluation, this work identifies emerging research trends, highlights current limitations, and outlines promising future directions for the development of scalable and reliable RL-based control strategies for RES-integrated multi-layered smart grids. The following Table 2 provides a comparison of contributions between the current review and previous works considering the different evaluation attributes.

1.4. Paper Structure

The overall paper structure is illustrated in Figure 1. More specifically:
  • Section 1 introduces the background and motivation of the study, reviews the most relevant existing surveys on RL applications in smart grids, identifies the main gaps in the literature, and presents the novelty and contributions of this work.
  • Section 2 describes the methodology followed in this systematic review, including the literature search strategy, article retrieval process, filtering and selection criteria, attribute extraction, quality assessment, and synthesis framework used for the analysis of the selected studies.
  • Section 3 provides an overview of RES energy integration in modern smart grid architectures. It first introduces the main renewable energy technologies commonly deployed in smart grid infrastructures, then discusses the key operational environments in which they are integrated, namely, power grids, microgrids, and building energy systems, together with their main control objectives.
  • Section 4 presents the mathematical foundations of RL-based control, including the basic RL formulation, multi-agent formulations, and aspects of RL and the main algorithmic families applied in energy systems, such as value-based, policy-based, and actor–critic methods.
  • Section 5 organizes the selected RL-based applications in RES-integrated energy systems into structured summary tables. The studies are grouped by application domain, including power grids, microgrids, and building energy systems, while emphasizing the main characteristics of each work.
  • Section 6 provides a detailed evaluation of the reviewed studies across multiple analytical dimensions, including algorithmic trends, agent architectures, reward design, baseline comparisons, renewable energy, and other integrated technologies and control objectives across the different smart grid domains.
  • Section 7 discusses the main trends identified through the review and outlines promising future research directions for reinforcement learning applications in RES-integrated smart grids, including power grids, microgrids, and buildings.
  • Section 8 concludes the paper by summarizing the key findings of the review, highlighting emerging trends, identifying current limitations, and presenting future directions for reinforcement learning in RES-integrated smart grid systems.

2. Methodology

The methodological steps for conducting the current review may be summarized as follows:
  • Article Search and Retrieval: A systematic literature search was carried out using major academic databases that capture a significant portion of high-impact peer-reviewed research in energy systems and artificial intelligence, specifically Scopus and Web of Science (WoS). An appropriate search strategy was designed to identify studies conducting RL applications on RES-integrated smart grid environments including power grids, microgrids, and building energy systems. To access these studies, RL-related keywords were combined with keywords relating to renewable energy technologies and energy management infrastructures. The main query was constructed as follows:
    ["Reinforcement Learning" OR "Deep Reinforcement Learning" OR "RL" OR "DRL"] AND ["Renewable Energy" OR RES OR "Solar PV" OR "Wind Power" OR "Distributed Generation"] AND ["Smart Grid" OR "Power System" OR "Distribution Network" OR Microgrid OR "Energy Management System" OR "Building Energy Management" OR BEMS].
    The initial search provided several hundred publications, including journal articles and conference proceedings from 2020 to 2025, which corresponds to the time window in which RL has seen significantly expanded applications in energy research. Titles, abstracts, and keywords were then screened to identify studies that characterized RL-based control strategies in renewable-integrated energy systems.
  • Filtering and Selection Criteria: After the initial search, filtering was applied to eliminate duplicate or irrelevant low-quality contributions. Duplicate records were first removed in the selected databases, after which the studies were evaluated based on their relevance to RL applications in renewable-integrated energy systems. A minimum citation threshold of 35 (not including self-citations) was adopted to guarantee the inclusion of influential works. For recent publications (e.g., 2025) where the number of citations may still be restricted, this threshold was relaxed to 25 citations if the work was published in a reputable peer-reviewed forum and its methodological contribution was clear. Peer-reviewed journal articles and highly cited conference papers were included if the paper presented RL-based techniques for the operation, control, or optimization of energy systems with a central presence of renewable energy sources and did not treat the RL method in a theoretical manner.
  • Data Extraction and Attribute Collection: A detailed set of methodological and system-level attributes were extracted for each included study to enable systematic comparison. The collected attributes included: (a) RL methodology, (b) agent architecture (single-agent or multi-agent), (c) forecast horizon and control timestep time intervals, (d) baseline control strategies used for benchmarking, (e) renewable energy technologies involved (PV, wind turbines, hydropower, biomass, hybrid systems, etc.), (f) other integrated control energy assets (e.g., generators, energy storage systems, EVs, flexible loads, HVAC systems), (g) smart grid framework (power grid, microgrid, or building energy system); (h) type of power grid, microgrid, or building energy framework, (i) real-world or simulation data utilization, (j) reward design, and (k) control objectives (e.g., voltage regulation, frequency control, energy scheduling, cost optimization, comfort, demand response). Special emphasis was placed on the treatment of studies reporting quantitative performance gains such as reduced operational cost, increased renewable utilization, or grid stability improvement.
  • Quality Assessment: All included studies were evaluated based on reporting the level of methodological soundness, clarity of the RL formulation, quality of the experimental evaluation, and availability of the reference list and supplementary materials. Priority was given to studies that clearly defined their state, action, and reward structure, provided sufficient details about training process and learning architectures, treated baseline strategies separately from RL-based solutions, and presented quantitative results as measures demonstrating the value of proposed RL-based controllers. Additionally, studies from established journals and conferences from recognized publishers such as Elsevier, IEEE, Springer, and MDPI were prioritized. In particular, the focus was on papers which presented a complete pipeline, from system modeling to control design to performance evaluation.
  • Data Synthesis and Comparative Analysis: The final set of included studies was organized systematically based on the relevant smart grid framework (power grid, microgrid, building), RL type, RES technologies, control objectives, and energy subsystems. Overall, the review integrated 84 highly-cited works, including 29 concerning applications of RL in power grid frameworks, 30 concerning RL applications in microgrid frameworks, and 25 concerning applications of RL in building energy frameworks. A summarized organization provides a wider view of the RL application environment of the included studies. Next, detailed statistics and comparison tables were developed to visualize trends in terms of RL method choice, specific optimal configurations, and performance results. A discussion of RL application behavior for RES-integrated grid systems and future RL exploiting opportunities is then provided based on this synthesis. The PRISMA diagram portraying the overall methodology is illustrated in Figure 2.

3. RES Integration in Smart Grid Architectures

Smart grids represent the next generation of electricity networks, and are designed to enhance the efficiency, reliability, and sustainability of modern power systems through the integration of advanced communication, sensing, and control technologies [72,73]. In contrast to conventional grids characterized by centralized generation and unidirectional power flow from generation units to end users, smart grids enable bidirectional exchange of both energy and information across the entire network infrastructure [74]. This paradigm shift facilitates the seamless integration of diverse distributed energy resources (DERs), including renewable energy sources such as photovoltaic (PV) systems, wind turbines (WTs), and other decentralized generation technologies [75,76,77,78]. By leveraging digital monitoring systems, intelligent control algorithms, advanced metering infrastructure, and distributed energy management platforms, smart grids can support real-time system visibility and enable adaptive coordination between supply and demand [79,80]. Such capabilities play a critical role in addressing the challenges associated with the increasing penetration of variable renewable energy sources. In particular, they contribute to improved grid stability, enhanced demand-side participation, and more efficient coordination of generation, storage, and consumption. This applies not only to large-scale interconnected power systems but also to localized energy environments, including microgrids and smart building infrastructure [81].

3.1. Types of RES in Smart Grid Environments

Modern smart grids incorporate a wide range of RES technologies, each characterized by distinct operational features that influence system planning, stability, and control strategies [76]. Among these technologies, solar photovoltaics and wind turbines dominate, accounting for a substantial portion of newly installed renewable capacity worldwide. In addition, other resources such as solar thermal, geothermal, biomass, and hydropower systems contribute to the overall energy mix, offering varying degrees of controllability and operational flexibility [82,83]. More specifically, the primary RES technologies integrated into smart grids may be summarized as follows [45]:
  • Photovoltaic Systems (PV): Such systems convert solar irradiance directly into electrical energy through semiconductor-based cells. Photovoltaics are deployed either as large-scale solar farms connected to transmission networks or as distributed rooftop installations within distribution grids and buildings [84,85,86,87]. The power output of PV systems follows clear daily and seasonal patterns and is highly sensitive to cloud coverage and atmospheric conditions, often resulting in rapid fluctuations that may lead to voltage deviations and power imbalances [88,89].
  • Wind Turbines (WT): Wind turbines generate electricity by converting the kinetic energy of wind into mechanical and subsequently electrical power through rotor–generator systems [90,91]. Although wind generation may occur continuously during both day and night, it remains strongly dependent on stochastic wind patterns and site-specific conditions [92]. In large-scale wind farms connected to transmission networks, sudden variations in wind speed can induce power ramping events, posing challenges for frequency regulation and requiring rapid balancing mechanisms [93,94].
  • Biomass Energy (BIO): Biomass-based power generation relies on the combustion, gasification, or biochemical conversion of organic materials, including agricultural residues, wood waste, and dedicated energy crops [95,96]. In contrast to intermittent RES, these systems exhibit behavior similar to conventional thermal power plants, offering relatively stable and controllable output [95,97]. Consequently, biomass is able to support baseload or dispatchable generation while also contributing to waste valorization and circular bio-economy strategies [97,98].
  • Solar Thermal Energy: Solar thermal technologies capture solar radiation and convert it into thermal energy, which can be used directly for heating or indirectly for electricity generation through thermodynamic cycles such as steam turbines [99]. Concentrated solar power (CSP) systems employ optical components to concentrate solar energy and often incorporate thermal storage, enabling energy dispatch even in the absence of solar radiation [100]. This capability partially mitigates the inherent intermittency of solar resources [101,102].
  • Geothermal Energy (GEO): Geothermal systems exploit thermal energy stored beneath the Earth’s surface by extracting high-temperature fluids from underground reservoirs [103]. This energy may be utilized for electricity generation or directly for heating and industrial processes [103,104]. Due to the stable nature of geothermal resources, such systems provide highly reliable and continuous renewable generation with minimal variability [105].
  • Hydropower (Hydro): Hydropower systems generate electricity by converting the kinetic and potential energy of water into mechanical energy through turbine–generator units [106]. Reservoir-based hydropower offers significant operational flexibility; stored water can be released on demand, enabling dispatchable generation and ancillary services such as frequency regulation, spinning reserve, and load following [107]. This high degree of controllability makes hydropower one of the most versatile RES technologies in power system operation [108,109].
Such technologies are deployed across multiple spatial scales, from centralized utility-scale plants to distributed units embedded in local energy networks. As their penetration increases, these systems introduce significant variability and uncertainty into grid operation [110], making traditional control strategies increasingly inadequate for ensuring reliable and efficient performance [28].

3.2. Smart Grid Types

The integration of RES takes place across several interconnected layers of modern energy infrastructure, forming a hierarchical architecture that spans large-scale power grids, localized microgrids, and distributed building energy systems [111]. These RES-integrated energy frameworks are among the most frequently examined in the literature, as they capture the full spectrum of energy system scales, ranging from centralized transmission networks to community-level systems and individual buildings [69,111]. Although each framework operates under distinct physical constraints and control requirements, all require effective coordination between RES generation and demand in order to maintain stable operation. More specifically:
  • Power Grids: At the largest scale, power grids comprise extensive transmission and distribution networks that deliver electricity from generation facilities to consumers over wide geographic regions [112]. In these systems, RES are typically integrated through utility-scale solar plants, onshore and offshore wind farms, large hydropower facilities connected to high-voltage networks, biomass plants, and concentrated solar power systems [113]. Under such conditions, system operators must continuously balance generation and demand while maintaining voltage and frequency within acceptable limits [114]. Because RES output can vary rapidly, the coordinated management of multiple resources is essential for preserving network stability [115,116,117].
  • Microgrids: At an intermediate scale, microgrids integrate distributed RES units such as rooftop PV, small wind turbines, biomass generators, and hybrid renewable systems with energy storage devices, controllable loads, and in some cases backup diesel generators or combined heat and power (CHP) units [118]. Microgrids may operate either in grid-connected mode, exchanging power with the main utility network [119], or islanded mode, where supply and demand must be balanced locally [120,121]. Within these environments, energy management systems are responsible for coordinating distributed resources while ensuring economic efficiency, reliability, and resilience [69].
  • Buildings: At the most distributed scale, building energy systems integrate RES directly on the demand side of the electricity network [122]. Such systems may include rooftop PV, solar thermal collectors, geothermal heat pumps (GHP), building-integrated photovoltaics (BIPV), small vertical wind turbines, and battery storage, thereby enabling local generation and energy management [123,124]. In such settings, renewable generation interacts with subsystems such as HVAC, domestic hot water, electric vehicle charging, and smart appliances [125]. Their operation requires continuous adjustment of energy consumption to align with renewable availability while maintaining indoor comfort and minimizing operating costs [43,45].
Table 3 summarizes renewable energy installations and their operational role across smart grid frameworks consisting of power grids, microgrids, and building energy systems. Together, such energy frameworks form a multi-layered smart energy ecosystem in which renewable resources are generated, distributed, and consumed. Effective coordination among these layers is essential for enabling high renewable penetration while maintaining system reliability and economic performance.

3.3. Control Objectives in RES-Integrated Smart Grids

The growing penetration of renewable energy resources is introducing substantial operational complexity into modern energy systems, primarily due to their stochastic generation patterns and limited dispatchability [45,126]. Conventional control approaches are often inadequate for capturing the dynamic and uncertain behavior of renewable-rich energy environments [45]. Within this context, reinforcement learning (RL) has emerged as a promising data-driven control paradigm capable of learning effective decision-making policies through continuous interaction with the energy system environment:
  • RL Control in RES-Integrated Power Grids: In large-scale power grid environments, RL-based controllers are primarily applied to system-level tasks that demand fast and adaptive decision-making. These applications include voltage regulation in distribution networks with high photovoltaic penetration [127], frequency stabilization in systems with large shares of wind generation [128,129], and economic dispatch under uncertain renewable output. RL agents can learn to adequately to coordinate grid-support resources such as on-load tap changers (OLTCs) [130], reactive power compensators [131], flexible loads [132], and battery energy storage systems [133,134] to maintain stable operating conditions while reducing operational costs.
  • RL Control in RES-Integrated Microgrids: In microgrid environments, the main role of RL is to optimize local energy management through the coordinated operation of distributed energy resources and storage systems [135,136]. Since microgrids often rely heavily on renewable generation, their operation requires continuous scheduling of energy flows among generation units, batteries, and loads [137]. RL algorithms enable controllers to learn strategies for managing renewable variability, reducing backup generator fuel consumption, lowering energy costs, and preserving stable operation, particularly under islanded conditions. In this setting, RL-based policies may dynamically determine when energy should be stored, consumed, or exchanged with the main grid.
  • RL Control in RES-Integrated Buildings: At the building energy system level, reinforcement learning is primarily applied to demand-side energy management problems [138,139]. Buildings equipped with renewable generation and storage must intelligently coordinate their energy consumption with local renewable production [139]. RL agents can learn to schedule HVAC operation, battery charging and discharging, appliance usage, and electric vehicle charging in ways that maximize the utilization of locally generated renewable energy while maintaining occupant comfort and reducing electricity costs [48]. Through continuous interaction with environmental conditions, occupancy patterns, and energy price signals, RL-based building controllers can progressively adapt their actions to achieve more efficient and flexible energy management [45].
Table 4 summarizes the typical control objectives of RL in smart grid environments across these three energy frameworks. Overall, RL offers a control paradigm that is well-suited to the uncertainty, nonlinearity, and multi-objective optimization challenges of renewable-integrated smart grids. By learning control policies directly from system data and operational feedback, RL-based control shows fruitful potential to improve the reliability, efficiency, and sustainability of future energy systems with high shares of renewable energy.

4. Mathematical Concepts of RL-Based Control for Smart Grids

Section 4 below outlines the fundamental principles of RL as applied to control in energy systems, covering its operational workflow, mathematical formulation, and algorithmic foundations. It provides a structured perspective on how RL formulates and solves decision-making problems, extends to multi-agent environments, and is implemented through different categories of learning algorithms. Collectively, these components establish the theoretical and practical framework necessary for applying RL in complex RES-integrated smart grid systems.

4.1. Steps of RL-Based Control

The RL control process in energy systems is formulated as an iterative interaction loop between an intelligent agent and its environment. A conceptual representation of this process is illustrated in Figure 3. The main steps of this interaction are summarized as follows [45]:
1.
Environment: The environment represents the energy system under control, such as a power grid, microgrid, or building energy management system. It comprises RES generators, energy storage systems, controllable loads, electricity market signals, and operational constraints. Such elements behave dynamically in response to external factors, including weather conditions, user behavior, and system demand.
2.
State Observation: Real-time data describing the current operating condition of the system are acquired through sensors, smart meters, and monitoring platforms. Typical variables include RES generation levels, energy demand, storage state-of-charge, electricity prices, grid voltage conditions, and environmental factors such as solar irradiance or temperature. Such measurements collectively define the state space perceived by the RL agent.
3.
RL Agent: The RL agent acts as the decision-making entity responsible for selecting actions that enhance system performance. Depending on the application, it may operate in a centralized manner (e.g., at the microgrid level) or within a decentralized multi-agent framework. Actions may involve generation dispatch adjustments, battery charge/discharge control, load modulation, or the regulation of voltage support devices.
4.
Control Action Execution: The selected action is implemented within the energy system through appropriate control interfaces or supervisory mechanisms. In power grids, this may involve reactive power support or transformer tap adjustments; in microgrids, generation and storage scheduling; and in buildings, HVAC control, battery management, or appliance scheduling.
5.
Environment Transition: Following the execution of the control action, the environment transitions to a new state. Renewable generation varies according to weather conditions, demand evolves over time, and storage levels are updated. This updated condition constitutes the next observed state for the RL agent.
6.
Reward Evaluation: A numerical reward is computed to evaluate the effectiveness of the selected action. The reward function reflects the operational objectives of the system. In power grids, this aspect may relate to voltage stability or congestion mitigation; in microgrids, to cost reduction or renewable utilization; and in buildings, to energy efficiency or occupant comfort.
7.
Policy Update: The RL algorithm uses the reward feedback to refine its decision policy. Through repeated interaction with the environment, the agent progressively learns strategies that maximize long-term performance under dynamic and uncertain operating conditions.
Through this continuous learning loop, RL-based controllers are able adapt to renewable variability, fluctuating demand, and dynamic market signals. As a result, RL offers a flexible control framework that can outperform conventional deterministic strategies in complex energy environments.

4.2. Mathematical Formulation of RL

The RL decision-making process is commonly modeled as a Markov Decision Process (MDP), which formalizes the interaction between the agent and its environment [140,141]. The MDP is defined by the tuple ( S , A , P , R , γ ) , where [140,141]:
  • S denotes the state space representing the possible operating conditions of the energy system (e.g., renewable generation levels, system demand, storage states, electricity prices).
  • A represents the action space, corresponding to the set of control decisions available to the agent (e.g., dispatch control, storage operation, load scheduling, voltage regulation).
  • P ( s | s , a ) defines the transition probability describing how the system evolves from state s to state s after action a is applied.
  • R ( s , a ) represents the reward obtained after executing action a in state s.
  • γ [ 0 , 1 ] is the discount factor that determines the relative importance of future rewards.
The objective of the RL agent is to learn an optimal policy π ( a | s ) that maximizes the expected cumulative reward [140,141]:
G t = k = 0 γ k R t + k .
The expected return associated with a state under policy π is expressed through the value function
V π ( s ) = E [ G t s t = s ] .
Similarly, the action-value function evaluates the quality of selecting action a in state s:
Q π ( s , a ) = E [ G t s t = s , a t = a ] .
In many model-free RL approaches, the Q-function is iteratively updated using temporal difference learning:
Q ( s , a ) Q ( s , a ) + α R + γ max a Q ( s , a ) Q ( s , a )
where α denotes the learning rate controlling how new information updates the value estimates [140,141].

4.3. Multi-Agent RL

Many RES-integrated energy systems involve multiple interacting decision-making entities, including distributed generators, storage units, grid operators, and building controllers. In such settings, the control problem can be naturally formulated within the framework of multi-agent reinforcement learning (MARL), where multiple RL agents learn cooperative or competitive behaviors in a shared environment [71,142].
In MARL, the system is commonly modeled as a stochastic game defined by the tuple ( S , { A i } , P , { R i } , γ ) , where S denotes the global state space, A i denotes the action space of agent i, P describes the state transition dynamics, and R i represents the reward assigned to each agent [142]. Each agent aims to maximize its expected return:
J i ( π i ) = E π 1 , , π N t = 0 γ t R i ( s t , a 1 , t , , a N , t ) .
Since all agents update their policies simultaneously, the environment becomes non-stationary from the perspective of each individual agent. As a result, effective coordination mechanisms and appropriate training strategies become essential.
Within MARL, the actor serves as the decision-making component that learns a policy π i ( a i | o i ) , mapping observations to actions; that is, it determines which action each agent should take [43,45]. In contrast, the critic acts as the evaluation component by estimating a value function, such as Q i ( s , a 1 , , a N ) or V i ( s ) , and assessing the quality of actions in terms of expected future rewards. In this sense, the actor performs policy optimization, whereas the critic performs value estimation, guiding the agent toward improved behavior through gradient-based updates. Building on this principle, the effectiveness of MARL in smart grid applications depends on three closely related design dimensions (see also Figure 4): the network architecture, the learning paradigm, and the coordination mechanism among agents [43,45].
  • Agents Structure: The network architecture determines how actor and critic functions are distributed across agents and how global system interactions are represented. In centralized critic–decentralized actor schemes, each agent learns an individual policy π i ( a i | o i ) , while a shared critic Q i ( s , a 1 , , a N ) evaluates joint actions using global information; this approach supports stable learning in strongly coupled systems such as power grids. By contrast, fully decentralized architectures rely on independent critics and actors, i.e., Q i ( o i , a i ) , which improves scalability but also increases non-stationarity because the agents operate without global awareness. Fully shared architectures assume homogeneous agents and learn a common policy π ( a | o ) , whereas hybrid structures combine shared and local components, such as a shared critic with partially shared actors or graph-based critics; this offers a practical compromise between scalability and coordination quality in large-scale smart grid settings [43,45].
  • Agents Training: The learning paradigm defines how information is used during policy optimization and plays a central role in addressing MARL non-stationarity. Centralized training with decentralized execution (CTDE) is the most widely adopted approach, in which agents are trained using global state–action information but execute their policies using only local observations, i.e., π i ( a i | o i ) , enabling both coordination and practical deployment. Fully decentralized training eliminates global information exchange and allows agents to update independently, which is attractive for privacy-sensitive or communication-constrained energy systems, although it often results in slower convergence. Fully centralized training learns a joint policy π ( a 1 , , a N | s ) , effectively reducing the problem to a single-agent formulation, albeit with limited scalability [43,45]. Hybrid approaches, including federated or periodically synchronized learning, seek to balance data locality, communication efficiency, and convergence stability in distributed smart grid applications.
  • Agents Coordination: The coordination mechanism determines how agents align their decisions within a shared and physically coupled environment. In implicit coordination, agents are linked indirectly through shared rewards or shared critics, e.g., R i = R global , which promotes system-level objectives such as cost minimization or power balancing without requiring explicit communication. Explicit coordination, on the other hand, involves direct information exchange such as message passing or shared latent representations, enabling agents to condition their actions on the states or intentions of others. This is particularly beneficial in tightly coupled tasks such as voltage regulation or congestion management [43,45]. Emergent coordination arises when cooperative behavior develops naturally through interaction, even in the absence of explicit coordination mechanisms. Hierarchical coordination, in turn, decomposes the system into multiple control layers such as aggregator–prosumer or grid–microgrid structures, offering improved scalability while reflecting the inherently multi-level organization of smart grids.
Overall, these three dimensions are closely interconnected rather than independent [45]. The selected architecture influences which training strategies are feasible, while the coordination mechanism determines the extent to which inter-agent coupling is captured by the critic. In smart grid applications, where distributed energy resources are physically and economically interconnected, the careful design of these elements is essential for achieving stability, scalability, and strong system-wide performance.
Although MARL is well suited to distributed smart-grid control, deploying it in large-scale systems with high renewable penetration is still difficult [143]. One challenge is non-stationarity; from the viewpoint of each agent, the environment keeps changing because the other agents are also updating their policies [144]. Another issue is scalability, since the state and action dimensions increase quickly when adding more PV inverters, EV chargers, buildings, prosumers, or microgrids. In addition, full communication between all agents is rarely practical because of bandwidth limits, latency, cybersecurity risks, and privacy concerns. For these reasons, recent MARL studies have increasingly used scalable coordination mechanisms such as attention-based critics, graph neural networks, graph attention networks, federated critics, hierarchical control, and mean-field or collective-strategy approximations [145,146].

4.4. Common RL Algorithms

Reinforcement learning algorithms used in RES-integrated energy systems are generally grouped into three main families: value-based methods, policy-based methods, and actor–critic approaches [43,45]; see also Figure 5. These paradigms differ in how they learn the optimal control policy and solve the underlying optimization problem [147]. Their suitability depends on the nature of the task, including whether the action space is discrete or continuous and whether the control structure is centralized or distributed [147].
In smart grid applications, value-based algorithms are commonly used for discrete decisions such as switching and scheduling [45]. Policy-based methods are more appropriate for continuous control tasks such as adjusting power dispatch or battery charging rates. Actor–critic methods combine both ideas, and are widely used in complex energy management problems with high-dimensional state spaces and nonlinear dynamics [43,45].

4.4.1. Value-Based Algorithms

Value-based RL methods estimate the expected cumulative reward of taking a given action in a given state. Based on these estimates, the agent selects actions that maximize long-term return and gradually improves its policy through repeated interaction with the environment. The literature includes many value-based approaches; the following are among the most representative in energy system control [148]:
  • Q-learning (Q-l): One of the most fundamental methods in this family is Q-learning [149]. Its objective is to estimate the optimal action-value function Q ( s , a ) , which represents the maximum expected return obtained by taking action a in state s and then following the optimal policy. The optimal Q-function satisfies the Bellman optimality equation [148,149]:
    Q ( s , a ) = E R ( s , a ) + γ max a Q ( s , a )
    where s is the next state reached after taking action a in state s. The Q-values are updated iteratively using the temporal-difference rule [149]:
    Q ( s , a ) Q ( s , a ) + α R + γ max a Q ( s , a ) Q ( s , a )
    where α is the learning rate. In energy applications, Q-learning has been used for switching decisions in distribution networks, storage scheduling in microgrids, and appliance coordination in building energy management systems [148].
  • Deep Q-Networks (DQNs): To overcome the limitations of tabular Q-learning in large state spaces, DQNs approximate the Q-function with a neural network parameterized by θ [149]:
    Q ( s , a ; θ ) Q ( s , a ) .
    The network is trained by minimizing the temporal-difference loss [149]:
    L ( θ ) = E R + γ max a Q ( s , a ; θ ) Q ( s , a ; θ ) 2
    where θ denotes the target-network parameters used to stabilize learning. DQNs have been widely applied in smart grids for tasks such as voltage regulation in distribution networks and energy scheduling in renewable-dominated microgrids [150].

4.4.2. Policy-Based Algorithms

Unlike value-based methods, policy-based RL algorithms optimize the control policy directly without explicitly learning a value function. The policy is parameterized as π θ ( a | s ) , where θ denotes the policy parameters and the objective is to maximize the expected cumulative reward [151]:
J ( θ ) = E π θ t = 0 γ t R ( s t , a t ) .
The policy parameters are updated using the policy gradient theorem [152]:
θ J ( θ ) = E π θ θ log π θ ( a | s ) Q π ( s , a ) ,
which allows the agent to improve its policy in the direction of higher expected reward.
  • Proximal Policy Optimization (PPO): A widely used method in this family is PPO, which introduces a clipped surrogate objective to ensure stable policy updates. Its objective is given by [152]
    L C L I P ( θ ) = E min r t ( θ ) A t , clip ( r t ( θ ) , 1 ϵ , 1 + ϵ ) A t ,
    where
    r t ( θ ) = π θ ( a t | s t ) π θ o l d ( a t | s t )
    and A t is the advantage function, which estimates the relative benefit of selecting action a t in state s t .
Policy-based methods are especially useful in renewable-integrated energy systems with continuous control variables, such as power flow regulation, battery dispatch, or HVAC modulation in buildings [153].

4.4.3. Actor–Critic Algorithms

Actor–critic algorithms combine value-based and policy-based learning through two interacting components: an actor, which defines the control policy, and a critic, which evaluates the quality of the selected actions. The actor updates the policy parameters according to the policy gradient [151]:
θ J ( θ ) = E θ log π θ ( a | s ) A ( s , a )
where the advantage function A ( s , a ) measures the relative benefit of choosing action a over the average policy behavior. The critic estimates the value function using temporal-difference learning [151]:
V ( s ) V ( s ) + α R + γ V ( s ) V ( s ) .
Several actor–critic variants have been proposed for continuous control tasks:
  • Deterministic Policy Gradient (DDPG): DDPG extends deterministic policy gradients using deep neural networks. The actor produces deterministic actions [154]:
    a = μ ( s | θ μ )
    while the critic evaluates the action-value function:
    Q ( s , a | θ Q ) .
    The actor parameters are updated using the gradient [154]:
    θ μ J = E a Q ( s , a | θ Q ) θ μ μ ( s | θ μ ) .
  • Soft Actor-Critic (SAC): Another widely used algorithm is SAC, which includes entropy regularization to promote exploration. Its objective becomes [155]
    J ( π ) = E t R ( s t , a t ) + α H ( π ( · | s t ) ) ,
    where H is the entropy of the policy distribution and α controls the exploration level.
  • Twin Delayed Deep Deterministic Policy Gradient (TD3): TD3 improves DDPG by addressing overestimation bias and function-approximation error in actor–critic learning. It uses two critic networks, delayed actor updates, and target policy smoothing to stabilize continuous-control learning [156]. The critic target is computed using the minimum of two target critics:
    y = R + γ min i = 1 , 2 Q θ i ( s , a )
    where Q θ 1 and Q θ 2 are the two target critics. The smoothed target action is obtained as
    a = μ θ ( s ) + ϵ , ϵ clip ( N ( 0 , σ ) , c , c ) .
    This mechanism reduces optimistic value estimates and improves policy stability in continuous-control problems such as inverter regulation, ESS dispatch, frequency control, and HVAC modulation.
Actor–critic methods are particularly well suited to large-scale renewable-integrated energy systems, where continuous actions and high-dimensional state spaces must be handled simultaneously [45]. For this reason, they have been widely used in smart grids for optimal power dispatch, in microgrids for energy management, and in buildings for intelligent energy control.

4.4.4. Emerging RL Techniques: Distributional RL and Evolutionary Policy Optimization

Beyond the classical families of value-based, policy-based, and actor–critic algorithms, recent RL research has increasingly moved toward methods that are better suited to uncertainty, risk sensitivity, and more difficult control landscapes. One important direction is distributional RL, which goes beyond estimating only the expected return of an action and instead learns the full distribution of possible returns [157]. This is particularly meaningful for RES-integrated smart grids, where future rewards can vary widely because of stochastic PV and wind generation, uncertain demand, electricity price volatility, and occasional disruptive events. In this context, methods such as C51, QR-DQN, IQN, and Rainbow DQN can support more risk-aware decision making in applications including storage scheduling, EV charging, market bidding, voltage regulation, and resilience-oriented microgrid control [157,158,159]. In the literature, Rainbow DQN has already appeared in microgrid battery scheduling problems under uncertain demand, renewable generation, and price conditions, showing improved performance compared with conventional DQN and continuous-control benchmarks.
Another emerging direction is evolutionary policy optimization, in which policies are improved through population-based or gradient-free search rather than relying only on gradient-based updates [160]. Examples include evolution strategies, genetic policy search, CMA-ES, neuroevolution, and population-based training [161]. These methods can be especially useful when the reward landscape is non-smooth, multi-objective, highly constrained, or difficult to differentiate, which is often the case in smart grid control problems involving discrete switching actions, nonlinear device constraints, degradation costs, and competing technical and economic goals. Although evolutionary policy optimization is still less common than DQN, DDPG, SAC, PPO, and TD3 in RES-integrated energy applications, it offers promising potential for controller tuning, policy initialization, hybrid RL–optimization schemes, and multi-objective coordination of DERs, ESS, EVs, and flexible loads. More broadly, these emerging RL approaches can complement conventional DRL by improving uncertainty representation, robustness, and exploration in high-dimensional smart-grid environments [160].

5. Key Attributes and Summaries of RL-Based Applications

5.1. Tables Description

Before presenting detailed summaries of the different applications, this section introduces the key attributes used to categorize each reviewed study according to the smart grid domain under consideration, namely, power grids in Table 5 with their Summaries Table 6; Microgrids in Table 7 with their summaries in Table 8; and buildings in Table 9 with their summaries in Table 10. These tables provide a systematic classification of the selected studies based on several core characteristics, including the reference, year of publication, RL method, agent type, RES type, controlled equipment, grid subtype, data type, and number of citations. This structured organization supports a comprehensive understanding of the methodological approaches reported in the literature. More specifically, the summarized table columns are defined as follows:
  • Ref.: Indicates the reference of the corresponding application in the first column.
  • Year: Indicates the publication year of each RL application.
  • Method: Indicates the specific RL algorithmic methodology employed in each study.
  • Agent: Indicates the agent type associated with the corresponding RL methodology, i.e., single-agent or multi-agent RL approach.
  • FH/TS: Indicates the forecast horizon (FH) and control time step (TS) of the RL approach, expressed in hours (h) and minutes (m).
  • Baseline: Indicates the baseline control approaches used in each study to evaluate the RL method, such as optimal power flow (OPF), droop control (Droop), model predictive control (MPC), fixed control strategy (Fixed), rule-based control (RBC), particle swarm optimization (PSO), genetic algorithm (GA), proportional–integral–derivative (PID) control, stochastic programming (SP), second-order cone programming (SOCP), dynamic programming (DP), artificial neural network (ANN), multiple linear regression (MLR), mixed-integer linear programming (MILP), and alternating-direction method of multipliers (ADMM), among others.
  • RES: Indicates the type of integrated RES technology considered in the corresponding smart grid environment, such as photovoltaic (PV), wind turbine (WT), biomass (BIO), solar water heating (SWH), and hydroelectric Energy (Hydro).
  • Integration: Indicates the controllable equipment involved in each application within the smart grid environment. These may include distributed generation (DG) (e.g., photovoltaic systems, wind turbines, biomass units, and conventional generators), energy storage systems (ESS) (e.g., battery storage units and thermal storage systems), electrical loads (Load) (e.g., residential, commercial, and industrial demand), electric vehicles (EVs), or electric vehicle charging stations (EVCS) (e.g., EV fleets and charging infrastructure). The grid infrastructure (Grid) (e.g., transmission/distribution lines, buses, substations, and power flow constraints) is considered together with control-oriented devices, which are divided into power electronics (PE) (e.g., inverters, SVC, STATCOM, and converter-based devices) and grid control equipment (GCE) (e.g., OLTC transformers, capacitor banks, voltage regulators, and frequency control mechanisms).
  • Type: Indicates the subtype of the smart grid environment. For power grids, this may include active distribution networks (ADNs) and transmission systems; for microgrids, grid-connected, networked, community, and islanded microgrids; and for buildings, residential, office, laboratory, industrial, district-scale, and community-scale systems.
  • Citations: Indicates the number of citations of each work according to Scopus.
The abbreviation “N/A” denotes elements that were not identified in the corresponding tables and figures. In addition, a brief summary of each application is provided in Table 6, Table 8, and Table 10 for power grid, microgrid, and building RL applications, respectively. These tables include the following columns:
  • Author: Indicates the author and reference of each application in the first column.
  • Summary: Provides a brief description of the methodology and main outcomes of the corresponding application.
Each application is presented in the same order and with the same reference number as in the corresponding key attribute tables, allowing for straightforward cross-referencing.

5.2. RL for Power Grid Applications

The key attributes of RL applications in power grids are presented in Table 5, while their corresponding summaries are provided in Table 6.

5.3. RL for Microgrid Applications

The key attributes of RL applications in Microgrids are presented in Table 7, while their corresponding summaries are provided in Table 8.

5.4. RL for Building Applications

The key attributes of RL applications in buildings are presented in Table 9, while their corresponding summaries are provided in Table 10.

6. Evaluation

To provide a structured and comparative overview of existing reinforcement RL applications in RES-integrated smart grid environments, including power grids, microgrids, and buildings, this review examines the literature through essential key attributes:
1.
RL Methods: Examines the main RL algorithms used in the literature, highlighting their core structures, features, and emerging trends.
2.
Agent Architectures: Investigates the dominant agent design approaches, with particular attention to multi-agent RL and its dimensions.
3.
Reward Functions: Analyzes how reward functions are designed across RL-based control applications in smart grid systems.
4.
Baseline Control: Reviews the conventional or benchmark control strategies used to evaluate RL performance.
5.
Utilized Data: Examines the types of data used to develop, train, and test RL-based frameworks.
6.
Control Objectives: Identifies the most common control goals in RL-based smart grid applications and discusses their main characteristics.
7.
RES Types and Integrated Equipment: Explores the trends associated with different renewable energy sources and other integrated technologies considered in RL applications.
8.
Grid Types: Reviews the main types of power grids, microgrids, and building energy systems considered in the literature, highlighting common trends and limitations.
9.
Performance Comparison: Reviews the integrated research works that offer the most significant contributions to the field based on performance and other aspects in a comparative manner.
10.
Simulation Tools: Reviews the simulation and computational tools that practitioners have utilized to provide results in energy management applications for power grids, microgrids, and buildings.
These dimensions capture the main factors involved in the design, implementation, and evaluation of RL-based controllers. Examining the literature through these perspectives helps to reveal how different RL approaches match specific system configurations, operational constraints, and control objectives while providing a clearer understanding of current research trends and remaining challenges. The evaluation of each key attribute is provided sequentially for power grid, microgrid and building RL implementations.

6.1. RL Methods

RL applications show clear similarities in algorithm selection and control formulation across power grids, microgrids, and building energy systems. Actor–critic methods, mainly DDPG, SAC, and TD3, dominate the reviewed studies (Figure 6 Left and Right). Such approaches are well suited to continuous control of DERs, HVAC systems, and grid devices under nonlinear and uncertain conditions [169,200,225]. In particular, DDPG and SAC appear most frequently due to their compatibility with continuous action spaces, while TD3 is often used to reduce overestimation bias and PPO is preferred when stable policy updates and easier tuning are more important [168,175,237]. Value-based methods such as DQN, D3QN, and Rainbow RL remain relevant mainly for discrete or hybrid decisions, including appliance scheduling and capacitor switching [201,224,232]. Classical Q-learning is still used in simpler and lower-dimensional problems, especially in microgrids and residential energy applications where interpretability and low computational cost are useful [191,206,222]. Model-based and planning-oriented methods such as MCTS and MuZero-inspired approaches appear less frequently, and are mainly used in more complex uncertainty-driven settings [194,198]. Overall, the reviewed literature does not point to a single universal RL method but rather to a domain-dependent selection of algorithms.

6.1.1. RL in Power Grid Applications

In power grid applications, the literature shows a clear shift from simple value-learning methods towards actor–critic and policy gradient algorithms. This trend occurs mainly because modern grid control problems involve continuous, nonlinear, and constraint-sensitive actions. DDPG-type methods were among the earliest dominant choices due to their ability to directly handle generator set points, inverter reactive power, SVC actions, ESS dispatch, and active/reactive power regulation without discretizing the action space [164,170,177,178]. This is important because coarse discrete actions may reduce voltage quality, increase losses, or limit the flexibility of inverters and storage systems [164,177,178]. However, DDPG is sensitive to exploration noise, critic instability, and overestimation bias, which explains the later use of TD3 and SAC as more stable alternatives [168,175,181,183]. TD3 improves reliability through twin critics and delayed policy updates, making it useful for stability-sensitive tasks such as load frequency control and voltage regulation [175,181]. In contrast, SAC uses entropy-regularized exploration, which is valuable under stochastic RES generation and uncertain load patterns [169,183]. PPO-based methods form a separate group in power grid applications. They are mainly used when stable learning and bounded policy updates are more important than fine continuous-control precision, such as in OPF approximation, outage management, and voltage regulation under uncertainty [162,181,182,185]. The clipped update mechanism of PPO is useful in power grid settings, where unstable policy changes may produce infeasible or unsafe behavior during training [162,185]. From this evaluation of the literature it appears that DDPG, TD3, and SAC are mainly preferred for high-resolution continuous actuation, while PPO is often selected for robust and more general policy learning [162,164,168,183,185]. Value-based RL approaches remain important, though mostly for discrete operational decision-making. DQN-type methods seem better suited to switching, capacitor bank control, network reconfiguration, topology restoration, braking resistor activation, and market/action selection, where the action space is naturally discrete [165,172,185,188]. Such methods are simpler and more interpretable than actor–critic approaches, but scale poorly in continuous, high-dimensional, or combinatorial action spaces [165,172]. Thus, DQN is not obsolete but has become more specialized for the discrete control layer of smart grid operation, while actor–critic methods dominate continuous grid control tasks [164,165,170,172,177,178,188]. A final trend in power grid research applications concerns the growing use of safe and constrained RL. Recent studies increasingly evaluate RL not only by cost, voltage quality, or speed but by feasibility and operational admissibility. Safety layers, action masking, constrained policy optimization, stability-constrained neural policies, and physics-informed penalties are widely used to reduce infeasible or unstable actions [164,171,174,177,184]. This is a significant trend, since reward penalties may reduce violations statistically but do not always guarantee safe behavior in critical infrastructure [171,174]. Constrained policy methods explicitly separate reward maximization from constraint satisfaction, while action masking and physics-guided learning embed operational limits more directly into the learning process [171,177,184]. Although these approaches are more complex, they are closer to the requirements of real deployment [164,171,174,184].

6.1.2. RL in Microgrid Applications

In microgrid applications, the RL landscape illustrates larger heterogeneity than in power grid voltage control. This is because microgrids combine scheduling, storage dispatch, frequency regulation, converter control, market participation, and multi-energy coordination; as a result, no single RL family seems to dominate the microgrid energy framework. Value-based methods are mainly used for discrete scheduling, DDPG-type methods for continuous dispatch, TD3 and SAC for more uncertain or dynamic control, and PPO/A3C-type methods for stable policy learning under stochastic operation [191,192,193,200,201,212,219]. This diversity reflects the fact that microgrid controllers must handle RES uncertainty, storage degradation, grid exchange, temporal energy balance, and user-side flexibility at the same time [197,203,208,217,220]. Q-learning, DQN, and enhanced value-based RL remain useful when the control problem can be expressed through discrete actions such as battery scheduling, appliance control, trading, or community energy management [192,206,209,216]. Their advantages are simplicity, interpretability, and low computational cost, which are attractive in residential and community microgrids [206,216]. However, as in the case of power grids, they become less suitable for continuous or high-dimensional dispatch because discretization reduces control precision and increases the complexity of the Q-function [193,200,213]. Enhanced variants such as Rainbow DQN offer improved robustness under stochastic prices and RES generation, but remain best suited to discrete scheduling layers [201]. DDPG-type actor–critic methods are preferred when microgrid operation requires continuous dispatch of batteries, diesel generators, microturbines, converters, or flexible resources [193,196,199,208,215]. They avoid coarse action discretization and directly produce continuous control signals, which is important for SOC management, BESS power control, DG dispatch, and secondary voltage/frequency regulation [193,199]. Still, DDPG can be sensitive to hyperparameters, uncertainty, and critic overestimation; more recent studies often address this by adding prioritized replay, recurrent structures, automated tuning, or hierarchical training logic [193,196,200,210,213]. DDPG is generally less robust than TD3, SAC, or PPO unless supported by some of these stabilization mechanisms [200,213,219]. Thus, TD3 and SAC represent more mature actor–critic options for uncertain or fast microgrid dynamics. TD3 is useful in tasks where overestimation control and stable continuous actions are important, such as frequency regulation, virtual synchronous generator tuning, and converter duty cycle control [212,213]. Moreover, SAC seems more suitable when uncertainty and exploration are central, since entropy maximization supports smoother adaptation under variable RES, prices, and loads [203,204,214,217,220]. Compared with DDPG, SAC generally provides more robust exploration and policy adaptation, as seen in [203,204,214]. PPO and A3C-type methods are mainly used where stable learning and reliable policy improvement are more important than aggressive value exploitation. PPO appears in microgrid scheduling, MT/ESS energy management, VPP coordination, and real-time operation with PV, wind, ESS, diesel units, EVs, and grid transactions [200,218,219]. Its clipped update mechanism helps to avoid destructive policy changes that could lead to infeasible SOC trajectories or inefficient RES utilization [200,219]. A3C-type learning is also relevant for risk-sensitive energy management in microgrids, where the goal includes both cost reduction and shortfall risk mitigation [197]. Thus, PPO/A3C methods are especially attractive for higher-level scheduling under uncertainty [197,200,218,219]. Finally, planning-oriented and structure-preserving methods such as Dyna-Q, MuZero, and MPC-parameterized RL appear in problems where model-free learning alone is too sample-inefficient or insufficiently aware of temporal constraints [191,194,207]. Such approaches combine learning with planning, prediction, or embedded control structure, making them better aligned with forecast dependence, time-varying storage constraints, and energy balance feasibility [194,207].

6.1.3. RL in Building Applications

In building energy systems, RL methods are shaped by energy cost, PV self-consumption, thermal comfort, storage operation, EV charging, appliance flexibility, and stochastic weather or occupancy. The literature separates into value-based RL for discrete scheduling, DDPG/TD3 for continuous HVAC and storage control, SAC for uncertainty-aware coordination, and PPO/MAPPO for stable high-level scheduling [222,223,224,225,227,240,241,243]. Compared with power grids, RL in buildings places more emphasis on delayed thermal dynamics, comfort, storage constraints, and demand response [223,228,229,238]. Q-learning, DQN, and D3QN-type methods remain relevant when the problem is naturally discrete, for instance appliance scheduling, EV charging, DHW operation, flexible-load activation, or household energy management [222,224,228,232,242]. These methods offer interpretable action selection and align well with user-side decisions [224,242]; however, they are less suitable for continuous actions such as HVAC setpoints, battery power, thermal storage levels, or voltage gain signals [223,226,237]. Advanced variants such as D3QN with PER and feasibility screening can improve stability and constraint awareness, but remain better suited to discretized control layers [232]. DDPG is widely used for continuous HVAC, ESS, BESS, and thermal storage control because it directly outputs real-valued actions without discretizing comfort or storage decisions [223,230,234,236]. This is important because coarse actions may cause discomfort, excessive battery cycling, or inefficient PV use [223,230]. Still, DDPG may become unstable under stochastic occupancy, PV variability, and long-delay thermal dynamics, which motivates the use of TD3 and SAC in more demanding settings [227,240,241]. TD3 improves DDPG-like control by reducing overestimation bias, but may be less attractive when broad exploration is needed [227,241]. SAC is particularly suitable for buildings with uncertain RES generation, thermal storage, HVAC flexibility, and multi-building coordination. Its entropy-regularized policy supports broader exploration under variable PV, occupancy, weather, prices, and comfort constraints [221,225,226,235,237,239]. SAC is generally more robust to stochastic operation than DDPG, handles continuous control more naturally than DQN, and is often more specialized for smooth energy resource coordination than PPO [225,226,235,237], which explains its frequent use in district- or cluster-level building control [221,226,235,239]. PPO is also common in residential HEMS and resilient building scheduling because it provides stable policy updates across such mixed objectives as cost, comfort, PV self-consumption, EV constraints, battery operation, and power quality [238,240,241,243,245]. Although it may be less naturally exploratory than SAC, it is easier to stabilize and has shown strong performance against DQN, DDPG, A2C, TD3, SAC, MILP, and other baselines [240,241,243]. PPO is especially useful for high-level scheduling and resilient HEMS operation [241,243,245]. Finally, A2C/federated and MAPPO-type approaches reflect the growing importance of privacy, scalability, and policy-sharing efficiency across residential communities and building clusters [231,234,238]. These methods are more complex than centralized SAC or PPO, but address the issue that building data often cannot be pooled into a single controller due to being distributed across many homes, occupants, and prosumers [231,234].

6.2. Hybrid RL Methodologies

Across the reviewed hybrid RL applications, the most common hybridization category is RL combined with mathematical optimization or planning. Such combinations appear in many cases, including convex OPF, SOCP, MISOCP-MPC, MILP, MPC-parameterized RL, evolutionary optimization, and Lagrangian or safety-constrained optimization [165,173,190,194,196,203,205,207,211,217]; see Figure 7. This form is common because power grids and microgrid frameworks consist of strongly constrained systems in which voltage, power flow, line capacity, SOC, dispatch, and reserve constraints cannot be left to trial-and-error learning. In these methods, RL handles sequential adaptation and uncertainty, while optimization preserves feasibility and scheduling structure.
The second commonly deployed hybridization group in smart grid frameworks consists of RL combined with classical control, physics-based control, or domain rule embedding [176,180,191,215,236]. Here, RL does not replace conventional control but instead tunes PI/PID gains, optimizes Q-V droop parameters, refines rule-based battery dispatch, or uses PoPA/pinch analysis logic to guide energy targets. These approaches are often easier to justify in deployment because the final action remains connected to known engineering logic.
A third commonly encountered group combines RL with forecasting, surrogate models, or data-driven modeling [178,186,196,205,214]. Such methods use DNN surrogate power flow models, demand/PV forecasting, interval prediction, probabilistic wavelet fuzzy neural networks, or broader DNN–IoT structures. Their main role is to improve anticipation, sample efficiency, and uncertainty representation. Some of these examples are direct learning–control hybrids, while broader DNN–IoT frameworks are better viewed as AI–framework hybrids [186].
A smaller hybrid RL group combines RL with safety or imitation-learning mechanisms [162,204,238]. These practices address unsafe exploration and poor sample efficiency by using expert initialization, action guidance, or safety evaluation networks. Overall, it seems that power grid hybridization mainly targets feasibility and network security, microgrid hybridization combines dispatch optimality with uncertainty adaptation, and building hybridization is mainly used to preserve expert comfort logic or accelerate multi-agent learning.

Critical Comparison of RL Methods

The target of this comparison is to clarify which RL method families are most appropriate for different RES-integrated smart grid applications depending on the action-space structure, uncertainty, coordination needs, and operational safety. Overall, value-based methods are best suited to discrete scheduling problems such as appliance operation, capacitor switching, OLTC/capacitor coordination, and simple BESS charge/discharge logic, since their decision structure naturally matches finite action spaces [165,201,206,224,232,233]. However, they become less attractive when the control problem involves continuous physical variables such as reactive power, HVAC setpoints, battery power, SVC outputs, M-SOP dispatch, or EV charging rates, where actor–critic methods provide a more natural formulation [164,175,177,223,227]. This explains why DQN-type approaches appear mainly in capacitor scheduling, appliance coordination, and discrete battery EMS applications, while being less dominant in fast inverter-based voltage control.
By contrast, DDPG, TD3, and SAC-type actor–critic methods are better suited to continuous real-time control. DDPG and TD3 are particularly appropriate when the control action is deterministic and physical, such as BESS power, SVC reactive power, M-SOP dispatch, HVAC/battery setpoints, or frequency control parameter tuning [164,175,177,223,227]. TD3 is generally preferable to vanilla DDPG when stability and overestimation bias are concerns, as observed in renewable-integrated frequency control and building/off-grid storage operation [175,227,241]. SAC is more suitable when the environment is strongly stochastic, since its entropy-regularized exploration improves robustness against RES intermittency, uncertain demand, price fluctuations, and flexible-load variability [203,204,217,221,235,237].
For large RES-rich distribution networks, building districts, and microgrid clusters, the evidence points clearly toward MARL rather than single-agent RL. A single centralized agent faces dimensionality, privacy, and scalability problems when many PV inverters, homes, greenhouses, EVs, or microgrids act simultaneously. MARL frameworks such as MADDPG, MATD3, MASAC, MAAC, and MAPPO are more suitable in this context because they allow each controllable asset or subsystem to act as an agent while still learning coordinated behavior through centralized critics, shared rewards, attention mechanisms, collective strategy models, or partially-shared information [163,168,183,184,220,239,245].
A critical distinction is that safe and constrained RL is not only an algorithmic refinement but a deployment requirement for power system applications. In voltage regulation, frequency control, OPF, and multi-energy microgrid operation, constraint violations can be unacceptable; therefore, CPO, action masking, safety layers, Lipschitz-constrained policies, and Lagrangian SAC are more relevant for grid-critical control than unconstrained DRL, even when the latter achieves high rewards [164,171,174,217,229]. These methods are especially important when the RL agent directly controls grid devices such as PV inverters, OLTCs, SVCs, BESS units, M-SOPs, CHP units, or gas/thermal network components, rather than only recommending schedules.
Finally, the strongest practical direction appears to be hybrid RL, where RL is combined with OPF, MPC, forecasting, MCTS, surrogate power flow models, physics-guided embeddings, or imitation learning. Pure model-free RL is attractive because it avoids explicit models, but many real smart grid problems still require feasibility guarantees, safety margins, and economic optimality. Hybrid schemes often offer the best compromise, with RL providing fast adaptive decisions under uncertainty while optimization, forecasting, or physics-guided modules preserve operational credibility. This is particularly visible in OPF acceleration, voltage control with two time scales, community flexibility in smart buildings, VPP scheduling, and coordination of multi-energy microgrids [162,173,184,189,190,194,203,207,218].
As summarized in Table 11, the suitability of a given RL family is closely related to factors such as the nature of the action space, the degree of uncertainty, the need for safety and constraint satisfaction, and the level of coordination required in each RES-integrated application.

6.3. Agent Architectures

MARL accounts for approximately 48% of the reviewed studies, showing substantial but not dominant adoption compared with single-agent RL (Figure 8-Right). Its use reflects the distributed nature of smart grid control, in which DERs, microgrids, buildings, and prosumers may all act as interacting decision-making agents. Across power grids, microgrids, and buildings, the literature shows a clear tendency towards multi-agent actor–critic methods, consistent with broader MARL developments [71,142]. MARL DDPG-based approaches are the most common thanks to their support for continuous control tasks such as voltage regulation and energy scheduling [163,166,167,202,208]. TD3 and SAC are also increasingly used due to their improved stability and robustness [168,175,183,235]. Value-based methods such as Q-learning and DQN remain relevant for discrete decision-making and market-oriented problems [192,221,232], while PPO-based MARL is beginning to appear in scalable coordination settings [238] (see Figure 8-Left).
In power grid applications, MARL is mainly based on CTDE actor–critic architectures (See Figure 9-Left). This is especially common in volt/VAR control, frequency regulation, and coordinated energy scheduling, where agents are geographically distributed but coupled through network constraints. Recent studies usually preserve decentralized execution at the level of inverters, generators, or area controllers while using centralized or partially centralized critics to capture power flow and inter-area dependencies [163,166,167,168,173,176,179,183,184]. Newer works also improve the critic structure through attention mechanisms [163,168,173,183], surrogate or Gaussian process models for faster power flow evaluation [167,169], and graph- or topology-aware encoders for representing network structure [184].
MARL architectures appear to be more diverse in microgrid frameworks, where the control objectives include market participation, uncertainty management, multi-energy coordination, and community-level scheduling (See Figure 9-Center). While CTDE remains common [202,204,208,211], several studies also incorporate risk-aware learning, explicit communication, game theory, or evolutionary optimization. For example, risk-sensitive MARL addresses tail-risk operation under renewable uncertainty [197], while locality-aware communication and spatial discounting reduce the need for a single centralized critic [195]. Other works combine MARL with game-theoretic or multi-objective optimization structures, showing that microgrid coordination often requires broader decision-making frameworks beyond pure MARL [210,211].
In building applications, MARL is shaped less by network centralization and more by privacy, autonomy, and low communication requirements (See Figure 9-Right). This explains the frequent use of decentralized, federated, and sequential coordination schemes in residential and community settings [209,221,222,231,233]. CTDE-based building MARL is also expanding [232,234,235,238], but is usually adapted to the building context through value decomposition for heterogeneous agents [232], attention-based critics for scalable inter-building coordination [235], or imitation-enhanced MAPPO for faster convergence [238]. Since building agents often represent households, appliances, or prosumers with private preferences, federated aggregation [231,234], priority-based scheduling [222,233], and market-based cooperation [198,209] are especially relevant.
Agent Scalability and Communication Challenges: The reviewed MARL applications highlight three major barriers to large-scale deployment: non-stationarity, scalability, and communication limits. Non-stationarity appears because each agent learns while the other agents are also changing their policies, so the environment is not fixed from the viewpoint of any individual controller. Scalability becomes a concern when many distributed assets, such as PV inverters, EV chargers, buildings, prosumers, or microgrids, are modeled as separate agents. Communication is another practical limitation, since real smart grids cannot always exchange full observations, actions, or gradients among all controllers because of bandwidth, latency, cybersecurity, and privacy constraints. The most common solution deployed in the reviewed literature concerns CTDE. In this setting, centralized critics use global information during training while decentralized actors make local decisions during deployment; according to evaluations, this structure is widely used in voltage control and microgrid coordination, where it helps reduce non-stationarity while preserving distributed operation [163,168,173,183,204,234,235]. However, it needs to be mentioned that standard CTDE becomes harder to scale as the number of agents increases, since the centralized critic must process high-dimensional joint state–action information. For this reason, more recent studies use attention mechanisms, graph-based learning, federated aggregation, and collective strategy approximations to improve scalability.
Attention-based MARL is another important direction. Instead of treating all agents as equally relevant, the critic learns which neighboring or interacting agents are most important for the current decision. This is useful in voltage control, demand response, and building cluster coordination, where interactions are usually sparse rather than fully connected. Attention-enhanced MADDPG, MATD3, MASAC, and MAAC methods have been used to weight inter-agent dependencies selectively and improve coordination without requiring full all-to-all communication [163,168,173,183,235,239]. In this way, attention helps to reduce the effective interaction space to the most relevant agent relationships.
A second direction is the use of graph neural networks (GNNs) and graph attention networks (GATs). Unlike generic attention mechanisms, graph-based MARL can directly use the physical topology of the power network, where buses, feeders, inverters, loads, and DERs are naturally represented as nodes and edges. For example, graph-enhanced MARL frameworks embed grid topology and electrical sensitivity information into the critic, allowing the controller to learn which agents are physically coupled through the network [184,189]. This is especially useful for large active distribution networks, where it reduces the need for manually designed global features and can support generalization across feeders with different topologies. In this sense, GNN/GAT critics provide a scalable middle ground between fully centralized critics and purely local decentralized agents.
A third direction includes federated/hierarchical coordination. Federated MARL reduces direct data sharing by aggregating model parameters or abstracted system-level information instead of raw local states and actions [231,234]. Hierarchical schemes reduce coordination complexity by assigning different roles to aggregators, VPPs, DSOs, microgrids, or device-level agents [210,218,233,245]. Such structures are particularly useful for EV charging stations, prosumer communities, and VPPs, where direct coordination among every EV or household is not practical. Mean-field or collective strategy methods follow a similar idea by replacing the explicit modeling of all neighboring agents with an aggregate description of their average behavior. In the reviewed studies, this appears through local collective strategy models in which each residential microgrid estimates the behavior of nearby microgrids without accessing their private observations [220]. This is promising for large EV charging and prosumer networks, as each agent can respond to aggregate congestion, price, or demand signals instead of communicating with all other agents.
Table 12 indicates that the field is gradually moving beyond standard independent-agent MARL towards topology-aware, communication-efficient, and population-scalable coordination. For moderate-size systems, CTDE and attention-based critics are currently the most common choices. For large active distribution networks, GNN/GAT-based MARL appears especially suitable because it uses the physical graph structure of the grid. For very large populations of flexible edge devices, such as EV chargers, buildings, and prosumers, mean-field-style and collective strategy methods are promising because they avoid explicit all-to-all communication. However, true mean-field game formulations are still underrepresented in RES-integrated smart-grid RL, which makes them an important direction for future work on scalable EV charging coordination, transactional energy communities, and VPP-level flexibility aggregation.

6.4. Reward Functions

Reward design reflects how RL is used in smart grid control. Most studies do not optimize a single objective, instead combining several competing goals into one scalar reward through weighting factors [164,200,208,225,238] (see Figure 10-Right). This is expected, since RES-rich energy systems involve trade-offs between cost, security, flexibility, comfort, self-consumption, and reliability. In multi-agent settings, shared system-level rewards are common when agents operate in physically coupled environments such as grids or districts, where independent objectives may lead to poor coordination [163,167,179,204,221]. Local rewards appear less frequently and are mainly used in market-oriented or prosumer-centered settings, where individual profit, satisfaction, or marginal contribution is more important [192,202,209]. Multi-objective rewards dominate the field, while hierarchical and risk-aware formulations remain less common but increasingly relevant [165,197,203].
In terms of reward content, economic terms are the most frequent; cost, profit, electricity purchase, fuel use, or energy trading is usually the core of the reward definition [162,193,194,200,206,222,223,234] (see Figure 10-Left). Technical terms, especially voltage-related penalties, form the second main group, and are particularly common in power grid studies [163,164,165,167,168,177,182,184]. Safety and feasibility terms are also widely used to penalize constraint violations, instability, or unsafe operating margins [162,164,171,181,204,227]. Comfort, load satisfaction, and user inconvenience form a separate cluster, mainly in building and demand response applications [222,223,224,225,228,229,238]. Other recurring terms include frequency regulation, power loss, state of charge, equipment aging, and degradation [166,173,175,176,180,197,199,203]. Emissions, reliability, and forecast- or risk-sensitive rewards are less frequent but are becoming more visible [196,208,227].
In power grid applications, rewards mainly approximate secure network operation. Reward structures commonly combine voltage deviation, power loss, curtailment, switching effort, and control effort into one scalar objective [164,167,168,178,182]. Since grid control is constraint-sensitive, some studies also use quadratic, exponential, or barrier-shaped penalties to increase sensitivity near unsafe operating regions [166,170,175,183]. Shared rewards also seem to be frequent in MARL settings because voltage and frequency are system-level quantities affected by multiple agents [163,167,168,179,183,184]. More advanced studies separate reward maximization from safety through CMDP or CPO formulations [162,164,171].
In microgrid applications rewards appear to be more techno-economic, since the RL controller must coordinate RES, storage, dispatchable units, market exchange, and local demand. Many studies define rewards around total operating cost, including grid purchase, fuel use, DER degradation, and curtailment [193,194,200,206]. More recent works also include resilience, emissions, reliability, power quality, and user-oriented objectives [208,209,210,211]. Agent-specific rewards appear more commonly in decentralized microgrids, peer-to-peer trading, and community coordination [192,202,205]. Risk-aware or correction-based rewards remain less frequent, but appear in studies addressing mismatch penalties, shortfall risk, or second-stage compensation [197,203,204].
In building energy frameworks, rewards usually combine electricity cost with comfort, dissatisfaction, and storage use penalties [222,223,229,231,233]. Peak shaving, demand ramp reduction, PV self-consumption, under-generation penalties, and district-level smoothing are also increasingly used [221,224,225,226,232,235,238]. Smoother reward functions, including quadratic or Gaussian-shaped terms, are also becoming more common to avoid abrupt penalties and improve learning stability [226,227,238].
Soft and Hard Constraint Handling: It should be also noted that the reviewed studies handle operational constraints in different ways. Most use soft constraints that include voltage limits, frequency deviations, SoC bounds, comfort violations, power losses, curtailment, or user dissatisfaction as penalty terms in the reward function [163,164,165,166,167,170,183,199,200,223,227]. This is a simple approach that can be used with different RL algorithms, but does not strictly prevent infeasible actions during training or unseen operating conditions. A smaller group of studies uses hard or admissibility-oriented constraints in which feasibility is enforced through constrained RL, safety layers, action masking, physics-informed neural architectures, or optimization-based correction mechanisms [164,165,171,173,177,180,190,194,203,204,207,229]. Such approaches require more computation and modeling effort but offer stronger operational safety, especially in applications such as voltage control, frequency regulation, OPF, converter control, and storage scheduling. Table 13 provides a comparative table of soft vs. hard constraint-handling strategies in RL-based smart-grid applications, with the reviewed papers explicitly categorized according to whether they handle constraints through soft reward penalties or hard/admissibility-oriented mechanisms, including constrained RL, safety layers, action masking, physics-informed neural architectures, and optimization-guided correction.
As shown in Table 13, reward–penalty approaches are the most common because they are easy to implement and may be combined with almost any RL algorithm; however, their effect on safety is still indirect, as the agent might still explore infeasible actions. By contrast, hard or admissibility-oriented mechanisms are used to restrict, correct, or guide the action before it reaches the physical system. Although such methods can add computational cost, they appear to be better suited for deployment-oriented RL in safety-critical smart grid applications.

6.5. Baseline Control

The reviewed studies use a broad range of baseline controllers to compare RL performance. Overall, RL methods are often compared against other RL variants, showing that the field is increasingly focused on methodological refinement within RL itself (Figure 11-Left and Right). Optimization-based methods also remain common, since they provide strong references for operational efficiency and constraint handling. Rule-based control is another frequent baseline, especially because it reflects current practice in many real systems. Classical controllers such as MILP, PID, and MPC appear less often but remain useful depending on the application.
More specifically, as concerns power grid RL applications, the most common baselines are physics-based optimization methods, including OPF, ACOPF, SOCP/MISOCP, and MPC [162,170,171,179,184] (see Figure 12-Left). These approaches explicitly consider network constraints and are widely used in voltage regulation, volt/VAR control and economic dispatch. Classical local controllers such as droop control, greedy VVC, RBC, PID/FOPI, and linear control also appear frequently [163,165,167,174,175,183] in power grid contexts. Such approaches adequately reflect practical grid operation; however, they are usually limited when coordinating distributed DERs. RL-against-RL benchmarking in power grid applications appears frequently, with newer methods being compared against DDPG, MADDPG, SAC, DDQN, and related DRL variants in a mature comparison practice [168,169,173,176,183]. According to this analysis, power grid RL has moved from demonstrating feasibility toward comparing advanced RL architectures for safety, coordination, observability, and scalability.
In microgrid studies, baseline selection seems more heterogeneous due to the fact that microgrid operation includes scheduling, storage control, market participation, and uncertainty management. Many microgrid works compare RL with classical heuristics and fixed EMS strategies such as RBC, fixed dispatch, droop control, PID, and deterministic scheduling [191,198,199,215] (see Figure 12-Center). Optimization baselines such as MPC, DP, Lyapunov optimization, MILP, and deterministic planning are also widely used [194,200,203,206,208]; these provide strong references because of their explicit handling of constraints and scheduling objectives. RL-versus-RL comparisons are also frequent in microgrids, especially in multi-agent and risk-aware studies, where proposed RL methods are benchmarked against Q-learning, DQN, DDPG, PPO, and other RL variants [192,193,195,204,210]. In prediction-aided scheduling, forecasting baselines such as ANN, MLR, and ARIMA may also be used [196]. Overall, strong microgrid RL studies are expected to outperform simple heuristics while remaining competitive with optimization-based methods under uncertainty.
In building applications, RBC is the dominant baseline (Figure 12-Right), reflecting the continued use of fixed schedules, thermostat logic, priority rules, and heuristic supervisory control in real buildings [221,223,224,225,226,228,230,232,236,237]. This makes RBC a practically meaningful benchmark for building energy systems: if RL outperforms it, the method shows direct value for cost reduction, PV self-consumption, peak shaving, and comfort management. Optimization baselines such as MPC, MILP/MINLP, ADMM, and GA are also used [222,229,230,231,234]; these provide stronger methodological references than RBC due to the fact that they are able they optimize over forecasts and constraints. Finally, more mature RL-versus-RL benchmarking is increasingly common with MARL, federated RL, and hybrid RL compared to single-agent RL or earlier DRL methods [221,227,229,234,235]. This trend underlines a shift from simple heuristic and RBC-based comparison towards evaluating which RL architecture best handles coordination, privacy, mixed action spaces, and comfort–energy trade-offs.

6.6. Data Utilization

The reviewed studies show that RL-based smart grid control relies mainly on operational and context-aware time series data. Load demand and RES generation are the most common inputs, since they are central to supply–demand balancing under uncertainty (Figure 13-Right, Center and Left). Storage states, weather variables, electrical measurements, and market data are also frequently used, especially in scheduling, forecasting, flexibility management, and grid-level control. More specialized inputs such as frequency signals and EV-related states appear mainly in targeted applications.
In power grid applications, data are usually structured as cyber–physical time series. Typical inputs include voltages, active/reactive demand, RES generation, power injections, device states, and in some cases market signals [163,165,171,179] (see Figure 13 Left). Training of RL agents is mostly simulation-based, since unrestricted exploration is not feasible in real grids [162,170,182]. Is is also noticeable that many studies improve realism by embedding real PV, wind, load, and price data into physics-based environments [164,165,171,181,183]. Recent works also use local measurements, topology-aware features, surrogate model outputs, and device-coupled variables to support partial observability and physics-consistent learning [167,169,173,174,179,180,184].
In microgrids, data are more heterogeneous because operation depends on RES variability, demand, storage, and market interaction (Figure 13-Center). Most studies use PV/wind generation, load demand, storage states, and electricity prices [191,193,200,202]. Real or semi-real datasets, including weather, demand, and price traces, are often embedded into simulation environments [191,197,200,201,204], while synthetic stochastic scenarios are common in planning and market studies [192,198,209,210]. Compared with power grids, RL in microgrids utilize more device-level information, including batteries, hybrid storage, hydrogen components, EVs, and controllable loads [191,203,208,211]. Forecast-informed and history-augmented features also appear commonly in recurrent or prediction-aided frameworks [194,196,205].
In building applications, agents typically combine exogenous data such as weather, irradiance, prices, and PV generation with endogenous variables such as indoor temperature, storage states, appliance status, and load demand [221,223,226,230] (see Figure 13-Right). Many studies rely on simulation platforms such as EnergyPlus or CityLearn enriched with realistic operating conditions [221,225,235]. Other practitioners use real-world sensor, smart-meter, campus, or building datasets, suggesting movement towards digital twin-like frameworks [222,224,228,236,237]. Forecast-based inputs are also common because HVAC control, storage scheduling, and demand shifting are anticipatory tasks [222,225,233].

6.7. Control Objectives

Control objectives illustrate how smart grid RL has evolved from classical regulation and dispatch tasks toward flexibility, resilience, comfort, and market participation. This shift reflects the growing complexity of RES-rich energy systems, where controllers must handle technical constraints, economic goals, and user-centered requirements at the same time. More specifically, in power grid applications, the dominant objective concerns voltage regulation and volt–VAR control, especially in active distribution networks with high PV penetration (Figure 14-Left). RL is mainly used to keep voltages within safe limits while coordinating PV inverters, SVCs, OLTCs, capacitors, and ESS under uncertainty [163,164,165,167,168,169,173,177,178,179,180,181,182,183,184]. Frequency regulation and load-frequency control are also evident in research as control objectives, especially in multi-area systems [166,175,176]. Other works address economic dispatch and secure OPF, often with additional objectives such as loss minimization, curtailment reduction, switching effort, SOC preservation, and operational safety [162,169,171,172,173,174,178,180,181]. In microgrids, the main objective concerns energy scheduling with economic optimization (Figure 14-Center). Most studies coordinate RES generation, storage, dispatchable units, and grid exchange to reduce operating cost while maintaining supply–demand balance under uncertainty [191,193,194,196,200,201,203,206,208]. This highlights the role of RL in handling temporal coupling, RES intermittency, and forecast errors [193,194,200,203]. Market-oriented objectives such as peer-to-peer trading, pricing, community coordination, and profit allocation are common in the research [192,198,202,205,209,210]. Other studies focus on RES utilization, risk-aware operation, corrective control, voltage/frequency stability, multi-energy coordination, emissions, reliability, and planning-level decisions such as DER or EVCS sizing and placement [195,197,199,203,204,208,211,215].
In building RL energy applications, the main objective concerns the joint optimization of energy cost and occupant comfort (Figure 14-Right). RL is used to coordinate HVAC, appliances, storage, EVs, and sometimes DHW systems under changing prices, weather, and RES availability [222,223,224,229,230,231,232,238]. Grid-interactive objectives are also increasingly important, including peak reduction, demand response, load shaping, and RES self-consumption [221,225,226,233,235,236,237]. It is also evident that multi-building and community-level coordination is often addressed through MARL or federated RL, especially when privacy and communication constraints are present [221,231,234,235]. Recent studies include battery safety, hygiene constraints, carbon cost, and user dissatisfaction as well, showing a move toward more realistic and service-oriented objective design [227,228,234].

6.8. RES Types and Integrated Equipment

Across the reviewed smart grid applications, PV is clearly the dominant RES technology (Figure 15-Left, Center, and Right). This reflects its wide deployment, modularity, and easier integration across power grids, microgrids, and building energy systems. Wind appears as the second most common RES, mainly in larger-scale and hybrid systems. Biomass, hydropower, and solar water heating are much less frequent, showing that RL research has mainly focused on variable and operationally challenging RES.
In power-grid applications, RES integration is mostly PV-centered, with wind appearing less often (Figure 15-Left). The main focus is on the controllable equipment coordinated with RES. Most distribution-level studies consider inverter-interfaced devices such as PV inverters, smart inverters, SVCs, STATCOM-like devices, SOPs, LPCs, and HDTs, as they provide fast and continuous control for voltage support [163,165,167,168,174,177,178,183,184]. When slower or discrete devices such as OLTCs, capacitor banks, and voltage regulators are included, the problem shifts to one of coordinating multiple time scales [169,171,173,179]. BESS and EV charging further extend the control problem by adding active-power flexibility and temporal coupling [167,171,173,175,179,181]. Transmission and AGC-oriented studies are more generator-centered, with RL mainly addressing RES uncertainty through balancing and frequency regulation [162,166,170,176,182].
In microgrid studies, the RES mix is broader because renewable sources are treated as local supply resources rather than external disturbances. PV remains the dominant RES technology, and is often combined with wind turbines, while biomass technologies appear only occasionally [191,192,194,198,208,211] (Figure 15 Center). ESS/BESS is the most common complementary asset to RES, supporting smoothing, arbitrage, reserve provision, and reliability [191,193,196,200,206]. Dispatchable units, including diesel generators, microturbines, fuel cells, and synchronous DGs, remain important in islanded and reliability-oriented settings [192,193,198,204,215]. Recent studies also include hydrogen systems, CHP, thermal storage, gas coupling, electric–thermal loads, EVs, appliances, peer-to-peer trading, and community exchanges, showing a shift toward multi-energy and market-aware microgrid control [191,192,196,206,208,209,210,211].
In building applications, RL mainly focuses on behind-the-meter flexibility rather than direct grid stabilization. PV is again the dominant RES, but is usually coordinated with HVAC, storage, appliances, EVs, and community-level interactions [223,230,235,238] (Figure 15 Right). Building RL commonly uses on-site PV together with thermal and electrical flexibility to improve self-consumption, reduce cost, and preserve comfort [221,222,223,224,229,230,235,238]. HVAC and thermally coupled systems are central because buildings provide much of their flexibility through thermal inertia [221,222,224,225,226,228,232]. TSS technologies may also support load shifting, while batteries and EVs add electrical flexibility [221,223,224,225,229,230,230,233,234,235,237]. Appliances, HEMS/IBEMS structures, and district/community coordination further extend building RL toward user-centered and grid-interactive operation [221,222,226,227,228,229,231,233,234,235].

6.9. Grid Types

In power grid studies, the dominant setting concern the active distribution networks and secondary transmission system testbeds (Figure 16 Left). This is where high PV penetration, inverter-based DERs, BESS, OLTCs, capacitor banks, SVCs, and other flexible devices create nonlinear and highly coupled control problems suited to RL [163,164,165,167,168,169,171,173,174,177,178,179,180,181,182,183,184]. This explains the dominance of voltage regulation and volt/VAR control as control objectives, as stated in the previous subsection. Compared with distribution-level RL studies, fewer works investigate transmission grid applications such as AC optimal power flow, automatic generation control, and load-frequency control, where RL is mainly used to accelerate optimization or coordinate frequency support [162,170,175,176]. Most validations remain simulation-based, often using IEEE benchmark feeders with realistic PV, load, or price data [163,165,171,180,183,184]. Hardware-in-the-loop (HIL), cyber–physical validation, and utility-scale deployment remain limited [162,164,169,178,179,183,184,246].
In microgrid studies, the testbeds are more diverse, including islanded hybrid systems, grid-connected residential and community microgrids, hybrid-ESS systems, multi-energy microgrids and networked multi-microgrid settings [191,192,193,194,195,198,201,206,208,210]. Grid-connected microgrids are increasingly common, with RL used for economic dispatch, storage scheduling, trading, real-time correction, and market-aware coordination under uncertain prices and RES output [194,200,202,203,209] (Figure 16 Center). Islanded and inverter-dominated microgrids remain important for voltage/frequency regulation, hybrid energy storage system coordination, and resilience under low-inertia conditions [191,193,195,199]. Newer works cover community, multi-energy, and networked microgrids involving peer-to-peer trading, electric–gas coordination, and DSO–microgrid interaction [198,205,208,210]. Most studies are still simulation-based, although real irradiance, load, price, and mobility datasets are often used to improve realism [191,193,197,200,211].
In building studies, the dominant setting concerns not a network topology but an occupant- and asset- centered energy environment. Applications range from smart homes and commercial buildings to multi-building districts, campuses, communities, and zero-energy or DC building systems (Figure 16 Right) [222,223,224,226,228,229,230,232,234,238,238]. Residential HEMS illustrate the most common subtype, where RL coordinates PV, batteries, HVAC, DHW, appliances, and EV charging under dynamic tariffs and comfort constraints [222,224,228,229,230,231,236]. District and community settings are also growing, especially in CityLearn-type studies, where RL supports peak shaving, ramp mitigation, shared storage coordination, and net-load shaping [221,225,226,233,234,235]. Studies involving real commercial buildings remain rare, with FLEXLAB application standing out as one of the few experimental examples [223,232]. Although most work is still simulation-based, building studies often include realistic weather, occupancy, thermal dynamics, tariff, and PV data [223,231,234,238].

6.10. Performance Comparison

This subsection summarizes the most notable performance results reported in the reviewed RL-based smart-grid studies. The selection is based not only on the largest numerical improvement but also on methodological relevance, meaningful baselines, realistic constraints, real-time applicability, and validation under uncertainty, safety, comfort, or grid support requirements.
In power grid applications, the strongest results show that RL is adequate to support fast and constraint-aware operation in RES-rich networks: Zhou et al. [162] demonstrated that imitation-initialized PPO can solve AC-OPF with only about 0.6% cost deviation from interior-point solvers while being around seven times faster and maintaining near-perfect success rates. Wu et al. [179] reported that a MADDPG-based multi-timescale framework keeps voltages within 0.99–1.03 p.u., removes previous violations up to 1.062 p.u., and reduces computation time from 812 s to 0.177 s. Li et al. [171] showed that CPO-based safe RL achieves major cost reductions compared with DDPG, PPO, and SAC while keeping constraint violations near zero and remaining only about 11% away from MISOCP. Cao et al. [163] further showed that attention-enhanced MADDPG reduces voltage deviation from 0.46% to 0.12%, lowers the maximum voltage rise from 3.48% to 1.01%, and reaches 95.1% of centralized optimal performance with millisecond-level execution. Kou et al. [164] also demonstrated safe continuous-control RL in ADNs, reducing losses by around 45% while keeping voltages within 0.95–1.05 p.u. Other notable results include Luo et al. [189], with 80.7% fewer voltage violation occurrences; Huang et al. [183], with about 30% lower economic cost than droop control; and Mamodiya et al. [187], with 41.4% higher annual PV energy yield.
In microgrid applications, the strongest results indicate that RL is useful for dispatch, planning, market coordination, and stability-critical control. Abid et al. [211] reported an 85.1% improvement over MOAVOA-COMA and EV SOC increases of up to 154%, showing the value of MARL for multi-objective planning of RES, BESS, and EV charging infrastructure. Barbalho et al. [199] achieved up to 86% reduction in cumulative voltage/frequency error and a 4.7× improvement over droop control, which is especially relevant for islanded microgrids. Xiong et al. [205] showed that trust-region RL increases profit from $4,190 with DQN to $29,511, while interval prediction almost doubles profit compared with point forecasting. Li et al. [218] demonstrated that hierarchical PPO provides a 12.4% higher reward and roughly 98% faster decisions than optimization-based scheduling. Xiong et al. [219] also showed that PPO improves operator profit by 33.74% and self-balance by 22.97% in a PV/wind/ESS/EV/V2G microgrid. Additional strong results include Li et al. [196], with up to 38.4% operating-cost reduction and spinning-reserve reduction in 66.7% of periods; Zhou et al. [206], with 61.17% lower computation time at only 3.13% higher cost than MILP; and Wang et al. [220], with 22.6% lower peak transformer loading in residential microgrid clusters.
In building applications, the best results show that RL has moved beyond simple load shifting towards comfort-aware, grid-interactive, and multi-device energy management. Heidari et al. [228] reported 7–60% savings compared with energy-efficient RBC and 28–75% compared with conventional RBC while respecting thermal comfort and Legionella-related DHW hygiene constraints. Touzani et al. [223] provided one of the strongest practical demonstrations, deploying DDPG in a real commercial test facility and achieving up to 39.6% cost savings while coordinating HVAC setpoints, battery dispatch, and PV generation. Aldahmashi et al. [240] showed that PPO could reduce electricity bills by 31.5% and improve average power factor from 0.44 to 0.901. Deng et al. [237] reported 94.6% PV self-sufficiency and 74.1% PV absorption in stochastic residential DC microgrids, outperforming RBC and MPC. Talihati et al. [245] showed that MARL could increase PV self-consumption by 66.41%, cover 38.68% of EV demand, lower community costs, and generate ES-PV operator profit. Other notable results include Gao et al. [227], with a 70–80% improvement in off-grid deviation and near-zero battery safety violations; Shen et al. [232], with an 84% reduction in discomfort duration and 43% reduction in renewable curtailment; and Wang et al. [238], with a 34.86–46.10% improvement in self-sufficiency.

6.11. Simulation Tools

For power grid applications, the reviewed RL literature shows that most studies rely on custom simulation environments built around established power system solvers and IEEE or utility benchmark networks rather than on fully standardized RL benchmark suites. Several works create OpenAI Gym-style wrappers around power flow or OPF engines. For example, Zhou et al. [162] used PYPOWER for AC power flow and OPF studies on the IEEE 14-bus and Illinois 200-bus systems, while Zhang et al. [170] combined PYPOWER, OpenAI Gym, and TensorFlow for simulation-based RL training. At the distribution level, many studies use IEEE 33-, 34-, 123-, 141-, and 342-node feeders. These are implemented through different toolchains, including MATLAB/Reinforcement Learning Toolbox, MATLAB/Python co-simulation, Python/TensorFlow, Keras, PyTorch, YALMIP/Gurobi, Pandapower/PandaPower, and OpenDSS [163,164,167,168,169,171,173,178,180,183,184,185,189]. Other studies focus on automatic generation control, load-frequency control, or virtual power plant-oriented simulators, including numerical multi-generator load-frequency control models, MATLAB/Simulink two-area systems, China Southern Grid four-area simulations, power systems computer-aided design/electromagnetic transients (including DC-based virtual power plant simulations), and general IoT-enabled smart grid environments [166,175,176,186,188]. Overall, while open-source solvers such as PYPOWER, Pandapower, and OpenDSS are becoming more common, the reviewed power grid studies rarely use dedicated RL benchmark environments such as Gym-ANM, Grid2Op, PowerGym, or OpenDSS-Gym. Instead, they usually take standard electrical test systems as benchmarks, then build problem-specific RL environments around them.
For microgrid RL applications, most studies use custom testbeds developed in MATLAB/Simulink, Python/TensorFlow, Python/PyTorch, or hybrid MATLAB–Python environments. These testbeds are often driven by real or public datasets for PV generation, wind production, load demand, electricity prices, and weather conditions [191,192,193,194,196,200,201,203,205,208,210,219,220]. A smaller number of works provide more reusable or practitioner-oriented platforms. For example, Chen et al. [195] introduced the open-source PGSim simulator for PowerNet using IEEE-34-derived microgrid cases; Barbalho et al. [199] released MATLAB/Simulink files through Mendeley Data; and Barros et al. [216] combined GridLAB-D for residential microgrid energy simulation with OMNeT++ for communication network modeling. However, explicit sim-to-real methodologies are still rare, and most studies do not report transfer learning, domain randomization, HIL testing, or field deployment. The main exceptions are Rajamallaiah et al. [213], who validated a modified TD3 controller on OPAL-RT OP4512 HIL, and Sepehrzad et al. [214], who compiled a MATLAB networked microgrid simulator into RT-LAB/HYPERSIM for OPAL-RT HIL validation. In both cases, the transition towards practical validation was based on real-time HIL testing of the trained controller rather than an explicit transfer learning or domain randomization pipeline.
For building applications, the reviewed RL literature indicate maturity in terms of reusable benchmarking compared to the literature on power grids and microgrids. This is mainly because several studies use CityLearn, an open-source OpenAI Gym-compatible benchmark for multi-building demand response and grid-interactive building control [221,225,226,235]. Beyond CityLearn, some works develop custom OpenAI Gym-compatible environments linked to building simulators or real datasets. Examples include the EnergyPlus/SCooDER/FMI/FMUs/PyTorch FLEXLAB environment used by Touzani et al. [223], custom Gym–Stable-Baselines3 office building environments [227], OpenAI Gym environments with RC thermal models [238], and Gym/Pyomo/TensorFlow-based building or community simulations [198,242]. Many building control studies also rely on custom environments considering MATLAB/Simulink, Python/PyTorch, Python/TensorFlow, TRNSYS, IESVE-calibrated grey-box, or physics-based BES/HEMS simulators, usually using real PV, weather, load, price, occupancy, campus, or building datasets [222,224,228,229,231,232,234,237,240,241,243,244]. Regarding sim-to-real transition, there are significant limitations in transfer learning, domain randomization, HIL validation, and field deployment. Touzani et al. [223] provided one of the strongest simulator-to-field examples by training offline in a calibrated FLEXLAB simulator and then deploying the controller in the real testbed, while Heidari et al. [228] used stochastic offline pretraining across different Swiss climates, system sizes, and hot water behaviors as a transferability strategy. Wang et al. [238] used imitation learning from a pretrained HVAC policy to accelerate MAPPO training. Most other studies are fully offline and simulation- or data-driven, without any formal domain randomization, HIL validation, or real-field deployment.
Table 14 summarizes the simulator practices used in RL studies for power grids, microgrids, and buildings. It is evident that most studies rely on custom environments based on MATLAB, Simulink, Python, or deep learning, while the use of standardized benchmarks is limited. CityLearn is the clearest recurring benchmark in buildings, whereas power grid and microgrid applications mostly depend on IEEE test systems, custom environments, and selected open-source solvers such as PYPOWER, PandaPower, OpenDSS, GridLAB-D, and PGSim.

7. Discussion

The Section 7 synthesizes the main trends identified in the literature from the perspective of increasing RES penetration in smart grid applications. Based on these trends, it further outlines practical future directions for researchers and practitioners in the areas of RES-integrated power grids, microgrids, and building energy systems.

7.1. Current Trends

Viewed through the lens of RES integration, the reviewed RL literature reveals a coherent technological shift rather than a loose collection of algorithmic preferences. The common driver is the same across the smart grid continuum: as renewable penetration rises, operation becomes more uncertain, temporally coupled, and geographically distributed. What changes from one application layer to another is primarily the form that this pressure takes: at the network level, the main concern is secure electrical coordination under fast disturbances; in localized energy systems, the emphasis moves toward balancing, autonomy, and resilience; closer to the end user, the problem becomes one of orchestrating flexibility without undermining comfort or service quality. This distinction is important because it explains why the literature follows a shared direction while still producing different control architectures and problem formulations.
  • A first major trend is that RL becomes relevant precisely when RES integration transforms control from a static optimization task into a sequential decision problem under uncertainty. With high shares of solar and wind, system states are no longer smooth or easily predictable: voltage can fluctuate rapidly, storage adequacy becomes path-dependent, and local demand must be matched against variable generation over time [163,165,167,183,193,200,203,221,225,230]. In this setting, the value of RL is not just in its being data-driven but in being used to learn policies where the quality depends on how present decisions reshape future operating conditions. Therefore, the rise of RL in RES-rich systems should be understood as a consequence of temporal coupling, not merely one of algorithmic fashion.
  • A clear dominance of actor–critic methods is observed, which follows from the physical nature of RES flexibility itself. In many modern applications, control variables are continuous: inverter setpoints, storage dispatch, generator output adjustments, HVAC modulation, and EV charging rates all vary over a continuum rather than through a few discrete choices [168,169,175,196,199,211,223,230,237]. Under such conditions, actor–critic methods provide a practical balance between expressive policy learning and tractable optimization. At the same time, lighter value-based approaches remain useful where the action space is discrete, low-dimensional, or strongly structured, such as switching, scheduling, or simplified EMS tasks [162,164,191,192,206,222,224,232]. According to our evaluation, the research is not converging towards one universally superior RL family but towards a task-dependent hierarchy in which algorithm choice is shaped by the granularity and physics of the control problem.
  • Another trend concerns the move away from unconstrained black-box policies toward structured and physically meaningful action spaces. As renewable-rich systems become more complex, raw action outputs are increasingly difficult to train safely, interpret reliably, or deploy with confidence. This explains the growing use of parameterized or supervisory RL formulations in which the agent learns droop gains, target SoC values, controller parameters, or higher-level scheduling decisions rather than direct low-level commands [174,175,180,194,205,210,230,237]. The current shift reflects a deeper methodological maturity in which learning is increasingly aligned with engineering abstractions that operators already understand and trust.
  • Closely connected to the previous trend, an extensive deployment of hybrid RL approaches is evident. In practice, RES-integrated systems combine hard constraints, mixed time scales, heterogeneous devices, and multiple objectives, making pure RL increasingly insufficient as realism grows. The literature shows a steady movement toward architectures in which learning is combined with optimization layers, forecasting modules, surrogate models, or classical control laws [165,176,184,194,196,211,225,229,236]. In this context, hybridization should be viewed not as a compromise but as a recognition that intelligent control in RES-integrated environments requires both adaptation and structure. In this sense, the field is moving from pure model-free learning to a more disciplined approach in which RL is embedded within broader decision architectures.
  • Expansion of multi-agent RL is apparent, driven by the spatial and organizational distribution of RES flexibility. As controllable assets multiply and become more decentralized, coordination can no longer be treated as a side issue. In electrically coupled settings, one agent’s action may affect the feasible operating region of others, while in market-oriented settings multiple decision-makers need to respond to shared uncertainty; likewise, in privacy-sensitive environments it is important for coordination to occur without full state sharing [163,168,179,184,195,197,208,231,234,235]. In this context, the growth of MARL in recent highly cited research on smart grids reflects a more specific trend than the popularity of distributed learning, instead reflecting the structural decentralization introduced by RES integration itself.
  • The reviewed studies show that RL design is increasingly shaped by the physical time scale of the RES-related phenomenon under control. Fast feeder disturbances, slower storage scheduling decisions, and even slower thermal dynamics do not admit the same control architectures or training formulations [171,179,183,195,199,204,221,225,230]. The recent literature spans minute-level corrective control, hourly scheduling, and multi-stage supervisory strategies. An important implication is that temporal structure is no longer a secondary modeling detail but is becoming one of the main design pillars of RL-based RES management.
  • A final trend concerns the increasing role of reward engineering as an encoded statement of integration quality. As RES integration becomes more multi-objective, the reward is no longer a simple numerical signal but a compact expression of what the system is actually expected to value. Security, renewable utilization, comfort, resilience, cost, degradation, curtailment avoidance, and flexibility provision are all being translated into weighted reward structures that shape the learned behavior [164,167,178,193,197,208,221,226,228]. In this sense, reward design increasingly acts as the bridge between engineering priorities and algorithmic behavior. The more RES-integrated and multi-functional a smart grid ecosystem becomes, the more central this bridge becomes to the success of RL itself.
Taken together, the above trends illustrate that the field is moving from isolated device-level control toward orchestration of system-level RES flexibility. In power grids, the focus is shifting from standalone inverter support to coordinated control of PV, ESS, OLTCs, capacitors, and EV-related assets [169,171,179]. In microgrids, the literature is expanding from PV, diesel, and battery settings into multi-energy systems involving hydrogen, CHP, thermal storage, and community coordination [200,206,211]. In buildings, RL is increasingly used to integrate PV with HVAC, TES, BESS, appliances, EVs, and district-level interactions [225,234,238]. Despite such methodological sophistication, one common limitation across domains is that deployment still lags behind. Most studies remain simulation-based, even when driven by realistic exogenous data [183,200,223]. While the field is advancing rapidly in algorithmic sophistication, the main unresolved bottleneck remains robust validation and trustworthy transfer to real-world operation.

7.2. Future Directions

Future RL research for RES-integrated smart grids, including power grids, microgrids, and buildings, should move beyond simulation-based performance gains and focus more explicitly on deployment-ready intelligence. The central question is no longer whether RL can control RES-rich grid frameworks but whether it is able to do so in a safe, robust, and transparent manner under the uncertainty, coupling, and heterogeneity introduced by distributed RES.
  • A first and foundational direction is the development of safety-aware RL. As RES penetration rises, the consequences of poor control also intensify: voltage violations, instability, unsafe switching, storage misuse, and comfort rules can no longer be treated as secondary issues. Future work should further develop constrained and shielded RL, CMDP/CPO formulations, Lyapunov-based critics, and explicit safety filters that enforce electrical, thermal, and operational limits during both training and deployment, a trend that is only sporadically detected in recent research [162,164,171,204,223]. Safety needs to become a built-in design principle rather than an add-on.
  • Another key direction concerns the broader adoption of physics-informed and model-assisted learning. Purely model-free RL remains sample-inefficient and often difficult to interpret in RES-rich grids. More promising paths include graph-aware encoders, differentiable power flow surrogates, thermal models, digital twins, and learned world models that explicitly capture RES variability and multi-energy coupling. Such approaches could improve generalization, reduce training cost, and align learned policies more closely with real system behavior. Moreover, RES integration simultaneously creates slow scheduling problems, medium-horizon coordination tasks, and fast corrective control requirements. In many cases, a single flat policy is unable to handle all of these. To this end, future architectures should separate long-horizon planning, mid-level coordination, and fast local actuation, especially in systems that combine markets, storage, flexible loads, and inverter-dominated control. Such an approach would better reflect the real temporal organization of smart grid operation.
  • Another meaningful future direction is to enable scalable coordination mechanisms. As RES assets become more distributed, RL research need to focus on how agents communicate, what information they share, and how coordination remains effective under privacy, latency, and infrastructure limitations. This makes practices such as graph-based MARL, sparse attention, federated RL, event-triggered coordination, and leader/follower formulations especially relevant for future architectures. The central question is no longer centralized versus decentralized learning but how to efficiently achieve coordination under realistic communication conditions.
  • One of the field’s main bottlenecks is that most RES-oriented RL training and testing still takes place within closed simulation loops. Future progress will depend on better use of historical operational data, expert trajectories, and digital twins through offline RL, imitation learning, transfer learning, and meta-RL, followed by limited and safe online adaptation. This is particularly important in infrastructures where online trial-and-error is costly, unsafe, or simply unacceptable.
  • Many current studies still optimize expected performance while underrepresenting renewable ramps, forecast errors, rare contingencies, and uncertainties pertaining to user behavior. Future work should expand robust RL, distributional RL, adversarial training, CVaR-aware learning, and probabilistic scenario conditioning so that policies remain reliable under rare but realistic RES-induced disturbances [183,197,201,204]. In renewable-rich systems, average performance alone is not enough; reliability under stress is equally important.
  • Considering the adoption of broader objective formulations, future RL should move beyond single-objective cost minimization and learn policies that jointly optimize RES utilization, emissions, flexibility provision, reliability, battery degradation, comfort, and network support value. Multi-objective RL, preference-conditioned policies, and Pareto-based learning are especially relevant here, since RES integration is fundamentally a trade-off problem rather than a task of optimizing one metric. Following this direction would better align RL research with decarbonization goals and the operational reality of grid-interactive energy systems.
  • Another important future direction concerns the integration of newer RL methods into practical smart grid control frameworks. Distributional RL appears especially promising for RES-rich systems, since it can capture the variability of future returns rather than focusing only on their average value. This makes it particularly relevant for risk-aware applications such as storage scheduling, market bidding, EV charging, voltage regulation, and resilience management under rare but critical disturbances. Evolutionary policy optimization and population-based training may also become valuable in smart grid problems where the reward landscape is non-smooth, highly constrained, or inherently multi-objective, for example DER planning, controller tuning, hybrid RL–optimization schemes, and coordinated flexibility management. At the same time, graph-based RL and graph neural networks offer strong potential for improving scalability in distribution networks and interconnected microgrids by explicitly capturing electrical topology, while advanced MARL can support coordination among DERs, microgrids, buildings, aggregators, and EV fleets. To become truly practical, these techniques need to be combined with safe, physics-informed, federated, offline, and transfer learning mechanisms so that advanced RL can better handle high-dimensional uncertainty, reduce training effort, preserve privacy, and move closer to reliable deployment in real smart grid environments.
  • Finally, validation practice itself must evolve. RL for RES-integrated smart grids needs stronger assessment pipelines based on co-simulation, digital twins, hardware-in-the-loop, controller-in-the-loop, and staged pilot deployment rather than simulation alone. At the same time, explainability, operator trust, fallback logic, and human-in-the-loop supervision should become formal research targets, since real adoption in the future will depend not only on numerical performance but also on transparency and operational acceptance.
Overall, the most impactful future work will be that which treats RL not simply as an optimization tool but as a safe, physically grounded, uncertainty-aware, and deployable coordination layer for RES integration. The next frontier consists of better algorithms together with better alignment between algorithms and infrastructure, resulting in learning architectures that match the physics, data availability, time scales, and implementation constraints of renewable-rich smart grids in practice.

8. Conclusions

The current review has examined the role of RL in RES-integrated smart grid systems, with a focus on three major application domains: power grids, microgrids, and building energy systems. Its main contribution lies in the systematic organization and comparative analysis of recent RL-based studies across several key dimensions, including algorithm family, agent architecture, forecast horizon and timestep, baseline methods, RES technologies, integrated energy assets, testbed characteristics, reward design, and control objectives. This structured analysis offers a unified view of how RL is currently being used to support renewable integration across the different layers of modern energy systems.
According to the reviewed literature, applications that involve continuous control of inverters, storage systems, HVAC units, EVs, and distributed energy resources are dominated by actor–critic methods, particularly DDPG, SAC, TD3, PPO, and their multi-agent extensions. Value-based methods such as Q-learning, DQN, D3QN, and Rainbow DQN remain important in discrete or hybrid scheduling tasks, including appliance control, capacitor switching, and storage dispatch. From an application perspective, power grid studies mainly concentrate on voltage regulation, volt/VAR control, frequency regulation, and secure operation in active distribution and transmission networks. Microgrid applications are more diverse, often addressing energy scheduling, cost minimization, storage coordination, market participation, resilience, and islanded operation. In buildings, the focus is primarily on reducing energy cost, preserving comfort, PV self-consumption, demand response, peak shaving, and multi-building coordination.
Despite this progress, several important gaps remain. From a computational perspective, many RL methods still face challenges related to high training cost, sample inefficiency, sensitivity to hyperparameters, and limited scalability in large RES-rich systems. These limitations point to the need for more efficient learning strategies, including offline RL, transfer learning, meta-RL, imitation learning, and hierarchical RL. From an economic perspective, many studies still rely on relatively narrow cost minimization objectives; future work should better capture market participation, flexibility services, degradation costs, carbon costs, and the trade-offs among economy, reliability, comfort, and renewable utilization. From a standardization perspective, the field still lacks common benchmarks, reproducible datasets, unified evaluation metrics, and agreed-upon validation protocols, making direct comparison between RL approaches difficult. Finally, from a deployment perspective, most studies remain simulation-based, while real-world demonstrations, hardware-in-the-loop validation, controller-in-the-loop testing, and pilot-scale implementation are still limited.
Looking ahead, future research should move toward RL methods that are safer, more scalable, more interpretable, and more compatible with real smart grid operation. Safe and constrained RL could improve feasibility and operational reliability by explicitly accounting for voltage, frequency, storage, comfort, and device constraints. Physics-informed and model-assisted RL can reduce sample requirements and improve trust by embedding power flow equations, thermal dynamics, or digital twin models directly into the learning process. Graph-based RL and graph neural networks appear especially promising for distribution networks and multi-node energy systems, where topology and spatial coupling strongly influence control performance. Multi-agent RL can also support distributed coordination among DERs, buildings, microgrids, aggregators, and grid operators, while federated and privacy-preserving RL could enable learning across decentralized assets without requiring full data sharing. Hybrid RL–optimization frameworks represent a particularly practical direction, as they combine the adaptability of RL with the feasibility and transparency of model-based optimization.
Overall, RL is emerging as a powerful and flexible control paradigm for RES-integrated smart grids. Its future impact will depend on how successfully the field can bridge the gap between algorithmic performance and practical deployment. In this sense, the next generation of RL-based smart grid controllers will need to combine learning capability with physical awareness, safety guarantees, economic relevance, standardized evaluation, and credible real-world validation.

Author Contributions

Conceptualization, P.M.; methodology, P.M. and I.M.; software, P.M. and F.M.; validation, all authors; formal analysis, all authors; investigation, P.M., F.M. and H.H.C.; resources, all authors; writing—original draft preparation, P.M.; writing—review and editing, P.M., I.M. and H.H.C.; visualization, P.M.; supervision, P.M. and E.K. All authors have read and agreed to the published version of the manuscript.

Funding

This work has been supported by the HYPER-AI project, funded by the European Commission under Grant Agreement 101135982 through the Horizon Europe research and innovation program (https://hyper-ai-project.eu/, accessed on 14 March 2026).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest. None of the reviewed application studies were conducted by members of the HYPER-AI project, nor by direct collaborators of the authors within that project. Therefore, no project-related inclusion bias was present in the study selection process. Although the authors are involved in Special Issues in Energies, this does not represent a direct relationship with the current manuscript, which is submitted to Infrastructures. The authors have no editorial, financial, or institutional conflict related to this submission.

Abbreviations

A2CAdvantage Actor–Critic
A3CAsynchronous Advantage Actor–Critic
ACAlternating Current
ACOPFAlternating Current Optimal Power Flow
ADMMAlternating Direction Method of Multipliers
ADNActive Distribution Network
AGCAutomatic Generation Control
ANNArtificial Neural Network
APSOAdaptive Particle Swarm Optimization
ARIMAAutoregressive Integrated Moving Average
BEMSBuilding Energy Management System
BESSBattery Energy Storage System
BIPVBuilding-Integrated Photovoltaics
BIOBiomass
CBCapacitor Bank
CHPCombined Heat and Power
CMDPConstrained Markov Decision Process
CNNConvolutional Neural Network
COPCoefficient of Performance
CPLConstant Power Load
CPOConstrained Policy Optimization
CSPConcentrated Solar Power
CTDECentralized Training with Decentralized Execution
CVaRConditional Value at Risk
D3PGDistributed Distributional Deep Deterministic Policy Gradient
D3QNDueling Double Deep Q-Network
DCDirect Current
DDPGDeep Deterministic Policy Gradient
DDQNDouble Deep Q-Network
DERDistributed Energy Resource
DGDistributed Generation
DHWDomestic Hot Water
DLDeep Learning
DNNDeep Neural Network
DPDynamic Programming
DPGDeterministic Policy Gradient
DQNDeep Q-Network
DRDemand Response
DRLDeep Reinforcement Learning
DSODistribution System Operator
EA-MAACEvolutionary Attention-based Multi-Agent Actor–Critic
ELElectrolyzer
EMSEnergy Management System
ESSEnergy Storage System
EVElectric Vehicle
EVCSElectric Vehicle Charging Station
FASFeasible Action Screening
FCFuel Cell
FHForecast Horizon
FLCFuzzy Logic Controller
FOPIFractional-Order Proportional–Integral
GAGenetic Algorithm
GATGraph Attention Network
GCEGrid Control Equipment
GHPGeothermal Heat Pump
GNNGraph Neural Network
GPGaussian Process
HCNGHydrogen-Compressed Natural Gas
HEMSHome Energy Management System
HESSHybrid Energy Storage System
HILHardware-in-the-Loop
HTHydrogen Tank
HVACHeating, Ventilation, and Air Conditioning
IAEIntegral Absolute Error
ILImitation Learning
IPSInterior Point Solver
ITSEIntegral of Time-multiplied Squared Error
LFCLoad Frequency Control
LSTMLong Short-Term Memory
MAPEMean Absolute Percentage Error
MCTSMonte Carlo Tree Search
MDPMarkov Decision Process
MESMulti-Energy System
MGMicrogrid
MILPMixed-Integer Linear Programming
MINLPMixed-Integer Nonlinear Programming
MLRMultiple Linear Regression
MPCModel Predictive Control
MPPTMaximum Power Point Tracking
N/ANot Available/Not Applicable
NMGNetworked Microgrid
OLTCOn-Load Tap Changer
OPFOptimal Power Flow
P2PPeer-to-Peer
PARPeak-to-Average Ratio
PEPower Electronics
PERPrioritized Experience Replay
PIDProportional–Integral–Derivative
PINNPhysics-Informed Neural Network
PPOProximal Policy Optimization
PSOParticle Swarm Optimization
PVPhotovoltaic
QLQ-Learning
RBCRule-Based Control
RESRenewable Energy Sources
RLReinforcement Learning
RMSERoot Mean Square Error
SACSoft Actor–Critic
SCStochastic Control
SDAEStacked Denoising Autoencoder
SGSynchronous Generator
SoCState of Charge
SOCPSecond-Order Cone Programming
SOEState of Energy
SOHState of Health
SPStochastic Programming
SVCStatic VAR Compensator
SWHSolar Water Heating
TD3Twin Delayed Deep Deterministic Policy Gradient
TOUTime-of-Use
TR-RLTrust-Region Reinforcement Learning
TSTimestep
TSSThermal Storage System
UCBUpper Confidence Bound
V2GVehicle-to-Grid
V2HVehicle-to-Home
VDNValue Decomposition Network
VPPVirtual Power Plant
VSGVirtual Synchronous Generator
VSIVoltage Source Inverter
VVOVolt/VAR Optimization
WTWind Turbine

References

  1. Nguyen, B.N.; Ogliari, E.; Pafumi, E.; Alberti, D.; Leva, S.; Duong, M.Q. Forecasting generating power of sun tracking PV plant using long-short term memory neural network model: A case study in Ninh Thuan-Vietnam. In Proceedings of the 2024 Tenth International Conference on Communications and Electronics (ICCE); IEEE: Piscataway, NJ, USA, 2024; pp. 333–338. [Google Scholar]
  2. Hassan, Q.; Viktor, P.; Al-Musawi, T.J.; Ali, B.M.; Algburi, S.; Alzoubi, H.M.; Al-Jiboory, A.K.; Sameen, A.Z.; Salman, H.M.; Jaszczur, M. The renewable energy role in the global energy Transformations. Renew. Energy Focus 2024, 48, 100545. [Google Scholar] [CrossRef] [Scilit]
  3. Seetharaman; Moorthy, K.; Patwa, N.; Saravanan; Gupta, Y. Breaking barriers in deployment of renewable energy. Heliyon 2019, 5, e01166. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Singh, S. Energy crisis and climate change: Global concerns and their solutions. In Energy: Crises, Challenges and Solutions; John Wiley & Sons Ltd.: Hoboken, NJ, USA, 2021; pp. 1–17. [Google Scholar]
  5. Nanaki, E.A.; Xydis, G.A. Deployment of renewable energy systems: Barriers, challenges, and opportunities. In Advances in Renewable Energies and Power Technologies; Elsevier: Amsterdam, The Netherlands, 2018; pp. 207–229. [Google Scholar]
  6. Stephanie, F.; Karl, L. Incorporating renewable energy systems for a new era of grid stability. Fusion Multidiscip. Res. Int. J. 2020, 1, 37–49. [Google Scholar] [CrossRef] [Scilit]
  7. Reddy, V.J.; Hariram, N.; Ghazali, M.F.; Kumarasamy, S. Pathway to sustainability: An overview of renewable energy integration in building systems. Sustainability 2024, 16, 638. [Google Scholar] [CrossRef] [Scilit]
  8. D’Agostino, D.; Mazzella, S.; Minelli, F.; Minichiello, F. Obtaining the NZEB target by using photovoltaic systems on the roof for multi-storey buildings. Energy Build. 2022, 267, 112147. [Google Scholar] [CrossRef] [Scilit]
  9. Fetting, C. The European green deal. ESDN Rep. Dec. 2020, 2, 53. [Google Scholar] [CrossRef] [Scilit]
  10. Jäger-Waldau, A.; Bodis, K.; Kougias, I.; Szabo, S. The New European Renewable Energy Directive-Opportunities and Challenges for Photovoltaics. In Proceedings of the 2019 IEEE 46th Photovoltaic Specialists Conference (PVSC); IEEE: Piscataway, NJ, USA, 2019; pp. 0592–0594. [Google Scholar]
  11. Shivakumar, A.; Dobbins, A.; Fahl, U.; Singh, A. Drivers of renewable energy deployment in the EU: An analysis of past trends and projections. Energy Strategy Rev. 2019, 26, 100402. [Google Scholar] [CrossRef] [Scilit]
  12. Sinsel, S.R.; Riemke, R.L.; Hoffmann, V.H. Challenges and solution technologies for the integration of variable renewable energy sources—a review. Renew. Energy 2020, 145, 2271–2285. [Google Scholar] [CrossRef] [Scilit]
  13. Prasad, M.B.; Ganesh, P.; Kumar, K.V.; Mohanarao, P.; Swathi, A.; Manoj, V. Renewable energy integration in modern power systems: Challenges and opportunities. E3S Web Conf. 2024, 591, 03002. [Google Scholar] [CrossRef] [Scilit]
  14. Erdiwansyah; Mahidin; Husin, H.; Nasaruddin; Zaki, M.; Muhibbuddin. A critical review of the integration of renewable energy sources with various technologies. Prot. Control Mod. Power Syst. 2021, 6, 3. [Google Scholar] [CrossRef] [Scilit]
  15. Bâra, A.; Oprea, S.v. Photovoltaic systems forecast using machine learning algorithms and recurrent neural networks. Sci. Bull. Nav. Acad. 2023, 26, 178–185. [Google Scholar]
  16. Kurucan, M.; Michailidis, P.; Michailidis, I.; Minelli, F. A Modular Hybrid SOC-Estimation Framework with a Supervisor for Battery Management Systems Supporting Renewable Energy Integration in Smart Buildings. Energies 2025, 18, 4537. [Google Scholar] [CrossRef] [Scilit]
  17. Kurucan, M.; Özbaltan, M.; Yetgin, Z.; Alkaya, A. Applications of artificial neural network based battery management systems: A literature review. Renew. Sustain. Energy Rev. 2024, 192, 114262. [Google Scholar] [CrossRef] [Scilit]
  18. Coban, H.H.; Lewicki, W. Flexibility in power systems of integrating variable renewable energy sources. J. Adv. Res. Nat. Appl. Sci. 2023, 9, 190–204. [Google Scholar] [CrossRef] [Scilit]
  19. Eltigani, D.; Masri, S. Challenges of integrating renewable energy sources to smart grids: A review. Renew. Sustain. Energy Rev. 2015, 52, 770–780. [Google Scholar] [CrossRef] [Scilit]
  20. Michailidis, P.; Minelli, F.; Michailidis, I.; Kurucan, M.; Coban, H.H.; Kosmatopoulos, E. Machine learning for energy management in buildings: A systematic review on real-world applications. Energies 2025, 19, 219. [Google Scholar] [CrossRef] [Scilit]
  21. Shafiullah, G.; Oo, A.M.; Jarvis, D.; Ali, A.S.; Wolfs, P. Potential challenges: Integrating renewable energy with the smart grid. In Proceedings of the 2010 20th Australasian Universities Power Engineering Conference; IEEE: Piscataway, NJ, USA, 2010; pp. 1–6. [Google Scholar]
  22. Eissa, M. Protection techniques with renewable resources and smart grids—A survey. Renew. Sustain. Energy Rev. 2015, 52, 1645–1667. [Google Scholar] [CrossRef] [Scilit]
  23. Shahzad, S.; Abbasi, M.A.; Ali, H.; Iqbal, M.; Munir, R.; Kilic, H. Possibilities, challenges, and future opportunities of microgrids: A review. Sustainability 2023, 15, 6366. [Google Scholar] [CrossRef] [Scilit]
  24. Iwuanyanwu, O.; Gil-Ozoudeh, I.; Okwandu, A.C.; Ike, C.S. The integration of renewable energy systems in green buildings: Challenges and opportunities. Int. J. Appl. Res. Soc. Sci. 2022, 4, 431–450. [Google Scholar] [CrossRef] [Scilit]
  25. Alfadil, M.O.; A. Kassem, M.; A. Al-Mansob, R. Renewable Energy Integration for Net-Zero Buildings: Challenges, Opportunities, and Strategic Pathways. Buildings 2026, 16, 879. [Google Scholar] [CrossRef] [Scilit]
  26. D’Agostino, D.; Minelli, F.; D’Urso, M.; Minichiello, F. Fixed and tracking PV systems for Net Zero Energy Buildings: Comparison between yearly and monthly energy balance. Renew. Energy 2022, 195, 809–824. [Google Scholar] [CrossRef] [Scilit]
  27. Kurucan, M.; Gencal, M.C.; Michailidis, P.; Minelli, F. Applying Symbolic Discrete Controller Synthesis Technique for Energy Management and Thermal Comfort Optimization in HVAC Systems. Sustainability 2026, 18, 2615. [Google Scholar] [CrossRef] [Scilit]
  28. Michailidis, P.; Michailidis, I.; Minelli, F.; Coban, H.H.; Kosmatopoulos, E. Model Predictive Control for Smart Buildings: Applications and Innovations in Energy Management. Buildings 2025, 15, 3298. [Google Scholar] [CrossRef] [Scilit]
  29. Michailidis, P.; Pelitaris, P.; Korkas, C.; Michailidis, I.; Baldi, S.; Kosmatopoulos, E. Enabling optimal energy management with minimal IoT requirements: A legacy A/C case study. Energies 2021, 14, 7910. [Google Scholar] [CrossRef] [Scilit]
  30. Coban, H.H.; Michailidis, P.; Yildirim, Y.A.; Minelli, F. Flattening Winter Peaks with Dynamic Energy Storage: A Neighborhood Case Study in the Cold Climate of Ardahan, Turkey. Sustainability 2026, 18, 761. [Google Scholar] [CrossRef] [Scilit]
  31. Ahsan, F.; Dana, N.H.; Sarker, S.K.; Li, L.; Muyeen, S.; Ali, M.F.; Tasneem, Z.; Hasan, M.M.; Abhi, S.H.; Islam, M.R.; et al. Data-driven next-generation smart grid towards sustainable energy evolution: Techniques and technology review. Prot. Control Mod. Power Syst. 2023, 8, 43. [Google Scholar] [CrossRef] [Scilit]
  32. Amasyali, K.; El-Gohary, N.M. A review of data-driven building energy consumption prediction studies. Renew. Sustain. Energy Rev. 2018, 81, 1192–1205. [Google Scholar] [CrossRef] [Scilit]
  33. Michailidis, P.; Michailidis, I.; Gkelios, S.; Kosmatopoulos, E. Artificial neural network applications for energy management in buildings: Current trends and future directions. Energies 2024, 17, 570. [Google Scholar] [CrossRef] [Scilit]
  34. Bâra, A.; Oprea, S.V. Machine learning algorithms for power system sign classification and a multivariate stacked LSTM model for predicting the electricity imbalance volume. Int. J. Comput. Intell. Syst. 2024, 17, 80. [Google Scholar] [CrossRef] [Scilit]
  35. Bâra, A.; Oprea, S.V. Large Language Models and Transformer Architecture in Electricity Price Forecasting. Ovidius Univ. Ann. Econ. Sci. Ser. 2025, 25, 126–135. [Google Scholar] [CrossRef] [Scilit]
  36. Zhang, Y.; Gatsis, N.; Giannakis, G.B. Robust Energy Management for Microgrids With High-Penetration Renewables. IEEE Trans. Sustain. Energy 2013, 4, 944–953. [Google Scholar] [CrossRef] [Scilit]
  37. Yang, J.; Su, C. Robust optimization of microgrid based on renewable distributed power generation and load demand uncertainty. Energy 2021, 223, 120043. [Google Scholar] [CrossRef] [Scilit]
  38. Norouzi, M.; Aghaei, J.; Pirouzi, S.; Niknam, T.; Fotuhi-Firuzabad, M.; Shafie-khah, M. Hybrid stochastic/robust flexible and reliable scheduling of secure networked microgrids with electric springs and electric vehicles. Appl. Energy 2021, 300, 117395. [Google Scholar] [CrossRef] [Scilit]
  39. Mesbah, A. Stochastic Model Predictive Control: An Overview and Perspectives for Future Research. IEEE Control Syst. Mag. 2016, 36, 30–44. [Google Scholar] [CrossRef] [Scilit]
  40. Zhang, H.; Seal, S.; Wu, D.; Boulet, B.; Bouffard, F.; Joos, G. Data-Driven Model Predictive and Reinforcement Learning Based Control for Building Energy Management: A Survey. arXiv 2021, arXiv:2106.14450. [Google Scholar]
  41. Parisio, A.; Rikos, E.; Glielmo, L. A Model Predictive Control Approach to Microgrid Operation Optimization. IEEE Trans. Control Syst. Technol. 2014, 22, 1813–1827. [Google Scholar] [CrossRef] [Scilit]
  42. Morstyn, T.; Hredzak, B.; Aguilera, R.P.; Agelidis, V.G. Model Predictive Control for Distributed Microgrid Battery Energy Storage Systems. IEEE Trans. Control Syst. Technol. 2018, 26, 1107–1114. [Google Scholar] [CrossRef] [Scilit]
  43. Michailidis, P.; Michailidis, I.; Kosmatopoulos, E. Reinforcement learning for electric vehicle charging management: Theory and applications. Energies 2025, 18, 5225. [Google Scholar] [CrossRef] [Scilit]
  44. Gautam, M. Deep Reinforcement learning for resilient power and energy systems: Progress, prospects, and future avenues. Electricity 2023, 4, 336–380. [Google Scholar] [CrossRef] [Scilit]
  45. Michailidis, P.; Michailidis, I.; Kosmatopoulos, E. Reinforcement learning for optimizing renewable energy utilization in buildings: A review on applications and innovations. Energies 2025, 18, 1724. [Google Scholar] [CrossRef] [Scilit]
  46. François-Lavet, V.; Taralla, D.; Ernst, D.; Fonteneau, R. Deep Reinforcement Learning Solutions for Energy Microgrids Management. In Proceedings of the 12th European Workshop on Reinforcement Learning, Lille, France, 10–11 July 2015. [Google Scholar]
  47. Mocanu, E.; Mocanu, D.C.; Nguyen, P.H.; Liotta, A.; Webber, M.E.; Gibescu, M.; Slootweg, J.G. On-Line Building Energy Optimization Using Deep Reinforcement Learning. IEEE Trans. Smart Grid 2019, 10, 3698–3708. [Google Scholar] [CrossRef] [Scilit]
  48. Lazaridis, C.R.; Michailidis, I.; Karatzinis, G.; Michailidis, P.; Kosmatopoulos, E. Evaluating reinforcement learning algorithms in residential energy saving and comfort management. Energies 2024, 17, 581. [Google Scholar] [CrossRef] [Scilit]
  49. García, J.; Fernández, F. A Comprehensive Survey on Safe Reinforcement Learning. J. Mach. Learn. Res. 2015, 16, 1437–1480. [Google Scholar]
  50. Achiam, J.; Held, D.; Tamar, A.; Abbeel, P. Constrained Policy Optimization. In Proceedings of the 34th International Conference on Machine Learning; PMLR: New York, NY, USA, 2017; Volume 70, pp. 22–31. [Google Scholar]
  51. Banerjee, C.; Nguyen, K.; Fookes, C.; Raissi, M. A Survey on Physics Informed Reinforcement Learning: Review and Open Problems. Expert Syst. Appl. 2025, 287, 128166. [Google Scholar] [CrossRef] [Scilit]
  52. Yu, P.; Zhang, H.; Song, Y.; Wang, Z.; Dong, H.; Ji, L. Safe Reinforcement Learning for Power System Control: A Review. Renew. Sustain. Energy Rev. 2025, 223, 116022. [Google Scholar] [CrossRef] [Scilit]
  53. Villalva, M.; De Siqueira, T.; Ruppert, E. Voltage regulation of photovoltaic arrays: Small-signal analysis and control design. IET Power Electron. 2010, 3, 869–880. [Google Scholar] [CrossRef] [Scilit]
  54. Cui, W.; Jiang, Y.; Zhang, B. Reinforcement learning for optimal primary frequency control: A Lyapunov approach. IEEE Trans. Power Syst. 2022, 38, 1676–1688. [Google Scholar] [CrossRef] [Scilit]
  55. Adibi, M.; Van Der Woude, J. Secondary frequency control of microgrids: An online reinforcement learning approach. IEEE Trans. Autom. Control 2022, 67, 4824–4831. [Google Scholar] [CrossRef] [Scilit]
  56. Fu, Y.; Guo, X.; Mi, Y.; Yuan, M.; Ge, X.; Su, X.; Li, Z. The distributed economic dispatch of smart grid based on deep reinforcement learning. IET Gener. Transm. Distrib. 2021, 15, 2645–2658. [Google Scholar] [CrossRef] [Scilit]
  57. Rizki, A.; Touil, A.; Echchatbi, A.; Oucheikh, R. A reinforcement learning approach based on group relative policy optimization for economic dispatch in smart grids. Electricity 2025, 6, 49. [Google Scholar] [CrossRef] [Scilit]
  58. Singh, S.; Singh, S. Advancements and challenges in integrating renewable energy sources into distribution grid systems: A comprehensive review. J. Energy Resour. Technol. 2024, 146, 090801. [Google Scholar] [CrossRef] [Scilit]
  59. Coban, H.H.; Lewicki, W. Assessing the Efficiency of Hybrid Energy Facilities for Electric Vehicle Charging; Scientific Papers of Silesian University of Technology; Organization and Management Series; Silesian University of Technology Publishing House: Gliwice, Poland, 2023. [Google Scholar]
  60. Shang, Y.; Wu, W.; Guo, J.; Ma, Z.; Sheng, W.; Lv, Z.; Fu, C. Stochastic dispatch of energy storage in microgrids: An augmented reinforcement learning approach. Appl. Energy 2020, 261, 114423. [Google Scholar] [CrossRef] [Scilit]
  61. Phan, B.C.; Lai, Y.C. Control strategy of a hybrid renewable energy system based on reinforcement learning approach for an isolated microgrid. Appl. Sci. 2019, 9, 4001. [Google Scholar] [CrossRef] [Scilit]
  62. Bâra, A.; Oprea, S.V. Trading strategies on local electricity markets using Agent-Based modelling and Reinforcement learning: Vectors to expand energy Communities. Syst. Res. Behav. Sci. 2026, 43, 1188–1211. [Google Scholar] [CrossRef] [Scilit]
  63. Mbuwir, B.V.; Geysen, D.; Spiessens, F.; Deconinck, G. Reinforcement learning for control of flexibility providers in a residential microgrid. IET Smart Grid 2020, 3, 98–107. [Google Scholar]
  64. Hao, J.; Gao, D.W.; Zhang, J.J. Reinforcement learning for building energy optimization through controlling of central HVAC system. IEEE Open Access J. Power Energy 2020, 7, 320–328. [Google Scholar] [CrossRef] [Scilit]
  65. Cao, D.; Hu, W.; Zhao, J.; Zhang, G.; Zhang, B.; Liu, Z.; Chen, Z.; Blaabjerg, F. Reinforcement learning and its applications in modern power and energy systems: A review. J. Mod. Power Syst. Clean Energy 2020, 8, 1029–1042. [Google Scholar] [CrossRef] [Scilit]
  66. Erick, A.O.; Folly, K.A. Reinforcement learning approaches to power management in grid-tied microgrids: A review. In Proceedings of the 2020 Clemson University Power Systems Conference (PSC); IEEE: Piscataway, NJ, USA, 2020; pp. 1–6. [Google Scholar]
  67. Li, M.; Mour, N.; Smith, L. Machine learning based on reinforcement learning for smart grids: Predictive analytics in renewable energy management. Sustain. Cities Soc. 2024, 109, 105510. [Google Scholar] [CrossRef] [Scilit]
  68. Smart, E.E.; Olanrewaju, L.O.; Usman, J.; Otaru, K.; Umar, D. Artificial Intelligence (AI) in renewable energy forecasting and optimization. Renew. Energy 2025, 10, 11. [Google Scholar]
  69. Vamvakas, D.; Michailidis, P.; Korkas, C.; Kosmatopoulos, E. Review and evaluation of reinforcement learning frameworks on smart grid applications. Energies 2023, 16, 5326. [Google Scholar] [CrossRef] [Scilit]
  70. Michailidis, P.; Michailidis, I.; Vamvakas, D.; Kosmatopoulos, E. Model-free HVAC control in buildings: A review. Energies 2023, 16, 7124. [Google Scholar] [CrossRef] [Scilit]
  71. Michailidis, P.; Michailidis, I.; Kosmatopoulos, E. Review and evaluation of multi-agent control applications for energy management in buildings. Energies 2024, 17, 4835. [Google Scholar] [CrossRef] [Scilit]
  72. Salkuti, S.R.; Ray, P.; Pagidipala, S. Overview of next generation smart grids. In Next Generation Smart Grids: Modeling, Control and Optimization; Springer: Singapore, 2022; pp. 1–28. [Google Scholar]
  73. Butt, O.M.; Zulqarnain, M.; Butt, T.M. Recent advancement in smart grid technology: Future prospects in the electrical power network. Ain Shams Eng. J. 2021, 12, 687–695. [Google Scholar] [CrossRef] [Scilit]
  74. Moreno Escobar, J.J.; Morales Matamoros, O.; Tejeida Padilla, R.; Lina Reyes, I.; Quintana Espinosa, H. A comprehensive review on smart grids: Challenges and opportunities. Sensors 2021, 21, 6978. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  75. Nguyen, V.T.; Truong, T.B.T.; Nguyen, H.V.P.; Nguyen, B.N.; Truong, Q.V.; Nguyen, T.Q.; Le, M.T. Optimizing Wind Turbine Performance Using Fuzzy Logic-Controlled MPPT Algorithms. In Proceedings of the 2025 Asia Meeting on Environment and Electrical Engineering (EEE-AM); IEEE: Piscataway, NJ, USA, 2025; pp. 1–6. [Google Scholar]
  76. Ourahou, M.; Ayrir, W.; Hassouni, B.E.; Haddi, A. Review on smart grid control and reliability in presence of renewable energies: Challenges and prospects. Math. Comput. Simul. 2020, 167, 19–31. [Google Scholar] [CrossRef] [Scilit]
  77. Ali, S.S.; Choi, B.J. State-of-the-art artificial intelligence techniques for distributed smart grids: A review. Electronics 2020, 9, 1030. [Google Scholar] [CrossRef] [Scilit]
  78. Nam, N.B.; Ogliari, E.; Leva, S.; Pafumi, E.; Alberti, D.; Duong, M.Q. Comparative analysis of conformal prediction techniques and machine learning models for very short-term solar power forecasting. Energy AI 2025, 21, 100573. [Google Scholar] [CrossRef] [Scilit]
  79. Kakran, S.; Chanana, S. Smart operations of smart grids integrated with distributed generation: A review. Renew. Sustain. Energy Rev. 2018, 81, 524–535. [Google Scholar] [CrossRef] [Scilit]
  80. Abdullah, A.A.; Hassan, T.M. Smart grid (SG) properties and challenges: An overview. Discov. Energy 2022, 2, 8. [Google Scholar] [CrossRef] [Scilit]
  81. Salkuti, S.R. Challenges, issues and opportunities for the development of smart grid. Int. J. Electr. Comput. Eng. (IJECE) 2020, 10, 1179–1186. [Google Scholar] [CrossRef] [Scilit]
  82. Phuangpornpitak, N.; Tia, S. Opportunities and challenges of integrating renewable energy in smart grid system. Energy Procedia 2013, 34, 282–290. [Google Scholar] [CrossRef] [Scilit]
  83. Bose, B.K. Power electronics, smart grid, and renewable energy systems. Proc. IEEE 2017, 105, 2011–2018. [Google Scholar] [CrossRef] [Scilit]
  84. Minelli, F.; D’Agostino, D.; Reddy, V.J.; Michailidis, P. Towards Resilient Cities: Robust Selection of Rooftop Renewable Energy Technologies in Mediterranean Multifamily Buildings. Energy Eng. J. Assoc. Energy Eng. 2026, 123, 20. [Google Scholar] [CrossRef] [Scilit]
  85. Varma, R.K.; Salama, M. Large-scale photovoltaic solar power integration in transmission and distribution networks. In Proceedings of the 2011 IEEE Power and Energy Society General Meeting; IEEE: Piscataway, NJ, USA, 2011; pp. 1–4. [Google Scholar]
  86. Shapsough, S.; Takrouri, M.; Dhaouadi, R.; Zualkernan, I.A. Using IoT and smart monitoring devices to optimize the efficiency of large-scale distributed solar farms. Wirel. Netw. 2021, 27, 4313–4329. [Google Scholar]
  87. Meena, M.R.S.; Rathore, M.J.S.; Johri, M.S. Grid connected roof top solar power generation: A review. Int. J. Eng. Dev. Res. 2015, 3, 325–330. [Google Scholar]
  88. Yan, R.; Saha, T.K. Voltage variation sensitivity analysis for unbalanced distribution networks due to photovoltaic power fluctuations. IEEE Trans. Power Syst. 2012, 27, 1078–1089. [Google Scholar] [CrossRef] [Scilit]
  89. Minhas, D.M.; Khalid, R.R.; Frey, G. Real-time power balancing in photovoltaic-integrated smart micro-grid. In Proceedings of the IECON 2017—43rd Annual Conference of the IEEE Industrial Electronics Society; IEEE: Piscataway, NJ, USA, 2017; pp. 7469–7474. [Google Scholar]
  90. Hansen, M. Aerodynamics of Wind Turbines; Routledge: Abingdon, UK, 2015. [Google Scholar]
  91. Berg, D.E. Wind energy conversion. In Energy Conversion; CRC Press: Boca Raton, FL, USA, 2017; pp. 851–895. [Google Scholar]
  92. Tavner, P.; Edwards, C.; Brinkman, A.; Spinato, F. Influence of wind speed on wind turbine reliability. Wind Eng. 2006, 30, 55–72. [Google Scholar] [CrossRef] [Scilit]
  93. Boyle, J.; Littler, T.; Foley, A. Review of frequency stability services for grid balancing with wind generation. J. Eng. 2018, 2018, 1061–1065. [Google Scholar] [CrossRef] [Scilit]
  94. Feilat, E.; Azzam, S.; Al-Salaymeh, A. Impact of large PV and wind power plants on voltage and frequency stability of Jordan’s national grid. Sustain. Cities Soc. 2018, 36, 257–271. [Google Scholar] [CrossRef] [Scilit]
  95. Alidrisi, H.; Demirbas, A. Enhanced electricity generation using biomass materials. Energy Sources Part A Recovery Util. Environ. Eff. 2016, 38, 1419–1427. [Google Scholar] [CrossRef] [Scilit]
  96. Haq, Z. Biomass for Electricity Generation; Energy Information Administration: Washington, DC, USA, 2002.
  97. Avwioroko, A.; Ibegbulam, C.; Afriyie, I.; Fesomade, A.T. Smart grid integration of solar and biomass energy sources. Eur. J. Comput. Sci. Inf. Technol. 2024, 12, 1–14. [Google Scholar] [CrossRef] [Scilit]
  98. Perez-Navarro, A.; Alfonso, D.; Álvarez, C.; Ibáñez, F.; Sanchez, C.; Segura, I. Hybrid biomass-wind power plant for reliable energy generation. Renew. Energy 2010, 35, 1436–1443. [Google Scholar] [CrossRef] [Scilit]
  99. Yang, L.; Entchev, E.; Rosato, A.; Sibilio, S. Smart thermal grid with integration of distributed and centralized solar energy systems. Energy 2017, 122, 471–481. [Google Scholar] [CrossRef] [Scilit]
  100. Yousefzadeh, M.; Lenzen, M. Performance of concentrating solar power plants in a whole-of-grid context. Renew. Sustain. Energy Rev. 2019, 114, 109342. [Google Scholar] [CrossRef] [Scilit]
  101. Ogunmodimu, O.; Okoroigwe, E.C. Solar thermal electricity in Nigeria: Prospects and challenges. Energy Policy 2019, 128, 440–448. [Google Scholar] [CrossRef] [Scilit]
  102. Navarro, A.A.; Ramírez, L.; Domínguez, P.; Blanco, M.; Polo, J.; Zarza, E. Review and validation of Solar Thermal Electricity potential methodologies. Energy Convers. Manag. 2016, 126, 42–50. [Google Scholar] [CrossRef] [Scilit]
  103. Kabeyi, M.J.B.; Olanrewaju, O.A. Geothermal wellhead technology power plants in grid electricity generation: A review. Energy Strategy Rev. 2022, 39, 100735. [Google Scholar] [CrossRef] [Scilit]
  104. Acharjya, K.; Jain, N.K.; A, P.; Saraswat, B. Exploring Geothermal Energy’s Potential in Smart Cities Building Climate Control. E3S Web Conf. 2024, 540, 13007. [Google Scholar] [CrossRef] [Scilit]
  105. Nasruddin; Alhamid, M.I.; Daud, Y.; Surachman, A.; Sugiyono, A.; Aditya, H.; Mahlia, T. Potential of geothermal energy for electricity generation in Indonesia: A review. Renew. Sustain. Energy Rev. 2016, 53, 733–740. [Google Scholar] [CrossRef] [Scilit]
  106. Mekonnen, M.; Hoekstra, A.Y. The Water Footprint of Electricity from Hydropower; UNESCO-IHE Institute for Water Education: Delft, The Netherlands, 2011. [Google Scholar]
  107. Sternberg, R. Hydropower’s future, the environment, and global electricity systems. Renew. Sustain. Energy Rev. 2010, 14, 713–723. [Google Scholar] [CrossRef] [Scilit]
  108. Forouzandehmehr, N.; Han, Z.; Zheng, R. Stochastic dynamic game between hydropower plant and thermal power plant in smart grid networks. IEEE Syst. J. 2014, 10, 88–96. [Google Scholar] [CrossRef] [Scilit]
  109. De Silva, T.; Jorgenson, J.; Macknick, J.; Keohan, N.; Miara, A.; Jager, H.; Pracheil, B. Hydropower operation in future power grid with various renewable power integration. Renew. Energy Focus 2022, 43, 329–339. [Google Scholar] [CrossRef] [Scilit]
  110. Bessa, R.; Moreira, C.; Silva, B.; Matos, M. Handling renewable energy variability and uncertainty in power system operation. In Advances in Energy Systems: The Large-Scale Renewable Energy Integration Challenge; John Wiley & Sons Ltd.: Hoboken, NJ, USA, 2019; pp. 1–26. [Google Scholar]
  111. Meliani, M.; Barkany, A.E.; Abbassi, I.E.; Darcherif, A.M.; Mahmoudi, M. Energy management in the smart grid: State-of-the-art and future trends. Int. J. Eng. Bus. Manag. 2021, 13, 18479790211032920. [Google Scholar] [CrossRef] [Scilit]
  112. Keyhani, A. Smart power grids. In Smart Power Grids 2011; Springer: Berlin/Heidelberg, Germany, 2011; pp. 1–25. [Google Scholar]
  113. Eluri, H.B.; Naik, M.G. Challenges of res with integration of power grids, control strategies, optimization techniques of microgrids: A review. Int. J. Renew. Energy Res. (IJRER) 2021, 11, 1–19. [Google Scholar] [CrossRef] [Scilit]
  114. Lewicki, W.; Coban, H.H.; Minelli, F.; Michailidis, P. Freezers in Residential Buildings as a Source of Power Grid Frequency Regulation in Response to the Demand for Innovation Within the Smart City Concept: Thermal–Electric Modeling, Technical Potential and Operational Challenges. Energies 2026, 19, 1608. [Google Scholar] [CrossRef] [Scilit]
  115. Gandoman, F.H.; Ahmadi, A.; Sharaf, A.M.; Siano, P.; Pou, J.; Hredzak, B.; Agelidis, V.G. Review of FACTS technologies and applications for power quality in smart grids with renewable energy systems. Renew. Sustain. Energy Rev. 2018, 82, 502–514. [Google Scholar] [CrossRef] [Scilit]
  116. Gajduk, A.; Todorovski, M.; Kocarev, L. Stability of power grids: An overview. Eur. Phys. J. Spec. Top. 2014, 223, 2387–2409. [Google Scholar] [CrossRef] [Scilit]
  117. Alizadeh, M.I.; Moghaddam, M.P.; Amjady, N.; Siano, P.; Sheikh-El-Eslami, M.K. Flexibility in future power systems with high renewable penetration: A review. Renew. Sustain. Energy Rev. 2016, 57, 1186–1193. [Google Scholar] [CrossRef] [Scilit]
  118. García Vera, Y.E.; Dufo-López, R.; Bernal-Agustín, J.L. Energy management in microgrids with renewable energy sources: A literature review. Appl. Sci. 2019, 9, 3854. [Google Scholar] [CrossRef] [Scilit]
  119. Islam, M.; Yang, F.; Amin, M. Control and optimisation of networked microgrids: A review. IET Renew. Power Gener. 2021, 15, 1133–1148. [Google Scholar] [CrossRef] [Scilit]
  120. Anderson, A.A.; Suryanarayanan, S. Review of energy management and planning of islanded microgrids. CSEE J. Power Energy Syst. 2019, 6, 329–343. [Google Scholar] [CrossRef] [Scilit]
  121. Liu, X.; Su, B. Microgrids—an integration of renewable energy technologies. In Proceedings of the 2008 China International Conference on Electricity Distribution; IEEE: Piscataway, NJ, USA, 2008; pp. 1–7. [Google Scholar]
  122. Rezaie, B.; Esmailzadeh, E.; Dincer, I. Renewable energy options for buildings: Case studies. Energy Build. 2011, 43, 56–65. [Google Scholar] [CrossRef] [Scilit]
  123. Kaya, G.N.; Beyhan, F. A Study on the Potential of Photovoltaic Panels in Existing Buildings: Housing Example in the Mediterranean. Online J. Art Des. 2024, 12, 73–87. [Google Scholar] [CrossRef] [Scilit]
  124. Kaya, G.; BEYHAN, F. Integrated design approaches with photovoltaic panel and solar collectors in building envelope. World J. Environ. Res. 2021, 11, 49–61. [Google Scholar] [CrossRef] [Scilit]
  125. Manic, M.; Wijayasekara, D.; Amarasinghe, K.; Rodriguez-Andina, J.J. Building energy management systems: The age of intelligent and adaptive buildings. IEEE Ind. Electron. Mag. 2016, 10, 25–39. [Google Scholar] [CrossRef] [Scilit]
  126. Talari, S.; Shafie-Khah, M.; Osório, G.J.; Aghaei, J.; Catalão, J.P. Stochastic modelling of renewable energy sources from operators’ point-of-view: A survey. Renew. Sustain. Energy Rev. 2018, 81, 1953–1965. [Google Scholar] [CrossRef] [Scilit]
  127. Ge, L.; Li, J.; Hou, L.; Lai, J. Autonomous Voltage Regulation for Smart Distribution Network with High-Proportion PVs: A Graph Meta-Reinforcement Learning Approach. IEEE Trans. Sustain. Energy 2025, 16, 2768–2781. [Google Scholar] [CrossRef] [Scilit]
  128. Gao, W.; Fan, R.; Qiao, W.; Wang, S.; Gao, D.W. Deep Reinforcement Learning Based Control of Wind Turbines for Fast Frequency Response. IEEE Trans. Ind. Appl. 2025, 61, 8640–8649. [Google Scholar] [CrossRef] [Scilit]
  129. Jiang, Y.; Cui, W.; Zhang, B.; Cortés, J. Stable reinforcement learning for optimal frequency control: A distributed averaging-based integral approach. IEEE Open J. Control Syst. 2022, 1, 194–209. [Google Scholar] [CrossRef] [Scilit]
  130. Bosisio, A.; Soldan, F.; Pisani, M.; Bionda, E.; Belloni, F.; Morotti, A. A Q-Learning algorithm for optimizing On-Load tap changer operation and voltage control in distribution networks with high integration of renewable energy sources. J. Mod. Power Syst. Clean Energy 2025, 13, 2063–2073. [Google Scholar]
  131. Vlachogiannis, J.G.; Hatziargyriou, N.D. Reinforcement learning for reactive power control. IEEE Trans. Power Syst. 2004, 19, 1317–1325. [Google Scholar] [CrossRef] [Scilit]
  132. Rocchetta, R.; Bellani, L.; Compare, M.; Zio, E.; Patelli, E. A reinforcement learning framework for optimal operation and maintenance of power grids. Appl. Energy 2019, 241, 291–301. [Google Scholar] [CrossRef] [Scilit]
  133. Mussi, M.; Pellegrino, L.; Pindaro, O.F.; Restelli, M.; Trovò, F. A Reinforcement Learning controller optimizing costs and battery State of Health in smart grids. J. Energy Storage 2024, 82, 110572. [Google Scholar] [CrossRef] [Scilit]
  134. Kurucan, M. Accurate Frequency Control of Energy Storage Systems with a Symbolic Game Theory Approach. Çukurova Üniversitesi Mühendislik Fakültesi Derg. 2026, 41, 195–212. [Google Scholar] [CrossRef] [Scilit]
  135. Kuznetsova, E.; Li, Y.F.; Ruiz, C.; Zio, E.; Ault, G.; Bell, K. Reinforcement learning for microgrid energy management. Energy 2013, 59, 133–146. [Google Scholar] [CrossRef] [Scilit]
  136. Foruzan, E.; Soh, L.K.; Asgarpoor, S. Reinforcement learning approach for optimal distributed energy management in a microgrid. IEEE Trans. Power Syst. 2018, 33, 5749–5758. [Google Scholar] [CrossRef] [Scilit]
  137. Alshahr, S.; Alshahir, A.; Alnuman, H.; Alanazi, M.D.; Yousef, A.; Abbas, G. Dynamic renewable energy integration for EV charging via model-based reinforcement learning. Ain Shams Eng. J. 2026, 17, 104040. [Google Scholar] [CrossRef] [Scilit]
  138. Bashyal, A.; Alnahas, H.; Boroukhian, T.; Wicaksono, H. Demand response based industrial energy management with focus on consumption of renewable energy: A deep reinforcement learning approach. Procedia Comput. Sci. 2025, 253, 1442–1451. [Google Scholar] [CrossRef] [Scilit]
  139. Li, Z.; Sun, Z.; Meng, Q.; Wang, Y.; Li, Y. Reinforcement learning of room temperature set-point of thermal storage air-conditioning system with demand response. Energy Build. 2022, 259, 111903. [Google Scholar] [CrossRef] [Scilit]
  140. Puterman, M.L. Markov Decision Processes: Discrete Stochastic Dynamic Programming; John Wiley & Sons: Hoboken, NJ, USA, 2014. [Google Scholar]
  141. Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction; MIT Press: Cambridge, MA, USA, 1998; Volume 1. [Google Scholar]
  142. Zhang, K.; Yang, Z.; Başar, T. Multi-agent reinforcement learning: A selective overview of theories and algorithms. In Handbook of Reinforcement Learning and Control; Springer: Cham, Switzerland, 2021; pp. 321–384. [Google Scholar]
  143. Wang, J.; Xu, W.; Gu, Y.; Song, W.; Green, T.C. Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution Networks. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34. [Google Scholar]
  144. Nekoei, H.; Badrinaaraayanan, A.; Sinha, A.; Amini, M.; Rajendran, J.; Mahajan, A.; Chandar, S. Dealing with Non-Stationarity in Decentralized Cooperative Multi-Agent Deep Reinforcement Learning via Multi-Timescale Learning. In Proceedings of the 2nd Conference on Lifelong Learning Agents; PMLR: New York, NY, USA, 2023. [Google Scholar]
  145. Liu, Z.; Zhang, J.; Shi, E.; Liu, Z.; Niyato, D.; Ai, B.; Shen, X. Graph Neural Network Meets Multi-Agent Reinforcement Learning: Fundamentals, Applications, and Future Directions. IEEE Wirel. Commun. 2024, 31, 39–47. [Google Scholar] [CrossRef] [Scilit]
  146. Yang, Y.; Luo, R.; Li, M.; Zhou, M.; Zhang, W.; Wang, J. Mean Field Multi-Agent Reinforcement Learning. In Proceedings of the 35th International Conference on Machine Learning; PMLR: New York, NY, USA, 2018; Volume 80, pp. 5571–5580. [Google Scholar]
  147. Byeon, H. Advances in value-based, policy-based, and deep learning-based reinforcement learning. Int. J. Adv. Comput. Sci. Appl. 2023, 14, 348–354. [Google Scholar] [CrossRef] [Scilit]
  148. Watkins, C.J.; Dayan, P. Q-learning. Mach. Learn. 1992, 8, 279–292. [Google Scholar] [CrossRef] [Scilit]
  149. Fan, J.; Wang, Z.; Xie, Y.; Yang, Z. A theoretical analysis of deep Q-learning. In Proceedings of the Learning for Dynamics and Control; PMLR: New York, NY, USA, 2020; pp. 486–489. [Google Scholar]
  150. Alabdullah, M.H.; Abido, M.A. Microgrid energy management using deep Q-network reinforcement learning. Alex. Eng. J. 2022, 61, 9069–9078. [Google Scholar] [CrossRef] [Scilit]
  151. Sutton, R.S.; McAllester, D.; Singh, S.; Mansour, Y. Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 1999; Volume 12. [Google Scholar]
  152. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
  153. Ioannou, I.; Javaid, S.; Tan, Y.; Vassiliou, V. Autonomous reinforcement learning for intelligent and sustainable autonomous microgrid energy management. Electronics 2025, 14, 2691. [Google Scholar] [CrossRef] [Scilit]
  154. Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep reinforcement learning. arXiv 2015, arXiv:1509.02971. [Google Scholar]
  155. Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the International Conference on Machine Learning; PMLR: New York, NY, USA, 2018; pp. 1861–1870. [Google Scholar]
  156. Fujimoto, S.; van Hoof, H.; Meger, D. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the P35th International Conference on Machine Learning; PMLR: New York, NY, USA, 2018; pp. 1587–1596. [Google Scholar]
  157. Bellemare, M.G.; Dabney, W.; Munos, R. A Distributional Perspective on Reinforcement Learning. In Proceedings of the 34th International Conference on Machine Learning; PMLR: New York, NY, USA, 2017; Volume 70, pp. 449–458. [Google Scholar]
  158. Dabney, W.; Rowland, M.; Bellemare, M.G.; Munos, R. Distributional Reinforcement Learning with Quantile Regression. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Washington, DC, USA, 2018; Volume 32, pp. 2892–2901. [Google Scholar]
  159. Hessel, M.; Modayil, J.; van Hasselt, H.; Schaul, T.; Ostrovski, G.; Dabney, W.; Horgan, D.; Piot, B.; Azar, M.; Silver, D. Rainbow: Combining Improvements in Deep Reinforcement Learning. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Washington, DC, USA, 2018; Volume 32, pp. 3215–3222. [Google Scholar]
  160. Khadka, S.; Tumer, K. Evolution-Guided Policy Gradient in Reinforcement Learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2018; Volume 31. [Google Scholar]
  161. Hansen, N.; Ostermeier, A. Completely Derandomized Self-Adaptation in Evolution Strategies. Evol. Comput. 2001, 9, 159–195. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  162. Zhou, Y.; Zhang, B.; Xu, C.; Lan, T.; Diao, R.; Shi, D.; Wang, Z.; Lee, W.J. A data-driven method for fast ac optimal power flow solutions via deep reinforcement learning. J. Mod. Power Syst. Clean Energy 2020, 8, 1128–1139. [Google Scholar] [CrossRef] [Scilit]
  163. Cao, D.; Hu, W.; Zhao, J.; Huang, Q.; Chen, Z.; Blaabjerg, F. A multi-agent deep reinforcement learning based voltage regulation using coordinated PV inverters. IEEE Trans. Power Syst. 2020, 35, 4120–4123. [Google Scholar] [CrossRef] [Scilit]
  164. Kou, P.; Liang, D.; Wang, C.; Wu, Z.; Gao, L. Safe deep reinforcement learning-based constrained optimal control scheme for active distribution networks. Appl. Energy 2020, 264, 114772. [Google Scholar] [CrossRef] [Scilit]
  165. Yang, Q.; Wang, G.; Sadeghi, A.; Giannakis, G.B.; Sun, J. Two-timescale voltage control in distribution grids using deep reinforcement learning. IEEE Trans. Smart Grid 2019, 11, 2313–2323. [Google Scholar] [CrossRef] [Scilit]
  166. Rozada, S.; Apostolopoulou, D.; Alonso, E. Load frequency control: A deep multi-agent reinforcement learning approach. In Proceedings of the 2020 IEEE Power & Energy Society General Meeting (PESGM); IEEE: Piscataway, NJ, USA, 2020; pp. 1–5. [Google Scholar]
  167. Cao, D.; Zhao, J.; Hu, W.; Ding, F.; Huang, Q.; Chen, Z.; Blaabjerg, F. Data-driven multi-agent deep reinforcement learning for distribution system decentralized voltage control with high penetration of PVs. IEEE Trans. Smart Grid 2021, 12, 4137–4150. [Google Scholar] [CrossRef] [Scilit]
  168. Cao, D.; Zhao, J.; Hu, W.; Ding, F.; Huang, Q.; Chen, Z. Attention enabled multi-agent DRL for decentralized volt-VAR control of active distribution system using PV inverters and SVCs. IEEE Trans. Sustain. Energy 2021, 12, 1582–1592. [Google Scholar] [CrossRef] [Scilit]
  169. Cao, D.; Zhao, J.; Hu, W.; Yu, N.; Ding, F.; Huang, Q.; Chen, Z. Deep reinforcement learning enabled physical-model-free two-timescale voltage control method for active distribution systems. IEEE Trans. Smart Grid 2021, 13, 149–165. [Google Scholar]
  170. Zhang, X.; Liu, Y.; Duan, J.; Qiu, G.; Liu, T.; Liu, J. DDPG-based multi-agent framework for SVC tuning in urban power grid with renewable energy resources. IEEE Trans. Power Syst. 2021, 36, 5465–5475. [Google Scholar] [CrossRef] [Scilit]
  171. Li, H.; He, H. Learning to operate distribution networks with safe deep reinforcement learning. IEEE Trans. Smart Grid 2022, 13, 1860–1872. [Google Scholar] [CrossRef] [Scilit]
  172. Bui, V.H.; Su, W. Real-time operation of distribution network: A deep reinforcement learning-based reconfiguration approach. Sustain. Energy Technol. Assess. 2022, 50, 101841. [Google Scholar] [CrossRef] [Scilit]
  173. Hu, D.; Ye, Z.; Gao, Y.; Ye, Z.; Peng, Y.; Yu, N. Multi-agent deep reinforcement learning for voltage control with coordinated active and reactive power optimization. IEEE Trans. Smart Grid 2022, 13, 4873–4886. [Google Scholar] [CrossRef] [Scilit]
  174. Cui, W.; Li, J.; Zhang, B. Decentralized safe reinforcement learning for inverter-based voltage control. Electr. Power Syst. Res. 2022, 211, 108609. [Google Scholar] [CrossRef] [Scilit]
  175. Khalid, J.; Ramli, M.A.; Khan, M.S.; Hidayat, T. Efficient load frequency control of renewable integrated power system: A twin delayed DDPG-based deep reinforcement learning approach. IEEE Access 2022, 10, 51561–51574. [Google Scholar] [CrossRef] [Scilit]
  176. Li, J.; Yu, T.; Zhang, X. Coordinated automatic generation control of interconnected power system with imitation guided exploration multi-agent deep reinforcement learning. Int. J. Electr. Power Energy Syst. 2022, 136, 107471. [Google Scholar] [CrossRef] [Scilit]
  177. Li, P.; Wei, M.; Ji, H.; Xi, W.; Yu, H.; Wu, J.; Yao, H.; Chen, J. Deep reinforcement learning-based adaptive voltage control of active distribution networks with multi-terminal soft open point. Int. J. Electr. Power Energy Syst. 2022, 141, 108138. [Google Scholar] [CrossRef] [Scilit]
  178. Cao, D.; Zhao, J.; Hu, W.; Ding, F.; Yu, N.; Huang, Q.; Chen, Z. Model-free voltage control of active distribution system with PVs using surrogate model-based deep reinforcement learning. Appl. Energy 2022, 306, 117982. [Google Scholar] [CrossRef] [Scilit]
  179. Wu, Z.; Li, Y.; Gu, W.; Dong, Z.; Zhao, J.; Liu, W.; Zhang, X.P.; Liu, P.; Sun, Q. Multi-timescale voltage control for distribution system based on multi-agent deep reinforcement learning. Int. J. Electr. Power Energy Syst. 2023, 147, 108830. [Google Scholar] [CrossRef] [Scilit]
  180. Li, P.; Shen, J.; Wu, Z.; Yin, M.; Dong, Y.; Han, J. Optimal real-time Voltage/Var control for distribution network: Droop-control based multi-agent deep reinforcement learning. Int. J. Electr. Power Energy Syst. 2023, 153, 109370. [Google Scholar] [CrossRef] [Scilit]
  181. Petrusev, A.; Putratama, M.A.; Rigo-Mariani, R.; Debusschere, V.; Reignier, P.; Hadjsaid, N. Reinforcement learning for robust voltage control in distribution grids under uncertainties. Sustain. Energy Grids Netw. 2023, 33, 100959. [Google Scholar] [CrossRef] [Scilit]
  182. Rehman, A.U.; Ullah, Z.; Qazi, H.S.; Hasanien, H.M.; Khalid, H.M. Reinforcement learning-driven proximal policy optimization-based voltage control for PV and WT integrated power system. Renew. Energy 2024, 227, 120590. [Google Scholar] [CrossRef] [Scilit]
  183. Huang, J.; Zhang, H.; Tian, D.; Zhang, Z.; Yu, C.; Hancke, G.P. Multi-agent deep reinforcement learning with enhanced collaboration for distribution network voltage control. Eng. Appl. Artif. Intell. 2024, 134, 108677. [Google Scholar] [CrossRef] [Scilit]
  184. Zhang, B.; Cao, D.; Hu, W.; Ghias, A.M.; Chen, Z. Physics-Informed Multi-Agent deep reinforcement learning enabled distributed voltage control for active distribution network using PV inverters. Int. J. Electr. Power Energy Syst. 2024, 155, 109641. [Google Scholar] [CrossRef] [Scilit]
  185. Jacob, R.A.; Paul, S.; Chowdhury, S.; Gel, Y.R.; Zhang, J. Real-time outage management in active distribution networks using reinforcement learning over graphs. Nat. Commun. 2024, 15, 4766. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  186. Singh, A.R.; Sujatha, M.; Kadu, A.D.; Bajaj, M.; Addis, H.K.; Sarada, K. A deep learning and IoT-driven framework for real-time adaptive resource allocation and grid optimization in smart energy systems. Sci. Rep. 2025, 15, 19309. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  187. Mamodiya, U.; Kishor, I.; Garine, R.; Ganguly, P.; Naik, N. Artificial intelligence based hybrid solar energy systems with smart materials and adaptive photovoltaics for sustainable power generation. Sci. Rep. 2025, 15, 17370. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  188. Tang, X.; Wang, J. Deep reinforcement learning-based multi-objective optimization for virtual power plants and smart grids: Maximizing renewable energy integration and grid efficiency. Processes 2025, 13, 1809. [Google Scholar] [CrossRef] [Scilit]
  189. Luo, F.; Wang, S.; Lv, Y.; Mu, R.; Fo, J.; Zhang, T.; Xu, J.; Wang, C. Domain knowledge-enhanced graph reinforcement learning method for Volt/Var control in distribution networks. Appl. Energy 2025, 398, 126409. [Google Scholar] [CrossRef] [Scilit]
  190. Faghri, S.; Tahami, H.; Amini, R.; Katiraee, H.; Langeroudi, A.S.G.; Alinejad, M.; Nejati, M.G. Real-time energy flexibility optimization of grid-connected smart building communities with deep reinforcement learning. Sustain. Cities Soc. 2025, 119, 106077. [Google Scholar] [CrossRef] [Scilit]
  191. Nyong-Bassey, B.E.; Giaouris, D.; Patsios, C.; Papadopoulou, S.; Papadopoulos, A.I.; Walker, S.; Voutetakis, S.; Seferlis, P.; Gadoue, S. Reinforcement learning based adaptive power pinch analysis for energy management of stand-alone hybrid energy storage systems considering uncertainty. Energy 2020, 193, 116622. [Google Scholar] [CrossRef] [Scilit]
  192. Samadi, E.; Badri, A.; Ebrahimpour, R. Decentralized multi-agent based energy management of microgrid using reinforcement learning. Int. J. Electr. Power Energy Syst. 2020, 122, 106211. [Google Scholar] [CrossRef] [Scilit]
  193. Lei, L.; Tan, Y.; Dahlenburg, G.; Xiang, W.; Zheng, K. Dynamic energy dispatch based on deep reinforcement learning in IoT-driven smart isolated microgrids. IEEE Internet Things J. 2020, 8, 7938–7953. [Google Scholar]
  194. Shuai, H.; He, H. Online scheduling of a residential microgrid via Monte-Carlo tree search and a learned model. IEEE Trans. Smart Grid 2020, 12, 1073–1087. [Google Scholar] [CrossRef] [Scilit]
  195. Chen, D.; Chen, K.; Li, Z.; Chu, T.; Yao, R.; Qiu, F.; Lin, K. Powernet: Multi-agent deep reinforcement learning for scalable powergrid control. IEEE Trans. Power Syst. 2021, 37, 1007–1017. [Google Scholar]
  196. Li, Y.; Wang, R.; Yang, Z. Optimal scheduling of isolated microgrids using automated reinforcement learning-based multi-period forecasting. IEEE Trans. Sustain. Energy 2021, 13, 159–169. [Google Scholar]
  197. Munir, M.S.; Abedin, S.F.; Tran, N.H.; Han, Z.; Huh, E.N.; Hong, C.S. Risk-aware energy scheduling for edge computing with microgrid: A multi-agent deep reinforcement learning approach. IEEE Trans. Netw. Serv. Manag. 2021, 18, 3476–3497. [Google Scholar] [CrossRef] [Scilit]
  198. Tomin, N.; Shakirov, V.; Kozlov, A.; Sidorov, D.; Kurbatsky, V.; Rehtanz, C.; Lora, E.E. Design and optimal energy management of community microgrids with flexible renewable energy sources. Renew. Energy 2022, 183, 903–921. [Google Scholar] [CrossRef] [Scilit]
  199. Barbalho, P.I.d.N.; Lacerda, V.; Fernandes, R.; Coury, D.V. Deep reinforcement learning-based secondary control for microgrids in islanded mode. Electr. Power Syst. Res. 2022, 212, 108315. [Google Scholar] [CrossRef] [Scilit]
  200. Guo, C.; Wang, X.; Zheng, Y.; Zhang, F. Real-time optimal energy management of microgrid with uncertainties based on deep reinforcement learning. Energy 2022, 238, 121873. [Google Scholar] [CrossRef] [Scilit]
  201. Harrold, D.J.; Cao, J.; Fan, Z. Data-driven battery operation for energy arbitrage using rainbow deep reinforcement learning. Energy 2022, 238, 121958. [Google Scholar] [CrossRef] [Scilit]
  202. Harrold, D.J.; Cao, J.; Fan, Z. Renewable energy integration and microgrid energy trading using multi-agent deep reinforcement learning. Appl. Energy 2022, 318, 119151. [Google Scholar] [CrossRef] [Scilit]
  203. Hu, C.; Cai, Z.; Zhang, Y.; Yan, R.; Cai, Y.; Cen, B. A soft actor-critic deep reinforcement learning method for multi-timescale coordinated operation of microgrids. Prot. Control Mod. Power Syst. 2022, 7, 29. [Google Scholar] [CrossRef] [Scilit]
  204. Xia, Y.; Xu, Y.; Wang, Y.; Mondal, S.; Dasgupta, S.; Gupta, A.K.; Gupta, G.M. A safe policy learning-based method for decentralized and economic frequency control in isolated networked-microgrid systems. IEEE Trans. Sustain. Energy 2022, 13, 1982–1993. [Google Scholar] [CrossRef] [Scilit]
  205. Xiong, L.; Tang, Y.; Mao, S.; Liu, H.; Meng, K.; Dong, Z.; Qian, F. A two-level energy management strategy for multi-microgrid systems with interval prediction and reinforcement learning. IEEE Trans. Circuits Syst. I Regul. Pap. 2022, 69, 1788–1799. [Google Scholar] [CrossRef] [Scilit]
  206. Zhou, K.; Zhou, K.; Yang, S. Reinforcement learning-based scheduling strategy for energy storage in microgrid. J. Energy Storage 2022, 51, 104379. [Google Scholar] [CrossRef] [Scilit]
  207. Cai, W.; Kordabad, A.B.; Gros, S. Energy management in residential microgrid using model predictive control-based reinforcement learning and Shapley value. Eng. Appl. Artif. Intell. 2023, 119, 105793. [Google Scholar] [CrossRef] [Scilit]
  208. Monfaredi, F.; Shayeghi, H.; Siano, P. Multi-agent deep reinforcement learning-based optimal energy management for grid-connected multiple energy carrier microgrids. Int. J. Electr. Power Energy Syst. 2023, 153, 109292. [Google Scholar] [CrossRef] [Scilit]
  209. Binyamin, S.S.; Slama, S.A.B.; Zafar, B. Artificial intelligence-powered energy community management for developing renewable energy systems in smart homes. Energy Strategy Rev. 2024, 51, 101288. [Google Scholar] [CrossRef] [Scilit]
  210. Hedayatnia, A.; Ghafourian, J.; Sepehrzad, R.; Al-Durrad, A.; Anvari-Moghaddam, A. Two-stage data-driven optimal energy management and dynamic real-time operation in networked microgrid based on a deep reinforcement learning approach. Int. J. Electr. Power Energy Syst. 2024, 160, 110142. [Google Scholar] [CrossRef] [Scilit]
  211. Abid, M.S.; Apon, H.J.; Hossain, S.; Ahmed, A.; Ahshan, R.; Lipu, M.H. A novel multi-objective optimization based multi-agent deep reinforcement learning approach for microgrid resources planning. Appl. Energy 2024, 353, 122029. [Google Scholar] [CrossRef] [Scilit]
  212. Benhmidouch, Z.; Moufid, S.; Ait-Omar, A.; Abbou, A.; Laabassi, H.; Kang, M.; Chatri, C.; Ali, I.H.O.; Bouzekri, H.; Baek, J. A novel reinforcement learning policy optimization based adaptive VSG control technique for improved frequency stabilization in AC microgrids. Electr. Power Syst. Res. 2024, 230, 110269. [Google Scholar] [CrossRef] [Scilit]
  213. Rajamallaiah, A.; Karri, S.P.K.; Shankar, Y.R. Deep reinforcement learning based control strategy for voltage regulation of dc-dc buck converter feeding cpls in dc microgrid. IEEE Access 2024, 12, 17419–17430. [Google Scholar] [CrossRef] [Scilit]
  214. Sepehrzad, R.; Langeroudi, A.S.G.; Khodadadi, A.; Adinehpour, S.; Al-Durra, A.; Anvari-Moghaddam, A. An applied deep reinforcement learning approach to control active networked microgrids in smart cities with multi-level participation of battery energy storage system and electric vehicles. Sustain. Cities Soc. 2024, 107, 105352. [Google Scholar] [CrossRef] [Scilit]
  215. Abouzeid, S.I.; Chen, Y.; Zaery, M.; Abido, M.A.; Raza, A.; Abdelhameed, E.H. Load frequency control based on reinforcement learning for microgrids under false data attacks. Comput. Electr. Eng. 2025, 123, 110093. [Google Scholar] [CrossRef] [Scilit]
  216. Barros, E.B.C.; Souza, W.O.; Costa, D.G.; Rocha Filho, G.; Figueiredo, G.B.; Peixoto, M.L.M. Energy management in smart grids: An Edge-Cloud Continuum approach with Deep Q-learning. Future Gener. Comput. Syst. 2025, 165, 107599. [Google Scholar] [CrossRef] [Scilit]
  217. Jia, X.; Xia, Y.; Yan, Z.; Gao, H.; Qiu, D.; Guerrero, J.M.; Li, Z. Coordinated operation of multi-energy microgrids considering green hydrogen and congestion management via a safe policy learning approach. Appl. Energy 2025, 401, 126611. [Google Scholar] [CrossRef] [Scilit]
  218. Li, Y.; Chang, W.; Yang, Q. Deep reinforcement learning based hierarchical energy management for virtual power plant with aggregated multiple heterogeneous microgrids. Appl. Energy 2025, 382, 125333. [Google Scholar] [CrossRef] [Scilit]
  219. Xiong, B.; Zhang, L.; Hu, Y.; Fang, F.; Liu, Q.; Cheng, L. Deep reinforcement learning for optimal microgrid energy management with renewable energy and electric vehicle integration. Appl. Soft Comput. 2025, 176, 113180. [Google Scholar] [CrossRef] [Scilit]
  220. Wang, C.; Wang, M.; Wang, A.; Zhang, X.; Zhang, J.; Ma, H.; Yang, N.; Zhao, Z.; Lai, C.S.; Lai, L.L. Multiagent deep reinforcement learning-based cooperative optimal operation with strong scalability for residential microgrid clusters. Energy 2025, 314, 134165. [Google Scholar] [CrossRef] [Scilit]
  221. Vazquez-Canteli, J.R.; Henze, G.; Nagy, Z. MARLISA: Multi-agent reinforcement learning with iterative sequential action selection for load shaping of grid-interactive connected buildings. In Proceedings of the 7th ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation; Association for Computing Machinery: New York, NY, USA, 2020; pp. 170–179. [Google Scholar]
  222. Xu, X.; Jia, Y.; Xu, Y.; Xu, Z.; Chai, S.; Lai, C.S. A multi-agent reinforcement learning-based data-driven method for home energy management. IEEE Trans. Smart Grid 2020, 11, 3201–3211. [Google Scholar] [CrossRef] [Scilit]
  223. Touzani, S.; Prakash, A.K.; Wang, Z.; Agarwal, S.; Pritoni, M.; Kiran, M.; Brown, R.; Granderson, J. Controlling distributed energy resources via deep reinforcement learning for load flexibility and energy efficiency. Appl. Energy 2021, 304, 117733. [Google Scholar] [CrossRef] [Scilit]
  224. Lissa, P.; Deane, C.; Schukat, M.; Seri, F.; Keane, M.; Barrett, E. Deep reinforcement learning for home energy management system control. Energy AI 2021, 3, 100043. [Google Scholar] [CrossRef] [Scilit]
  225. Pinto, G.; Deltetto, D.; Capozzoli, A. Data-driven district energy management with surrogate models and deep reinforcement learning. Appl. Energy 2021, 304, 117642. [Google Scholar] [CrossRef] [Scilit]
  226. Pinto, G.; Piscitelli, M.S.; Vázquez-Canteli, J.R.; Nagy, Z.; Capozzoli, A. Coordinated energy management for a cluster of buildings through deep reinforcement learning. Energy 2021, 229, 120725. [Google Scholar] [CrossRef] [Scilit]
  227. Gao, Y.; Matsunami, Y.; Miyata, S.; Akashi, Y. Operational optimization for off-grid renewable building energy system using deep reinforcement learning. Appl. Energy 2022, 325, 119783. [Google Scholar] [CrossRef] [Scilit]
  228. Heidari, A.; Maréchal, F.; Khovalyg, D. Reinforcement Learning for proactive operation of residential energy systems by learning stochastic occupant behavior and fluctuating solar energy: Balancing comfort, hygiene and energy use. Appl. Energy 2022, 318, 119206. [Google Scholar] [CrossRef] [Scilit]
  229. Huang, C.; Zhang, H.; Wang, L.; Luo, X.; Song, Y. Mixed deep reinforcement learning considering discrete-continuous hybrid action space for smart home energy management. J. Mod. Power Syst. Clean Energy 2022, 10, 743–754. [Google Scholar] [CrossRef] [Scilit]
  230. Langer, L.; Volling, T. A reinforcement learning approach to home energy management for modulating heat pumps and photovoltaic systems. Appl. Energy 2022, 327, 120020. [Google Scholar] [CrossRef] [Scilit]
  231. Lee, S.; Choi, D.H. Federated reinforcement learning for energy management of multiple smart homes with distributed energy resources. IEEE Trans. Ind. Inform. 2020, 18, 488–497. [Google Scholar] [CrossRef] [Scilit]
  232. Shen, R.; Zhong, S.; Wen, X.; An, Q.; Zheng, R.; Li, Y.; Zhao, J. Multi-agent deep reinforcement learning optimization framework for building energy system with renewable energy. Appl. Energy 2022, 312, 118724. [Google Scholar] [CrossRef] [Scilit]
  233. Lai, B.C.; Chiu, W.Y.; Tsai, Y.P. Multiagent reinforcement learning for community energy management to mitigate peak rebounds under renewable energy uncertainty. IEEE Trans. Emerg. Top. Comput. Intell. 2022, 6, 568–579. [Google Scholar] [CrossRef] [Scilit]
  234. Qiu, D.; Xue, J.; Zhang, T.; Wang, J.; Sun, M. Federated reinforcement learning for smart building joint peer-to-peer energy and carbon allowance trading. Appl. Energy 2023, 333, 120526. [Google Scholar] [CrossRef] [Scilit]
  235. Xie, J.; Ajagekar, A.; You, F. Multi-agent attention-based deep reinforcement learning for demand response in grid-responsive buildings. Appl. Energy 2023, 342, 121162. [Google Scholar] [CrossRef] [Scilit]
  236. Zhou, X.; Du, H.; Sun, Y.; Ren, H.; Cui, P.; Ma, Z. A new framework integrating reinforcement learning, a rule-based expert system, and decision tree analysis to improve building energy flexibility. J. Build. Eng. 2023, 71, 106536. [Google Scholar] [CrossRef] [Scilit]
  237. Deng, X.; Zhang, Y.; Jiang, Y.; Qi, H. A novel operation method for renewable building by combining distributed DC energy system and deep reinforcement learning. Appl. Energy 2024, 353, 122188. [Google Scholar] [CrossRef] [Scilit]
  238. Wang, Z.; Xiao, F.; Ran, Y.; Li, Y.; Xu, Y. Scalable energy management approach of residential hybrid energy system using multi-agent deep reinforcement learning. Appl. Energy 2024, 367, 123414. [Google Scholar] [CrossRef] [Scilit]
  239. Ajagekar, A.; Decardi-Nelson, B.; You, F. Energy management for demand response in networked greenhouses with multi-agent deep reinforcement learning. Appl. Energy 2024, 355, 122349. [Google Scholar] [CrossRef] [Scilit]
  240. Aldahmashi, J.; Ma, X. Real-time energy management in smart homes through deep reinforcement learning. IEEE Access 2024, 12, 43155–43172. [Google Scholar] [CrossRef] [Scilit]
  241. Kang, H.; Jung, S.; Kim, H.; Jeoung, J.; Hong, T. Reinforcement learning-based optimal scheduling model of battery energy storage system at the building level. Renew. Sustain. Energy Rev. 2024, 190, 114054. [Google Scholar] [CrossRef] [Scilit]
  242. Kumari, A.; Kakkar, R.; Tanwar, S.; Garg, D.; Polkowski, Z.; Alqahtani, F.; Tolba, A. Multi-agent-based decentralized residential energy management using deep reinforcement learning. J. Build. Eng. 2024, 87, 109031. [Google Scholar] [CrossRef] [Scilit]
  243. Wang, J.; Wang, Y.; Qiu, D.; Su, H.; Strbac, G.; Gao, Z. Resilient energy management of a multi-energy building under low-temperature district heating: A deep reinforcement learning approach. Appl. Energy 2025, 378, 124780. [Google Scholar] [CrossRef] [Scilit]
  244. Lami, B.; Alsolami, M.; Alferidi, A.; Slama, S.B. A smart microgrid platform integrating AI and deep reinforcement learning for sustainable energy management. Energies 2025, 18, 1157. [Google Scholar] [CrossRef] [Scilit]
  245. Talihati, B.; Fu, S.; Zhang, B.; Zhao, Y.; Wang, Y.; Sun, Y. Community shared ES-PV system for managing electric vehicle loads via multi-agent reinforcement learning. Appl. Energy 2025, 380, 125039. [Google Scholar] [CrossRef] [Scilit]
  246. Liu, H.; Wu, W. Federated reinforcement learning for decentralized voltage control in distribution networks. IEEE Trans. Smart Grid 2022, 13, 3840–3843. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Paper structure.
Figure 1. Paper structure.
Infrastructures 11 00240 g001
Figure 2. Methodology of this review in a PRISMA-type diagram.
Figure 2. Methodology of this review in a PRISMA-type diagram.
Infrastructures 11 00240 g002
Figure 3. Step-by-step operation of RL control in smart grid environments.
Figure 3. Step-by-step operation of RL control in smart grid environments.
Infrastructures 11 00240 g003
Figure 4. MARL dimensions: structure, training and coordination.
Figure 4. MARL dimensions: structure, training and coordination.
Infrastructures 11 00240 g004
Figure 5. Common RL algorithms in smart grid applications.
Figure 5. Common RL algorithms in smart grid applications.
Infrastructures 11 00240 g005
Figure 6. Left: Occurrence of RL methods in overall smart grid applications. Right: Percentage (%) of value-based, policy-based, and actor–critic approaches in overall smart grid applications.
Figure 6. Left: Occurrence of RL methods in overall smart grid applications. Right: Percentage (%) of value-based, policy-based, and actor–critic approaches in overall smart grid applications.
Infrastructures 11 00240 g006
Figure 7. Left: Hybrid RL Schemes Occurrence in the Overall Smart Grid Applications; Right: Percentage (%) of Hybrid RL applications in the Overall Smart Grid Applications.
Figure 7. Left: Hybrid RL Schemes Occurrence in the Overall Smart Grid Applications; Right: Percentage (%) of Hybrid RL applications in the Overall Smart Grid Applications.
Infrastructures 11 00240 g007
Figure 8. Left: MARL occurrence in overall smart grid applications. Right: MARL percentage (%) in overall smart grid applications.
Figure 8. Left: MARL occurrence in overall smart grid applications. Right: MARL percentage (%) in overall smart grid applications.
Infrastructures 11 00240 g008
Figure 9. Left: Occurrence of structure schemes in MARL smart grid applications. Center: Occurrence of training schemes in MARL smart grid applications. Right: Occurrence of coordination schemes in MARL smart grid applications.
Figure 9. Left: Occurrence of structure schemes in MARL smart grid applications. Center: Occurrence of training schemes in MARL smart grid applications. Right: Occurrence of coordination schemes in MARL smart grid applications.
Infrastructures 11 00240 g009
Figure 10. Left: Occurrence of reward terms across smart grid frameworks. Right: Occurrence (%) of reward structure types across smart grid frameworks.
Figure 10. Left: Occurrence of reward terms across smart grid frameworks. Right: Occurrence (%) of reward structure types across smart grid frameworks.
Infrastructures 11 00240 g010
Figure 11. Left: Occurrence of different baseline types across smart grid applications. Right: Percentage of baseline types across smart grid applications.
Figure 11. Left: Occurrence of different baseline types across smart grid applications. Right: Percentage of baseline types across smart grid applications.
Infrastructures 11 00240 g011
Figure 12. Left: Baseline types occurrence in power grid applications. Center: Baseline types occurrence in microgrid applications. Right: Baseline types occurrence in building applications.
Figure 12. Left: Baseline types occurrence in power grid applications. Center: Baseline types occurrence in microgrid applications. Right: Baseline types occurrence in building applications.
Infrastructures 11 00240 g012
Figure 13. Left: Occurrence of data types in power grid applications. Center: Occurrence of data types in microgrid applications. Right: Occurrence of data types in building applications.
Figure 13. Left: Occurrence of data types in power grid applications. Center: Occurrence of data types in microgrid applications. Right: Occurrence of data types in building applications.
Infrastructures 11 00240 g013
Figure 14. Left: Occurrence of control objectives in power grid applications. Center: Occurrence of control objectives in microgrid applications. Right: Occurrence of control objectives in building applications.
Figure 14. Left: Occurrence of control objectives in power grid applications. Center: Occurrence of control objectives in microgrid applications. Right: Occurrence of control objectives in building applications.
Infrastructures 11 00240 g014
Figure 15. Left: RES type percentage (%) in power grid applications. Center: RES type percentage (%) in microgrid applications. Right: RES type percentage (%) in building applications.
Figure 15. Left: RES type percentage (%) in power grid applications. Center: RES type percentage (%) in microgrid applications. Right: RES type percentage (%) in building applications.
Infrastructures 11 00240 g015
Figure 16. Left: Occurrence of different grid types in power grid applications. Center: Occurrence of grid types in microgrid applications. Right: Occurrence of grid types in building applications.
Figure 16. Left: Occurrence of different grid types in power grid applications. Center: Occurrence of grid types in microgrid applications. Right: Occurrence of grid types in building applications.
Infrastructures 11 00240 g016
Table 1. Comparison of RL with classical and hybrid optimization frameworks.
Table 1. Comparison of RL with classical and hybrid optimization frameworks.
FrameworkStrengthsLimitationsKey Positioning
Stochastic OptimizationHandles uncertainty through scenarios or probabilistic forecasts.Requires accurate probability models; can be computationally heavy.Strong for uncertainty-aware scheduling with known probability models.
Robust OptimizationEnsures feasibility under worst-case uncertainty.Can be overly conservative and dependent on the uncertainty set.Strong for conservative operation under bounded RES/load uncertainty.
MPC/MILP/OPFUses explicit models, constraints, and operational limits.Requires accurate models and repeated online optimization.Strong for interpretable and constraint-aware model-based control.
Hybrid Data–Model MethodsCombine data, forecasts, and physical optimization.Depend on model structure and forecast quality.Bridge data adaptivity with model-based feasibility and transparency.
RL/DRL/MARLLearns adaptive policies for nonlinear and dynamic systems.Training can be unstable; feasibility is not inherent.Strong for adaptive, sequential, and fast real-time decision-making.
Table 2. Comparison of contributions between the current review and previous works considering different evaluation attributes.
Table 2. Comparison of contributions between the current review and previous works considering different evaluation attributes.
Evaluation Aspect[65][66][67][68]Current
RL Trends and Statusxx x
RL Typesxx x
Agent Architecturesx x
Reward Design x x
Baseline Approachesxxxxx
Input Data xxxx
Control Objectivesxxxxx
RES Types x xx
Grid Typesxx x
Performance Comparison xxxx
Simulation Tools x
Power Gridsx x
Microgridsxx xx
Buildingsx x x
Papers Evaluation Scale6713-Narrative84
Table 3. Typical renewable energy installations and operational roles across smart grid frameworks.
Table 3. Typical renewable energy installations and operational roles across smart grid frameworks.
FrameworkTypical RESCapacityOperational Role of RES
Power GridsUtility-scale solar PV farms, onshore/offshore wind farms, hydropower plants, biomass power plants, concentrated solar power (CSP)MW–GW scaleBulk electricity generation for national or regional networks, contribution to energy market supply, provision of ancillary services, reduction of fossil fuel generation, grid decarbonization
MicrogridsDistributed PV arrays, small wind turbines, biomass generators, small hydropower plants, hybrid RES systemskW–tens of MWsLocal electricity generation, energy autonomy and resilience, coordination with ESS for balancing RES intermittency, demand-side management, participation in local energy markets
BuildingsRooftop solar PVs, solar thermal collectors, building-integrated photovoltaics, geothermal heat pumps, small wind turbineskW–hundreds of kWsOn-site RES energy production, reduction of electricity demand, demand-response support, integration with HVAC and ESS to improve energy efficiency and flexibility
Table 4. Typical control objectives in smart grid environments.
Table 4. Typical control objectives in smart grid environments.
FrameworkPrimary Control ObjectivesControlled Components
Power GridsVoltage regulation, frequency stabilization, economic dispatch, coordinating RES generation, congestion managementReactive power from RES inverters, transformer tap changers (OLTC), ESS, conventional generators, flexible loads
MicrogridsEnergy scheduling, cost optimization, maximizing RES utilization, storage coordination, islanding management, resilience improvementBattery storage systems, diesel generators, CHP units, EV charging, flexible loads
BuildingsEnergy cost minimization, maximizing RES self-consumption, HVAC comfort control, appliance scheduling, demand response participationHVAC systems, battery storage, EV chargers, smart appliances, rooftop PV coordination
Table 5. Key attributes of RL applications in power grids.
Table 5. Key attributes of RL applications in power grids.
Ref.YearMethodAgentFH/TSBaselinesRESIntegrationTypeDataCit.
[162]2020PPO/ILSingleN/AOPFWTDG/GCE/GridTransmissionReal55
[163]2020DPPGMulti24 h/1 hRL/DroopPVDG/PE//Load/GridADNReal246
[164]2020DPPGSingle24 h/1 hMPCWTDG/GRID/PE/GCEADNSynth135
[165]2020DQN/SOCPSingle24 h/1 hFixed/GreedyPVPE/ GCE/GridADNReal286
[166]2020DDPGMultiN/ADroopPVDG/GCE/GridTransmissionSynth34
[167]2021DDPGMulti24 h/1 hRBCPVDG/ESS/ PE/Load/GridADNSynth197
[168]2021TD3Multi-/1 hRL/SPPVDG/PE/Load/GridADNReal156
[169]2021SACMultiN/ARL/SPPVDG/PE/GCE/GridADNReal153
[170]2021DDPGMultiN/ARL/ACOPFPV/WTDG/PE/Load/GridTransmissionSynth63
[171]2022CPOSingle24 h/1 hRL/SOCPPV/WTDG, ESS/GCE/GridADNReal144
[172]2022DQN/DDPGSingle24 h/1 hGA/PSO/OPFPV/WTDG/ESS/Grid/LoadsADNSynth54
[173]2022EA-MAAC/MPCMulti24 h/1 hRL/SOCPPVESS/EV/GCEADNSynth163
[174]2022REINFORCEMultiN/ARL/LinearPVDG/PE/GridADNSynth40
[175]2022TD3MultiN/ARL/PSO/GA/PIDPV/WT/HydroDG/EV/ESS/GCE/GridTransmissionSynth97
[176]2022TD3/PIDMultiN/ARL/PID/FOPIWTDG/GCE/GridTransmissionSynth43
[177]2022DDPGSingleN/AOPFPV/WTPE/DG/GridADNSynth55
[178]2022DDPG/ANNSingleN/ARL/MPC/SPPVDG/PE/GridTransmissionSynth94
[179]2023DDPGMulti30 m/5 mRL/SOCPPVGCE/ESS/DG/PEADNReal35
[180]2023DDPG/DroopMulti24 h/15 mRBC/RLPVDG/Load/GridADNSynth20
[181]2023TD3/PPOSingle24 h/30 mOPFPVESS/DG/PE/GridADNReal22
[182]2024PPOSingleN/ARBC/MPCPV/WTDG/PE/GridTransmissionSynth56
[183]2024SACMulti12 h/3 mRL/DroopPVDG/PE/Load/GridADNReal20
[184]2024DDPGMulti24 h/3 mRL/OPFPVDG/GridADNReal41
[185]2024PPOSingleN/ARL/SOCPPVDER/Load/GridADNSynth55
[186]2025UnspecifiedMultiN/ARL/ANN/Conv.PV/WTESS/Load/Grid/BiomassDistributionSynth47
[187]2025UnspecifiedSingle-/1 mANN/Conv.PVESS/GridGrid-ConnectedReal37
[188]2025DQNSingle24 h/1 hGA/PSO/MPC/ConvPV/WTDG/ESS/Load/GridDistributionReal33
[189]2025DDPGMulti-/3 mRLPV/WTDG/Load/GridADNReal29
[190]2025TD3/OPFSingle24 h/5 mRL/OPFPVDG/Load/GridADNReal25
Table 6. Summary of RL applications in power grids.
Table 6. Summary of RL applications in power grids.
AuthorSummary
Zhou et al. [162]This study formulates AC OPF as an MDP in which the state includes load (P/Q) and generator setpoints, while the action space consists of continuous adjustments of generator active power and voltage. A PPO actor–critic agent, initialized through imitation learning based on OPF solver outputs, learns a policy that minimizes generation cost while penalizing constraint violations. The reward combines economic efficiency with feasibility. Evaluated on transmission grids, the method achieves approximately 0.6% cost deviation, nearly 100% feasibility, around 7× faster execution than IPS, and single-step convergence in about 99% of cases.
Cao et al. [163]This work formulates voltage regulation as a Markov game in which each PV inverter agent observes local states, including load P/Q and PV generation, and outputs reactive power actions. A MADDPG actor–critic framework with attention enables coordination through centralized training and decentralized execution. The reward is defined on the basis of global voltage deviation, thereby promoting system-wide voltage stability. On IEEE distribution systems, the method reduces average voltage deviation by about 74% compared with droop control (from 0.46% to 0.12%) and reaches 95.1% of centralized optimal performance while maintaining real-time capability.
Kou et al. [164]A centralized DDPG-based safe RL agent is developed for voltage control in an active distribution network. The state includes bus voltages and DG outputs, while the action corresponds to reactive power adjustments of HDTs/LPCs. The reward penalizes voltage violations, power losses, and control effort, and a safety layer is introduced to enforce operational constraints. The method achieves about 45% loss reduction (197.23 → 109.03 kW), maintains voltages within 0.95–1.05 p.u., and converges in roughly 200 episodes.
Yang et al. [165]This study proposes a two-timescale RL-based Volt/VAR control strategy for distribution grids. A DQN agent manages capacitor switching at the slow timescale, whereas smart inverters solve a convex optimization problem for fast reactive power control. The MDP state includes load levels and capacitor status, the actions correspond to discrete capacitor settings, and the reward minimizes cumulative voltage deviation. Validation on a real 47-bus feeder and the IEEE 123-bus system with real PV/load data shows a voltage deviation reduction of about 20–40% relative to fixed, random, and greedy baselines.
Rozada et al. [166]The paper formulates load-frequency control as a multi-agent MDP in which each generator agent observes local states, specifically frequency deviation Δ ω and control signal z i , and produces continuous adjustments Δ z i through MADDPG actor networks with LSTM memory. The reward is defined as an exponential penalty on frequency deviation. Centralized critics use global information during training, while execution remains decentralized. On simulated grids with 2–8 generators, the method achieves near-optimal cumulative rewards (about 900/1000) and restores frequency rapidly after disturbances, outperforming conventional centralized approaches.
Cao et al. [167]This study proposes a multi-agent actor–critic DRL framework for voltage control, in which agents observe local states, including load P/Q, PV generation, and voltages, and output continuous control actions for PVs, BESS, and SVCs. A shared global reward penalizes voltage deviations and constraint violations through a surrogate power-flow model under centralized training and decentralized execution. Tests on IEEE distribution feeders show complete elimination of voltage violations and approximately 30–50% lower voltage deviations than conventional control, with stable performance under stochastic RES conditions.
Cao et al. [168]This work formulates Volt-VAR control as a multi-agent Markov game. Each agent observes local states, including load P/Q and PV output, and outputs continuous reactive power actions for PV inverters and SVCs. An attention-enhanced TD3 framework is adopted under centralized training and decentralized execution. The reward minimizes global voltage deviation while penalizing constraint violations. On IEEE feeders, the method achieves about 0.13% voltage deviation and 96.4% optimality, outperforming MADDPG and decentralized RL baselines.
Cao et al. [169]The paper models voltage control as a Markov game in which each subnetwork agent observes local voltages, PV outputs, and reactive power states and produces continuous inverter control actions, while a higher-level SAC agent schedules OLTC and capacitor operations using global states. The reward penalizes voltage deviations and excessive switching and is computed through a GP-based surrogate power-flow model. Under centralized training and decentralized execution, the method attains near-centralized performance and substantially reduces voltage violations on IEEE feeders compared with stochastic and DRL baselines.
Zhang et al. [170]This study introduces a hierarchical multi-agent RL framework in which system agents (MADDPG) determine voltage references and SVC agents (DDPG) regulate reactive power locally. States include local voltages and global bus voltages for the critic, while actions correspond to reactive power injections and voltage setpoints. The reward penalizes voltage deviations through a dynamic exponential function. Tested on large transmission grids with stochastic PV/wind scenarios, the method fully eliminates voltage violations (within 0.05 p.u.) and achieves orders-of-magnitude faster computation (about 0.002–0.033 s compared with seconds for ACOPF).
Li et al. [171]The work formulates distribution network operation as a CMDP, where the agent observes historical load, RES generation, prices, and BESS states, and outputs hybrid actions controlling DG power, storage, and voltage devices. A CPO-based actor–critic policy enforces constraints through a separate cost function while maximizing economic reward. The reward minimizes procurement and generation costs, whereas constraints penalize voltage, current, and operational violations. On IEEE feeders with real data, the method achieves up to 32.5% cost reduction relative to DDPG, near-zero constraint violations, and performance close to optimal MISOCP (about 11% gap).
Bui et al. [172]This study presents a three-stage DRL framework in which a single agent observes grid states, including DG outputs, loads, switch status, SOC, and prices, and outputs both discrete switching actions and continuous generation/BESS setpoints. The reward penalizes power losses, load shedding, and operating cost through power-flow simulation. Training is carried out offline in a MATLAB-based environment, allowing real-time inference without re-optimization. On microgrid and IEEE 33-bus systems, the method delivers faster decisions and improved operational efficiency compared with conventional optimization, particularly under fault conditions.
Hu et al. [173]The paper develops a multi-agent actor–critic RL framework (EA-MAAC) under CTDE, in which each agent observes local voltages and grid states and outputs continuous P/Q setpoints for PV inverters and EV loads. The reward combines line-loss minimization with voltage-violation penalties, thereby supporting coordinated voltage regulation. Centralized critics with attention enhance scalability, whereas decentralized execution enables real-time operation. On IEEE feeders, the method reduces voltage violations and improves returns relative to MADDPG and SAC, approaching optimal VVO performance with substantially lower computation time.
Cui et al. [174]This work proposes a decentralized policy-gradient RL framework (REINFORCE) that trains local neural-network controllers for each bus, using local voltage deviations as states and reactive power injections as actions. The reward minimizes voltage deviation and control effort, while Lipschitz constraints embedded in the neural networks ensure closed-loop stability. On a simulated IEEE 33-bus distribution grid, the approach achieves about 18.18% cost reduction relative to linear control and 5.26% relative to standard safe RL, while avoiding the instability observed in unconstrained RL.
Khalid et al. [175]A multi-agent TD3 actor–critic framework is employed in which each agent observes local ACE-based states, namely proportional, integral, and derivative components, and outputs continuous PID gain actions. The reward minimizes frequency deviation and tie-line power error, thereby promoting effective load-frequency control. The system is tested on a two-area renewable-integrated transmission grid with PV, wind, EVs, BESS, and thermal units. It achieves approximately 50–60% faster settling time (around 7.5 s versus more than 14 s) and markedly lower oscillations than DDPG, PSO, and GA.
Li et al. [176]The paper formulates multi-area AGC as a continuous multi-agent RL problem, in which each agent observes local frequency deviation, tie-line power, and system states and outputs tuning actions for PI controller gains. The IGE-MATD3 actor–critic framework employs twin critics, delayed updates, and imitation-guided exploration to improve stability and coordination. The reward combines frequency deviation, control effort, and market performance. On a simulated multi-area transmission system, the method reduces frequency deviation by about 30–60% and lowers control cost compared with MADDPG, MATD3, and heuristic PI baselines.
Li et al. [177]This study formulates voltage control as an MDP in which the state includes node voltages and power injections, while the actions correspond to active/reactive power setpoints of M-SOP converters. A DDPG actor–critic agent learns a continuous control policy, and the reward minimizes normalized voltage deviations. A key contribution is an action-masking layer that enforces physical constraints such as power balance and capacity coupling. On an IEEE 33-node active distribution network, the method achieves about 61% reduction in voltage deviation and 90.65% optimality relative to centralized control, demonstrating near-optimal adaptive performance under RES uncertainty.
Cao et al. [178]A DDPG-based actor–critic agent observes node voltages, PV power, and load injections and outputs continuous control actions for PV reactive power, SVC control, and PV curtailment. The reward penalizes voltage deviation and curtailment through a surrogate power-flow model. Trained on simulated distribution networks, the policy achieves about 77% reduction in voltage deviation (3.64% → 0.82%), outperforming MPC and DDQN while enabling real-time control at around 0.001 s.
Wu et al. [179]The work formulates voltage control as an MDP in which multiple agents (OLTC, CBs, PV, ESS) observe grid states such as voltages, loads, SOC, and previous actions, and output coordinated control signals across two timescales. A MADDPG actor–critic framework with a centralized critic supports cooperation, while Gumbel-Softmax is used to handle discrete actions. The reward penalizes global voltage deviation. On an IEEE 123-bus active distribution network, the method maintains voltages within approximately 0.99–1.03 p.u. and achieves performance comparable to optimal control with more than 4500× faster computation (0.177 s versus 812 s), outperforming DDPG/DQN under RES variability.
Li et al. [180]The paper formulates Volt/VAR control as a POMG-based multi-agent actor–critic problem, in which each agent observes local states, including PV generation, load P/Q, node voltages, and neighbor information, and outputs droop-curve parameters rather than direct control actions. The reward minimizes power losses and voltage violations, with an additional linear penalty introduced for training stability. Using MADDPG with centralized training and decentralized execution, agents coordinate through limited communication across subnetworks. On an IEEE 123-bus distribution grid, the method achieves about 25% loss reduction relative to droop control, eliminates voltage violations, and supports real-time inference of roughly 20 ms.
Petrusev et al. [181]This study applies a single-agent actor–critic DRL framework (TD3PG/PPO) in which the agent observes grid voltages, battery P/Q, SOC, and time, and outputs continuous inverter setpoints (P, Q) for ESS. The reward penalizes voltage violations outside 0.95–1.05 p.u. and SOC deviations, thereby maintaining both grid stability and storage availability. A two-stage policy combining offline pre-training with online adaptation improves robustness to unknown or drifting line impedances. On distribution grids, PPO reduces voltage violations from about 19.5% to 1.85%, and under uncertainty yields about 50% improvement over OPF with much faster execution.
Rehman et al. [182]A single-agent PPO actor–critic framework is employed to learn reactive power control policies for PV inverters and SVCs from system states such as bus voltages and RES output variability. The reward penalizes voltage deviations and power losses, thereby encouraging stable operation. Tested on an IEEE 33-bus distribution network, the method achieves ±5% voltage regulation, 68% controllability, and a violation ratio of 0.044, outperforming conventional control methods.
Huang et al. [183]The study formulates voltage control as a Markov game in which each PV inverter agent observes local states (PL, QL, PPV, QPV) and outputs continuous reactive power actions. It adopts an attention-enhanced MASAC framework under CTDE, where critics use global information and attention-weighted inter-agent interactions, while actors operate locally. The reward combines penalties on voltage deviation and reactive power loss. On IEEE distribution systems, the method achieves about 96–97% controllable rate and roughly 30% cost reduction relative to droop control, outperforming MADDPG, MATD3, and SAC variants.
Zhang et al. [184]A multi-agent GT-MADDPG framework is proposed in which each PV inverter observes local states, including voltages and topology embeddings obtained from a GNN, and outputs continuous reactive power actions. The reward penalizes voltage deviations and power losses and is further strengthened through physics-guided constraints introduced via a PINN loss. Training follows CTDE with replay buffers and graph-based representations. On IEEE 33- and 141-bus systems, the method achieves about 20–30% lower voltage deviation than DDPG/MADDPG and runs orders of magnitude faster than OPF, while remaining robust under uncertainty.
Jacob et al. [185]Propose a graph-based PPO framework for outage management in active distribution networks, where a centralized agent learns topology-aware switching and load-shedding actions. States include load, DER generation, voltages, branch flows, topology, and outage information, while rewards maximize supplied energy and penalize voltage violations. Tested on modified IEEE 13-, 34-, and 123-bus systems in OpenDSS, the method achieved near-optimal restoration compared with MISOCP/BPSO, while operating in milliseconds and reducing outage-related energy loss.
Singh et al. [186]Proposes the ORA-DL hybrid DRL framework, which combines DNN-based demand forecasting with multi-agent reinforcement learning for adaptive smart-grid resource allocation. Agents use demand, renewable generation, ESS, and IoT sensor data to make dynamic energy-distribution decisions aimed at improving grid stability, resource utilization, and operating cost. In simulation, ORA-DL outperformed conventional ML/IoT/DRL approaches in prediction accuracy, allocation efficiency, energy wastage, and cost reduction.
Mamodiya et al. [187]Propose a single-agent RL framework for dual-axis solar tracking, combined with CNN-LSTM irradiance forecasting and Edge AI control. The agent observes irradiance forecasts, environmental conditions, and panel orientation, and adjusts azimuth and elevation angles to maximize captured solar energy. Applied to adaptive PV modules with hybrid storage in a real smart-grid solar testbed, the framework improved annual energy yield, spectral absorption, panel temperature, and battery lifetime compared with conventional fixed-tilt and MPPT approaches.
Tang et al. [188]proposes a single-agent DQN-based DRL framework for energy scheduling in a smart-grid-integrated virtual power plant (VPP) with PV, wind, ESS, loads, and grid-market interaction. Using demand, SOC, renewable generation, and price data, the agent optimizes storage and market decisions. Tested on real Shenzhen VPP data, it reduced grid losses and improved renewable utilization compared with conventional optimization methods.
Luo et al. [189]The paper proposes a GAT-enhanced MADDPG framework for Volt/Var control in PV-rich active distribution networks. PV inverter agents use local load, PV generation, reactive power, voltage, and phase-angle information to regulate reactive power and jointly reduce voltage violations and losses. Tested on IEEE 33-bus and IEEE 141-bus systems with real PV/load data, GAMARL reduced voltage-violation frequency from 2.5% to 0.8%, cut violation occurrences by 80.7%, and lowered network losses by 4.7% compared with the best MARL baseline.
Faghri et al. [190]The paper proposes a hybrid TD3 and convex OPF framework for real-time scheduling in grid-connected smart-building communities with PV, EV parking lots, loads, and a distribution network. Using real CAISO 5-min data, the method reduced online computation time by 82.5% compared with OPF alone, while EV flexibility lowered distribution-network operating cost by 5.15% versus uncontrolled charging.
Table 7. Key attributes of RL applications in microgrids.
Table 7. Key attributes of RL applications in microgrids.
Ref.YearMethodAgentFH/TSBaselinesRESIntegrationTypeDataCit.
[191]2020Dyna-QLSingle72 h/-RBC/FixedPVDG/ESSIslandedReal54
[192]2020QLMulti-/1 hRLPV/WTDG/ESS/Load/GridGrid-connectedSynth144
[193]2020DPPGSingle24 h/1 hRLPVDG/ESSIslandedReal100
[194]2020MuZero/OPFSingleN/ARL/MPC/DPPV/WTDG/ESS/Load/GridIslandedSynth93
[195]2021A2CMulti1 s/0.05 sRL/MPCPV/WTDG/GCE/Load/GridIslandedSynth82
[196]2021DDPG/MILPSingle24 h/1 hRL/ANN/MLRPV/WTDG/ESS/LoadIslandedReal239
[197]2021A3CMulti24 h/15 mRLPV/WTDG/ESS/Load/GridGrid-connectedReal55
[198]2022MCTSMultiN/ADeterministicPV/WT/BIODG/ESS/GridNetworkedSynth151
[199]2022DDPGSingle-/50 msDroop/RBCPVDG/ESS/Load/GridIslandedSynth50
[200]2022PPOSingle24 h/1 hRL/SP/MPCPVDG/ESS/Load/GridGrid-connectedReal238
[201]2022DQNSingle-/1 hRL/MPCPV/WTESS/PE/GCE/Load/GridGrid-connectedReal53
[202]2022DDPGMulti24 h/1 hRLPVDG/ESS/Load/GridGrid-connectedReal97
[203]2022SAC/MPCSingle48 h/1 hRL/MPCPVESS/Load/GridGrid-connectedBoth59
[204]2022SACMulti-/30 sRLPVDG/Load/GridIslandedReal61
[205]2022TR-RL/ANNSingle24 h/1 hRLPVDG/Load/GridNetworkedBoth68
[206]2023QLsingle24 h/1 hMILPPVDG/ESS/Load/GridGrid-connectedReal64
[207]2023DPG/MPCSingle30 d/12 hMPCPVDG/Load/GridCommunityReal65
[208]2023DDPGMulti24 h/1 hRL/PSO/MILPPV/WTDG/ESS/Load/GridCommunityReal50
[209]2024QL/ANNMultiN/ARL/MILPPVDG/ESS/EV/Load/GridCommunitySynth50
[210]2024DDPGMulti24 h/1 hRLPV/WTDG/ESS/Load/GridNetworkedSynth88
[211]2024DDPGMultiN/ARL/HybridPV/WTRES/ESS/EVCS/GridGrid-connectedBoth79
[212]2024TD3Single1 s/-RL/Conv.PV/WTDG/ESS/EVCSIslandedSynth43
[213]2024TD3SingleN/ARLPV/WTPE/CPL/DC busIslandedReal *57
[214]2024SAC/ANNMulti5 msPID/FLC/ANNPVDG/ESS/EV/Load/GridIslandedSynth58
[215]2025DDPG/PIDSingleN/APIDPV/WTDG/ESS/EVCSOtherSynth21
[216]2025DQNSingle15 d/1 hGA/ConvPV/WTDG/Load/GridCommunitySynth25
[217]2025SACSingle7 d/1 hRLPV/WTCHP/GT/Load/GridGrid-ConnectedReal36
[218]2025PPOMulti24 h/30 mRLPV/WTDG/ESS/Load/GridGrid-ConnectedReal34
[219]2025PPOSingle24 h/1 hRL/MILP/PSO/RBCPV/WTDG/ESS/EV/Load/GridGrid-ConnectedReal28
[220]2025SACMulti24 h/1 hRL/MPCPVDG/ESS/EV/HVAC/Load/GridCommunityReal50
Real *: Validation using simulation or hardware-in-the-loop results driven by real-world measurement data.
Table 8. Summary of RL applications in microgrids.
Table 8. Summary of RL applications in microgrids.
AuthorSummary
Nyong et al. [191]This study employs a Dyna-Q agent that observes battery SOC and pinch-based energy targets and selects dispatch actions for HESS components, including BAT, FC, EL, and HT. The reward penalizes storage-limit violations and inefficient energy use while encouraging coordinated operation under uncertainty. Applied to an islanded PV + HESS microgrid using real load and irradiance data, the method outperforms deterministic and adaptive baselines in constraint handling and energy management, although no explicit numerical improvements are reported.
Samadi et al. [192]This work applies decentralized multi-agent Q-learning, where each DER and consumer observes local states such as time, prices, and operational constraints, including SOC and ramp limits, and selects discrete generation, consumption, or bidding actions. Rewards reflect individual profit or cost, while a hierarchical EMS clears the global market. Agents learn independently through ϵ -greedy, softmax, or UCB policies without direct communication. Simulation results show that the softmax strategy delivers the best overall performance, increasing DER profit, reducing customer cost, and lowering grid dependence.
Lei et al. [193]The paper formulates microgrid dispatch as a finite-horizon POMDP, in which the agent observes load, PV generation, and battery states and outputs continuous dispatch actions for DG and ESS units. The reward penalizes generation cost and power imbalance. Two actor–critic variants, FH-DDPG and recurrent FH-RDPG, are trained on real microgrid data to address uncertainty under partial observability. Results demonstrate clear cost reductions and improved power balance, with FH-RDPG showing the strongest robustness under stochastic conditions.
Shuai et al. [194]This study develops a model-based RL approach of the MuZero type, where the state includes past RES generation, load, and system variables such as SoC, encoded through LSTM, while the actions correspond to ESS, DG, and grid-exchange dispatch. The reward minimizes operating cost, including fuel, degradation, and curtailment terms. Policy optimization is carried out through MCTS over a learned dynamics model combined with OPF constraints. On a grid-connected residential microgrid, the method achieves near-optimal cost performance without relying on explicit forecasting models.
Chen et al. [195]The work formulates secondary voltage control as a decentralized MARL POMDP. Each DG observes local electrical states, including P/Q, voltages, currents, and neighbor messages, and then selects discrete voltage setpoints between 1.0 and 1.14 pu. The reward penalizes voltage deviations, with stronger penalties applied outside safe limits, while a spatially discounted global reward enhances coordination. Using an actor–critic architecture with LSTM-based communication encoding and action smoothing, PowerNet achieves about 36% better performance than PPO and the highest rewards (approximately 0.74 versus 0.66 for the baselines) on IEEE-based microgrids, while also showing strong scalability.
Li et al. [196]This study proposes a DDPG-based PER-AutoRL framework in which the agent maps historical PV, WT, and load time series to continuous forecasting outputs, with rewards defined through forecasting-error metrics such as MAPE and RMSE. Prioritized replay accelerates learning, while Bayesian optimization is used to tune the architecture and hyperparameters. These RL-based forecasts are then incorporated into a chance-constrained MILP scheduler for an islanded microgrid with PV, WT, MT, ESS, and DR. The method reduces operating cost by up to 38.4% (36.7% at 95% confidence) and lowers spinning reserve requirements in 66.7% of periods, outperforming MLR, ARIMA, LSTM, and RDPG baselines.
Munir et al. [197]This paper presents a multi-agent A3C framework in which agents observe stochastic energy states, including generation, demand, and storage levels, and output scheduling actions for microgrid resources and MEC loads. A global critic with shared networks supports coordinated learning under a CVaR-based reward that penalizes energy mismatch and shortage risk. Applied to a grid-connected microgrid with PV, storage, and grid dispatch, the method achieves about 9.5% lower training loss and roughly 2% higher accuracy than single-agent RL, while also improving stability under uncertainty.
Tomin et al. [198]The study formulates microgrid energy management as an MDP solved through MCTS, where the agent observes PV/wind generation, load, and storage SOC and selects battery charge/discharge and dispatchable generation actions. The reward maximizes economic profit by penalizing curtailment and load shedding, while UCB balances exploration and exploitation. Evaluated on realistic community microgrids with RES, ESS, and grid interaction, the method achieves substantial profit gains, in some cases exceeding 100%, compared with non-cooperative operation.
Barbalho et al. [199]The paper employs a centralized DDPG actor–critic agent whose state includes node voltages at the PCC and BESS buses, together with current and previous system frequency, while the actions are continuous active and reactive power setpoints for BESS units. The reward penalizes voltage and frequency deviations as well as unstable operating regions. Trained on an islanded AC microgrid in Simulink with SG–BESS–PV interactions, the method achieves up to 86% cumulative error reduction and as much as 4.7× improvement over droop control, with near-zero voltage deviation in several cases.
Guo et al. [200]The work formulates microgrid optimal energy management as an MDP in which the state includes PV/WT output, load, electricity price, and ESS SOC, while the actions control MT output and ESS charge/discharge. A PPO actor–critic policy is trained using clipped loss and GAE, with a reward that minimizes operating cost while penalizing constraint violations. In a grid-connected microgrid, the method achieves about 5–6% lower cost than MPC under uncertainty and converges faster (150 versus 250 episodes), while remaining robust to forecast errors.
Harrold et al. [201]This study models battery scheduling as a partially observable MDP. The state includes SOC, demand, RES generation, price, and, optionally, ANN-based forecasts, while the actions correspond to discrete charge/discharge levels. Rainbow DQN, incorporating distributional learning, prioritized replay, and a dueling architecture, optimizes a reward based on normalized cost savings and constraint penalties. Applied to a grid-connected microgrid with PV, WT, and dynamic pricing, the method achieves up to 20.22% higher savings than DQN and outperforms DDPG and LP/MPC-type baselines, reaching approximately £91.5k in total savings.
Harrold et al. [202]This work employs a multi-agent actor–critic framework (MADDPG/TD3/D3PG), in which each agent observes local states such as demand, prices, RES output, and ESS conditions, and outputs continuous charging, discharging, or trading actions. Rewards are agent-specific and based on marginal contribution, thereby aligning local decisions with global profit. Under centralized training and decentralized execution, the framework coordinates heterogeneous storage units in a grid-connected microgrid with HESS and energy trading. MARL achieves higher profits than single-agent control, while energy trading performs better than standard grid export.
Hu et al. [203]The study applies a SAC-based actor–critic agent to the second-stage MDP, where the state includes net load, battery SOC, and supercapacitor SOC, and the actions correct grid power, battery dispatch, and supercapacitor output around first-stage references. The reward penalizes schedule deviations, SOC imbalance, and degradation cost, while entropy regularization improves stability. In a grid-connected microgrid with PV, WT, and hybrid ESS, the method provides real-time correction of scheduling decisions, achieving low operating cost (264 USD total), fast convergence, and sub-second control time (about 0.156 s), thereby outperforming single-stage approaches in robustness.
Xia et al. [204]A multi-agent SAC framework is proposed in which each microgrid agent observes local frequency deviation and historical PV/load profiles and outputs continuous actions for DG setpoints and ESS power through SOC-dependent parameterization. A global reward penalizes frequency deviations and quadratic generation costs, while a safety layer filters unsafe actions. Under centralized training and local execution, the method converges stably in about 6000 episodes, removes unsafe actions after around 3000 episodes, maintains frequency within ±0.04 Hz, and achieves lower operating cost than standard multi-agent SAC and DDPG methods.
Xiong et al. [205]The paper formulates electricity pricing as an MDP in which the state includes time, wholesale price, and interval forecasts of PV and load, while the actions are continuous retail prices for each microgrid. A policy-gradient actor with a Gaussian policy is enhanced by a differentiable trust-region layer (TRLF/TRLW) to stabilize updates under uncertainty. The reward maximizes DSO profit, and a neural-network surrogate replaces direct interaction with the microgrid. On multi-microgrid IEEE systems, TRLF achieves 85.8% higher profit than DQN and about 3× that of PPO, while interval forecasting nearly doubles cumulative return compared with point forecasts.
Zhou et al. [206]The study formulates energy storage scheduling through Q-learning, where the agent observes SOC, PV generation, load demand, and electricity and gas prices and selects discrete battery charge/discharge actions. The reward is defined as the negative total operating cost, including energy, battery, and gas costs. Applied to a grid-connected integrated microgrid with PV and BESS at 1-h resolution, the RL approach achieves 61.17% faster computation than MILP with only 3.13% cost deviation, demonstrating strong real-time potential.
Cai et al. [207]The work formulates energy management as an episodic RL problem in which the state includes household demand, PV generation, and electricity prices, while the actions are energy-trading decisions derived from MPC optimization. The policy is represented by a parameterized MPC controller trained using least-squares Deterministic Policy Gradient to minimize cumulative monthly cost without discounting. A centralized multi-agent framework coordinates prosumers, while Shapley values ensure fair cost allocation. On a residential microgrid with real data, the method reduces collective electricity cost relative to standard MPC and enables efficient cooperative trading.
Monfaredi et al. [208]This study employs a multi-agent DDPG framework with SDAE-based state representation. Agents representing DERs, storage units, and grid interfaces observe high-dimensional states, including generation, load, storage levels, and network constraints, and output continuous scheduling actions. The reward jointly minimizes operating cost and emissions while satisfying system constraints under centralized training and decentralized execution. Applied to a grid-connected multi-energy microgrid with electric and gas networks, the method achieves about 2.9–4.25% cost reduction and 5.2–7.98% emission reduction relative to APSO, DDPG, DQN, and SO/MILP baselines.
Binyamin et al. [209]This paper develops a multi-agent Q-learning framework enhanced with ANN, in which each prosumer or consumer observes states such as SOC, demand, price, and PV availability and selects actions including trading, charging/discharging, and appliance scheduling. The reward captures both cost reduction and user satisfaction, leading to optimal decentralized policies for a P2P energy market. In a grid-connected residential microgrid with PV, ESS, EV loads, and smart appliances, the method achieves 9.3–16.07% higher rewards than the baselines, together with clear user-level cost savings.
Hedayatnia et al. [210]The study proposes a multi-agent DRL framework based on DDPG and DDQN within a Stackelberg-game structure, where the DSO acts as leader and the MGs act as followers. Agents exchange actions and rewards based on load demand, RES output, prices, and ESS states, and learn optimal power exchange and resource dispatch through a global reward combining economic and reliability penalties. Tested on networked grid-connected microgrids over IEEE feeders, the method achieves about 13–17% lower computational burden and 17–26% faster execution than DDQN/DQN, while also improving voltage and frequency stability.
Abid et al. [211]The paper introduces a hybrid multi-agent MADDPG framework embedded within a multi-objective MOAVOA loop. Agents coordinate EV charging and DER allocation using states such as load, RES output, network conditions, and EV SOC. Actions cover the placement, sizing, and operation of RES, BESS, and EV charging stations, while the reward aggregates cost, losses, emissions, and voltage stability into a Pareto-based objective. With centralized critic and decentralized actor learning on realistic distribution simulations, the method improves performance by about 85% over COMA-based baselines and increases EV SOC by up to 154%, indicating strong Pareto efficiency.
Benhmidouch et al. [212]proposes a TD3-based Actor–Critic adaptive VSG controller for frequency regulation in an islanded AC microgrid with inverter-based RES. The agent observes frequency, RoCoF, and a novel frequency-direction state ( Ψ ) and continuously adjusts virtual inertia (J) and damping coefficient (Dp), while rewards penalize frequency nadir, RoCoF, and transient frequency deviations. Evaluated on a 20 kVA grid-forming VSI microgrid, the TD3 Ψ controller achieved a 94.6% training success rate, improved frequency nadir from 49.773 Hz to 49.848 Hz, reduced RoCoF by 30.7% (−3.905 → −2.705 Hz/s), and lowered power overshoot from 6.50% to 1.99% (69.4% reduction) versus conventional VSG.
Raja et al. [213]proposes a modified TD3 actor–critic DRL framework for direct duty-cycle control of a DC–DC buck converter supplying CPLs in a DC microgrid. The agent observes voltage error, voltage states, and their derivatives/integrals, outputs continuous duty-ratio actions, and receives rewards based on voltage-tracking error minimization. Validation was performed on a real-time OPAL-RT HIL testbed under load changes, parameter uncertainty, and large disturbances. Compared with DQN and DDPG, TD3 achieved the best performance, reducing settling time to 1 ms (vs. 12 ms and 8 ms), limiting overshoot to 1.2%, and achieving the lowest IAE (0.0422) and ITSE (0.0009) values.
Sepehrzad et al. [214]develops a hybrid Multi-Agent SAC-based DRL and IPWFNN framework for frequency regulation in an islanded networked microgrid consisting of four interconnected MGs with PV generation, synchronous generators, BESSs, EV batteries, and loads. States include frequency deviations, load demand histories, PV generation data, and BESS operating conditions, while actions regulate generator control signals and BESS charging/discharging participation factors. The reward minimizes frequency deviations and generation costs while enforcing SOC/SOH and operational constraints. Tested on a 4-MG islanded NMG and validated through an OPAL-RT real-time platform, the method achieved >98% accuracy, 61.1% lower computation time, and 7.82% lower computational burden than ANN, fuzzy, and PID controllers while maintaining frequency stability under severe PV and load uncertainties.
Abouzeid et al. [215]A DDPG actor–critic agent is used to learn PID gains by observing frequency error and its integral and optimizing a reward that penalizes squared deviations while encouraging tight regulation. The policy outputs continuous control parameters for the LFC loop of a RES-integrated microgrid with PV, WT, ESS, diesel, EV, and SMES. Training is performed offline under stochastic disturbances and cyber-attacks using replay buffers and target networks. The method achieves better frequency stability and robustness than classical LFC approaches, with faster damping and lower deviations under attack scenarios.
Barros et al. [216]proposes a single-agent DQN framework for residential energy scheduling in a renewable-powered community microgrid. The agent uses appliance states, consumption classes, and production–consumption balance indicators to decide appliance operation, aiming to improve renewable utilization, reduce costs, and preserve comfort. In GridLAB simulations of a 20-house PV/wind microgrid, the method reduced peak demand, lowered household bills, and delivered faster decisions than cloud-based baselines.
Jia et al. [217]proposes HSC-SAC, a single-agent actor–critic framework for the coordinated control of a grid-connected multi-energy microgrid with PV, wind, CHP, gas turbines, electrolyzers, thermal networks, HCNG infrastructure, and SVCs. Using load, RES, price, weather, and network states, the agent minimizes operating cost while enforcing safety constraints. Tested on a modified IEEE 33-bus system, the method reduced cost, removed voltage and congestion violations, and nearly eliminated unsafe exploration after training.
Li et al. [218]The paper proposes a hierarchical multi-agent PPO framework for coordinating four heterogeneous grid-connected microgrids within a VPP, integrating PV, wind turbines, ESS, microturbines, and flexible loads. Using market, load, storage, and device states, the framework optimizes internal pricing, local scheduling, and centralized ESS dispatch. Tested on a real-data VPP setup, it improved reward, reduced operating costs, and achieved much faster decision-making than non-hierarchical PPO and optimization-based baselines.
Li et al. [219]proposes a single-agent PPO actor–critic framework for real-time energy scheduling in a grid-connected microgrid with PV, wind, diesel generators, ESS, EVs with V2G, loads, and grid transactions. Using price, generation, demand, and storage/EV states, the agent optimizes dispatch and charging decisions. Tested with real hourly data, the method outperformed TD3, MILP, PSO, and SC in profit and self-balance performance.
Li et al. [220]develops a decentralized MASAC-based MARL framework for cooperative energy scheduling in grid-connected residential microgrid clusters with PV, ESS, EVs, HVAC systems, loads, and grid interaction. Using real Pecan Street data, the method coordinates storage, EV charging, and HVAC control while reducing energy cost, discomfort, range anxiety, and transformer stress. Compared with MPC and other MARL baselines, it achieved the best cumulative reward and reduced peak transformer demand by 22.6%.
Table 9. Key attributes of RL applications in building energy systems.
Table 9. Key attributes of RL applications in building energy systems.
Ref.YearMethodAgentFH/TSBaselinesRESIntegrationTypeDataCit.
[221]2020SACMulti24 h/1 hRL/RBCPVDG/TSS/HVAC/Load/GridDistrictReal67
[222]2020QlMulti24 h/1 hGAPVDG/HVAC/EVCS/Load/GridResidentReal283
[223]2021DDPGSingleN/ARBCPVDG/HVAC/ESS/GridLabReal *90
[224]2021DQNSingle-/5 mRBCPVDG/HVAC/TSS/Load/GridResidentReal132
[225]2021SACSingle92 d/1 hRBCPVDG/HVACs/TSS/Load/GridDistrictSynth112
[226]2021SACSingle-/5 mRBCPVDG/HVACs/TSS/Load/GridDistrictSynth88
[227]2022TD3Single7 d/1 hRL/RBCPV/BIODG/ESS/GridIndustrialReal61
[228]2022DQNSingleN/ARBCPV/SWHDG/TSS/HVAC/GridResidentReal50
[229]2022DQN/DDPGSingle24 h/1 hRL/RBC/MILPPVDG/ESS/HVAC/Load/GridResidentReal73
[230]2022DDPGSingle1 y/1 hRBC/MPCPVDG/ESS/TSS/HVAC/GridResidentReal54
[231]2022A2CMulti24 h/1 hRL/RBC/MILPPVDG/ESS/HVAC/Load/GridResidentsSynth211
[232]2022D3QNMulti3 d/1 mRBCPV/WTDG/ESS/HVAC/GridOfficeSynth120
[233]2022QlMulti92 d/1 hRBCPVDG/ESS/EVCS/Load/GridCommunityReal59
[234]2023DDPGMulti24 h/1 hRL/ADMMPVDG/ESS/TSS/HVAC/Load/GridCommunityReal90
[235]2023SACMulti1 y/1 hRL/RBCPVDG/ESS/TSS/HVAC/Load/GridDistrictReal80
[236]2023DDPG/RBCSingle-/1 hRBCPVDG/ESS/GridResidentReal54
[237]2024SACSingleN/ARBCPVDG/ESS/EV/LoadResidentReal60
[238]2024PPO/ILMultiN/ARBCPVDG/ESS/HVAC/Load/GridResidentReal60
[239]2024SACMulti-/1 mRL/RBCPVDG/ESS/HVAC/Load/GridGreenhouseReal69
[240]2024PPOSingle24 h/15 mRL/RBC/MILPPVDG/ESS/EV/HVAC/Load/GridResidentReal50
[241]2024PPOSingle24 h/1 hRL/MILPPVDG/ESS/Load/GridResidentReal *64
[242]2024DQNMulti24 h/1 hHeuristicPVEV/Load/GridResidentReal57
[243]2025PPOSingle24 h/1 hRL/MILPPVESS/DHW/HVAC/Load/GridResidentReal *25
[244]2025UnspecifiedSingle24 h/1 hStatic/ConvPV/WTDG/ESS/EV/Load/GridCommunityReal25
[245]2025PPOMulti7 d/1 hRL/MILPPVDG/ESS/EV/Load/GridCommunityReal25
Real *: Cases of validation in real-world settings.
Table 10. Summary of RL applications in building energy systems.
Table 10. Summary of RL applications in building energy systems.
AuthorSummary
Vazquez et al. [221]This study proposes MARLISA, a multi-agent SAC framework in which each building agent observes weather conditions, loads, storage SoC, and PV generation, and outputs continuous storage-control actions. The reward combines building-level and district-level objectives, while coordination is achieved through sequential action selection and shared demand predictions. Applied to PV-equipped buildings with thermal storage in CityLearn, the method achieves approximately 15% peak reduction, 35% ramping reduction, and 10% load-factor improvement relative to RBC, while converging within 1–2 years.
Xu et al. [222]This work applies multi-agent tabular Q-learning to a residential HEMS with PV, HVAC, appliances, and EVs. Each agent observes forecasted electricity prices and PV generation and selects discrete appliance actions such as ON/OFF states, power levels, and EV charging rates. The reward balances electricity cost and user discomfort. Using real PJM data, the method reduces cost by about 44.6% compared with the absence of demand response and clearly outperforms GA with nearly 40× faster computation.
Touzani et al. [223]The study adopts a DDPG actor–critic framework in which the state includes indoor temperature, PV generation, battery state, and electricity prices, while the actions control HVAC setpoints and battery dispatch. The reward penalizes energy cost, comfort violations, and battery-limit violations. Implemented in the real commercial FLEXLAB testbed, the approach improves operational efficiency and achieves up to 39.6% cost reduction compared with RBC while maintaining similar comfort levels.
Lissa et al. [224]This work employs a DQN agent for a residential HEMS, with states including indoor and outdoor temperature, DHW temperature, PV production, and time, and actions corresponding to heat-pump operating modes. The reward penalizes comfort deviations and grid electricity use while encouraging PV-based operation. Compared with RBC, the method achieves average savings of about 8%, with gains reaching up to 16% in some cases, together with 9.5% higher PV utilization and 10% load shifting, while keeping comfort violations below 1%.
Pinto et al. [225]A centralized SAC agent maps a high-dimensional state space, including weather forecasts, prices, SoC, indoor temperature, and loads, to continuous control actions for heat-pump output and storage charging/discharging. The reward combines comfort penalties, storage incentives, and peak-shaving objectives. Trained in CityLearn with LSTM-based building dynamics, the framework coordinates four commercial buildings and achieves 23% peak reduction, 20% PAR reduction, and about 3% cost savings relative to RBC while maintaining near-comfort conditions.
Pinto et al. [226]This study develops a centralized SAC-based controller for a commercial-building cluster in CityLearn. The agent observes weather forecasts, prices, loads, PV generation, and storage SoCs, and outputs continuous actions for eight thermal storage units. The reward promotes both cost minimization and peak reduction. The method learns to shift cooling demand using COP–temperature dynamics and price signals, achieving about 4% cost reduction and up to 12% peak reduction relative to RBC, together with improved PAR and robustness across climates.
Gao et al. [227]The work formulates building energy management as a continuous-action MDP in which a TD3 agent controls power exchange among PV, bioenergy, BESS, and the grid. The state includes time, battery SoC, and generation/load conditions, while the reward penalizes off-grid deviation and unsafe SoC levels. In a simulated renewable building energy system, TD3 improves off-grid deviation and grid-export reduction by about 70–80% compared with DDPG and RBC, while maintaining nearly zero battery safety violations.
Heidari et al. [228]This study uses a DQN agent for residential HVAC and DHW control. The state includes indoor temperature, tank temperature, PV production, weather conditions, and hot-water demand, while the actions correspond to discrete heating decisions. The reward minimizes energy use while enforcing comfort and hygiene constraints, including Legionella-risk considerations. Using a two-stage offline–online training procedure with real residential data, the method achieves 7–60% and 28–75% energy savings relative to rule-based baselines while preserving comfort and safety.
Huang et al. [229]The paper models HEMS as an MDP in which the state includes time, SoC, indoor and outdoor temperature, PV/load history, and appliance states. Actions combine discrete appliance scheduling with continuous HVAC/BESS control. A hybrid MDRL framework uses DQN for discrete actions and DDPG for continuous ones. Applied to a grid-connected residential building with PV, BESS, and HVAC, the method achieves 25.8% cost reduction relative to RBC and about 7% relative to DDPG, while the safe-MDRL variant reduces comfort violations by roughly 80%.
Langer et al. [230]A DDPG-based policy is trained on an MDP whose state includes battery and thermal SoC, demand, PV generation, time features, and temperature. The actions define target SoCs for battery and thermal storage, thereby indirectly controlling the heat pump and energy flows. The reward balances electricity cost and comfort violations while encouraging PV self-consumption and load shifting. Tested on a residential SHEMS with real data, the method reaches about 75% self-sufficiency, clearly outperforms RBC, and approaches MPC performance without requiring forecasts.
Lee et al. [231]The study formulates a distributed MDP in which each home employs three A2C agents for AC, washing machine, and ESS control. States include electricity price, temperature, PV generation, and ESS SOE, while actions schedule appliance energy use. Rewards combine cost minimization with comfort penalties. A federated learning loop aggregates local models through FedSGD without sharing raw data. The approach achieves about 15–35% cost reduction compared with MILP and BEopt, converges faster than standalone DRL, and scales effectively to a larger number of households.
Shen et al. [232]This work formulates building energy control as a multi-agent MDP in which HVAC-zone and battery agents observe local temperatures, RES generation, SoC, and prices, and select discrete airflow or charge/discharge actions. A D3QN framework with PER, FAS, and VDN improves sample efficiency, feasibility, and cooperation. The reward combines energy cost, RES curtailment, and comfort violations. In a simulated commercial building, the method improves comfort by 84%, RES utilization by 43%, and reduces cost by 8% relative to RBC.
Lai et al. [233]The paper proposes a bi-level MARL Stackelberg framework in which a community aggregator uses Q-learning for ESS dispatch, while multiple appliance-level agents apply sequential Q-learning for scheduling. The state includes time, demand, and storage conditions, and the reward combines electricity cost with user dissatisfaction. PV uncertainty is handled through LSTM forecasting. In a residential community, the method achieves 30.8% peak reduction and 9.68% cost savings relative to single-agent RL, as well as 2.7% lower cost than MINLP under uncertainty.
Qiu et al. [234]The study develops a multi-agent DDPG framework with a federated centralized critic. Each building observes local prices, loads, PV generation, temperatures, and ESS/TES states, and outputs continuous HVAC, storage, and trading actions. The reward combines electricity, gas, and carbon costs. A key contribution is the abstracted centralized critic, which relies on community net demand and emissions to preserve privacy while stabilizing learning. In a three-building community MES, the proposed Fed-JPC policy achieves about 6–8% cost reduction compared with MARL baselines and remains within 1.88% of the centralized optimum, with about 100× faster computation.
Xie et al. [235]This work formulates a multi-agent SAC control problem in which each building agent observes local thermal states, PV generation, ESS SoC, prices, weather, and shared building signals, and outputs continuous HVAC and battery actions. The reward combines electricity cost with a global net-load penalty to encourage cooperation. In a district-scale CityLearn testbed with nine mixed residential and commercial buildings, independent SAC does not outperform RBC, whereas the proposed MAAC achieves more than 8% net-load reduction and stronger peak shaving, demonstrating the value of coordinated MARL.
Zhou et al. [236]A DDPG actor–critic agent is combined with a rule-based energy system (RBES) to control continuous battery charge/discharge actions. The state includes PV generation, building demand, electricity prices, and time features, while the reward promotes lower electricity cost and improved PV utilization. The RBES enforces physical heuristics, and RL refines grid–battery interaction. Applied to a real-data PV-equipped building, the method reduces cost by 4.6–7.0% and improves self-consumption by up to 10.6% relative to rule-based baselines.
Deng et al. [237]The study formulates a continuous-control MDP in which a SAC agent observes PV generation, load demand, battery SoC, EV SoC, and power mismatch, and outputs a continuous gain coefficient to regulate DC-bus voltage and indirectly coordinate loads and storage. The reward penalizes power mismatch and user dissatisfaction. In a residential DC microgrid, the method increases PV self-sufficiency by about 15–20% and PV absorption by about 18% compared with MPC and RBC, while also reducing mismatch and discomfort.
Wang et al. [238]This work develops a multi-agent MAPPO framework enhanced with imitation learning. Agents control HVAC and battery dispatch using states that include PV generation, load demand, SoC, TOU prices, and time features. The shared reward balances energy-cost reduction and thermal comfort, thereby encouraging cooperation. Applied to a residential zero-energy home, the method improves PV utilization and storage scheduling. Compared with PI control, it increases self-sufficiency by 34.86–46.10%, self-consumption by 15.78–18.47%, and improves temperature regulation by 1.33 °C toward the setpoint, while converging rapidly in about 50 episodes.
Ajagekar et al. [239]Present an attention-based multi-agent SAC framework for five grid-interactive greenhouses with PV and BESS. Each agent observes climate, weather, PV output, battery SOC, demand, and time variables, and controls battery charging and discharging to reduce net grid demand. Using simulated NYC greenhouses with real NSRDB climate data, the method achieved at least 28% lower grid demand than RBC, with stronger gains in winter, summer, and fall.
Aldahmashi et al. [240]Propose a PPO-based DRL framework for a residential HEMS coordinating PV, ESS, EV, HVAC, and flexible appliances, while jointly managing active and reactive power. The agent uses price, PV, temperature, ESS/EV states, appliance status, and power factor to optimize scheduling, comfort, and power quality. Tested on a simulated smart home with real Australian PV and weather data, PPO reduced electricity cost by 31.5%, outperforming DQN and DDPG, while also improving the average power factor.
Kang et al. [241]Develop a PPO-based RL framework for scheduling a grid-connected residential PV-BESS system. The agent observes PV generation, electricity demand, battery SOC, month, and hour, and controls continuous battery charging and discharging. The reward jointly maximizes self-sufficiency and reduces daily peak load. Using real residential data from South Korea, PPO showed the best stability among A2C, TD3, and SAC, and performed close to the MILP optimum.
Kumari et al. [242]Propose a multi-agent DQN-based residential HEMS in which separate agents manage non-shiftable, shiftable, and controllable household loads, including EVs and PV-equipped homes. Using appliance-demand profiles and hourly electricity prices, the framework learns scheduling decisions that reduce energy use, optimize cost, and increase profit. Tested with real OpenEI load data and PJM prices, it achieved 86% peak-hour energy savings, outperforming the heuristic benchmark (79%) and approaching the ideal benchmark (92%).
Wang et al. [243]The paper develops a PPO-based DRL framework for resilient energy scheduling in a grid-connected residential multi-energy building connected to a low-temperature district heating network, with PV, battery and thermal storage, heat pump, DHW, space heating, and grid interaction. Using real data from Copenhagen, the method reduced daily operating cost from £136.2 to £122.3 compared with DDPG and cut computation time from 8.86 s (MILP) to about 15 ms, showing strong potential for real-time control.
Lami et al. [244]proposes a single-agent DRL framework for energy management in a grid-connected community building microgrid with PV, wind, battery storage, EVs with V2H, P2P trading, loads, and grid interaction. Using real hourly Saudi Arabian data, the method reduced energy costs by 23%, lowered grid dependence by 40%, achieved 85% renewable utilization, and improved overall energy efficiency from 85% to 92%.
Talihati et al. [245]develops a PPO-based MARL framework for a residential community energy system with PV, shared battery storage, EV charging, and grid interaction. Three agents coordinate battery dispatch, EV charging, and ES-PV pricing using community load, PV output, prices, battery SOC, and EV demand. Tested with real Australian data on an IEEE-14 residential network, the method increased PV self-consumption by 66.41%, covered 38.68% of EV demand, reduced community electricity costs by 7.73%, and generated €51,924.65 profit for the ES-PV operator.
Table 11. Critical comparison of RL methods.
Table 11. Critical comparison of RL methods.
RL FamilyBest Suited ForStrengthsLimitationsRepresentative Cases
QL/DQNDiscrete scheduling: appliances, OLTCs/capacitors, BESS modes, finite EMS actionsSimple; interpretable; natural for finite action spacesWeak for continuous control; limited scalability[165,201,206,224,232,233]
PPO/A2CStable EMS, OPF approximation, BESS–EV–HVAC scheduling, resilient operationStable convergence; reliable policy updatesLower sample efficiency; longer training[162,200,241,243]
DDPG/TD3Continuous physical control: PV inverters, SVCs, BESS, M-SOPs, HVAC, EVs, LFCHandles continuous actions; TD3 improves stabilitySensitive training; unsafe exploration without constraints[164,175,177,223,227]
SAC/MASACStochastic RES-rich EMS, storage coordination, building clusters, voltage/frequency controlStrong exploration; robust under uncertaintyComputationally heavier; reward-sensitive[203,204,217,221,235,237]
MARLDistributed PV inverters, prosumers, EVs, buildings, greenhouses, microgrid clustersScalable coordination, decentralized executionNon-stationarity, coordination/privacy complexity[163,168,183,184,220,239,245]
Safe RLVoltage, frequency, OPF, SOC, line-flow, converter, comfort, security-constrained controlSafer operation, better deployabilityMore complex, may be conservative[164,171,174,217,229]
Hybrid RLOPF acceleration, two-timescale VVC, smart building flexibility, VPP scheduling, multi-energy MGsAdaptive and feasible, closer to deploymentRequires models, solvers, forecasts, or surrogates[162,173,184,189,190,194,203,207,218]
Table 12. Scalability mechanisms in MARL for smart grids.
Table 12. Scalability mechanisms in MARL for smart grids.
MechanismMain Challenge AddressedRepresentative PapersMain Advantage
CTDENon-stationarity[163,167,204,234]Stable training with local execution
Attention criticScalability of joint critics[168,173,183,235,239]Selects relevant agent interactions
GNN/GAT criticLarge network topology[184,189]Learns grid-aware coupling structure
Federated MARLPrivacy and data sharing[231,234]Coordinates without raw data exchange
Hierarchical MARLMulti-layer control complexity[177,210,233,245]Separates aggregator and device roles
Mean-field/collective strategyMassive agent populations[220]Replaces all-to-all interaction with aggregate behavior
Explicit communicationLocal observability limits[180,195,214]Exchanges selected neighbor/system information
Fully decentralized MARLCommunication constraints[174,175,192,209]Avoids centralized critics and full communication
Table 13. Soft versus hard constraint-handling strategies in RL-based smart grid applications.
Table 13. Soft versus hard constraint-handling strategies in RL-based smart grid applications.
TypeMechanismRepresentative PapersOverheadSafety
Soft penalty rewardLinear/weighted penalties[163,164,165,167,200,223]Low-MediumLow-Medium
Soft nonlinear rewardQuadratic/exponential/barrier penalties[166,170,183,199,227]MediumMedium
Soft risk-aware rewardCVaR/reliability/mismatch penalties[197,208,210,211]MediumMedium
Hard constrained RLCMDP/CPO/Lagrangian RL[171]Medium-HighHigh
Hard safety layerAction filtering/correction[164,204,229]MediumHigh
Hard action-space designAction masking/physical parameterization[177,180,204]Low-MediumMedium-High
Hard architecture designStability/physics-informed networks[174,184]Medium-HighMedium-High
Hybrid hard constraintOPF/MPC/SOCP correction[165,173,190,194,203,207]HighHigh-Extreme
Table 14. Summary of simulator practices across RL applications in power grids, microgrids, and buildings.
Table 14. Summary of simulator practices across RL applications in power grids, microgrids, and buildings.
SimulatorShort DescriptionPower Grid CasesMicrogrid CasesBuilding CasesTotal
Custom environmentsMATLAB, Simulink, Python, TensorFlow/PyTorch, Keras, PSCAD, TRNSYS, HOMER, or other custom simulators.[163,164,165,166,167,168,169,171,172,173,175,176,177,178,179,182,186,187,188,189][191,192,193,194,196,197,200,201,202,203,204,205,206,208,209,210,211,212,215,217,218,219,220][198,209,222,224,228,229,231,232,233,234,236,237,240,241,243,244]59
Open-source solversReusable solvers such as PYPOWER, PandaPower, OpenDSS, GridLAB-D, OMNeT++, or PGSim.[162,170,180,181,183,184,185][195,199,216][245]11
OpenAI GymGym-compatible or benchmark-style RL environments, including CityLearn and FLEXLAB Gym-style platforms.[190][221,223,225,226,227,235,238,239,242]10
HILOPAL-RT, RT-LAB/HYPERSIM, real testbed, or real-field validation beyond offline simulation.[187][213,214][223]4
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Michailidis, P.; Minelli, F.; Coban, H.H.; Michailidis, I.; Kosmatopoulos, E. Reinforcement Learning for Optimizing Renewable Energy Utilization in Smart Grids: Recent Advances in Power Grids, Microgrids, and Building Energy Systems. Infrastructures 2026, 11, 240. https://doi.org/10.3390/infrastructures11070240

AMA Style

Michailidis P, Minelli F, Coban HH, Michailidis I, Kosmatopoulos E. Reinforcement Learning for Optimizing Renewable Energy Utilization in Smart Grids: Recent Advances in Power Grids, Microgrids, and Building Energy Systems. Infrastructures. 2026; 11(7):240. https://doi.org/10.3390/infrastructures11070240

Chicago/Turabian Style

Michailidis, Panagiotis, Federico Minelli, Hasan Huseyin Coban, Iakovos Michailidis, and Elias Kosmatopoulos. 2026. "Reinforcement Learning for Optimizing Renewable Energy Utilization in Smart Grids: Recent Advances in Power Grids, Microgrids, and Building Energy Systems" Infrastructures 11, no. 7: 240. https://doi.org/10.3390/infrastructures11070240

APA Style

Michailidis, P., Minelli, F., Coban, H. H., Michailidis, I., & Kosmatopoulos, E. (2026). Reinforcement Learning for Optimizing Renewable Energy Utilization in Smart Grids: Recent Advances in Power Grids, Microgrids, and Building Energy Systems. Infrastructures, 11(7), 240. https://doi.org/10.3390/infrastructures11070240

Article Metrics

Back to TopTop