Reinforcement Learning for Optimizing Renewable Energy Utilization in Smart Grids: Recent Advances in Power Grids, Microgrids, and Building Energy Systems
Abstract
1. Introduction
1.1. General
1.2. Previous Works
1.3. Novelty and Contribution
1.4. Paper Structure
- Section 1 introduces the background and motivation of the study, reviews the most relevant existing surveys on RL applications in smart grids, identifies the main gaps in the literature, and presents the novelty and contributions of this work.
- Section 2 describes the methodology followed in this systematic review, including the literature search strategy, article retrieval process, filtering and selection criteria, attribute extraction, quality assessment, and synthesis framework used for the analysis of the selected studies.
- Section 3 provides an overview of RES energy integration in modern smart grid architectures. It first introduces the main renewable energy technologies commonly deployed in smart grid infrastructures, then discusses the key operational environments in which they are integrated, namely, power grids, microgrids, and building energy systems, together with their main control objectives.
- Section 4 presents the mathematical foundations of RL-based control, including the basic RL formulation, multi-agent formulations, and aspects of RL and the main algorithmic families applied in energy systems, such as value-based, policy-based, and actor–critic methods.
- Section 5 organizes the selected RL-based applications in RES-integrated energy systems into structured summary tables. The studies are grouped by application domain, including power grids, microgrids, and building energy systems, while emphasizing the main characteristics of each work.
- Section 6 provides a detailed evaluation of the reviewed studies across multiple analytical dimensions, including algorithmic trends, agent architectures, reward design, baseline comparisons, renewable energy, and other integrated technologies and control objectives across the different smart grid domains.
- Section 7 discusses the main trends identified through the review and outlines promising future research directions for reinforcement learning applications in RES-integrated smart grids, including power grids, microgrids, and buildings.
- Section 8 concludes the paper by summarizing the key findings of the review, highlighting emerging trends, identifying current limitations, and presenting future directions for reinforcement learning in RES-integrated smart grid systems.
2. Methodology
- Article Search and Retrieval: A systematic literature search was carried out using major academic databases that capture a significant portion of high-impact peer-reviewed research in energy systems and artificial intelligence, specifically Scopus and Web of Science (WoS). An appropriate search strategy was designed to identify studies conducting RL applications on RES-integrated smart grid environments including power grids, microgrids, and building energy systems. To access these studies, RL-related keywords were combined with keywords relating to renewable energy technologies and energy management infrastructures. The main query was constructed as follows:["Reinforcement Learning" OR "Deep Reinforcement Learning" OR "RL" OR "DRL"] AND ["Renewable Energy" OR RES OR "Solar PV" OR "Wind Power" OR "Distributed Generation"] AND ["Smart Grid" OR "Power System" OR "Distribution Network" OR Microgrid OR "Energy Management System" OR "Building Energy Management" OR BEMS].The initial search provided several hundred publications, including journal articles and conference proceedings from 2020 to 2025, which corresponds to the time window in which RL has seen significantly expanded applications in energy research. Titles, abstracts, and keywords were then screened to identify studies that characterized RL-based control strategies in renewable-integrated energy systems.
- Filtering and Selection Criteria: After the initial search, filtering was applied to eliminate duplicate or irrelevant low-quality contributions. Duplicate records were first removed in the selected databases, after which the studies were evaluated based on their relevance to RL applications in renewable-integrated energy systems. A minimum citation threshold of 35 (not including self-citations) was adopted to guarantee the inclusion of influential works. For recent publications (e.g., 2025) where the number of citations may still be restricted, this threshold was relaxed to 25 citations if the work was published in a reputable peer-reviewed forum and its methodological contribution was clear. Peer-reviewed journal articles and highly cited conference papers were included if the paper presented RL-based techniques for the operation, control, or optimization of energy systems with a central presence of renewable energy sources and did not treat the RL method in a theoretical manner.
- Data Extraction and Attribute Collection: A detailed set of methodological and system-level attributes were extracted for each included study to enable systematic comparison. The collected attributes included: (a) RL methodology, (b) agent architecture (single-agent or multi-agent), (c) forecast horizon and control timestep time intervals, (d) baseline control strategies used for benchmarking, (e) renewable energy technologies involved (PV, wind turbines, hydropower, biomass, hybrid systems, etc.), (f) other integrated control energy assets (e.g., generators, energy storage systems, EVs, flexible loads, HVAC systems), (g) smart grid framework (power grid, microgrid, or building energy system); (h) type of power grid, microgrid, or building energy framework, (i) real-world or simulation data utilization, (j) reward design, and (k) control objectives (e.g., voltage regulation, frequency control, energy scheduling, cost optimization, comfort, demand response). Special emphasis was placed on the treatment of studies reporting quantitative performance gains such as reduced operational cost, increased renewable utilization, or grid stability improvement.
- Quality Assessment: All included studies were evaluated based on reporting the level of methodological soundness, clarity of the RL formulation, quality of the experimental evaluation, and availability of the reference list and supplementary materials. Priority was given to studies that clearly defined their state, action, and reward structure, provided sufficient details about training process and learning architectures, treated baseline strategies separately from RL-based solutions, and presented quantitative results as measures demonstrating the value of proposed RL-based controllers. Additionally, studies from established journals and conferences from recognized publishers such as Elsevier, IEEE, Springer, and MDPI were prioritized. In particular, the focus was on papers which presented a complete pipeline, from system modeling to control design to performance evaluation.
- Data Synthesis and Comparative Analysis: The final set of included studies was organized systematically based on the relevant smart grid framework (power grid, microgrid, building), RL type, RES technologies, control objectives, and energy subsystems. Overall, the review integrated 84 highly-cited works, including 29 concerning applications of RL in power grid frameworks, 30 concerning RL applications in microgrid frameworks, and 25 concerning applications of RL in building energy frameworks. A summarized organization provides a wider view of the RL application environment of the included studies. Next, detailed statistics and comparison tables were developed to visualize trends in terms of RL method choice, specific optimal configurations, and performance results. A discussion of RL application behavior for RES-integrated grid systems and future RL exploiting opportunities is then provided based on this synthesis. The PRISMA diagram portraying the overall methodology is illustrated in Figure 2.
3. RES Integration in Smart Grid Architectures
3.1. Types of RES in Smart Grid Environments
- Photovoltaic Systems (PV): Such systems convert solar irradiance directly into electrical energy through semiconductor-based cells. Photovoltaics are deployed either as large-scale solar farms connected to transmission networks or as distributed rooftop installations within distribution grids and buildings [84,85,86,87]. The power output of PV systems follows clear daily and seasonal patterns and is highly sensitive to cloud coverage and atmospheric conditions, often resulting in rapid fluctuations that may lead to voltage deviations and power imbalances [88,89].
- Wind Turbines (WT): Wind turbines generate electricity by converting the kinetic energy of wind into mechanical and subsequently electrical power through rotor–generator systems [90,91]. Although wind generation may occur continuously during both day and night, it remains strongly dependent on stochastic wind patterns and site-specific conditions [92]. In large-scale wind farms connected to transmission networks, sudden variations in wind speed can induce power ramping events, posing challenges for frequency regulation and requiring rapid balancing mechanisms [93,94].
- Biomass Energy (BIO): Biomass-based power generation relies on the combustion, gasification, or biochemical conversion of organic materials, including agricultural residues, wood waste, and dedicated energy crops [95,96]. In contrast to intermittent RES, these systems exhibit behavior similar to conventional thermal power plants, offering relatively stable and controllable output [95,97]. Consequently, biomass is able to support baseload or dispatchable generation while also contributing to waste valorization and circular bio-economy strategies [97,98].
- Solar Thermal Energy: Solar thermal technologies capture solar radiation and convert it into thermal energy, which can be used directly for heating or indirectly for electricity generation through thermodynamic cycles such as steam turbines [99]. Concentrated solar power (CSP) systems employ optical components to concentrate solar energy and often incorporate thermal storage, enabling energy dispatch even in the absence of solar radiation [100]. This capability partially mitigates the inherent intermittency of solar resources [101,102].
- Geothermal Energy (GEO): Geothermal systems exploit thermal energy stored beneath the Earth’s surface by extracting high-temperature fluids from underground reservoirs [103]. This energy may be utilized for electricity generation or directly for heating and industrial processes [103,104]. Due to the stable nature of geothermal resources, such systems provide highly reliable and continuous renewable generation with minimal variability [105].
- Hydropower (Hydro): Hydropower systems generate electricity by converting the kinetic and potential energy of water into mechanical energy through turbine–generator units [106]. Reservoir-based hydropower offers significant operational flexibility; stored water can be released on demand, enabling dispatchable generation and ancillary services such as frequency regulation, spinning reserve, and load following [107]. This high degree of controllability makes hydropower one of the most versatile RES technologies in power system operation [108,109].
3.2. Smart Grid Types
- Power Grids: At the largest scale, power grids comprise extensive transmission and distribution networks that deliver electricity from generation facilities to consumers over wide geographic regions [112]. In these systems, RES are typically integrated through utility-scale solar plants, onshore and offshore wind farms, large hydropower facilities connected to high-voltage networks, biomass plants, and concentrated solar power systems [113]. Under such conditions, system operators must continuously balance generation and demand while maintaining voltage and frequency within acceptable limits [114]. Because RES output can vary rapidly, the coordinated management of multiple resources is essential for preserving network stability [115,116,117].
- Microgrids: At an intermediate scale, microgrids integrate distributed RES units such as rooftop PV, small wind turbines, biomass generators, and hybrid renewable systems with energy storage devices, controllable loads, and in some cases backup diesel generators or combined heat and power (CHP) units [118]. Microgrids may operate either in grid-connected mode, exchanging power with the main utility network [119], or islanded mode, where supply and demand must be balanced locally [120,121]. Within these environments, energy management systems are responsible for coordinating distributed resources while ensuring economic efficiency, reliability, and resilience [69].
- Buildings: At the most distributed scale, building energy systems integrate RES directly on the demand side of the electricity network [122]. Such systems may include rooftop PV, solar thermal collectors, geothermal heat pumps (GHP), building-integrated photovoltaics (BIPV), small vertical wind turbines, and battery storage, thereby enabling local generation and energy management [123,124]. In such settings, renewable generation interacts with subsystems such as HVAC, domestic hot water, electric vehicle charging, and smart appliances [125]. Their operation requires continuous adjustment of energy consumption to align with renewable availability while maintaining indoor comfort and minimizing operating costs [43,45].
3.3. Control Objectives in RES-Integrated Smart Grids
- RL Control in RES-Integrated Power Grids: In large-scale power grid environments, RL-based controllers are primarily applied to system-level tasks that demand fast and adaptive decision-making. These applications include voltage regulation in distribution networks with high photovoltaic penetration [127], frequency stabilization in systems with large shares of wind generation [128,129], and economic dispatch under uncertain renewable output. RL agents can learn to adequately to coordinate grid-support resources such as on-load tap changers (OLTCs) [130], reactive power compensators [131], flexible loads [132], and battery energy storage systems [133,134] to maintain stable operating conditions while reducing operational costs.
- RL Control in RES-Integrated Microgrids: In microgrid environments, the main role of RL is to optimize local energy management through the coordinated operation of distributed energy resources and storage systems [135,136]. Since microgrids often rely heavily on renewable generation, their operation requires continuous scheduling of energy flows among generation units, batteries, and loads [137]. RL algorithms enable controllers to learn strategies for managing renewable variability, reducing backup generator fuel consumption, lowering energy costs, and preserving stable operation, particularly under islanded conditions. In this setting, RL-based policies may dynamically determine when energy should be stored, consumed, or exchanged with the main grid.
- RL Control in RES-Integrated Buildings: At the building energy system level, reinforcement learning is primarily applied to demand-side energy management problems [138,139]. Buildings equipped with renewable generation and storage must intelligently coordinate their energy consumption with local renewable production [139]. RL agents can learn to schedule HVAC operation, battery charging and discharging, appliance usage, and electric vehicle charging in ways that maximize the utilization of locally generated renewable energy while maintaining occupant comfort and reducing electricity costs [48]. Through continuous interaction with environmental conditions, occupancy patterns, and energy price signals, RL-based building controllers can progressively adapt their actions to achieve more efficient and flexible energy management [45].
4. Mathematical Concepts of RL-Based Control for Smart Grids
4.1. Steps of RL-Based Control
- 1.
- Environment: The environment represents the energy system under control, such as a power grid, microgrid, or building energy management system. It comprises RES generators, energy storage systems, controllable loads, electricity market signals, and operational constraints. Such elements behave dynamically in response to external factors, including weather conditions, user behavior, and system demand.
- 2.
- State Observation: Real-time data describing the current operating condition of the system are acquired through sensors, smart meters, and monitoring platforms. Typical variables include RES generation levels, energy demand, storage state-of-charge, electricity prices, grid voltage conditions, and environmental factors such as solar irradiance or temperature. Such measurements collectively define the state space perceived by the RL agent.
- 3.
- RL Agent: The RL agent acts as the decision-making entity responsible for selecting actions that enhance system performance. Depending on the application, it may operate in a centralized manner (e.g., at the microgrid level) or within a decentralized multi-agent framework. Actions may involve generation dispatch adjustments, battery charge/discharge control, load modulation, or the regulation of voltage support devices.
- 4.
- Control Action Execution: The selected action is implemented within the energy system through appropriate control interfaces or supervisory mechanisms. In power grids, this may involve reactive power support or transformer tap adjustments; in microgrids, generation and storage scheduling; and in buildings, HVAC control, battery management, or appliance scheduling.
- 5.
- Environment Transition: Following the execution of the control action, the environment transitions to a new state. Renewable generation varies according to weather conditions, demand evolves over time, and storage levels are updated. This updated condition constitutes the next observed state for the RL agent.
- 6.
- Reward Evaluation: A numerical reward is computed to evaluate the effectiveness of the selected action. The reward function reflects the operational objectives of the system. In power grids, this aspect may relate to voltage stability or congestion mitigation; in microgrids, to cost reduction or renewable utilization; and in buildings, to energy efficiency or occupant comfort.
- 7.
- Policy Update: The RL algorithm uses the reward feedback to refine its decision policy. Through repeated interaction with the environment, the agent progressively learns strategies that maximize long-term performance under dynamic and uncertain operating conditions.
4.2. Mathematical Formulation of RL
- S denotes the state space representing the possible operating conditions of the energy system (e.g., renewable generation levels, system demand, storage states, electricity prices).
- A represents the action space, corresponding to the set of control decisions available to the agent (e.g., dispatch control, storage operation, load scheduling, voltage regulation).
- defines the transition probability describing how the system evolves from state s to state after action a is applied.
- represents the reward obtained after executing action a in state s.
- is the discount factor that determines the relative importance of future rewards.
4.3. Multi-Agent RL
- Agents Structure: The network architecture determines how actor and critic functions are distributed across agents and how global system interactions are represented. In centralized critic–decentralized actor schemes, each agent learns an individual policy , while a shared critic evaluates joint actions using global information; this approach supports stable learning in strongly coupled systems such as power grids. By contrast, fully decentralized architectures rely on independent critics and actors, i.e., , which improves scalability but also increases non-stationarity because the agents operate without global awareness. Fully shared architectures assume homogeneous agents and learn a common policy , whereas hybrid structures combine shared and local components, such as a shared critic with partially shared actors or graph-based critics; this offers a practical compromise between scalability and coordination quality in large-scale smart grid settings [43,45].
- Agents Training: The learning paradigm defines how information is used during policy optimization and plays a central role in addressing MARL non-stationarity. Centralized training with decentralized execution (CTDE) is the most widely adopted approach, in which agents are trained using global state–action information but execute their policies using only local observations, i.e., , enabling both coordination and practical deployment. Fully decentralized training eliminates global information exchange and allows agents to update independently, which is attractive for privacy-sensitive or communication-constrained energy systems, although it often results in slower convergence. Fully centralized training learns a joint policy , effectively reducing the problem to a single-agent formulation, albeit with limited scalability [43,45]. Hybrid approaches, including federated or periodically synchronized learning, seek to balance data locality, communication efficiency, and convergence stability in distributed smart grid applications.
- Agents Coordination: The coordination mechanism determines how agents align their decisions within a shared and physically coupled environment. In implicit coordination, agents are linked indirectly through shared rewards or shared critics, e.g., , which promotes system-level objectives such as cost minimization or power balancing without requiring explicit communication. Explicit coordination, on the other hand, involves direct information exchange such as message passing or shared latent representations, enabling agents to condition their actions on the states or intentions of others. This is particularly beneficial in tightly coupled tasks such as voltage regulation or congestion management [43,45]. Emergent coordination arises when cooperative behavior develops naturally through interaction, even in the absence of explicit coordination mechanisms. Hierarchical coordination, in turn, decomposes the system into multiple control layers such as aggregator–prosumer or grid–microgrid structures, offering improved scalability while reflecting the inherently multi-level organization of smart grids.
4.4. Common RL Algorithms
4.4.1. Value-Based Algorithms
- Q-learning (Q-l): One of the most fundamental methods in this family is Q-learning [149]. Its objective is to estimate the optimal action-value function , which represents the maximum expected return obtained by taking action a in state s and then following the optimal policy. The optimal Q-function satisfies the Bellman optimality equation [148,149]:where is the next state reached after taking action a in state s. The Q-values are updated iteratively using the temporal-difference rule [149]:where is the learning rate. In energy applications, Q-learning has been used for switching decisions in distribution networks, storage scheduling in microgrids, and appliance coordination in building energy management systems [148].
- Deep Q-Networks (DQNs): To overcome the limitations of tabular Q-learning in large state spaces, DQNs approximate the Q-function with a neural network parameterized by [149]:The network is trained by minimizing the temporal-difference loss [149]:where denotes the target-network parameters used to stabilize learning. DQNs have been widely applied in smart grids for tasks such as voltage regulation in distribution networks and energy scheduling in renewable-dominated microgrids [150].
4.4.2. Policy-Based Algorithms
- Proximal Policy Optimization (PPO): A widely used method in this family is PPO, which introduces a clipped surrogate objective to ensure stable policy updates. Its objective is given by [152]whereand is the advantage function, which estimates the relative benefit of selecting action in state .
4.4.3. Actor–Critic Algorithms
- Deterministic Policy Gradient (DDPG): DDPG extends deterministic policy gradients using deep neural networks. The actor produces deterministic actions [154]:while the critic evaluates the action-value function:The actor parameters are updated using the gradient [154]:
- Soft Actor-Critic (SAC): Another widely used algorithm is SAC, which includes entropy regularization to promote exploration. Its objective becomes [155]where is the entropy of the policy distribution and controls the exploration level.
- Twin Delayed Deep Deterministic Policy Gradient (TD3): TD3 improves DDPG by addressing overestimation bias and function-approximation error in actor–critic learning. It uses two critic networks, delayed actor updates, and target policy smoothing to stabilize continuous-control learning [156]. The critic target is computed using the minimum of two target critics:where and are the two target critics. The smoothed target action is obtained asThis mechanism reduces optimistic value estimates and improves policy stability in continuous-control problems such as inverter regulation, ESS dispatch, frequency control, and HVAC modulation.
4.4.4. Emerging RL Techniques: Distributional RL and Evolutionary Policy Optimization
5. Key Attributes and Summaries of RL-Based Applications
5.1. Tables Description
- Ref.: Indicates the reference of the corresponding application in the first column.
- Year: Indicates the publication year of each RL application.
- Method: Indicates the specific RL algorithmic methodology employed in each study.
- Agent: Indicates the agent type associated with the corresponding RL methodology, i.e., single-agent or multi-agent RL approach.
- FH/TS: Indicates the forecast horizon (FH) and control time step (TS) of the RL approach, expressed in hours (h) and minutes (m).
- Baseline: Indicates the baseline control approaches used in each study to evaluate the RL method, such as optimal power flow (OPF), droop control (Droop), model predictive control (MPC), fixed control strategy (Fixed), rule-based control (RBC), particle swarm optimization (PSO), genetic algorithm (GA), proportional–integral–derivative (PID) control, stochastic programming (SP), second-order cone programming (SOCP), dynamic programming (DP), artificial neural network (ANN), multiple linear regression (MLR), mixed-integer linear programming (MILP), and alternating-direction method of multipliers (ADMM), among others.
- RES: Indicates the type of integrated RES technology considered in the corresponding smart grid environment, such as photovoltaic (PV), wind turbine (WT), biomass (BIO), solar water heating (SWH), and hydroelectric Energy (Hydro).
- Integration: Indicates the controllable equipment involved in each application within the smart grid environment. These may include distributed generation (DG) (e.g., photovoltaic systems, wind turbines, biomass units, and conventional generators), energy storage systems (ESS) (e.g., battery storage units and thermal storage systems), electrical loads (Load) (e.g., residential, commercial, and industrial demand), electric vehicles (EVs), or electric vehicle charging stations (EVCS) (e.g., EV fleets and charging infrastructure). The grid infrastructure (Grid) (e.g., transmission/distribution lines, buses, substations, and power flow constraints) is considered together with control-oriented devices, which are divided into power electronics (PE) (e.g., inverters, SVC, STATCOM, and converter-based devices) and grid control equipment (GCE) (e.g., OLTC transformers, capacitor banks, voltage regulators, and frequency control mechanisms).
- Type: Indicates the subtype of the smart grid environment. For power grids, this may include active distribution networks (ADNs) and transmission systems; for microgrids, grid-connected, networked, community, and islanded microgrids; and for buildings, residential, office, laboratory, industrial, district-scale, and community-scale systems.
- Citations: Indicates the number of citations of each work according to Scopus.
- Author: Indicates the author and reference of each application in the first column.
- Summary: Provides a brief description of the methodology and main outcomes of the corresponding application.
5.2. RL for Power Grid Applications
5.3. RL for Microgrid Applications
5.4. RL for Building Applications
6. Evaluation
- 1.
- RL Methods: Examines the main RL algorithms used in the literature, highlighting their core structures, features, and emerging trends.
- 2.
- Agent Architectures: Investigates the dominant agent design approaches, with particular attention to multi-agent RL and its dimensions.
- 3.
- Reward Functions: Analyzes how reward functions are designed across RL-based control applications in smart grid systems.
- 4.
- Baseline Control: Reviews the conventional or benchmark control strategies used to evaluate RL performance.
- 5.
- Utilized Data: Examines the types of data used to develop, train, and test RL-based frameworks.
- 6.
- Control Objectives: Identifies the most common control goals in RL-based smart grid applications and discusses their main characteristics.
- 7.
- RES Types and Integrated Equipment: Explores the trends associated with different renewable energy sources and other integrated technologies considered in RL applications.
- 8.
- Grid Types: Reviews the main types of power grids, microgrids, and building energy systems considered in the literature, highlighting common trends and limitations.
- 9.
- Performance Comparison: Reviews the integrated research works that offer the most significant contributions to the field based on performance and other aspects in a comparative manner.
- 10.
- Simulation Tools: Reviews the simulation and computational tools that practitioners have utilized to provide results in energy management applications for power grids, microgrids, and buildings.
6.1. RL Methods
6.1.1. RL in Power Grid Applications
6.1.2. RL in Microgrid Applications
6.1.3. RL in Building Applications
6.2. Hybrid RL Methodologies
Critical Comparison of RL Methods
6.3. Agent Architectures
6.4. Reward Functions
6.5. Baseline Control
6.6. Data Utilization
6.7. Control Objectives
6.8. RES Types and Integrated Equipment
6.9. Grid Types
6.10. Performance Comparison
6.11. Simulation Tools
7. Discussion
7.1. Current Trends
- A first major trend is that RL becomes relevant precisely when RES integration transforms control from a static optimization task into a sequential decision problem under uncertainty. With high shares of solar and wind, system states are no longer smooth or easily predictable: voltage can fluctuate rapidly, storage adequacy becomes path-dependent, and local demand must be matched against variable generation over time [163,165,167,183,193,200,203,221,225,230]. In this setting, the value of RL is not just in its being data-driven but in being used to learn policies where the quality depends on how present decisions reshape future operating conditions. Therefore, the rise of RL in RES-rich systems should be understood as a consequence of temporal coupling, not merely one of algorithmic fashion.
- A clear dominance of actor–critic methods is observed, which follows from the physical nature of RES flexibility itself. In many modern applications, control variables are continuous: inverter setpoints, storage dispatch, generator output adjustments, HVAC modulation, and EV charging rates all vary over a continuum rather than through a few discrete choices [168,169,175,196,199,211,223,230,237]. Under such conditions, actor–critic methods provide a practical balance between expressive policy learning and tractable optimization. At the same time, lighter value-based approaches remain useful where the action space is discrete, low-dimensional, or strongly structured, such as switching, scheduling, or simplified EMS tasks [162,164,191,192,206,222,224,232]. According to our evaluation, the research is not converging towards one universally superior RL family but towards a task-dependent hierarchy in which algorithm choice is shaped by the granularity and physics of the control problem.
- Another trend concerns the move away from unconstrained black-box policies toward structured and physically meaningful action spaces. As renewable-rich systems become more complex, raw action outputs are increasingly difficult to train safely, interpret reliably, or deploy with confidence. This explains the growing use of parameterized or supervisory RL formulations in which the agent learns droop gains, target SoC values, controller parameters, or higher-level scheduling decisions rather than direct low-level commands [174,175,180,194,205,210,230,237]. The current shift reflects a deeper methodological maturity in which learning is increasingly aligned with engineering abstractions that operators already understand and trust.
- Closely connected to the previous trend, an extensive deployment of hybrid RL approaches is evident. In practice, RES-integrated systems combine hard constraints, mixed time scales, heterogeneous devices, and multiple objectives, making pure RL increasingly insufficient as realism grows. The literature shows a steady movement toward architectures in which learning is combined with optimization layers, forecasting modules, surrogate models, or classical control laws [165,176,184,194,196,211,225,229,236]. In this context, hybridization should be viewed not as a compromise but as a recognition that intelligent control in RES-integrated environments requires both adaptation and structure. In this sense, the field is moving from pure model-free learning to a more disciplined approach in which RL is embedded within broader decision architectures.
- Expansion of multi-agent RL is apparent, driven by the spatial and organizational distribution of RES flexibility. As controllable assets multiply and become more decentralized, coordination can no longer be treated as a side issue. In electrically coupled settings, one agent’s action may affect the feasible operating region of others, while in market-oriented settings multiple decision-makers need to respond to shared uncertainty; likewise, in privacy-sensitive environments it is important for coordination to occur without full state sharing [163,168,179,184,195,197,208,231,234,235]. In this context, the growth of MARL in recent highly cited research on smart grids reflects a more specific trend than the popularity of distributed learning, instead reflecting the structural decentralization introduced by RES integration itself.
- The reviewed studies show that RL design is increasingly shaped by the physical time scale of the RES-related phenomenon under control. Fast feeder disturbances, slower storage scheduling decisions, and even slower thermal dynamics do not admit the same control architectures or training formulations [171,179,183,195,199,204,221,225,230]. The recent literature spans minute-level corrective control, hourly scheduling, and multi-stage supervisory strategies. An important implication is that temporal structure is no longer a secondary modeling detail but is becoming one of the main design pillars of RL-based RES management.
- A final trend concerns the increasing role of reward engineering as an encoded statement of integration quality. As RES integration becomes more multi-objective, the reward is no longer a simple numerical signal but a compact expression of what the system is actually expected to value. Security, renewable utilization, comfort, resilience, cost, degradation, curtailment avoidance, and flexibility provision are all being translated into weighted reward structures that shape the learned behavior [164,167,178,193,197,208,221,226,228]. In this sense, reward design increasingly acts as the bridge between engineering priorities and algorithmic behavior. The more RES-integrated and multi-functional a smart grid ecosystem becomes, the more central this bridge becomes to the success of RL itself.
7.2. Future Directions
- A first and foundational direction is the development of safety-aware RL. As RES penetration rises, the consequences of poor control also intensify: voltage violations, instability, unsafe switching, storage misuse, and comfort rules can no longer be treated as secondary issues. Future work should further develop constrained and shielded RL, CMDP/CPO formulations, Lyapunov-based critics, and explicit safety filters that enforce electrical, thermal, and operational limits during both training and deployment, a trend that is only sporadically detected in recent research [162,164,171,204,223]. Safety needs to become a built-in design principle rather than an add-on.
- Another key direction concerns the broader adoption of physics-informed and model-assisted learning. Purely model-free RL remains sample-inefficient and often difficult to interpret in RES-rich grids. More promising paths include graph-aware encoders, differentiable power flow surrogates, thermal models, digital twins, and learned world models that explicitly capture RES variability and multi-energy coupling. Such approaches could improve generalization, reduce training cost, and align learned policies more closely with real system behavior. Moreover, RES integration simultaneously creates slow scheduling problems, medium-horizon coordination tasks, and fast corrective control requirements. In many cases, a single flat policy is unable to handle all of these. To this end, future architectures should separate long-horizon planning, mid-level coordination, and fast local actuation, especially in systems that combine markets, storage, flexible loads, and inverter-dominated control. Such an approach would better reflect the real temporal organization of smart grid operation.
- Another meaningful future direction is to enable scalable coordination mechanisms. As RES assets become more distributed, RL research need to focus on how agents communicate, what information they share, and how coordination remains effective under privacy, latency, and infrastructure limitations. This makes practices such as graph-based MARL, sparse attention, federated RL, event-triggered coordination, and leader/follower formulations especially relevant for future architectures. The central question is no longer centralized versus decentralized learning but how to efficiently achieve coordination under realistic communication conditions.
- One of the field’s main bottlenecks is that most RES-oriented RL training and testing still takes place within closed simulation loops. Future progress will depend on better use of historical operational data, expert trajectories, and digital twins through offline RL, imitation learning, transfer learning, and meta-RL, followed by limited and safe online adaptation. This is particularly important in infrastructures where online trial-and-error is costly, unsafe, or simply unacceptable.
- Many current studies still optimize expected performance while underrepresenting renewable ramps, forecast errors, rare contingencies, and uncertainties pertaining to user behavior. Future work should expand robust RL, distributional RL, adversarial training, CVaR-aware learning, and probabilistic scenario conditioning so that policies remain reliable under rare but realistic RES-induced disturbances [183,197,201,204]. In renewable-rich systems, average performance alone is not enough; reliability under stress is equally important.
- Considering the adoption of broader objective formulations, future RL should move beyond single-objective cost minimization and learn policies that jointly optimize RES utilization, emissions, flexibility provision, reliability, battery degradation, comfort, and network support value. Multi-objective RL, preference-conditioned policies, and Pareto-based learning are especially relevant here, since RES integration is fundamentally a trade-off problem rather than a task of optimizing one metric. Following this direction would better align RL research with decarbonization goals and the operational reality of grid-interactive energy systems.
- Another important future direction concerns the integration of newer RL methods into practical smart grid control frameworks. Distributional RL appears especially promising for RES-rich systems, since it can capture the variability of future returns rather than focusing only on their average value. This makes it particularly relevant for risk-aware applications such as storage scheduling, market bidding, EV charging, voltage regulation, and resilience management under rare but critical disturbances. Evolutionary policy optimization and population-based training may also become valuable in smart grid problems where the reward landscape is non-smooth, highly constrained, or inherently multi-objective, for example DER planning, controller tuning, hybrid RL–optimization schemes, and coordinated flexibility management. At the same time, graph-based RL and graph neural networks offer strong potential for improving scalability in distribution networks and interconnected microgrids by explicitly capturing electrical topology, while advanced MARL can support coordination among DERs, microgrids, buildings, aggregators, and EV fleets. To become truly practical, these techniques need to be combined with safe, physics-informed, federated, offline, and transfer learning mechanisms so that advanced RL can better handle high-dimensional uncertainty, reduce training effort, preserve privacy, and move closer to reliable deployment in real smart grid environments.
- Finally, validation practice itself must evolve. RL for RES-integrated smart grids needs stronger assessment pipelines based on co-simulation, digital twins, hardware-in-the-loop, controller-in-the-loop, and staged pilot deployment rather than simulation alone. At the same time, explainability, operator trust, fallback logic, and human-in-the-loop supervision should become formal research targets, since real adoption in the future will depend not only on numerical performance but also on transparency and operational acceptance.
8. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| A2C | Advantage Actor–Critic |
| A3C | Asynchronous Advantage Actor–Critic |
| AC | Alternating Current |
| ACOPF | Alternating Current Optimal Power Flow |
| ADMM | Alternating Direction Method of Multipliers |
| ADN | Active Distribution Network |
| AGC | Automatic Generation Control |
| ANN | Artificial Neural Network |
| APSO | Adaptive Particle Swarm Optimization |
| ARIMA | Autoregressive Integrated Moving Average |
| BEMS | Building Energy Management System |
| BESS | Battery Energy Storage System |
| BIPV | Building-Integrated Photovoltaics |
| BIO | Biomass |
| CB | Capacitor Bank |
| CHP | Combined Heat and Power |
| CMDP | Constrained Markov Decision Process |
| CNN | Convolutional Neural Network |
| COP | Coefficient of Performance |
| CPL | Constant Power Load |
| CPO | Constrained Policy Optimization |
| CSP | Concentrated Solar Power |
| CTDE | Centralized Training with Decentralized Execution |
| CVaR | Conditional Value at Risk |
| D3PG | Distributed Distributional Deep Deterministic Policy Gradient |
| D3QN | Dueling Double Deep Q-Network |
| DC | Direct Current |
| DDPG | Deep Deterministic Policy Gradient |
| DDQN | Double Deep Q-Network |
| DER | Distributed Energy Resource |
| DG | Distributed Generation |
| DHW | Domestic Hot Water |
| DL | Deep Learning |
| DNN | Deep Neural Network |
| DP | Dynamic Programming |
| DPG | Deterministic Policy Gradient |
| DQN | Deep Q-Network |
| DR | Demand Response |
| DRL | Deep Reinforcement Learning |
| DSO | Distribution System Operator |
| EA-MAAC | Evolutionary Attention-based Multi-Agent Actor–Critic |
| EL | Electrolyzer |
| EMS | Energy Management System |
| ESS | Energy Storage System |
| EV | Electric Vehicle |
| EVCS | Electric Vehicle Charging Station |
| FAS | Feasible Action Screening |
| FC | Fuel Cell |
| FH | Forecast Horizon |
| FLC | Fuzzy Logic Controller |
| FOPI | Fractional-Order Proportional–Integral |
| GA | Genetic Algorithm |
| GAT | Graph Attention Network |
| GCE | Grid Control Equipment |
| GHP | Geothermal Heat Pump |
| GNN | Graph Neural Network |
| GP | Gaussian Process |
| HCNG | Hydrogen-Compressed Natural Gas |
| HEMS | Home Energy Management System |
| HESS | Hybrid Energy Storage System |
| HIL | Hardware-in-the-Loop |
| HT | Hydrogen Tank |
| HVAC | Heating, Ventilation, and Air Conditioning |
| IAE | Integral Absolute Error |
| IL | Imitation Learning |
| IPS | Interior Point Solver |
| ITSE | Integral of Time-multiplied Squared Error |
| LFC | Load Frequency Control |
| LSTM | Long Short-Term Memory |
| MAPE | Mean Absolute Percentage Error |
| MCTS | Monte Carlo Tree Search |
| MDP | Markov Decision Process |
| MES | Multi-Energy System |
| MG | Microgrid |
| MILP | Mixed-Integer Linear Programming |
| MINLP | Mixed-Integer Nonlinear Programming |
| MLR | Multiple Linear Regression |
| MPC | Model Predictive Control |
| MPPT | Maximum Power Point Tracking |
| N/A | Not Available/Not Applicable |
| NMG | Networked Microgrid |
| OLTC | On-Load Tap Changer |
| OPF | Optimal Power Flow |
| P2P | Peer-to-Peer |
| PAR | Peak-to-Average Ratio |
| PE | Power Electronics |
| PER | Prioritized Experience Replay |
| PID | Proportional–Integral–Derivative |
| PINN | Physics-Informed Neural Network |
| PPO | Proximal Policy Optimization |
| PSO | Particle Swarm Optimization |
| PV | Photovoltaic |
| QL | Q-Learning |
| RBC | Rule-Based Control |
| RES | Renewable Energy Sources |
| RL | Reinforcement Learning |
| RMSE | Root Mean Square Error |
| SAC | Soft Actor–Critic |
| SC | Stochastic Control |
| SDAE | Stacked Denoising Autoencoder |
| SG | Synchronous Generator |
| SoC | State of Charge |
| SOCP | Second-Order Cone Programming |
| SOE | State of Energy |
| SOH | State of Health |
| SP | Stochastic Programming |
| SVC | Static VAR Compensator |
| SWH | Solar Water Heating |
| TD3 | Twin Delayed Deep Deterministic Policy Gradient |
| TOU | Time-of-Use |
| TR-RL | Trust-Region Reinforcement Learning |
| TS | Timestep |
| TSS | Thermal Storage System |
| UCB | Upper Confidence Bound |
| V2G | Vehicle-to-Grid |
| V2H | Vehicle-to-Home |
| VDN | Value Decomposition Network |
| VPP | Virtual Power Plant |
| VSG | Virtual Synchronous Generator |
| VSI | Voltage Source Inverter |
| VVO | Volt/VAR Optimization |
| WT | Wind Turbine |
References
- Nguyen, B.N.; Ogliari, E.; Pafumi, E.; Alberti, D.; Leva, S.; Duong, M.Q. Forecasting generating power of sun tracking PV plant using long-short term memory neural network model: A case study in Ninh Thuan-Vietnam. In Proceedings of the 2024 Tenth International Conference on Communications and Electronics (ICCE); IEEE: Piscataway, NJ, USA, 2024; pp. 333–338. [Google Scholar]
- Hassan, Q.; Viktor, P.; Al-Musawi, T.J.; Ali, B.M.; Algburi, S.; Alzoubi, H.M.; Al-Jiboory, A.K.; Sameen, A.Z.; Salman, H.M.; Jaszczur, M. The renewable energy role in the global energy Transformations. Renew. Energy Focus 2024, 48, 100545. [Google Scholar] [CrossRef] [Scilit]
- Seetharaman; Moorthy, K.; Patwa, N.; Saravanan; Gupta, Y. Breaking barriers in deployment of renewable energy. Heliyon 2019, 5, e01166. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Singh, S. Energy crisis and climate change: Global concerns and their solutions. In Energy: Crises, Challenges and Solutions; John Wiley & Sons Ltd.: Hoboken, NJ, USA, 2021; pp. 1–17. [Google Scholar]
- Nanaki, E.A.; Xydis, G.A. Deployment of renewable energy systems: Barriers, challenges, and opportunities. In Advances in Renewable Energies and Power Technologies; Elsevier: Amsterdam, The Netherlands, 2018; pp. 207–229. [Google Scholar]
- Stephanie, F.; Karl, L. Incorporating renewable energy systems for a new era of grid stability. Fusion Multidiscip. Res. Int. J. 2020, 1, 37–49. [Google Scholar] [CrossRef] [Scilit]
- Reddy, V.J.; Hariram, N.; Ghazali, M.F.; Kumarasamy, S. Pathway to sustainability: An overview of renewable energy integration in building systems. Sustainability 2024, 16, 638. [Google Scholar] [CrossRef] [Scilit]
- D’Agostino, D.; Mazzella, S.; Minelli, F.; Minichiello, F. Obtaining the NZEB target by using photovoltaic systems on the roof for multi-storey buildings. Energy Build. 2022, 267, 112147. [Google Scholar] [CrossRef] [Scilit]
- Fetting, C. The European green deal. ESDN Rep. Dec. 2020, 2, 53. [Google Scholar] [CrossRef] [Scilit]
- Jäger-Waldau, A.; Bodis, K.; Kougias, I.; Szabo, S. The New European Renewable Energy Directive-Opportunities and Challenges for Photovoltaics. In Proceedings of the 2019 IEEE 46th Photovoltaic Specialists Conference (PVSC); IEEE: Piscataway, NJ, USA, 2019; pp. 0592–0594. [Google Scholar]
- Shivakumar, A.; Dobbins, A.; Fahl, U.; Singh, A. Drivers of renewable energy deployment in the EU: An analysis of past trends and projections. Energy Strategy Rev. 2019, 26, 100402. [Google Scholar] [CrossRef] [Scilit]
- Sinsel, S.R.; Riemke, R.L.; Hoffmann, V.H. Challenges and solution technologies for the integration of variable renewable energy sources—a review. Renew. Energy 2020, 145, 2271–2285. [Google Scholar] [CrossRef] [Scilit]
- Prasad, M.B.; Ganesh, P.; Kumar, K.V.; Mohanarao, P.; Swathi, A.; Manoj, V. Renewable energy integration in modern power systems: Challenges and opportunities. E3S Web Conf. 2024, 591, 03002. [Google Scholar] [CrossRef] [Scilit]
- Erdiwansyah; Mahidin; Husin, H.; Nasaruddin; Zaki, M.; Muhibbuddin. A critical review of the integration of renewable energy sources with various technologies. Prot. Control Mod. Power Syst. 2021, 6, 3. [Google Scholar] [CrossRef] [Scilit]
- Bâra, A.; Oprea, S.v. Photovoltaic systems forecast using machine learning algorithms and recurrent neural networks. Sci. Bull. Nav. Acad. 2023, 26, 178–185. [Google Scholar]
- Kurucan, M.; Michailidis, P.; Michailidis, I.; Minelli, F. A Modular Hybrid SOC-Estimation Framework with a Supervisor for Battery Management Systems Supporting Renewable Energy Integration in Smart Buildings. Energies 2025, 18, 4537. [Google Scholar] [CrossRef] [Scilit]
- Kurucan, M.; Özbaltan, M.; Yetgin, Z.; Alkaya, A. Applications of artificial neural network based battery management systems: A literature review. Renew. Sustain. Energy Rev. 2024, 192, 114262. [Google Scholar] [CrossRef] [Scilit]
- Coban, H.H.; Lewicki, W. Flexibility in power systems of integrating variable renewable energy sources. J. Adv. Res. Nat. Appl. Sci. 2023, 9, 190–204. [Google Scholar] [CrossRef] [Scilit]
- Eltigani, D.; Masri, S. Challenges of integrating renewable energy sources to smart grids: A review. Renew. Sustain. Energy Rev. 2015, 52, 770–780. [Google Scholar] [CrossRef] [Scilit]
- Michailidis, P.; Minelli, F.; Michailidis, I.; Kurucan, M.; Coban, H.H.; Kosmatopoulos, E. Machine learning for energy management in buildings: A systematic review on real-world applications. Energies 2025, 19, 219. [Google Scholar] [CrossRef] [Scilit]
- Shafiullah, G.; Oo, A.M.; Jarvis, D.; Ali, A.S.; Wolfs, P. Potential challenges: Integrating renewable energy with the smart grid. In Proceedings of the 2010 20th Australasian Universities Power Engineering Conference; IEEE: Piscataway, NJ, USA, 2010; pp. 1–6. [Google Scholar]
- Eissa, M. Protection techniques with renewable resources and smart grids—A survey. Renew. Sustain. Energy Rev. 2015, 52, 1645–1667. [Google Scholar] [CrossRef] [Scilit]
- Shahzad, S.; Abbasi, M.A.; Ali, H.; Iqbal, M.; Munir, R.; Kilic, H. Possibilities, challenges, and future opportunities of microgrids: A review. Sustainability 2023, 15, 6366. [Google Scholar] [CrossRef] [Scilit]
- Iwuanyanwu, O.; Gil-Ozoudeh, I.; Okwandu, A.C.; Ike, C.S. The integration of renewable energy systems in green buildings: Challenges and opportunities. Int. J. Appl. Res. Soc. Sci. 2022, 4, 431–450. [Google Scholar] [CrossRef] [Scilit]
- Alfadil, M.O.; A. Kassem, M.; A. Al-Mansob, R. Renewable Energy Integration for Net-Zero Buildings: Challenges, Opportunities, and Strategic Pathways. Buildings 2026, 16, 879. [Google Scholar] [CrossRef] [Scilit]
- D’Agostino, D.; Minelli, F.; D’Urso, M.; Minichiello, F. Fixed and tracking PV systems for Net Zero Energy Buildings: Comparison between yearly and monthly energy balance. Renew. Energy 2022, 195, 809–824. [Google Scholar] [CrossRef] [Scilit]
- Kurucan, M.; Gencal, M.C.; Michailidis, P.; Minelli, F. Applying Symbolic Discrete Controller Synthesis Technique for Energy Management and Thermal Comfort Optimization in HVAC Systems. Sustainability 2026, 18, 2615. [Google Scholar] [CrossRef] [Scilit]
- Michailidis, P.; Michailidis, I.; Minelli, F.; Coban, H.H.; Kosmatopoulos, E. Model Predictive Control for Smart Buildings: Applications and Innovations in Energy Management. Buildings 2025, 15, 3298. [Google Scholar] [CrossRef] [Scilit]
- Michailidis, P.; Pelitaris, P.; Korkas, C.; Michailidis, I.; Baldi, S.; Kosmatopoulos, E. Enabling optimal energy management with minimal IoT requirements: A legacy A/C case study. Energies 2021, 14, 7910. [Google Scholar] [CrossRef] [Scilit]
- Coban, H.H.; Michailidis, P.; Yildirim, Y.A.; Minelli, F. Flattening Winter Peaks with Dynamic Energy Storage: A Neighborhood Case Study in the Cold Climate of Ardahan, Turkey. Sustainability 2026, 18, 761. [Google Scholar] [CrossRef] [Scilit]
- Ahsan, F.; Dana, N.H.; Sarker, S.K.; Li, L.; Muyeen, S.; Ali, M.F.; Tasneem, Z.; Hasan, M.M.; Abhi, S.H.; Islam, M.R.; et al. Data-driven next-generation smart grid towards sustainable energy evolution: Techniques and technology review. Prot. Control Mod. Power Syst. 2023, 8, 43. [Google Scholar] [CrossRef] [Scilit]
- Amasyali, K.; El-Gohary, N.M. A review of data-driven building energy consumption prediction studies. Renew. Sustain. Energy Rev. 2018, 81, 1192–1205. [Google Scholar] [CrossRef] [Scilit]
- Michailidis, P.; Michailidis, I.; Gkelios, S.; Kosmatopoulos, E. Artificial neural network applications for energy management in buildings: Current trends and future directions. Energies 2024, 17, 570. [Google Scholar] [CrossRef] [Scilit]
- Bâra, A.; Oprea, S.V. Machine learning algorithms for power system sign classification and a multivariate stacked LSTM model for predicting the electricity imbalance volume. Int. J. Comput. Intell. Syst. 2024, 17, 80. [Google Scholar] [CrossRef] [Scilit]
- Bâra, A.; Oprea, S.V. Large Language Models and Transformer Architecture in Electricity Price Forecasting. Ovidius Univ. Ann. Econ. Sci. Ser. 2025, 25, 126–135. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Gatsis, N.; Giannakis, G.B. Robust Energy Management for Microgrids With High-Penetration Renewables. IEEE Trans. Sustain. Energy 2013, 4, 944–953. [Google Scholar] [CrossRef] [Scilit]
- Yang, J.; Su, C. Robust optimization of microgrid based on renewable distributed power generation and load demand uncertainty. Energy 2021, 223, 120043. [Google Scholar] [CrossRef] [Scilit]
- Norouzi, M.; Aghaei, J.; Pirouzi, S.; Niknam, T.; Fotuhi-Firuzabad, M.; Shafie-khah, M. Hybrid stochastic/robust flexible and reliable scheduling of secure networked microgrids with electric springs and electric vehicles. Appl. Energy 2021, 300, 117395. [Google Scholar] [CrossRef] [Scilit]
- Mesbah, A. Stochastic Model Predictive Control: An Overview and Perspectives for Future Research. IEEE Control Syst. Mag. 2016, 36, 30–44. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Seal, S.; Wu, D.; Boulet, B.; Bouffard, F.; Joos, G. Data-Driven Model Predictive and Reinforcement Learning Based Control for Building Energy Management: A Survey. arXiv 2021, arXiv:2106.14450. [Google Scholar]
- Parisio, A.; Rikos, E.; Glielmo, L. A Model Predictive Control Approach to Microgrid Operation Optimization. IEEE Trans. Control Syst. Technol. 2014, 22, 1813–1827. [Google Scholar] [CrossRef] [Scilit]
- Morstyn, T.; Hredzak, B.; Aguilera, R.P.; Agelidis, V.G. Model Predictive Control for Distributed Microgrid Battery Energy Storage Systems. IEEE Trans. Control Syst. Technol. 2018, 26, 1107–1114. [Google Scholar] [CrossRef] [Scilit]
- Michailidis, P.; Michailidis, I.; Kosmatopoulos, E. Reinforcement learning for electric vehicle charging management: Theory and applications. Energies 2025, 18, 5225. [Google Scholar] [CrossRef] [Scilit]
- Gautam, M. Deep Reinforcement learning for resilient power and energy systems: Progress, prospects, and future avenues. Electricity 2023, 4, 336–380. [Google Scholar] [CrossRef] [Scilit]
- Michailidis, P.; Michailidis, I.; Kosmatopoulos, E. Reinforcement learning for optimizing renewable energy utilization in buildings: A review on applications and innovations. Energies 2025, 18, 1724. [Google Scholar] [CrossRef] [Scilit]
- François-Lavet, V.; Taralla, D.; Ernst, D.; Fonteneau, R. Deep Reinforcement Learning Solutions for Energy Microgrids Management. In Proceedings of the 12th European Workshop on Reinforcement Learning, Lille, France, 10–11 July 2015. [Google Scholar]
- Mocanu, E.; Mocanu, D.C.; Nguyen, P.H.; Liotta, A.; Webber, M.E.; Gibescu, M.; Slootweg, J.G. On-Line Building Energy Optimization Using Deep Reinforcement Learning. IEEE Trans. Smart Grid 2019, 10, 3698–3708. [Google Scholar] [CrossRef] [Scilit]
- Lazaridis, C.R.; Michailidis, I.; Karatzinis, G.; Michailidis, P.; Kosmatopoulos, E. Evaluating reinforcement learning algorithms in residential energy saving and comfort management. Energies 2024, 17, 581. [Google Scholar] [CrossRef] [Scilit]
- García, J.; Fernández, F. A Comprehensive Survey on Safe Reinforcement Learning. J. Mach. Learn. Res. 2015, 16, 1437–1480. [Google Scholar]
- Achiam, J.; Held, D.; Tamar, A.; Abbeel, P. Constrained Policy Optimization. In Proceedings of the 34th International Conference on Machine Learning; PMLR: New York, NY, USA, 2017; Volume 70, pp. 22–31. [Google Scholar]
- Banerjee, C.; Nguyen, K.; Fookes, C.; Raissi, M. A Survey on Physics Informed Reinforcement Learning: Review and Open Problems. Expert Syst. Appl. 2025, 287, 128166. [Google Scholar] [CrossRef] [Scilit]
- Yu, P.; Zhang, H.; Song, Y.; Wang, Z.; Dong, H.; Ji, L. Safe Reinforcement Learning for Power System Control: A Review. Renew. Sustain. Energy Rev. 2025, 223, 116022. [Google Scholar] [CrossRef] [Scilit]
- Villalva, M.; De Siqueira, T.; Ruppert, E. Voltage regulation of photovoltaic arrays: Small-signal analysis and control design. IET Power Electron. 2010, 3, 869–880. [Google Scholar] [CrossRef] [Scilit]
- Cui, W.; Jiang, Y.; Zhang, B. Reinforcement learning for optimal primary frequency control: A Lyapunov approach. IEEE Trans. Power Syst. 2022, 38, 1676–1688. [Google Scholar] [CrossRef] [Scilit]
- Adibi, M.; Van Der Woude, J. Secondary frequency control of microgrids: An online reinforcement learning approach. IEEE Trans. Autom. Control 2022, 67, 4824–4831. [Google Scholar] [CrossRef] [Scilit]
- Fu, Y.; Guo, X.; Mi, Y.; Yuan, M.; Ge, X.; Su, X.; Li, Z. The distributed economic dispatch of smart grid based on deep reinforcement learning. IET Gener. Transm. Distrib. 2021, 15, 2645–2658. [Google Scholar] [CrossRef] [Scilit]
- Rizki, A.; Touil, A.; Echchatbi, A.; Oucheikh, R. A reinforcement learning approach based on group relative policy optimization for economic dispatch in smart grids. Electricity 2025, 6, 49. [Google Scholar] [CrossRef] [Scilit]
- Singh, S.; Singh, S. Advancements and challenges in integrating renewable energy sources into distribution grid systems: A comprehensive review. J. Energy Resour. Technol. 2024, 146, 090801. [Google Scholar] [CrossRef] [Scilit]
- Coban, H.H.; Lewicki, W. Assessing the Efficiency of Hybrid Energy Facilities for Electric Vehicle Charging; Scientific Papers of Silesian University of Technology; Organization and Management Series; Silesian University of Technology Publishing House: Gliwice, Poland, 2023. [Google Scholar]
- Shang, Y.; Wu, W.; Guo, J.; Ma, Z.; Sheng, W.; Lv, Z.; Fu, C. Stochastic dispatch of energy storage in microgrids: An augmented reinforcement learning approach. Appl. Energy 2020, 261, 114423. [Google Scholar] [CrossRef] [Scilit]
- Phan, B.C.; Lai, Y.C. Control strategy of a hybrid renewable energy system based on reinforcement learning approach for an isolated microgrid. Appl. Sci. 2019, 9, 4001. [Google Scholar] [CrossRef] [Scilit]
- Bâra, A.; Oprea, S.V. Trading strategies on local electricity markets using Agent-Based modelling and Reinforcement learning: Vectors to expand energy Communities. Syst. Res. Behav. Sci. 2026, 43, 1188–1211. [Google Scholar] [CrossRef] [Scilit]
- Mbuwir, B.V.; Geysen, D.; Spiessens, F.; Deconinck, G. Reinforcement learning for control of flexibility providers in a residential microgrid. IET Smart Grid 2020, 3, 98–107. [Google Scholar]
- Hao, J.; Gao, D.W.; Zhang, J.J. Reinforcement learning for building energy optimization through controlling of central HVAC system. IEEE Open Access J. Power Energy 2020, 7, 320–328. [Google Scholar] [CrossRef] [Scilit]
- Cao, D.; Hu, W.; Zhao, J.; Zhang, G.; Zhang, B.; Liu, Z.; Chen, Z.; Blaabjerg, F. Reinforcement learning and its applications in modern power and energy systems: A review. J. Mod. Power Syst. Clean Energy 2020, 8, 1029–1042. [Google Scholar] [CrossRef] [Scilit]
- Erick, A.O.; Folly, K.A. Reinforcement learning approaches to power management in grid-tied microgrids: A review. In Proceedings of the 2020 Clemson University Power Systems Conference (PSC); IEEE: Piscataway, NJ, USA, 2020; pp. 1–6. [Google Scholar]
- Li, M.; Mour, N.; Smith, L. Machine learning based on reinforcement learning for smart grids: Predictive analytics in renewable energy management. Sustain. Cities Soc. 2024, 109, 105510. [Google Scholar] [CrossRef] [Scilit]
- Smart, E.E.; Olanrewaju, L.O.; Usman, J.; Otaru, K.; Umar, D. Artificial Intelligence (AI) in renewable energy forecasting and optimization. Renew. Energy 2025, 10, 11. [Google Scholar]
- Vamvakas, D.; Michailidis, P.; Korkas, C.; Kosmatopoulos, E. Review and evaluation of reinforcement learning frameworks on smart grid applications. Energies 2023, 16, 5326. [Google Scholar] [CrossRef] [Scilit]
- Michailidis, P.; Michailidis, I.; Vamvakas, D.; Kosmatopoulos, E. Model-free HVAC control in buildings: A review. Energies 2023, 16, 7124. [Google Scholar] [CrossRef] [Scilit]
- Michailidis, P.; Michailidis, I.; Kosmatopoulos, E. Review and evaluation of multi-agent control applications for energy management in buildings. Energies 2024, 17, 4835. [Google Scholar] [CrossRef] [Scilit]
- Salkuti, S.R.; Ray, P.; Pagidipala, S. Overview of next generation smart grids. In Next Generation Smart Grids: Modeling, Control and Optimization; Springer: Singapore, 2022; pp. 1–28. [Google Scholar]
- Butt, O.M.; Zulqarnain, M.; Butt, T.M. Recent advancement in smart grid technology: Future prospects in the electrical power network. Ain Shams Eng. J. 2021, 12, 687–695. [Google Scholar] [CrossRef] [Scilit]
- Moreno Escobar, J.J.; Morales Matamoros, O.; Tejeida Padilla, R.; Lina Reyes, I.; Quintana Espinosa, H. A comprehensive review on smart grids: Challenges and opportunities. Sensors 2021, 21, 6978. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nguyen, V.T.; Truong, T.B.T.; Nguyen, H.V.P.; Nguyen, B.N.; Truong, Q.V.; Nguyen, T.Q.; Le, M.T. Optimizing Wind Turbine Performance Using Fuzzy Logic-Controlled MPPT Algorithms. In Proceedings of the 2025 Asia Meeting on Environment and Electrical Engineering (EEE-AM); IEEE: Piscataway, NJ, USA, 2025; pp. 1–6. [Google Scholar]
- Ourahou, M.; Ayrir, W.; Hassouni, B.E.; Haddi, A. Review on smart grid control and reliability in presence of renewable energies: Challenges and prospects. Math. Comput. Simul. 2020, 167, 19–31. [Google Scholar] [CrossRef] [Scilit]
- Ali, S.S.; Choi, B.J. State-of-the-art artificial intelligence techniques for distributed smart grids: A review. Electronics 2020, 9, 1030. [Google Scholar] [CrossRef] [Scilit]
- Nam, N.B.; Ogliari, E.; Leva, S.; Pafumi, E.; Alberti, D.; Duong, M.Q. Comparative analysis of conformal prediction techniques and machine learning models for very short-term solar power forecasting. Energy AI 2025, 21, 100573. [Google Scholar] [CrossRef] [Scilit]
- Kakran, S.; Chanana, S. Smart operations of smart grids integrated with distributed generation: A review. Renew. Sustain. Energy Rev. 2018, 81, 524–535. [Google Scholar] [CrossRef] [Scilit]
- Abdullah, A.A.; Hassan, T.M. Smart grid (SG) properties and challenges: An overview. Discov. Energy 2022, 2, 8. [Google Scholar] [CrossRef] [Scilit]
- Salkuti, S.R. Challenges, issues and opportunities for the development of smart grid. Int. J. Electr. Comput. Eng. (IJECE) 2020, 10, 1179–1186. [Google Scholar] [CrossRef] [Scilit]
- Phuangpornpitak, N.; Tia, S. Opportunities and challenges of integrating renewable energy in smart grid system. Energy Procedia 2013, 34, 282–290. [Google Scholar] [CrossRef] [Scilit]
- Bose, B.K. Power electronics, smart grid, and renewable energy systems. Proc. IEEE 2017, 105, 2011–2018. [Google Scholar] [CrossRef] [Scilit]
- Minelli, F.; D’Agostino, D.; Reddy, V.J.; Michailidis, P. Towards Resilient Cities: Robust Selection of Rooftop Renewable Energy Technologies in Mediterranean Multifamily Buildings. Energy Eng. J. Assoc. Energy Eng. 2026, 123, 20. [Google Scholar] [CrossRef] [Scilit]
- Varma, R.K.; Salama, M. Large-scale photovoltaic solar power integration in transmission and distribution networks. In Proceedings of the 2011 IEEE Power and Energy Society General Meeting; IEEE: Piscataway, NJ, USA, 2011; pp. 1–4. [Google Scholar]
- Shapsough, S.; Takrouri, M.; Dhaouadi, R.; Zualkernan, I.A. Using IoT and smart monitoring devices to optimize the efficiency of large-scale distributed solar farms. Wirel. Netw. 2021, 27, 4313–4329. [Google Scholar]
- Meena, M.R.S.; Rathore, M.J.S.; Johri, M.S. Grid connected roof top solar power generation: A review. Int. J. Eng. Dev. Res. 2015, 3, 325–330. [Google Scholar]
- Yan, R.; Saha, T.K. Voltage variation sensitivity analysis for unbalanced distribution networks due to photovoltaic power fluctuations. IEEE Trans. Power Syst. 2012, 27, 1078–1089. [Google Scholar] [CrossRef] [Scilit]
- Minhas, D.M.; Khalid, R.R.; Frey, G. Real-time power balancing in photovoltaic-integrated smart micro-grid. In Proceedings of the IECON 2017—43rd Annual Conference of the IEEE Industrial Electronics Society; IEEE: Piscataway, NJ, USA, 2017; pp. 7469–7474. [Google Scholar]
- Hansen, M. Aerodynamics of Wind Turbines; Routledge: Abingdon, UK, 2015. [Google Scholar]
- Berg, D.E. Wind energy conversion. In Energy Conversion; CRC Press: Boca Raton, FL, USA, 2017; pp. 851–895. [Google Scholar]
- Tavner, P.; Edwards, C.; Brinkman, A.; Spinato, F. Influence of wind speed on wind turbine reliability. Wind Eng. 2006, 30, 55–72. [Google Scholar] [CrossRef] [Scilit]
- Boyle, J.; Littler, T.; Foley, A. Review of frequency stability services for grid balancing with wind generation. J. Eng. 2018, 2018, 1061–1065. [Google Scholar] [CrossRef] [Scilit]
- Feilat, E.; Azzam, S.; Al-Salaymeh, A. Impact of large PV and wind power plants on voltage and frequency stability of Jordan’s national grid. Sustain. Cities Soc. 2018, 36, 257–271. [Google Scholar] [CrossRef] [Scilit]
- Alidrisi, H.; Demirbas, A. Enhanced electricity generation using biomass materials. Energy Sources Part A Recovery Util. Environ. Eff. 2016, 38, 1419–1427. [Google Scholar] [CrossRef] [Scilit]
- Haq, Z. Biomass for Electricity Generation; Energy Information Administration: Washington, DC, USA, 2002.
- Avwioroko, A.; Ibegbulam, C.; Afriyie, I.; Fesomade, A.T. Smart grid integration of solar and biomass energy sources. Eur. J. Comput. Sci. Inf. Technol. 2024, 12, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Perez-Navarro, A.; Alfonso, D.; Álvarez, C.; Ibáñez, F.; Sanchez, C.; Segura, I. Hybrid biomass-wind power plant for reliable energy generation. Renew. Energy 2010, 35, 1436–1443. [Google Scholar] [CrossRef] [Scilit]
- Yang, L.; Entchev, E.; Rosato, A.; Sibilio, S. Smart thermal grid with integration of distributed and centralized solar energy systems. Energy 2017, 122, 471–481. [Google Scholar] [CrossRef] [Scilit]
- Yousefzadeh, M.; Lenzen, M. Performance of concentrating solar power plants in a whole-of-grid context. Renew. Sustain. Energy Rev. 2019, 114, 109342. [Google Scholar] [CrossRef] [Scilit]
- Ogunmodimu, O.; Okoroigwe, E.C. Solar thermal electricity in Nigeria: Prospects and challenges. Energy Policy 2019, 128, 440–448. [Google Scholar] [CrossRef] [Scilit]
- Navarro, A.A.; Ramírez, L.; Domínguez, P.; Blanco, M.; Polo, J.; Zarza, E. Review and validation of Solar Thermal Electricity potential methodologies. Energy Convers. Manag. 2016, 126, 42–50. [Google Scholar] [CrossRef] [Scilit]
- Kabeyi, M.J.B.; Olanrewaju, O.A. Geothermal wellhead technology power plants in grid electricity generation: A review. Energy Strategy Rev. 2022, 39, 100735. [Google Scholar] [CrossRef] [Scilit]
- Acharjya, K.; Jain, N.K.; A, P.; Saraswat, B. Exploring Geothermal Energy’s Potential in Smart Cities Building Climate Control. E3S Web Conf. 2024, 540, 13007. [Google Scholar] [CrossRef] [Scilit]
- Nasruddin; Alhamid, M.I.; Daud, Y.; Surachman, A.; Sugiyono, A.; Aditya, H.; Mahlia, T. Potential of geothermal energy for electricity generation in Indonesia: A review. Renew. Sustain. Energy Rev. 2016, 53, 733–740. [Google Scholar] [CrossRef] [Scilit]
- Mekonnen, M.; Hoekstra, A.Y. The Water Footprint of Electricity from Hydropower; UNESCO-IHE Institute for Water Education: Delft, The Netherlands, 2011. [Google Scholar]
- Sternberg, R. Hydropower’s future, the environment, and global electricity systems. Renew. Sustain. Energy Rev. 2010, 14, 713–723. [Google Scholar] [CrossRef] [Scilit]
- Forouzandehmehr, N.; Han, Z.; Zheng, R. Stochastic dynamic game between hydropower plant and thermal power plant in smart grid networks. IEEE Syst. J. 2014, 10, 88–96. [Google Scholar] [CrossRef] [Scilit]
- De Silva, T.; Jorgenson, J.; Macknick, J.; Keohan, N.; Miara, A.; Jager, H.; Pracheil, B. Hydropower operation in future power grid with various renewable power integration. Renew. Energy Focus 2022, 43, 329–339. [Google Scholar] [CrossRef] [Scilit]
- Bessa, R.; Moreira, C.; Silva, B.; Matos, M. Handling renewable energy variability and uncertainty in power system operation. In Advances in Energy Systems: The Large-Scale Renewable Energy Integration Challenge; John Wiley & Sons Ltd.: Hoboken, NJ, USA, 2019; pp. 1–26. [Google Scholar]
- Meliani, M.; Barkany, A.E.; Abbassi, I.E.; Darcherif, A.M.; Mahmoudi, M. Energy management in the smart grid: State-of-the-art and future trends. Int. J. Eng. Bus. Manag. 2021, 13, 18479790211032920. [Google Scholar] [CrossRef] [Scilit]
- Keyhani, A. Smart power grids. In Smart Power Grids 2011; Springer: Berlin/Heidelberg, Germany, 2011; pp. 1–25. [Google Scholar]
- Eluri, H.B.; Naik, M.G. Challenges of res with integration of power grids, control strategies, optimization techniques of microgrids: A review. Int. J. Renew. Energy Res. (IJRER) 2021, 11, 1–19. [Google Scholar] [CrossRef] [Scilit]
- Lewicki, W.; Coban, H.H.; Minelli, F.; Michailidis, P. Freezers in Residential Buildings as a Source of Power Grid Frequency Regulation in Response to the Demand for Innovation Within the Smart City Concept: Thermal–Electric Modeling, Technical Potential and Operational Challenges. Energies 2026, 19, 1608. [Google Scholar] [CrossRef] [Scilit]
- Gandoman, F.H.; Ahmadi, A.; Sharaf, A.M.; Siano, P.; Pou, J.; Hredzak, B.; Agelidis, V.G. Review of FACTS technologies and applications for power quality in smart grids with renewable energy systems. Renew. Sustain. Energy Rev. 2018, 82, 502–514. [Google Scholar] [CrossRef] [Scilit]
- Gajduk, A.; Todorovski, M.; Kocarev, L. Stability of power grids: An overview. Eur. Phys. J. Spec. Top. 2014, 223, 2387–2409. [Google Scholar] [CrossRef] [Scilit]
- Alizadeh, M.I.; Moghaddam, M.P.; Amjady, N.; Siano, P.; Sheikh-El-Eslami, M.K. Flexibility in future power systems with high renewable penetration: A review. Renew. Sustain. Energy Rev. 2016, 57, 1186–1193. [Google Scholar] [CrossRef] [Scilit]
- García Vera, Y.E.; Dufo-López, R.; Bernal-Agustín, J.L. Energy management in microgrids with renewable energy sources: A literature review. Appl. Sci. 2019, 9, 3854. [Google Scholar] [CrossRef] [Scilit]
- Islam, M.; Yang, F.; Amin, M. Control and optimisation of networked microgrids: A review. IET Renew. Power Gener. 2021, 15, 1133–1148. [Google Scholar] [CrossRef] [Scilit]
- Anderson, A.A.; Suryanarayanan, S. Review of energy management and planning of islanded microgrids. CSEE J. Power Energy Syst. 2019, 6, 329–343. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Su, B. Microgrids—an integration of renewable energy technologies. In Proceedings of the 2008 China International Conference on Electricity Distribution; IEEE: Piscataway, NJ, USA, 2008; pp. 1–7. [Google Scholar]
- Rezaie, B.; Esmailzadeh, E.; Dincer, I. Renewable energy options for buildings: Case studies. Energy Build. 2011, 43, 56–65. [Google Scholar] [CrossRef] [Scilit]
- Kaya, G.N.; Beyhan, F. A Study on the Potential of Photovoltaic Panels in Existing Buildings: Housing Example in the Mediterranean. Online J. Art Des. 2024, 12, 73–87. [Google Scholar] [CrossRef] [Scilit]
- Kaya, G.; BEYHAN, F. Integrated design approaches with photovoltaic panel and solar collectors in building envelope. World J. Environ. Res. 2021, 11, 49–61. [Google Scholar] [CrossRef] [Scilit]
- Manic, M.; Wijayasekara, D.; Amarasinghe, K.; Rodriguez-Andina, J.J. Building energy management systems: The age of intelligent and adaptive buildings. IEEE Ind. Electron. Mag. 2016, 10, 25–39. [Google Scholar] [CrossRef] [Scilit]
- Talari, S.; Shafie-Khah, M.; Osório, G.J.; Aghaei, J.; Catalão, J.P. Stochastic modelling of renewable energy sources from operators’ point-of-view: A survey. Renew. Sustain. Energy Rev. 2018, 81, 1953–1965. [Google Scholar] [CrossRef] [Scilit]
- Ge, L.; Li, J.; Hou, L.; Lai, J. Autonomous Voltage Regulation for Smart Distribution Network with High-Proportion PVs: A Graph Meta-Reinforcement Learning Approach. IEEE Trans. Sustain. Energy 2025, 16, 2768–2781. [Google Scholar] [CrossRef] [Scilit]
- Gao, W.; Fan, R.; Qiao, W.; Wang, S.; Gao, D.W. Deep Reinforcement Learning Based Control of Wind Turbines for Fast Frequency Response. IEEE Trans. Ind. Appl. 2025, 61, 8640–8649. [Google Scholar] [CrossRef] [Scilit]
- Jiang, Y.; Cui, W.; Zhang, B.; Cortés, J. Stable reinforcement learning for optimal frequency control: A distributed averaging-based integral approach. IEEE Open J. Control Syst. 2022, 1, 194–209. [Google Scholar] [CrossRef] [Scilit]
- Bosisio, A.; Soldan, F.; Pisani, M.; Bionda, E.; Belloni, F.; Morotti, A. A Q-Learning algorithm for optimizing On-Load tap changer operation and voltage control in distribution networks with high integration of renewable energy sources. J. Mod. Power Syst. Clean Energy 2025, 13, 2063–2073. [Google Scholar]
- Vlachogiannis, J.G.; Hatziargyriou, N.D. Reinforcement learning for reactive power control. IEEE Trans. Power Syst. 2004, 19, 1317–1325. [Google Scholar] [CrossRef] [Scilit]
- Rocchetta, R.; Bellani, L.; Compare, M.; Zio, E.; Patelli, E. A reinforcement learning framework for optimal operation and maintenance of power grids. Appl. Energy 2019, 241, 291–301. [Google Scholar] [CrossRef] [Scilit]
- Mussi, M.; Pellegrino, L.; Pindaro, O.F.; Restelli, M.; Trovò, F. A Reinforcement Learning controller optimizing costs and battery State of Health in smart grids. J. Energy Storage 2024, 82, 110572. [Google Scholar] [CrossRef] [Scilit]
- Kurucan, M. Accurate Frequency Control of Energy Storage Systems with a Symbolic Game Theory Approach. Çukurova Üniversitesi Mühendislik Fakültesi Derg. 2026, 41, 195–212. [Google Scholar] [CrossRef] [Scilit]
- Kuznetsova, E.; Li, Y.F.; Ruiz, C.; Zio, E.; Ault, G.; Bell, K. Reinforcement learning for microgrid energy management. Energy 2013, 59, 133–146. [Google Scholar] [CrossRef] [Scilit]
- Foruzan, E.; Soh, L.K.; Asgarpoor, S. Reinforcement learning approach for optimal distributed energy management in a microgrid. IEEE Trans. Power Syst. 2018, 33, 5749–5758. [Google Scholar] [CrossRef] [Scilit]
- Alshahr, S.; Alshahir, A.; Alnuman, H.; Alanazi, M.D.; Yousef, A.; Abbas, G. Dynamic renewable energy integration for EV charging via model-based reinforcement learning. Ain Shams Eng. J. 2026, 17, 104040. [Google Scholar] [CrossRef] [Scilit]
- Bashyal, A.; Alnahas, H.; Boroukhian, T.; Wicaksono, H. Demand response based industrial energy management with focus on consumption of renewable energy: A deep reinforcement learning approach. Procedia Comput. Sci. 2025, 253, 1442–1451. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.; Sun, Z.; Meng, Q.; Wang, Y.; Li, Y. Reinforcement learning of room temperature set-point of thermal storage air-conditioning system with demand response. Energy Build. 2022, 259, 111903. [Google Scholar] [CrossRef] [Scilit]
- Puterman, M.L. Markov Decision Processes: Discrete Stochastic Dynamic Programming; John Wiley & Sons: Hoboken, NJ, USA, 2014. [Google Scholar]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction; MIT Press: Cambridge, MA, USA, 1998; Volume 1. [Google Scholar]
- Zhang, K.; Yang, Z.; Başar, T. Multi-agent reinforcement learning: A selective overview of theories and algorithms. In Handbook of Reinforcement Learning and Control; Springer: Cham, Switzerland, 2021; pp. 321–384. [Google Scholar]
- Wang, J.; Xu, W.; Gu, Y.; Song, W.; Green, T.C. Multi-Agent Reinforcement Learning for Active Voltage Control on Power Distribution Networks. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34. [Google Scholar]
- Nekoei, H.; Badrinaaraayanan, A.; Sinha, A.; Amini, M.; Rajendran, J.; Mahajan, A.; Chandar, S. Dealing with Non-Stationarity in Decentralized Cooperative Multi-Agent Deep Reinforcement Learning via Multi-Timescale Learning. In Proceedings of the 2nd Conference on Lifelong Learning Agents; PMLR: New York, NY, USA, 2023. [Google Scholar]
- Liu, Z.; Zhang, J.; Shi, E.; Liu, Z.; Niyato, D.; Ai, B.; Shen, X. Graph Neural Network Meets Multi-Agent Reinforcement Learning: Fundamentals, Applications, and Future Directions. IEEE Wirel. Commun. 2024, 31, 39–47. [Google Scholar] [CrossRef] [Scilit]
- Yang, Y.; Luo, R.; Li, M.; Zhou, M.; Zhang, W.; Wang, J. Mean Field Multi-Agent Reinforcement Learning. In Proceedings of the 35th International Conference on Machine Learning; PMLR: New York, NY, USA, 2018; Volume 80, pp. 5571–5580. [Google Scholar]
- Byeon, H. Advances in value-based, policy-based, and deep learning-based reinforcement learning. Int. J. Adv. Comput. Sci. Appl. 2023, 14, 348–354. [Google Scholar] [CrossRef] [Scilit]
- Watkins, C.J.; Dayan, P. Q-learning. Mach. Learn. 1992, 8, 279–292. [Google Scholar] [CrossRef] [Scilit]
- Fan, J.; Wang, Z.; Xie, Y.; Yang, Z. A theoretical analysis of deep Q-learning. In Proceedings of the Learning for Dynamics and Control; PMLR: New York, NY, USA, 2020; pp. 486–489. [Google Scholar]
- Alabdullah, M.H.; Abido, M.A. Microgrid energy management using deep Q-network reinforcement learning. Alex. Eng. J. 2022, 61, 9069–9078. [Google Scholar] [CrossRef] [Scilit]
- Sutton, R.S.; McAllester, D.; Singh, S.; Mansour, Y. Policy gradient methods for reinforcement learning with function approximation. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 1999; Volume 12. [Google Scholar]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
- Ioannou, I.; Javaid, S.; Tan, Y.; Vassiliou, V. Autonomous reinforcement learning for intelligent and sustainable autonomous microgrid energy management. Electronics 2025, 14, 2691. [Google Scholar] [CrossRef] [Scilit]
- Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep reinforcement learning. arXiv 2015, arXiv:1509.02971. [Google Scholar]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the International Conference on Machine Learning; PMLR: New York, NY, USA, 2018; pp. 1861–1870. [Google Scholar]
- Fujimoto, S.; van Hoof, H.; Meger, D. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the P35th International Conference on Machine Learning; PMLR: New York, NY, USA, 2018; pp. 1587–1596. [Google Scholar]
- Bellemare, M.G.; Dabney, W.; Munos, R. A Distributional Perspective on Reinforcement Learning. In Proceedings of the 34th International Conference on Machine Learning; PMLR: New York, NY, USA, 2017; Volume 70, pp. 449–458. [Google Scholar]
- Dabney, W.; Rowland, M.; Bellemare, M.G.; Munos, R. Distributional Reinforcement Learning with Quantile Regression. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Washington, DC, USA, 2018; Volume 32, pp. 2892–2901. [Google Scholar]
- Hessel, M.; Modayil, J.; van Hasselt, H.; Schaul, T.; Ostrovski, G.; Dabney, W.; Horgan, D.; Piot, B.; Azar, M.; Silver, D. Rainbow: Combining Improvements in Deep Reinforcement Learning. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Washington, DC, USA, 2018; Volume 32, pp. 3215–3222. [Google Scholar]
- Khadka, S.; Tumer, K. Evolution-Guided Policy Gradient in Reinforcement Learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2018; Volume 31. [Google Scholar]
- Hansen, N.; Ostermeier, A. Completely Derandomized Self-Adaptation in Evolution Strategies. Evol. Comput. 2001, 9, 159–195. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, Y.; Zhang, B.; Xu, C.; Lan, T.; Diao, R.; Shi, D.; Wang, Z.; Lee, W.J. A data-driven method for fast ac optimal power flow solutions via deep reinforcement learning. J. Mod. Power Syst. Clean Energy 2020, 8, 1128–1139. [Google Scholar] [CrossRef] [Scilit]
- Cao, D.; Hu, W.; Zhao, J.; Huang, Q.; Chen, Z.; Blaabjerg, F. A multi-agent deep reinforcement learning based voltage regulation using coordinated PV inverters. IEEE Trans. Power Syst. 2020, 35, 4120–4123. [Google Scholar] [CrossRef] [Scilit]
- Kou, P.; Liang, D.; Wang, C.; Wu, Z.; Gao, L. Safe deep reinforcement learning-based constrained optimal control scheme for active distribution networks. Appl. Energy 2020, 264, 114772. [Google Scholar] [CrossRef] [Scilit]
- Yang, Q.; Wang, G.; Sadeghi, A.; Giannakis, G.B.; Sun, J. Two-timescale voltage control in distribution grids using deep reinforcement learning. IEEE Trans. Smart Grid 2019, 11, 2313–2323. [Google Scholar] [CrossRef] [Scilit]
- Rozada, S.; Apostolopoulou, D.; Alonso, E. Load frequency control: A deep multi-agent reinforcement learning approach. In Proceedings of the 2020 IEEE Power & Energy Society General Meeting (PESGM); IEEE: Piscataway, NJ, USA, 2020; pp. 1–5. [Google Scholar]
- Cao, D.; Zhao, J.; Hu, W.; Ding, F.; Huang, Q.; Chen, Z.; Blaabjerg, F. Data-driven multi-agent deep reinforcement learning for distribution system decentralized voltage control with high penetration of PVs. IEEE Trans. Smart Grid 2021, 12, 4137–4150. [Google Scholar] [CrossRef] [Scilit]
- Cao, D.; Zhao, J.; Hu, W.; Ding, F.; Huang, Q.; Chen, Z. Attention enabled multi-agent DRL for decentralized volt-VAR control of active distribution system using PV inverters and SVCs. IEEE Trans. Sustain. Energy 2021, 12, 1582–1592. [Google Scholar] [CrossRef] [Scilit]
- Cao, D.; Zhao, J.; Hu, W.; Yu, N.; Ding, F.; Huang, Q.; Chen, Z. Deep reinforcement learning enabled physical-model-free two-timescale voltage control method for active distribution systems. IEEE Trans. Smart Grid 2021, 13, 149–165. [Google Scholar]
- Zhang, X.; Liu, Y.; Duan, J.; Qiu, G.; Liu, T.; Liu, J. DDPG-based multi-agent framework for SVC tuning in urban power grid with renewable energy resources. IEEE Trans. Power Syst. 2021, 36, 5465–5475. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; He, H. Learning to operate distribution networks with safe deep reinforcement learning. IEEE Trans. Smart Grid 2022, 13, 1860–1872. [Google Scholar] [CrossRef] [Scilit]
- Bui, V.H.; Su, W. Real-time operation of distribution network: A deep reinforcement learning-based reconfiguration approach. Sustain. Energy Technol. Assess. 2022, 50, 101841. [Google Scholar] [CrossRef] [Scilit]
- Hu, D.; Ye, Z.; Gao, Y.; Ye, Z.; Peng, Y.; Yu, N. Multi-agent deep reinforcement learning for voltage control with coordinated active and reactive power optimization. IEEE Trans. Smart Grid 2022, 13, 4873–4886. [Google Scholar] [CrossRef] [Scilit]
- Cui, W.; Li, J.; Zhang, B. Decentralized safe reinforcement learning for inverter-based voltage control. Electr. Power Syst. Res. 2022, 211, 108609. [Google Scholar] [CrossRef] [Scilit]
- Khalid, J.; Ramli, M.A.; Khan, M.S.; Hidayat, T. Efficient load frequency control of renewable integrated power system: A twin delayed DDPG-based deep reinforcement learning approach. IEEE Access 2022, 10, 51561–51574. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Yu, T.; Zhang, X. Coordinated automatic generation control of interconnected power system with imitation guided exploration multi-agent deep reinforcement learning. Int. J. Electr. Power Energy Syst. 2022, 136, 107471. [Google Scholar] [CrossRef] [Scilit]
- Li, P.; Wei, M.; Ji, H.; Xi, W.; Yu, H.; Wu, J.; Yao, H.; Chen, J. Deep reinforcement learning-based adaptive voltage control of active distribution networks with multi-terminal soft open point. Int. J. Electr. Power Energy Syst. 2022, 141, 108138. [Google Scholar] [CrossRef] [Scilit]
- Cao, D.; Zhao, J.; Hu, W.; Ding, F.; Yu, N.; Huang, Q.; Chen, Z. Model-free voltage control of active distribution system with PVs using surrogate model-based deep reinforcement learning. Appl. Energy 2022, 306, 117982. [Google Scholar] [CrossRef] [Scilit]
- Wu, Z.; Li, Y.; Gu, W.; Dong, Z.; Zhao, J.; Liu, W.; Zhang, X.P.; Liu, P.; Sun, Q. Multi-timescale voltage control for distribution system based on multi-agent deep reinforcement learning. Int. J. Electr. Power Energy Syst. 2023, 147, 108830. [Google Scholar] [CrossRef] [Scilit]
- Li, P.; Shen, J.; Wu, Z.; Yin, M.; Dong, Y.; Han, J. Optimal real-time Voltage/Var control for distribution network: Droop-control based multi-agent deep reinforcement learning. Int. J. Electr. Power Energy Syst. 2023, 153, 109370. [Google Scholar] [CrossRef] [Scilit]
- Petrusev, A.; Putratama, M.A.; Rigo-Mariani, R.; Debusschere, V.; Reignier, P.; Hadjsaid, N. Reinforcement learning for robust voltage control in distribution grids under uncertainties. Sustain. Energy Grids Netw. 2023, 33, 100959. [Google Scholar] [CrossRef] [Scilit]
- Rehman, A.U.; Ullah, Z.; Qazi, H.S.; Hasanien, H.M.; Khalid, H.M. Reinforcement learning-driven proximal policy optimization-based voltage control for PV and WT integrated power system. Renew. Energy 2024, 227, 120590. [Google Scholar] [CrossRef] [Scilit]
- Huang, J.; Zhang, H.; Tian, D.; Zhang, Z.; Yu, C.; Hancke, G.P. Multi-agent deep reinforcement learning with enhanced collaboration for distribution network voltage control. Eng. Appl. Artif. Intell. 2024, 134, 108677. [Google Scholar] [CrossRef] [Scilit]
- Zhang, B.; Cao, D.; Hu, W.; Ghias, A.M.; Chen, Z. Physics-Informed Multi-Agent deep reinforcement learning enabled distributed voltage control for active distribution network using PV inverters. Int. J. Electr. Power Energy Syst. 2024, 155, 109641. [Google Scholar] [CrossRef] [Scilit]
- Jacob, R.A.; Paul, S.; Chowdhury, S.; Gel, Y.R.; Zhang, J. Real-time outage management in active distribution networks using reinforcement learning over graphs. Nat. Commun. 2024, 15, 4766. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Singh, A.R.; Sujatha, M.; Kadu, A.D.; Bajaj, M.; Addis, H.K.; Sarada, K. A deep learning and IoT-driven framework for real-time adaptive resource allocation and grid optimization in smart energy systems. Sci. Rep. 2025, 15, 19309. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mamodiya, U.; Kishor, I.; Garine, R.; Ganguly, P.; Naik, N. Artificial intelligence based hybrid solar energy systems with smart materials and adaptive photovoltaics for sustainable power generation. Sci. Rep. 2025, 15, 17370. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tang, X.; Wang, J. Deep reinforcement learning-based multi-objective optimization for virtual power plants and smart grids: Maximizing renewable energy integration and grid efficiency. Processes 2025, 13, 1809. [Google Scholar] [CrossRef] [Scilit]
- Luo, F.; Wang, S.; Lv, Y.; Mu, R.; Fo, J.; Zhang, T.; Xu, J.; Wang, C. Domain knowledge-enhanced graph reinforcement learning method for Volt/Var control in distribution networks. Appl. Energy 2025, 398, 126409. [Google Scholar] [CrossRef] [Scilit]
- Faghri, S.; Tahami, H.; Amini, R.; Katiraee, H.; Langeroudi, A.S.G.; Alinejad, M.; Nejati, M.G. Real-time energy flexibility optimization of grid-connected smart building communities with deep reinforcement learning. Sustain. Cities Soc. 2025, 119, 106077. [Google Scholar] [CrossRef] [Scilit]
- Nyong-Bassey, B.E.; Giaouris, D.; Patsios, C.; Papadopoulou, S.; Papadopoulos, A.I.; Walker, S.; Voutetakis, S.; Seferlis, P.; Gadoue, S. Reinforcement learning based adaptive power pinch analysis for energy management of stand-alone hybrid energy storage systems considering uncertainty. Energy 2020, 193, 116622. [Google Scholar] [CrossRef] [Scilit]
- Samadi, E.; Badri, A.; Ebrahimpour, R. Decentralized multi-agent based energy management of microgrid using reinforcement learning. Int. J. Electr. Power Energy Syst. 2020, 122, 106211. [Google Scholar] [CrossRef] [Scilit]
- Lei, L.; Tan, Y.; Dahlenburg, G.; Xiang, W.; Zheng, K. Dynamic energy dispatch based on deep reinforcement learning in IoT-driven smart isolated microgrids. IEEE Internet Things J. 2020, 8, 7938–7953. [Google Scholar]
- Shuai, H.; He, H. Online scheduling of a residential microgrid via Monte-Carlo tree search and a learned model. IEEE Trans. Smart Grid 2020, 12, 1073–1087. [Google Scholar] [CrossRef] [Scilit]
- Chen, D.; Chen, K.; Li, Z.; Chu, T.; Yao, R.; Qiu, F.; Lin, K. Powernet: Multi-agent deep reinforcement learning for scalable powergrid control. IEEE Trans. Power Syst. 2021, 37, 1007–1017. [Google Scholar]
- Li, Y.; Wang, R.; Yang, Z. Optimal scheduling of isolated microgrids using automated reinforcement learning-based multi-period forecasting. IEEE Trans. Sustain. Energy 2021, 13, 159–169. [Google Scholar]
- Munir, M.S.; Abedin, S.F.; Tran, N.H.; Han, Z.; Huh, E.N.; Hong, C.S. Risk-aware energy scheduling for edge computing with microgrid: A multi-agent deep reinforcement learning approach. IEEE Trans. Netw. Serv. Manag. 2021, 18, 3476–3497. [Google Scholar] [CrossRef] [Scilit]
- Tomin, N.; Shakirov, V.; Kozlov, A.; Sidorov, D.; Kurbatsky, V.; Rehtanz, C.; Lora, E.E. Design and optimal energy management of community microgrids with flexible renewable energy sources. Renew. Energy 2022, 183, 903–921. [Google Scholar] [CrossRef] [Scilit]
- Barbalho, P.I.d.N.; Lacerda, V.; Fernandes, R.; Coury, D.V. Deep reinforcement learning-based secondary control for microgrids in islanded mode. Electr. Power Syst. Res. 2022, 212, 108315. [Google Scholar] [CrossRef] [Scilit]
- Guo, C.; Wang, X.; Zheng, Y.; Zhang, F. Real-time optimal energy management of microgrid with uncertainties based on deep reinforcement learning. Energy 2022, 238, 121873. [Google Scholar] [CrossRef] [Scilit]
- Harrold, D.J.; Cao, J.; Fan, Z. Data-driven battery operation for energy arbitrage using rainbow deep reinforcement learning. Energy 2022, 238, 121958. [Google Scholar] [CrossRef] [Scilit]
- Harrold, D.J.; Cao, J.; Fan, Z. Renewable energy integration and microgrid energy trading using multi-agent deep reinforcement learning. Appl. Energy 2022, 318, 119151. [Google Scholar] [CrossRef] [Scilit]
- Hu, C.; Cai, Z.; Zhang, Y.; Yan, R.; Cai, Y.; Cen, B. A soft actor-critic deep reinforcement learning method for multi-timescale coordinated operation of microgrids. Prot. Control Mod. Power Syst. 2022, 7, 29. [Google Scholar] [CrossRef] [Scilit]
- Xia, Y.; Xu, Y.; Wang, Y.; Mondal, S.; Dasgupta, S.; Gupta, A.K.; Gupta, G.M. A safe policy learning-based method for decentralized and economic frequency control in isolated networked-microgrid systems. IEEE Trans. Sustain. Energy 2022, 13, 1982–1993. [Google Scholar] [CrossRef] [Scilit]
- Xiong, L.; Tang, Y.; Mao, S.; Liu, H.; Meng, K.; Dong, Z.; Qian, F. A two-level energy management strategy for multi-microgrid systems with interval prediction and reinforcement learning. IEEE Trans. Circuits Syst. I Regul. Pap. 2022, 69, 1788–1799. [Google Scholar] [CrossRef] [Scilit]
- Zhou, K.; Zhou, K.; Yang, S. Reinforcement learning-based scheduling strategy for energy storage in microgrid. J. Energy Storage 2022, 51, 104379. [Google Scholar] [CrossRef] [Scilit]
- Cai, W.; Kordabad, A.B.; Gros, S. Energy management in residential microgrid using model predictive control-based reinforcement learning and Shapley value. Eng. Appl. Artif. Intell. 2023, 119, 105793. [Google Scholar] [CrossRef] [Scilit]
- Monfaredi, F.; Shayeghi, H.; Siano, P. Multi-agent deep reinforcement learning-based optimal energy management for grid-connected multiple energy carrier microgrids. Int. J. Electr. Power Energy Syst. 2023, 153, 109292. [Google Scholar] [CrossRef] [Scilit]
- Binyamin, S.S.; Slama, S.A.B.; Zafar, B. Artificial intelligence-powered energy community management for developing renewable energy systems in smart homes. Energy Strategy Rev. 2024, 51, 101288. [Google Scholar] [CrossRef] [Scilit]
- Hedayatnia, A.; Ghafourian, J.; Sepehrzad, R.; Al-Durrad, A.; Anvari-Moghaddam, A. Two-stage data-driven optimal energy management and dynamic real-time operation in networked microgrid based on a deep reinforcement learning approach. Int. J. Electr. Power Energy Syst. 2024, 160, 110142. [Google Scholar] [CrossRef] [Scilit]
- Abid, M.S.; Apon, H.J.; Hossain, S.; Ahmed, A.; Ahshan, R.; Lipu, M.H. A novel multi-objective optimization based multi-agent deep reinforcement learning approach for microgrid resources planning. Appl. Energy 2024, 353, 122029. [Google Scholar] [CrossRef] [Scilit]
- Benhmidouch, Z.; Moufid, S.; Ait-Omar, A.; Abbou, A.; Laabassi, H.; Kang, M.; Chatri, C.; Ali, I.H.O.; Bouzekri, H.; Baek, J. A novel reinforcement learning policy optimization based adaptive VSG control technique for improved frequency stabilization in AC microgrids. Electr. Power Syst. Res. 2024, 230, 110269. [Google Scholar] [CrossRef] [Scilit]
- Rajamallaiah, A.; Karri, S.P.K.; Shankar, Y.R. Deep reinforcement learning based control strategy for voltage regulation of dc-dc buck converter feeding cpls in dc microgrid. IEEE Access 2024, 12, 17419–17430. [Google Scholar] [CrossRef] [Scilit]
- Sepehrzad, R.; Langeroudi, A.S.G.; Khodadadi, A.; Adinehpour, S.; Al-Durra, A.; Anvari-Moghaddam, A. An applied deep reinforcement learning approach to control active networked microgrids in smart cities with multi-level participation of battery energy storage system and electric vehicles. Sustain. Cities Soc. 2024, 107, 105352. [Google Scholar] [CrossRef] [Scilit]
- Abouzeid, S.I.; Chen, Y.; Zaery, M.; Abido, M.A.; Raza, A.; Abdelhameed, E.H. Load frequency control based on reinforcement learning for microgrids under false data attacks. Comput. Electr. Eng. 2025, 123, 110093. [Google Scholar] [CrossRef] [Scilit]
- Barros, E.B.C.; Souza, W.O.; Costa, D.G.; Rocha Filho, G.; Figueiredo, G.B.; Peixoto, M.L.M. Energy management in smart grids: An Edge-Cloud Continuum approach with Deep Q-learning. Future Gener. Comput. Syst. 2025, 165, 107599. [Google Scholar] [CrossRef] [Scilit]
- Jia, X.; Xia, Y.; Yan, Z.; Gao, H.; Qiu, D.; Guerrero, J.M.; Li, Z. Coordinated operation of multi-energy microgrids considering green hydrogen and congestion management via a safe policy learning approach. Appl. Energy 2025, 401, 126611. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Chang, W.; Yang, Q. Deep reinforcement learning based hierarchical energy management for virtual power plant with aggregated multiple heterogeneous microgrids. Appl. Energy 2025, 382, 125333. [Google Scholar] [CrossRef] [Scilit]
- Xiong, B.; Zhang, L.; Hu, Y.; Fang, F.; Liu, Q.; Cheng, L. Deep reinforcement learning for optimal microgrid energy management with renewable energy and electric vehicle integration. Appl. Soft Comput. 2025, 176, 113180. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Wang, M.; Wang, A.; Zhang, X.; Zhang, J.; Ma, H.; Yang, N.; Zhao, Z.; Lai, C.S.; Lai, L.L. Multiagent deep reinforcement learning-based cooperative optimal operation with strong scalability for residential microgrid clusters. Energy 2025, 314, 134165. [Google Scholar] [CrossRef] [Scilit]
- Vazquez-Canteli, J.R.; Henze, G.; Nagy, Z. MARLISA: Multi-agent reinforcement learning with iterative sequential action selection for load shaping of grid-interactive connected buildings. In Proceedings of the 7th ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation; Association for Computing Machinery: New York, NY, USA, 2020; pp. 170–179. [Google Scholar]
- Xu, X.; Jia, Y.; Xu, Y.; Xu, Z.; Chai, S.; Lai, C.S. A multi-agent reinforcement learning-based data-driven method for home energy management. IEEE Trans. Smart Grid 2020, 11, 3201–3211. [Google Scholar] [CrossRef] [Scilit]
- Touzani, S.; Prakash, A.K.; Wang, Z.; Agarwal, S.; Pritoni, M.; Kiran, M.; Brown, R.; Granderson, J. Controlling distributed energy resources via deep reinforcement learning for load flexibility and energy efficiency. Appl. Energy 2021, 304, 117733. [Google Scholar] [CrossRef] [Scilit]
- Lissa, P.; Deane, C.; Schukat, M.; Seri, F.; Keane, M.; Barrett, E. Deep reinforcement learning for home energy management system control. Energy AI 2021, 3, 100043. [Google Scholar] [CrossRef] [Scilit]
- Pinto, G.; Deltetto, D.; Capozzoli, A. Data-driven district energy management with surrogate models and deep reinforcement learning. Appl. Energy 2021, 304, 117642. [Google Scholar] [CrossRef] [Scilit]
- Pinto, G.; Piscitelli, M.S.; Vázquez-Canteli, J.R.; Nagy, Z.; Capozzoli, A. Coordinated energy management for a cluster of buildings through deep reinforcement learning. Energy 2021, 229, 120725. [Google Scholar] [CrossRef] [Scilit]
- Gao, Y.; Matsunami, Y.; Miyata, S.; Akashi, Y. Operational optimization for off-grid renewable building energy system using deep reinforcement learning. Appl. Energy 2022, 325, 119783. [Google Scholar] [CrossRef] [Scilit]
- Heidari, A.; Maréchal, F.; Khovalyg, D. Reinforcement Learning for proactive operation of residential energy systems by learning stochastic occupant behavior and fluctuating solar energy: Balancing comfort, hygiene and energy use. Appl. Energy 2022, 318, 119206. [Google Scholar] [CrossRef] [Scilit]
- Huang, C.; Zhang, H.; Wang, L.; Luo, X.; Song, Y. Mixed deep reinforcement learning considering discrete-continuous hybrid action space for smart home energy management. J. Mod. Power Syst. Clean Energy 2022, 10, 743–754. [Google Scholar] [CrossRef] [Scilit]
- Langer, L.; Volling, T. A reinforcement learning approach to home energy management for modulating heat pumps and photovoltaic systems. Appl. Energy 2022, 327, 120020. [Google Scholar] [CrossRef] [Scilit]
- Lee, S.; Choi, D.H. Federated reinforcement learning for energy management of multiple smart homes with distributed energy resources. IEEE Trans. Ind. Inform. 2020, 18, 488–497. [Google Scholar] [CrossRef] [Scilit]
- Shen, R.; Zhong, S.; Wen, X.; An, Q.; Zheng, R.; Li, Y.; Zhao, J. Multi-agent deep reinforcement learning optimization framework for building energy system with renewable energy. Appl. Energy 2022, 312, 118724. [Google Scholar] [CrossRef] [Scilit]
- Lai, B.C.; Chiu, W.Y.; Tsai, Y.P. Multiagent reinforcement learning for community energy management to mitigate peak rebounds under renewable energy uncertainty. IEEE Trans. Emerg. Top. Comput. Intell. 2022, 6, 568–579. [Google Scholar] [CrossRef] [Scilit]
- Qiu, D.; Xue, J.; Zhang, T.; Wang, J.; Sun, M. Federated reinforcement learning for smart building joint peer-to-peer energy and carbon allowance trading. Appl. Energy 2023, 333, 120526. [Google Scholar] [CrossRef] [Scilit]
- Xie, J.; Ajagekar, A.; You, F. Multi-agent attention-based deep reinforcement learning for demand response in grid-responsive buildings. Appl. Energy 2023, 342, 121162. [Google Scholar] [CrossRef] [Scilit]
- Zhou, X.; Du, H.; Sun, Y.; Ren, H.; Cui, P.; Ma, Z. A new framework integrating reinforcement learning, a rule-based expert system, and decision tree analysis to improve building energy flexibility. J. Build. Eng. 2023, 71, 106536. [Google Scholar] [CrossRef] [Scilit]
- Deng, X.; Zhang, Y.; Jiang, Y.; Qi, H. A novel operation method for renewable building by combining distributed DC energy system and deep reinforcement learning. Appl. Energy 2024, 353, 122188. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Xiao, F.; Ran, Y.; Li, Y.; Xu, Y. Scalable energy management approach of residential hybrid energy system using multi-agent deep reinforcement learning. Appl. Energy 2024, 367, 123414. [Google Scholar] [CrossRef] [Scilit]
- Ajagekar, A.; Decardi-Nelson, B.; You, F. Energy management for demand response in networked greenhouses with multi-agent deep reinforcement learning. Appl. Energy 2024, 355, 122349. [Google Scholar] [CrossRef] [Scilit]
- Aldahmashi, J.; Ma, X. Real-time energy management in smart homes through deep reinforcement learning. IEEE Access 2024, 12, 43155–43172. [Google Scholar] [CrossRef] [Scilit]
- Kang, H.; Jung, S.; Kim, H.; Jeoung, J.; Hong, T. Reinforcement learning-based optimal scheduling model of battery energy storage system at the building level. Renew. Sustain. Energy Rev. 2024, 190, 114054. [Google Scholar] [CrossRef] [Scilit]
- Kumari, A.; Kakkar, R.; Tanwar, S.; Garg, D.; Polkowski, Z.; Alqahtani, F.; Tolba, A. Multi-agent-based decentralized residential energy management using deep reinforcement learning. J. Build. Eng. 2024, 87, 109031. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Wang, Y.; Qiu, D.; Su, H.; Strbac, G.; Gao, Z. Resilient energy management of a multi-energy building under low-temperature district heating: A deep reinforcement learning approach. Appl. Energy 2025, 378, 124780. [Google Scholar] [CrossRef] [Scilit]
- Lami, B.; Alsolami, M.; Alferidi, A.; Slama, S.B. A smart microgrid platform integrating AI and deep reinforcement learning for sustainable energy management. Energies 2025, 18, 1157. [Google Scholar] [CrossRef] [Scilit]
- Talihati, B.; Fu, S.; Zhang, B.; Zhao, Y.; Wang, Y.; Sun, Y. Community shared ES-PV system for managing electric vehicle loads via multi-agent reinforcement learning. Appl. Energy 2025, 380, 125039. [Google Scholar] [CrossRef] [Scilit]
- Liu, H.; Wu, W. Federated reinforcement learning for decentralized voltage control in distribution networks. IEEE Trans. Smart Grid 2022, 13, 3840–3843. [Google Scholar] [CrossRef] [Scilit]
















| Framework | Strengths | Limitations | Key Positioning |
|---|---|---|---|
| Stochastic Optimization | Handles uncertainty through scenarios or probabilistic forecasts. | Requires accurate probability models; can be computationally heavy. | Strong for uncertainty-aware scheduling with known probability models. |
| Robust Optimization | Ensures feasibility under worst-case uncertainty. | Can be overly conservative and dependent on the uncertainty set. | Strong for conservative operation under bounded RES/load uncertainty. |
| MPC/MILP/OPF | Uses explicit models, constraints, and operational limits. | Requires accurate models and repeated online optimization. | Strong for interpretable and constraint-aware model-based control. |
| Hybrid Data–Model Methods | Combine data, forecasts, and physical optimization. | Depend on model structure and forecast quality. | Bridge data adaptivity with model-based feasibility and transparency. |
| RL/DRL/MARL | Learns adaptive policies for nonlinear and dynamic systems. | Training can be unstable; feasibility is not inherent. | Strong for adaptive, sequential, and fast real-time decision-making. |
| Evaluation Aspect | [65] | [66] | [67] | [68] | Current |
|---|---|---|---|---|---|
| RL Trends and Status | x | x | x | ||
| RL Types | x | x | x | ||
| Agent Architectures | x | x | |||
| Reward Design | x | x | |||
| Baseline Approaches | x | x | x | x | x |
| Input Data | x | x | x | x | |
| Control Objectives | x | x | x | x | x |
| RES Types | x | x | x | ||
| Grid Types | x | x | x | ||
| Performance Comparison | x | x | x | x | |
| Simulation Tools | x | ||||
| Power Grids | x | x | |||
| Microgrids | x | x | x | x | |
| Buildings | x | x | x | ||
| Papers Evaluation Scale | 67 | 13 | - | Narrative | 84 |
| Framework | Typical RES | Capacity | Operational Role of RES |
|---|---|---|---|
| Power Grids | Utility-scale solar PV farms, onshore/offshore wind farms, hydropower plants, biomass power plants, concentrated solar power (CSP) | MW–GW scale | Bulk electricity generation for national or regional networks, contribution to energy market supply, provision of ancillary services, reduction of fossil fuel generation, grid decarbonization |
| Microgrids | Distributed PV arrays, small wind turbines, biomass generators, small hydropower plants, hybrid RES systems | kW–tens of MWs | Local electricity generation, energy autonomy and resilience, coordination with ESS for balancing RES intermittency, demand-side management, participation in local energy markets |
| Buildings | Rooftop solar PVs, solar thermal collectors, building-integrated photovoltaics, geothermal heat pumps, small wind turbines | kW–hundreds of kWs | On-site RES energy production, reduction of electricity demand, demand-response support, integration with HVAC and ESS to improve energy efficiency and flexibility |
| Framework | Primary Control Objectives | Controlled Components |
|---|---|---|
| Power Grids | Voltage regulation, frequency stabilization, economic dispatch, coordinating RES generation, congestion management | Reactive power from RES inverters, transformer tap changers (OLTC), ESS, conventional generators, flexible loads |
| Microgrids | Energy scheduling, cost optimization, maximizing RES utilization, storage coordination, islanding management, resilience improvement | Battery storage systems, diesel generators, CHP units, EV charging, flexible loads |
| Buildings | Energy cost minimization, maximizing RES self-consumption, HVAC comfort control, appliance scheduling, demand response participation | HVAC systems, battery storage, EV chargers, smart appliances, rooftop PV coordination |
| Ref. | Year | Method | Agent | FH/TS | Baselines | RES | Integration | Type | Data | Cit. |
|---|---|---|---|---|---|---|---|---|---|---|
| [162] | 2020 | PPO/IL | Single | N/A | OPF | WT | DG/GCE/Grid | Transmission | Real | 55 |
| [163] | 2020 | DPPG | Multi | 24 h/1 h | RL/Droop | PV | DG/PE//Load/Grid | ADN | Real | 246 |
| [164] | 2020 | DPPG | Single | 24 h/1 h | MPC | WT | DG/GRID/PE/GCE | ADN | Synth | 135 |
| [165] | 2020 | DQN/SOCP | Single | 24 h/1 h | Fixed/Greedy | PV | PE/ GCE/Grid | ADN | Real | 286 |
| [166] | 2020 | DDPG | Multi | N/A | Droop | PV | DG/GCE/Grid | Transmission | Synth | 34 |
| [167] | 2021 | DDPG | Multi | 24 h/1 h | RBC | PV | DG/ESS/ PE/Load/Grid | ADN | Synth | 197 |
| [168] | 2021 | TD3 | Multi | -/1 h | RL/SP | PV | DG/PE/Load/Grid | ADN | Real | 156 |
| [169] | 2021 | SAC | Multi | N/A | RL/SP | PV | DG/PE/GCE/Grid | ADN | Real | 153 |
| [170] | 2021 | DDPG | Multi | N/A | RL/ACOPF | PV/WT | DG/PE/Load/Grid | Transmission | Synth | 63 |
| [171] | 2022 | CPO | Single | 24 h/1 h | RL/SOCP | PV/WT | DG, ESS/GCE/Grid | ADN | Real | 144 |
| [172] | 2022 | DQN/DDPG | Single | 24 h/1 h | GA/PSO/OPF | PV/WT | DG/ESS/Grid/Loads | ADN | Synth | 54 |
| [173] | 2022 | EA-MAAC/MPC | Multi | 24 h/1 h | RL/SOCP | PV | ESS/EV/GCE | ADN | Synth | 163 |
| [174] | 2022 | REINFORCE | Multi | N/A | RL/Linear | PV | DG/PE/Grid | ADN | Synth | 40 |
| [175] | 2022 | TD3 | Multi | N/A | RL/PSO/GA/PID | PV/WT/Hydro | DG/EV/ESS/GCE/Grid | Transmission | Synth | 97 |
| [176] | 2022 | TD3/PID | Multi | N/A | RL/PID/FOPI | WT | DG/GCE/Grid | Transmission | Synth | 43 |
| [177] | 2022 | DDPG | Single | N/A | OPF | PV/WT | PE/DG/Grid | ADN | Synth | 55 |
| [178] | 2022 | DDPG/ANN | Single | N/A | RL/MPC/SP | PV | DG/PE/Grid | Transmission | Synth | 94 |
| [179] | 2023 | DDPG | Multi | 30 m/5 m | RL/SOCP | PV | GCE/ESS/DG/PE | ADN | Real | 35 |
| [180] | 2023 | DDPG/Droop | Multi | 24 h/15 m | RBC/RL | PV | DG/Load/Grid | ADN | Synth | 20 |
| [181] | 2023 | TD3/PPO | Single | 24 h/30 m | OPF | PV | ESS/DG/PE/Grid | ADN | Real | 22 |
| [182] | 2024 | PPO | Single | N/A | RBC/MPC | PV/WT | DG/PE/Grid | Transmission | Synth | 56 |
| [183] | 2024 | SAC | Multi | 12 h/3 m | RL/Droop | PV | DG/PE/Load/Grid | ADN | Real | 20 |
| [184] | 2024 | DDPG | Multi | 24 h/3 m | RL/OPF | PV | DG/Grid | ADN | Real | 41 |
| [185] | 2024 | PPO | Single | N/A | RL/SOCP | PV | DER/Load/Grid | ADN | Synth | 55 |
| [186] | 2025 | Unspecified | Multi | N/A | RL/ANN/Conv. | PV/WT | ESS/Load/Grid/Biomass | Distribution | Synth | 47 |
| [187] | 2025 | Unspecified | Single | -/1 m | ANN/Conv. | PV | ESS/Grid | Grid-Connected | Real | 37 |
| [188] | 2025 | DQN | Single | 24 h/1 h | GA/PSO/MPC/Conv | PV/WT | DG/ESS/Load/Grid | Distribution | Real | 33 |
| [189] | 2025 | DDPG | Multi | -/3 m | RL | PV/WT | DG/Load/Grid | ADN | Real | 29 |
| [190] | 2025 | TD3/OPF | Single | 24 h/5 m | RL/OPF | PV | DG/Load/Grid | ADN | Real | 25 |
| Author | Summary |
|---|---|
| Zhou et al. [162] | This study formulates AC OPF as an MDP in which the state includes load (P/Q) and generator setpoints, while the action space consists of continuous adjustments of generator active power and voltage. A PPO actor–critic agent, initialized through imitation learning based on OPF solver outputs, learns a policy that minimizes generation cost while penalizing constraint violations. The reward combines economic efficiency with feasibility. Evaluated on transmission grids, the method achieves approximately 0.6% cost deviation, nearly 100% feasibility, around 7× faster execution than IPS, and single-step convergence in about 99% of cases. |
| Cao et al. [163] | This work formulates voltage regulation as a Markov game in which each PV inverter agent observes local states, including load P/Q and PV generation, and outputs reactive power actions. A MADDPG actor–critic framework with attention enables coordination through centralized training and decentralized execution. The reward is defined on the basis of global voltage deviation, thereby promoting system-wide voltage stability. On IEEE distribution systems, the method reduces average voltage deviation by about 74% compared with droop control (from 0.46% to 0.12%) and reaches 95.1% of centralized optimal performance while maintaining real-time capability. |
| Kou et al. [164] | A centralized DDPG-based safe RL agent is developed for voltage control in an active distribution network. The state includes bus voltages and DG outputs, while the action corresponds to reactive power adjustments of HDTs/LPCs. The reward penalizes voltage violations, power losses, and control effort, and a safety layer is introduced to enforce operational constraints. The method achieves about 45% loss reduction (197.23 → 109.03 kW), maintains voltages within 0.95–1.05 p.u., and converges in roughly 200 episodes. |
| Yang et al. [165] | This study proposes a two-timescale RL-based Volt/VAR control strategy for distribution grids. A DQN agent manages capacitor switching at the slow timescale, whereas smart inverters solve a convex optimization problem for fast reactive power control. The MDP state includes load levels and capacitor status, the actions correspond to discrete capacitor settings, and the reward minimizes cumulative voltage deviation. Validation on a real 47-bus feeder and the IEEE 123-bus system with real PV/load data shows a voltage deviation reduction of about 20–40% relative to fixed, random, and greedy baselines. |
| Rozada et al. [166] | The paper formulates load-frequency control as a multi-agent MDP in which each generator agent observes local states, specifically frequency deviation and control signal , and produces continuous adjustments through MADDPG actor networks with LSTM memory. The reward is defined as an exponential penalty on frequency deviation. Centralized critics use global information during training, while execution remains decentralized. On simulated grids with 2–8 generators, the method achieves near-optimal cumulative rewards (about 900/1000) and restores frequency rapidly after disturbances, outperforming conventional centralized approaches. |
| Cao et al. [167] | This study proposes a multi-agent actor–critic DRL framework for voltage control, in which agents observe local states, including load P/Q, PV generation, and voltages, and output continuous control actions for PVs, BESS, and SVCs. A shared global reward penalizes voltage deviations and constraint violations through a surrogate power-flow model under centralized training and decentralized execution. Tests on IEEE distribution feeders show complete elimination of voltage violations and approximately 30–50% lower voltage deviations than conventional control, with stable performance under stochastic RES conditions. |
| Cao et al. [168] | This work formulates Volt-VAR control as a multi-agent Markov game. Each agent observes local states, including load P/Q and PV output, and outputs continuous reactive power actions for PV inverters and SVCs. An attention-enhanced TD3 framework is adopted under centralized training and decentralized execution. The reward minimizes global voltage deviation while penalizing constraint violations. On IEEE feeders, the method achieves about 0.13% voltage deviation and 96.4% optimality, outperforming MADDPG and decentralized RL baselines. |
| Cao et al. [169] | The paper models voltage control as a Markov game in which each subnetwork agent observes local voltages, PV outputs, and reactive power states and produces continuous inverter control actions, while a higher-level SAC agent schedules OLTC and capacitor operations using global states. The reward penalizes voltage deviations and excessive switching and is computed through a GP-based surrogate power-flow model. Under centralized training and decentralized execution, the method attains near-centralized performance and substantially reduces voltage violations on IEEE feeders compared with stochastic and DRL baselines. |
| Zhang et al. [170] | This study introduces a hierarchical multi-agent RL framework in which system agents (MADDPG) determine voltage references and SVC agents (DDPG) regulate reactive power locally. States include local voltages and global bus voltages for the critic, while actions correspond to reactive power injections and voltage setpoints. The reward penalizes voltage deviations through a dynamic exponential function. Tested on large transmission grids with stochastic PV/wind scenarios, the method fully eliminates voltage violations (within 0.05 p.u.) and achieves orders-of-magnitude faster computation (about 0.002–0.033 s compared with seconds for ACOPF). |
| Li et al. [171] | The work formulates distribution network operation as a CMDP, where the agent observes historical load, RES generation, prices, and BESS states, and outputs hybrid actions controlling DG power, storage, and voltage devices. A CPO-based actor–critic policy enforces constraints through a separate cost function while maximizing economic reward. The reward minimizes procurement and generation costs, whereas constraints penalize voltage, current, and operational violations. On IEEE feeders with real data, the method achieves up to 32.5% cost reduction relative to DDPG, near-zero constraint violations, and performance close to optimal MISOCP (about 11% gap). |
| Bui et al. [172] | This study presents a three-stage DRL framework in which a single agent observes grid states, including DG outputs, loads, switch status, SOC, and prices, and outputs both discrete switching actions and continuous generation/BESS setpoints. The reward penalizes power losses, load shedding, and operating cost through power-flow simulation. Training is carried out offline in a MATLAB-based environment, allowing real-time inference without re-optimization. On microgrid and IEEE 33-bus systems, the method delivers faster decisions and improved operational efficiency compared with conventional optimization, particularly under fault conditions. |
| Hu et al. [173] | The paper develops a multi-agent actor–critic RL framework (EA-MAAC) under CTDE, in which each agent observes local voltages and grid states and outputs continuous P/Q setpoints for PV inverters and EV loads. The reward combines line-loss minimization with voltage-violation penalties, thereby supporting coordinated voltage regulation. Centralized critics with attention enhance scalability, whereas decentralized execution enables real-time operation. On IEEE feeders, the method reduces voltage violations and improves returns relative to MADDPG and SAC, approaching optimal VVO performance with substantially lower computation time. |
| Cui et al. [174] | This work proposes a decentralized policy-gradient RL framework (REINFORCE) that trains local neural-network controllers for each bus, using local voltage deviations as states and reactive power injections as actions. The reward minimizes voltage deviation and control effort, while Lipschitz constraints embedded in the neural networks ensure closed-loop stability. On a simulated IEEE 33-bus distribution grid, the approach achieves about 18.18% cost reduction relative to linear control and 5.26% relative to standard safe RL, while avoiding the instability observed in unconstrained RL. |
| Khalid et al. [175] | A multi-agent TD3 actor–critic framework is employed in which each agent observes local ACE-based states, namely proportional, integral, and derivative components, and outputs continuous PID gain actions. The reward minimizes frequency deviation and tie-line power error, thereby promoting effective load-frequency control. The system is tested on a two-area renewable-integrated transmission grid with PV, wind, EVs, BESS, and thermal units. It achieves approximately 50–60% faster settling time (around 7.5 s versus more than 14 s) and markedly lower oscillations than DDPG, PSO, and GA. |
| Li et al. [176] | The paper formulates multi-area AGC as a continuous multi-agent RL problem, in which each agent observes local frequency deviation, tie-line power, and system states and outputs tuning actions for PI controller gains. The IGE-MATD3 actor–critic framework employs twin critics, delayed updates, and imitation-guided exploration to improve stability and coordination. The reward combines frequency deviation, control effort, and market performance. On a simulated multi-area transmission system, the method reduces frequency deviation by about 30–60% and lowers control cost compared with MADDPG, MATD3, and heuristic PI baselines. |
| Li et al. [177] | This study formulates voltage control as an MDP in which the state includes node voltages and power injections, while the actions correspond to active/reactive power setpoints of M-SOP converters. A DDPG actor–critic agent learns a continuous control policy, and the reward minimizes normalized voltage deviations. A key contribution is an action-masking layer that enforces physical constraints such as power balance and capacity coupling. On an IEEE 33-node active distribution network, the method achieves about 61% reduction in voltage deviation and 90.65% optimality relative to centralized control, demonstrating near-optimal adaptive performance under RES uncertainty. |
| Cao et al. [178] | A DDPG-based actor–critic agent observes node voltages, PV power, and load injections and outputs continuous control actions for PV reactive power, SVC control, and PV curtailment. The reward penalizes voltage deviation and curtailment through a surrogate power-flow model. Trained on simulated distribution networks, the policy achieves about 77% reduction in voltage deviation (3.64% → 0.82%), outperforming MPC and DDQN while enabling real-time control at around 0.001 s. |
| Wu et al. [179] | The work formulates voltage control as an MDP in which multiple agents (OLTC, CBs, PV, ESS) observe grid states such as voltages, loads, SOC, and previous actions, and output coordinated control signals across two timescales. A MADDPG actor–critic framework with a centralized critic supports cooperation, while Gumbel-Softmax is used to handle discrete actions. The reward penalizes global voltage deviation. On an IEEE 123-bus active distribution network, the method maintains voltages within approximately 0.99–1.03 p.u. and achieves performance comparable to optimal control with more than 4500× faster computation (0.177 s versus 812 s), outperforming DDPG/DQN under RES variability. |
| Li et al. [180] | The paper formulates Volt/VAR control as a POMG-based multi-agent actor–critic problem, in which each agent observes local states, including PV generation, load P/Q, node voltages, and neighbor information, and outputs droop-curve parameters rather than direct control actions. The reward minimizes power losses and voltage violations, with an additional linear penalty introduced for training stability. Using MADDPG with centralized training and decentralized execution, agents coordinate through limited communication across subnetworks. On an IEEE 123-bus distribution grid, the method achieves about 25% loss reduction relative to droop control, eliminates voltage violations, and supports real-time inference of roughly 20 ms. |
| Petrusev et al. [181] | This study applies a single-agent actor–critic DRL framework (TD3PG/PPO) in which the agent observes grid voltages, battery P/Q, SOC, and time, and outputs continuous inverter setpoints (P, Q) for ESS. The reward penalizes voltage violations outside 0.95–1.05 p.u. and SOC deviations, thereby maintaining both grid stability and storage availability. A two-stage policy combining offline pre-training with online adaptation improves robustness to unknown or drifting line impedances. On distribution grids, PPO reduces voltage violations from about 19.5% to 1.85%, and under uncertainty yields about 50% improvement over OPF with much faster execution. |
| Rehman et al. [182] | A single-agent PPO actor–critic framework is employed to learn reactive power control policies for PV inverters and SVCs from system states such as bus voltages and RES output variability. The reward penalizes voltage deviations and power losses, thereby encouraging stable operation. Tested on an IEEE 33-bus distribution network, the method achieves ±5% voltage regulation, 68% controllability, and a violation ratio of 0.044, outperforming conventional control methods. |
| Huang et al. [183] | The study formulates voltage control as a Markov game in which each PV inverter agent observes local states (PL, QL, PPV, QPV) and outputs continuous reactive power actions. It adopts an attention-enhanced MASAC framework under CTDE, where critics use global information and attention-weighted inter-agent interactions, while actors operate locally. The reward combines penalties on voltage deviation and reactive power loss. On IEEE distribution systems, the method achieves about 96–97% controllable rate and roughly 30% cost reduction relative to droop control, outperforming MADDPG, MATD3, and SAC variants. |
| Zhang et al. [184] | A multi-agent GT-MADDPG framework is proposed in which each PV inverter observes local states, including voltages and topology embeddings obtained from a GNN, and outputs continuous reactive power actions. The reward penalizes voltage deviations and power losses and is further strengthened through physics-guided constraints introduced via a PINN loss. Training follows CTDE with replay buffers and graph-based representations. On IEEE 33- and 141-bus systems, the method achieves about 20–30% lower voltage deviation than DDPG/MADDPG and runs orders of magnitude faster than OPF, while remaining robust under uncertainty. |
| Jacob et al. [185] | Propose a graph-based PPO framework for outage management in active distribution networks, where a centralized agent learns topology-aware switching and load-shedding actions. States include load, DER generation, voltages, branch flows, topology, and outage information, while rewards maximize supplied energy and penalize voltage violations. Tested on modified IEEE 13-, 34-, and 123-bus systems in OpenDSS, the method achieved near-optimal restoration compared with MISOCP/BPSO, while operating in milliseconds and reducing outage-related energy loss. |
| Singh et al. [186] | Proposes the ORA-DL hybrid DRL framework, which combines DNN-based demand forecasting with multi-agent reinforcement learning for adaptive smart-grid resource allocation. Agents use demand, renewable generation, ESS, and IoT sensor data to make dynamic energy-distribution decisions aimed at improving grid stability, resource utilization, and operating cost. In simulation, ORA-DL outperformed conventional ML/IoT/DRL approaches in prediction accuracy, allocation efficiency, energy wastage, and cost reduction. |
| Mamodiya et al. [187] | Propose a single-agent RL framework for dual-axis solar tracking, combined with CNN-LSTM irradiance forecasting and Edge AI control. The agent observes irradiance forecasts, environmental conditions, and panel orientation, and adjusts azimuth and elevation angles to maximize captured solar energy. Applied to adaptive PV modules with hybrid storage in a real smart-grid solar testbed, the framework improved annual energy yield, spectral absorption, panel temperature, and battery lifetime compared with conventional fixed-tilt and MPPT approaches. |
| Tang et al. [188] | proposes a single-agent DQN-based DRL framework for energy scheduling in a smart-grid-integrated virtual power plant (VPP) with PV, wind, ESS, loads, and grid-market interaction. Using demand, SOC, renewable generation, and price data, the agent optimizes storage and market decisions. Tested on real Shenzhen VPP data, it reduced grid losses and improved renewable utilization compared with conventional optimization methods. |
| Luo et al. [189] | The paper proposes a GAT-enhanced MADDPG framework for Volt/Var control in PV-rich active distribution networks. PV inverter agents use local load, PV generation, reactive power, voltage, and phase-angle information to regulate reactive power and jointly reduce voltage violations and losses. Tested on IEEE 33-bus and IEEE 141-bus systems with real PV/load data, GAMARL reduced voltage-violation frequency from 2.5% to 0.8%, cut violation occurrences by 80.7%, and lowered network losses by 4.7% compared with the best MARL baseline. |
| Faghri et al. [190] | The paper proposes a hybrid TD3 and convex OPF framework for real-time scheduling in grid-connected smart-building communities with PV, EV parking lots, loads, and a distribution network. Using real CAISO 5-min data, the method reduced online computation time by 82.5% compared with OPF alone, while EV flexibility lowered distribution-network operating cost by 5.15% versus uncontrolled charging. |
| Ref. | Year | Method | Agent | FH/TS | Baselines | RES | Integration | Type | Data | Cit. |
|---|---|---|---|---|---|---|---|---|---|---|
| [191] | 2020 | Dyna-QL | Single | 72 h/- | RBC/Fixed | PV | DG/ESS | Islanded | Real | 54 |
| [192] | 2020 | QL | Multi | -/1 h | RL | PV/WT | DG/ESS/Load/Grid | Grid-connected | Synth | 144 |
| [193] | 2020 | DPPG | Single | 24 h/1 h | RL | PV | DG/ESS | Islanded | Real | 100 |
| [194] | 2020 | MuZero/OPF | Single | N/A | RL/MPC/DP | PV/WT | DG/ESS/Load/Grid | Islanded | Synth | 93 |
| [195] | 2021 | A2C | Multi | 1 s/0.05 s | RL/MPC | PV/WT | DG/GCE/Load/Grid | Islanded | Synth | 82 |
| [196] | 2021 | DDPG/MILP | Single | 24 h/1 h | RL/ANN/MLR | PV/WT | DG/ESS/Load | Islanded | Real | 239 |
| [197] | 2021 | A3C | Multi | 24 h/15 m | RL | PV/WT | DG/ESS/Load/Grid | Grid-connected | Real | 55 |
| [198] | 2022 | MCTS | Multi | N/A | Deterministic | PV/WT/BIO | DG/ESS/Grid | Networked | Synth | 151 |
| [199] | 2022 | DDPG | Single | -/50 ms | Droop/RBC | PV | DG/ESS/Load/Grid | Islanded | Synth | 50 |
| [200] | 2022 | PPO | Single | 24 h/1 h | RL/SP/MPC | PV | DG/ESS/Load/Grid | Grid-connected | Real | 238 |
| [201] | 2022 | DQN | Single | -/1 h | RL/MPC | PV/WT | ESS/PE/GCE/Load/Grid | Grid-connected | Real | 53 |
| [202] | 2022 | DDPG | Multi | 24 h/1 h | RL | PV | DG/ESS/Load/Grid | Grid-connected | Real | 97 |
| [203] | 2022 | SAC/MPC | Single | 48 h/1 h | RL/MPC | PV | ESS/Load/Grid | Grid-connected | Both | 59 |
| [204] | 2022 | SAC | Multi | -/30 s | RL | PV | DG/Load/Grid | Islanded | Real | 61 |
| [205] | 2022 | TR-RL/ANN | Single | 24 h/1 h | RL | PV | DG/Load/Grid | Networked | Both | 68 |
| [206] | 2023 | QL | single | 24 h/1 h | MILP | PV | DG/ESS/Load/Grid | Grid-connected | Real | 64 |
| [207] | 2023 | DPG/MPC | Single | 30 d/12 h | MPC | PV | DG/Load/Grid | Community | Real | 65 |
| [208] | 2023 | DDPG | Multi | 24 h/1 h | RL/PSO/MILP | PV/WT | DG/ESS/Load/Grid | Community | Real | 50 |
| [209] | 2024 | QL/ANN | Multi | N/A | RL/MILP | PV | DG/ESS/EV/Load/Grid | Community | Synth | 50 |
| [210] | 2024 | DDPG | Multi | 24 h/1 h | RL | PV/WT | DG/ESS/Load/Grid | Networked | Synth | 88 |
| [211] | 2024 | DDPG | Multi | N/A | RL/Hybrid | PV/WT | RES/ESS/EVCS/Grid | Grid-connected | Both | 79 |
| [212] | 2024 | TD3 | Single | 1 s/- | RL/Conv. | PV/WT | DG/ESS/EVCS | Islanded | Synth | 43 |
| [213] | 2024 | TD3 | Single | N/A | RL | PV/WT | PE/CPL/DC bus | Islanded | Real * | 57 |
| [214] | 2024 | SAC/ANN | Multi | 5 ms | PID/FLC/ANN | PV | DG/ESS/EV/Load/Grid | Islanded | Synth | 58 |
| [215] | 2025 | DDPG/PID | Single | N/A | PID | PV/WT | DG/ESS/EVCS | Other | Synth | 21 |
| [216] | 2025 | DQN | Single | 15 d/1 h | GA/Conv | PV/WT | DG/Load/Grid | Community | Synth | 25 |
| [217] | 2025 | SAC | Single | 7 d/1 h | RL | PV/WT | CHP/GT/Load/Grid | Grid-Connected | Real | 36 |
| [218] | 2025 | PPO | Multi | 24 h/30 m | RL | PV/WT | DG/ESS/Load/Grid | Grid-Connected | Real | 34 |
| [219] | 2025 | PPO | Single | 24 h/1 h | RL/MILP/PSO/RBC | PV/WT | DG/ESS/EV/Load/Grid | Grid-Connected | Real | 28 |
| [220] | 2025 | SAC | Multi | 24 h/1 h | RL/MPC | PV | DG/ESS/EV/HVAC/Load/Grid | Community | Real | 50 |
| Author | Summary |
|---|---|
| Nyong et al. [191] | This study employs a Dyna-Q agent that observes battery SOC and pinch-based energy targets and selects dispatch actions for HESS components, including BAT, FC, EL, and HT. The reward penalizes storage-limit violations and inefficient energy use while encouraging coordinated operation under uncertainty. Applied to an islanded PV + HESS microgrid using real load and irradiance data, the method outperforms deterministic and adaptive baselines in constraint handling and energy management, although no explicit numerical improvements are reported. |
| Samadi et al. [192] | This work applies decentralized multi-agent Q-learning, where each DER and consumer observes local states such as time, prices, and operational constraints, including SOC and ramp limits, and selects discrete generation, consumption, or bidding actions. Rewards reflect individual profit or cost, while a hierarchical EMS clears the global market. Agents learn independently through -greedy, softmax, or UCB policies without direct communication. Simulation results show that the softmax strategy delivers the best overall performance, increasing DER profit, reducing customer cost, and lowering grid dependence. |
| Lei et al. [193] | The paper formulates microgrid dispatch as a finite-horizon POMDP, in which the agent observes load, PV generation, and battery states and outputs continuous dispatch actions for DG and ESS units. The reward penalizes generation cost and power imbalance. Two actor–critic variants, FH-DDPG and recurrent FH-RDPG, are trained on real microgrid data to address uncertainty under partial observability. Results demonstrate clear cost reductions and improved power balance, with FH-RDPG showing the strongest robustness under stochastic conditions. |
| Shuai et al. [194] | This study develops a model-based RL approach of the MuZero type, where the state includes past RES generation, load, and system variables such as SoC, encoded through LSTM, while the actions correspond to ESS, DG, and grid-exchange dispatch. The reward minimizes operating cost, including fuel, degradation, and curtailment terms. Policy optimization is carried out through MCTS over a learned dynamics model combined with OPF constraints. On a grid-connected residential microgrid, the method achieves near-optimal cost performance without relying on explicit forecasting models. |
| Chen et al. [195] | The work formulates secondary voltage control as a decentralized MARL POMDP. Each DG observes local electrical states, including P/Q, voltages, currents, and neighbor messages, and then selects discrete voltage setpoints between 1.0 and 1.14 pu. The reward penalizes voltage deviations, with stronger penalties applied outside safe limits, while a spatially discounted global reward enhances coordination. Using an actor–critic architecture with LSTM-based communication encoding and action smoothing, PowerNet achieves about 36% better performance than PPO and the highest rewards (approximately 0.74 versus 0.66 for the baselines) on IEEE-based microgrids, while also showing strong scalability. |
| Li et al. [196] | This study proposes a DDPG-based PER-AutoRL framework in which the agent maps historical PV, WT, and load time series to continuous forecasting outputs, with rewards defined through forecasting-error metrics such as MAPE and RMSE. Prioritized replay accelerates learning, while Bayesian optimization is used to tune the architecture and hyperparameters. These RL-based forecasts are then incorporated into a chance-constrained MILP scheduler for an islanded microgrid with PV, WT, MT, ESS, and DR. The method reduces operating cost by up to 38.4% (36.7% at 95% confidence) and lowers spinning reserve requirements in 66.7% of periods, outperforming MLR, ARIMA, LSTM, and RDPG baselines. |
| Munir et al. [197] | This paper presents a multi-agent A3C framework in which agents observe stochastic energy states, including generation, demand, and storage levels, and output scheduling actions for microgrid resources and MEC loads. A global critic with shared networks supports coordinated learning under a CVaR-based reward that penalizes energy mismatch and shortage risk. Applied to a grid-connected microgrid with PV, storage, and grid dispatch, the method achieves about 9.5% lower training loss and roughly 2% higher accuracy than single-agent RL, while also improving stability under uncertainty. |
| Tomin et al. [198] | The study formulates microgrid energy management as an MDP solved through MCTS, where the agent observes PV/wind generation, load, and storage SOC and selects battery charge/discharge and dispatchable generation actions. The reward maximizes economic profit by penalizing curtailment and load shedding, while UCB balances exploration and exploitation. Evaluated on realistic community microgrids with RES, ESS, and grid interaction, the method achieves substantial profit gains, in some cases exceeding 100%, compared with non-cooperative operation. |
| Barbalho et al. [199] | The paper employs a centralized DDPG actor–critic agent whose state includes node voltages at the PCC and BESS buses, together with current and previous system frequency, while the actions are continuous active and reactive power setpoints for BESS units. The reward penalizes voltage and frequency deviations as well as unstable operating regions. Trained on an islanded AC microgrid in Simulink with SG–BESS–PV interactions, the method achieves up to 86% cumulative error reduction and as much as 4.7× improvement over droop control, with near-zero voltage deviation in several cases. |
| Guo et al. [200] | The work formulates microgrid optimal energy management as an MDP in which the state includes PV/WT output, load, electricity price, and ESS SOC, while the actions control MT output and ESS charge/discharge. A PPO actor–critic policy is trained using clipped loss and GAE, with a reward that minimizes operating cost while penalizing constraint violations. In a grid-connected microgrid, the method achieves about 5–6% lower cost than MPC under uncertainty and converges faster (150 versus 250 episodes), while remaining robust to forecast errors. |
| Harrold et al. [201] | This study models battery scheduling as a partially observable MDP. The state includes SOC, demand, RES generation, price, and, optionally, ANN-based forecasts, while the actions correspond to discrete charge/discharge levels. Rainbow DQN, incorporating distributional learning, prioritized replay, and a dueling architecture, optimizes a reward based on normalized cost savings and constraint penalties. Applied to a grid-connected microgrid with PV, WT, and dynamic pricing, the method achieves up to 20.22% higher savings than DQN and outperforms DDPG and LP/MPC-type baselines, reaching approximately £91.5k in total savings. |
| Harrold et al. [202] | This work employs a multi-agent actor–critic framework (MADDPG/TD3/D3PG), in which each agent observes local states such as demand, prices, RES output, and ESS conditions, and outputs continuous charging, discharging, or trading actions. Rewards are agent-specific and based on marginal contribution, thereby aligning local decisions with global profit. Under centralized training and decentralized execution, the framework coordinates heterogeneous storage units in a grid-connected microgrid with HESS and energy trading. MARL achieves higher profits than single-agent control, while energy trading performs better than standard grid export. |
| Hu et al. [203] | The study applies a SAC-based actor–critic agent to the second-stage MDP, where the state includes net load, battery SOC, and supercapacitor SOC, and the actions correct grid power, battery dispatch, and supercapacitor output around first-stage references. The reward penalizes schedule deviations, SOC imbalance, and degradation cost, while entropy regularization improves stability. In a grid-connected microgrid with PV, WT, and hybrid ESS, the method provides real-time correction of scheduling decisions, achieving low operating cost (264 USD total), fast convergence, and sub-second control time (about 0.156 s), thereby outperforming single-stage approaches in robustness. |
| Xia et al. [204] | A multi-agent SAC framework is proposed in which each microgrid agent observes local frequency deviation and historical PV/load profiles and outputs continuous actions for DG setpoints and ESS power through SOC-dependent parameterization. A global reward penalizes frequency deviations and quadratic generation costs, while a safety layer filters unsafe actions. Under centralized training and local execution, the method converges stably in about 6000 episodes, removes unsafe actions after around 3000 episodes, maintains frequency within ±0.04 Hz, and achieves lower operating cost than standard multi-agent SAC and DDPG methods. |
| Xiong et al. [205] | The paper formulates electricity pricing as an MDP in which the state includes time, wholesale price, and interval forecasts of PV and load, while the actions are continuous retail prices for each microgrid. A policy-gradient actor with a Gaussian policy is enhanced by a differentiable trust-region layer (TRLF/TRLW) to stabilize updates under uncertainty. The reward maximizes DSO profit, and a neural-network surrogate replaces direct interaction with the microgrid. On multi-microgrid IEEE systems, TRLF achieves 85.8% higher profit than DQN and about 3× that of PPO, while interval forecasting nearly doubles cumulative return compared with point forecasts. |
| Zhou et al. [206] | The study formulates energy storage scheduling through Q-learning, where the agent observes SOC, PV generation, load demand, and electricity and gas prices and selects discrete battery charge/discharge actions. The reward is defined as the negative total operating cost, including energy, battery, and gas costs. Applied to a grid-connected integrated microgrid with PV and BESS at 1-h resolution, the RL approach achieves 61.17% faster computation than MILP with only 3.13% cost deviation, demonstrating strong real-time potential. |
| Cai et al. [207] | The work formulates energy management as an episodic RL problem in which the state includes household demand, PV generation, and electricity prices, while the actions are energy-trading decisions derived from MPC optimization. The policy is represented by a parameterized MPC controller trained using least-squares Deterministic Policy Gradient to minimize cumulative monthly cost without discounting. A centralized multi-agent framework coordinates prosumers, while Shapley values ensure fair cost allocation. On a residential microgrid with real data, the method reduces collective electricity cost relative to standard MPC and enables efficient cooperative trading. |
| Monfaredi et al. [208] | This study employs a multi-agent DDPG framework with SDAE-based state representation. Agents representing DERs, storage units, and grid interfaces observe high-dimensional states, including generation, load, storage levels, and network constraints, and output continuous scheduling actions. The reward jointly minimizes operating cost and emissions while satisfying system constraints under centralized training and decentralized execution. Applied to a grid-connected multi-energy microgrid with electric and gas networks, the method achieves about 2.9–4.25% cost reduction and 5.2–7.98% emission reduction relative to APSO, DDPG, DQN, and SO/MILP baselines. |
| Binyamin et al. [209] | This paper develops a multi-agent Q-learning framework enhanced with ANN, in which each prosumer or consumer observes states such as SOC, demand, price, and PV availability and selects actions including trading, charging/discharging, and appliance scheduling. The reward captures both cost reduction and user satisfaction, leading to optimal decentralized policies for a P2P energy market. In a grid-connected residential microgrid with PV, ESS, EV loads, and smart appliances, the method achieves 9.3–16.07% higher rewards than the baselines, together with clear user-level cost savings. |
| Hedayatnia et al. [210] | The study proposes a multi-agent DRL framework based on DDPG and DDQN within a Stackelberg-game structure, where the DSO acts as leader and the MGs act as followers. Agents exchange actions and rewards based on load demand, RES output, prices, and ESS states, and learn optimal power exchange and resource dispatch through a global reward combining economic and reliability penalties. Tested on networked grid-connected microgrids over IEEE feeders, the method achieves about 13–17% lower computational burden and 17–26% faster execution than DDQN/DQN, while also improving voltage and frequency stability. |
| Abid et al. [211] | The paper introduces a hybrid multi-agent MADDPG framework embedded within a multi-objective MOAVOA loop. Agents coordinate EV charging and DER allocation using states such as load, RES output, network conditions, and EV SOC. Actions cover the placement, sizing, and operation of RES, BESS, and EV charging stations, while the reward aggregates cost, losses, emissions, and voltage stability into a Pareto-based objective. With centralized critic and decentralized actor learning on realistic distribution simulations, the method improves performance by about 85% over COMA-based baselines and increases EV SOC by up to 154%, indicating strong Pareto efficiency. |
| Benhmidouch et al. [212] | proposes a TD3-based Actor–Critic adaptive VSG controller for frequency regulation in an islanded AC microgrid with inverter-based RES. The agent observes frequency, RoCoF, and a novel frequency-direction state () and continuously adjusts virtual inertia (J) and damping coefficient (Dp), while rewards penalize frequency nadir, RoCoF, and transient frequency deviations. Evaluated on a 20 kVA grid-forming VSI microgrid, the TD3 controller achieved a 94.6% training success rate, improved frequency nadir from 49.773 Hz to 49.848 Hz, reduced RoCoF by 30.7% (−3.905 → −2.705 Hz/s), and lowered power overshoot from 6.50% to 1.99% (69.4% reduction) versus conventional VSG. |
| Raja et al. [213] | proposes a modified TD3 actor–critic DRL framework for direct duty-cycle control of a DC–DC buck converter supplying CPLs in a DC microgrid. The agent observes voltage error, voltage states, and their derivatives/integrals, outputs continuous duty-ratio actions, and receives rewards based on voltage-tracking error minimization. Validation was performed on a real-time OPAL-RT HIL testbed under load changes, parameter uncertainty, and large disturbances. Compared with DQN and DDPG, TD3 achieved the best performance, reducing settling time to 1 ms (vs. 12 ms and 8 ms), limiting overshoot to 1.2%, and achieving the lowest IAE (0.0422) and ITSE (0.0009) values. |
| Sepehrzad et al. [214] | develops a hybrid Multi-Agent SAC-based DRL and IPWFNN framework for frequency regulation in an islanded networked microgrid consisting of four interconnected MGs with PV generation, synchronous generators, BESSs, EV batteries, and loads. States include frequency deviations, load demand histories, PV generation data, and BESS operating conditions, while actions regulate generator control signals and BESS charging/discharging participation factors. The reward minimizes frequency deviations and generation costs while enforcing SOC/SOH and operational constraints. Tested on a 4-MG islanded NMG and validated through an OPAL-RT real-time platform, the method achieved >98% accuracy, 61.1% lower computation time, and 7.82% lower computational burden than ANN, fuzzy, and PID controllers while maintaining frequency stability under severe PV and load uncertainties. |
| Abouzeid et al. [215] | A DDPG actor–critic agent is used to learn PID gains by observing frequency error and its integral and optimizing a reward that penalizes squared deviations while encouraging tight regulation. The policy outputs continuous control parameters for the LFC loop of a RES-integrated microgrid with PV, WT, ESS, diesel, EV, and SMES. Training is performed offline under stochastic disturbances and cyber-attacks using replay buffers and target networks. The method achieves better frequency stability and robustness than classical LFC approaches, with faster damping and lower deviations under attack scenarios. |
| Barros et al. [216] | proposes a single-agent DQN framework for residential energy scheduling in a renewable-powered community microgrid. The agent uses appliance states, consumption classes, and production–consumption balance indicators to decide appliance operation, aiming to improve renewable utilization, reduce costs, and preserve comfort. In GridLAB simulations of a 20-house PV/wind microgrid, the method reduced peak demand, lowered household bills, and delivered faster decisions than cloud-based baselines. |
| Jia et al. [217] | proposes HSC-SAC, a single-agent actor–critic framework for the coordinated control of a grid-connected multi-energy microgrid with PV, wind, CHP, gas turbines, electrolyzers, thermal networks, HCNG infrastructure, and SVCs. Using load, RES, price, weather, and network states, the agent minimizes operating cost while enforcing safety constraints. Tested on a modified IEEE 33-bus system, the method reduced cost, removed voltage and congestion violations, and nearly eliminated unsafe exploration after training. |
| Li et al. [218] | The paper proposes a hierarchical multi-agent PPO framework for coordinating four heterogeneous grid-connected microgrids within a VPP, integrating PV, wind turbines, ESS, microturbines, and flexible loads. Using market, load, storage, and device states, the framework optimizes internal pricing, local scheduling, and centralized ESS dispatch. Tested on a real-data VPP setup, it improved reward, reduced operating costs, and achieved much faster decision-making than non-hierarchical PPO and optimization-based baselines. |
| Li et al. [219] | proposes a single-agent PPO actor–critic framework for real-time energy scheduling in a grid-connected microgrid with PV, wind, diesel generators, ESS, EVs with V2G, loads, and grid transactions. Using price, generation, demand, and storage/EV states, the agent optimizes dispatch and charging decisions. Tested with real hourly data, the method outperformed TD3, MILP, PSO, and SC in profit and self-balance performance. |
| Li et al. [220] | develops a decentralized MASAC-based MARL framework for cooperative energy scheduling in grid-connected residential microgrid clusters with PV, ESS, EVs, HVAC systems, loads, and grid interaction. Using real Pecan Street data, the method coordinates storage, EV charging, and HVAC control while reducing energy cost, discomfort, range anxiety, and transformer stress. Compared with MPC and other MARL baselines, it achieved the best cumulative reward and reduced peak transformer demand by 22.6%. |
| Ref. | Year | Method | Agent | FH/TS | Baselines | RES | Integration | Type | Data | Cit. |
|---|---|---|---|---|---|---|---|---|---|---|
| [221] | 2020 | SAC | Multi | 24 h/1 h | RL/RBC | PV | DG/TSS/HVAC/Load/Grid | District | Real | 67 |
| [222] | 2020 | Ql | Multi | 24 h/1 h | GA | PV | DG/HVAC/EVCS/Load/Grid | Resident | Real | 283 |
| [223] | 2021 | DDPG | Single | N/A | RBC | PV | DG/HVAC/ESS/Grid | Lab | Real * | 90 |
| [224] | 2021 | DQN | Single | -/5 m | RBC | PV | DG/HVAC/TSS/Load/Grid | Resident | Real | 132 |
| [225] | 2021 | SAC | Single | 92 d/1 h | RBC | PV | DG/HVACs/TSS/Load/Grid | District | Synth | 112 |
| [226] | 2021 | SAC | Single | -/5 m | RBC | PV | DG/HVACs/TSS/Load/Grid | District | Synth | 88 |
| [227] | 2022 | TD3 | Single | 7 d/1 h | RL/RBC | PV/BIO | DG/ESS/Grid | Industrial | Real | 61 |
| [228] | 2022 | DQN | Single | N/A | RBC | PV/SWH | DG/TSS/HVAC/Grid | Resident | Real | 50 |
| [229] | 2022 | DQN/DDPG | Single | 24 h/1 h | RL/RBC/MILP | PV | DG/ESS/HVAC/Load/Grid | Resident | Real | 73 |
| [230] | 2022 | DDPG | Single | 1 y/1 h | RBC/MPC | PV | DG/ESS/TSS/HVAC/Grid | Resident | Real | 54 |
| [231] | 2022 | A2C | Multi | 24 h/1 h | RL/RBC/MILP | PV | DG/ESS/HVAC/Load/Grid | Residents | Synth | 211 |
| [232] | 2022 | D3QN | Multi | 3 d/1 m | RBC | PV/WT | DG/ESS/HVAC/Grid | Office | Synth | 120 |
| [233] | 2022 | Ql | Multi | 92 d/1 h | RBC | PV | DG/ESS/EVCS/Load/Grid | Community | Real | 59 |
| [234] | 2023 | DDPG | Multi | 24 h/1 h | RL/ADMM | PV | DG/ESS/TSS/HVAC/Load/Grid | Community | Real | 90 |
| [235] | 2023 | SAC | Multi | 1 y/1 h | RL/RBC | PV | DG/ESS/TSS/HVAC/Load/Grid | District | Real | 80 |
| [236] | 2023 | DDPG/RBC | Single | -/1 h | RBC | PV | DG/ESS/Grid | Resident | Real | 54 |
| [237] | 2024 | SAC | Single | N/A | RBC | PV | DG/ESS/EV/Load | Resident | Real | 60 |
| [238] | 2024 | PPO/IL | Multi | N/A | RBC | PV | DG/ESS/HVAC/Load/Grid | Resident | Real | 60 |
| [239] | 2024 | SAC | Multi | -/1 m | RL/RBC | PV | DG/ESS/HVAC/Load/Grid | Greenhouse | Real | 69 |
| [240] | 2024 | PPO | Single | 24 h/15 m | RL/RBC/MILP | PV | DG/ESS/EV/HVAC/Load/Grid | Resident | Real | 50 |
| [241] | 2024 | PPO | Single | 24 h/1 h | RL/MILP | PV | DG/ESS/Load/Grid | Resident | Real * | 64 |
| [242] | 2024 | DQN | Multi | 24 h/1 h | Heuristic | PV | EV/Load/Grid | Resident | Real | 57 |
| [243] | 2025 | PPO | Single | 24 h/1 h | RL/MILP | PV | ESS/DHW/HVAC/Load/Grid | Resident | Real * | 25 |
| [244] | 2025 | Unspecified | Single | 24 h/1 h | Static/Conv | PV/WT | DG/ESS/EV/Load/Grid | Community | Real | 25 |
| [245] | 2025 | PPO | Multi | 7 d/1 h | RL/MILP | PV | DG/ESS/EV/Load/Grid | Community | Real | 25 |
| Author | Summary |
|---|---|
| Vazquez et al. [221] | This study proposes MARLISA, a multi-agent SAC framework in which each building agent observes weather conditions, loads, storage SoC, and PV generation, and outputs continuous storage-control actions. The reward combines building-level and district-level objectives, while coordination is achieved through sequential action selection and shared demand predictions. Applied to PV-equipped buildings with thermal storage in CityLearn, the method achieves approximately 15% peak reduction, 35% ramping reduction, and 10% load-factor improvement relative to RBC, while converging within 1–2 years. |
| Xu et al. [222] | This work applies multi-agent tabular Q-learning to a residential HEMS with PV, HVAC, appliances, and EVs. Each agent observes forecasted electricity prices and PV generation and selects discrete appliance actions such as ON/OFF states, power levels, and EV charging rates. The reward balances electricity cost and user discomfort. Using real PJM data, the method reduces cost by about 44.6% compared with the absence of demand response and clearly outperforms GA with nearly 40× faster computation. |
| Touzani et al. [223] | The study adopts a DDPG actor–critic framework in which the state includes indoor temperature, PV generation, battery state, and electricity prices, while the actions control HVAC setpoints and battery dispatch. The reward penalizes energy cost, comfort violations, and battery-limit violations. Implemented in the real commercial FLEXLAB testbed, the approach improves operational efficiency and achieves up to 39.6% cost reduction compared with RBC while maintaining similar comfort levels. |
| Lissa et al. [224] | This work employs a DQN agent for a residential HEMS, with states including indoor and outdoor temperature, DHW temperature, PV production, and time, and actions corresponding to heat-pump operating modes. The reward penalizes comfort deviations and grid electricity use while encouraging PV-based operation. Compared with RBC, the method achieves average savings of about 8%, with gains reaching up to 16% in some cases, together with 9.5% higher PV utilization and 10% load shifting, while keeping comfort violations below 1%. |
| Pinto et al. [225] | A centralized SAC agent maps a high-dimensional state space, including weather forecasts, prices, SoC, indoor temperature, and loads, to continuous control actions for heat-pump output and storage charging/discharging. The reward combines comfort penalties, storage incentives, and peak-shaving objectives. Trained in CityLearn with LSTM-based building dynamics, the framework coordinates four commercial buildings and achieves 23% peak reduction, 20% PAR reduction, and about 3% cost savings relative to RBC while maintaining near-comfort conditions. |
| Pinto et al. [226] | This study develops a centralized SAC-based controller for a commercial-building cluster in CityLearn. The agent observes weather forecasts, prices, loads, PV generation, and storage SoCs, and outputs continuous actions for eight thermal storage units. The reward promotes both cost minimization and peak reduction. The method learns to shift cooling demand using COP–temperature dynamics and price signals, achieving about 4% cost reduction and up to 12% peak reduction relative to RBC, together with improved PAR and robustness across climates. |
| Gao et al. [227] | The work formulates building energy management as a continuous-action MDP in which a TD3 agent controls power exchange among PV, bioenergy, BESS, and the grid. The state includes time, battery SoC, and generation/load conditions, while the reward penalizes off-grid deviation and unsafe SoC levels. In a simulated renewable building energy system, TD3 improves off-grid deviation and grid-export reduction by about 70–80% compared with DDPG and RBC, while maintaining nearly zero battery safety violations. |
| Heidari et al. [228] | This study uses a DQN agent for residential HVAC and DHW control. The state includes indoor temperature, tank temperature, PV production, weather conditions, and hot-water demand, while the actions correspond to discrete heating decisions. The reward minimizes energy use while enforcing comfort and hygiene constraints, including Legionella-risk considerations. Using a two-stage offline–online training procedure with real residential data, the method achieves 7–60% and 28–75% energy savings relative to rule-based baselines while preserving comfort and safety. |
| Huang et al. [229] | The paper models HEMS as an MDP in which the state includes time, SoC, indoor and outdoor temperature, PV/load history, and appliance states. Actions combine discrete appliance scheduling with continuous HVAC/BESS control. A hybrid MDRL framework uses DQN for discrete actions and DDPG for continuous ones. Applied to a grid-connected residential building with PV, BESS, and HVAC, the method achieves 25.8% cost reduction relative to RBC and about 7% relative to DDPG, while the safe-MDRL variant reduces comfort violations by roughly 80%. |
| Langer et al. [230] | A DDPG-based policy is trained on an MDP whose state includes battery and thermal SoC, demand, PV generation, time features, and temperature. The actions define target SoCs for battery and thermal storage, thereby indirectly controlling the heat pump and energy flows. The reward balances electricity cost and comfort violations while encouraging PV self-consumption and load shifting. Tested on a residential SHEMS with real data, the method reaches about 75% self-sufficiency, clearly outperforms RBC, and approaches MPC performance without requiring forecasts. |
| Lee et al. [231] | The study formulates a distributed MDP in which each home employs three A2C agents for AC, washing machine, and ESS control. States include electricity price, temperature, PV generation, and ESS SOE, while actions schedule appliance energy use. Rewards combine cost minimization with comfort penalties. A federated learning loop aggregates local models through FedSGD without sharing raw data. The approach achieves about 15–35% cost reduction compared with MILP and BEopt, converges faster than standalone DRL, and scales effectively to a larger number of households. |
| Shen et al. [232] | This work formulates building energy control as a multi-agent MDP in which HVAC-zone and battery agents observe local temperatures, RES generation, SoC, and prices, and select discrete airflow or charge/discharge actions. A D3QN framework with PER, FAS, and VDN improves sample efficiency, feasibility, and cooperation. The reward combines energy cost, RES curtailment, and comfort violations. In a simulated commercial building, the method improves comfort by 84%, RES utilization by 43%, and reduces cost by 8% relative to RBC. |
| Lai et al. [233] | The paper proposes a bi-level MARL Stackelberg framework in which a community aggregator uses Q-learning for ESS dispatch, while multiple appliance-level agents apply sequential Q-learning for scheduling. The state includes time, demand, and storage conditions, and the reward combines electricity cost with user dissatisfaction. PV uncertainty is handled through LSTM forecasting. In a residential community, the method achieves 30.8% peak reduction and 9.68% cost savings relative to single-agent RL, as well as 2.7% lower cost than MINLP under uncertainty. |
| Qiu et al. [234] | The study develops a multi-agent DDPG framework with a federated centralized critic. Each building observes local prices, loads, PV generation, temperatures, and ESS/TES states, and outputs continuous HVAC, storage, and trading actions. The reward combines electricity, gas, and carbon costs. A key contribution is the abstracted centralized critic, which relies on community net demand and emissions to preserve privacy while stabilizing learning. In a three-building community MES, the proposed Fed-JPC policy achieves about 6–8% cost reduction compared with MARL baselines and remains within 1.88% of the centralized optimum, with about 100× faster computation. |
| Xie et al. [235] | This work formulates a multi-agent SAC control problem in which each building agent observes local thermal states, PV generation, ESS SoC, prices, weather, and shared building signals, and outputs continuous HVAC and battery actions. The reward combines electricity cost with a global net-load penalty to encourage cooperation. In a district-scale CityLearn testbed with nine mixed residential and commercial buildings, independent SAC does not outperform RBC, whereas the proposed MAAC achieves more than 8% net-load reduction and stronger peak shaving, demonstrating the value of coordinated MARL. |
| Zhou et al. [236] | A DDPG actor–critic agent is combined with a rule-based energy system (RBES) to control continuous battery charge/discharge actions. The state includes PV generation, building demand, electricity prices, and time features, while the reward promotes lower electricity cost and improved PV utilization. The RBES enforces physical heuristics, and RL refines grid–battery interaction. Applied to a real-data PV-equipped building, the method reduces cost by 4.6–7.0% and improves self-consumption by up to 10.6% relative to rule-based baselines. |
| Deng et al. [237] | The study formulates a continuous-control MDP in which a SAC agent observes PV generation, load demand, battery SoC, EV SoC, and power mismatch, and outputs a continuous gain coefficient to regulate DC-bus voltage and indirectly coordinate loads and storage. The reward penalizes power mismatch and user dissatisfaction. In a residential DC microgrid, the method increases PV self-sufficiency by about 15–20% and PV absorption by about 18% compared with MPC and RBC, while also reducing mismatch and discomfort. |
| Wang et al. [238] | This work develops a multi-agent MAPPO framework enhanced with imitation learning. Agents control HVAC and battery dispatch using states that include PV generation, load demand, SoC, TOU prices, and time features. The shared reward balances energy-cost reduction and thermal comfort, thereby encouraging cooperation. Applied to a residential zero-energy home, the method improves PV utilization and storage scheduling. Compared with PI control, it increases self-sufficiency by 34.86–46.10%, self-consumption by 15.78–18.47%, and improves temperature regulation by 1.33 °C toward the setpoint, while converging rapidly in about 50 episodes. |
| Ajagekar et al. [239] | Present an attention-based multi-agent SAC framework for five grid-interactive greenhouses with PV and BESS. Each agent observes climate, weather, PV output, battery SOC, demand, and time variables, and controls battery charging and discharging to reduce net grid demand. Using simulated NYC greenhouses with real NSRDB climate data, the method achieved at least 28% lower grid demand than RBC, with stronger gains in winter, summer, and fall. |
| Aldahmashi et al. [240] | Propose a PPO-based DRL framework for a residential HEMS coordinating PV, ESS, EV, HVAC, and flexible appliances, while jointly managing active and reactive power. The agent uses price, PV, temperature, ESS/EV states, appliance status, and power factor to optimize scheduling, comfort, and power quality. Tested on a simulated smart home with real Australian PV and weather data, PPO reduced electricity cost by 31.5%, outperforming DQN and DDPG, while also improving the average power factor. |
| Kang et al. [241] | Develop a PPO-based RL framework for scheduling a grid-connected residential PV-BESS system. The agent observes PV generation, electricity demand, battery SOC, month, and hour, and controls continuous battery charging and discharging. The reward jointly maximizes self-sufficiency and reduces daily peak load. Using real residential data from South Korea, PPO showed the best stability among A2C, TD3, and SAC, and performed close to the MILP optimum. |
| Kumari et al. [242] | Propose a multi-agent DQN-based residential HEMS in which separate agents manage non-shiftable, shiftable, and controllable household loads, including EVs and PV-equipped homes. Using appliance-demand profiles and hourly electricity prices, the framework learns scheduling decisions that reduce energy use, optimize cost, and increase profit. Tested with real OpenEI load data and PJM prices, it achieved 86% peak-hour energy savings, outperforming the heuristic benchmark (79%) and approaching the ideal benchmark (92%). |
| Wang et al. [243] | The paper develops a PPO-based DRL framework for resilient energy scheduling in a grid-connected residential multi-energy building connected to a low-temperature district heating network, with PV, battery and thermal storage, heat pump, DHW, space heating, and grid interaction. Using real data from Copenhagen, the method reduced daily operating cost from £136.2 to £122.3 compared with DDPG and cut computation time from 8.86 s (MILP) to about 15 ms, showing strong potential for real-time control. |
| Lami et al. [244] | proposes a single-agent DRL framework for energy management in a grid-connected community building microgrid with PV, wind, battery storage, EVs with V2H, P2P trading, loads, and grid interaction. Using real hourly Saudi Arabian data, the method reduced energy costs by 23%, lowered grid dependence by 40%, achieved 85% renewable utilization, and improved overall energy efficiency from 85% to 92%. |
| Talihati et al. [245] | develops a PPO-based MARL framework for a residential community energy system with PV, shared battery storage, EV charging, and grid interaction. Three agents coordinate battery dispatch, EV charging, and ES-PV pricing using community load, PV output, prices, battery SOC, and EV demand. Tested with real Australian data on an IEEE-14 residential network, the method increased PV self-consumption by 66.41%, covered 38.68% of EV demand, reduced community electricity costs by 7.73%, and generated €51,924.65 profit for the ES-PV operator. |
| RL Family | Best Suited For | Strengths | Limitations | Representative Cases |
|---|---|---|---|---|
| QL/DQN | Discrete scheduling: appliances, OLTCs/capacitors, BESS modes, finite EMS actions | Simple; interpretable; natural for finite action spaces | Weak for continuous control; limited scalability | [165,201,206,224,232,233] |
| PPO/A2C | Stable EMS, OPF approximation, BESS–EV–HVAC scheduling, resilient operation | Stable convergence; reliable policy updates | Lower sample efficiency; longer training | [162,200,241,243] |
| DDPG/TD3 | Continuous physical control: PV inverters, SVCs, BESS, M-SOPs, HVAC, EVs, LFC | Handles continuous actions; TD3 improves stability | Sensitive training; unsafe exploration without constraints | [164,175,177,223,227] |
| SAC/MASAC | Stochastic RES-rich EMS, storage coordination, building clusters, voltage/frequency control | Strong exploration; robust under uncertainty | Computationally heavier; reward-sensitive | [203,204,217,221,235,237] |
| MARL | Distributed PV inverters, prosumers, EVs, buildings, greenhouses, microgrid clusters | Scalable coordination, decentralized execution | Non-stationarity, coordination/privacy complexity | [163,168,183,184,220,239,245] |
| Safe RL | Voltage, frequency, OPF, SOC, line-flow, converter, comfort, security-constrained control | Safer operation, better deployability | More complex, may be conservative | [164,171,174,217,229] |
| Hybrid RL | OPF acceleration, two-timescale VVC, smart building flexibility, VPP scheduling, multi-energy MGs | Adaptive and feasible, closer to deployment | Requires models, solvers, forecasts, or surrogates | [162,173,184,189,190,194,203,207,218] |
| Mechanism | Main Challenge Addressed | Representative Papers | Main Advantage |
|---|---|---|---|
| CTDE | Non-stationarity | [163,167,204,234] | Stable training with local execution |
| Attention critic | Scalability of joint critics | [168,173,183,235,239] | Selects relevant agent interactions |
| GNN/GAT critic | Large network topology | [184,189] | Learns grid-aware coupling structure |
| Federated MARL | Privacy and data sharing | [231,234] | Coordinates without raw data exchange |
| Hierarchical MARL | Multi-layer control complexity | [177,210,233,245] | Separates aggregator and device roles |
| Mean-field/collective strategy | Massive agent populations | [220] | Replaces all-to-all interaction with aggregate behavior |
| Explicit communication | Local observability limits | [180,195,214] | Exchanges selected neighbor/system information |
| Fully decentralized MARL | Communication constraints | [174,175,192,209] | Avoids centralized critics and full communication |
| Type | Mechanism | Representative Papers | Overhead | Safety |
|---|---|---|---|---|
| Soft penalty reward | Linear/weighted penalties | [163,164,165,167,200,223] | Low-Medium | Low-Medium |
| Soft nonlinear reward | Quadratic/exponential/barrier penalties | [166,170,183,199,227] | Medium | Medium |
| Soft risk-aware reward | CVaR/reliability/mismatch penalties | [197,208,210,211] | Medium | Medium |
| Hard constrained RL | CMDP/CPO/Lagrangian RL | [171] | Medium-High | High |
| Hard safety layer | Action filtering/correction | [164,204,229] | Medium | High |
| Hard action-space design | Action masking/physical parameterization | [177,180,204] | Low-Medium | Medium-High |
| Hard architecture design | Stability/physics-informed networks | [174,184] | Medium-High | Medium-High |
| Hybrid hard constraint | OPF/MPC/SOCP correction | [165,173,190,194,203,207] | High | High-Extreme |
| Simulator | Short Description | Power Grid Cases | Microgrid Cases | Building Cases | Total |
|---|---|---|---|---|---|
| Custom environments | MATLAB, Simulink, Python, TensorFlow/PyTorch, Keras, PSCAD, TRNSYS, HOMER, or other custom simulators. | [163,164,165,166,167,168,169,171,172,173,175,176,177,178,179,182,186,187,188,189] | [191,192,193,194,196,197,200,201,202,203,204,205,206,208,209,210,211,212,215,217,218,219,220] | [198,209,222,224,228,229,231,232,233,234,236,237,240,241,243,244] | 59 |
| Open-source solvers | Reusable solvers such as PYPOWER, PandaPower, OpenDSS, GridLAB-D, OMNeT++, or PGSim. | [162,170,180,181,183,184,185] | [195,199,216] | [245] | 11 |
| OpenAI Gym | Gym-compatible or benchmark-style RL environments, including CityLearn and FLEXLAB Gym-style platforms. | [190] | – | [221,223,225,226,227,235,238,239,242] | 10 |
| HIL | OPAL-RT, RT-LAB/HYPERSIM, real testbed, or real-field validation beyond offline simulation. | [187] | [213,214] | [223] | 4 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Michailidis, P.; Minelli, F.; Coban, H.H.; Michailidis, I.; Kosmatopoulos, E. Reinforcement Learning for Optimizing Renewable Energy Utilization in Smart Grids: Recent Advances in Power Grids, Microgrids, and Building Energy Systems. Infrastructures 2026, 11, 240. https://doi.org/10.3390/infrastructures11070240
Michailidis P, Minelli F, Coban HH, Michailidis I, Kosmatopoulos E. Reinforcement Learning for Optimizing Renewable Energy Utilization in Smart Grids: Recent Advances in Power Grids, Microgrids, and Building Energy Systems. Infrastructures. 2026; 11(7):240. https://doi.org/10.3390/infrastructures11070240
Chicago/Turabian StyleMichailidis, Panagiotis, Federico Minelli, Hasan Huseyin Coban, Iakovos Michailidis, and Elias Kosmatopoulos. 2026. "Reinforcement Learning for Optimizing Renewable Energy Utilization in Smart Grids: Recent Advances in Power Grids, Microgrids, and Building Energy Systems" Infrastructures 11, no. 7: 240. https://doi.org/10.3390/infrastructures11070240
APA StyleMichailidis, P., Minelli, F., Coban, H. H., Michailidis, I., & Kosmatopoulos, E. (2026). Reinforcement Learning for Optimizing Renewable Energy Utilization in Smart Grids: Recent Advances in Power Grids, Microgrids, and Building Energy Systems. Infrastructures, 11(7), 240. https://doi.org/10.3390/infrastructures11070240

