Next Article in Journal
Data-Driven ANN Model Development for Maximum Power Point Estimation in PV Panel Under Partial Shading Conditions
Previous Article in Journal
Physics-Informed Deep Reinforcement Learning for Compact VBT Farms: Integration, Power Quality, and Economics
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Proceeding Paper

Physics-Constrained Multi-Agent Deep Reinforcement Learning for Real-Time Energy Management of a Saharan Hybrid Microgrid †

by
Redouane Mihramane
*,
S. Salah Ech-Charqaouy
,
Abdelkader Boulezhar
,
Amjad Ech-Charqaouy
and
Nizar Ech-Charqaouy
Faculty of Sciences Ain Chock, Hassan II University, Casablanca 20100, Morocco
*
Author to whom correspondence should be addressed.
Presented at the 2nd International Conference on Sciences and Techniques for Renewable Energy and the Environment, Al Hoceima, Morocco, 28–30 April 2026.
Eng. Proc. 2026, 144(1), 9; https://doi.org/10.3390/engproc2026144009
Published: 25 June 2026

Abstract

This paper addresses the challenge of ensuring physically feasible and reliable real-time control of hybrid microgrids in harsh desert environments. A physics-constrained multi-agent Deep Q-Network (MA-DQN) is proposed for energy management of a grid-interactive microgrid in the Moroccan Sahara. The method embeds operational constraints directly into learning through action filtering, penalty-aware rewards, and coordinated PCC control. The results show a reduction in operational cost from 1250 MAD to 1120 MAD (−10.4%) and CO2 emissions from 318.9 kg to 272.5 kg (−14.6%), while maintaining voltage within ±10% limits and eliminating PCC oscillations. The framework delivers stable, reliable, and deployment-ready control.

1. Introduction

The increasing penetration of renewable energy sources in power systems has led to the widespread development of hybrid microgrids as a promising solution for ensuring reliable and sustainable electricity supply, particularly in remote and isolated regions [1,2,3,4]. In this context, desert environments, such as those found in the Moroccan Sahara, offer significant potential for renewable energy exploitation due to high solar irradiation and favorable wind conditions. However, the operation of such systems remains highly challenging due to harsh climatic conditions, resource intermittency, and the intrinsic fragility of distribution networks.
Hybrid microgrids integrating photovoltaic (PV), wind, diesel generators, and battery storage systems have been widely studied for off-grid and weak-grid applications [5,6,7]. These systems must ensure a continuous balance between generation and demand while maintaining operational constraints such as voltage stability, power flow limits, and battery state-of-charge (SOC) management. In practice, maintaining voltage within acceptable limits (typically ±10%) and ensuring stable operation at the point of common coupling (PCC) are critical requirements for system reliability and service continuity [8,9].
Traditional energy management approaches, including rule-based strategies and optimization techniques such as particle swarm optimization (PSO), have been extensively applied to microgrid control problems [10,11]. Previous work has demonstrated the effectiveness of PSO for offline optimization of hybrid microgrids, providing satisfactory solutions in terms of cost and emission reduction under predefined operating conditions [11,12]. However, such approaches generally operate in an offline or semi-static framework and lack the adaptability required to cope with highly dynamic and uncertain environments. Moreover, they often fail to capture the complex interactions between distributed energy resources and the network, particularly under real-time operating conditions.
To overcome these limitations, reinforcement learning-based approaches have been progressively introduced [13,14]. In particular, a first Deep Q-Network (DQN)-based strategy has shown improved adaptability and enhanced operational performance, including reductions in operational cost and CO2 emissions, while enabling real-time decision-making capabilities [15,16]. Nevertheless, this class of approaches remains limited when dealing with large-scale systems involving multiple interacting components and strict operational constraints.
In recent years, deep reinforcement learning (DRL) has emerged as a powerful paradigm for sequential decision-making problems in energy systems. Several studies have demonstrated the potential of DRL for real-time energy management in microgrids, enabling adaptive control strategies that respond to stochastic variations in load demand and renewable generation [15,16]. Furthermore, the extension toward multi-agent reinforcement learning (MARL) allows for decentralized and scalable control architectures, where multiple agents coordinate to manage distributed energy resources efficiently [17,18].
Despite these advances, a critical limitation remains: most DRL-based approaches do not explicitly enforce physical and operational constraints during the learning process. As a result, the learned policies may generate infeasible or unsafe control actions, leading to voltage violations, instability at the PCC, or excessive battery degradation. This issue has been widely recognized in the broader field of safe reinforcement learning, where the need to incorporate constraints directly into policy learning has been emphasized [18,19]. However, the integration of such concepts into real-world microgrid applications remains limited.
To address these challenges, this paper proposes a physics-constrained multi-agent Deep Q-Network (MA-DQN) framework for the real-time energy management of a grid-interactive hybrid microgrid located in the Moroccan Sahara (Boujdour). The proposed approach embeds physical constraints directly into the learning process through a constrained action space, penalty-aware reward formulation, and coordinated control of the PCC. This design ensures compliance with key operational limits, including SOC bounds, voltage regulation, and exclusive import/export behavior at the grid interface [15,16,19].
Unlike conventional DRL approaches that rely on post-processing or soft constraint enforcement, the proposed framework guarantees physically consistent and deployment-ready control policies. By explicitly integrating engineering constraints into the learning architecture, the method bridges the gap between theoretical DRL models and practical microgrid operation. The effectiveness of the proposed approach is validated through a realistic case study based on a radial low-voltage microgrid, demonstrating improved stability, robustness, and operational reliability under dynamic conditions [12,15,16].

2. Microgrid System and Problem Formulation

2.1. Microgrid Architecture

The considered system is a grid-interactive hybrid microgrid located in a Saharan environment, designed to ensure reliable power supply under harsh climatic conditions and variable renewable energy availability [1,2,3]. The microgrid integrates multiple distributed energy resources, including photovoltaic (PV) generation, a wind turbine, a diesel generator, and several battery energy storage systems (BESS), all interconnected through a radial low-voltage distribution network [5,6].
The architecture is organized around a point of common coupling (PCC), which enables bidirectional power exchange with the main grid. This interface plays a critical role in maintaining system stability and ensuring operational flexibility, particularly under supply–demand imbalance conditions [20,21]. The radial structure reflects typical configurations encountered in remote or weak-grid areas, where simplicity and robustness are essential design criteria [4].
Each component of the microgrid fulfills a specific operational role. Renewable sources (PV and wind) provide primary energy generation but are inherently intermittent. The diesel generator ensures backup supply and enhances system reliability during periods of low renewable production. Battery energy storage systems act as dynamic buffers, mitigating power fluctuations, supporting load demand, and maintaining system stability through controlled charge–discharge cycles [5,22]. The coordinated interaction among these elements is essential to ensure power balance and compliance with network constraints.
A critical operational feature of the system is the PCC constraint, which enforces exclusive import or export at any given time. This requirement prevents oscillatory power exchanges and guarantees stable interaction with the upstream network, especially under real-time control conditions [21].
Figure 1 illustrates the architecture of the grid-interactive Saharan hybrid microgrid.

2.2. Mathematical Formulation

The operation of the microgrid is governed by a set of physical and operational constraints that must be satisfied at each time step. The fundamental requirement is the real-time power balance between generation, storage, and demand:
P P V ( t ) + P W T ( t ) + P D G ( t ) + P B S S ( t ) + P g r i d ( t ) = P l o a d ( t )
where P P V , P W T , P D G , P B S S , and P g r i d denote the power contributions of photovoltaic generation, wind turbine, diesel generator, battery storage systems, and grid exchange, respectively.

2.3. Battery State-of-Charge Constraint

The battery dynamics are constrained by state-of-charge (SOC) limits to ensure safe and sustainable operation:
S O C m i n S O C ( t ) S O C m a x
where SOC(t) denotes the state of charge of the battery at time step t, S O C m i n and S O C m a x represent the minimum and maximum allowable SOC limits, respectively, which define the safe operating range of the battery.
This constraint prevents overcharging and deep discharging, thereby preserving battery lifetime and ensuring operational reliability [5].

2.4. Voltage Constraint

Voltage levels across the distribution network must remain within acceptable limits to ensure power quality and equipment safety:
V m i n V i ( t ) V m a x
where V i ( t ) denotes the voltage magnitude at node i at time step t, V m i n and V m a x represent the minimum and maximum allowable voltage limits, respectively, and V n o m is the nominal voltage of the distribution network.
The voltage limits are defined as:
With V m i n = 0.9 V n o m , V m a x = 1.1 V n o m
Maintaining voltage within ±10% of the nominal value is a standard requirement in distribution systems and is particularly critical in radial microgrids with high renewable penetration [4,8,9].

2.5. PCC Operational Constraint

To ensure stable interaction with the main grid, the power exchange at the PCC must satisfy an exclusivity condition:
P g r i d i m p o r t ( t ) · P g r i d e x p o r t ( t ) = 0
where P g r i d i m p o r t ( t ) denotes the power imported from the main grid at time step t , and P g r i d e x p o r t ( t ) represents the power exported to the main grid at the same time step.
This constraint guarantees that the microgrid cannot simultaneously import and export power, thereby eliminating oscillatory behavior and ensuring consistent grid operation [21].

2.6. Objective Function

m i n ( C o p + λ C O 2 )
where C o p represents the operational cost, including fuel consumption and grid energy exchange, while C O 2 denotes the associated carbon emissions. The weighting factor λ enables a trade-off between economic performance and environmental impact [10,11].
To provide a clear and structured overview of the operational constraints governing the microgrid, Table 1 summarizes the main physical and operational requirements considered in the proposed framework. These constraints ensure real-time power balance, safe battery operation, voltage regulation within acceptable limits, and stable interaction with the main grid.

3. PSO-Based Baseline

Particle Swarm Optimization (PSO) is adopted in this study as a reference optimization method for microgrid energy management. PSO is a population-based metaheuristic algorithm widely used for solving nonlinear and multi-objective optimization problems in power systems due to its simplicity and fast convergence characteristics [10,11]. In hybrid microgrids, PSO has been extensively applied to determine optimal dispatch strategies that minimize operational costs while satisfying system constraints such as power balance, battery state-of-charge limits, and voltage stability requirements [5,10].
In previous work, PSO has demonstrated its effectiveness for the offline optimization of hybrid microgrids, achieving significant reductions in operational cost and carbon emissions under predefined operating scenarios [11,12]. However, this approach operates within an offline optimization framework, where control actions are computed based on known load and generation profiles over a fixed time horizon. As a result, it lacks adaptability to real-time variations and cannot effectively handle the dynamic uncertainties inherent to renewable energy systems [23,24].
Moreover, the integration of physical and operational constraints within PSO is typically handled through penalty-based formulations or post-processing mechanisms. Such approaches do not guarantee constraint satisfaction at every decision step and may lead to suboptimal or infeasible solutions, particularly under rapidly changing operating conditions [10,11]. These limitations motivate the transition toward learning-based approaches capable of real-time adaptation and explicit constraint handling, as highlighted in recent studies on advanced energy management systems [13,14].
Table 2 presents a comparative overview between the PSO-based approach and the proposed DRL-based framework, highlighting key differences in terms of adaptability, constraint handling, and computational performance.

4. Proposed MA-DQN Framework

4.1. RL Formulation

The energy management problem is formulated as a sequential decision-making process, where optimal control actions must be determined at each time step under dynamic operating conditions. This problem is modeled using a Deep Q-Network (DQN) framework, in which the microgrid is represented as an environment interacting with a learning agent [15,16].
The system state s(t) captures the essential information required for decision-making, including load demand, renewable generation levels, battery state-of-charge (SOC), and nodal voltage profiles:
s ( t ) = [ P l o a d ( t ) , P P V ( t ) , P W T ( t ) , S O C ( t ) , V i ( t ) ]
where s ( t ) denotes the state vector of the microgrid at time step t, P l o a d ( t ) represents the total load demand, P P V ( t ) and P W T ( t ) denote the power generated by the photovoltaic system and the wind turbine, respectively, S O C ( t ) is the state of charge of the battery energy storage system, and V i ( t ) represents the voltage magnitude at node i of the distribution network.
This state representation enables the agent to simultaneously perceive energy balance conditions and network constraints in real time, ensuring that both operational and physical aspects of the system are considered [24].
The action space a ( t ) consists of dispatch decisions applied to controllable components, including diesel generator output, battery charging/discharging power, and power exchange at the PCC. These actions directly influence system dynamics and must satisfy operational constraints.
The objective is to learn an optimal policy that maximizes the expected cumulative reward. The reward function is designed to reflect both economic and operational criteria, including cost minimization, emission reduction, and strict adherence to system constraints. The Q-function is updated iteratively according to the standard DQN formulation [15].

4.2. Multi-Agent Architecture

To address the distributed and heterogeneous nature of the microgrid, a multi-agent reinforcement learning (MARL) architecture is adopted. In this framework, each energy resource is controlled by an independent agent, enabling decentralized decision-making while maintaining coordinated system performance [17,18].
The proposed architecture includes a photovoltaic (PV) agent responsible for managing solar generation, a wind agent dedicated to handling wind energy integration, a battery agent controlling charge and discharge cycles, a diesel agent ensuring backup generation and system reliability, and a PCC coordinator agent regulating power exchange with the main grid.
This distributed structure enhances scalability and flexibility, as each agent operates based on local observations while contributing to global system objectives. Such coordination is particularly effective for handling complex interactions between heterogeneous energy sources [17].
Figure 2 illustrates the multi-agent deep reinforcement learning architecture for microgrid control.

4.3. Centralized Training and Decentralized Execution (CTDE)

The proposed MA-DQN framework follows a centralized training and decentralized execution (CTDE) paradigm. During training, agents have access to global system information, enabling them to learn coordinated policies that account for interdependencies between microgrid components.
Once training is completed, each agent operates independently using only local observations, allowing real-time implementation without requiring full system observability. This ensures both scalability and practical deployability.
The CTDE paradigm provides an effective trade-off between coordination and decentralization, making it particularly suitable for complex energy systems operating under uncertainty [17,18].

4.4. Constraint Handling (Key Contribution)

A key contribution of this work lies in the explicit integration of physical and operational constraints into the reinforcement learning process. Unlike conventional DRL approaches that rely on post-processing or soft penalties, the proposed framework enforces constraint compliance directly within the learning architecture [19].
First, a constraint-aware action filtering mechanism is introduced to ensure that only physically feasible actions are considered. Actions violating power balance, SOC limits, or PCC operational rules are removed from the action space prior to execution.
Second, the reward function incorporates penalty terms associated with constraint violations, including voltage deviations beyond acceptable limits, battery overcharge or deep discharge, and violations of PCC import/export exclusivity. These penalties guide the learning process toward safe and stable operating policies, consistent with recent developments in safe reinforcement learning for energy systems [19].
Third, the PCC behavior is explicitly modeled to eliminate oscillatory power exchanges. The exclusivity constraint ensures that import and export actions cannot occur simultaneously, thereby improving grid interaction stability.
This physics-constrained learning strategy aligns with recent advances in safe reinforcement learning, where constraint satisfaction is embedded directly into policy optimization [18,19]. By integrating engineering constraints into the decision-making process, the proposed framework ensures physically consistent, robust, and deployment-ready control policies.

5. Results and Discussion

5.1. Performance Comparison

The performance of the proposed MA-DQN framework is evaluated against the PSO-based baseline in terms of economic cost, environmental impact, and operational stability.
The results show that the MA-DQN approach achieves a significant reduction in operational cost compared to PSO. Specifically, the operational cost decreases from 1250 MAD (PSO) to 1120 MAD (MA-DQN), corresponding to a reduction of approximately 10.4%. This improvement is mainly attributed to the ability of the MA-DQN to adapt dynamically to variations in load demand and renewable generation, in contrast to the offline PSO, which is known to be less responsive to real-time fluctuations [11,24].
In addition, the proposed method significantly reduces CO2 emissions from 318.9 kg (PSO) to 272.5 kg (MA-DQN), representing a reduction of approximately 14.6%. This reduction is achieved through better prioritization of renewable energy sources and optimized coordination of storage systems, consistent with recent DRL-based microgrid management strategies reported in [15,16].
From an operational perspective, the MA-DQN framework demonstrates improved system stability. By embedding physical constraints directly into the learning process, the controller avoids infeasible operating conditions that may arise in PSO-based solutions, in line with the principles of constrained reinforcement learning discussed in [18,19].
Table 3 presents a quantitative comparison between the PSO-based and MA-DQN approaches, highlighting improvements in cost, emissions, and operational robustness.

5.2. Voltage Stability

Voltage stability is a critical performance indicator in low-voltage radial microgrids, particularly under high penetration of renewable energy sources [4,8,9].
The simulation results show that the proposed MA-DQN framework maintains voltage levels within the acceptable range of ±10% of the nominal value across all nodes of the network. This is achieved through the integration of voltage constraints into both the state representation and the reward function, enabling the agent to anticipate and prevent violations in real time [19].
In contrast, the PSO-based approach exhibits occasional voltage deviations beyond the acceptable limits, particularly during periods of high renewable generation or rapid load variation. These violations highlight the limitations of offline optimization in capturing real-time network dynamics, as also observed in [11,24].
Figure 3 illustrates the voltage profile comparison between the PSO and MA-DQN approaches.

5.3. PCC Behavior

A distinctive feature of the proposed framework is the explicit modeling and control of the point of common coupling (PCC), which plays a crucial role in ensuring stable interaction between the microgrid and the main grid [21].
The results demonstrate that the MA-DQN successfully eliminates oscillatory power exchanges between import and export modes. The exclusivity constraint enforced during training ensures stable and consistent system operation, preventing rapid switching that could degrade power quality.
Conversely, the PSO-based approach may lead to oscillatory behavior at the PCC, particularly under fluctuating renewable generation conditions. These oscillations arise from the absence of explicit constraints governing import/export behavior in the optimization process, a limitation which has also been highlighted in previous studies on classical optimization methods [10,11].
The coordinated action of the PCC agent ensures smooth transitions and stable interaction with the main grid, which is essential for real-world deployment.
Figure 4 presents the PCC power exchange dynamics for both approaches.

5.4. Real-Time Capability

One of the key advantages of the MA-DQN framework lies in its real-time applicability, which is a known limitation of traditional metaheuristic optimization techniques such as PSO [11].
The computational analysis indicates that the PSO-based method requires approximately 12.8 s to compute optimal solutions for a given time step, due to its iterative optimization process. This makes it unsuitable for real-time deployment in dynamic environments.
In contrast, the MA-DQN framework requires approximately 2.1 s per decision cycle after training, corresponding to a reduction of nearly 84% in computation time. This efficiency is consistent with the fast inference capabilities of deep reinforcement learning models reported in [15,16].
As a result, the MA-DQN enables real-time control of the microgrid, allowing rapid adaptation to variations in load demand and renewable generation.
Table 4 summarizes the computational performance of both methods.

6. Conclusions

This paper presented a physics-constrained multi-agent Deep Q-Network (MA-DQN) framework for real-time energy management of a grid-interactive hybrid microgrid operating in a Saharan environment. The proposed approach ensures strict compliance with physical and operational constraints while enabling adaptive and reliable decision-making under dynamic conditions.
The results demonstrate significant improvements in performance, with a reduction in operational cost (−10.4%) and CO2 emissions (−14.6%), while maintaining voltage stability within ±10% limits and ensuring stable PCC operation. In addition, the framework achieves real-time capability with a substantial reduction in computation time.
By promoting the efficient integration of renewable energy sources and reducing dependence on diesel generation, the proposed method contributes to more sustainable and environmentally friendly microgrid operation, particularly in remote and energy-constrained regions.
Overall, this work highlights the potential of physics-constrained deep reinforcement learning as a practical and scalable solution for next-generation intelligent energy management systems in support of renewable energy integration and sustainable development.

Author Contributions

Conceptualization, R.M. and S.S.E.-C.; methodology, R.M. and S.S.E.-C.; software, R.M., A.E.-C. and N.E.-C.; validation, R.M., S.S.E.-C. and A.B.; formal analysis, A.E.-C. and N.E.-C.; investigation, N.E.-C.; resources, A.B. and A.E.-C.; data curation, N.E.-C.; writing—original draft preparation, R.M. and S.S.E.-C.; writing—review and editing, A.E.-C. and N.E.-C.; visualization, R.M. and S.S.E.-C.; supervision, S.S.E.-C. and A.B.; project administration, A.B. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Saha, D.; Bazmohammadi, N.; Vasquez, J.C.; Guerrero, J.M. Multiple Microgrids: A Review of Architectures and Operation and Control Strategies. Energies 2023, 16, 600. [Google Scholar] [CrossRef] [Scilit]
  2. Uddin, M.; Mo, H.; Dong, D.; Elsawah, S.; Zhu, J.; Guerrero, J.M. Microgrids: A Review, Outstanding Issues and Future Trends. Energy Strategy Rev. 2023, 49, 101127. [Google Scholar] [CrossRef] [Scilit]
  3. Li, S.; Oshnoei, A.; Blaabjerg, F.; Anvari-Moghaddam, A. Hierarchical Control for Microgrids: A Survey on Classical and Machine Learning-Based Methods. Sustainability 2023, 15, 8952. [Google Scholar] [CrossRef] [Scilit]
  4. Shirkhani, M.; Tavoosi, J.; Danyali, S.; Sarvenoee, A.K.; Abdali, A.; Mohammadzadeh, A.; Zhang, C. A review on microgrid decentralized energy/voltage control structures and methods. Energy Rep. 2023, 10, 368–380. [Google Scholar] [CrossRef] [Scilit]
  5. Zahraoui, Y.; Alhamrouni, I.; Mekhilef, S.; Khan, M.R.B.; Seyedmahmoudian, M.; Stojcevski, A.; Horan, B. Energy Management System in Microgrids: A Comprehensive Review. Sustainability 2021, 13, 10492. [Google Scholar] [CrossRef] [Scilit]
  6. Albarakati, A.J.; Boujoudar, Y.; Azeroual, M.; Eliysaouy, L.; Kotb, H.; Aljarbouh, A.; Alkahtani, H.K.; Mostafa, S.M.; Tassaddiq, A.; Pupkov, A. Microgrid Energy Management and Monitoring Systems: A Comprehensive Review. Front. Energy Res. 2022, 10, 1097858. [Google Scholar] [CrossRef] [Scilit]
  7. Kassab, F.A.; Rodriguez, R.; Celik, B.; Locment, F.; Sechilariu, M. A Comprehensive Review of Sizing and Energy Management Strategies for Optimal Planning of Microgrids with PV and Other Renewable Integration. Appl. Sci. 2024, 14, 10479. [Google Scholar] [CrossRef] [Scilit]
  8. Ech-Charqaouy, S.S.; Saifaoui, D.; Benzohra, O.; Lebsir, A. Integration of Decentralized Generations into the Distribution Network Smart Grid Downstream of the Meter. IJSmartGrid 2020, 4, 17–27. [Google Scholar] [CrossRef] [Scilit]
  9. Ech-Charqaouy, S.S.; Saifaoui, D.; Benzohra, O.; Lebsir, A. Impact of Integrating Renewable Energies into Distribution Networks on the Voltage Profile. Int. J. Renew. Energy Res. 2020, 10, 143–154. [Google Scholar] [CrossRef] [Scilit]
  10. Thirunavukkarasu, G.S.; Seyedmahmoudian, M.; Jamei, E.; Horan, B.; Mekhilef, S.; Stojcevski, A. Role of optimization techniques in microgrid energy management systems—A review. Energy Strategy Rev. 2022, 43, 100899. [Google Scholar] [CrossRef] [Scilit]
  11. Esparza, A.; Blondin, M.; Trovão, J.P.F. A Review of Optimization Strategies for Energy Management in Microgrids. Energies 2025, 18, 3245. [Google Scholar] [CrossRef] [Scilit]
  12. Mihramane, R.L.; Ech-Charqaouy, S.S.; Saifaoui, D.; Ech-Charqaouy, N.; Ech-Charqaouy, A. Management and Optimization of a Renewable Energy Hybrid System Integrated into a Microgrid. Int. J. Renew. Energy Res. 2025, 15, 213–225. [Google Scholar] [CrossRef] [Scilit]
  13. Trivedi, R.; Khadem, S. Implementation of artificial intelligence techniques in microgrid control environment. Energy AI 2022, 8, 100147. [Google Scholar] [CrossRef] [Scilit]
  14. Joshi, A.; Capezza, S.; Alhaji, A.; Chow, M.-Y. Survey on AI and Machine Learning Techniques for Microgrid Energy Management Systems. IEEE/CAA J. Autom. Sin. 2023, 10, 1513–1529. [Google Scholar] [CrossRef] [Scilit]
  15. Nakabi, T.; Toivanen, P. Deep Reinforcement Learning for Energy Management in a Microgrid with Flexible Demand. Sustain. Energy Grids Netw. 2021, 25, 100413. [Google Scholar] [CrossRef] [Scilit]
  16. Upadhyay, S.; Ahmed, I.; Mihet-Popa, L. Energy Management System for an Industrial Microgrid Using Reinforcement Learning. Energies 2024, 17, 3898. [Google Scholar] [CrossRef] [Scilit]
  17. Nguyen, T.; Nguyen, N.D.; Nahavandi, S. Deep Reinforcement Learning for Multiagent Systems. IEEE Trans. Cybern. 2020, 50, 3826–3839. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Zhang, X.; Wang, Q.; Yu, J.; Sun, Q.; Hu, H.; Liu, X. Multi-Agent Deep-Reinforcement-Learning-Based Strategy for Energy Scheduling. Electronics 2023, 12, 4763. [Google Scholar] [CrossRef] [Scilit]
  19. Ye, Y.; Wang, H.; Chen, P.; Yi, T.; Strbac, G. Safe Deep Reinforcement Learning for Microgrid Energy Management. IEEE Trans. Smart Grid 2023, 14, 3759–3775. [Google Scholar] [CrossRef] [Scilit]
  20. Hu, J.; Shan, Y.; Cheng, K.W.; Islam, S. Overview of Power Converter Control in Microgrids. IEEE Trans. Power Electron. 2022, 37, 9907–9922. [Google Scholar] [CrossRef] [Scilit]
  21. Espina, E.; Llanos, J.; Burgos-Mellado, C.; Cardenas-Dobson, R.; Martinez-Gomez, M.; Saez, D. Distributed Control Strategies for Microgrids. IEEE Access 2020, 8, 193412–193448. [Google Scholar] [CrossRef] [Scilit]
  22. Maher, K.; Kubiak, P.; Cen, Z. Integrating Battery Energy Storage Systems in Hot Desert Regions. In Proceedings of the REPE, Beijing, China, 15–17 September 2023. [Google Scholar] [CrossRef] [Scilit]
  23. Juma, S.A.; Ayeng’O, S.P.; Kimambo, C.Z.M. Review of Control Strategies for Optimized Microgrid Operations. IET Renew. Power Gener. 2024, 18, 2785–2818. [Google Scholar] [CrossRef] [Scilit]
  24. Halev, A.; Liu, Y.; Liu, X. Microgrid Control under Uncertainty. Eng. Appl. Artif. Intell. 2024, 138, 109360. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Architecture of the Grid-Interactive Saharan Hybrid Microgrid.
Figure 1. Architecture of the Grid-Interactive Saharan Hybrid Microgrid.
Engproc 144 00009 g001
Figure 2. Multi-Agent Deep Reinforcement Learning Architecture for Microgrid Control.
Figure 2. Multi-Agent Deep Reinforcement Learning Architecture for Microgrid Control.
Engproc 144 00009 g002
Figure 3. Voltage Profile Comparison.
Figure 3. Voltage Profile Comparison.
Engproc 144 00009 g003
Figure 4. PCC Power Exchange Dynamics.
Figure 4. PCC Power Exchange Dynamics.
Engproc 144 00009 g004
Table 1. Summary of Operational Constraints.
Table 1. Summary of Operational Constraints.
ConstraintMathematical ExpressionDescription
Power balance P = P l o a d Real-time equilibrium
SOC limits S O C m i n S O C S O C m a x Battery protection
Voltage limits 0.9 V n o m V 1.1 V n o m Power quality
PCC constraint P i m p P e x p = 0 No simultaneous exchange
Table 2. Comparison Between PSO and DRL-Based Approaches.
Table 2. Comparison Between PSO and DRL-Based Approaches.
FeaturePSO-Based MethodDRL-Based Method (Proposed)
Optimization typeOfflineReal-time
AdaptabilityLimitedHigh
Handling of uncertaintyWeakStrong
Constraint integrationPenalty-basedEmbedded (physics-constrained)
Computational speedModerate to slowFast (after training)
ScalabilityLimitedHigh (multi-agent)
Table 3. Performance Comparison: PSO vs. MA-DQN.
Table 3. Performance Comparison: PSO vs. MA-DQN.
MetricPSO-BasedMA-DQN (Proposed)Improvement
Operational Cost1250 MAD1120 MAD−10.4%
CO2 Emissions318.9 kg272.5 kg−14.6%
AdaptabilityLowHigh
Constraint SatisfactionPartialFull (embedded)
StabilityModerateHigh
Table 4. Computational Performance.
Table 4. Computational Performance.
MethodOptimization TypeComputation TimeReal-Time Suitability
PSOOffline/iterative12.8 sNo
MA-DQNOnline inference2.1 sYes
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Mihramane, R.; Ech-Charqaouy, S.S.; Boulezhar, A.; Ech-Charqaouy, A.; Ech-Charqaouy, N. Physics-Constrained Multi-Agent Deep Reinforcement Learning for Real-Time Energy Management of a Saharan Hybrid Microgrid. Eng. Proc. 2026, 144, 9. https://doi.org/10.3390/engproc2026144009

AMA Style

Mihramane R, Ech-Charqaouy SS, Boulezhar A, Ech-Charqaouy A, Ech-Charqaouy N. Physics-Constrained Multi-Agent Deep Reinforcement Learning for Real-Time Energy Management of a Saharan Hybrid Microgrid. Engineering Proceedings. 2026; 144(1):9. https://doi.org/10.3390/engproc2026144009

Chicago/Turabian Style

Mihramane, Redouane, S. Salah Ech-Charqaouy, Abdelkader Boulezhar, Amjad Ech-Charqaouy, and Nizar Ech-Charqaouy. 2026. "Physics-Constrained Multi-Agent Deep Reinforcement Learning for Real-Time Energy Management of a Saharan Hybrid Microgrid" Engineering Proceedings 144, no. 1: 9. https://doi.org/10.3390/engproc2026144009

APA Style

Mihramane, R., Ech-Charqaouy, S. S., Boulezhar, A., Ech-Charqaouy, A., & Ech-Charqaouy, N. (2026). Physics-Constrained Multi-Agent Deep Reinforcement Learning for Real-Time Energy Management of a Saharan Hybrid Microgrid. Engineering Proceedings, 144(1), 9. https://doi.org/10.3390/engproc2026144009

Article Metrics

Back to TopTop