1. Introduction
The decarbonization of power systems often necessitates the deployment of hybrid energy microgrids to mitigate the dependence on fossil fuels. These microgrids integrate renewable and conventional energy resources with energy storage capabilities, connecting directly to consumers and regulated by hybrid renewable energy systems (HRESs) to ensure efficient energy utilization and storage.
A parameter sensitivity analysis serves as a critical methodology for the design, optimization, and comprehensive assessment of HRES viability. This analytical approach facilitates the identification of parameters that exert significant influence on the technical, economic, and environmental performance of these systems. Recent scholarly investigations have focused on the impact of various hybrid configurations on the technical and financial viability of HRESs, frequently employing simulation models to address inherent uncertainties [
1,
2,
3,
4,
5]. The primary sources of uncertainty within HRESs include the following:
The intrinsic variability in renewable resources, such as the solar irradiance and wind speed [
1,
2].
Fluctuations in the load demand [
3,
6].
Financial variables, including the cost of capital and the energy discount rate [
2,
3,
4,
7].
The operational lifespan of system equipment [
2,
3,
4].
Variations in energy pricing [
8,
9].
Equipment failure rates [
1,
2,
8,
9].
Analyses of renewable resources indicate that variations in solar insolation and wind velocity significantly impact the system configuration and, consequently, the levelized cost of energy (LCOE) [
1,
2,
9]. Furthermore, research has consistently demonstrated that the lifespan of system components, the annual average wind speed, and the average annual solar irradiance have a significant impact on the energy costs and system reliability. There are also reports highlighting that the integration of diverse renewable sources increases the share of renewable energy and reduces the reliance on diesel generators, thereby contributing to an enhanced HRES reliability [
4,
7,
8,
9].
Regarding the economic aspects and profitability of HRESs, it has been established that a reduction in the proportion of renewable resources can decrease the LCOE [
3,
6,
8]. Conversely, an increase in the load demand or ambient temperatures tends to escalate the LCOE [
6]. These studies collectively suggest that discount rates and investment costs exert a substantial influence on the profitability and sizing of HRESs. Notably, 100% renewable systems have demonstrated the potential for cost reductions exceeding 30% when compared to traditional power system configurations [
5,
7,
9].
Table 1 summarizes the sensitive parameters commonly analyzed in the numerous studies presented in the literature.
A parameter sensitivity analysis is essential for ensuring the viability of HRESs. It accounts for uncertainties such as resource variability and fluctuating load demands while also considering specific requirements, including critical load priorities and desired energy autonomy levels [
5,
7,
8]. This analysis can also facilitate the sale of surplus energy to the grid, helping to maximize the renewable penetration and profitability [
3,
6]. However, the current approaches often overlook two key aspects:
The HRES sensitivity to emergency energy constraints: To minimize the carbon footprint, it is crucial to limit reliance on generators or the main grid. In many deep reinforcement learning (DRL) applications, this limit is often implicitly below 5% or entirely absent [
11]. A thorough exploration is needed to understand how system policies adapt when this hyperparameter is relaxed during scaling.
Dynamic economic sizing and associated risks: The real-time spot price signals and the inherent volatility of renewable resources demand risk-aware planning. While recent research has proposed risk-sensitive deep reinforcement learning using the conditional value at risk (CVaR) [
12], its applications within HRES contexts remain limited.
Traditionally, HRES sizing relied on meta-heuristics such as multi-objective particle swarm optimization (MOPSO) and the non-dominated sorting genetic algorithm II (NSGA-II). However, recent advancements in DRL are revolutionizing this field. DRL enables real-time policy learning that inherently accounts for system dynamics, shifting control from rigid rule-based or forecast-driven approaches to autonomous, adaptive decision-making, even in complex and uncertain environments.
Several DRL algorithms have become prominent in this domain. The twin delayed deep deterministic policy gradient (TD3) aims to prevent the overestimation of action values through twin Q-networks and delayed policy updates [
13]. Soft actor–critic (SAC), on the other hand, prioritizes maximizing the policy entropy to encourage the exploration of diverse control strategies [
14]. A hierarchical deep Q-network (HDQN) stands out with its hierarchical policy structure, making it particularly well-suited for learning across different decision levels, such as distinguishing between operational control and strategic system sizing [
15].
Despite these advancements, a critical gap remains in the research: the sensitivity of HRESs to emergency energy constraints. While DRL applications for HRESs have largely focused on cost reduction and renewable energy integration, the impact of a permissible load demand from generators and the main grid on system design is often overlooked. It is crucial to understand how varying levels of contribution from diesel generators (DG) and the electrical power grid (GRID) affect the reliability and cost of HRESs, especially for designing resilient systems. Unfortunately, this vital constraint is frequently neglected, particularly in models that lack dynamic penalties for insufficient backup power. Furthermore, while classical sensitivity methods are commonly applied in power grids, their use in hybrid systems remains limited [
16,
17]. This study aimed to address this critical gap by comprehensively analyzing the sensitivity of HRESs to backup-energy constraints. Our goal was to identify crucial trade-offs and anticipate the risks associated with resource variability and operational constraints.
This manuscript is related to our earlier paper [
18], but its objective is different and more specific. The previous study primarily established the feasibility of using a DRL-based framework for hybrid renewable energy system sizing under a general multi-criteria setting. In that earlier work, backup sources such as the diesel generator and the utility grid were part of the overall sizing environment, but they were not the central analytical axis of the study.
In contrast, the present manuscript focuses specifically on the sensitivity of HRES design to admissible backup-energy contributions. More precisely, the main originality of this paper lies in treating the generator and grid contribution ratios, denoted by and , as explicit design-control parameters and in analyzing how these parameters reshape the trade-offs among LCOE, LPSP, and REF. Accordingly, the contribution here is not simply an additional application of DRL, but a constraint-aware sensitivity analysis framework, complemented by a structured benchmark against MOPSO and NSGA-II, with the aim of extracting design-oriented insights on the role of backup-energy flexibility in HRES planning.
Our key contributions are as follows:
A sensitivity analysis framework in which tau and gamma are modeled as explicit admissible contribution ratios for the generator and the grid, allowing for backup-energy dependence to be studied as a design variable.
A DRL-based constrained optimization approach with a cumulative reward that jointly accounts for LCOE minimization, LPSP minimization, REF maximization, and penalties associated with excessive reliance on backup sources.
A structured multi-objective benchmark against MOPSO and NSGA-II, enabling a comparative discussion of where DRL is advantageous, where evolutionary baselines remain more conservative, and how admissible backup constraints reshape the compromise between the cost, reliability, and renewable integration.
A parameter sensitivity analysis of DG/GRID contributions is crucial for optimizing the sizing of HRESs. It allows us to test the robustness of a given configuration under real-world uncertainties. This analysis helps identify key parameters, such as the allowable percentage of energy input from the generator and the main grid, which significantly influence the values of the LCOE, LPSP, and REF, especially under various constraints. Ultimately, it highlights how resource variability and economic factors shape the design choices for a truly sustainable energy transition.
This article is structured as follows:
Section 2 details the study’s methodology, including a comprehensive presentation of the proposed DRL approach, comparative methods, and experimental design.
Section 3 presents the study’s results and compares the different methodologies. The article concludes with a summary and outlines future research directions.
2. System Modeling
This research includes a sensitivity analysis of HRESs. This section establishes the conceptual and experimental framework for the analysis of diesel generator and grid configurations in hybrid microgrids. We structured it into three subsections: the formulation of the optimization problem, modeling using DRL, and the design of the numerical experiment. The core objective of this study was to leverage an intelligent agent to discover beneficial energy distribution strategies. These strategies were designed to satisfy critical reliability constraints, such as limiting the emergency energy supply from the generator and the main grid. Simultaneously, the goal was to maximize renewable energy penetration and minimize the operating costs. A key part of this study also involved examining how different allowable limits for these emergency sources impact the system’s overall performance.
2.1. HRES Configurations
For clarity,
Figure 1 was created for this manuscript to reflect the updated simulated system design; the earlier configuration is described in [
18].
In the context of this sensitivity analysis, the study supports the following hypothesis: the constituents of the system that meet one need differ from those that meet another. These systems consist of photovoltaic (PV) panels, wind turbines (WTs), a battery energy storage system (BESS), and possibly a diesel generator (DG).
Indeed, Kushwaha and Bhattacharjee [
19] highlighted the importance of choosing the right HRES components. The results of their study on the comparison of different types of systems that include PV panels, WTs, a BESS, a DG, and/or a biogas generator (BG) revealed the PV + WT + BESS + BG + DG system to be best suited in their context. Similarly, the study by Mokhtara, Negrou et al. [
20] showed that the most suitable system for electrifying a residential building in a rural environment is PV + WT + BESS + DG.
Therefore, to ensure an exhaustive analysis, this study considered four (4) types of hybrid renewable energy systems:
PV panels coupled with a BESS: PV + BESS.
PV panels coupled with a BESS and a DG: PV + BESS + DG.
PV panels coupled with WTs and a BESS: PV + WT + BESS.
PV panels coupled with WTs, a BESS, and a DG: PV + WT + BESS + DG.
Table 2 summarizes the technical and techno-economic parameters used for the main HRES components. These values follow the component assumptions adopted in the present study and are consistent with the system modeling framework reported in our earlier DRL work [
18].
Figure 1 illustrates an HRES that integrates photovoltaic panels, wind turbines, battery energy storage systems, and a generator. A bidirectional inverter facilitates the conversion of battery energy to loads and the storage of surplus energy, serving a dual purpose. A hybrid inverter converts DC energy from the PV panels into usable AC power for the loads. The system also connects to a generator, providing power during periods of the underproduction of renewable energy.
The power grid maintains a bidirectional connection to the loads, enabling two functions: supplying energy when the demand exceeds the renewable production and exporting surplus energy to the grid when storage is not feasible.
At each hourly time step, the energy balance is enforced by prioritizing renewable generation and then activating storage and backup sources when needed. When the renewable production exceeds the load demand, the surplus is first directed to battery charging subject to the converter efficiency and storage limits; if the battery is already full, the remaining surplus may be exported to the grid when grid export is allowed. Conversely, when the renewable production is lower than the demand, the deficit is covered sequentially by the battery and then by the diesel generator and/or the grid within the admissible limits defined by tau and gamma. Conversion losses are accounted for through the inverter and storage efficiencies, while any remaining unmet demand contributes directly to the LPSP metric.
The core challenge of this study lies in balancing two conflicting objectives: minimizing the reliance on external sources for decarbonization while maintaining an acceptable level of supply security, especially during renewable energy deficits. Backup energy bridges the gap between the demand and the available renewable sources (PV + WT + BESS). However, this backup must remain below a critical threshold to prevent excessive fossil fuel use and foster resilient, autonomous systems.
Therefore, this study addressed a dynamic optimization problem of energy dispatch between various sources. The primary goal was to minimize the LCOE while adhering to strict constraints on the DG and GRID contributions and ensuring system robustness. This optimization is complex due to the inherent uncertainty of renewable generation (PV + WT), a fluctuating demand, and nonlinear interactions between microgrid components.
2.2. System Performance Metrics
The levelized cost of energy (LCOE), renewable energy fraction (REF), and loss of power supply probability (LPSP) are often used as HRES performance metrics. Their role is as follows:
The LCOE measures the economic performance of an HRES [
21,
22];
The REF quantifies the level of renewable integration within the microgrid [
23];
The LPSP indicates the reliability of an HRES [
23,
24].
This study considered these three performance indicators for their ability to offer a balanced approach to HRES evaluation. More specifically, the LCOE provides a summary of all the investment, operating, and maintenance costs over the system’s lifetime. The LPSP provides a quantitative measure of the probability of energy failure. This measure ensures that the critical needs remain covered, even in the event of fluctuating resources. Finally, the REF serves as an indicator of the proportion of energy from renewable sources. The REF thus provides a measure of a system’s effectiveness in minimizing the dependence on fossil fuels and reducing its carbon footprint.
2.2.1. Levelized Cost of Energy
The levelized cost of energy is a tool used as an indicator for comparing the lifetime costs of the electricity production of different energy technologies
. This measure is widely used in the literature to assess the economic competitiveness of renewable and conventional energy [
21,
22,
25,
26].
The LCOE is highly dependent on the discount rate, initial capital costs, operating costs, effective lifetime of the installation
, and energy efficiency [
21,
25]. The LCOE, expressed in
$/kWh, is represented by Equation (
1) [
21]:
with
where
represents the total investment cost over the entire project life. This cost includes the initial investment cost
, replacement costs
, maintenance and operation costs
, and the costs of purchasing energy from the grid
subtracted from the costs of selling energy on the grid
.
and
are, respectively, the total quantities of energy purchased and sold on the electricity grid.
represents the rate of purchase of energy, and
is the rate of energy sales.
Table 3 presents the values of the parameters used in this study.
2.2.2. Renewable Energy Fraction
The REF, usually expressed as a percentage (%), is an indicator that measures the share of energy produced from renewable sources in each energy system. For this study, the REF focuses on the amount of energy produced by the PV source in PV + BESS and PV + BESS + DG systems, as well as by the PV and WT sources for PV + WT + BESS and PV + WT + BESS + DG systems. The REF enables the quantification of progress toward more sustainable energy goals [
23].
Considering a PV + WT + BESS + DG system attached to the grid, the REF is expressed by Equation (
3). Combining the REF with the total system efficiency helps assess how an added renewable capacity reduces fossil fuel use [
23]. This study specifically reflects the reduction in diesel generator energy input.
and represent the production from wind turbines and the production from PV panels, respectively, at time t. is the efficiency of the PV panels. This is a value provided by the manufacturer. and represent the amount of power generated by the grid and the diesel generator at the same time t.
2.2.3. Loss of Power Supply Probability
The loss of power supply probability is a crucial indicator in electricity distribution and transmission systems, especially for integrating renewable energy and variable loads. The optimization of the LPSP enables the optimization of HRES management. It also improves the reliability of the system by incorporating the uncertainty related to the variability in loads and energy sources [
24,
27,
28]. The LPSP is expressed in Equation (
4) below [
29]:
with
representing the total amount of energy supplied by the sources at time
t. Thus,
represents the energy loss recorded in time
t.
2.3. Optimization Model
The optimization of the HRES in the context of this sensitivity analysis study is a multi-criteria optimization problem under uncertainties and constraints. The objectives of the optimization problem are the LCOE, the REF, and the LPSP, as described in the previous sections. Equation (
5) presents the optimization problem to be solved under the constraints in Equation (
6) [
30]:
with
(
Table 4) representing the percentages of allowable energy input via the grid and the diesel generator, respectively.
represents the maximum acceptable value for the HRES reliability criterion.
The selected tau and gamma ranges were chosen to cover five interpretable operating regimes, from restrictive backup usage (0.1) to relatively permissive backup usage (0.5), while preserving a manageable sensitivity grid for comparative analysis. Values below 0.1 were considered overly restrictive for the selected study setting because they would sharply limit the corrective role of the backup sources, whereas values above 0.5 would reduce the practical meaning of the low-carbon design objective by allowing excessive dependence on external support. The adopted range therefore provides a balanced compromise between interpretability, feasibility, and computational tractability.
To analyze the sensitivity of HRESs to external power inputs, this study presents an enhanced experimental design. We examined the impact of contributions from diesel generators and the main grid on key metrics, including the LCOE, REF, and LPSP. This section begins by detailing the specific parameters under investigation and their varied ranges, as summarized in
Table 4. Following this, we introduced the three primary evaluation methodologies utilized: DRL, MOPSO, and the NSGA-II.
2.4. Load Profiles
This study considered three load demand profiles (P1, P2, and P3) with significant variations ranging from 21.9% to 74.6% (
Figure 2). These profiles, which exhibit high variation, represent the demands of three different types of buildings from the National Renewable Energy Laboratory (NREL) database. The aim of selecting these three buildings with contrasting profiles was to provide a rigorous analysis of the influence of load variability on the technical and economic performance of the HRES. This approach also ensures the optimal representativeness of real operating scenarios. However, in the context of this study, the same locality was considered. Thus, the climate data were the typical meteorological year (TMY) data of the said locality. The wind speed, solar irradiance, and temperature of this locality are shown in
Figure 3. This figure illustrates the significant seasonal variability in wind and solar resources and temperatures throughout the year. This analysis of the source variability makes it possible to accurately assess the climatic impact on the HRES performance and reliability.
To improve clarity on the input data pipeline, all load and climatic series were organized on an hourly basis over one representative year (8760 time steps). The preprocessing stage consisted of aligning the NREL demand profiles and the TMY climatic variables on the same temporal grid, checking the unit consistency, and directly using the resulting hourly profiles as inputs to the HRES evaluation model. No ad hoc artificial measurements were introduced; instead, the reported results are simulation-derived values obtained from the adopted physical and techno-economic modeling framework.
In addition, this study evaluated four (4) types of HRESs: PV + BESS, PV + WT + BESS, PV + BESS + DG, and PV + WT + BESS + DG. The systems were evaluated for each type of profile to analyze the impact of each system type on the profiles and the associated analysis factors.
Table 5 shows the total demand for each profile.
For clarity,
Figure 2 and
Figure 3 were generated specifically for this manuscript from the study datasets and simulation setup (
Section 2.5); earlier work [
18] is cited for additional methodological details.
2.5. Solution Approach
This section outlines the methodologies employed to determine the best HRES configurations based on the LCOE, REF, and LPSP performance criteria.
The following subsections present the optimization and evaluation procedures adopted in the present study; additional methodological details are available in [
18].
The manuscript emphasizes that the objective here is a sensitivity analysis rather than exhaustive algorithmic retuning. For this reason, the DRL training setting was kept unchanged across the tested scenarios so that the reported differences can be attributed to the load profiles, HRES configurations, and admissible backup-energy constraints rather than to changing the learning settings. Replay memory and a target network were used throughout the optimization to reduce the instability caused by sequentially correlated updates.
2.5.1. Deep Reinforcement Learning
This study presents a multi-step DRL technique (
Figure 4). This technique relies on a replay-memory and target-network training procedure designed to improve learning stability during sequential HRES evaluations. It begins with an initialization phase, where key system parameters are configured to establish the initial state vector at time step zero. As the process progresses, these parameter values are updated to determine the new state. The HRES state, denoted as
S, is defined in Equation (
7):
where
and
are the number of PV panels and wind turbines to be considered in each agent’s decision.
and
represent the storage system’s capacity and the diesel generator’s capacity to be considered, respectively.
and
are the amount of energy not supplied and the amount of energy lost over the entire evaluation period, i.e., over the lifetime of the HRES. Finally, LCOE, LPSP, and REF are the values of the factors measuring the quality of the DRL agent’s decision.
Following the initialization phase, the agent transitions into the decision phase. In this stage, it dynamically selects updated values for critical HRES parameters, including the number of PV panels and wind turbines, the battery energy storage system capacity, and the diesel generator size. These selections were made by considering the HRES’s previous state and aimed to optimize the system within the permissible energy limits from the DG and the grid.
To define the HRES’s subsequent state, a complete analysis spanning the system’s total lifespan is performed after each agent decision. This analysis yields new values for the objective functions and parameters that characterize the updated state. The agent’s new state and its corresponding decision are logged in the replay buffer. This buffer serves as a training ground, enabling the agent to continuously learn from its past decisions and the resulting impacts. For each agent training iteration, a minibatch of T random transactions is drawn from the replay buffer, as shown in Equation (
8):
Each element of the minibatch of sampled transitions (Equation (
8)) is composed of five parameters:
, the current state at time
j;
, the action taken from state
;
, the reward received after taking action
;
, the next state of the system; and
, a binary indicator signaling whether
is a terminal state or not. For each element of the transaction T, the target value
is calculated using Equation (
9):
with
, and
representing the discount rate, the action of the agent, the state of the HRES obtained after applying the action in the environment, the reward received, and the weight of the target network (the lagged copy of
) respectively. Thus,
is the target action value parameterized by
.
and
(Equation (
10)) are the learning rate and the minimization function of the loss function, respectively, with respect to the parameter
(Equation (
11)).
In addition to the target network used in this study to stabilize learning, the replay buffer allows correlations between sequential data to be broken [
31]. Through this decision-making and learning approach, the agent enhances its understanding of the system and the environment, adapting its decisions to seek a better reward.
Finally, the application of DRL in this multi-objective optimization study requires the definition of the reward function to be maximized by the DRL agent during the process. This reward function can be written as shown in Equation (
12).
where
n represents the number of iterations of the study,
is the weighting of the objective function
, and
is the penalty function with respect to the defined constraints and is expressed by Equation (
13). This study considered
,
, and
, respectively, for the LCOE, LPSP, and REF objective functions.
The objective weights were chosen to give comparable importance to the three performance dimensions during learning rather than to define a convex combination summing to one. Since each improvement term is already normalized by the largest absolute improvement at a given iteration, the role of the weights is to preserve a balanced influence of cost, reliability, and renewable integration in the reward signal. Equal weights were therefore used to avoid introducing an a priori preference for one objective over the others in this sensitivity-oriented study.
is the maximum permissible value via the grid or diesel generator. It is a function of
,
, and the load profile (Equation (
5)).
represents the total quantity admitted via the grid or diesel generator at iteration
i.
represents the value obtained from the LPSP at iteration
i.
The reward function was designed to reflect the multi-objective nature of the sizing problem while preserving a simple and stable optimization signal for the DRL agent. The first term rewards improvements with respect to the best objective values found so far, thereby encouraging the progressive refinement of the solution. In this formulation, the LCOE and LPSP are minimized, whereas the REF is maximized; the sign of each term is therefore chosen according to the direction of improvement associated with each objective. The normalization by the largest absolute improvement avoids the disproportionate dominance of one objective solely because of scale differences.
The penalty term has a complementary role. Its purpose is not to replace the objective terms but to enforce feasibility with respect to the backup-energy constraints and the reliability requirement. In practice, penalizes configurations that exceed the admissible generator or grid contributions and penalizes solutions that do not satisfy the prescribed LPSP threshold. This structure was chosen to maintain a clear separation between performance-seeking behavior and constraint-satisfaction behavior. As a result, the reward remains interpretable: the agent is encouraged to improve the techno-economic objectives while being discouraged from relying excessively on external backup sources.
2.5.2. Implementation Details
To improve reproducibility, the main DRL implementation variables are summarized in
Table 6. The agent uses a TD3-inspired training procedure with replay memory and a target network to stabilize updates. At each training step, a minibatch of transitions
is sampled from the replay buffer, the temporal-difference target is computed, and the network parameters are updated through gradient descent on the mean-squared loss. The state vector includes the main design variables and the resulting performance indicators, namely
,
,
,
, the unmet energy, the lost energy, the LCOE, the LPSP, and the REF.
The key training hyperparameters used in the study include the learning rate , the discount factor , the replay buffer minibatch size T, and the stop condition based on the maximum number of training iterations. In addition, the target network and replay memory were used to reduce the instability commonly associated with sequentially correlated updates. Since the objective of this manuscript was a sensitivity analysis rather than algorithmic hyperparameter optimization, the DRL hyperparameters were kept fixed across the tested scenarios so that the reported differences can be attributed to the , , load profiles, and HRES configurations rather than to changing training settings.
The action space of the agent corresponds to candidate updates of the main sizing variables, namely the number of PV panels, the number of wind turbines, the battery capacity, and the diesel generator capacity. After each action, the resulting HRES configuration is fully evaluated over the considered simulation horizon, and the corresponding next state is computed. A state is treated as terminal when the stopping condition of the current optimization episode is reached, namely the prescribed maximum number of training iterations.
The replay buffer was used to store sampled transitions and reduce the bias induced by sequential correlation. The target network was used as a delayed reference network to stabilize the target-value computation during learning. In addition, the same training settings were maintained across all studied scenarios so that the reported differences can be attributed to the sensitivity parameters tau and gamma, the load profiles, and the HRES configurations rather than to changes in the learning setup.
2.5.3. Multi-Objective Particle Swarm Optimization
We employed MOPSO, as depicted in
Figure 5, as a key comparison method. MOPSO is a powerful algorithm specifically designed to tackle complex multi-objective optimization problems.
Unlike standard PSO, it explicitly handles conflicting objectives, such as minimizing the LCOE, maximizing the REF, and minimizing the LPSP. It achieves this by utilizing a Pareto dominance mechanism and an external archive to store a diverse set of non-dominated solutions on the Pareto front [
32,
33].
MOPSO has a strong track record in real-world applications, from industrial optimization to the design of distributed energy systems. It is particularly effective at balancing exploration and exploitation, leading to rapid convergence towards high-quality solutions [
32,
34].
For this research, MOPSO models a complex decision space where each “particle” represents a unique technological configuration of the hybrid renewable energy system. These particles evolve through the decision space, guided by learning rules derived from their own history and the best archived solutions, with updates influenced by a leader selection mechanism.
This allows us to generate representative Pareto fronts, which are crucial for comparatively evaluating the DRL performance under various stress scenarios. Therefore, MOPSO serves as a robust reference method in this study, valued for the quality of its generated solutions and the diversity of the energy trade-offs that it identifies [
32].
Table 7 presents the values of the MOPSO parameters considered in this study [
33,
35].
2.5.4. Non-Dominated Sorted Genetic Algorithm II
The NSGA-II is a widely used, population-based, multi-objective evolutionary algorithm in the field of optimization due to its ability to produce well-distributed Pareto fronts while maintaining high-quality solutions efficiently [
36,
37,
38,
39].
The NSGA-II has been widely applied in sectors as varied as water management, energy architecture, and biological systems, with convincing results on the robustness and diversity of solutions [
36].
Each individual in the population represents a specific microgrid technology configuration, comprising the number of PVs, the number of WTs, the BESS capacity, and the DG capacity. These individuals are assessed simultaneously according to three conflicting objectives: the LCOE, the LPSP, and the REF (
Figure 6).
The NSGA-II employs an evolutionary process that includes selection, crossover, and mutation, followed by the classification of solutions based on their Pareto dominance and density metric, which preserves diversity without requiring partition parameters. Moreover, this approach includes a mechanism of elitism that merges the parental and child populations to ensure that the best solutions are maintained from one generation to the next [
38,
40].
The non-dominated sorting algorithm used in the NSGA-II exhibits an optimized computational complexity of , where M represents the number of objectives and N denotes the population size. This enables the NSGA-II to efficiently identify reliable trade-offs between the economic, energy, and environmental performance in complex systems, such as HRESs. The NSGA-II is a relevant comparative method for assessing the quality of solutions generated by DRL-based approaches.
Table 8 presents the values of the NSGA-II parameters assumed in this study [
41,
42].
4. Conclusions
This study presents a sensitivity analysis framework for HRES sizing in which tau and gamma are treated as explicit admissible backup-energy ratios for the generator and the grid. The analysis was conducted across four HRES configurations and three contrasting load profiles with the objective of understanding how these constraints reshape the compromise between three conflicting criteria: the LCOE, LPSP, and REF.
The experimental results show that the main value of the proposed DRL approach lies not in universal dominance, but in its ability to uncover adaptive trade-offs in complex scenarios, especially when the demand variability and production structure are more difficult to control. The analyses revealed that various factors impact the sizing process of hybrid renewable energy systems. First, the results show that the type of system plays a central role in the overall performance. Configurations that integrate multiple sources, including PV panels, WTs, a BESS, and a DG, offer the best trade-off between the LCOE, LPSP, and REF. Specifically, for a PV + WT + BESS + DG system, the DRL approach can identify parameters that ensure a robust system performance, achieving a low LPSP (below 0.3%) and a competitive LCOE (under $0.4/kWh). Compared with MOPSO and NSGA-II, DRL also maintains a high REF exceeding 60%.
Dimensioning also requires that the technical constraints be carefully calibrated. A value that is too low (<0.1) limits the generator’s backup supply, increasing the LPSP, while a value that is too high (>0.4) penalizes the REF and increases the LCOE. The results suggest that an average (∼0.25–0.3) allows the DRL method to use the generator only when economically and ecologically justified. As for , an intermediate level (∼0.3) seems more suitable. This intermediate value provides sufficient flexibility without inducing excessive dependence on the grid. This approach maximizes the return on investments while controlling costs.
Compared with the MOPSO and NSGA-II methods, the results also show that the adopted DRL approach can adapt to changing multi-objective constraints, even if the LPSP or LCOE performance remains more dispersed in some contexts, depending on the specific conditions. In contrast, MOPSO and NSGA-II offered greater stability and reliability in standard scenarios, at the cost of less use of renewable sources.
Through these analyses, this study suggests utilizing DRL in dynamic contexts, where the objective is to enhance energy autonomy and integrate renewable energies while maintaining strategic adaptability to operational constraints. Thus, a hybrid combination with more conservative methods can be considered for systems subjected to strict requirements for the continuity of supply.
Overall, the results suggest that intermediate admissible backup thresholds offer better compromises than extreme values. Moderate tau and gamma values provide enough operational flexibility to preserve reliability without inducing excessive dependence on fossil-based or grid-supplied energy. In that sense, the framework provides design-oriented guidance for resilient and low-carbon HRES planning.
As a result, the following avenues of research for improving the DRL approach are presented:
Hybridization of the DRL method with conservative approaches to improve the accuracy between exploration and exploitation.
Adaptation of the DRL to dynamic electricity markets.
Coupling the DRL with weather and load forecasts for long-term predictive optimization.
In short, by strengthening the robustness of the DRL, it could position itself as a technological pillar of the smart energy transition.