1. Introduction
The increasing penetration of renewable energy sources and the growing demand for reliable off-grid power supply have placed direct current microgrids (DCMGs) at the forefront of modern power systems [
1,
2]. DCMGs offer inherent advantages such as higher efficiency, simpler integration of storage devices, and absence of reactive power and synchronization issues compared to AC microgrids [
3,
4]. A typical DCMG comprises photovoltaic (PV) panels, battery energy storage, supercapacitors (SCs), and variable loads, all interconnected through DC-DC converters [
5]. However, the intermittent nature of solar generation and the stochastic behavior of load demand introduce significant challenges in maintaining DC bus voltage stability and ensuring optimal power sharing among storage units [
6,
7].
To address these challenges, various control strategies have been proposed. Linear controllers such as proportional-integral (PI) regulators are widely adopted due to their simplicity, but they often fail under large disturbances or parameter variations [
8]. Nonlinear SMC has emerged as a robust alternative, offering finite-time convergence and insensitivity to matched uncertainties [
9,
10]. In the context of DCMGs, SMC has been successfully applied to regulate converter currents and bus voltage [
11,
12]. Nevertheless, classical SMC relies on fixed gains and may exhibit chattering, which can degrade actuator lifetime. Moreover, the absence of adaptation mechanisms limits its performance under rapidly changing operating conditions.
To overcome the limitations of fixed-gain controllers, fuzzy logic systems (FLS) have been integrated into adaptive control strategies. FLS can approximate unknown nonlinear functions using linguistic rules, thereby reducing the need for accurate system models [
13,
14]. For instance, a hierarchical energy management system combining SMC with fuzzy logic was proposed for a DCMG, achieving improved dynamic response compared to linear controllers [
15]. However, fuzzy-based strategies remain heuristic and lack learning capabilities; their performance heavily depends on the quality of predefined rule bases, which may not cover extreme or unforeseen scenarios [
16].
Recent advances in adaptive fuzzy control have addressed challenging issues such as event-triggered mechanisms, dynamic quantization, and deception attacks in networked control systems [
17,
18]. For instance, adaptive event-triggered tracking control strategies have been developed for nonlinear networked systems with dynamic quantization and deception attacks [
17], while quantized fuzzy guaranteed cost control has been proposed for electric vehicles with uncertain parameters [
18]. These techniques provide valuable insights for enhancing robustness and communication efficiency in complex control systems.
More recently, artificial intelligence techniques, particularly deep reinforcement learning (DRL), have attracted attention for energy management in microgrids [
19,
20]. DQL, a representative DRL algorithm, enables an agent to learn optimal policies by interacting with the environment [
21]. A hybrid control architecture combining SMC for low-level regulation and DQL for high-level decision making (charging, discharging, load shedding) was recently demonstrated for a DCMG [
22]. The results showed improved voltage stability and reduced battery stress compared to fuzzy logic control. Nevertheless, this architecture still has notable shortcomings: the DQL agent uses only instantaneous state information without any predictive capability, the SMC gains are fixed, and no mechanism compensates for unmodeled nonlinearities or external disturbances.
In our previous work [
20], we introduced a DQL-based sliding mode control scheme for DC microgrids. However, that approach used only instantaneous state information and fixed SMC gains. The present paper significantly extends this prior contribution by incorporating: (i) predictive information into the DQL state space, (ii) online adaptation of SMC gains, and (iii) a fuzzy compensator for unknown nonlinearities. These extensions lead to substantially improved robustness and voltage regulation performance.
In parallel, a data-driven predictive fuzzy adaptive control approach was developed for nonlinearly parameterized systems [
23]. That work introduced a data-driven predictive model built from historical input-output data to forecast future system behavior. The controller gains were optimized online by minimizing a predictive performance index, and fuzzy logic systems were employed to approximate lumped unknown functions. Stability was proven via a composite Lyapunov function, and simulation results demonstrated significant improvements in tracking accuracy and disturbance rejection. However, that approach was designed for general nonlinear systems and has not been applied to DCMGs, nor combined with deep reinforcement learning.
Motivated by the above observations, this paper proposes a data driven predictive hybrid control strategy for DCMGs that synergistically integrates three complementary techniques: adaptive fuzzy sliding mode control, a historical data based predictive model, and a DQL agent with augmented state. This synergy enables proactive energy management while maintaining stability guarantees for the low level AFSMC inner loop, with overall system boundedness ensured in a cascade sense under bounded reference currents.
The main contributions are as follows:
A predictive model is constructed online from stored input–output data to forecast DC bus voltage and PV power over a finite horizon. This model provides prediction errors that are used both to optimize SMC gains and to enrich the state space of the DQL agent.
The SMC gains are recursively updated in real time by minimizing a predictive performance criterion, ensuring fast voltage regulation under varying conditions without manual retuning.
Fuzzy logic compensators are embedded into the SMC laws to approximate unknown nonlinearities and disturbances, eliminating the need for an exact system model and significantly reducing chattering.
The DQL agent’s state vector is augmented with prediction errors and future estimates, enabling it to learn anticipative policies that proactively manage battery and SC operations and load shedding decisions.
A composite Lyapunov analysis is provided to prove uniform ultimate boundedness of the low level AFSMC inner loop, assuming that the reference currents generated by the DQL agent are bounded and piecewise continuous a condition physically guaranteed by converter limits and the agent’s bounded action space. The boundedness of the inner loop tracking errors then implies boundedness of the DC bus voltage dynamics, ensuring overall stability in a cascade sense.
The comparison is performed against a conventional DQL agent that relies solely on instantaneous state information and uses fixed gain SMC, as proposed in [
22], in contrast to our augmented state and online gain optimization strategy.
The remainder of the paper is organized as follows.
Section 2 presents the dynamic model of the DCMG.
Section 3 describes the data-driven predictive model and the online gain optimization algorithm.
Section 4 details the adaptive fuzzy sliding mode controllers.
Section 5 presents the augmented DQL agent and the integrated decision-making strategy.
Section 6 provides the stability analysis.
Section 7 discusses simulation results and comparisons.
Section 8 concludes the paper.
3. Data Driven Predictive Model and Online Gain Optimization Algorithm
This section describes how historical input output data are exploited to build a predictive model of the DC microgrid. This model is then used to optimize the gains of the sliding mode controllers online, thereby enhancing voltage regulation and disturbance rejection.
3.1. Predictive Model Construction with Recursive Gain Optimization
Let the discrete time measurements be taken at sampling instants
k = 0,1,2,… with a fixed step Δ
t. For each subsystem, the control input (duty cycle) and the measured output (e.g., DC bus voltage, inductor current) are stored. Define the following data matrices over a window of length
L:
where
u(⋅) represents the vector of duty cycles
and
y(⋅) represents the measured outputs (typically
vDC,
Ipv,
Ib,
ISC, and). The goal is to predict the future output
y(
k +
p) based on past data and a sequence of future control inputs.
Following the subspace identification philosophy, a linear predictor can be constructed as:
where
p is the prediction horizon (
p = 1, 2,…,
Np),
is the sequence of future control inputs (to be optimized) and
Hu,
Hy, Hf are constant matrices (of appropriate dimensions) that are updated online using recursive least squares or a similar adaptive identification method.
In practice, for a DC microgrid, the most critical variable to predict is the DC bus voltage
vDC. Therefore, we focus on a single output predictor for
vDC, although the same formalism applies to other variables (e.g., PV power). The predictor takes the form:
where
θu,
θy,
θf are parameter vectors estimated from data. Detailed implementation parameters, including the prediction horizon, data window length, forgetting factor, matrix dimensions, RLS initialization, and persistency of excitation verification, are provided in
Section 3.3.
Define the regression vector at sampling instant
k as:
where
is the vector of duty cycles, and
is the DC bus voltage. The predictor output is:
where
is the parameter vector estimated online via RLS:
with forgetting factor λ ∈ (0,1].
Lemma 1 (Bounded prediction error). Under the assumption of persistency of excitation and for a sufficiently large data window L, the prediction error is bounded, i.e., for some positive constant .
The prediction error will be used later both as a compensation term in the SMC laws and as an additional input to the DQL agent.
In the proposed strategy, the SMC gains are made time varying and optimized online. Let the vector of tunable gains be denoted by: .
These time varying gains directly influence the SMC laws derived in
Section 4. In order to obtain the most suitable gain vector k(
k) at each sampling instant, we introduce a predictive performance index tailored to the regulation objectives of the DC microgrid. This index is designed to capture the future evolution of the DC bus voltage while penalising abrupt changes in the control gains. Specifically, the cost function combines a weighted quadratic term of the predicted voltage error over a finite horizon and a regularisation term that limits the step to step variation of the gains. The expression of this index is given by:
Here, k denotes the discrete sampling instant, and denotes the j-step ahead prediction of the bus voltage at sampling instant k, obtained from the discrete-time predictor (7, is the reference value, Δk(k) = k(k) − k(k − 1) is the gain increment, and Q > 0, R > 0 are tuning weights. The first term enforces accurate voltage tracking, while the second term prevents excessively aggressive gain updates that could destabilise the system. The minimisation of Jp (k) with respect to k(k) is performed online using a recursive gradient based algorithm, as detailed in the following subsection.
The first term drives the predicted bus voltage towards its reference, while the second term prevents aggressive gain fluctuations that could destabilize the system or cause excessive chattering. At each sampling instant k, the optimal gain vector k∗(k) is obtained by minimizing Jp (k) subject to constraints that guarantee closed loop stability. Because the predictor (7) is linear in the future control inputs (which themselves depend on the gains through the SMC laws), a direct minimization over k(k) can be performed using a gradient based recursive method. We adopt a RLS approach with a forgetting factor λ ∈ (0,1] to adapt to changing system dynamics. Define the gradient vector of the performance index with respect to the gains: .
An explicit expression for
ϕ(
k) can be derived by propagating the sensitivity of the predicted output through the SMC law and the predictor. In practice, a numerical approximation using finite differences is also feasible given the low dimensionality of
k(k). The gain update law is then:
where
is the RLS gain matrix computed as:
The matrix P(k) is the covariance estimate, initialized as P(0) = αI with α > 0 large.
To ensure stability and respect hardware limits, the optimized gains are projected onto a feasible region:
where
and
are lower and upper bounds determined from the system’s physical constraints (e.g., positive gains, maximum allowable switching frequency). The optimized gains
are fed directly into the SMC laws described in
Section 4. Importantly, the gain update frequency can be chosen lower than the control sampling frequency to reduce computational burden. In this work, we update the gains every
Nopt samples while the SMC runs at the converter switching frequency. Moreover, the prediction error
computed from the data driven predictor is used as a feedforward compensation term in the SMC laws. This combination of online gain optimization and predictive compensation significantly improves the transient response and robustness of the microgrid.
3.2. Proposed Algorithm
The proposed algorithm implements a recursive, online procedure for optimizing control gains within a data driven predictive strategy. The process begins with an initialization phase, during which the initial control gains, covariance matrix, forgetting factor, prediction horizon, weighting matrices, and admissible gain limits are established. Following this, the algorithm operates iteratively at each sampling instant.
At every iteration, current system measurements are collected and incorporated into updated data windows.
These data sets can be used, if necessary, to refine the predictor parameters through a RLS scheme. Using the updated predictor, future system outputs are estimated over the specified prediction horizon, allowing the computation of the instantaneous prediction error.
This prediction error serves as a basis for calculating the gradient of a defined performance criterion, which drives the adaptive adjustment of control gains via an update law that involves both the gain matrix and the covariance matrix. To ensure stability and boundedness, the updated gains are projected onto predefined feasible intervals. The optimized gains, along with the corresponding prediction error, are then provided to the SMC layer.
This procedure repeats at every sampling instant, enabling the controller to continuously adapt to changing system dynamics and operating conditions. A graphical representation of this entire process is presented in the flowchart of
Figure 2, which clearly illustrates each step of the online gain optimization method.
3.3. Implementation Details of the Data Driven Predictor and Gain Optimization
This subsection provides the detailed implementation parameters required for full reproducibility of the data-driven predictor and the online gain optimization algorithm described above. The prediction horizon is set to Np = 3 steps, with a sampling interval Δt = 0.1 s. This short horizon is chosen to capture the dominant dynamics of the DC microgrid while maintaining computational efficiency. The data window length is L = 20 samples, corresponding to 2 s of historical data, which provides sufficient information for the RLS estimation while remaining responsive to time-varying system behavior. The forgetting factor is set to λ = 0.99, which offers a trade-off between tracking capability (fast adaptation to changes) and noise rejection (smoothing of parameter estimates). This value is typical for RLS applications in power systems and was found to work well in our simulations. The matrix dimensions are specified as follows: Up ∈ R20×3: past control input matrix (20 samples × 3 duty cycles) Yp ∈ R20×1: past output matrix (20 samples × 1 bus voltage), Hu ∈ R1×60: input-to-output predictor matrix, Hy ∈ R1×20: output-to-output predictor matrix and Hf ∈R1×3: future input predictor matrix.
The RLS covariance matrix is initialized as P(0) = 103 I, where I is the identity matrix of appropriate dimension. This initialization ensures rapid initial convergence of the parameter estimates. To ensure persistency of excitation (PE) and prevent estimator divergence, we monitor the condition number of the covariance matrix P(k). The condition number is defined as κ(P) = λmax (P)/λmin (P), where λmax and λmin are the maximum and minimum eigenvalues of P, respectively. If the condition number exceeds a threshold of 104, we partially reinitialize the covariance matrix as P(k) = 103 I to prevent numerical instability and divergence of the estimates. This PE check is performed at each sampling instant before the gain update. The gain vector k(k) = [kpv (k), kb (k), kSC (k)]T is updated every Nopt = 2 samples (i.e., every 0.2 s) to reduce computational burden, while the SMC runs at the converter switching frequency. The gain bounds are set to ki,min = 25 and ki,max = 70 for all subsystems, based on physical constraints (maximum allowable switching frequency, actuator saturation limits). The performance index weights are chosen as Q = 10 (penalizing voltage deviations) and R = 0.1 (penalizing gain variations), providing a balance between fast voltage regulation and smooth gain adaptation. These weights were tuned empirically through preliminary simulations.
A sensitivity analysis was conducted to assess the impact of the forgetting factor λ on the predictor performance. Values of λ in the range [0.95, 0.995] were tested. The value λ = 0.99 was found to provide the best trade-off, yielding a prediction RMSE of 0.42 V compared to 0.51 V for λ = 0.95 and 0.48 V for λ = 0.995.
Algorithm summarizes the complete online gain optimization procedure.
Initialize:
Initial gains: k(0) = [40, 40, 40]T
Covariance matrix: P(0) = 103 I
Forgetting factor: λ = 0.99
Prediction horizon: Np = 3
Data window length: L = 20
Gain bounds: kmin = 25, kmax = 70
Performance index weights: Q = 10, R = 0.1
At each sampling instant k (every Δt = 0.1 s):
Collect measurements:
Update data windows:
Update predictor parameters using RLS with forgetting factor λ:
- ○
Compute prediction error:
- ○
Update parameter vector θ(k) using RLS update equations
Check persistency of excitation:
- ○
Compute condition number κ(P(k))
- ○
If κ(P(k)) > 104, reinitialize P(k) = 103 I
If k mod Nopt = 0 (gain update instant):
- ○
Compute predictions: for
- ○
Compute gradient: (finite differences or analytical)
- ○
Update RLS gain:
- ○
Update covariance:
- ○
Update gains: k(k) = k(k − 1) − K(k)ϕ(k)
- ○
Project gains: ki (k) = min(max(ki (k), ki,min), ki,max) for i ∈ {pv, b, SC}
Output: Optimized gains k(k) and prediction error ep (k) to AFSMC layer
The above parameters and procedures ensure that the data driven predictor and gain optimization algorithm can be reproduced with minimal ambiguity.
4. Adaptive Fuzzy Sliding Mode Control
This section develops the low level control layer of the proposed hybrid strategy. For each subsystem (PV, battery, SC), a sliding mode controller is designed to track a current reference provided by the EMS. To compensate for unknown nonlinearities, external disturbances, and modeling errors, fuzzy logic systems are embedded into the SMC laws. Moreover, the control gains are those optimized online by the data driven predictor (
Section 3), and a prediction error feedforward term is added to enhance robustness.
4.1. Fuzzy Logic System Approximator
A fuzzy logic system with center average defuzzifier, product inference, singleton fuzzifier, and Gaussian membership functions can approximate any continuous function on a compact set to arbitrary accuracy.
The FLS output is expressed as:
where
is the vector of adjustable consequent parameters, and
is the vector of fuzzy basis functions defined by:
with
. For each subsystem i, the fuzzy logic system uses n = 5 Gaussian membership functions per input variable (negative large, negative small, zero, positive small, positive large), with centers
uniformly distributed over the operating range. The number of fuzzy rules is
, where
nz is the dimension of
zi. For the battery subsystem,
, giving N = 25 rules.
The universal approximation property guarantees that for any continuous function h(z) on a compact set and any ϵ > 0, there exists an FLS such that where is an optimal parameter vector.
4.2. Predictive Fuzzy Sliding Mode Control Laws
For each converter, we define a current tracking error and a sliding surface. Let the reference current for subsystem i be (provided by the EMS or MPPT). The tracking error is: .
The sliding surface is chosen as . The control objective is to force Si → 0 in finite time, which implies perfect current tracking.
The dynamics of
Si are obtained from the converter model (1)–(3). For example, for the PV boost converter:
The same structure holds for the battery and SC, with their respective voltages and inductances.
In practice, the converter models contain unknown functions due to parameter uncertainties, unmodeled nonlinearities, and external disturbances.
For the PV subsystem, we lump all unknown terms into a continuous function
, where
may include states, parameters, and time. According to the universal approximation theorem, there exists an FLS such that:
where
is the approximation error. Similarly, for the battery and SC, we define
and
.
Because the optimal parameters θ∗ are unknown, we use online estimates and design adaptive laws. The fuzzy compensator output is .
The proposed control law for each converter consists of three parts: an equivalent control term (based on the nominal model), a switching term with the optimized gain, a fuzzy adaptive compensation term, and a prediction error feedforward term. For the PV converter:
where
is the time varying gain optimized by Algorithm (
Figure 2),
is the saturation function replacing the sign to reduce chattering, with
ϕ > 0 the boundary layer thickness,
is the estimate of the fuzzy parameters,
is a predictive compensation gain,
is the one step ahead prediction error of the DC bus voltage (from
Section 3). Analogous control laws are derived for the battery and SC:
The predictive term anticipates future voltage deviations and proactively adjusts the duty cycle, thereby improving transient response.
To update the fuzzy parameter estimates, we use a gradient descent approach with a σ-modification term to ensure boundedness. The adaptive law for
is:
where
is a positive definite adaptation gain matrix, and
is a small constant (σ-modification). The same structure applies to
and
. The term
comes from the Lyapunov design.
The integration of the sliding mode laws (11)–(13) with the adaptive fuzzy compensator (14) yields several synergistic benefits. The sliding mode component guarantees finite time convergence of the current tracking error despite matched uncertainties, providing inherent robustness. Meanwhile, the fuzzy compensator learns unknown nonlinear dynamics online, which reduces the required switching gain magnitude and effectively attenuates chattering. A predictive feedforward term, derived from the data driven model, anticipates future voltage deviations and pre-emptively corrects them. Finally, the time varying gains ki (k) are optimized online, ensuring that the controller maintains peak performance even as operating conditions change.
4.3. Reference Current Generation
The reference currents () are determined as follows:
A conventional perturb-and-observe maximum power point tracking algorithm provides to maximise the photovoltaic power output.
The high level DQL agent described in
Section 5 supplies the battery and SC references
and
.
The agent selects appropriate charging, discharging and load shedding commands while ensuring that the combined power from the battery and SC together with the PV source and the load satisfies the bus voltage dynamics given in Equation (4).
All current references are limited to respect physical constraints such as maximum allowable battery current and SC voltage range.
4.4. Summary of the AFSMC Layer
At each sampling instant, the AFSMC layer performs the following steps:
Step 1: Receive the optimized gains
from Algorithm (
Figure 2) and the prediction error
from the data driven predictor.
Step 2: Measure the actual currents Ii and compute the tracking errors .
Step 3: Update the fuzzy parameter estimates using (14).
Step 4: Compute the duty cycles di using (11)–(13) with the saturation function.
Step 5: Apply the duty cycles to the corresponding converters.
5. Augmented DQL Agent for High Level Energy Management
This section describes the learning based decision layer that coordinates the charging/discharging of the battery and SC, as well as load shedding actions. Unlike conventional DQL implementations that rely solely on instantaneous measurements [
18], the proposed agent receives an augmented state vector enriched with predictive information from the data driven model. This augmentation enables anticipative policies that proactively counteract future disturbances.
In a DC microgrid subject to rapidly varying solar irradiance and load demand, decisions based only on current measurements are inherently reactive. By the time a voltage deviation is detected, the system may already be experiencing significant transients. The data driven predictor provides short term forecasts of the DC bus voltage and PV power, as well as a prediction error. Incorporating these forecasts into the agent’s state allows it to learn actions that prevent or mitigate future voltage drops or over voltages, thereby improving stability and reducing battery stress.
5.1. Extended State Representation
At each decision step, the agent perceives an augmented state vector
s(t) ∈
S defined as:
where
is the current DC bus voltage,
is the current PV power,
is the current load demand,
is the battery state of charge,
is the SC state of charge,
is the one step prediction error of the bus voltage (computed by the data driven predictor),
is the one step ahead prediction of the bus voltage,
is the one step ahead prediction of the PV power.
All components are normalized to the interval [0, 1] using min-max scaling to facilitate neural network training. The inclusion of provides feedback on the predictor’s accuracy, while and give direct foresight of near future conditions.
The agent selects a discrete action at each decision step. The action space is designed to reflect the main energy management decisions:
where:
abat ∈ {charge, idle, discharge}—battery mode. The corresponding current reference is then set to a predefined value (e.g., 0.2C rate) or zero, scaled by the agent’s confidence (further refined by a lower level PI or SMC).
aSC ∈ {charge, idle, discharge} SC mode. The current reference is similarly determined.
ashed ∈ {no shed, partial shed, full shed} load shedding level. Partial shed reduces the load by a fixed percentage (e.g., 30%), while full shed disconnects non critical loads.
The action space consists of 3 modes for the battery modes for the SC , and 3 load shedding levels . The total number of discrete actions is 3 × 3 × 3 = 27. The agent can output a single integer action that encodes the combination, or use a multi-head architecture. In this work, we adopt a flat action space with 27 actions.
5.2. Reward Function Design
The reward function guides the agent toward desirable behaviors: maintaining voltage stability, preserving battery health, minimizing load shedding, and using the SC for transient peaks. The instantaneous reward
r(
t) is defined as:
Each component is detailed below. The Voltage Regulation Reward
where
αv > 0 is a weight. This term penalizes deviations from the nominal bus voltage.
The Battery SoC Preservation Reward: To avoid deep discharges and overcharging, a quadratic penalty is applied when SoC leaves a safe band
:
The Load Shedding Penalty: Load shedding is allowed but penalized to encourage energy balancing without disconnection:
where
l is the indicator function. Partial shedding incurs a smaller penalty than full shedding (by scaling
αshed accordingly).
The Battery Wear Reduction: High frequency charge/discharge cycles accelerate battery aging. To discourage unnecessary cycling, we penalize the absolute change in battery current reference:
The Predictive Term: A novel component that uses the prediction error and future voltage forecast to reward anticipative actions:
This term encourages the agent to keep both the current prediction error and the predicted future voltage small, thereby learning to act before a voltage deviation actually occurs. All weights αv, αsoc, αshed, αwear, αpred, αfut are positive constants chosen empirically to balance the objectives.
5.3. Deep Q-Learning Architecture and Training Algorithm
The agent uses a Deep Q-Network (DQN) to approximate the action value function , where θ are the network weights. The network architecture consists of:
Input layer: 8 neurons (one per augmented state component).
Two hidden layers, each with 128 neurons and ReLU activation.
Output layer: 27 neurons (one per action) with linear activation.
The DQN is trained using experience replay and a target network to stabilize learning. At each decision step, the agent selects an action according to an ϵ-greedy policy: with probability ϵ it explores uniformly, otherwise it chooses the action with the highest Q-value.
We adopt the standard DQN architecture [
19], where the action-value function is approximated by a neural network. Throughout the paper, we refer to this approach as DQL for consistency with the reinforcement learning literature.
The DQL agent is trained offline in a high fidelity MATLAB (2023 b)/Simulink simulator that implements the full DCMG model (5), the AFSMC layer, the data driven predictor, and the gain optimization algorithm. Training episodes are designed to cover a wide range of operating conditions:
PV power profiles: sinusoidal variations (diurnal cycles), step changes (cloud passages), and random fluctuations (stochastic weather).
Load profiles: slow ramps, abrupt steps, and pseudo random sequences.
Initial conditions: random SoCbat and SoCSC within [0.2, 0.9] and [0.3, 0.8] respectively.
The main hyperparameters used in the DQL model are as follows: the learning rate is set to 0.001, and the discount factor (γ) is fixed at 0.95. The exploration strategy begins with an initial exploration rate (ϵ0) of 1.0, which gradually decays to a final value (ϵmin) of 0.01 using a decay rate of 0.995 per episode. The experience replay buffer has a capacity of 50,000 transitions, from which mini-batches of size 64 are sampled during training. Additionally, the target network is updated every 100 episodes to stabilize learning. The model is trained over a total of 3000 episodes.
The training proceeds as follows:
Step 1: Initialize the Q-network with random weights, and copy them to the target network.
Step 2: For each episode:
- ○
Reset the simulation environment to a random initial state.
- ○
For each decision step t:
- ▪
Observe augmented state s(t).
- ▪
Select action a(t) using ϵ-greedy.
- ▪
Apply the action to the AFSMC layer (which translates it into current references and load shedding command).
- ▪
Run the microgrid simulation for the decision interval (0.1 s).
- ▪
Compute reward r(t) and observe next state s(t + 1).
- ▪
Store transition (s(t), a(t), r(t), s(t + 1)) in replay buffer.
- ▪
Sample a random mini-batch from the buffer and update the Q-network by minimizing the temporal difference loss:
- ○
End episode.
- ○
Decay ϵ.
- ○
Every 100 episodes, update target network weights.
After convergence (typically around 2500–3000 episodes), the agent’s policy is frozen and used for validation.
5.4. Integration with Lower Control Layers
The DQL agent outputs a discrete action
a(
t) = (
abat,
aSC,
ashed) at each decision step. The mapping from DQL actions to reference currents is defined as follows (
Table 1):
The reference currents are further adjusted based on the current SOC to prevent overcharging or deep discharge:
where
is the maximum battery current. The SC current reference is similarly adjusted with
:
The AFSMC layer then tracks these current references using the sliding mode control laws (11)–(13), which guarantee finite-time convergence of the actual currents
Ii to their references
. The load shedding command is implemented by reducing the load power according to the selected level:
where non critical loads are disconnected first in the case of partial shedding.
At each decision step, the agent outputs:
Battery mode: the corresponding reference current is set to a predefined value.
SC mode: similarly, is set to a fixed magnitude.
Load shedding command: the load is reduced by the specified amount.
These references are then tracked by the AFSMC layer. Importantly, the SC is primarily used for fast transients, while the battery supplies sustained power. The agent learns to coordinate both to keep the bus voltage stable and the SoC levels balanced.
The proposed high level layer is a DQL agent whose state is augmented with one step predictions and the prediction error of the DC bus voltage and PV power. It learns to select discrete actions (battery mode, SC mode, load shedding level) that optimize a multiobjective reward function. The agent is trained offline in a detailed simulation environment and then deployed in conjunction with the AFSMC and the online gain optimizer. This hierarchical, predictive, learning based architecture constitutes the main novelty of the paper.
7. Simulation Results
This section presents a comprehensive simulation study to validate the proposed data driven predictive hybrid control strategy. The performance of the proposed approach is compared against two strategies: (i) a conventional fuzzy logic based energy management system coupled with a fixed gain SMC (FLC method) [
15] and (ii) a standard DQL agent with only instantaneous state information combined with fixed gain SMC (DQL method) [
20]. The two strategies (FLC and DQL) both employ fixed- gain sliding mode controllers, whereas our method benefits from online gain optimization and an augmented DQL state.
All simulations are carried out in MATLAB using the averaged model of the DC microgrid described in
Section 2. The microgrid parameters are as follows: nominal DC bus voltage
vDCref = 48 V; PV boost converter
Lpv = 3 mH,
Cpv = 1000 μF; battery bidirectional converter
Lb = 3 mH,
Cb = 1000 μF, battery nominal voltage
vb = 36 V, capacity 100 Ah; SC converter
LSC = 2 mH,
CSC = 2000 μF, SC bank rated at 30 F; DC bus capacitance
CDC = 2200 μF; load resistive, variable from 0 to 2 kW.
For the proposed method, the adaptive fuzzy SMC gains are initialised to kpv (0) = kb (0) = kSC (0) = 40 with bounds [25, 70]. The saturation boundary layer thickness is ϕ = 0.15, predictive compensation gains are kp,pv = kp,b = kp,SC = 2. The data driven predictor uses a window length L = 20, prediction horizon Np = 3 and forgetting factor λ = 0.99. The DQL agent is trained for 3000 episodes.
The three test scenarios are concatenated on a single time axis: Scenario A (0–20 s, sinusoidal variations), Scenario B (20–40 s, step changes) and Scenario C (40–60 s, stochastic fluctuations). An additional extreme disturbance scenario (Scenario D) is simulated separately for robustness assessment.
Scenario A: Smooth Sinusoidal Variations.
The performance under slowly varying PV power and load. The proposed method reduces the voltage RMSE from 1.24 V (FLC) and 0.91 V (standard DQL) to only 0.42 V, an improvement of 66% and 54% respectively. The maximum voltage deviation is cut by more than half, from 2.7 V to 0.9 V. Battery stress (integral of |Ib|) is reduced by 62% compared to the fuzzy, and load shedding is completely eliminated. These gains stem from the adaptive fuzzy compensation and the predictive feedforward term, which anticipate the slow oscillations and keep the bus voltage tightly regulated.
Scenario B: Step Changes (PV drop and load step).
Figure 3 shows the DC bus voltage response of the three controllers during the concatenated scenarios, with a particular focus on the step disturbances at
t = 25 s (PV drop from 1000 W to 300 W) and
t = 30 s (load step from 500 W to 1200 W). The proposed method exhibits the smallest voltage dip (2.1 V vs. 4.3 V for FLC and 3.0 V for standard DQL) and the fastest settling time (0.22 s vs. 0.65 s and 0.38 s).
The improved performance is attributed to the online gain optimisation (
Figure 4) and the prediction error feedforward. As seen in
Figure 5, the battery SMC gain
kb (
t) rises from 20 to 28 within 0.04 s after the PV drop and further to 35 during the load step, providing extra robustness exactly when needed.
The battery current profiles (
Figure 5) confirm that the proposed controller uses the SC to handle transient peaks, reducing the peak battery current by 42% compared to FLC and by 25% compared to standard DQL. Moreover, the fuzzy compensator eliminates the chattering visible in the fixed gain SMC responses.
Scenario C: Stochastic Fluctuations.
Under random variations, the advantage of the augmented DQL agent becomes most evident. The proposed method reduces the voltage RMSE to 0.58 V (against 1.67 V for FLC and 1.22 V for standard DQL). Deep discharge events (SoC < 20%) are cut from 8.2 to 0.9, a reduction of 89%. Load shedding ratio drops from 9.8% to 1.7%, meaning the microgrid rarely needs to disconnect loads despite the highly variable generation and demand. The standard DQL without prediction still outperforms the fuzzy but cannot anticipate the rapid fluctuations, leading to occasional voltage sags and more frequent load shedding.
Scenario D: Extreme Disturbance (PV loss and heavy load step).
This scenario tests the robustness limits. At
t = 10 s the PV generator is disconnected, and at
t = 30 s the load steps from 800 W to 1800 W.
Figure 6 shows the DC bus voltage and the battery SoC. The proposed method keeps the voltage above 46 V (minimum 46.5 V) and recovers to 48 V within 0.25 s, while FLC drops to 43 V and takes 0.9 s to recover. The battery SoC decreases more slowly under the proposed control because the agent reduces unnecessary charging and prioritises the SC for transient support. Notably, the proposed method never activates load shedding during this scenario, whereas FLC exhibits a load shedding ratio of 8% and DQL of 3%.
Figure 7 zooms on the interval [24 s, 26 s] to show the evolution of
kb (
t) during the PV drop. The gain smoothly increases from 20 to 28, reaching its peak exactly at the moment of the disturbance (25 s). This behaviour confirms that the predictor successfully detects the impending voltage drop and triggers the gain update, which is then projected to stay within safe bounds [25, 70]. The forgetting factor
λ = 0.99 allows the predictor to track slow system changes without becoming overly sensitive to noise.
Figure 8 presents the learning curves (average reward per episode) over 3000 training episodes. The augmented DQL agent (with predictive state information) converges after approximately 2200 episodes, whereas the standard DQL agent requires 2600 episodes. The final average reward of the augmented agent is about 15% higher, indicating that the predictive information accelerates learning and leads to a better policy. After convergence, the policy is fixed and used for testing in Scenarios A–D.
The improvements of the proposed method. The voltage RMSE is reduced by 64% (vs. FLC) and 52% (vs. standard DQL); maximum voltage deviation is reduced by 58% and 43%; deep discharge events are reduced by 89% and 78%; battery cycling stress by 69% and 54%; load shedding ratio by 82% and 73%; settling time by 65% and 45%. These figures demonstrate the clear superiority of combining data driven prediction, adaptive fuzzy SMC, and augmented deep reinforcement learning.
Scenario E: Robustness Assessment under Unseen Conditions and Uncertainties.
To further validate the practical applicability of the proposed method, we conducted an additional robustness assessment under operating conditions not encountered during the DQL training phase. This scenario is designed to test the method’s resilience to uncertainties, measurement imperfections, and parameter variations that are typical in real-world microgrid applications.
The robustness assessment consists of 50 independent Monte Carlo simulations, each lasting 40 s, under the following challenging conditions:
(i) Unseen irradiance profile: The PV irradiance profile used in this scenario was deliberately constructed to be different from any profile seen during standard DQL training. It combines rapid fluctuations (simulating intermittent cloud cover) with sudden step changes (simulating weather fronts). The profile includes:
Initial steady state at 800 W/m2 (0–5 s)
Rapid fluctuations between 200–900 W/m2 with 0.5 Hz variations (5–15 s)
Abrupt drop from 900 to 300 W/m2 at t = 15 s
Stochastic variations with random amplitude (15–25 s)
Gradual increase from 300 to 700 W/m2 with superimposed noise (25–35 s)
Final step to 1000 W/m2 at t = 35 s
(ii) Measurement noise: Gaussian white noise with 0.5% standard deviation of the nominal value was added to all sensor measurements, including DC bus voltage, inductor currents, and PV power. This noise level is representative of typical measurement uncertainties in practical power systems.
(iii) Parameter variations: To test robustness against model uncertainties, the converter inductances were varied by ±20% from their nominal values:
Lpv ∈ [2.4,3.6] mH (nominal: 3.0 mH)
Lb ∈ [2.4,3.6] mH (nominal: 3.0 mH)
LSC ∈ [1.6,2.4] mH (nominal: 2.0 mH)
These parameter variations were randomly sampled for each Monte Carlo run and kept constant throughout the simulation.
(iv) Random initial SoC conditions: To evaluate the method’s ability to handle different initial energy storage states, the initial SoC values were randomly sampled from wide ranges: Battery SoC: uniformly distributed in [0.15, 0.95] and SC SoC: uniformly distributed in [0.2, 0.9].
These ranges cover both stressed conditions (low SoC) and fully charged conditions (high SoC), testing the energy management system’s decision-making capabilities across the full operating envelope. The load profile was kept identical across all 50 runs and consisted of a variable resistive load with step changes from 500 W to 1200 W at t = 20 s, and from 1200 W to 800 W at t = 30 s, representing typical household demand variations.
Figure 9 presents the box plots comparing the performance of the three methods (FLC, DQL, and the proposed method) over the 50 Monte Carlo runs. Two key metrics are shown:
(a) Voltage RMSE: The proposed method exhibits the lowest median RMSE (0.51 V) with the smallest interquartile range (IQR = 0.18 V), compared to FLC (median = 1.72 V, IQR = 0.42 V) and standard DQL (median = 1.15 V, IQR = 0.31 V). The maximum RMSE observed for the proposed method (0.89 V) remains below the minimum RMSE of FLC (1.32 V), demonstrating consistent superior performance even under worst-case conditions.
(b) Deep discharge events (SoC < 20%): The proposed method shows a median of 0.8 deep discharge events per simulation, with 90% of runs having fewer than 2 events. In contrast, FLC exhibits a median of 7.5 events (90th percentile: 12 events), and standard DQL shows a median of 3.2 events (90th percentile: 6 events). This represents a reduction of 89% compared to FLC and 75% compared to DQL.
The box plots clearly demonstrate that the proposed method maintains its performance advantages even under significant uncertainties. The lower dispersion of results indicates that the method is more robust to variations in operating conditions and parameter uncertainties. This robustness stems from three factors: (i) the online adaptation of SMC gains that automatically adjusts to changing dynamics, (ii) the fuzzy compensator that approximates unknown nonlinearities, and (iii) the augmented DQL agent that learns policies resilient to different initial conditions.
Notably, the proposed method never triggered load shedding in any of the 50 Monte Carlo runs, whereas FLC triggered load shedding in 38% of the runs (19 out of 50) and standard DQL in 14% of the runs (7 out of 50). This further confirms the robustness and reliability of our approach.
The average settling time after disturbances (defined as the time to recover within 2% of the reference voltage) for the proposed method was 0.28 s (standard deviation: 0.06 s), compared to 0.71 s (std: 0.15 s) for FLC and 0.42 s (std: 0.11 s) for standard DQL. These results indicate that the proposed method not only achieves better steady-state performance but also recovers faster from disturbances, even under uncertain conditions.
The computational cost of the proposed method was evaluated in terms of training time and online execution time. The DQL agent was trained for 3000 episodes, taking approximately 4.5 h. Regarding the online execution time per decision step (0.1 s), the data-driven predictor (RLS update) requires 0.8 ms, the gain optimization Algorithm requires 1.2 ms, the AFSMC layer requires 0.5 ms, and the DQL policy inference requires 0.3 ms. This results in a total online execution time of 2.8 ms per step, which is well below the 100 ms decision interval, confirming the feasibility of real-time implementation.
The simulation results clearly demonstrate that the proposed method outperforms the standard DQL baseline (non augmented, fixed-gain SMC) and the FLC method across all evaluated metrics. It achieves precise DC bus voltage regulation under sinusoidal, step, stochastic, and extreme scenarios, while significantly reducing deep battery discharges, cycling stress, and load shedding events. Thanks to online SMC gain adaptation and the integration of predictive information into the DQL agent, our approach responds faster to disturbances and learns anticipatory energy management policies. These advantages confirm the clear superiority of our predictive hybrid control strategy.