Next Article in Journal
Stabilizing Chaotic Food Supply Chains: A Four-Tier Nonlinear Control Framework for Sustainability Outcomes
Next Article in Special Issue
A DMAIC-Based Technology–Organization–Environment (TOE) Framework for Sustainable Industry 4.0 Adoption
Previous Article in Journal
Occurrence Patterns and Pollution Risk of Microplastics in Surface Sediments and Sediment Cores of the Three Gorges Reservoir, China
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Collaborative Energy Management and Price Prediction Framework for Multi-Microgrid Aggregated Virtual Power Plants

by
Muhammad Waqas Khalil
1,
Syed Ali Abbas Kazmi
1,*,
Mustafa Anwar
1,
Mahesh Kumar Rathi
2,
Fahim Ahmed Ibupoto
3 and
Mukesh Kumar Maheshwari
4
1
U.S.-Pakistan Center for Advanced Studies in Energy (USPCAS-E), National University of Sciences and Technology (NUST), H-12, Islamabad 44000, Pakistan
2
Department of Electrical Engineering, Mehran University of Engineering & Technology, Jamshoro 76062, Pakistan
3
Department of Mining Engineering, Balochistan University of Information Technology, Engineering and Management Sciences, Quetta 87300, Pakistan
4
Department of Electrical Engineering, Bahria University, Karachi Campus, Karachi 75260, Pakistan
*
Author to whom correspondence should be addressed.
Sustainability 2026, 18(1), 275; https://doi.org/10.3390/su18010275
Submission received: 12 November 2025 / Revised: 16 December 2025 / Accepted: 22 December 2025 / Published: 26 December 2025

Abstract

Rapid integration of renewable energy sources poses a serious problem to the functionality of microgrids since they are characterized by underlying uncertainties and variability. This paper proposes a multi-stage approach to energy management to overcome these issues in a virtual power plant that combines heterogeneous microgrids. The solution is based on multi-agent deep reinforcement learning to coordinate internal energy pricing, microgrid scheduling, and virtual power plant-level energy storage system management. The proposed model autonomously learns the optimal dynamic pricing strategies based on load and generation dynamics, which is efficient in dealing with operational uncertainties and maintaining microgrid privacy due to its decentralized structure. The efficiency of the proposed solution is tested on comparative simulations based on real-world data, which prove the superiority of the framework to the traditional operation modes, which are isolated microgrids and the energy sharing scenarios. The findings prove that the suggested solution has a dual beneficial impact on both virtual power plant operators and involved microgrids, as it leads to profit enhancement and, at the same time, system stability. This process facilitates the successful balancing of conflicting interests among the stakeholders at a time when the operation is low-carbon. The study offers an overall solution to dealing with complicated multi-microgrids and brings substantial changes in the integration of renewable energy, as well as the distributed management of energy resources. The framework is a scalable model that can be used in the future perspective of power systems with high-renewable penetration to address both economic and operational issues of the contemporary energy grids.

1. Introduction

The rising need for a sustainable and reliable energy supply and the worldwide push for decarbonization are part of heightened energy management difficulties. A microgrid (MG) operates independently or in cooperation with the main grid and satisfies the energy demands of residential and non-residential users. However, the integration of multi-microgrids (MMGs) in settlements and along networks of roads raises a number of particularly significant difficulties, such as the synchronization between the local energy source and storage systems with their load. This can be achieved through the creation of virtual power plants (VPPs) that work as aggregates of distributed energy resources (DERs) and optimize the working of the MGs in the interconnected grid system. Global climate change and greenhouse gas (GHG) emissions are compelling a transition towards decarbonized energy systems through a wider utilization of the distributed energy resources, but a significant barrier to this is the impracticality of managing and integrating directly too many DERs, which are geographically widespread and have diverse economic and social particularities, in the energy market. The virtual power plant represents one solution for transregional aggregation of the diverse DERs, such as renewable energy sources (RESs), energy storage systems (ESSs), flexible loads and electric vehicles (EVs) to facilitate their management in a tidy portfolio. The quest for sustainable growth in the 21st century has added a lot of complexity to energy planning, which now needs to optimize a lot of different techno-economic and environmental requirements at the same time [1].
An essential part of the contemporary energy landscape is the strategic planning of the infrastructure for electric vehicle charging. The authors in [2] proposed an integrated approach to EV charging station design that specifically takes into consideration the intricate influence of vehicle route patterns on the best location for stations. It was observed that their method, which makes use of an advanced two-stage adaptive cooperative evolutionary clustering algorithm, may successfully lower overall building costs without sacrificing service coverage. At the same time, MGs are becoming more and more important for maximizing the use of dispersed renewable energy sources. The two main hierarchical steps of this optimization process are individual MG optimization, which focuses on internal energy balance and economic dispatch, and MMG optimization, which controls power flows and transactions among connected MGs. In these systems, efficient and intelligent energy scheduling is crucial to maximizing the use of distributed energy resources. These scheduling strategies have a direct and significant impact on important performance metrics, including overall energy efficiency, supply reliability during grid failures, and maintaining power quality for sensitive loads [3]. The natural insecurities in renewable energy are being solved using advanced computing methods. In [4], the authors established an ideal scheduling platform of the MGs through the integration of a robust optimization platform and probabilistic prediction using machine learning. The hybrid solution solved severe problems of uncertainty in the periodicity of the renewable energy supply and sudden rises and falls in the energy market prices, enhancing the resiliency of the system in terms of both operations and finances. To further enhance real-time control, authors in [5] proposed a novel supervised learning approach for optimal power scheduling in isolated MGs. The optimal charge–discharge choices for battery energy storage systems were successfully simulated and replicated by this method, which significantly lowers daily operating expenses. Isolated, single-MG systems have inherent constraints and are frequently insufficient to provide the whole gamut of contemporary energy demands, especially at larger scales, as research and real-world applications in distributed energy have developed. The creation and spread of MMG systems were directly influenced by this realization. These MMG systems significantly increase the quality and penetration of renewable energy integration by facilitating power exchange and ancillary service support amongst nearby MGs. Compared to isolated MGs or conventional centralized grid systems, this interconnected architecture increases the diversity, resilience, and environmental sustainability of regional electricity consumption [6]. The authors in [7] presented a real-time optimization control technique based on the collaborative energy storage concept in order to handle the heightened complexity of these systems. This approach improves the MMG system’s overall resilience and dependability, making it more capable of withstanding and recovering from both external grid disturbances and internal errors.
In VPPs with many MGs, energy management is a complex area at the nexus of smart grids, control theory, optimization, and artificial intelligence (AI), particularly for energy price prediction. The state-of-the-art solutions make extensive use of algorithmic and system-level innovations. VPPs integrate distributed energy resources, such as MGs, energy storage, renewable energy sources, and programmable loads, to operate as a unified entity inside energy markets. The heterogeneity of MGs, which differ in resource types, controllability, and operational constraints, needs sophisticated management and coordination frameworks [8]. These systems address the stochastic character of renewable energy, flexibility in responding to market signals, and the attainment of common objectives such as cost reduction and resilience [9]. To maximize local and global goals (cost, reliability, and emission reduction), energy management systems (EMSs) for VPPs use distributed or hierarchical coordination techniques [10]. Primary, secondary, and tertiary levels are all integrated in hierarchical control, which handles market involvement, stability, and local optimization, respectively [11]. With established savings and stability benefits, distributed control, which is frequently based on cooperative game theory and multi-agent systems, allows for decentralized decision-making while promoting coalition building and prosumer engagement [12]. To improve operational and financial performance, modern EMSs also use electric vehicles, demand response (DR), and sophisticated supply and demand forecasting [13].
Smart contracts and market-based coordination facilitate local energy trade across and within MGs, including peer-to-peer and transactive energy markets [14,15]. In decentralized operations, blockchain and IoT technologies are being utilized more and more to guarantee traceability, transparency, and safe data exchange [16,17]. MG cluster management frequently entails optimizing multi-timescale scheduling, real-time, intra-day, and day-ahead using distributed or rolling horizon methodologies [18]. In order to participate in the market, scheduling algorithms must incorporate locational marginal prices, congestion control, and auxiliary service offering [19,20]. Predicting energy prices is essential because it allows for the best possible scheduling, bidding, and involvement in both day-ahead and real-time markets. Although they have been used historically for price forecasting, methods such as regression models, ARIMA, and exponential smoothing may have trouble with volatility and non-linearity [11]. Predictive accuracy is improved by supervised learning techniques like random forests, support vector machines (SVMs), gradient boosting, and ensemble approaches (including meta-ensemble learning) [21,22], especially when working with high-dimensional, non-linear data. For prosumer-based management systems, it has been demonstrated that ensemble approaches, which combine several weak learners, improve load and price prediction [23]. The capacity of recurrent neural networks (RNNs), particularly Long Short-Term Memory (LSTM) networks, to identify long-term dependencies in time series data has led to their widespread adoption. LSTMs perform better than classical approaches, especially when addressing fluctuating renewables and predicting under uncertainty [21,22]. Often utilized as a preprocessing step for forecasting, unsupervised methods such as K-means clustering are employed for load and prosumer classification, scenario reduction, and anomaly detection [24]. To better adapt to various MG characteristics and market structures, hybrid frameworks include optimization, machine learning, and clustering [11]. Recent developments train agents to react dynamically to market conditions by using reinforcement learning (RL) for adaptive pricing and bidding strategies [25]. While maintaining operational trust and data privacy, VPPs are able to predict and react to real-time price signals on their own [26].
In the electricity market, a VPP operator acts as an intermediary, facilitating communication between participating MGs and the main grid. Exclusive communication between each MG and the VPP strikes a balance between the interests of each MG and successfully protects their operational privacy [27]. A VPP, in contrast to independently run MGs, offers coordinated energy management for several MGs. It does this by managing a varied portfolio of dispersed energy resources, such as demand response assets, photovoltaics, wind turbines, and energy storage systems, and by encouraging operational synergies between them. The VPP has to balance the interests of all parties involved in the aggregation framework, in addition to overseeing energy exchanges with the main grid. Operational cost reduction, privacy protection, and operational simplification are the main goals for individual MGs. On the other hand, VPP operators are primarily driven by the revenue they make by enabling energy transactions between the main grid and the MG group [28]. There are several methods for this cooperation in the literature. A data-driven decision-making methodology was suggested in [29] to optimize scheduling for the following day based on past activity trends. An architecture for ensuring day-ahead commitments for MG support services was proposed by [30], who expanded on hierarchical control in VPPs. Similarly, using a lower-level tracking error model to improve trajectory accuracy and an upper-level quadratic program to reduce the rebound effect, the authors developed a reliable technique for aggregating thermostatically regulated loads [31]. In [32], a model-free deep reinforcement learning (DRL) method for controlling regenerative braking energy storage in railway systems was presented. Their approach, which was developed as a Markov Decision Process (MDP) with a multistage reward function, outperformed conventional techniques by increasing key energy objectives by more than 5%. In order to maximize renewable integration, flexible loads, and low-carbon economic operation while guaranteeing electric vehicle charging satisfaction, authors in [33] presented an improved Dueling Double Deep Q-Network for MG energy management that uses a mixed penalty function. Recent studies have further expanded the operational complexity and optimization needs of MMG systems. For example, integrated underfrequency load shedding strategies have been proposed to enhance stability in islanded MGs containing diverse load categories. At the same time, other works have explored bidirectional time-of-use electricity pricing through multilevel game-theoretic formulations for MMG environments [34,35]. Hierarchical and multi-agent coordination frameworks of operational optimization of an MMG system and VPPs have been explored in recent studies of DRL-based systems. Nevertheless, the majority of the current strategies concentrate on one of the coordination levels, either centralized VPP-level dispatch or microgrid-level scheduling, thus restricting the system-wide coordination [36,37]. Existing energy trading and coordination studies have mainly focused on the short-term operational goals of reducing costs and market efficiency using pricing and dispatching mechanisms, whereas longer-term techno-economic and environmental effects are usually handled separately [38].
With these developments, evident gaps exist in current DRL-based VPP studies. To begin with, internal price prediction and real-time MG dispatch are not jointly modeled in the majority of studies in a common learning environment. Second, preservation of privacy is usually mandated implicitly as opposed to explicitly by the decentralized observation and learning processes. Third, the coordination of heterogeneous MGs with different resource compositions under a single VPP structure remains insufficiently explored. Lastly, techno-economic and environmental effects are not usually considered simultaneously as part of DRL-based energy management models. Inspired by such constraints, this paper provides a three-step, multi-agent DRL model that combines internal pricing, distributed MG scheduling, and coordination of VPP-level energy storage and maintains MG privacy and allows conducting a comprehensive techno-economic–environmental performance analysis. The explicit modeling of the sequential and interrelated nature of decisions in a multi-stakeholder system is a significant divergence from conventional approaches. Second, the study presents a complex multi-agent DRL solution in which each MG and the VPP are represented as independent agents. The conflicting financial interests of the participating MGs and the VPP operator are automatically balanced by this method. Its model-free design, which functions without explicit system models, improves scalability and practicality. Third, to quantitatively illustrate the additional operational and economic advantages of the fully integrated framework, the work offers a comprehensive analysis against three precisely defined benchmark cases: isolated MGs operation and a centralized storage sharing model, and integrated VPP operation. Lastly, the suggested approach addresses a significant issue in real-world MG aggregation by providing a privacy-preserving solution since the DRL agents can develop efficient tactics based on local observations and minimal information exchange.
The main contributions of this study are presented below.
1.
A three-stage DRL-based framework for VPP energy management that coordinates internal pricing, MG dispatch, and centralized storage operations in a sequential, closed-loop process.
2.
A multi-agent DRL methodology that autonomously learns to balance the economic objectives of the VPP operator and the individual MGs, deriving a mutually beneficial scheduling scheme without relying on explicit system models.
3.
The enhanced performance and incremental value of the suggested coordinated strategy are demonstrated by an extensive performance assessment that compares three operational cases: isolated MGs, centralized storage sharing, and integrated VPP operation.
4.
A practical and expandable solution for MMG systems that improves financial results while protecting each participant’s privacy and independence.
5.
An integrated decision-making framework that combines sustainability assessment with techno-economic and emission-reduction objectives.
The sections of this paper are as follows: The system framework and models are introduced in Section 2 The MMG’s energy management strategies are formulated in Section 3. The case study and an analysis of the numerical results are covered in Section 4. The paper concludes with Section 5, which provides a comprehensive summary and directions for future research.

2. System Framework and Modeling

2.1. System Overview

This study examines a distribution market architecture that consists of a VPP aggregator and many geographically separated MGs. Each MG has its own set of distributed energy resources, including a diesel generator, a local energy storage system, residential and non-residential loads, and renewable generation (photovoltaics and wind turbines). On the other hand, the VPP aggregator is controlled by a central controller and runs a centralized energy storage system. MGs are only connected to the VPP, which acts as an intermediary for transactions with the wholesale power market. They do not trade energy directly. The given framework is based on a two-layered analysis structure. The first tier deals with the short-term operational coordination by means of the suggested multi-stage energy management framework by DRL, which decides upon hourly dispatch decisions, internal price signals, energy exchanges, and system operation profiles. The second layer is the assessment of the long-term techno-economic and environmental performance based on the operational results produced by the DRL coordination layer.

2.2. VPP Architecture Model

Figure 1 illustrates the architecture of the proposed VPP, which aggregates MGs. Each MG possesses an energy management system (EMS) for optimizing its local distributed energy resources, encompassing renewable sources, a local energy storage system, flexible loads, and electric vehicles. By combining these MGs, the VPP operator serves as the intermediary and takes part in the wholesale electricity market on their behalf. The VPP increases its own profitability by controlling the power trading between the MGs and taking advantage of their complementary generation, consumption, and storage profiles. This allows individual MGs, who are too small to participate directly in the market, to take advantage of their prosumer capabilities.

2.3. Equipment Modeling

This study develops models for the core components of the MMG system: PV, WT, DG, and ESS. These models are designed to accurately capture the operational characteristics and performance of each unit. In this study, the power output of renewable energy sources, such as WT and PV panels, is modeled as non-dispatchable and must be fully consumed. Consequently, the operational cost of RES generation is neglected [39].

2.3.1. PV Model

The output power of a photovoltaic system, denoted as P P V , is typically calculated using [40]:
P P V = P p v r I ( φ , m , d , h ) 1000 1 + θ T ( T p v T r )
where P p v r is the photovoltaic (PV) system’s rated power, I ( φ , m , d , h ) is the solar irradiance on the surface, θ T is the temperature coefficient, and T p v and T r are the PV system’s actual and rated temperatures, respectively. The rated temperature is set to be 25 °C.

2.3.2. WT Model

The output power of the wind turbine, denoted as P W T , is determined as [40]:
P W T = 0 , if V < V c i or V > V c o P w t r V r 3 V c i 3 P w t r V r 3 V c i 3 V c i 3 , if V c i V V r P w t r , if V r V V c o
where P w t r is the rated power of the wind turbine (WT), and V represents the wind speed, with the subscripts c i , c o , and r denoting the cut-in, cut-out, and rated wind speeds, respectively.

2.3.3. DG Model

A diesel generator (DG) system’s power is determined by its efficiency, fuel consumption, and machine features. The following expression defines it [41].
P D G = F d g A g × P d g . o u t B g
where F d g denotes the fuel consumption rate, P d g . o u t is the diesel generator’s output power (kW), and the coefficients A g and B g characterize the fuel consumption per unit of output power (L/kWh). The power of the diesel generator (DG) in MG i is constrained as follows:
P i D G m i n P i D G ( t ) P i D G m a x

2.3.4. ESS Model

The battery model is characterized by its state of charge (SOC) formulated as [42]:
SOC i ( t ) = SOC i ( t Δ t ) P i B ( t ) Δ t η i B E i if P i B ( t ) 0 SOC i ( t Δ t ) η i B P i B ( t ) Δ t E i if P i B ( t ) < 0
In the context of the battery’s operation, SOC i represents the state of charge of the battery in MG i. The power of the battery P i B can either be positive, indicating discharge, or negative, indicating charging. The term η i B refers to the efficiency of the battery, while E i represents the nominal capacity of the battery in MG i. Lastly, Δ t denotes the time step. The state of charge S O C i of the battery is limited to a given range to prevent over-stress and early aging. This limit is defined as follows:
SOC i , min SOC i ( t ) SOC i , max

2.4. Load Modeling

This study categorizes electricity demand into residential and non-residential loads. The analysis focuses on a region where residential demand, encompassing essential household services such as lighting, cooking, and water pumping, constitutes the vast majority of total consumption. The non-residential sector, including facilities like EV charging stations, parks, service areas, and toll stations, accounts for a minor share, reflecting the area’s sparse population and the consequent lack of large-scale commercial or industrial infrastructure. This work models two distinct demand classes: inelastic baseline load and flexible load. The study specifically focuses on transferable load (TL) as the flexible component, which can be temporally shifted at constant power. TL operation is governed by contracts that obligate the MG to pay reimbursements to users for any scheduled load shifts. The total load demand P L D is defined as
P L D = P H + P S + P T + P E V

2.4.1. Residential Load Model

P H = h = 1 n i = 1 m P i , h × D i , h
where P H represents the total residential load in a community, P i , h is the power rating of the i -th device in home h, D i , h is the duty cycle of the i -th device in home h, n is the number of homes, and m is the number of devices in each home.

2.4.2. Service Area Load Model

P S = t = 1 T i = 1 m q i 1 ( t ) + j = 1 n q j 2 ( t )
where P S is the energy utilized by the service areas, q i 1 ( t ) is the energy utilized by the i -th service area for one hour (kWh), q j 2 ( t ) is the energy consumption of the j -th parking area for one hour (kWh), m and n are the number of service and parking areas of the highway networks, and T represents the total number of hours in a year [43].

2.4.3. Toll Station Load Model

P T = P t 1 × T t 1 × n t
where P T is the energy consumption of the highway toll station, P t 1 is the load of the toll station, at 0.04 kW; T t 1 is the operation time of the toll station (hour) and n t is the total number of toll stations [43].

2.4.4. Electric Vehicle Charging Load Model

The electric vehicle charging load at time t of the charging station j is [44]:
p i , t = 0.67 p i , 0 + 9.15 S O C i , t 0 S O C i , t < 0.2 p i , 0 0.2 S O C i , t < 0.8 p i , 0 + 4.55 ( S O C i , t 0.8 ) 0.8 S O C i , t 1.0
P E V j , t = i = 1 h j p i , t j
where p i , 0 denotes the rated charging power of the i -th EV (kWh), S O C i , t represents its state of charge at time t, p i , t is its charging power at time t (kW). The charging power for the i -th EV at time t located at station j is specified as p i , t j (kW). Let h be the total number of EVs scheduled for charging at station j. And P E V j , t is the aggregate charging load at time t of the charging station j (kW).

2.5. Energy Price Modeling

To ensure financial viability, the VPP operator constrains the internal energy price for each MG within a bounded range, determined by prevailing system conditions according to Equation (13). To foster participation, MGs may accept an internal price that is lower than the external market rate, a condition formalized in Equation (14) [36].
ε · E x t p r i c e m e a n I n t p r i c e E x t p r i c e E x t p r i c e m e a n · ε
ξ t T E x t t p r i c e = t T I n t t p r i c e
where I n t p r i c e is price of internal electricity, E x t p r i c e is price of external or wholesale market electricity, and E x t p r i c e m e a n is the average external price. ε is the parameter that defines the allowable fluctuation of the internal electricity price, and ξ is the internal pricing preference factor.

2.6. Economic Modeling

Vital financial parameters are used to find out the economic feasibility of the proposed hybrid system. The key ones are the net present cost (NPC) and the levelized cost of energy (LCOE), which cannot be separated concerning performance metrics such as the payback time (TBP), internal rate of return (IRR), and the return on investment (ROI).

2.6.1. Net Present Cost

The main economic objective of this research is to reduce the net present cost (NPC). The NPC includes the overall costs of the project over the course of its life span, 25 years, and can be presented as shown in the following Equation [41]:
N P C = C I n v + C O M + C R e p + F C d g
where C I n v is the investment cost, C O M is the operation and maintenance cost, C R e p is the replacement cost and F C d g is the fuel cost.

2.6.2. Cost of Energy

The levelized cost of energy (LCOE) is an important economic indicator that calculates the average cost per kilowatt-hour (kWh) of generated energy [41]:
L C O E = N P C × C R F t = 1 T P load ( t )
where CRF stands for the capital recovery factor (converting the initial cost to annual capital cost), and P l o a d represents the power load. The CRF is calculated as [41]:
C R F ( i r , n ) = i r × ( 1 + i r ) n ( 1 + i r ) n 1
where i r is the interest or discount rate and n is the number of periods. This study uses a discount rate of 9.75% [45].

2.6.3. Performance Parameters

The key financial objectives of this analysis are net present cost (NPC) and levelized cost of energy (LCOE). These measures are closely related to important performance indicators such as payback period (TBP), internal rate of return (IRR), and return on investment (ROI) [46]. In this study, the ideal decision is one that minimizes both the levelized cost of energy (LCOE) and the net present cost (NPC). As a result, these economic indicators are highly correlated with total system performance. The payback period (TBP) is the number of years required to repay the project’s initial expenditure. Equation (18) describes the TBP computation procedure.
T B P = C i n + t = 1 T C a n n ( t ) ( 1 + i r ) t · t = 1 T B a n n ( 1 + i r ) t × T
In this formulation, T represents the system payback period, which is defined as the minimum number of years needed for total income to exceed the initial investment. C i n is the initial cost of investment, C a n n refers to the total annual cost, and B a n n represents the annual net cash flow.
The internal rate of return (IRR) is the discount rate that results in a net present value of zero for all cash flows in an investment, as stated in Equation (19). It is a fundamental metric for determining investment profitability.
I R R ( % ) = C . f l o w ( 1 + i r ) t C c a p e x × 100
where initial investment in the project is represented by C c a p e x and C . f l o w is net cash flow for a particular period. Similarly, the return on investment ROI, defined in Equation (20), is calculated as the annual change in nominal cash flows relative to the capital cost difference. This metric reflects the extent of long-term cost savings in comparison to the initial investment.
R O I ( % ) = 1 n · C c a p n · C c a p , r e f i = 0 n C i , r e f C i × 100
where C c a p e x , r e f and C c a p e x denote the initial costs of the base case and optimized system, respectively. The variable n is the project lifetime, while C i , r e f and C i represent the annual cash flows for the base-case and optimized systems, respectively.

2.7. Environmental Modeling

The analysis primarily focuses on CO2 emissions, as they have the greatest impact on the overall greenhouse gas (GHG) emission factors. However, the estimation of carbon emissions also incorporates components related to fuel consumption and system modeling. The mathematical formulation for carbon emission calculation is presented in Equation (21).
E = E f c × A r × ( 1 η E R ) 100
where E f c denotes the emission factor, A r represents the activity rate, and η E R signifies the overall emission reduction efficiency.

3. Energy Management Strategy for MMG

The suggested framework involves a coordination mechanism based on a sequence of operations, which involves internal price setting, distributed MG dispatch, and centralized storage optimization. This multi-agent, model-free, deep reinforcement learning algorithm allows the system to self-manage the competing goals of the VPP and the MGs. The privacy of the suggested framework is achieved by the localized observation and decentralized decision-making. The non-sensitive aggregated variables (net load and price-related feedback) are exchanged by each MG only, and no internal data (generation schedules, DG fuel characteristics, ESS parameters, EV charging behaviors, flexible load attributes, and cost structures) is shared with other parties. The VPP agent does not have any form of interaction with the MG states but only with internal price signals, so that proprietary technical or economic information is not exposed. Such a design allows the DRL agents to be informed about policies that only consider information pertinent to their own environment, which ensures the privacy of MG without disorganized work.

3.1. Multi-Agent-Based Deep Reinforcement Learning Strategy

This multi-stage framework is structured around a deep reinforcement learning (DRL) process comprising three sequential phases: retail price setting, MG optimization, and VPP energy storage system scheduling. The model assumes perfect foresight of daily average external market prices and the daily average load profiles for each MG as predetermined inputs. The decision-making problem for each agent is formally defined through a Markov decision process (MDP) characterized by the following components.

3.1.1. Reinforcing VPP Efficiency: MDP-Based Internal Pricing Strategy

VPP Agent 1 establishes internal retail prices for the MGs by dynamically reconciling their specific demand profiles with prevailing external market prices.
  • State space model:
In the proposed reinforcement learning model for energy management, the state at time t captures the condition of the MG, including the state of the internal pricing system. Guided by the electricity price set by the VPP, the MG will optimize its electricity consumption strategy. It will then feed back the optimized electricity demand (price-dependent) to the VPP.
S i , t = t , E x t t p r i c e η i , t , t E x t t p r i c e I n t t p r i c e
η i , t = P i , t L D P i , m e a n L D
where η i , t represents the scheduled load normalized by its average value for MG i at time t, and t E x t t p r i c e I n t t p r i c e defines the net cumulative internal price adjustment across the scheduling horizon.
  • Action space model:
The variable A i , t represents the normalized internal electricity price for MG i at time t, a formulation that enhances neural network training. The actual price is recovered via denormalization in Equation (25).
A i , t = A i , t i n , A i , t i n [ 1 , 1 ]
I n t t p r i c e = E x t t p r i c e · 1 + A i , t i n · ϵ
  • Reward function model:
The internal pricing mechanism aims to maximize profit from electricity transactions. However, considerable daily fluctuation in MG net loads causes high volatility in the VPP’s reward signal, potentially destabilizing the training process. To address this instability, the following reward function is proposed:
R t = P i , t n e t l o a d · I n t t p r i c e P i , m e a n L D · I n t m e a n p r i c e
The reward signal is normalized with a baseline correction term, which is calculated as the product of the average load and average price. This modification seeks to reduce the volatility in rewards caused by daily fluctuation in load and pricing.

3.1.2. MDP-Driven Demand Response Strategy

The MG operator must determine the operational schedule for all controllable assets; consequently, this study formulates their decision-making process as an MDP.
  • State space model:
The state of the MG at time t is observed as follows:
S i , t = t , I n t t p r i c e , η i , t , S O C i , t 1 , P i , t 1 D G λ t D G t P t T L
The state includes the load ratio η i , t , (planned to average), the previous energy storage state of charge S O C i , t 1 , the prior diesel generator output P i , t 1 D G , and the cumulative amount of transferred load t P t T L .
  • Action space model:
The action A i , t defines the dispatch setpoints for all schedulable resources within MG i at time t.
A i , t = A i , t E S S , A i , t D G , A i , t T L , A i , t E S S , A i , t D G , A i , t T L [ 1 , 1 ]
where A i , t E S S , A i , t D G and A i , t T L denote the normalized control actions for the energy storage system, diesel generator, and transferable load, respectively.
P t E S S = A i , t E S S · δ c · P m a x c + δ d · P m a x d P t D G = P t 1 D G + A i , t D G · δ u p · R u p + δ d o w n · R d o w n P t T L = A i , t T L · P t L D · δ + · γ T L + + δ · γ T L
  • Reward function model:
The objective function minimizes the total system operating cost. Consequently, the reward is defined as the negative value of the MG’s total operational expenditure. To stabilize training against daily load and price volatility, subtract a baseline compensation term from the reward. This term is calculated as the product of the average load and the average price.
R i , t = C t b u y C t D G C t T L + P i , m e a n l o a d · I m e a n p r i c e
C t b u y = I n t t p r i c e · P i , t n e t l o a d C t T L = λ T L · | P i , t T L | C t D G = λ t d i e s e l · P i , t D G P i , t n e t l o a d = P i , t l o a d + P t E S S P t D G + P t T L
where C t b u y , C t D G , and C t T L denote the costs associated with electricity procurement, DG operation, and transferable load scheduling, respectively. Compensation of transferable load λ T L is 0.05 $/kWh [28].

3.1.3. Reinforcing VPP Efficiency: MDP-Based ESS Strategy

After determining the aggregate MG demand, the VPP operator develops a dispatch plan for its centralized energy storage system (ESS). This function relates to VPP Agent 2.
  • State space model:
The system state observed at time t is defined as follows:
S t = t , E x t t p r i c e , η t , S O C i , t 1
where η t denotes the normalized total load, calculated as the ratio of the load at time t to the mean load.
  • Action space model:
A fundamental limitation for action denormalization is that the ESS’s discharge power, as determined by A t , must be limited to the MMG’s entire net load at time t.
A t = A V P P , t E S S , A V P P , t E S S [ 1 , 1 ]
P V P P , t E S S = max i P i , t n e t l o a d , A V P P , t E S S · δ c · P m a x c + δ d · P m a x d
  • Reward function model:
The goal of the VPP’s energy storage system dispatch is to meet the aggregate power demand of the MGs at the lowest cost. As a result, the reward function is expressed as follows:
R t = E x t t p r i c e · P T n e t l o a d + P V P P , T E S s + i P i , m e a n l o a d · E x t m e a n p r i c e
The functions of reward in Equations (26), (30) and (35) will be designed so that they are the inverse of the total operation cost to be minimized by each agent. In the case of the MG agents, the reward is combined linearly without any extra weighting factors on the cost of electricity purchasing, diesel generator (DG) fuel cost, and transferable load (TL) compensation cost. This takes care of the fact that every component of cost will contribute proportionately based on the actual economic value of the same. The baseline terms that are incorporated in the reward formation are merely presented as a variance-reduction mechanism to reduce the volatility of rewards in the training procedure and do not affect the fundamental optimization objective. Any physical or operational constraints, such as energy storage state-of-charge (SOC) limits, DG power output and ramp-rate constraints, and transferable load (TL) limits, are imposed by explicit action-space limitation. The infeasible actions produced by the DRL agents are limited before execution, such that the state of the systems never leaves the realistic operating range. The method does not involve any fines or rewards in the form of a penalty and enhances the stability of training, besides ensuring physically possible control measures.

3.2. Implementation of Internal Price Setting

The internal electricity price is modeled as a decision variable that is set by the VPP using a DRL-based internal price setting mechanism. The VPP pricing agent chooses the internal price at every time step to maximize long-term operational goals given predefined pricing constraints, and does not approximate any external or internal price signal. To align the VPP operator and participating MGs, internal electricity pricing must follow the constraint provided in Equation (14). Non-compliance with this regulation would result in hefty financial penalties for the VPP operator. Since only the enforcement with rewards cannot guarantee strict adherence, this work uses a powerful constraint-handling framework where the direct-action space mapping is utilized.
The cross-temporal character of the internal price restriction poses a substantial hurdle for direct implementation within the DRL algorithm. To overcome this, the restriction in Equation (14) is reformed into Equation (36), integrating the average external price. This modified version looks at the cumulative change in prices only, and a solution can be easily managed. These three steps are then followed to ensure rigorous constraint satisfaction at every operating interval.
t T E x t t p r i c e I n t t p r i c e = 1 ξ · E x t m e a n p r i c e · T
  • Step 1: Determine the final feasible target
Compute the total cumulative price adjustment permitted up to the final time period T. This establishes the global constraint that the entire sequence of adjustments must satisfy.
  • Step 2: Compute a feasible action corridor
Working backwards from T recursively calculate the feasible price range for each time step. This ensures every interim adjustment remains within the limits (Equation (13)) and can still achieve the final target from Step 1.
  • Step 3: Project agent actions onto the corridor
At each time step, intersect the DRL agent’s proposed action with the corresponding feasible range from Step 2. The final executed action is the agent’s raw action projected onto this admissible set, ensuring all constraints are met.
This study employs the proximal policy optimization (PPO) algorithm, a deep reinforcement learning (DRL) method, to solve the formulated problem. PPO’s core principle involves constraining the divergence between successive policy updates to prevent destabilizing large policy shifts, typically achieved through a specialized clipping mechanism. In reinforcement learning, a policy defines the agent’s strategy for interacting with its environment. Formally, it is a function that maps states from the state space to actions from the action space, dictating which action to take in any given situation. Figure 2 shows the PPO architecture. This design effectively balances the exploitation of improvements from a new policy with the stability inherited from the previous one. Unlike value-based approaches such as Deep Q-Network (DQN), which derive policies from learned value functions, policy-based methods like PPO directly optimize the action-selection policy. This is implemented by representing the policy π θ with parameters θ and optimizing these parameters via gradient ascent. The PPO algorithm updates θ by combining two key objective functions: one for policy loss and another for value function loss.
L C L I P ( θ ) = E ^ t min r t ( θ ) · A t , clip r t ( θ ) , 1 ϵ , 1 + ϵ · A t
L V F ( θ ) = E ^ t V target V ( s t ; θ v ) 2
The objective is optimized using the probability ratio r t ( θ ) , which compares the current policy to the previous policy. This ratio is weighted by the advantage estimate A t and constrained by a clipping function dependent on a predefined threshold ϵ . The expectation E is computed over the current sample batch. Simultaneously, the value function is updated by minimizing the error between the estimated value V ( s t ; θ v ) and the target value V t a r g e t . In the PPO algorithm, the policy and value function parameters are updated by optimizing their respective gradients. By interacting with the environment over multiple iterations and continuously refining the policy, the algorithm gradually converges to the optimal action strategy.

3.3. Model Training

Several DRL agents are used in the multi-stage energy management framework, which calls for a unique training process. An overview of the proposed methodology is presented below.

3.3.1. MG Agent Training Phase

In the first training phase, when internal price setting is not established, MG agents are trained using external electricity prices. VPP agents are not involved in the process and are still inactive at this point. Based on internal operational data and external price signals, each MG agent creates scheduling decisions. It then carries out the appropriate activities for its controllable devices. Through reinforcement learning, the results of these activities yield incentives that are utilized to adjust the parameters of the corresponding MG agent. Each MG agent trains alone in its own environment during this phase since MGs are not able to communicate with one another. As shown in Figure 3, this isolation leads to the creation of unique operating methods suited to the unique device configurations and features of each MG.

3.3.2. VPP Agent Training Phase

The MG agents are included as part of the environment for later VPP agent training after they have been trained initially. Two independent agents share the responsibilities of the VPP aggregator in order to manage different functions: Agent 1 uses external price data to determine internal electricity price settings, and Agent 2 manages the energy storage system (ESS) dispatch based on external prices, net power balance, and total load. In order to preserve environmental stability during the training phase, these VPP agents go through independent training procedures without direct inter-agent contact. Figure 4 illustrates how MG agents train VPP agents.

3.3.3. MG Agent Training Phase by VPP

Upon completion of VPP agent training, these agents are incorporated into the environment for fine-tuning the MG agents’ parameters. In this phase, MG agents utilize the internal price settings provided by VPP Agent 1 as a primary input to determine their operational schedules. The resulting demand profiles from all MGs are subsequently aggregated and fed to VPP Agent 2 as part of its input state, completing the integrated feedback loop illustrated in Figure 5.

3.3.4. Convergence

Repeat steps Section 3.3.2 and Section 3.3.3, alternately training the MG and VPP agents. After each training round, test the agents and record the results. MG agent parameters are fixed when training VPP agents and vice versa. Monitor reward variation to determine convergence. If the difference in average rewards between rounds is less than 5%, training is considered converged and ends; otherwise, training continues. The pseudo-code for the multi-agent three-stage DRL training procedure is listed as Algorithm 1. A structured multi-agent three-stage learning process is proposed to train the framework, which includes (i) independent pre-training of MG agents, (ii) independent pre-training of the VPP internal price setting agent and the VPP ESS agent and (iii) alternating joint fine-tuning of all the agents in the integrated environment. The individual agents of the MGs are initially trained to serve optimal local scheduling with the help of external electricity costs, but not with VPP coordination. Equally, the VPP pricing and ESS agents are trained with aggregated responses of MG. In the joint fine-tuning stage, every agent communicates in a sequence in every episode by using internal price setting, distributed MG scheduling and VPP ESS control. A moving-average reward criterion is used to measure convergence, i.e., convergence is attained once the relative change in average reward is less than a set threshold across successive episodes.
Algorithm 1 Multi-agent three-stage DRL training procedure.
  1:
Initialize MG agents { π M G , i } , VPP pricing agent π price and VPP ESS agent π ESS
  2:
Initialize policy network π θ and value network θ v
  3:
Set hyperparameters: learning rate α , discount factor γ , clipping factor ϵ , and number of episodes
⁠ 
  4:
Phase I: Independent Pre-Training
  5:
for each MG agent i do
  6:
    for episode = 1 to 200,000 do
  7:
        Observe local MG state S i , t
  8:
        Receive external market price E x t t Price
  9:
        Select scheduling action A i , t π M G , i
10:
      Execute MG operation and observe reward R i , t
11:
      Update policy π M G , i using PPO
12:
    end for
13:
end for
14:
for episode = 1 to 200,000 do
15:
     Observe aggregated MG demand and external price E x t t Price
16:
     Select internal electricity price Int t Price π price
17:
     Observe reward R i , t
18:
     Update policy π price using PPO
19:
end for
20:
for episode = 1 to 200,000 do
21:
    Observe aggregated net load and external price E x t t Price
22:
    Select ESS action A t ESS π ESS
23:
    Observe ESS reward R t ESS
24:
    Update policy π ESS using PPO
25:
end for
⁠ 
26:
Phase II: Alternating Joint Fine-Tuning
27:
for episode = 1 to 20,000 do
28:
    Stage 1: Internal Price Setting
29:
    Set internal electricity price Int t Price π price
30:
    Stage 2: Distributed MG Scheduling
31:
    for each MG agent i do
32:
           Observe state S i , t and price I n t t Price
33:
           Select scheduling action A i , t π M G , i
34:
           Execute MG operation
35:
    end for
36:
    Stage 3: VPP ESS Scheduling
37:
    Observe total MG net load
38:
    Select ESS action A t ESS π ESS
39:
    Compute rewards for all agents
40:
    Update { π M G , i } , π price , and π ESS using PPO
41:
end for
⁠ 
42:
Convergence Criterion
43:
Monitor moving-average rewards for all agents
44:
Terminate training if relative reward change < ϵ for K consecutive episodes

4. Case Study and Numerical Results

This section presents a comprehensive performance evaluation of the proposed framework through numerical simulations. A workstation with an Intel i7-8700 processor and 16 GB of RAM was used to develop and run the full model in a coordinated environment using MATLAB 2024a and Python 3.10. A 60-min resolution is used in the scheduling horizon. As shown in Figure 6, eight typical sites from western Pakistan were chosen for case studies in order to verify the efficacy of the model. The corresponding meteorological data [47], including average temperature, solar radiation, and wind speed profiles, are presented in Figure 7, while Figure 8 illustrates the annual energy demands. Energy load profiles were developed using standardized procedures based on consumption characteristics [48].

4.1. Simulation Settings

This study simulates a virtual power plant comprising eight heterogeneous MGs, with their respective configurations detailed in Table 1. Table 2 specifies the simulation parameters for the proposed strategy. Wholesale electricity prices are based on 2024 trading data from the PJM Interconnection [49]. The dataset is partitioned into 80% for training and 20% for testing. The multi-agent deep reinforcement learning (DRL) models were trained on the training set and validated on the testing set, while all comparative benchmark methods were evaluated directly on the testing data.
This study proposes a multi-stage DRL-based framework for VPP energy management. To thoroughly evaluate its performance, three distinct operational cases are analyzed and compared.
  • Case 1 (baseline-isolated MGs): involves neither coordination nor sharing of energy. Every MG functions independently, balancing supply and demand using just its own resources. Every MG has its own self-reliant operation, where none of them is coordinated at VPP or even with internal price settings. Every MG has the same DRL-based local scheduling agent as that of the proposed framework, except that the agent does not react to an internal price setting, but to the external price of the electricity market itself. There is no information sharing and coordination between MGs, and VPP does not affect MG functions. The case is an example of an intelligent yet non-coordinated benchmark, which makes it possible to evaluate the contribution of VPP-based coordination instead of the sophistication of controllers.
  • Case 2 (centralized storage sharing): In this case, MGs function autonomously without coordinated dispatch or internal pricing setting. To enable aggregated power balancing, a centralized ESS is implemented at the VPP level. To maintain methodological consistency, the centralized ESS is managed using the same DRL-based control technique as in the suggested framework. MGs continue to optimize locally depending on external prices since they do not receive internal price signals. Internal price-based coordination is not included in this scenario, which isolates the impact of centralized storage sharing.
  • Case 3 (integrated VPP operation): Include internal energy price settings for MGs established by the VPP agent, and it represents the entire suggested structure. The multi-stage method in this case allows for completely coordinated functioning.
The core of Case 3 is the following sequential decision-making process:
  • Stage 1 (VPP pricing): Based on operational information gathered from each MG, including load, generation, and storage status, as well as the wholesale energy price, the VPP agent calculates the internal retail electricity price for the MGs.
  • Stage 2 (MG scheduling): Each MG’s energy management system modifies the operation of its internal resources, such as local ESS, dispatchable generators, and any flexible loads, in response to the retail price. The VPP agent then receives the revised net load data.
  • Stage 3 (Centralized ESS & market clearing): Based on the total net load of all MGs, the VPP agent determines the charging/discharging schedule for its centralized ESS. The net energy balance, which is determined by subtracting the total MG load from the central ESS dispatch, is then used by the VPP to conduct wholesale market transactions.
The goal of this model-free DRL framework is to automatically balance the conflicting goals of the MGs and the VPP in order to arrive at an ideal scheduling plan that benefits both parties. The incremental advantages of the suggested coordinated technique over conventional isolated operation and simpler storage-sharing models will be illustrated through a comparative study of these three scenarios. The system’s economic performance, self-sufficiency, and MG interaction power are evaluated across several scenarios to validate the proposed technique. The research assumes that there is enough redundancy in the energy transmission system to reliably support the necessary inter-MG energy exchanges.

4.2. MG Performance

The scheduling results for the non-cooperative MG configuration in Case 1 are shown in Figure 9, emphasizing the difficult energy balancing dynamics. To handle the varying load demand, the MG is forced to rely solely on its internal resources, mainly limited energy storage (ESS), a diesel engine, and intermittent PV and wind power. The ESS absorbs extra energy during times of high renewable generation, but the MG is still susceptible to supply and demand imbalances. The activation of the DG highlights the system’s need for costly and carbon-intensive production to maintain stability, particularly during periods of low renewable energy or load peaks. This emphasizes the disadvantages of isolated operation and reinforces the need for coordinated energy management by resulting in wasteful resource consumption, increased operating costs, and a larger reduction in the amount of renewable energy that is available.
Figure 10 displays the net energy profiles of the eight MGs (MG 1–8) before coordination. Notable oscillations and periods of both surplus (positive) and deficit (negative) energy are characteristics of these profiles. This variation highlights the inherent instability of solitary operation. The stabilizing effect of Case 2, however, is depicted in Figure 11, where energy sharing is made possible by the VPP’s centralized energy storage system (ESS). The aggregate net energy profile is successfully smoothed by the VPP’s optimization of the central ESS charging and discharging schedules. The centralized ESS serves as a buffer, absorbing collective surpluses and discharging during collective shortfalls, even as individual MG imbalances continue. As a result, net energy fluctuations are less severe, and system stability is improved. The lack of dynamic internal price setting, however, restricts the capacity to actively influence the behavior of individual MGs; in other words, the VPP largely responds to imbalances rather than actively averting them through price signals. When compared to the isolated baseline (Case 1), this case shows that centralized storage sharing alone can greatly increase resource utilization and dependability; however, more complex coordination methods may be able to achieve even greater improvements.

4.3. VPPs Performance

The VPP uses a decision-making process to formulate its decision-making process. Real-time external market prices and operational information from participating MGs, including load and generation characteristics, are incorporated into the state space. The VPP agent calculates an ideal internal price I n t p r i c e for the MGs based on this state. Within a restricted flexibility factor ε , this price is allowed to vary from the external price E x t p r i c e . This enables the model to dynamically balance demand response, grid conditions, and the economic goals of the VPP. The MG optimizes its operations for cost-effectiveness in response to internal pricing signals, which are generated from external market prices and internal load levels. In order to avoid costly grid electricity at times of high prices, it deliberately turns on its distributed generation (DG) or discharges its energy storage system. On the other hand, it takes advantage of temporal price arbitrage by charging the ESS during times of low prices.
This section compares the energy trading performance of many MGs in Case 3, weighing the financial advantages of direct wholesale market participation vs. an internal pricing scheme for VPPs over 24 h with real-time decision-making. The VPP’s dynamic pricing strategy, which tailors internal rates for each MG according to their unique operational features and current situations, is illustrated in Figure 12. The findings show that VPP aggregation benefits participating MGs financially. The findings highlight the ability of the VPP operator to establish advantageous internal trading circumstances for its aggregated MGs, which is a major advantage of the suggested hierarchical energy management system. The VPP lowers its net vulnerability to the unstable wholesale market by strategically planning the centralized energy storage system and using excess energy from one MG to fill in the gaps of another. The VPP operator can offer internal prices that are steadier and often lower than peak external prices due to this internal balancing act. The graph shows that the internal price stays below the external price during important buying times for MGs.
The cost-saving benefit for MG1 from the VPP’s internal pricing structure over direct market trading is shown in Figure 13a. Because of this, MG1’s internal energy purchase price is cheaper, as shown by the fact that its total cost through the VPP is about $40.67 as opposed to about $43.42 on the open market. As demonstrated by the 6.4% decrease in energy prices for MG1, the VPP serves as a strategic aggregator, protecting MGs from high market volatility and fostering a more favorable internal market, which leads to observable cost reductions. This demonstrates how well the collaborative model works to increase the distributed energy resources’ profitability. Figure 13b indicates the savings in costs MG2 achieves based on the internal pricing structure of the VPP. Trade is facilitated through VPP whereby the cost of energy procurement to MG2 is lowered to a definite economic advantage of $11.90 compared to the market rate of $13.07. This cost efficiency is attributed to the optimized internal pricing system that uses coordinated storage operation and MMG complementarity that is provided by the VPP. These results indicate that the proposed structure can be scaled to different MGs and consumption patterns.
As illustrated in Figure 14a, the internal pricing of the VPP helps to cushion the PV-deficient MG3 against the peak prices in the market and hence its costs of energy are minimized by approximately 2.4% in the time periods of peak demand. Under VPP membership, MG3, with high reliance on external electricity because of the absence of local PV production, reduces its energy costs by approximately 2.4%, dropping the cost per kilowatt-hour to a prohibitive market price of $157.80 to $154.02. The VPP can use its resources in a better manner to prevent the fluctuation of wholesale prices, and this is the critical role that it serves in assisting the MGs with limited resources by ensuring that the whole process is carried out carefully to supply the MGs with limited resources between 11:00 and 16:00, which is during the peak hours. The result confirms that the framework may ensure economic viability and operational stability for the most vulnerable participants in the energy ecosystem. As seen in Figure 14b, the VPP framework offers a notable 13.50% decrease in energy purchasing prices, demonstrating great economic efficiency for the participating MG4. Figure 15a illustrates the 2.25% cost savings achieved through VPP mediation, proving the model’s ability to generate value even for MG5 with lower marginal gains. The internal pricing method reduces energy costs by 7.90% for MG6, as shown in Figure 15b, demonstrating the VPP’s reliable performance under a range of operational conditions. Similarly, Figure 16 shows that energy costs are decreased by 11.6% for MG7 and 7.54% for MG8, respectively, showing the scalability and great economic advantage of the VPP aggregation approach for this configuration. The outcomes of energy trading prices through the wholesale market and VPP are shown in Table 3.
The VPP centralized energy storage system’s real-time, 24 h response to electricity market pricing is depicted in Figure 17. The graph clearly illustrates a conventional price-driven arbitrage strategy. A growing state of charge indicates that ESS charging mainly occurs during periods of low electricity costs, storing cheap energy. However, by discharging (as shown by a decreasing state of charge) when prices are high to supply the combined MGs, the ESS eliminates the need to purchase expensive power from the wholesale market. This planned cycle of charging and discharging lowers the net load on the main grid and gives the VPP significant financial rewards by utilizing pricing differences. The inverse relationship between the market price and the ESS energy level validates the effectiveness of the proposed DRL agent in managing the storage asset optimally to lower operational expenses.

4.4. Performance Comparison of Internal Pricing Models for VPPs

Three cutting-edge deep reinforcement learning algorithms: proximal policy optimization (PPO), advantage actor–critic (A2C), and soft actor-critic (SAC) were used in a comparative study to assess the effectiveness of the suggested methodology. A thorough performance evaluation utilizing mean absolute error (MAE) and mean absolute percentage error (MAPE) across eight MG systems served as the basis for choosing the best algorithm for internal price setting. They are used to quantify the deviation between the internally set prices and the corresponding external market prices. These metrics therefore serve as indicators of pricing stability and bounded alignment with market signals.
With the lowest MAE in six of the eight test cases and a 75% success rate, PPO was found to be the best algorithm by the evaluation, while SAC outperformed it in the other two systems. PPO performed consistently across a variety of MG configurations and showed very good accuracy in MG3 (MAE: 0.7856). The acquired MAPE values for the eight MGs are 8.46%, 13.3%, 8.28%, 11.8%, 9.22%, 13.4%, 11.2%, and 9.85%, which further validate this dominant performance. The overall framework’s outstanding accuracy is confirmed by these low error rates. PPO was the most dependable option for the hierarchical deep reinforcement learning framework because of its superior and consistent error measures as well as its built-in stability mechanisms, which are essential for managing market price volatility.

4.5. Techno-Economic Performance

The techno-economic analyses presented in this section are derived from the operational results of the DRL-based coordination framework. By analyzing component sizing, cost, and reliability, the techno-economic optimization approach determined the best system configuration for every location. Based on site-specific factors such as local demand, solar irradiation, and wind speed, four different configurations were chosen to minimize the levelized cost of electricity (LCOE) and net present cost (NPC) while guaranteeing a small capacity shortfall. The chosen architectures show how capital investment and operational performance are traded off in various geographic settings. The primary input parameters for modeling and optimizing the MMG system are listed in Table 4 [45]. Table 5 summarizes the economic analysis results with excess energy and renewable fraction of the selected sites.

4.5.1. PV, WT, and DG with ESS

The techno-economic analysis shows that both MG1 and MG2 achieve the same levelized cost of energy (LCOE) of about $0.144/kWh, indicating high efficiency and economic viability. This suggests similar operational effectiveness in providing reasonably priced power. But a significant distinction is the initial outlay of funds: With a capital expenditure (CAPEX) of about $15.8 million and a net present cost (NPC) of $25.9 million, MG1 is more costly than MG2, which has an NPC of $22 million and CAPEX of $13.4 million. This disparity suggests that MG1 is more extensive or has more advanced technology. Despite its greater initial capital cost, MG1’s competitive LCOE confirms its long-term economic sustainability. By finding the perfect balance between capital expenditure and operational effectiveness, both configurations’ exceptional performance demonstrates their potential to provide reliable, renewable energy solutions while preserving a solid value offer.

4.5.2. WT and DG with ESS

Techno-economic analysis indicates that a WT-DG-ESS is the optimal configuration for MG3. This hybrid system’s levelized cost of energy is $0.146/kWh, which is a little higher than prior setups and reflects the cost structure of mixing conventional and renewable energy sources. The system requires a large initial investment because it is larger in size and requires more technical expertise to balance wind power, storage, and diesel backup, as it is manifested by its larger net present cost of $44.8 million and capital expenditure (CAPEX) of $24.7 million. Although the capital cost of the arrangement is high, it provides a good alternative in areas with intermittent wind resources, which guarantee reliability in energy supply and minimize reliance on fuels in the long run. The economic feasibility of the system can be reflected in the fact that it utilizes renewable energy to create a steady power supply, which accounts for the increased initial costs in terms of the operating stability and long-term economic viability.

4.5.3. PV and WT with ESS

The levelized cost of energy of MG4, which is a PV-WT-ESS hybrid system, is about $0.164/kWh. This higher price in comparison with other schemes represents the gigantic capital density of incorporating and harmonizing multiple renewable sources of generation and a massive energy depository. The system is also very large, as the net cost today is approximately $46 million. Although this setup is more expensive at the start, it provides greater energy resilience and high levels of penetration of renewable energy, and it will be especially appropriate in regions where the solar and wind resources are complementary. The system’s architecture guarantees long-term operational sustainability and safeguards against changes in fuel prices by decreasing reliance on conventional fuel-based power generation, which helps to justify the initial capital investment.

4.5.4. PV and DG with ESS

The techno-economic performance of the PV-DG-ESS combinations spanning MG5 to MG8 demonstrates varying cost–efficiency trade-offs. Although MG5 has the largest capital expenditure, with a net present cost of about $65.5 million, it delivers a competitive levelized cost of energy of about $0.152/kWh. MG6 and MG7 demonstrate improved capital efficiency with LCOEs of $0.149/kWh and $0.150/kWh, respectively, and NPCs of $39.7 million and $36.5 million, demonstrating that comparable operating costs can be maintained with fewer initial inputs. MG8 balances initial investment and continuing operating expenses with an NPC of around $47.2 million and an LCOE of about $0.153/kWh. The variance in these outcomes emphasizes the adaptability of the PV-DG-ESS architecture, which can be adjusted to satisfy specific operational and financial objectives while consistently enhancing renewable integration and reducing reliance on conventional fuel sources. Figure 18 compares the economic performance of PV, WT, DG, and ESS technologies across all MGs.

4.6. Financial Performance

The financial study of the eight MG topologies (MG1–8) in Figure 18 demonstrates their high economic significance, with the majority of projects showing outstanding investment potential. With an IRR of 34–37% and return on investment (ROI) values of 30–33%, MG1, MG2, and MG3 demonstrate exceptional financial viability, which is bolstered by quick payback durations of 2.78–3.07 years. With constant payback times ranging from 3.08 to 3.19 years and IRRs of 33–34% and ROIs of 29–30%, MG5 through MG8 likewise offer very favorable and reliable returns. With a payback period of 3.62 years and a relatively low IRR of 28% and ROI of 23%, MG4’s performance is nonetheless financially feasible and represents the trade-offs of a larger-scale, high-renewable penetration system. The findings taken together highlight the strong financial returns and strong investment security provided by the suggested MG designs, most of which guarantee capital recovery in a short amount of time. The results can be found in Table 5.

4.7. Environmental Performance

The environmental analyses in this section are based on the operational outcomes of the DRL-based coordination framework. Carbon emissions are evaluated for both the conventional grid-supplied system and each MG configuration operating under its original, non-coordinated dispatch strategy. These reference scenarios serve as the baseline for emission assessment. Using identical emission factors, total emissions are then calculated for both the baseline and the optimized coordinated scenarios to ensure consistency in comparison. The resulting emission reductions reflect the environmental gains achieved through coordinated operation, greater utilization of renewable energy, and reduced reliance on grid electricity and diesel-based generation. Comparing the proposed MG configurations to traditional grid systems, the emission reduction (ER) analysis shows a revolutionary environmental performance. With ER values ranging from 91.47% (MG3) to a full 100% (MG4), these systems show that the local energy supply is almost entirely decarbonized. All other combinations (MG1, MG2, MG5, MG6, MG7, and MG8) continuously show ER percentages above 94%, demonstrating the optimized hybrid designs’ strong capacity to reduce emissions. This demonstration emphasizes how vital advanced MGs are to deep decarbonization ambitions as it sharply contrasts with present grid practice, which relies heavily on carbon-intensive generation. The findings illustrate the gains of MG systems as a basic strategy to provide sustainable, low-carbon energy infrastructure. The percentage % emission reductions at each site are shown in Figure 19.

5. Conclusions

The growing standardization of distributed energy sources has also posed opportunities and challenges to the contemporary power systems, which have required the adoption of advanced management techniques that enable full exploitation of MG cooperation. This study has addressed this pressing need by developing and approving a comprehensive low-carbon framework for VPPs that integrate multiple MGs. According to the study’s findings, MGs cooperating may significantly lower carbon emissions, protect the environment, and share resources while also boosting the economy. One significant advancement in VPP operating methodology is the multi-stage collaborative energy management approach that has been proposed. This method efficiently tackles the intricate uncertainties present in load demands and renewable energy generation while optimizing system self-containment by integrating an advanced internal price setting mechanism that was acquired by deep reinforcement learning (DRL). The load and generation patterns dictate the pricing structure.
The fully integrated VPP operation (Case 3) is preferable compared to storage sharing models (Case 2) and isolated MG scenarios (Case 1), according to a comparative analysis employing rigorous simulation based on real-world data. The results show that the three-stage DRL architecture successfully reconciles the competing interests of the VPP operator and participating MGs, resulting in significant improvements in operational flexibility and financial efficiency. Specifically, dynamic internal pricing and the coordinated energy storage system (ESS) dispatch mechanism reduce energy costs by 2.25% to 13.5% in different MG configurations while maintaining system stability. The model-free DRL methodology has demonstrated remarkable resilience to changing resource availability and market conditions in terms of autonomously learning effective scheduling strategies without the requirement for explicit system models. This feature makes the framework a scalable solution because grids that are mostly powered by renewable energy will always have some volatility and change.
This paper offers a strong theoretical basis and useful method for energy management in MMG systems, in addition to its direct technical contributions. With emission reduction rates ranging from 91.47% to 100%, the proposed approach has been successfully applied to eight different MG configurations, demonstrating its ability to hasten the shift to more sustainable and efficient energy communities. The financial sustainability of the setups, as evidenced by internal rates of return ranging from 28% to 37%, further supports the business case for employing such integrated methodologies. In conclusion, this study makes a substantial contribution to the field of distributed energy resource management by demonstrating how ingenious coordinating techniques can uncover the latent value in MMG systems. The proposed approach offers a scalable, privacy-preserving, and economically viable path to achieving global carbon neutrality goals while enhancing grid resilience and reliability. Future studies will focus on extending this approach to multi-energy systems and examining its use in other regulatory and market settings.

Author Contributions

Conceptualization, M.W.K., S.A.A.K., M.K.R. and F.A.I.; methodology, M.W.K. and S.A.A.K.; software, M.W.K.; validation, M.W.K., S.A.A.K., M.A., M.K.R., F.A.I. and M.K.M.; formal analysis, M.A., M.K.R., F.A.I. and M.K.M.; investigation, M.W.K. and S.A.A.K.; resources, S.A.A.K.; data curation, M.W.K. and M.A.; writing—original draft preparation, M.W.K. and S.A.A.K.; writing—review and editing, S.A.A.K., M.A., M.K.R., F.A.I. and M.K.M.; visualization, M.W.K.; supervision, S.A.A.K.; project administration, S.A.A.K., M.A., M.K.R., F.A.I. and M.K.M.; funding acquisition, S.A.A.K. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the HEC-Pakistan NRPU Project, grant number 15722.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

MGMicrogrid
MMGMulti-microgrid
VPPVirtual power plant
DERDistributed energy resource
DGDiesel generator
EMSEnergy management system
PVPhotovoltaic
WTWind turbine
ESSEnergy storage system
RESRenewable energy source
DRLDeep reinforcement learning
PPOProximal policy optimization
SOCState of charge
LCOELevelized cost of energy
NPCNet present cost
CAPEX     Initial investment cost
IRRInternal rate of return
ROIReturn on investment
TBPPayback period
EREmission reduction
EPEconomic parameter
FPFinancial parameter
EAEnergy analysis
EnvAEnvironmental analysis
EEExcess energy
RFRenewable fraction
TLTransferable load
P p v r PV system’s rated power
P w t r WT’s rated power
P d g . o u t DG’s output power
P i B Power of battery
η i B Efficiency of the battery
C a n n Aggregated annual cost
B a n n Annual net cash flow
I n t p r i c e Internal electricity price
E x t p r i c e External electricity price
ξ Pricing preferential factor
i r Interest rate
λ D G Fuel cost per unit of energy produced
δ c ESS charging δ d ESS discharging
δ u p Rise of DG power generation
δ d o w n Fall of DG power generation
P b u y Purchased power from/to the VPP
P s e l l Sold power from/to the VPP
I n t C _ P b u y Internal cost of buying power
I n t C _ P s e l l Internal cost of selling power
E x t C _ P b u y External cost of buying power
E x t C _ P s e l l External cost of selling power
R u p Ramp up limit of DG
R d o w n Ramp down limit of DG
δ + Increasing of TL
δ Decreasing of TL

References

  1. Khalil, M.W.; Altamimi, A.; Kazmi, S.A.A.; Khan, Z.A.; Shin, D.R. Integration of Distributed Generations in Smart Distribution Networks Using Multi-Criteria Based Sustainable Planning Approach. Sustainability 2022, 15, 384. [Google Scholar] [CrossRef]
  2. Wu, J.; Li, Q.; Bie, Y.; Zhou, W. Location-routing optimization problem for electric vehicle charging stations in an uncertain transportation network: An adaptive co-evolutionary clustering algorithm. Energy 2024, 304, 132142. [Google Scholar] [CrossRef]
  3. Singh, A.R.; Raju, D.K.; Raghav, L.P.; Kumar, R.S. State-of-the-art review on energy management and control of networked microgrids. Sustain. Energy Technol. Assess. 2023, 57, 103248. [Google Scholar] [CrossRef]
  4. Aguilar, D.; Quinones, J.J.; Pineda, L.R.; Ostanek, J.; Castillo, L. Optimal scheduling of renewable energy microgrids: A robust multi-objective approach with machine learning-based probabilistic forecasting. Appl. Energy 2024, 369, 123548. [Google Scholar] [CrossRef]
  5. Huy, T.H.B.; Le, T.D.; Van Phu, P.; Park, S.; Kim, D. Real-time power scheduling for an isolated microgrid with renewable energy and energy storage system via a supervised-learning-based strategy. J. Energy Storage 2024, 88, 111506. [Google Scholar] [CrossRef]
  6. Guo, C.; Wang, X.; Zheng, Y.; Zhang, F. Optimal energy management of multi-microgrids connected to distribution system based on deep reinforcement learning. Int. J. Electr. Power Energy Syst. 2021, 131, 107048. [Google Scholar] [CrossRef]
  7. Wu, N.; Xu, J.; Linghu, J.; Huang, J. Real-time optimal control and dispatching strategy of multi-microgrid energy based on storage collaborative. Int. J. Electr. Power Energy Syst. 2024, 160, 110063. [Google Scholar] [CrossRef]
  8. Goia, B.; Cioara, T.; Anghel, I. Virtual power plant optimization in smart grids: A narrative review. Future Internet 2022, 14, 128. [Google Scholar] [CrossRef]
  9. Kim, Y.M.; Jung, D.; Chang, Y.; Choi, D.H. Intelligent micro energy grid in 5G era: Platforms, business cases, testbeds, and next generation applications. Electronics 2019, 8, 468. [Google Scholar] [CrossRef]
  10. Mbungu, N.T.; Naidoo, R.M.; Bansal, R.C.; Vahidinasab, V. Overview of the optimal smart energy coordination for microgrid applications. IEEE Access 2019, 7, 163063–163084. [Google Scholar] [CrossRef]
  11. Villalón, A.; Rivera, M.; Salgueiro, Y.; Muñoz, J.; Dragičević, T.; Blaabjerg, F. Predictive control for microgrid applications: A review study. Energies 2020, 13, 2454. [Google Scholar] [CrossRef]
  12. Raju, L.; Morais, A.A. Multi-agent systems based advanced energy management of smart micro-grid. In Multi Agent Systems-Strategies and Applications; IntechOpen: London, UK, 2020. [Google Scholar]
  13. Khan, M.R.; Haider, Z.M.; Malik, F.H.; Almasoudi, F.M.; Alatawi, K.S.S.; Bhutta, M.S. A comprehensive review of microgrid energy management strategies considering electric vehicles, energy storage systems, and AI techniques. Processes 2024, 12, 270. [Google Scholar] [CrossRef]
  14. Zahraoui, Y.; Korotko, T.; Rosin, A.; Agabus, H. Market mechanisms and trading in microgrid local electricity markets: A comprehensive review. Energies 2023, 16, 2145. [Google Scholar] [CrossRef]
  15. Schwidtal, J.M.; Piccini, P.; Troncia, M.; Chitchyan, R.; Montakhabi, M.; Francis, C.; Gorbatcheva, A.; Capper, T.; Mustafa, M.A.; Andoni, M.; et al. Emerging business models in local energy markets: A systematic review of peer-to-peer, community self-consumption, and transactive energy models. Renew. Sustain. Energy Rev. 2023, 179, 113273. [Google Scholar] [CrossRef]
  16. Khan, H.; Masood, T. Impact of blockchain technology on smart grids. Energies 2022, 15, 7189. [Google Scholar] [CrossRef]
  17. Wu, Y.; Wu, Y.; Cimen, H.; Vasquez, J.C.; Guerrero, J.M. Towards collective energy Community: Potential roles of microgrid and blockchain to go beyond P2P energy trading. Appl. Energy 2022, 314, 119003. [Google Scholar] [CrossRef]
  18. Kochupurackal, A.; Pancholi, K.P.; Islam, S.N.; Anwar, A.; Oo, A.M.T. Rolling horizon optimisation based peer-to-peer energy trading under real-time variations in demand and generation. Energy Syst. 2023, 14, 541–565. [Google Scholar] [CrossRef]
  19. Amanbek, Y.; Kalakova, A.; Zhakiyeva, S.; Kayisli, K.; Zhakiyev, N.; Friedrich, D. Distribution locational marginal price based transactive energy management in distribution systems with smart prosumers—A multi-agent approach. Energies 2022, 15, 2404. [Google Scholar] [CrossRef]
  20. Reihani, E.; Siano, P.; Genova, M. A new method for peer-to-peer energy exchange in distribution grids. Energies 2020, 13, 799. [Google Scholar] [CrossRef]
  21. Touhs, H.; Temouden, A.; Khallaayoun, A.; Chraibi, M.; El Hafdaoui, H. A scheduling algorithm for appliance energy consumption optimization in a dynamic pricing environment. World Electr. Veh. J. 2023, 15, 1. [Google Scholar] [CrossRef]
  22. Mehta, Y.; Xu, R.; Lim, B.; Wu, J.; Gao, J. A review for green energy machine learning and AI services. Energies 2023, 16, 5718. [Google Scholar] [CrossRef]
  23. Chen, Y.; Fu, G.; Liu, X. Air-conditioning load forecasting for prosumer based on meta ensemble learning. IEEE Access 2020, 8, 123673–123682. [Google Scholar] [CrossRef]
  24. Miraftabzadeh, S.M.; Colombo, C.G.; Longo, M.; Foiadelli, F. K-means and alternative clustering methods in modern power systems. IEEE Access 2023, 11, 119596–119633. [Google Scholar] [CrossRef]
  25. Srinivasan, S.; Kumarasamy, S.; Andreadakis, Z.E.; Lind, P.G. Artificial intelligence and mathematical models of power grids driven by renewable energy sources: A survey. Energies 2023, 16, 5383. [Google Scholar] [CrossRef]
  26. Jurdak, R.; Dorri, A.; Vilathgamuwa, M. A trusted and privacy-preserving internet of mobile energy. IEEE Commun. Mag. 2021, 59, 89–95. [Google Scholar] [CrossRef]
  27. Mejia, C.; Kajikawa, Y. Emerging topics in energy storage based on a large-scale analysis of academic articles and patents. Appl. Energy 2020, 263, 114625. [Google Scholar] [CrossRef]
  28. Chang, W.; Yang, Q. Low carbon oriented collaborative energy management framework for multi-microgrid aggregated virtual power plant considering electricity trading. Appl. Energy 2023, 351, 121906. [Google Scholar] [CrossRef]
  29. Rosato, A.; Panella, M.; Andreotti, A.; Mohammed, O.A.; Araneo, R. Two-stage dynamic management in energy communities using a decision system based on elastic net regularization. Appl. Energy 2021, 291, 116852. [Google Scholar] [CrossRef]
  30. Bolzoni, A.; Parisio, A.; Todd, R.; Forsyth, A.J. Optimal virtual power plant management for multiple grid support services. IEEE Trans. Energy Convers. 2020, 36, 1479–1490. [Google Scholar] [CrossRef]
  31. Gong, X.; Castillo-Guerra, E.; Cardenas-Barrera, J.L.; Cao, B.; Saleh, S.A.; Chang, L. Robust hierarchical control mechanism for aggregated thermostatically controlled loads. IEEE Trans. Smart Grid 2020, 12, 453–467. [Google Scholar] [CrossRef]
  32. Chen, J.; Zhao, Y.; Wang, M.; Yang, K.; Ge, Y.; Wang, K.; Lin, H.; Pan, P.; Hu, H.; He, Z.; et al. Multi-timescale reward-based DRL energy management for regenerative braking energy storage system. IEEE Trans. Transp. Electrif. 2025, 11, 7488–7500. [Google Scholar] [CrossRef]
  33. Zhao, C.; Li, Y.; Zhang, Q.; Ren, L. Low Carbon Economic Energy Management Method in a Microgrid Based on Enhanced D3QN Algorithm with Mixed Penalty Function. IEEE Trans. Sustain. Energy 2025, 16, 1686–1696. [Google Scholar] [CrossRef]
  34. Wang, C.; Cheng, B.; He, X.; Xi, L.; Yang, N.; Zhao, Z.; Lai, C.S.; Lai, L.L. Integrated Underfrequency Load Shedding Strategy for Islanded Microgrids Integrating Multiclass Load-Related Factors. IEEE Trans. Smart Grid 2025, 16, 4305–4323. [Google Scholar] [CrossRef]
  35. Wang, C.; Liu, Y.; Zhang, Y.; Xi, L.; Yang, N.; Zhao, Z.; Lai, C.S.; Lai, L.L. Strategy for optimizing the bidirectional time-of-use electricity price in multi-microgrids coupled with multilevel games. Energy 2025, 323, 135731. [Google Scholar] [CrossRef]
  36. Li, Y.; Chang, W.; Yang, Q. Deep reinforcement learning based hierarchical energy management for virtual power plant with aggregated multiple heterogeneous microgrids. Appl. Energy 2025, 382, 125333. [Google Scholar] [CrossRef]
  37. Wang, Z.; Chen, B.; Wang, J.; kim, J. Decentralized energy management system for networked microgrids in grid-connected and islanded modes. IEEE Trans. Smart Grid 2015, 7, 1097–1105. [Google Scholar] [CrossRef]
  38. Tushar, W.; Saha, T.K.; Yuen, C.; Smith, D.; Poor, H.V. Peer-to-peer trading in electricity networks: An overview. IEEE Trans. Smart Grid 2020, 11, 3185–3200. [Google Scholar] [CrossRef]
  39. Chen, X.; Dong, W.; Yang, L.; Yang, Q. Scenario-based robust capacity planning of regional integrated energy systems considering carbon emissions. Renew. Energy 2023, 207, 359–375. [Google Scholar] [CrossRef]
  40. Zhang, L.; Shi, R.; Ma, X.; Jia, L.; Lee, K.Y. Highway self-contained multi-microgrid energy management strategy based on universal gravitation. Energy 2025, 327, 136430. [Google Scholar] [CrossRef]
  41. Kharrich, M.; Kamel, S.; Abdel-Akher, M.; Eid, A.; Zawbaa, H.M.; Kim, J. Optimization based on movable damped wave algorithm for design of photovoltaic/wind/diesel/biomass/battery hybrid energy systems. Energy Rep. 2022, 8, 11478–11491. [Google Scholar] [CrossRef]
  42. Merabet, A.; Al-Durra, A.; El-Fouly, T.; El-Saadany, E.F. Optimization and energy management for cluster of interconnected microgrids with intermittent non-polluting and diesel generators in off-grid communities. Electr. Power Syst. Res. 2025, 241, 111319. [Google Scholar] [CrossRef]
  43. Shi, R.; Gao, Y.; Ning, J.; Tang, K.; Jia, L. Research on highway self-consistent energy system planning with uncertain wind and photovoltaic power output. Sustainability 2023, 15, 3166. [Google Scholar] [CrossRef]
  44. Ge, X.; Shi, L.; Fu, Y.; Muyeen, S.; Zhang, Z.; He, H. Data-driven spatial-temporal prediction of electric vehicle load profile considering charging behavior. Electr. Power Syst. Res. 2020, 187, 106469. [Google Scholar] [CrossRef]
  45. Qureshi, Z.-A.; Kazmi, S.A.A.; Mushtaq, S.; Anwar, M. An integrated assessment framework of renewable based microgrid deployment for remote isolated area electrification across different climatic zones and future grid extensions. Sustain. Cities Soc. 2024, 101, 105069. [Google Scholar] [CrossRef]
  46. Peng, C.Y.; Kuo, C.C.; Tsai, C.T. Optimal configuration with capacity analysis of PV-PLUS-BESS for behind-the-meter application. Appl. Sci. 2021, 11, 7851. [Google Scholar] [CrossRef]
  47. NASA POWER. Prediction of Worldwide Energy Resources (POWER). 2022. Available online: https://power.larc.nasa.gov/ (accessed on 15 August 2025).
  48. National Transmission & Despatch Company Limited. Power System Statistics; National Transmission & Despatch Company Limited: Lahore, Pakistan, 2024. [Google Scholar]
  49. PJM ISO. Wholesale Electricity Prices. 2024. Available online: https://www.eia.gov/electricity/?rto=pjm (accessed on 5 August 2025).
Figure 1. VPP architecture model.
Figure 1. VPP architecture model.
Sustainability 18 00275 g001
Figure 2. PPO architecture.
Figure 2. PPO architecture.
Sustainability 18 00275 g002
Figure 3. MG agent training.
Figure 3. MG agent training.
Sustainability 18 00275 g003
Figure 4. VPP agent training.
Figure 4. VPP agent training.
Sustainability 18 00275 g004
Figure 5. MG agent training by VPP.
Figure 5. MG agent training by VPP.
Sustainability 18 00275 g005
Figure 6. Project map of selected areas in Pakistan for MGs 1–8.
Figure 6. Project map of selected areas in Pakistan for MGs 1–8.
Sustainability 18 00275 g006
Figure 7. Average temperature, radiation, and wind speed of selected regions.
Figure 7. Average temperature, radiation, and wind speed of selected regions.
Sustainability 18 00275 g007
Figure 8. Adapted load profiles: (a) Residential Load; (b) Non-Residential Load.
Figure 8. Adapted load profiles: (a) Residential Load; (b) Non-Residential Load.
Sustainability 18 00275 g008
Figure 9. Energy contribution of proposed MG configurations: (a) PV-WT-DG-ESS; (b) WT-DG-ESS; (c) PV-WT-ESS; (d) PV-DG-ESS.
Figure 9. Energy contribution of proposed MG configurations: (a) PV-WT-DG-ESS; (b) WT-DG-ESS; (c) PV-WT-ESS; (d) PV-DG-ESS.
Sustainability 18 00275 g009
Figure 10. Uncoordinated net energy exchange of each microgrid prior to VPP intervention.
Figure 10. Uncoordinated net energy exchange of each microgrid prior to VPP intervention.
Sustainability 18 00275 g010
Figure 11. Net energy for each microgrid with enabled centralized storage sharing.
Figure 11. Net energy for each microgrid with enabled centralized storage sharing.
Sustainability 18 00275 g011
Figure 12. Real-time electricity internal price setting for 24-h horizon.
Figure 12. Real-time electricity internal price setting for 24-h horizon.
Sustainability 18 00275 g012
Figure 13. Energy price trends over a 24-h horizon: (a) MG1; (b) MG2.
Figure 13. Energy price trends over a 24-h horizon: (a) MG1; (b) MG2.
Sustainability 18 00275 g013
Figure 14. Energy price trends over a 24-h horizon: (a) MG3; (b) MG4.
Figure 14. Energy price trends over a 24-h horizon: (a) MG3; (b) MG4.
Sustainability 18 00275 g014
Figure 15. Energy price trends over a 24-h horizon: (a) MG5 and (b) MG6.
Figure 15. Energy price trends over a 24-h horizon: (a) MG5 and (b) MG6.
Sustainability 18 00275 g015
Figure 16. Energy price trends over a 24-h horizon: (a) MG7; (b) MG8.
Figure 16. Energy price trends over a 24-h horizon: (a) MG7; (b) MG8.
Sustainability 18 00275 g016
Figure 17. Scheduling of ESS by VPP across 24-h horizon.
Figure 17. Scheduling of ESS by VPP across 24-h horizon.
Sustainability 18 00275 g017
Figure 18. Radar diagram for techno-economic and financial parameters.
Figure 18. Radar diagram for techno-economic and financial parameters.
Sustainability 18 00275 g018
Figure 19. Emission reduction for proposed MGs.
Figure 19. Emission reduction for proposed MGs.
Sustainability 18 00275 g019
Table 1. Configurations of MG framework.
Table 1. Configurations of MG framework.
DevicePVWTDGESS
MG1
MG2
MG3×
MG4×
MG5×
MG6×
MG7×
MG8×
Table 2. Simulation parameters.
Table 2. Simulation parameters.
ParameterValue
Charging efficiency0.8
discharging efficiency0.8
Internal price bound0.2
Preferential coefficient0.9
Initial training episodes200,000
Alternate training episodes20,000
Learning rate0.002
Gamma0.99
Batch size128
Clip range0.2
Discount factor0.99
Table 3. Cost comparison performance.
Table 3. Cost comparison performance.
Sr. No.Energy Trading Price ($) VPP Profit (%)
VPPWholesale Market
MG140.6743.426.40
MG211.9013.078.95
MG3154.02157.802.40
MG4158.74183.5213.5
MG5279.38285.822.25
MG663.3468.787.90
MG7164.93186.5611.60
MG8120.02129.807.54
Table 4. Data of MG simulation parameters.
Table 4. Data of MG simulation parameters.
Systems ParameterValueSystems ParameterValue
PV panel Wind turbine
Rated capacity (W)370Rated capacity (kW)800
Efficiency (%)19.1Capital cost ($/kW)1500
Capital cost ($/kW)700Replacement cost ($/kW)1250
Replacement cost ($/kW)500O&M cost ($/year)10
O&M cost ($/year)9Lifetime (year)20
Lifetime (year)25Energy storage system
Diesel generator Nominal capacity (kWh)1.02
Fuel price ($/L)0.96Capital cost ($/kW)200
Capital cost ($/kW)500Replacement cost ($/kW)200
Replacement cost ($/kW)450O&M cost ($/kW)10
O&M cost ($/kW)0.03Lifetime (year)5
Lifetime (hour)15,000Maximum state of charge (%)100
DG generation price ($/kWh)0.75Minimum state of charge (%)20
Table 5. Techno-economic assessment of sustainable solutions.
Table 5. Techno-economic assessment of sustainable solutions.
Sr. No.EP FP EA EnvA
LCOENPCCAPEXIRRROITBPEERFER
($)($M)($M)(%)(%)(Year)(%)(%)(%)
MG10.14425.915.834303.0734.395.295.17
MG20.1442213.435313.0135.494.894.79
MG30.14644.824.737332.7846.891.391.47
MG40.1644630.328233.6260.1100100
MG50.15265.538.933293.1636.295.495.40
MG60.14939.723.234303.0732.695.295.18
MG70.1536.521.334303.0832.595.295.17
MG80.15347.22732293.1936.295.295.20
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Khalil, M.W.; Kazmi, S.A.A.; Anwar, M.; Rathi, M.K.; Ibupoto, F.A.; Maheshwari, M.K. A Collaborative Energy Management and Price Prediction Framework for Multi-Microgrid Aggregated Virtual Power Plants. Sustainability 2026, 18, 275. https://doi.org/10.3390/su18010275

AMA Style

Khalil MW, Kazmi SAA, Anwar M, Rathi MK, Ibupoto FA, Maheshwari MK. A Collaborative Energy Management and Price Prediction Framework for Multi-Microgrid Aggregated Virtual Power Plants. Sustainability. 2026; 18(1):275. https://doi.org/10.3390/su18010275

Chicago/Turabian Style

Khalil, Muhammad Waqas, Syed Ali Abbas Kazmi, Mustafa Anwar, Mahesh Kumar Rathi, Fahim Ahmed Ibupoto, and Mukesh Kumar Maheshwari. 2026. "A Collaborative Energy Management and Price Prediction Framework for Multi-Microgrid Aggregated Virtual Power Plants" Sustainability 18, no. 1: 275. https://doi.org/10.3390/su18010275

APA Style

Khalil, M. W., Kazmi, S. A. A., Anwar, M., Rathi, M. K., Ibupoto, F. A., & Maheshwari, M. K. (2026). A Collaborative Energy Management and Price Prediction Framework for Multi-Microgrid Aggregated Virtual Power Plants. Sustainability, 18(1), 275. https://doi.org/10.3390/su18010275

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop