Next Article in Journal
Real-World Combustion Emissions, Engine Size and Vehicle Mass Versus Euro-Class Access Criteria in Low-Emission Zones
Previous Article in Journal
LCC-S vs. LCC-LCC: Efficient Wireless Charging for Underwater Drones Under Seawater Conditions
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Cooperative Game-Based Low-Carbon Optimal Operation Strategy for Multi-Microgrids Based on Multi-Agent Deep Reinforcement Learning

1
State Grid Shanghai Municipal Electric Power Company, Shanghai 200122, China
2
School of Mechanical Engineering, University of Shanghai for Science and Technology, Shanghai 200093, China
*
Author to whom correspondence should be addressed.
Energies 2026, 19(15), 3683; https://doi.org/10.3390/en19153683
Submission received: 1 June 2026 / Revised: 30 July 2026 / Accepted: 4 August 2026 / Published: 5 August 2026
(This article belongs to the Section B: Energy and Environment)

Abstract

Distributed integrated energy microgrids support the low-carbon transition of regional energy systems. However, multi-agent trading among microgrids still faces insufficient cross-market coordination, weak low-carbon incentives, and difficulties in fair benefit allocation. To address these issues, this paper proposes a cooperative-game-based low-carbon optimal operation strategy for multi-microgrids using multi-agent deep reinforcement learning. First, an energy-carbon-green certificate peer-to-peer coordinated trading mechanism and a green-carbon offsetting-based dual-incentive model are developed to link energy exchange, carbon quota adjustment, and green certificate circulation. Second, a Nash bargaining-based cooperative game model is formulated for multi-commodity P2P trading to maximize coalition benefits and ensure a fair allocation of surplus. Finally, the cooperative game is transformed into a Markov decision process, and a centralized training and decentralized execution framework with homogeneous agents is constructed based on the multi-agent soft actor-critic algorithm. Case studies using data from the Yangtze River Delta region of China show that the proposed method achieves a 1.25% optimality gap compared with the centralized MILP benchmark and reduces the coalition operating cost by 8.19% relative to independent operation. The carbon trading costs of the three microgrids are reduced by 37.27%, 40.13%, and 33.82%, respectively, verifying the economic applicability of the proposed method.

1. Introduction

Driven by national carbon peaking and carbon neutrality targets, energy systems are shifting toward clean, low-carbon, secure, and efficient operation. Microgrids serve as key carriers for integrating distributed renewable energy and coordinating source, grid, load, and storage resources. They therefore form an important part of new power systems. Conventional microgrids usually have simple structures and limited energy supply forms. Thus, they cannot effectively cope with the high randomness of wind and photovoltaic power generation or the variable load demand. To address these issues, distributed integrated energy microgrids (DIEMs) integrate multiple energy conversion technologies and carbon-resource response capabilities. DIEMs support the coordinated coupling, flexible conversion, and cascade utilization of different energy forms. Therefore, DIEMs provide a key technical approach for improving the flexibility, economic performance, and low-carbon performance of energy systems [1].
The operation mode of microgrids shifts from independent operation toward multi-agent interconnection and mutual support. By exploiting the complementary potential of multi-microgrids through cooperation and peer-to-peer interactions, this operation mode provides an important approach to efficient regional energy allocation and low-carbon operation [2,3]. In this context, multi-microgrids engage in direct energy exchange through peer-to-peer (P2P) trading. Trading quantities and prices are negotiated independently based on the supply–demand surplus or deficit for each microgrid. In this way, surplus renewable energy is locally consumed, and complementary load support among microgrids is achieved [4,5,6]. However, most existing studies focus primarily on energy interactions. The integrated design of green certificate consumption obligations and surplus–deficit adjustment of carbon emission quotas remains insufficient. This limitation limits the potential for low-carbon, cost-effective operation of multi-microgrid systems [7,8]. Meanwhile, P2P trading among multi-microgrids features differentiated interests, autonomous decision-making, and interactive game relationships. Each participant pursues individual benefit maximization, while the overall optimum is achieved through coalition cooperation. Therefore, cooperative game theory serves as a core theoretical tool for describing multi-agent collaborative decision-making, maximizing coalition benefits, and achieving fair benefit allocation [9,10]. In Ref. [11], cooperative game theory is combined with risk-averse stochastic programming to improve the collaborative benefits of multi-microgrids under carbon-reduction constraints. In Refs. [12,13,14,15], multi-microgrid trading optimization models are constructed based on Nash bargaining theory to achieve the fair allocation of multi-agent cooperative surplus. However, existing studies still have three limitations. First, most P2P trading studies focus primarily on electricity or multi-energy exchange, while the coordinated interactions among energy trading, carbon quota adjustment, and green certificate circulation have not been sufficiently modeled. Second, although Nash bargaining has been used for cooperative surplus allocation, its integration with energy-carbon-green certificate coordinated trading remains limited. Third, existing DRL-based methods primarily focus on energy scheduling performance, while their integration with cooperative-game payoff allocation and multi-commodity P2P settlement remains insufficient.
In terms of model solving, traditional methods mainly include optimization methods based on Karush–Kuhn–Tucker (KKT) conditions [16] and heuristic algorithms [17]. Although these methods are theoretically sound, they have several limitations in cooperative game problems with nonconvexity, nonlinearity, and strong uncertainty. These limitations include complex modeling, low solution efficiency, and insufficient robustness. These limitations make them difficult to adapt to dynamic multi-agent interactions and real-time decision-making requirements. Multi-agent deep reinforcement learning provides an efficient solution for complex cooperative game problems because it does not require an accurate model. It also adapts to environmental dynamics and supports high-dimensional continuous decision-making [18,19]. In Refs. [20,21], improved deep deterministic policy gradient (DDPG) algorithms show superior performance in power balance and energy scheduling. In Ref. [22], a multi-agent deep reinforcement learning algorithm is adopted to improve the energy utilization efficiency of multi-energy microgrids. A novel model-free, data-driven algorithm based on centralized training and decentralized execution is also developed. However, Ref. [22] mainly focuses on MASAC-based energy management and does not couple the learning framework with Nash bargaining-based surplus allocation or carbon-green certificate P2P settlement. Unlike Ref. [23], which emphasizes green-carbon offsetting in a Stackelberg game, this study integrates green-carbon offsetting into a cooperative multi-microgrid P2P trading framework and solves it using CTDE-MASAC. Notably, multi-microgrid cooperative games often exhibit nonconvexity and non-stationarity. Recent studies have applied MASAC and its improved variants to multi-microgrid scheduling and control problems, showing their applicability to multi-agent coordination under renewable energy uncertainty and dynamic operating conditions [24,25]. Compared with deterministic policy methods such as MADDPG, MASAC introduces a maximum-entropy objective and stochastic policy exploration, making it suitable for continuous action scheduling problems with multiple coupled agents [25]. Compared with on-policy methods such as MAPPO, its off-policy learning mechanism can reuse historical experiences and improve sample efficiency. Therefore, MASAC is adopted in this study as the solution framework, and its convergence and robustness are further evaluated through comparative experiments in Section 5.2. To further clarify the research gaps, Table 1 compares representative studies from five aspects: market type and trading mechanism, game model, solution method, uncertainty handling, and carbon/GCT mechanism. The comparison shows that existing studies typically address P2P energy trading, cooperative benefit allocation, and DRL-based scheduling separately, while the integration of these approaches with energy-carbon-green certificate coordination remains insufficient.
To address these gaps, this paper proposes a cooperative-game-based low-carbon optimal operation strategy for multi-microgrids that considers coordinated trading of energy-carbon-green certificates. The main contributions are summarized as follows:
(1) An energy-carbon-green certificate P2P coordinated trading mechanism is developed for multi-microgrids. By coupling energy exchange, carbon quota adjustment, and green certificate circulation, the proposed mechanism establishes a value transmission pathway among the energy, CET, and GCT markets. Furthermore, a dual-incentive pricing model based on green-carbon offsetting is introduced to convert surplus green certificates into carbon-offset value, thereby strengthening the incentives for renewable energy consumption and carbon emission reduction.
(2) A Nash bargaining-based cooperative game model is formulated for multi-microgrid P2P trading. The model considers multi-commodity transactions involving electricity, heat, hydrogen, carbon quotas, and green certificates. It maximizes coalition benefits while allocating cooperative surplus fairly among DIEMs, thereby improving both system-level efficiency and individual rationality.
(3) A CTDE-MASAC solution framework with homogeneous agents is designed for the proposed cooperative game. The optimization problem is formulated as a Markov decision process, and decentralized actors with a centralized critic are used to enable adaptive decision-making in continuous action spaces. This framework improves the applicability of multi-agent deep reinforcement learning to the coordinated, low-carbon operation of multiple microgrids.
This paper is organized as follows: Section 2 develops the energy-carbon-green certificate P2P coordinated trading mechanism and the green-carbon offsetting-based dual-incentive model. Section 3 formulates the cooperative-game-based optimal operation model and the Nash bargaining allocation mechanism. Section 4 presents the CTDE-MASAC solution framework. Section 5 reports the case studies and comparative results. Section 6 concludes the paper and discusses future research directions.

2. Bidirectional Interaction Mechanism for Multi-Microgrids Considering P2P Energy-Carbon-Green Certificate Coordinated Market Trading

2.1. Architectural Design of the Distributed Integrated Energy Microgrid

The overall architecture of the distributed integrated energy microgrid (DIEM) is shown in Figure 1. The system comprises an energy supply side, an energy conversion and coupling side, an energy demand side, and a green certificate-carbon market trading side. Compared with conventional microgrids, DIEM offers three key technical advantages: (1) enhanced on-site renewable energy consumption through electricity-heat-hydrogen multi-energy coupling; (2) a closed-loop carbon pathway enabled by integrated carbon capture, utilization, and storage (CCUS) and power-to-gas (P2G) technologies; and (3) coordinated participation in the electricity market, carbon trading market, and green certificate trading market, achieving joint optimization of economic performance and low-carbon operation.
Specifically, the DIEM architecture is described as follows. (1) Energy supply side. Wind power, photovoltaics (PV), the main grid, and natural gas serve as the core supply inputs, ensuring the microgrid’s basic energy supply while enhancing system resilience through a diversified energy portfolio. (2) Energy conversion and coupling side. As the core hub of the system, this side integrates combined heat and power (CHP) units, gas boilers (GB), power-to-gas (P2G) equipment, carbon capture, utilization, and storage (CCUS) devices, hydrogen fuel cells (HFC), and energy storage systems (ESS), enabling deep coupling and cascaded utilization of natural gas, electricity, heat, hydrogen, and carbon flows. Among these, the operational characteristics of key devices are as follows. The CHP unit consumes natural gas to simultaneously produce electricity and heat, with its power output P t CHP and heat output H t CHP satisfying the heat-to-power ratio constraint H t CHP = η CHP h / e P t CHP , where η CHP h / e is the heat-to-power ratio coefficient. The GB converts the chemical energy of natural gas into heat with a thermal efficiency of approximately 85–95%. The P2G device produces hydrogen via water electrolysis, and its hydrogen output P t P 2 G , H 2 and electrical input P t P 2 G satisfy the linear relationship P t P 2 G , H 2 = η P 2 G P t P 2 G , where η P 2 G is the power-to-hydrogen efficiency. The CCUS unit captures CO2 emitted from the system, with its captured amount E t CCUS being a quadratic function of its energy consumption E t CCUS = σ ( a CCUS ( P t CCUS ) 2 + b CCUS P t CCUS ) . The HFC converts hydrogen into electricity and heat, with electrical and thermal efficiencies of approximately 50–60% and 40–45%, respectively. The ESS enables multi-period shifting of electricity, heat, and hydrogen, with its state of charge S O C t governed by S O C t + 1 = S O C t + η cha P t cha P t dis / η dis , subject to charging/discharging power limits and capacity constraints. The mathematical operation models of these components are detailed in Ref. [23]. (3) Energy demand side. Energy is precisely supplied to electrical, thermal, and hydrogen loads through conversion and storage devices, thereby meeting diverse user demands and achieving efficient supply–demand matching. (4) “Green certificate-carbon market” trading side. Through bidirectional interfaces with the carbon trading market and the green certificate trading market, carbon emission costs and the environmental benefits of renewable energy are explicitly quantified, providing market-based support for low-carbon operation. In summary, the DIEM constructs an integrated generation-conversion-storage-trading operating framework through physical-layer multi-energy coupling and market-layer value linkage, supporting the low-carbon and efficient operation of microgrids in new power systems.

2.2. Design of the Bidirectional Interaction Mechanism for Multi-Microgrids

Building on the low-carbon operational architecture of the DIEM described above, a peer-to-peer (P2P) coordinated trading mechanism covering energy, carbon quotas, and green certificates is developed for multi-microgrid systems. With microgrid autonomy and mutual benefit as the core principles, each microgrid acts as an independent trading entity and conducts direct peer-to-peer transactions at the execution level. The distribution network operator, shown as a dashed box in Figure 2, does not participate in P2P price negotiation or in transaction volume determination. Its role is limited to market supervision, transaction compliance review, and distribution-network security verification. Therefore, the proposed mechanism preserves decentralized P2P trading among DIEMs while ensuring operational security and transaction compliance.
As shown in Figure 2, multi-microgrids form a three-dimensional coordinated trading system through the energy, carbon, and green certificate markets. Through autonomous negotiation and peer trading, energy consumption, carbon responsibility sharing, and green certificate circulation are accomplished simultaneously. This mechanism enhances on-site renewable energy utilization and operational flexibility while reducing microgrid electricity costs and grid dependence, thereby providing a direct trading scenario and market foundation for cooperative games among multi-microgrids. The coupling of energy, carbon, green certificate, and transaction flows forms a decentralized, efficient, and autonomous coordinated trading mode, further improving the economy, low-carbon performance, and reliability of the multi-microgrid system.

2.3. Dual-Incentive Mechanism Model for Green Certificate-Carbon Trading Based on Green-Carbon Offsetting

To incentivize distributed integrated energy microgrids to increase their renewable energy output share and reduce carbon emissions, a green certificate-carbon trading dual-incentive mechanism based on “green-carbon” offsetting is developed, as illustrated in Figure 3. The core of this mechanism is to exploit the low-carbon attribute value of green certificates and establish a bidirectional linkage between the green certificate trading (GCT) market and the carbon emission trading (CET) market.
A green certificate is an official certification of renewable energy generation fed into the grid. Each green certificate corresponds to an equivalent amount of carbon emission reduction from renewable energy generation, which can be quantified through carbon footprint tracking and carbon emission flow methods, and represents the low-carbon benefit relative to conventional coal-fired power [26]. Based on this quantified value, a bidirectional interaction channel is established between the green certificate market and the carbon market, with green certificate acquisition and carbon quota trading serving as the exchange media. Under this mechanism, a DIEM can convert surplus green certificates into carbon quotas to offset part of its own carbon emissions. These carbon quotas obtained through green certificate conversion are defined as green carbon. Through the “green-carbon” offsetting mechanism, a DIEM can simultaneously enhance renewable energy consumption and obtain additional carbon quota revenue, forming a positive incentive loop between green electricity production and carbon emission compliance.

2.3.1. Modeling of the Dynamic Reward-and-Penalty-Based Tiered Carbon Trading Mechanism

China’s carbon market is currently undergoing deepening development, and pilot projects have provided a foundation for distributed entities to participate in carbon trading. To clarify the low-carbon operational space for microgrids under regional carbon-emission constraints [27] and to promote deep coupling between carbon costs and operational strategies, a carbon-market participation model for DIEMs is established based on the consumption-side carbon-emission accounting principle, following the “polluter pays” rule. The carbon emission responsibility boundaries are defined as follows: (1) as end-use energy prosumers, DIEMs must bear direct and indirect carbon emission responsibilities throughout their operational lifecycle; (2) indirect carbon emissions from purchased electricity are uniformly quantified using the regional grid benchmark emission factors published by the national authority; (3) devices such as ESS and HFC, which only participate in energy conversion and storage processes without directly consuming fossil fuels, are excluded from the DIEM’s carbon emission accounting boundary. It should be clarified that the upstream carbon emissions associated with the electricity stored in the ESS are already accounted for in the “purchased electricity” term E e buy in Equation (2); similarly, the carbon footprint of hydrogen consumed by the HFC is already accounted for at the production stage (e.g., P2G hydrogen production or purchased hydrogen). Double-counting these emissions at the DIEM consumption side would overestimate actual carbon emissions and distort the incentive signals of the carbon trading mechanism. Therefore, in line with the “no double-counting” principle of carbon emissions accounting, this study excludes ESS and HFC from the terminal carbon emissions boundary.
To strengthen the incentive and constraint functions of the carbon market, a dynamic reward-and-penalty-based tiered carbon trading mechanism is introduced. This mechanism comprises three components: initial carbon quota allocation, actual carbon emission calculation, and tiered carbon trading cost accounting [28]. The main symbols and parameters used in the mathematical models are summarized in Appendix A, Table A1.
(1) Initial carbon quota model
To incentivize energy conservation and emission reduction, carbon quotas are allocated to DIEMs at no cost. Based on the carbon emission accounting boundary, DIEM carbon emissions mainly originate from two sources: (i) indirect carbon emissions from purchased electricity, which are accounted for on the consumption side; and (ii) direct carbon emissions from natural gas consumed by CHP and GB units. The initial carbon quota model is formulated as Equation (1):
E 0 = E e , 0 buy + E 0 CHP + E 0 GB E e , 0 buy = χ e t = 1 T P t buy E 0 CHP = χ h t = 1 T ( P t CHP + H t CHP ) E 0 GB = χ h t = 1 T H t GB
where E e , 0 buy is the carbon quota for purchased electricity; E 0 CHP and E 0 GB are the carbon quotas for CHP and GB, respectively; P t buy is the electricity purchased from the main grid at time slot t ; P t CHP is the CHP output power; H t CHP and H t GB are the heat outputs of CHP and GB at time t ; and T denotes the 24 h daily optimization horizon.
(2) Actual carbon emission model
Given that electricity purchased from the main grid is primarily supplied by coal-fired units, and that the system is equipped with CCUS devices capable of capturing CO2 and converting it into methane via P2G, the actual carbon emission model is formulated as Equation (2) based on the aforementioned accounting boundary:
E = E e buy + E total CHP , GB E CCUS E e buy = t = 1 T a 1 + b 1 P t buy + c 1 P t buy 2 E total CHP , GB = t = 1 T a 2 + b 2 P t CHP + H t CHP + H t GB + c 2 P t CHP + H t CHP + H t GB 2 E CCUS = t = 1 T ϖ P t CCUS
where E e buy is the indirect carbon emission from purchased electricity; E total CHP , GB is the direct carbon emission from CHP and GB; E CCUS is the CO2 captured by CCUS; ϖ is the capture parameter; P t CCUS is the total power consumption of CCUS at time t ; and a 1 , b 1 , c 1 and a 2 , b 2 , c 2 are the carbon emission calculation parameters for coal-fired and gas-fired units, respectively.
(3) Dynamic reward-and-penalty-based tiered carbon trading model
The carbon trading volume for the DIEM is given by Equation (3):
E t = E E 0
The tiered carbon trading cost model for the DIEM is expressed as [29]:
C C = α 0 C ( 1 + ε ) ( E t l ) α 0 C l , 2 l E t l α 0 C E t , l E t 0 α 0 C E t , 0 E t l α 0 C ( 1 + ν ) ( E t l ) + α 0 C l , l E t 2 l α 0 C ( 1 + 2 ν ) ( E t 2 l ) + α 0 C ( 2 + ν ) l , 2 l E t 3 l α 0 C ( 1 + 3 ν ) ( E t 3 l ) + α 0 C ( 3 + 3 ν ) l , 3 l E t 4 l α 0 C ( 1 + 4 ν ) ( E t 4 l ) + α 0 C ( 4 + 6 ν ) l , 4 l E t 5 l
When E t < 0 , the DIEM’s actual carbon emissions are below its allocated allowance, and the surplus carbon quota can be sold in the market. As the surplus increases, the selling price increases progressively in a tiered manner (each interval of length, with a price compensation coefficient ε ), rewarding deep emission reductions. When E t > 0 , the emissions exceed the allowance, and the DIEM needs to purchase additional quotas; as the purchase volume increases, the buying price increases progressively (with price growth coefficient ν ), penalizing high emissions. The first two piecewise segments in Equation (4) correspond to the “surplus selling” scenario ( 2 l E t 0 ), while the last five segments correspond to the “deficit purchasing” scenario ( 0 E t 5 l ). This tiered mechanism, through its nonlinear price gradients, provides positive incentives for low-carbon behavior and strong disincentives for high-carbon behavior.

2.3.2. Modeling of the Green Certificate-Carbon Trading Mechanism Based on “Green-Carbon” Offsetting

(1) Green Certificate Trading Mechanism
To quantify the environmental value of renewable energy generation in the DIEM and establish green certificate circulation rules aligned with compliance requirements, a basic green certificate trading model is developed. The model directly links renewable energy output to green certificate volume and defines the green certificate quota obligations for the DIEM, from which the green certificate trading volume is derived. A quantity-based Cournot model is then introduced to capture the price formation mechanism in the GCT market and dynamically reflect the market value of green certificates through supply–demand interactions, as expressed in Equations (5)–(7):
G r = α P t PV + P t WT 1000
G e = β L t e 1000
G GCT = G r G e
where P t PV and P t WT are the actual PV and wind power outputs at time slot t , respectively; L t e is the electrical load at time slot t . Since the scheduling interval is 1 h, dividing by 1000 converts kWh into MWh. One MWh of renewable electricity generation corresponds to one green certificate. α is the conversion coefficient from renewable electricity generation to green certificates, and β is the quota coefficient for green certificate demand.
The GCT price based on the quantity-competition Cournot model is given by Equations (8)–(10):
λ t GCT = k 1 k 2 G GCT
k 1 = λ 0 GCT
k 2 = 1 ζ λ 0 GCT / G e
(2) Green-Carbon Offsetting Mechanism
To fully leverage the low-carbon attributes of green certificates and realize value linkage between green certificates and carbon quotas, the green-carbon offset volume is quantified along three dimensions: renewable energy emission-reduction efficiency, renewable energy output share, and the ratio of actual carbon emissions to allocated carbon quotas, as given by Equations (11)–(13). A tiered green certificate quota reward-penalty function is also designed, linking the carbon trading volume after green-carbon offsetting to reward-and-penalty intensity through graduated coefficient adjustments to enhance emission reduction incentives.
E green = κ G green θ + ρ η green ξ ω green
η green = P t PV + P t WT P t PV + P t WT + P t buy + P t HFC , e + P t CHP + P t ESS , e , dis
ω green = E E 0
All three are normalized dimensionless indices (ranging from 0 to 1), quantifying the carbon reduction effect of renewable generation, the share of clean energy in the overall energy mix, and the contribution of carbon allowance surplus/deficit to the demand for green-carbon offsetting. The weighting coefficients ρ and ξ are determined by the analytic hierarchy process (AHP) combined with expert experience, with values set to ρ = 0.4 and ξ = 0.3 ; the fact that ρ > ξ reflects the relatively dominant contribution of renewable energy emission reduction efficiency to the green-carbon offset volume. The correction coefficient κ calibrates the value equivalence between the green certificate market and the carbon market; based on monthly trading data from the Yangtze River Delta region in 2023, it is fitted as κ = 1.05 . P t HFC , e is the electrical power generated by HFC at time t ; and P t ESS , e , dis is the discharge power of the electrical energy storage system.
The tiered green certificate quota reward-penalty function is expressed as Equations (14) and (15):
E t in = E E 0 E green
G CET = ω 1 + ω 2 l θ + ω 3 E t in + 2 l θ , E t in 2 l ω 1 l θ + ω 2 E t in + l θ , 2 l E t in l ω 1 E t in θ , l E t in l ω 1 l θ + ω 2 E t in l θ , l E t in 2 l ω 1 + ω 2 l θ + ω 3 E t in 2 l θ , E t in 2 l
where ω 1 , ω 2 , ω 3 are the green certificate quota reward-penalty coefficients, and they satisfy ω 1 < ω 2 < ω 3 .
The tiered green certificate quota reward-penalty volume model divides the system’s carbon quota trading volume into multiple intervals. When the system needs to purchase carbon quotas in the carbon market, additional green certificate quota constraints are imposed; conversely, when the system holds surplus carbon quotas available for sale, corresponding green certificate quota reductions are granted. Through this graduated quota adjustment, compliance pressure and incentive signals from the CET market are transmitted to the GCT market, forming bidirectional constraints and coordinated incentives between the two markets.
(3) Participation in Carbon Market Trading
On the basis of Equation (4), substituting Equation (14) into the carbon trading volume yields the tiered carbon trading model as Equation (16):
C C = α 0 C ( 1 + ε ) ( E t in l ) α 0 C l , 2 l E t in l α 0 C E t in , l E t in 0 α 0 C E t in , 0 E t in l α 0 C ( 1 + ν ) ( E t in l ) + α 0 C l , l E t in 2 l α 0 C ( 1 + 2 ν ) ( E t in 2 l ) + α 0 C ( 2 + ν ) l , 2 l E t in 3 l α 0 C ( 1 + 3 ν ) ( E t in 3 l ) + α 0 C ( 3 + 3 ν ) l , 3 l E t in 4 l α 0 C ( 1 + 4 ν ) ( E t in 4 l ) + α 0 C ( 4 + 6 ν ) l , 4 l E t in 5 l
Under the green certificate-carbon trading bidirectional interaction framework, an increase in renewable energy generation directly drives growth in green certificate holdings and correspondingly raises the green-carbon offset volume in the tiered CET mechanism, thereby incentivizing renewable energy production. Meanwhile, when the carbon quota trading volume enters a higher interval, the CET market price rises in a tiered manner, and the green certificate quota reward-penalty intensity also increases accordingly. If the system needs to purchase carbon quotas to fulfill compliance obligations, it will face green certificate quota penalties; otherwise, it will receive green certificate quota reductions. Through the dual transmission pathways of carbon market price signals and green certificate quota adjustments, this mechanism strengthens the coordinated linkage between the GCT and CET markets and effectively increases trading activity in both markets.

3. Cooperative-Game-Based Optimal Operation Model for Multi-Microgrid P2P Trading Considering Nash Bargaining

3.1. P2P Trading Model for Energy, Carbon Quotas, and Green Certificates Among Multi-Microgrids

3.1.1. Energy P2P Trading

The total P2P electricity trading volume of the i -th DIEM at time slot t is the sum of its P2P trading volumes with all other DIEMs, reflecting the net P2P electricity trading volume at that time slot, as expressed in Equation (17):
P t P 2 P , i , e = j ϕ i P t P 2 P , i j , e
where P t P 2 P , i , e is the total P2P electricity trading volume of the i -th DIEM at time slot t ; and ϕ i is the set of other DIEMs available for trading with DIEM   i . The trading direction is defined as follows: when P t P 2 P , i j , e > 0 , DIEM   i is the buyer and DIEM j is the seller; when P t P 2 P , i j , e < 0 , DIEM   i is the seller and DIEM j is the buyer.
The relationship between the sign of P2P trading volume and trading direction is defined as follows: for any DIEM   i , P t P 2 P , i , e > 0 indicates that DIEM   i is a net buyer (purchasing electricity), while P t P 2 P , i , e < 0 indicates that DIEM   i is a net seller (selling electricity). The sign conventions for carbon allowance trading E t P 2 P , i and green certificate trading G t P 2 P , i follow the same rule.
For the buyer and seller of the same P2P electricity transaction, a bilateral consistency constraint applies:
P t P 2 P , i j , e + P t P 2 P , j i , e = 0
For DIEM   i , the total P2P electricity trading cost is the sum of trading costs across all time slots, as in Equation (19):
C e P 2 P , i = t T j ϕ i P t P 2 P , i j , e λ t P 2 P , i j , e
where C e P 2 P , i is the P2P electricity trading cost.
Similarly, the P2P trading costs for heat and hydrogen are obtained as Equations (20) and (21):
C h P 2 P , i = t T j ϕ i P t P 2 P , i j , h λ t P 2 P , i j , h
C H 2 P 2 P , i = t T j ϕ i P t P 2 P , i j , H 2 λ t P 2 P , i j , H 2
where C h P 2 P , i and C H 2 P 2 P , i are the P2P trading costs for heat and hydrogen energy, respectively.
C energy P 2 P , i = C e P 2 P , i + C h P 2 P , i + C H 2 P 2 P , i

3.1.2. Carbon Quota P2P Trading

P2P carbon quota trading is a key pathway for optimizing carbon resource allocation across multiple microgrids. Its trading logic is consistent with P2P electricity trading and adjusts for surplus–deficit in carbon quotas through peer interactions. For the i -th DIEM, its total P2P carbon quota trading volume at time slot t is the sum of carbon quota trades with all other DIEMs, as in Equation (23):
E t P 2 P , i = j ϕ i E t P 2 P , i j
where E t P 2 P , i is the total P2P carbon quota trading volume of the i -th DIEM at time slot t . The trading direction convention for E t P 2 P , i j is identical to that of P2P electricity trading, following the direction definition in Equation (17).
To ensure transaction consistency and fairness, the buyer and seller of the same P2P carbon quota transaction must satisfy:
E t P 2 P , i j + E t P 2 P , j i = 0
For DIEM   i , the total P2P carbon quota trading cost is the sum of trading costs across all time slots:
C C P 2 P , i = t T j ϕ i E t P 2 P , i j λ t P 2 P , i j , C
where C C P 2 P , i is the P2P carbon quota trading cost.

3.1.3. Green Certificate P2P Trading

P2P green certificate trading is a key mechanism for promoting renewable energy consumption and transmitting green value signals. Its trading logic is also consistent with P2P electricity trading. For the i -th DIEM, its total P2P green certificate trading volume at time slot t is the sum of green certificate trades with all other DIEMs, as in Equation (26):
G t P 2 P , i = j ϕ i G t P 2 P , i j
where G t P 2 P , i is the total P2P green certificate trading volume of the i -th DIEM at time slot t . The trading direction convention is identical to that of P2P electricity trading. Refer to the definition of the trading direction in Equation (17).
For the buyer and seller of the same P2P green certificate transaction, a bilateral consistency constraint applies:
G t P 2 P , i j + G t P 2 P , j i = 0
For DIEM   i , the total P2P green certificate trading cost is the sum of trading costs across all time slots:
C G P 2 P , i = t T j ϕ i G t P 2 P , i j λ t P 2 P , i j , G
where C G P 2 P , i is the P2P green certificate trading cost.

3.2. Cooperative-Game-Based Optimal Operation Model for Multi-Microgrids Based on Nash Bargaining

3.2.1. Cooperative Operating Cost Model for Multi-Microgrids

(1) Objective Function
Each DIEM   i minimizes its total operating cost, which includes the interaction cost with the external grid, microgrid operation and maintenance costs, wind and solar curtailment costs, energy storage operation and maintenance costs, carbon trading costs, green certificate trading costs, and comprehensive P2P trading costs, as expressed in Equation (29):
C i DIEM = C utility , i + C units , i + C cut , i + C ESS , i + C C , i + C GCT , i + C total , i P 2 P
1) External interaction cost:
C utility , i = t = 1 T λ t CH 4 N t CH 4 , i + λ t Gb P t buy , i λ t Gs P t sell , i
2) Unit operation and maintenance cost:
C units , i = t = 1 T α 1 P t CHP , i + β 1 P t CHP , i 2 + α 2 H t GB , i + β 2 H t GB , i 2 + α 3 P t P 2 G , i + α 4 P t HFC , H 2 , i + α 5 P t CCUS , i + α 6
3) Wind and solar curtailment cost:
C cut , i = t = 1 T λ cur P t WT , cur , i + P t PV , cur , i
4) Energy storage operation and maintenance cost:
C ESS , i = t = 1 T ς P t ESS , e , cha , i + P t ESS , h , cha , i + P t ESS , H 2 , cha , i + P t ESS , e , dis , i + P t ESS , h , dis , i + P t ESS , H 2 , dis , i
5) Green certificate trading cost:
G ^ GCT , i = G r , i G e , i G CET , i G green , i
C GCT , i = t = 1 T G ^ GCT , i λ t GCT
6) P2P comprehensive trading cost:
C total , i P 2 P = C energy P 2 P , i + C C P 2 P , i + C G P 2 P , i
(2) Constraints
1) Energy supply–demand balance constraints:
P t PV , i + P t WT , i + P t buy , i + P t HFC , e , i + P t CHP , i + P t ESS , e , dis , i P t EL , i P t CCUS , i + j ϕ i P t P 2 P , i j , e = P t sell , i + P t ESS , e , cha , i + L t e , i
H t CHP , i + H t GB , i + H t HFC , h , i + P t ESS , h , dis , i + j ϕ i P t P 2 P , i j , h = P t ESS , h , cha , i + L t h , i
P t P 2 G , H 2 , i + P t ESS , H 2 , dis , i + j ϕ i P t P 2 P , i j , H 2 = P t MR , H 2 , i + P t HFC , H 2 , i + P t ESS , H 2 , cha , i + L t H 2 , i
P ^ t PV , i P t PV , i = P t PV , cur , i
P ^ t WT , i P t WT , i = P t WT , cur , i
P ^ t PV , i 1 δ PV P t PV , i P ^ t PV , i 1 + δ PV
P ^ t WT , i 1 δ WT P t WT , i P ^ t WT , i 1 + δ WT
2) Natural gas purchase constraint:
N min CH 4 , i N t CH 4 , i N max CH 4 , i
3) Natural gas flow balance constraint:
N t CH 4 , i + N t MR , CH 4 , i = N t CHP , i + N t GB , i
4) Carbon emission flow balance constraints:
E t CHP , C , i + E t GB , C , i E t CCUS , C , i + j ϕ i E t P 2 P , i j = E t out , i
E t CHP , C , i + E t GB , C , i E t P 2 G , C , i E t CCUS , s , i + j ϕ i E t P 2 P , i j = E t out , i
When N DIEMs form a cooperative coalition, the total operational cost of the multi-microgrid cooperative system is given by Equation (48):
C total DIEM = i = 1 N C i DIEM
where C total DIEM is the total operational cost of the multi-microgrid cooperative coalition.

3.2.2. Multi-Microgrid Cooperative Game Model Based on Nash Bargaining

In multi-microgrid cooperative operation, participants are concerned not only with their individual benefit gains but also with the fair and reasonable distribution of cooperative profits. Moreover, under the energy-carbon-green certificate coordinated trading mechanism, achieving supply–demand balance while reducing carbon emissions through cooperative operations is a key challenge. As an important branch of cooperative game theory, Nash bargaining theory is centered on collective rationality and social optimization, effectively characterizing multi-party cooperative interactions and providing a systematic analytical framework for balancing efficiency improvement with fair distribution. Unlike traditional optimization methods that focus solely on minimizing total cost, the Nash bargaining model introduces a disagreement point (i.e., the cost under non-cooperation) and a cooperative surplus-distribution mechanism, achieving Pareto optimality while satisfying individual rationality constraints.
The standard form of the Nash bargaining problem can be stated as follows: given N participants, each participant i has a disagreement point d i and a feasible payoff set U N , the Nash bargaining solution is the solution to the optimization problem in Equation (49):
max u U , u i d i i = 1 N u i d i w i
where u i is the payoff of DIEM   i under cooperation (the negative of operational cost, i.e., u i = C i DIEM ); d i is the payoff of DIEM   i without cooperation, i.e., the Nash disagreement point. Specifically, d i is obtained by solving an independent operation optimization problem for DIEM   i : setting all P2P trading variables to zero ( P t P 2 P , i , * = 0 , E t P 2 P , i = 0 , G t P 2 P , i = 0 ) in the objective function and constraints of Equations (29)–(47), retaining only each DIEM’s interactions with the external grid and natural gas network, and solving for the optimal individual operational cost C i DIEM , non . Then d i = C i DIEM , non . The term u i d i represents the incremental payoff (cooperative surplus) obtained by DIEM   i through cooperative operation. w i is the bargaining power weight of participant i , reflecting its contribution to the coalition.
In this study, equal bargaining weights are adopted, i.e., w 1 = w 2 = w 3 = 1 / 3 . The rationale is threefold: (1) the three DIEMs in the case study have similar installed capacities and load scales, belonging to homogeneous agents without significant differences in bargaining power; (2) the equal-weight setting reduces Equation (49) to the standard Nash bargaining solution, facilitating comparison with existing Nash bargaining studies [12,13,14,15] to validate the effectiveness of the model and algorithm; (3) in practical applications where significant differences exist in installed capacity, carbon emissions, or renewable penetration among DIEMs, the weights can be set proportionally to these indicators.
The Nash bargaining model above is a complex optimization problem with multiple variables, strong coupling, non-convexity, and nonlinearity. It involves multi-commodity trading decisions for electricity, heat, hydrogen, carbon quotas, and green certificates, as well as operational coupling constraints across DIEMs, making direct solution via traditional optimization methods infeasible. The original problem is therefore equivalently decomposed into two tractable subproblems: a cooperative coalition total-profit maximization subproblem and a multi-commodity P2P trading payment-negotiation subproblem.
Subproblem 1: cooperative profit maximization
u i = C i DIEM = C utility , i + C units , i + C cut , i + C ESS , i + C C , i + C GCT , i + C total , i P 2 P
max i = 1 N u i = C total DIEM
s.t.   Equations   37 47
Subproblem 2: multi-commodity P2P trading payment negotiation
max i = 1 N u i d i w i
To facilitate a solution, taking the logarithm of Equation (53) converts the product form into a summation form, as given by Equations (54)–(56):
max i = 1 N w i ln u i d i
max i = 1 N w i ln u i * j = 1 , j i N P t P 2 P , i j , e * λ t P 2 P , i j , e + P t P 2 P , i j , h * λ t P 2 P , i j , h + P t P 2 P , i j , H 2 * λ t P 2 P , i j , H 2 + E t P 2 P , i j * λ t P 2 P , i j , C + G t P 2 P , i j * λ t P 2 P , i j , G d i
  u i * j = 1 , j i N P t P 2 P , i j , e * λ t P 2 P , i j , e + P t P 2 P , i j , h * λ t P 2 P , i j , h + P t P 2 P , i j , H 2 * λ t P 2 P , i j , H 2 + E t P 2 P , i j * λ t P 2 P , i j , C + G t P 2 P , i j * λ t P 2 P , i j , G d i
In the proposed implementation, the Nash bargaining allocation is realized endogenously through the negotiated settlement prices for electricity, heat, hydrogen, carbon quotas, and green certificates. The resulting bilateral settlement payments are included in the comprehensive P2P trading cost C total , i P 2 P defined in Equation (36) and are therefore already reflected in the reported operating cost of each DIEM. No additional ex-post lump-sum side-payment variable is introduced in the proposed framework; therefore, separate side-payment values are not reported.

3.3. Nash Bargaining Equilibrium Conditions

The Nash bargaining solution is the unique solution satisfying a set of axiomatic conditions in cooperative game theory. Starting from the axiomatic theory of Nash bargaining, this section proves that the solution to the multi-microgrid cooperative game optimization model in Equation (53) satisfies the Nash bargaining equilibrium conditions.
For a cooperative game with N participants, given a disagreement point d i = ( d 1 , d 2 , , d N ) and a feasible payoff set U N , Nash proved that there exists a unique bargaining solution u i * = ( u 1 * , u 2 * , , u N * ) satisfying the following axioms:
(1) Individual Rationality:
u i * d i , i N , each participant obtains a payoff no lower than the disagreement point. Since the payoff is defined as the negative of the operating cost, i.e., u i = C i DIEM and d i = C i DIEM , non , the individual rationality condition u i d i is equivalently expressed in the cost domain as C i DIEM C i DIEM , non .
(2) Pareto Optimality:
By contradiction, suppose another feasible solution u U exists such that u i u i * for all i , with at least one strict inequality. Since the objective function i = 1 N u i d i w i is strictly increasing in each component, then i = 1 N u i d i w i > i = 1 N u i * d i w i , contradicting the fact that u i * is the optimal solution of the maximization problem in Equation (53). Therefore, u i * must be Pareto optimal.
(3) Symmetry:
If two DIEMs i and j have identical cost functions, disagreement points d i = d j , and bargaining weights w i = w j , the model objective function is symmetric with respect to u i and u j . If the optimal solution satisfies u i * u j * , swapping the values of i and j yields another optimal solution, contradicting the uniqueness of the solution. Therefore, u i * = u j * must hold.
(4) Invariance to Affine Transformations:
For any affine transformation u i = α i u i + β i   ( α i > 0 ) applied to the payoff of each DIEM   i , the disagreement point correspondingly becomes d i = α i d i + β i . The resulting objective function is expressed as Equation (57):
i = 1 N ( u i d i ) w i = i = 1 N ( α i u i + β i α i d i β i ) w i = i = 1 N α i w i ( u i d i ) w i = i = 1 N α i w i i = 1 N ( u i d i ) w i
Since i = 1 N α i w i is a positive constant, the solution to the optimization problem is unchanged, i.e., u i * = α i u i * + β i .
In summary, the solution of the above model satisfies the axiomatic conditions of Nash bargaining. This completes the proof.

4. Homogeneous Multi-Agent Deep Reinforcement Learning Solution Framework for Multi-Microgrids

To address the optimization challenges of the Nash bargaining cooperative game model under complex constraints, this paper proposes a multi-agent deep reinforcement learning method based on centralized training and decentralized execution. By modeling the cooperative game as a Markov decision process, this study designs a multi-agent deep reinforcement learning framework using the MASAC algorithm to establish collaborative interactions among multiple DIEMs. The framework maps the strategic interactions in the Nash bargaining game into a dynamic state-action-reward iterative optimization process. A centralized critic network approximates the weighted Nash product objective, while decentralized actor networks enable autonomous decision-making for each DIEM. This approach effectively overcomes the limitations of traditional methods in non-convex, non-linear, and multi-commodity trading scenarios, and it provides an end-to-end adaptive solution for the cooperative operation of multi-microgrids in a coordinated energy-carbon-green certificate market.

4.1. Markov Decision Process Modeling

The optimal operation problem of the cooperative multi-microgrid game is formulated as a Markov decision process, defined by the 5-tuple MDP = S , A , P , R , γ . By defining the state space S , action space A , state transition function P , reward function R , and discount factor γ , cooperative game strategies are mapped to sequential decisions. As homogeneous agents, the DIEMs share identical structures for their state space, action space, and reward functions, but are coupled through peer-to-peer transactions. Agent interactions are governed by a centralized training and decentralized execution framework. During training, a centralized critic network observes global information to evaluate the joint policy; during execution, each DIEM makes independent decisions based solely on local observations.
Let N = { 1 , 2 , , N } denote the set of agents, where each agent i corresponds to a DIEM and has a homogeneous state space, action space, and reward function structure.

4.1.1. State Space S

The local state s t i of agent i at time step t contains information reflecting its operating status and the external market environment:
s t i = L t e , i , L t h , i , L t H 2 , i , P t WT , i , P t PV , i , SOC t ESS , e , i , SOC t ESS , h , i , SOC t ESS , H 2 , i , P t sell , i , P t buy , i , λ t GCT , i , λ t P 2 P , i j
λ t P 2 P , i j = λ t P 2 P , i j , e , λ t P 2 P , i j , h , λ t P 2 P , i j , H 2 , λ t P 2 P , i j , C , λ t P 2 P , i j , G
s t = s t 1 , s t 2 , s t N
where λ t P 2 P , i j is the set of P2P transaction prices for agent i at time slot t , and s t is the joint state set of all agents at time slot t .

4.1.2. Action Space A

The continuous action a t i of agent i at time slot t comprises microgrid outputs, energy storage charging and discharging, and multi-commodity P2P trading volumes. The equipment in DIEM is relatively complex and can be categorized into energy conversion devices and energy storage systems.
a t i = P i , w , t , B i , w , t , P t P 2 P , i j , E t P 2 P , i j , G t P 2 P , i j
P t P 2 P , i j = P t P 2 P , i j , e , P t P 2 P , i j , h , P t P 2 P , i j , H 2
a t = a t 1 , a t 2 , a t N
where P i , w , t is the power output of the energy conversion device w for agent i at time slot t , B i , w , t is the energy output or input of the energy storage device w for agent i at time slot t , P t P 2 P , i j is the set of P2P electricity, heat, and hydrogen trading volumes for agent i at time slot t ; and a t is the joint action set of all agents at time slot t .
To ensure the feasibility of continuous actions generated by decentralized agents, constraint-handling and post-processing mechanisms are applied to the action outputs. First, local equality constraints, including the energy, gas, and carbon flow balance constraints in Equations (37)–(39) and (45)–(47), are handled using a penalty-based method during training. Any violation of these balance constraints introduces a large negative penalty into the reward function, thereby guiding the actor networks to learn feasible operating policies. In addition, a projection layer is applied to keep equipment outputs, storage charging and discharging powers, and trading quantities within their physical upper and lower bounds. For bilateral P2P transactions, the consistency condition P i j , t P 2 P + P j i , t P 2 P = 0 should be satisfied. To meet this condition in decentralized execution, the agents’ raw continuous outputs are treated as trading bids. Before physical dispatch, a lightweight bilateral matching rule is applied to determine the actually executed trading volume, where P i j , t P 2 P , actual = ( P i j , t P 2 P , bid P j i , t P 2 P , bid ) / 2 and P j i , t P 2 P , actual = P i j , t P 2 P , actual . This post-processing rule ensures consistency in bilateral trading while preserving the decentralized execution structure.

4.1.3. State Transition Function P

The system state transitions are driven by the dynamic characteristics of the DIEM and source-load uncertainties, formulated as s t + 1 ~ P ( s t + 1 | s t , a t ) , where a t = [ a t 1 , a t 2 , , a t N ] is the joint action. The state transition function is jointly determined by the DIEM’s physical dynamics and the evolution of external market prices.

4.1.4. Reward Function R

The weighted Nash bargaining model is integrated into the multi-agent reinforcement learning training objective. The immediate reward r t i for agent i at time slot t is defined as the negative of its individual operational cost C i , t DIEM , consistent with the individual cost objective rather than the coalition total. The detailed calculation is provided in Section 3.2.1:
r t i = C i , t DIEM
To achieve the Nash bargaining objective of the cooperative game, centralized training employs a global reward function that approximates the Nash product. During the early stages of training, exploratory and suboptimal policies may cause an agent’s payoff to fall below the disagreement point, resulting in a non-positive argument ( R i t d i 0 ), which would make the logarithm undefined. To handle this, a clipping mechanism with a small positive constant ε and a penalty term is introduced. The global reward function is formulated as follows:
r t global = i N w i ln max R i t d i , ε i N P i t
R i t = = 1 t r i ,
where r t global is the global reward for all agents at time slot t ; R i t represents the accumulated payoff of agent i up to time t ; d i is the disagreement point, namely the payoff under non-cooperation; and ε is a small positive constant, e.g., 10−3, used to keep the logarithm well defined. P i t is a large penalty applied when R i t d i 0 , discouraging policies that violate individual rationality. This formulation computes the global reward from individual costs and remains consistent with the decentralized execution architecture and the Nash bargaining model. Although R i t is accumulated progressively within one daily scheduling episode, while d i is obtained from the full-day independent-operation benchmark, this comparison is used as a reward-shaping mechanism during training. The individual rationality condition is finally evaluated at the terminal time t = T , where R i T and d i are defined over the same full-day horizon.

4.2. Principles and Framework of the Deep Reinforcement Learning Algorithm for Multi-Microgrids Based on CTDE-MASAC

To address non-convexity and non-stationarity in cooperative multi-microgrid games, deep reinforcement learning offers an effective solution, as it adapts well to complex, dynamic environments. For continuous action-space optimization, SAC, an off-policy actor-critic algorithm based on the maximum entropy framework, introduces a policy entropy term into the conventional cumulative reward objective. This constructs a joint optimization objective that maximizes both the expected cumulative reward and the policy entropy, effectively balancing decision-making performance with exploration. To this end, the SAC algorithm is extended to a multi-agent variant, namely MASAC, to meet the requirements of continuous action space optimization in cooperative multi-microgrid games. The effective operation of MASAC relies on a suitable multi-agent deep reinforcement learning framework, within which the centralized training and decentralized execution strategy is highly compatible.
The core principle of the CTDE framework is that agents are trained centrally but execute independently in a decentralized manner. During training, the centralized critic network accesses the states and actions of all agents and uses global information to evaluate the long-term value of the joint policy. This mechanism effectively mitigates the non-stationarity inherent in multi-agent environments. Because each agent’s policy updates alter the state distributions of others, global information is essential to accurately model this interdependence. However, during execution, agents rely solely on local observations and pre-trained decentralized actor networks to make independent decisions. This eliminates the need for real-time communication, thereby ensuring both privacy protection and communication efficiency in practical deployment [30].

4.2.1. MASAC Algorithm Principles [31,32]

The MASAC algorithm extends SAC to multi-agent environments, where each agent maintains an independent decentralized actor network and all agents share a centralized critic network. This centralized critic takes the global state and the joint actions of all agents as inputs to estimate the long-term cumulative reward.
(1) Maximum Entropy Reinforcement Learning Objective:
J ( π ) = t = 1 T E ( s t , a t ) ~ ρ π r ( s t , a t ) + α H ( π ( | s t ) )
where J ( π ) is the single-agent maximum entropy objective function, ρ π is the joint state-action distribution induced by the policy π , H ( π ( | s t ) ) = E a t ~ π log π ( a t | s t ) is the policy entropy, and α is the temperature parameter.
For multi-agent systems, the objective function of MASAC is extended to:
J ( π 1 , , π N ) = t = 0 E ( s t , a t ) ~ ρ π r ( s t , a t ) + α i = 1 N H ( π i ( | s t i ) )
(2) Definition of the Soft Value Function
The soft state value function represents the expected cumulative soft reward obtained at state s t when following policy π :
V ( s t ) = E a t ~ π Q ( s t , a t ) α log π ( a t | s t )
The soft action value function satisfies the soft Bellman equation:
Q ( s t , a t ) = r ( s t , a t ) + γ E s t + 1 ~ p V ( s t + 1 )
where α log π ( a t | s t ) is the entropy regularization term, γ is the discount factor, and p is the environmental state transition probability distribution.
(3) Critic Network Update:
MASAC employs two centralized Critic networks, Q θ 1 ( s , a 1 , , a N ) and Q θ 2 ( s , a 1 , , a N ) , along with their corresponding target networks, Q θ ¯ 1 and Q θ ¯ 2 . The double-Critic design estimates the target value by selecting the minimum of the two Q-values, thereby mitigating overestimation.
The loss function for the critic networks is defined as the soft Bellman residual:
L Q ( θ j ) = E ( s , a , r , s t + 1 ) ~ D Q θ j ( s t , a t ) y ¯ 2 , j = 1 , 2
The target value y ¯ is formulated as:
y ¯ = r + γ E a t + ! ~ π min j = 1 , 2 Q θ ¯ j ( s t + 1 , a t + 1 ) α i = 1 N log π i ( a t + 1 i | s t + 1 i )
where L Q ( θ j ) is the loss function of the j -th critic network, ( s , a , r , s t + 1 ) is a transition tuple sampled from historical interactions, y ¯ is the critic target value, min j = 1 , 2 Q θ ¯ j ( s t + 1 , a t + 1 ) is the minimum value between the two target critic networks, and θ ¯ j are the parameters of the target critic networks.
(4) Actor Network Update:
The Actor network π ϕ i ( a t i | s t i ) of each agent is updated by minimizing the KL divergence, with the objective of aligning the policy distribution with the exponentially transformed soft Q-value distribution:
J π ( ϕ i ) = E s ~ D , a i ~ π ϕ i α log π ϕ i ( a t i | s t i ) min j = 1 , 2 Q θ j ( s , a 1 , , a N )
In gradient computation, the stochastic policy is sampled using the reparameterization trick:
a t i = f ϕ i ( ϵ t i ; s t i )
Reparameterization enables gradients to be back-propagated directly through the sampling process:
^ ϕ i J π ( ϕ i ) = E ϵ ~ N α ϕ i log π ϕ i ( a t i | s t i ) + a i α log π ϕ i ( a t i | s t i ) a i min j Q θ j ( s t , a t ) ϕ i f ϕ i ( ϵ t i ; s t i )
where J π ( ϕ i ) is the objective function of the i -th agent’s actor network, f ϕ i is the reparameterized policy function, ϵ t i is a noise vector sampled from a fixed distribution, and ϕ i J π ( ϕ i ) is the gradient of the Actor objective function with respect to the parameters ϕ i .
(5) Adaptive Temperature Coefficient:
The temperature parameter α controls the intensity of entropy regularization. An excessively large value promotes over-exploration, whereas an excessively small value leads to insufficient exploration. To tackle this, SAC incorporates an adaptive adjustment mechanism that dynamically optimizes α by minimizing the following objective function:
L ( α ) = E a i ~ π i α log π i ( a i | s i ) α H ¯
where L ( α ) is the adaptive loss function for the temperature parameter α , and H ¯ is the target entropy.
(6) Soft Update of the Target Network:
To stabilize training, the parameters of the target networks are updated via a soft-update mechanism, smoothly tracking the current network parameters:
θ ¯ j τ θ j + ( 1 τ ) θ ¯ j , j = 1 , 2
where τ is the soft update coefficient.

4.2.2. Extraction of Marginal Contributions and Solving for P2P Transaction Payment Schemes

Once the MASAC algorithm converges, each DIEM’s actor network obtains an optimal cooperative operating policy. However, resolving the cooperative game further requires determining internal settlement prices to ensure a fair distribution of benefits. In economics, marginal contribution refers to the impact on total system revenue from adding a participant or a unit of a resource. The centralized critic network Q s , a ; θ approximates the long-term cumulative value of the weighted Nash product. According to the envelope theorem, at the optimal strategy, the partial derivative of the value function with respect to a parameter equals the partial derivative of the Lagrangian function with respect to its corresponding constraint. Therefore, the partial derivatives of the critic network carry the economic significance of shadow prices. By treating the critic network as a close approximation of the true value function, the nodal marginal values can be extracted by calculating the partial derivatives:
λ t i , e = Q ( s , a ) P t P 2 P , i , e ( s , a * )
According to the Nash bargaining solution, the transaction price must reflect the marginal valuations of both parties. The average marginal value of the two trading entities is adopted as the P2P settlement price:
λ t P 2 P , i j , e = λ t i , e + λ t j , e 2
Although the critic network is a neural approximation and its gradients may contain slight numerical inaccuracies, averaging the marginal valuations (as formulated in the settlement price equations) effectively mitigates unilateral approximation errors and balances the negotiation. The empirical stability of these derived prices is further examined in the subsequent case studies.

4.2.3. A Deep Reinforcement Learning Framework for Multi-Microgrids Based on CTDE-MASAC

Based on the above principles, the deep reinforcement learning framework for multi-microgrids using CTDE-MASAC is constructed as shown in Figure 4. This framework utilizes the environment module as the data hub to integrate all operational information from the multi-microgrid system and output local observations s t i for each DIEM. Based on these local observations, the actor network of each DIEM generates continuous action strategies a t i . These actions encompass P2P transaction volumes, microgrid power outputs, and energy storage charging/discharging states. Upon receiving the joint actions of all DIEMs, the environment module updates the system state, computes the global reward r ( s t , a t ) , and stores the transition tuple ( s , a , r , s t + 1 ) into the experience replay buffer. During the training phase, the centralized critic network has access to the states and actions of all DIEMs. It evaluates the long-term value of the joint policy by minimizing the soft Bellman residual. Concurrently, each actor network computes its policy gradients and updates its parameters using the critic’s value estimates and an entropy regularization term. Data decoupling through the experience replay buffer and the clipped double-critic mechanism effectively mitigates Q-value overestimation. Meanwhile, target networks smoothly track updates through a soft-update mechanism. Upon completing the training, the framework transitions to the decentralized execution phase. Each DIEM relies exclusively on its local observations and trained actor network to generate real-time dispatch commands independently, eliminating the need for online inter-agent communication. Through structured training iterations, this paradigm combines global guidance during centralized training with autonomous decision-making during decentralized execution, achieving adaptive convergence toward the Nash bargaining equilibrium. To further clarify the implementation process of the proposed CTDE-MASAC framework, the training and decentralized execution procedure is summarized in Algorithm 1.
Algorithm 1. CTDE-MASAC training and decentralized execution procedure
Input: Multi-microgrid environment; agent set N ; replay buffer D ; batch size M ; discount factor γ ; soft-update coefficient τ ; temperature coefficient α .
Output: Trained actor policies π ϕ i for all DIEM agents.
1.  Initialize actor networks π ϕ i for all DIEM agents.
2.  Initialize centralized critic networks Q θ 1 and Q θ 2 .
3.  Initialize target critic networks Q θ ¯ 1 and Q θ ¯ 2 , and initialize replay buffer D .
4.  for each training episode do
5.   Reset the environment and obtain the initial joint state s t .
6.   for each time step t do
7.    Each DIEM observes the local state s t i and generates an action a t i using π ϕ i .
8.    Construct the joint action a t = [ a t 1 , a t 2 , , a t N ] and execute a t .
9.    Calculate R t i , and r t according to Equations (64)–(66).
10.     Store transition ( s t , a t , r t , s t + 1 ) in D .
11.     Sample a mini-batch of M transitions from D .
12.     Update Q θ 1 and Q θ 2 by minimizing the critic loss in Equation (71).
13.     Update each actor policy π ϕ i according to Equation (73).
14.     Update the temperature coefficient α according to Equation (76).
15.     Softly update Q θ ¯ 1 and Q θ ¯ 2 according to Equation (77).
16.    end for
17. end for
18. During decentralized execution, each DIEM uses s t i and π ϕ i to generate a t i .

5. Case Study

5.1. Basic Data

A typical industrial park microgrid cluster in the Yangtze River Delta region is selected for the case study. This cluster comprises three distributed, integrated energy microgrids, denoted DIEM 1, DIEM 2, and DIEM 3, each connected to the 10 kV distribution network via a 10 kV/0.4 kV transformer. The energy load and renewable power generation forecast data for each microgrid are based on statistical analyses of actual operational data from the park over the past three years, as shown in Figure 5. The basic market, device, and trading parameters are listed in Table 2, including carbon trading parameters, green certificate trading parameters, and device operation coefficients.
The software platform uses Windows 11 and Python 3.10. Deep neural networks are built with PyTorch 2.0, and parallel computing acceleration is enabled using an NVIDIA RTX 4060 GPU and CUDA 12.7.
To ensure the reproducibility of the proposed deep reinforcement learning method, the detailed neural network architectures and key training hyperparameters of the CTDE-MASAC algorithm are explicitly summarized in Table 3. Both the actor and critic networks utilize Multi-Layer Perceptron (MLP) architectures with two hidden layers.
Justification of modeling assumptions and key parameters: The parameters utilized in this study are carefully selected to reflect realistic engineering scenarios. The baseline carbon trading price and benchmark green certificate price are set according to the parameter values in Table 2, based on recent historical average clearing prices in the Yangtze River Delta regional pilot carbon market and the national green electricity market in China. Additionally, a fundamental modeling assumption in the cooperative game is the equal market status of the participants. Consequently, the Nash bargaining weights are set symmetrically ( w 1 = w 2 = w 3 = 1 / 3 ), assuming that all three industrial park microgrids possess equal bargaining power in the decentralized P2P market without any monopolistic dominance. The capacities of the CHP, P2G, and ESS units are scaled proportionally to the actual peak load demands of typical medium-sized industrial parks to ensure the feasibility of the energy-carbon balancing mechanism.

5.2. Analysis of Algorithm Performance

To verify the adaptability of the proposed centralized training and decentralized execution model for multi-microgrid cooperative games in continuous action spaces, comparative experiments are conducted with three mainstream deep reinforcement learning algorithms: MADDPG, MASAC, and Multi-Agent Proximal Policy Optimization (MAPPO). Figure 6 shows the training reward curves of one representative random seed and is used to illustrate the convergence behavior of different algorithms in the continuous action space. The statistical results over five independent random seeds are reported in Table 4. The curves indicate that the MASAC algorithm demonstrates a clear performance advantage. Compared with the deterministic policy used in MADDPG, MASAC’s stochastic policy provides stronger exploration. Compared to MAPPO’s online policy, the off-policy nature of MASAC enables full reuse of historical experiences from the replay buffer, thereby significantly improving data utilization efficiency. Therefore, MASAC better adapts to the interaction characteristics and cooperative game requirements of multi-microgrids, providing reliable methodological support for coordinated system optimization.
To rigorously evaluate the optimality and statistical robustness of the proposed MASAC framework, a centralized Mixed-Integer Linear Programming (MILP) model with perfect foresight is introduced as the benchmark for the global theoretical optimum. To avoid drawing conclusions from a single training run, all DRL algorithms (MASAC, MADDPG, and MAPPO) are trained and tested across five independent random seeds. Table 4 presents the mean and standard deviation of the daily total operating cost.
To quantitatively highlight the advantages of the proposed approach, the results in Table 4 demonstrate that MASAC significantly outperforms existing DRL methods. Specifically, MASAC reduces the optimality gap to a mere 1.25% relative to the centralized MILP baseline, whereas MADDPG and MAPPO exhibit larger gaps of 4.02% and 3.11%, respectively. Furthermore, the standard deviation of MASAC (±142.50 CNY) is roughly 65% lower than that of MADDPG, proving superior algorithmic stability. The primary advantage of the proposed CTDE-MASAC approach lies in its ability to achieve near-optimal economic dispatch with millisecond-level execution time (0.12 s) in a fully decentralized manner. This completely avoids the heavy communication overhead and strict data privacy issues inherent in centralized optimization methods like MILP, making it highly practical for real-world multi-agent microgrid operations. It should be noted, however, that the centralized MILP serves as a perfect-foresight, in-sample benchmark. Thus, the 1.25% optimality gap demonstrates the algorithm’s convergence capacity under known deterministic conditions, rather than a guaranteed bound for out-of-sample performance.
Regarding agent configuration, the multi-microgrid system comprises N = 3 agents that are structurally homogeneous in terms of neural network topology and action dimensions, which simplifies real-world decentralized implementation. However, they are highly heterogeneous regarding their physical parameters, including distinct renewable energy penetration rates, industrial/commercial load profiles, and carbon emission boundaries. In terms of scalability, since the trained policies are executed locally by each agent, the online computational complexity scales linearly as O ( N ) . Although this theoretical O ( N ) scalability highlights the architectural advantage of decentralized execution, rigorous empirical testing on large-scale multi-microgrid clusters is reserved for future research.
To rigorously demonstrate the robustness of the proposed CTDE-MASAC framework against source-load uncertainties, an out-of-sample evaluation is conducted. Using Monte Carlo simulations, 10 unseen testing scenarios are generated by injecting random Gaussian noise within a ±15% fluctuation range into the baseline WT, PV, and electrical/thermal/hydrogen load profiles shown in Figure 5. The pre-trained MASAC policies are then executed directly on these 10 unseen scenarios in a decentralized manner without retraining.
The execution results show that the trained agents can dynamically adjust energy storage outputs and P2P trading volumes to absorb stochastic fluctuations. Throughout all out-of-sample testing scenarios, the energy supply–demand balance constraints are satisfied without load shedding. As shown in Table 5, compared with the baseline deterministic cost of 67,673.24 CNY, the out-of-sample costs range from 65,545.80 CNY to 69,815.30 CNY, with fluctuations between −3.14% and +3.17%. These results indicate that the proposed strategy maintains good generalization capability and robustness under source-load uncertainty.

5.3. Analysis of Optimized Operating Strategies

Figure 7 illustrates the electricity, heat, and hydrogen power supply–demand balance mechanisms for DIEM 1, DIEM 2, and DIEM 3 from a microgrid-level scheduling perspective. At the electric power level, each microgrid matches the electric load through the collaborative use of distributed energy resources, energy storage systems, and P2P transactions. During peak-load periods, local power sources and storage are prioritized for discharge. During off-peak periods, storage charging and power purchases are used for replenishment, thereby achieving a dynamic intra-day power balance. At the thermal energy level, microgrids rely primarily on CHP and GB as heat sources, combining thermal energy storage with P2P heat trading to form a coordinated heating system. In this system, DIEM 2 serves as the primary provider of a stable heat supply, while DIEM 1 and DIEM 3 meet their own heat demand through heat procurement. Overall, thermal output remains stable and closely aligns with the heat demand curve. At the hydrogen energy level, each microgrid constructs a supply–demand system based on P2G, hydrogen storage discharge, and P2P hydrogen trading. DIEM 2 serves as the primary hydrogen seller, while DIEM 1 acts as the primary buyer. DIEM 3 remains approximately balanced after a small initial hydrogen export. The hydrogen power output generally aligns with the hydrogen load curve. Through multi-energy complementarity and P2P transactions, microgrids effectively ensure stable operation and supply–demand balance across electricity, heat, and hydrogen systems.
Figure 8 illustrates the dynamic SOC variations in electrical energy storage in the multi-microgrid system. Between 1:00–10:00 and 15:00–20:00, the SOC ranges from 20% to 90%. Through frequent charging and discharging, the storage system responds rapidly to system power fluctuations and supports intra-day peak shaving and real-time energy balancing.
The interaction characteristics of energy trading power and prices among multi-microgrids are shown in Figure 9. From the perspective of electrical energy trading, the power exhibits high-frequency fluctuations between −100 kW and 100 kW, and the transaction price fluctuates synchronously with time-of-use (TOU) electricity prices. During the peak-load periods of 10:00–15:00 and 18:00–22:00, the P2P electricity trading prices rise toward the time-of-use purchase-price bound while remaining above the feed-in tariff, reflecting the price-guiding role of external electricity tariffs in P2P energy transactions. In terms of thermal energy trading, DIEM 2 consistently exports heat with negative power, whereas DIEM 3 imports heat with positive power. The transaction price fluctuates slightly within the upper and lower heat price limits, indicating a stable supply–demand matching relationship. In terms of hydrogen energy trading, DIEM 2 exports hydrogen with a constant negative power, while DIEM 1 purchases hydrogen with a constant positive power, and the transaction prices adjust smoothly within the upper and lower bounds of the hydrogen price. Overall, electricity, heat, and hydrogen exhibit different interaction characteristics in P2P trading. Electricity trading is most strongly affected by price signals, whereas heat and hydrogen trading show more stable supply–demand structures and pricing mechanisms.
To address potential concerns regarding the marginal values extracted from the neural approximation of the centralized critic network, the P2P settlement prices in Figure 9 are examined empirically. The derived trading prices for electricity, heat, and hydrogen do not exhibit divergent spikes and remain within their corresponding market bounds. These observations indicate that the gradients extracted from the converged centralized critic can provide stable approximate marginal-value signals for P2P settlement pricing.
Furthermore, a quantitative market-clearing check is conducted for decentralized P2P transactions. Based on the bilateral matching rule in Section 4.1.2, the net P2P trading volume among the three DIEMs is approximately zero at each time slot for electricity, heat, and hydrogen. The maximum absolute mismatch is below 10−6 kW, indicating that the total buying volume equals the total selling volume within numerical tolerance at the settled prices. In addition, the P2P carbon quota and green certificate trading volumes under Scenario 3 sum to zero across the three DIEMs. These results further confirm that the derived settlement prices support stable market clearing.

5.4. Analysis of Economic and Environmental Benefits

To further analyze the effects of the green certificate-tiered carbon trading mechanism on the economic and environmental performance of the multi-microgrid system, three comparative scenarios are established, as shown in Table 6.
Scenario 1: Neither the carbon trading mechanism nor the green certificate trading mechanism is considered, and only the total operating cost is minimized.
Scenario 2: The tiered carbon trading mechanism is considered, and the carbon trading cost is incorporated into the optimization model.
Scenario 3: The green certificate-tiered carbon trading mechanism is considered, with economic operating costs, carbon trading costs, and green certificate trading costs included simultaneously.
Table 7 presents the economic and environmental benefit results of the DIEMs under the three scenarios. Following the introduction of the tiered carbon trading mechanism in Scenario 2, carbon trading costs are higher than in Scenario 1. However, while carbon emissions are reduced by 27.91%, 22.04%, and 40.94%, respectively, the total operating costs increase by only 4.38%, 4.23%, and 2.61%, respectively, reflecting the constraining and guiding role of carbon cost internalization in shaping the system’s low-carbon operation strategies. Scenario 3 introduces a green certificate trading mechanism on top of tiered carbon trading, further reducing each DIEM’s carbon emissions. Although green certificate trading costs are introduced, Scenario 3 further reduces the carbon trading costs of the three DIEMs by 37.27%, 40.13%, and 33.82%, respectively, compared with Scenario 2, and reduces their total operating costs by 0.53%, 0.68%, and 0.46%, respectively. This demonstrates that the green certificate-tiered carbon trading mechanism can constrain carbon emissions through carbon trading and incentivize the use of clean energy through green certificate trading. While achieving deep carbon emission reductions, it effectively mitigates the impact of carbon costs on system economic efficiency, providing robust support for multi-microgrids to realize the synergistic optimization of economic and environmental benefits under low-carbon constraints.
Table 8 reports the non-cooperative operating costs and Nash bargaining outcomes under Scenario 3. To obtain the disagreement points, each DIEM is optimized independently under the same load profiles, renewable energy forecasts, device parameters, and external carbon and green certificate market settings as those adopted in Scenario 3. At the same time, all inter-DIEM P2P trading variables are fixed to zero. The resulting non-cooperative operating costs for DIEM 1, DIEM 2, and DIEM 3 are 26,368.07 CNY, 24,736.92 CNY, and 22,603.23 CNY, respectively, giving a total non-cooperative operating cost of 73,708.22 CNY. Under cooperative operation based on the Nash bargaining mechanism, the operating costs of DIEM 1, DIEM 2, and DIEM 3 decrease to 24,502.44 CNY, 22,619.51 CNY, and 20,551.29 CNY, respectively, and the total coalition operating cost decreases to 67,673.24 CNY.
The corresponding cooperative cost savings of DIEM 1, DIEM 2, and DIEM 3 are 1865.63 CNY, 2117.41 CNY, and 2051.94 CNY, representing cost reduction rates of 7.08%, 8.56%, and 9.08%, respectively. At the coalition level, cooperative operation yields a total cost saving of 6034.98 CNY and an overall cost reduction rate of 8.19%. In the proposed implementation, the Nash bargaining allocation is realized through the negotiated multi-commodity P2P settlement prices. The associated bilateral settlement payments are included in the comprehensive P2P trading cost C total , i P 2 P and are therefore already reflected in the final cooperative operating cost of each DIEM. No additional ex-post lump-sum side-payment variable is introduced; therefore, separate side-payment values are not reported.
According to the payoff definitions u i = C i DIEM and d i = C i DIEM , non , the individual rationality condition u i d i is equivalent to C i DIEM C i DIEM , non . As shown in Table 8, the final cooperative operating cost of each DIEM is lower than its corresponding independently optimized non-cooperative operating cost, and u i d i = C i DIEM , non C i DIEM > 0 holds for all three DIEMs. Therefore, the numerical results directly verify the individual rationality condition of the Nash bargaining framework. Although equal bargaining weights are adopted, the realized cost savings differ because the three DIEMs have different operating characteristics and feasible P2P trading strategies.
Beyond the verified cooperative cost savings and individual rationality results, the proposed framework also has practical implications for industrial microgrid clusters. First, the green-carbon offsetting mechanism provides an additional value-conversion pathway between renewable energy consumption and carbon compliance, thereby reducing the economic pressure associated with carbon trading. Second, P2P cooperative trading enables local surplus–deficit complementarity among DIEMs, thereby reducing dependence on the external grid during peak periods. Third, the tiered carbon trading mechanism provides a flexible market signal to guide low-carbon operations. These findings indicate that the proposed framework can support coordinated economic and low-carbon operation in regional multi-microgrid systems.
In the sensitivity analysis, the baseline operating point is set according to the parameters in Table 2. The carbon quota coefficient and green certificate quota coefficient are varied around their baseline values, while the remaining parameters are kept unchanged. The resulting carbon price and green certificate price are expressed in CNY/kg and CNY/kWh, respectively. To investigate the effects of carbon quota constraint intensity and renewable energy consumption obligations on the system’s environmental and economic performance, a sensitivity analysis is conducted on the carbon quota coefficient and the green certificate quota coefficient, as shown in Figure 10. As shown in Figure 10a, the carbon price increases as the carbon quota coefficient decreases, and decreases as the green certificate quota coefficient increases. This is because when carbon quotas are tightened, high-emission DIEMs must purchase more carbon quotas, whereas the total volume of sellable carbon quotas from low-emission DIEMs decreases accordingly; this creates a market pattern of increased demand and shrinking supply for carbon emission rights, thereby driving up the carbon price. When the carbon quota coefficient remains constant, an increase in the green certificate quota coefficient raises the required green certificate purchase volume for the coal-fired generators within each DIEM, elevating generation costs and reducing their power outputs; this leads to a decline in the overall demand for carbon emission rights, ultimately causing the carbon market price to drop. As shown in Figure 10b, the green certificate price rises as the green certificate quota coefficient increases, and falls as the carbon quota coefficient decreases. This is because an increase in the green certificate quota coefficient significantly boosts the green certificate demand from coal-fired generators in each DIEM, while the total supply of green certificates available from renewable energy generators decreases; the tightening supply–demand balance directly pushes up the market price. Conversely, tightening the carbon quota coefficient suppresses the equilibrium output of coal-fired units in each DIEM while increasing renewable energy output. This increases the supply of green certificates and reduces demand, ultimately leading to a decline in the green certificate market price.

6. Conclusions

This paper proposes a CTDE-MASAC-based cooperative-game low-carbon optimal operation method for multi-microgrids. The proposed method integrates energy-carbon-green certificate P2P coordinated trading, green-carbon offsetting-based dual incentives, and Nash bargaining-based cooperative surplus allocation. A centralized training and decentralized execution framework is further developed to address cooperative games in continuous action spaces. The main conclusions are as follows:
(1)
The proposed CTDE-MASAC framework achieves near-optimal and stable decision-making performance in multi-microgrid cooperative operation. Compared with the centralized MILP benchmark, the proposed MASAC method achieves an average daily operating cost of 67,673.24 CNY with an optimality gap of only 1.25%, while reducing the computation time to 0.12 s. Compared with MADDPG and MAPPO, MASAC achieves lower operating costs and a smaller standard deviation across independent random seeds, indicating better training stability. In addition, the out-of-sample Monte Carlo tests show that the trained policies maintain supply–demand balance without load shedding, and the daily operating cost fluctuates within ±3.2%. These results verify the adaptability and robustness of the proposed method under source-load uncertainty.
(2)
The green certificate-carbon trading dual-incentive mechanism effectively strengthens the coupling between renewable energy consumption and carbon emission reduction. Compared with the scenario that considers only tiered carbon trading, the proposed green-carbon offsetting mechanism further reduces the carbon trading costs for DIEM 1, DIEM 2, and DIEM 3 by 37.27%, 40.13%, and 33.82%, respectively. Meanwhile, the total operating costs decreased by 0.53%, 0.68%, and 0.46%, respectively. These results indicate that the proposed mechanism can convert the value of green certificates into carbon-offset benefits, reduce the economic pressure associated with carbon compliance, and promote the coordinated optimization of economic and environmental benefits.
(3)
The Nash bargaining-based P2P cooperative trading model improves coalition efficiency and individual rationality. Under the energy-carbon-green certificate coordinated trading mechanism, electricity, heat, hydrogen, carbon quotas, and green certificates are jointly traded among DIEMs, enabling complementarity between surpluses and deficits across multiple energy carriers and market resources. Under Nash bargaining-based cooperative operation, the operating cost of each DIEM is lower than its corresponding non-cooperative operating cost, with reductions of 7.08%, 8.56%, and 9.08%, respectively. The total coalition operating cost decreases from 73,708.22 CNY to 67,673.24 CNY, yielding a cooperative cost saving of 6034.98 CNY and an overall reduction of 8.19%. These results numerically verify that all three DIEMs satisfy the individual rationality condition and demonstrate the economic effectiveness of the proposed cooperative game mechanism.
Although the proposed framework demonstrates good economic, low-carbon, and algorithmic performance, this study is still limited by the three-DIEM case setting, homogeneous-agent structure, and equal bargaining weights. Future work will extend the framework to larger-scale heterogeneous multi-microgrid systems and further consider communication delays, privacy protection, and real-time market clearing in practical applications.

Author Contributions

Conceptualization, P.Z. and D.H.; Methodology, P.L. and D.H.; Software, P.L.; Validation, P.Z., P.L. and L.J.; Writing—original draft, P.L.; Writing—review and editing, P.Z., L.J. and D.H.; Supervision, D.H. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

Author Pengfei Zhang was employed by the company State Grid Shanghai Municipal Electric Power Company. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AbbreviationFull name
B-e-D i Electricity bought from DIEM i , i = 1, 2, 3
B-H2-D i Hydrogen bought from DIEM i , i = 1, 2, 3
B-h-D i Heat bought from DIEM i , i = 1, 2, 3
CCUSCarbon capture, utilization, and storage
CETCarbon emission trading
CH4Methane
CHPCombined heat and power
CNYChinese yuan
CO2Carbon dioxide
CTDECentralized training and decentralized execution
CUDACompute Unified Device Architecture
DDPGDeep deterministic policy gradient
DIEMDistributed integrated energy microgrid
D i DIEM i , i = 1, 2, 3
eElectricity or electrical energy, used in figure legends
e-ESS-chaElectrical energy storage charging
e-ESS-disElectrical energy storage discharging
e-loadElectrical load
ELElectrolyzer
ESSEnergy storage system
GBGas boiler
GCTGreen certificate trading
GPUGraphics processing unit
Grid-buyElectricity bought from the main grid
Grid-sellElectricity sold to the main grid
H2Hydrogen
H2-ESS-chaHydrogen energy storage charging
H2-ESS-disHydrogen energy storage discharging
H2-LBLower bound of the hydrogen trading price
H2-UBUpper bound of the hydrogen trading price
H2-loadHydrogen load
HFCHydrogen fuel cell
KKTKarush–Kuhn–Tucker
LBLower bound
MADDPGMulti-agent deep deterministic policy gradient
MAPPOMulti-agent proximal policy optimization
MASACMulti-agent soft actor-critic
MDPMarkov decision process
MRMethane reactor
O&MOperation and maintenance
P2GPower-to-gas
P2PPeer-to-peer trading
PVPhotovoltaics
S-e-D i Electricity sold to DIEM i , i = 1, 2, 3
S-H2-D i Hydrogen sold to DIEM i , i = 1, 2, 3
S-h-D i Heat sold to DIEM i , i = 1, 2, 3
SACSoft actor-critic
SOCState of charge
hHeat
h-ESS-chaHeat storage charging
h-ESS-disHeat storage discharging
h-LBLower bound of the heat trading price
h-UBUpper bound of the heat trading price
h-loadHeat load
TOUTime-of-use
UBUpper bound
WTWind power

Appendix A

Explanation of symbols and parameters used in the mathematical models.
Table A1. Description of mathematical symbols and parameters.
Table A1. Description of mathematical symbols and parameters.
SymbolsMeaning
E 0 , E Total carbon quota and actual carbon emission of a DIEM, respectively
E t , E t in , E green Carbon trading volume, carbon trading volume after green-carbon offsetting, and green-carbon offset volume, respectively
C C , C C i n Carbon trading cost and carbon trading cost after green-carbon offsetting, respectively
χ e , χ h Carbon quota coefficients for electricity generated by coal-fired units and natural gas consumed by gas-fired units, respectively
l , α , ε , ν Interval length, price coefficient, price compensation coefficient, and price growth coefficient in the tiered carbon trading model, respectively
G r , G e , G GCT Green certificates obtained from renewable energy generation, green certificate quota requirement, and green certificate trading volume, respectively
G green , G CET Green certificate volume used for green-carbon offsetting and tiered green certificate quota reward-penalty volume, respectively
λ t GCT , λ 0 GCT Green certificate price at time slot t and benchmark green certificate price, respectively
k 1 , k 2 , ζ Inverse price function parameters and price ratio coefficient in the GCT market
θ , η green , ω green Renewable energy emission reduction efficiency, renewable energy output share, and ratio of actual carbon emissions to carbon quota, respectively
κ Green-carbon correction coefficient
γ Discount factor in the reinforcement learning model
τ Target-network soft-update coefficient in the MASAC algorithm
ρ , ξ Weighting coefficients in the green-carbon offsetting model
ω 1 , ω 2 , ω 3 Green certificate quota reward-penalty coefficients
P t P 2 P , i j , e , P t P 2 P , i j , h , P t P 2 P , i j , H 2 P2P electricity, heat, and hydrogen trading volumes between DIEM i and DIEM j , respectively
E t P 2 P , i , G t P 2 P , i Total P2P carbon quota and green certificate trading volumes of DIEM i , respectively
λ t P 2 P , i j , e , λ t P 2 P , i j , h , λ t P 2 P , i j , H 2 λ t P 2 P , i j , C , λ t P 2 P , i j , G P2P trading prices of electricity, heat, hydrogen, carbon quotas, and green certificates between DIEM i and DIEM j , respectively
λ t i , e Marginal electricity value of DIEM i at time slot t extracted from the critic network
C i DIEM , C total DIEM Total operating cost of DIEM i and total operating cost of the multi-microgrid coalition, respectively
C utility , i , C units , i , C cut , i , C ESS , i External interaction cost, unit operation and maintenance cost, wind and solar curtailment cost, and energy storage operation and maintenance cost of DIEM i , respectively
C C , i , C GCT , i , C total , i P 2 P Carbon trading cost, green certificate trading cost, and comprehensive P2P trading cost of DIEM i , respectively
C i P 2 P , energy , C i P 2 P , C , C i P 2 P , G P2P energy trading cost, P2P carbon quota trading cost, and P2P green certificate trading cost of DIEM i respectively
λ t CH 4 , N t CH 4 , i Natural gas purchase price and natural gas purchase volume of DIEM i at time slot t , respectively
λ t Gb , λ t Gs Time-of-use electricity price and feed-in tariff, respectively
P t sell , i , P t P 2 G , i , P t HFC , H 2 , i Electricity sold to the main grid, P2G electrical input power, and hydrogen consumed by HFC in DIEM i , respectively
α 1 α 6 , β 1 β 2 Operating cost coefficients of energy conversion devices
λ cur , P t WT , cur , i , P t PV , cur , i Wind and solar curtailment cost coefficient, wind curtailment volume, and PV curtailment volume of DIEM i , respectively
ς Degradation cost per unit of charged or discharged energy
P t ESS , e , cha , i , P t ESS , h , cha , i , P t ESS , H 2 , cha , i Charging powers of electrical, heat, and hydrogen energy storage systems in DIEM i
P t ESS , e , dis , i , P t ESS , h , dis , i , P t ESS , H 2 , dis , i Discharging powers of electrical, heat, and hydrogen energy storage systems in DIEM i
G ^ i GCT Total green certificate trading volume of DIEM i
P ^ t PV , i , P ^ t WT , i , δ PV , δ WT Forecast PV and wind outputs of DIEM i and their maximum allowable fluctuation coefficients, respectively
P t EL , i , L t h , i , L t H 2 , i Electrical input power of the electrolyzer, heat load, and hydrogen load of DIEM i , respectively
H t HFC , h , i , P t P 2 G , H 2 , i , P t MR , H 2 , i Heat output of HFC, hydrogen output of P2G, and hydrogen output of the methane reactor in DIEM i , respectively
N max CH 4 , i , N min CH 4 , i , N t MR , CH 4 , i , N t CHP , i , N t GB , i Upper and lower limits of natural gas purchase, natural gas generated by MR, and natural gas consumption of CHP and GB in DIEM i , respectively
E t CHP , C , i , E t GB , C , i , E t CCUS , C , i , E t CCUS , s , i , E t P 2 G , C , i , E t out , i Carbon emissions from CHP and GB, CO2 captured by CCUS, CO2 stored by CCUS, CO2 used by P2G, and carbon emissions released into the atmosphere in DIEM i , respectively
u i , d i , w i Payoff, disagreement point, and bargaining power weight of DIEM i , respectively
u i * Optimal operational payoff of DIEM i obtained in Subproblem 1
P t P 2 P , i j , e , * , P t P 2 P , i j , h , * , P t P 2 P , i j , H 2 , * , E t P 2 P , i j , * , G t P 2 P , i j , * Optimal P2P electricity, heat, hydrogen, carbon quota, and green certificate trading quantities between DIEM i and DIEM j , respectively
S , A , P , R State space, action space, transition function, and reward function in the Markov decision process
s t i , a t i Local state and action of agent i at time slot t , respectively
R t i , r t Immediate reward of agent i and global reward of all agents at time slot t , respectively
π i Policy of agent i
Q ( s t , a t ) , V ( s t ) Soft action-value function and soft state-value function, respectively
θ j , ϕ i Parameter sets of the j -th critic network and the actor network of agent i , respectively
D Experience replay buffer
H ( ) Policy entropy
M Training hyperparameter listed in Table 3

References

  1. Li, H.; Zhu, J.; Dong, H. Two-stage distributionally robust optimization scheduling for multi-energy Microgrid Considering Covariate Factors. Proc. CSEE 2025, 45, 822–834. [Google Scholar]
  2. Zhong, X.; Zhong, W.; Liu, Y.; Yang, C.; Xie, S. Optimal energy management for multi-energy multi-microgrid networks considering carbon emission limitations. Energy 2022, 246, 123428. [Google Scholar] [CrossRef]
  3. Karimi, H. Optimal Operation Scheduling of Water-Energy Nexus Multi-Microgrid Systems Integrated with Energy Storage Systems and Renewable Energy. Energy 2026, 345, 140221. [Google Scholar] [CrossRef]
  4. Hussain, J.; Huang, Q.; Li, J.; Zhang, Z.; Hussain, F.; Ahmed, S.A.; Manzoor, K. Optimization of social welfare in P2P community microgrid with efficient decentralized energy management and communication-efficient power trading. J. Energy Storage 2024, 81, 110458. [Google Scholar] [CrossRef]
  5. Nouri, F.; Vahedipour-Dahraie, M.; Shariatinasab, R.; Siano, P. A decision-making framework for multi-microgrids scheduling considering joint P2P energy and reserve trading floor. Sustain. Energy Grids Netw. 2025, 42, 101685. [Google Scholar] [CrossRef]
  6. Mensin, Y.; Ketjoy, N.; Chamsa-Ard, W.; Kaewpanha, M.; Mensin, P. The P2P energy trading using maximized self-consumption priorities strategies for sustainable microgrid community. Energy Rep. 2022, 8, 14289–14303. [Google Scholar] [CrossRef]
  7. Wang, R.; Wen, X.; Wang, X.; Fu, Y.; Zhang, Y. Low carbon optimal operation of integrated energy system based on carbon capture technology, LCA carbon emissions and tiered carbon trading. Appl. Energy 2022, 311, 118664. [Google Scholar] [CrossRef]
  8. Liu, L.; Jiang, K.; Liu, N.; Zhang, Y. Multi-agent energy-carbon sharing mechanism for parks based on Stackelberg game. Proc. CSEE 2024, 44, 2119–2131. [Google Scholar]
  9. Jiang, Q.; Mu, Y.; Jia, H.; Cao, Y.; Wang, Z.; Wei, W.; Hou, K.; Yu, X. A Stackelberg Game-based planning approach for integrated community energy system considering multiple participants. Energy 2022, 258, 124802. [Google Scholar] [CrossRef]
  10. Cai, G.; Jiang, Y.; Huang, N.; Yang, D.; Pan, X.; Shang, W. Large-scale electric vehicles charging and discharging optimization scheduling based on multi-agent two-level game under electricity demand response mechanism. Proc. CSEE 2023, 43, 85–99. [Google Scholar]
  11. Wang, Z.; Hou, H.; Zhao, B.; Zhang, L.; Shi, Y.; Xie, C. Risk-averse stochastic capacity planning and P2P trading collaborative optimization for multi-energy microgrids considering carbon emission limitations: An asymmetric Nash bargaining approach. Appl. Energy 2024, 357, 122505. [Google Scholar] [CrossRef]
  12. Zhang, X.; Wang, P.; Guo, Z.; Zhang, K.; Xiong, P.; Wang, M.; Pan, F.; Li, C. Nash bargaining-based game for transactive energy of multi-microgrids with dynamic carbon emission factor. Int. J. Electr. Power Energy Syst. 2025, 173, 111367. [Google Scholar] [CrossRef]
  13. Gu, X.; Wang, Q.; Hu, Y.; Zhu, Y.; Ge, Z. Distributed low-carbon optimal operation strategy of multi-microgrids integrated energy system based on Nash bargaining. Power Syst. Technol. 2022, 46, 1464–1482. [Google Scholar]
  14. Duan, P.; Zhao, B.; Zhang, X.; Fen, M. A day-ahead optimal operation strategy for integrated energy systems in multi-public buildings based on cooperative game. Energy 2023, 275, 127395. [Google Scholar] [CrossRef]
  15. Zhang, K.; Chen, J.; Qi, X.; Zhang, W.; Wei, M.; Lin, D. Cooperative optimal operation of multi-microgrids and shared energy storage for voltage regulation of distribution networks based on improved Nash bargaining. Int. J. Electr. Power Energy Syst. 2025, 166, 110532. [Google Scholar] [CrossRef]
  16. Wang, Y.; Li, K.; Li, S.; Ma, X.; Zhang, C. A bi-level scheduling strategy for integrated energy systems considering integrated demand response and energy storage co-optimization. J. Energy Storage 2023, 66, 107508. [Google Scholar] [CrossRef]
  17. Sun, P.; Yun, T.; Chen, Z. Multi-objective robust optimization of multi-energy microgrid with waste treatment. Renew. Energy 2021, 178, 1198–1210. [Google Scholar] [CrossRef]
  18. Liu, J.; Chen, J.; Wang, X.; Zeng, J.; Huang, Q. Energy management and optimization of multi-energy grid based on deep reinforcement learning. Power Syst. Technol. 2020, 44, 3794–3803. [Google Scholar]
  19. Fan, H.; Duan, Z.; Chen, Z.; Zhu, S.; Liu, H.; Li, W.; Yang, Y. Two-layer Optimization Scheduling for Off-grid Microgrids Based on Multi-agent Deep Policy Gradient. Electr. Power 2025, 58, 11–20, 32. [Google Scholar]
  20. Du, Y.; Wu, D. Deep reinforcement learning from demonstrations to assist service restoration in islanded microgrids. IEEE Trans. Sustain. Energy 2022, 13, 1062–1072. [Google Scholar] [CrossRef]
  21. Li, Y.; Wang, R.; Yang, Z. Optimal scheduling of isolated microgrids using automated reinforcement learning-based multi-period forecasting. IEEE Trans. Sustain. Energy 2021, 13, 159–169. [Google Scholar] [CrossRef]
  22. Song, D.; Yan, L.; Dai, X.; Zhu, X.; Hagenmeyer, V.; Zhai, J. Low-carbon energy management for networked multi-energy microgrids using multi-agent soft actor-critic algorithm. Sustain. Energy Grids Netw. 2025, 43, 101821. [Google Scholar] [CrossRef]
  23. Hou, H.; Ge, X.; Yan, Y.; Lu, Y.; Zhang, J.; Dong, Z.Y. An integrated energy system “green-carbon” offset mechanism and optimization method with Stackelberg game. Energy 2024, 294, 130617. [Google Scholar] [CrossRef]
  24. Gao, J.; Li, Y.; Wang, B.; Wu, H. Multi-microgrid collaborative optimization scheduling using an improved multi-agent soft actor-critic algorithm. Energies 2023, 16, 3248. [Google Scholar] [CrossRef]
  25. Xie, L.L.; Li, Y.; Fan, P.; Wan, L.; Zhang, K.; Yang, J. Research on load frequency control of multi-microgrids in an isolated system based on the multi-agent soft actor-critic algorithm. IET Renew. Power Gener. 2024, 18, 1230–1246. [Google Scholar] [CrossRef]
  26. Hao, D.; Hu, Z.; Tan, Z.; Li, T.; Wang, Y.; Hu, H.; Deng, Z. Low-carbon economic scheduling of integrated energy system considering bidirectional interaction of green certificate-ladder carbon and carbon capture. Electr. Power Autom. Equip. 2025, 45, 69–77. [Google Scholar]
  27. Zhang, J.; Ren, Z.; Jiang, Y.; Feng, J.; Sun, Y. Committed carbon emission operation region of microgrids: Theory, construction and observation. Trans. China Electrotech. Soc. 2024, 39, 2342–2359. [Google Scholar]
  28. Lei, L.; Wu, N. An optimal scheduling strategy for electricity-thermal synergy and complementarity among multi-microgrid based on cooperative games. Renew. Energy 2024, 237, 121575. [Google Scholar] [CrossRef]
  29. Zhang, H.; Han, D.; Lu, Z.; Yan, Z. Optimized dispatch of mobile energy storage for low-carbon temporal-spatial management based on two-layer multi-agent deep reinforcement learning. Proc. CSEE 2025, 45, 7974–7986. [Google Scholar]
  30. Azar, B.M.; Kazemzadeh, R.; Oskouei, M.Z.; Mohammadi-Ivatloo, B. Smart prosumers management based on multi-agent deep reinforcement learning to participate in decentralized peer-to-peer market. Appl. Energy 2026, 412, 127650. [Google Scholar] [CrossRef]
  31. Wang, Z.; Li, Y.; Wu, F.; Shi, L.; Ding, R.; He, S. Deep reinforcement learning real-time dispatch approach for cascade hydropower with hybrid pumped-storage mitigating photovoltaic uncertainties. Appl. Energy 2026, 408, 127403. [Google Scholar] [CrossRef]
  32. Li, F.; Hou, H.; Ni, T.; Wang, P.; Wang, Y.; Li, Z.; Pouresmaeil, E. A novel deep reinforcement learning framework for optimal scheduling of low-carbon integrated energy systems considering battery degradation. Energy 2025, 341, 139476. [Google Scholar] [CrossRef]
Figure 1. Architecture of the distributed integrated energy microgrid.
Figure 1. Architecture of the distributed integrated energy microgrid.
Energies 19 03683 g001
Figure 2. Framework of the energy-carbon-green certificate P2P coordinated trading mechanism among multi-microgrids.
Figure 2. Framework of the energy-carbon-green certificate P2P coordinated trading mechanism among multi-microgrids.
Energies 19 03683 g002
Figure 3. Dual-incentive mechanism for green certificate-carbon trading.
Figure 3. Dual-incentive mechanism for green certificate-carbon trading.
Energies 19 03683 g003
Figure 4. CTDE-MASAC-based deep reinforcement learning framework for multi-microgrids.
Figure 4. CTDE-MASAC-based deep reinforcement learning framework for multi-microgrids.
Energies 19 03683 g004
Figure 5. Energy load and renewable energy output forecasts for multi-microgrids: (a) DIEM 1; (b) DIEM 2; (c) DIEM 3.
Figure 5. Energy load and renewable energy output forecasts for multi-microgrids: (a) DIEM 1; (b) DIEM 2; (c) DIEM 3.
Energies 19 03683 g005
Figure 6. Training reward curves of one representative random seed for different algorithms in the continuous action space.
Figure 6. Training reward curves of one representative random seed for different algorithms in the continuous action space.
Energies 19 03683 g006
Figure 7. Optimal energy output scheduling results for multi-microgrids: (a) electric power balance; (b) thermal power balance; (c) hydrogen power balance.
Figure 7. Optimal energy output scheduling results for multi-microgrids: (a) electric power balance; (b) thermal power balance; (c) hydrogen power balance.
Energies 19 03683 g007aEnergies 19 03683 g007b
Figure 8. SOC variations in electric energy storage systems in multi-microgrids.
Figure 8. SOC variations in electric energy storage systems in multi-microgrids.
Energies 19 03683 g008
Figure 9. Energy trading power and prices among multi-microgrids: (a) electricity trading power; (b) electricity trading price; (c) thermal energy trading power; (d) thermal energy trading price; (e) hydrogen energy trading power; (f) hydrogen energy trading price.
Figure 9. Energy trading power and prices among multi-microgrids: (a) electricity trading power; (b) electricity trading price; (c) thermal energy trading power; (d) thermal energy trading price; (e) hydrogen energy trading power; (f) hydrogen energy trading price.
Energies 19 03683 g009aEnergies 19 03683 g009b
Figure 10. Sensitivity analysis of quota coefficients: (a) effect of quota coefficients on the carbon price; (b) effect of quota coefficients on the green certificate price.
Figure 10. Sensitivity analysis of quota coefficients: (a) effect of quota coefficients on the carbon price; (b) effect of quota coefficients on the green certificate price.
Energies 19 03683 g010
Table 1. Comparison of related studies on multi-microgrid trading and low-carbon optimization.
Table 1. Comparison of related studies on multi-microgrid trading and low-carbon optimization.
ReferenceMarket TypeGame ModelAlgorithmUncertainty
Handling
Carbon/Green
Certificate
Mechanism
[11]P2P energy trading with carbon constraintsAsymmetric Nash bargainingStochastic programmingScenario-based uncertaintyCarbon constraints only
[12,13,14,15]Multi-microgrid energy tradingNash bargainingMathematical optimizationLimited considerationLimited carbon coupling; no GCT
[20,21]Microgrid energy schedulingNon-cooperative schedulingDDPG/automated RLData-driven forecastingNot considered
[22]Networked multi-energy microgrid managementMulti-agent coordinationMASACSource-load uncertainty learningLow-carbon operation; no GCT
[23]Integrated energy system with green-carbon interactionStackelberg gameMathematical optimizationDeterministic scenarioGreen-carbon offsetting
This studyEnergy-carbon-GCT P2P coordinated tradingNash bargaining cooperative gameCTDE-MASACWT/PV forecast boundsCET-GCT coupling with P2P trading
Table 2. Basic simulation parameters.
Table 2. Basic simulation parameters.
ParametersValueParametersValueParametersValue
λ 0 GCT /(CNY/kWh)0.15 ρ 0.4 λ 0.2
k 1 0.15 ξ 0.3 δ 0.25
k 2 0.25 κ 1.05 k chp 0.035 CNY/kWh
θ /(kg·kWh−1)0.96 α 3 × 10−4 k gb 0.02 CNY/kWh
ω 1 0.2 ζ 0.6 k p 2 g 0.015 CNY/kWh
ω 2 0.25 η ccus 0.9 a , b , c (Coal) 0.8, 0.05, 0.4; (Gas) 0.1, 0.002, 0.03
ω 3 0.3 l 2000 kg
Table 3. Neural network architecture and hyperparameters of the CTDE-MASAC algorithm.
Table 3. Neural network architecture and hyperparameters of the CTDE-MASAC algorithm.
ParametersValueParametersValue
Actor Network ArchitectureMLP: [State_dim, 256, 256, Action_dim]Batch Size4096
Critic Network ArchitectureMLP: [State_dim + Action_dim, 256, 256, 1]Replay Buffer Capacity100,000
Activation FunctionReLU (Hidden layers), Tanh (actor output)Target Entropy dim A
Maximum Training Episodes2000Discount Factor γ 0.995
Actor Learning Rate ( α π )1 × 10−4Soft-update coefficient τ 0.005
Critic Learning Rate ( α Q )1 × 10−3
Table 4. Comparison of operational cost, optimality gap, and computation time.
Table 4. Comparison of operational cost, optimality gap, and computation time.
AlgorithmAverage Daily Operational Cost (CNY)Optimality Gap (%)Computation Time (s)
Centralized MILP66,837.020.00%124.50
Proposed MASAC67,673.24 ± 142.501.25%0.12
MADDPG69,521.15 ± 412.304.02%0.15
MAPPO68,914.80 ± 385.103.11%0.11
Table 5. Out-of-sample testing results across 10 stochastic scenarios.
Table 5. Out-of-sample testing results across 10 stochastic scenarios.
ScenarioSource-Load FluctuationTotal Operational Cost (CNY)Fluctuation vs. Baseline (%)
BaselineDeterministic (0%)67,673.24
Scenario 1Random noise (±15%)67,125.40−0.81%
Scenario 2Random noise (±15%)68,532.18+1.27%
Scenario 3Random noise (±15%)66,210.55−2.16%
Scenario 4Random noise (±15%)69,815.30+3.17%
Scenario 5Random noise (±15%)67,890.12+0.32%
Scenario 6Random noise (±15%)65,545.80−3.14%
Scenario 7Random noise (±15%)68,102.75+0.63%
Scenario 8Random noise (±15%)67,012.30−0.98%
Scenario 9Random noise (±15%)69,105.45+2.12%
Scenario 10Random noise (±15%)66,515.20−1.71%
Table 6. Scenario settings.
Table 6. Scenario settings.
ScenarioCarbon Trading Market ConsideredGreen Certificate Trading Market Considered
Scenario 1NoNo
Scenario 2YesNo
Scenario 3YesYes
Table 7. Economic and environmental performance under different trading scenarios.
Table 7. Economic and environmental performance under different trading scenarios.
ScenarioDIEMCarbon
Emissions (kg)
P2P Carbon Quota Trading Volume (kg)P2P Green Certificate Trading Volume
(Certificates)
Carbon Trading Cost (CNY)Green Certificate Trading Cost (CNY)Total Operating Cost (CNY)
Scenario 1DIEM 18535.47////23,599.18
DIEM 27062.54////21,850.36
DIEM 35290.89////20,120.57
Scenario 2DIEM 16153.64324.54/1030.79/24,632.99
DIEM 25505.78−221.51/930.48/22,775.33
DIEM 33124.68−103.03/526.58/20,645.52
Scenario 3DIEM 15224.16210.59846.32646.57256.6924,502.44
DIEM 24463.99−134.36−351.29557.07212.0822,619.51
DIEM 32830.50−76.23−495.03348.5182.2120,551.29
Table 8. Non-cooperative operating costs and Nash bargaining outcomes under Scenario 3.
Table 8. Non-cooperative operating costs and Nash bargaining outcomes under Scenario 3.
Metric/ParticipantDIEM 1DIEM 2DIEM 3Total Coalition
Non-cooperative Operating Cost (CNY)26,368.0724,736.9222,603.2373,708.22
Final Cooperative Operating Cost (CNY)24,502.4422,619.5120,551.2967,673.24
Cooperative Cost Saving (CNY)1865.632117.412051.946034.98
Final Cost Reduction Rate (%)7.08%8.56%9.08%8.19%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhang, P.; Liu, P.; Jiang, L.; Han, D. A Cooperative Game-Based Low-Carbon Optimal Operation Strategy for Multi-Microgrids Based on Multi-Agent Deep Reinforcement Learning. Energies 2026, 19, 3683. https://doi.org/10.3390/en19153683

AMA Style

Zhang P, Liu P, Jiang L, Han D. A Cooperative Game-Based Low-Carbon Optimal Operation Strategy for Multi-Microgrids Based on Multi-Agent Deep Reinforcement Learning. Energies. 2026; 19(15):3683. https://doi.org/10.3390/en19153683

Chicago/Turabian Style

Zhang, Pengfei, Pan Liu, Li Jiang, and Dong Han. 2026. "A Cooperative Game-Based Low-Carbon Optimal Operation Strategy for Multi-Microgrids Based on Multi-Agent Deep Reinforcement Learning" Energies 19, no. 15: 3683. https://doi.org/10.3390/en19153683

APA Style

Zhang, P., Liu, P., Jiang, L., & Han, D. (2026). A Cooperative Game-Based Low-Carbon Optimal Operation Strategy for Multi-Microgrids Based on Multi-Agent Deep Reinforcement Learning. Energies, 19(15), 3683. https://doi.org/10.3390/en19153683

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop