Next Article in Journal
Autonomous Driving Vulnerability Analysis Under Mixed Traffic Conditions in a Simulated Living Laboratory Environment for Sustainable Smart Cities
Next Article in Special Issue
The Green Side of the Machine: Industrial Robots and Corporate Energy Efficiency in China
Previous Article in Journal
Environmental Implications of Reuse: A Case Study of Electrical and Electronic Devices in Slovenia
Previous Article in Special Issue
A Novel Hybrid Framework for Short-Term Carbon Emissions Forecasting in China: Aggregate and Sectoral Perspectives
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Multi-Agent Reinforcement Learning for Sustainable Integration of Heterogeneous Resources in a Double-Sided Auction Market with Power Balance Incentive Mechanism

1
State Grid Zhejiang Electric Power Company Lishui Power Supply Company, Lishui 323000, China
2
College of Electrical Engineering, Zhejiang University, Hangzhou 310027, China
3
Zhejiang Key Laboratory of Electrical Technology and System on Renewable Energy, Hangzhou 310027, China
*
Author to whom correspondence should be addressed.
Sustainability 2026, 18(1), 141; https://doi.org/10.3390/su18010141
Submission received: 23 September 2025 / Revised: 13 December 2025 / Accepted: 18 December 2025 / Published: 22 December 2025

Abstract

Traditional electricity market bidding typically focuses on unilateral structures, where independent energy storage units and flexible loads act merely as price takers. This reduces bidding motivation and weakens the balancing capability of regional power systems, thereby limiting the large-scale utilization of renewable energy. To address these challenges and support sustainable power system operation, this paper proposes a double-sided auction market strategy for heterogeneous multi-resource (HMR) participation based on multi-agent reinforcement learning (MARL). The framework explicitly considers the heterogeneous bidding and quantity reporting behaviors of renewable generation, flexible demand, and energy storage. An improved incentive mechanism is introduced to enhance real-time system power balance, thereby enabling higher renewable energy integration and reducing curtailment. To efficiently solve the market-clearing problem, an improved Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm is employed, along with a temporal-difference (TD) error-based prioritized experience replay mechanism to strengthen exploration. Case studies validate the effectiveness of the proposed approach in guiding heterogeneous resources toward cooperative bidding behaviors, improving market efficiency, and reinforcing the sustainable and resilient operation of future power systems.

1. Introduction

As the global energy transition accelerates, traditional power systems are gradually evolving toward market-oriented, flexible, and low-carbon directions [1,2]. The introduction of electricity markets has shifted electricity production, transmission, and consumption from a planned economy model to a more flexible competitive market model, which not only improves the efficiency of electricity resource allocation but also enhances the operational stability of power systems [3,4]. Early research primarily used supply–demand curves and classical equilibrium models to describe price formation, which, though simple and effective, did not fully consider the market’s dynamics and complexity, particularly the response of demand and the impact of other resources. Therefore, some studies add principal component analysis and other methods to the supply and demand curves to improve the reliability of the analysis [5,6]. With increasing competition in electricity markets, the unilateral bidding mechanism has revealed many shortcomings, especially in addressing the volatility of renewable energy and emerging resources such as energy storage systems, as traditional models struggle to handle the complex interactions between resources [7].
To address these shortcomings, researchers began to introduce game theory models to analyze the competitive strategies between power generation companies and extended these models that include demand response and energy storage behaviors [8,9]. Ref. [10] establishes a stochastic scheduling model to realize the synergetic optimization of day-ahead (DA) market bids. Ref. [11] presents both optimal bidding strategies and operating schedules of the integrated energy system with the consideration of the CO2 emission limits. The above research is mainly based on static game models. Dynamic game models are more suitable for describing long-term decisions and price fluctuations in electricity markets, taking into account feedback effects between participants. Hence, Ref. [12] presents a stochastic dynamic bi-level model for managing market participants. Ref. [13] proposes a bi-level model for non-cooperative game interactions among power generation companies. Bilateral bidding allows both electricity buyers and sellers to trade based on market conditions and their own needs, which better coordinates resource behaviors and optimizes market resource allocation. Hence, Ref. [14] considers bilateral contracting in electricity markets with demand response. It investigates how curtailment and shifting affect the energy quantity and energy cost of consumers that adopt a time-of-use tariff involving three block periods. Virtual computing needs can also become the subject of bilateral transactions. Ref. [15] uses the combination of barter and evolutionary game theory to model the exchange of instances among cloud service providers through an auction market. Peer-to-peer (P2P) energy markets are the other double-sided auction market. Ref. [16] proposes a bi-level optimization algorithm for trading quantity and surplus maximization in P2P energy trading, incorporating a double-sided carbon taxation scheme. The sheer number of participants in peer-to-peer transactions presents computational challenges. Consequently, game theory methods based on evolutionary algorithms have been extensively studied [17,18].
Nevertheless, game theory assumes that market participants behave rationally and have complete physical models, which limits its application in practice. In reality, there is often information asymmetry and decision-making bias. The operation of market participants themselves is difficult to model, and sometimes the reference for quotations can only be evaluated based on historical data [19,20]. In response to these complex market environments, reinforcement learning (RL), an intelligent algorithm capable of self-learning, has gradually become an important tool in electricity market research. Multi-agent reinforcement learning (MARL), an extension of RL, provides a new solution for electricity markets by simulating the interactions and competition between multiple agents in complex environments [21]. In the MARL framework, each participant continually adjusts its strategy to maximize its profits, thereby optimizing the entire electricity market [22,23]. In order to overcome challenges such as credit assignment and coordination, Ref. [24] proposes a parallelizable deep MT-MARL framework by incorporating multi-task learning (MTL) with MARL. Ref. [25] demonstrates that MARL can find superior bidding strategies for the market participants compared to traditional methods. Based on this, Ref. [26] presents the complete DA market using Multi-Agent Deep Deterministic Policy Gradient (MADDPG). Considering the shortcomings of MADDPG in terms of convergence reliability, Ref. [27] presents Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3), which is a state-of-the-art method.
However, current multi-agent reinforcement learning (MARL) approaches for bidding decisions suffer from two critical limitations. First, they often homogenize diverse controllable resources—such as cold, heat, and electricity—failing to leverage their distinct operational characteristics. Second, and more significantly, in bilateral markets with high penetration of renewable energy, existing reward functions are typically myopic and self-interested, focusing solely on individual agent profits. They largely neglect system-level objectives, such as incentivizing the balanced consumption of renewable generation against load demand. While some approaches attempt to address this by applying simplistic, fixed penalty terms for grid imbalances [28] or by relying on a centralized critic to indirectly guide agents [29], these methods often fail to create a strong and direct coupling between an agent’s profit-seeking behavior and the collective system goal. The incentive for an individual agent to deviate for selfish gain often outweighs the weak, externally imposed penalty. This fundamental misalignment between individual agent incentives and collective system goals leads to significant power imbalances in the final market-clearing outcomes, severely limiting the practical application of MARL in modern power markets.
To fill this gap in existing research, this paper proposes an improved incentive MARL method for double-sided market using an improved MATD3. The novelty of our approach compared to closely related MARL-based market strategies is explicitly detailed in our contributions and summarized in Table 1. The main contributions of this paper are as follows:
(1)
In direct contrast to prevailing models that use separate, often static, penalties [28], this paper designs a novel market-clearing incentive mechanism to resolve the misalignment between individual profits and system-wide stability. This mechanism is genuinely new in that it dynamically integrates the system-level power balance residual directly into the profit calculation of each agent’s reward function. This creates an intrinsic, market-driven incentive where agents learn that maximizing their financial return is directly contingent on contributing to grid stability and renewable energy consumption. Hence, the profit-seeking behavior of individual agents will cooperate with the collective goal of market equilibrium and sustainability.
(2)
Conventional methods [26,27] rely on uniform experience replay, which treats all market outcomes equally and thus learns inefficiently from rare but critical events (e.g., moments of extreme power imbalance or highly profitable trades). This paper introduces a TD-error-based priority experience weighted replay mechanism. While prioritized experience replay is an established technique in single-agent RL [29], its adaptation and application to stabilize training and accelerate convergence in the volatile, high-dimensional, and sparse-reward context of multi-agent bilateral electricity markets represents a novel contribution. Our specific design weights experiences based not only on error but also on their relevance to system-level objectives, enables agents to focus on high-impact, information-rich experiences, thereby significantly enhancing training efficiency and accelerating convergence to robust bidding strategies in a dynamic, high-dimensional market landscape.
As detailed in Table 1, our work distinguishes itself from seminal approaches like MADDPG and MATD3 not by replacing the core algorithm, but by fundamentally redesigning the incentive structure and learning process around it. While they provide the architectural foundation, our novel incentive mechanism and prioritized replay scheme address the critical, unresolved issues of goal misalignment and learning inefficiency that are prevalent in their standard implementations for power market applications. This article is organized as follows. Section 2 presents the mathematical modeling of individual DERs and bilateral market clearing. Section 3 details the algorithm process of MATD3. Section 4 validates the proposed method through different cases in comparison to existing methods. Finally, Section 5 concludes the work and outlines future research directions.
Table 1. Comparison between existing models and methods.
Table 1. Comparison between existing models and methods.
abcdefg
Refs. [13,30]
Refs. [31,32]
Refs. [21,26]
Refs. [33,34]
Ref. [25]
This paper
a: electricity market; b: heterogeneous resource; c: double-sided market; d: bidding price and quantity; e: RL decision framework; f: priority weighted experience replay mechanism; g: incentive mechanism.

2. Mathematical Modeling of Heterogeneous Multi-Resource and Bilateral Market Clearing

2.1. Mathematical Modeling of Heterogeneous Multi-Resource

In a bilateral bidding market, HMR can participate, which can be classified into generation, storage, and load sides. The generation side is divided into distributed photovoltaic (PV), wind power, and small hydropower, based on the type of operation. The load side is categorized into electrical load, cooling load, and heating load, depending on the type of demand. The storage side is represented by independent energy storage systems with charge and discharge capabilities, which can decide whether to charge or discharge based on current market conditions.
  • The generation sides
For the generation side, the generation capacity of PV systems is strongly correlated with current environmental parameters, including temperature, irradiance, and wind speed. Since PV power plants are connected to the grid through power electronic devices, the power output can change rapidly between adjacent time steps. Therefore, this paper does not consider the impact of the ramp rate of PV power plants on their participation in bidding. Similar to PV systems, wind power generation is strongly correlated with environmental parameters. Unlike traditional pumped storage hydropower, run-of-river small hydropower (SH) plants generate power based on seasonal rainfall and the cascade relationship, and do not possess energy storage characteristics. Therefore, the bidding quantity of generation sides can be expressed as the following formula at time slot t :
0 P t G P t G , max
C min market C t G C max market
where P t G is the bid quantity of the generation sides, P t G , max is the maximum power of the generation sides, C t G is the bid price of the generation sides, C max market is the maximum market price limitation, and C min market is minimum market price limitation.
2.
The load sides
On the load side, electrical load, cooling load, and heating load all have adjustable capabilities. However, since load-side resources do not solely depend on environmental parameters, they are also influenced by the current demand on the load side. At the same time, conventional heating and cooling loads, such as large air conditioners and boilers, do not have fast response capabilities. Therefore, when participating in bilateral market bidding, the load side must also consider its ramp-up rate. Based on the above, the bidding on the load side can be expressed by the following formula:
P t L , max P t L P t L , max
P t L P t 1 L P r t L , u p Δ t
P t 1 L P t L P r t L , down Δ t
C min market C t L C max market
where P t L is the bid quantity of the load sides, P t L , max is the maximum power of the load sides, C t L is the bid price of the generation sides, P r t L , u p is the ramp-up rate during the rising process, and P r t L , down is the ramp-up rate during the down process.
3.
The storage sides
For independent energy storage systems, they can either provide electricity as a generation side resource or absorb electricity as a load side resource. At the same time, due to the energy storage characteristics of storage plants, their participation in the electricity market must consider the energy consumption of the system during the power generation process. Therefore, the modeling of the storage side is as follows:
P t c h , max P t c h 0
0 P t dis P t dis , max
S O C t = S O C t 1 + ( P t c h η c h P t dis η dis ) Δ t
S O C min S O C t S O C max
P t c h P t dis = 0
P t S = P t c h + P t dis
where P t S is the bid quantity of the storage sides, P t ch , max and P t dis , max represent the maximum charge and discharge power, η c h and η dis represent the operation efficiency, S O C t is the state-of-charge for the storage sides, and S O C max and S O C min represent the maximum and minimum capacity.

2.2. The Double-Sided Auction Power Market Clearing Mechanism

In the bilateral bidding market, spot trading is the main trading method in the electricity market. To achieve effective clearing for all participants, this paper is set against the backdrop of a centralized clearing electricity spot market with the following assumptions:
(1)
Let there be N power generators, M load consumers, and Z energy storage systems participating in the market. Each market participant will submit fair bids and quantities based on publicly available market information.
(2)
Each power generator is assumed to be fully rational and risk-neutral, meaning they always strive to maximize their own profit.
(3)
The generation side, storage side, and load side engage in free bidding, meaning that for each trading period, each participant is required to submit a capacity–price pair.
Based on the above assumptions, the dispatching agency collects and calculates the capacity and corresponding bids. The clearing process is then performed according to the principle of “load with higher bids cleared first and generation with lower bids cleared first.” Compared to unilateral bidding markets, bilateral bidding markets not only need to ensure that each participant’s profit is maximized when participating in the electricity market but also require the balance of electricity after the clearing process. Therefore, this paper proposes an incentive method for each participant in the bilateral bidding clearing process in Table 2. The core of the incentive mechanism is the secondary adjustment of the profits of each market participant. Participants who contribute to the energy balance by compensating for the shortfall in the planned curve will receive additional rewards. Conversely, participants who are aware of the planned curve shortfall but fail to make targeted adjustments to their output will be penalized. These incentives will serve as the reward values for the reinforcement learning strategy updates of the corresponding participants. In the next chapter, this paper will focus on how to efficiently update the reinforcement learning agent strategies based on the aforementioned reward values.

3. Multi-Agent RL Method

In the previous section, the mathematical modeling of HMRs and the bilateral market clearing mechanism are presented. Traditional multi-agent bidding and pricing models in bilateral markets are typically based on game theory to establish a multi-agent bilateral bidding market clearing model for solution. However, these models overlook the dynamic characteristics of participants during actual operations. In reality, when the electricity market operates, participants will extensively learn from the behaviors of other participants and adjust their strategies accordingly. Traditional convex optimization-based modeling and solution methods are not well-suited to dynamically changing environments. To address this issue, this paper proposes a multi-agent reinforcement learning approach. In the following sections, we will introduce the action–state design and strategy parameter updates. The whole method construction and training process in actual operation is shown in Figure 1.

3.1. Action and State Space

The construction of the RL agent is specifically intended to determine the optimal bidding price and quantity for each entity in the market, which are evidently treated as the actions within the RL framework. The primary objective of bilateral power markets is to meet the power balance. Hence, the actions should align with the current power gap in the power system. Furthermore, states provide information for an agent to make decisions and evaluate the long-term benefits. Therefore, the state space should contain the data that are valuable for deciding on the next action. Specifically, for multi-agent reinforcement learning, each agent needs to learn the decisions of other agents. Therefore, the actions from the previous period and the market clearing results must be used as inputs for strategy learning.
The traditional multi-agent RL method employs a centralized training and decentralized execution approach, where during training, all agents can observe the actions and states of other agents, while during execution, each agent relies solely on its own information. However, in the context of a bilateral auction market, each participating agent can only access a subset of publicly available market information and cannot observe the states and actions of all other agents. Therefore, the proposed method in this paper introduces a synchronized training and decentralized execution mode based on partial market environment information. This modification allows each agent to observe publicly available market information during both training and execution. The consistency in state input for both training and execution helps accelerate convergence and enhances the stability of multi-agent training. Based on the analysis, we design the state space for each agent to include the previous period’s power balance gap, bid prices, bid quantities, awarded quantities, and clearing prices, as well as the power balance gap demand for the current period. Specifically, for the load side and generation side, the awarded quantities of each participant in the previous stage will be published as public market information, serving as a reference for other participants. The agent actions and states in time slot t are as follows:
A t G = [ P t G , C t G ] , G = P V , W T , and   S H
A t L = [ P t L , C t L ] , L = E L , H L , and   C L
A t s = [ P t s , C t s ] , S = E S
P t 1 G , c = [ P t 1 P V , c , P t 1 W T , c , P t 1 S H , c ]
P t 1 L , c = [ P t 1 E L , c , P t 1 H L , c , P t 1 C L , c ]
P t 1 s , c = [ P t 1 E S , c , P t 1 E S , c , P t 1 E S , c ]
S t G = [ P t 1 n , P t 1 G , c , C t 1 G , , C t 1 M , P t n ] , G = P V , W T , and   S H
S t L = [ P t 1 n , P t 1 L , C t 1 L , P t 1 L , c , C t 1 M , P t n ] , L = E L , H L , and   C L
S t s = [ P t 1 n , P t 1 s , C t 1 s , P t 1 s , c , C t 1 M , P t n ] , S = E S
where A t G and S t G are actions and states of the generation side, A t L and S t L are actions and states of the load side, and A t s and S t s are actions and states of the storage side.

3.2. RL Agent Optimization Algorithm

Based on the state space, an RL agent will decide the actions, and the environment will give back the final reward under the actions. Among the existing methods, the MADDPG algorithm is one of the most advanced algorithms [35,36]. However, MADDPG employs a centralized training and decentralized execution approach, which neglects the heterogeneity of individual agents and results in poor training stability in non-static and highly competitive multi-agent environments. To address these issues, this paper introduces an improved MATD3 algorithm. The core improvements of the proposed algorithm compared to MADDPG are as follows:
(1)
Double Q-networks
In MADDPG, the Q-value for a given state–action pair is updated based on the estimated value from the target Q-network. In the proposed algorithm, the double Q-learning is constructed to reduce overestimation bias by maintaining two separate Q-networks for each agent. The Q-value update rule in MATD3 is given by
y i = R t G + γ   min A t G >{ Q 1 ( S t + 1 G , A t + 1 G ) , Q 2 ( S t + 1 G , A t + 1 G ) }
where y i represents the target value for the Q-update, the expected future return of the agent based on its current state–action pair, γ is the discount factor for balancing immediate rewards with future rewards, and Q 1 and Q 2 are the double Q-networks.
(2)
Delayed Target Networks
The proposed algorithm improves the stability of the learning process by introducing delayed target networks, a technique borrowed from TD3. The target networks for both the critic and actor are updated less frequently than the main networks, preventing rapid changes in the value estimates. The target network update rule for the critic is
θ critic = τ θ critic + ( 1 τ ) θ critic
where θ critic is the parameter of critic target network, τ is the update rate for the target network (a small value like 0.005), and θ critic is the current critic network parameter.
(3)
Target policy construction
Like the design for the target network, the proposed method employs the use of target policies for both the critic and actor networks to prevent instability during training. These target policies are updated periodically, ensuring that the learning process remains smooth. The target policy update for the actor is
μ ( S t G ) = τ μ ( S t G ) + ( 1 τ ) μ ( S t G )
where μ ( S t G ) is the target policy for the agent of the generation side and μ ( S t G ) is the current policy, which is updated based on the gradient of the critic target network.

3.3. Priority Experience Weighted Replay Mechanism

The RL algorithm requires many samples to learn effective policies, which can increase training time and computational resource demands. In order to improve the exploring efficiency of the agents, this paper establishes a priority weighted experience replay (PWER) mechanism for the above optimization algorithm. This mechanism consists of two critical components. Firstly, a higher advantage function value indicates a greater potential for exploration associated with the corresponding exploring action. Hence, the samples generated are assigned a selection probability according to the above value. The larger the probability, the easier it is to be selected as an update sample; the smaller the probability, the less likely it is to be selected. Secondly, the trained samples are assigned the lower weight to reduce the effect on the training results, thereby further improving the efficiency of screening environmental samples.
It is important to note that the framework proposed in this paper is based on multi-agent reinforcement learning. Therefore, each agent maintains its own experience buffer, and the updates to its policy are solely based on the samples selected from its own experience pool. The complete pseudo code of the priority replay experience mechanism taking the generation agent as an example is shown in Table 3.

4. The Simulation Cases

To further demonstrate the advantages of the proposed method, this paper presents corresponding case studies for comparison. The bilateral bidding market considered in this study is the day-ahead market, where the participants are consistent with those shown in Figure 2. The generation side includes PV, wind, and SH, while the consumption side consists of electricity load, cooling load, and heating load. Energy storage is represented by an independent lithium-ion battery storage station. In the multi-agent framework, the PV is agent 1, the wind is agent 2, the SH is agent 3, the electrical load is agent 4, the cooling load is agent 5, the heating load is agent 6, and the energy storage is agent 7. Each type of resource participates in the market with price and energy volume point per 15 min under the proposed mechanism, totaling 96 points. The planned generation and consumption curves for agent 1 to agent 6 are shown in Figure 3. The operating parameters for the power market, the load, and the energy storage station are listed in Table 4. The neural network architectures and hyper-parameters for the MATD3 algorithm are presented in Table 5. The algorithm benchmarks used in this paper are MADDPG and the original MATD3, and the parameters of their action and evaluation networks are the same as those in Table 5. All experiments are run on an industrial-grade server equipped with Intel(R) Xeon(R) Gold 6330 CPU @ 2.00 GHz, 4*Nvidia GeForce RTX 4090 GPU and 256 GB RAM.

4.1. The Convergence Performance

RL-based decision-making requires an analysis of convergence performance. Therefore, this study evaluates the convergence performance of multiple algorithms: (1) MADDPG, (2) original MATD3, and (3) the proposed: MATD3 with PWER. In Figure 4, there is a significant gap in the convergence performance of different algorithms. The algorithm proposed in this paper demonstrates superior performance in both convergence speed and average reward value. The superior convergence performance stems from two key algorithmic enhancements. First, the incorporation of PWER significantly improves sample efficiency by focusing learning on transitions with high TD error. This is particularly beneficial in the sparse-reward environment of electricity markets, where critical states—such as those involving volatile renewable generation or strategic bidding—are rare but highly informative. As a result, the proposed method achieves faster initial convergence, as evidenced by the steeper ascent of the learning curve in early training stages. Secondly, the inherent stability of the MATD3 framework, reinforced by the twin-critic design, reduces value overestimation bias and leads to smoother policy updates. This is reflected in the lower reward variance of the proposed method after convergence, compared to the oscillatory behavior of MADDPG. Such stability not only aids training but also supports better generalization during testing, as the learned policy is less sensitive to noisy value estimates. In contrast, while MADDPG attains high rewards quickly, its pronounced fluctuations indicate sensitivity to overestimation bias. The baseline MATD3, lacking prioritized replay, converges more slowly due to less efficient use of past experiences. These observations, supported by ablation studies, confirm that both the PWER component and the underlying MATD3 structure contribute critically to the overall performance improvement.

4.2. The Market Clearing Conditions

Figure 5 illustrates the behaviors of the generation side, load side, and energy storage side when participating in the bilateral bidding process over a 24-h period. The left axis shows the bid price, while the right axis shows the energy quantity. The agent submits its bid price (orange line) and bid quantity (blue line) for each hour. The black line represents the actual cleared quantity of energy that was accepted by the market operator. The figure illustrates a diverse ecosystem of bidding strategies within a simulated electricity market, with each agent embodying a distinct strategic archetype. The generation side agents align its bid quantity with a predictable diurnal generation profile and employs a slightly nuanced pricing strategy, raising its bid price from zero during anticipated high-demand periods to capture better market prices. The load agents are price-inelastic, often bidding at the market price cap to guarantee their demand is met, prioritizing reliability above cost. The storage exemplifies a sophisticated energy arbitrageur and price-maker. It pursues a high-risk, high-reward strategy by bidding extremely high prices just below the market cap. This tactic is designed to capture maximal profit margins by becoming the marginal dispatch unit only during periods of extreme scarcity and peak pricing. Its sporadic cleared quantity is not a sign of failure but the intended outcome of a classic peak-shaving strategy. Collectively, these behaviors showcase a complex interplay between volume-driven renewables, profit-driven dispatchable units, reliability-focused loads, and arbitrage-seeking storage assets.
Figure 6 shows the profits obtained by each agent in the bilateral bidding market. From the perspective of the equilibrium in cooperative games, the profits of the agents are relatively balanced, and no market monopoly has occurred. On the one hand, this indicates that the strategies of the agents are independent of each other, and none of the agents have adopted the strategies of their competitors. On the other hand, this demonstrates the success of the incentive mechanism in the bilateral bidding market, effectively achieving the goal of fostering both cooperation and competitive equilibrium among the agents. The reward curves for both the generation and load sides reveal that, because the generation level does not correlate with load changes, in actual bilateral transactions, there are instances where the generation side continues to supply electricity despite knowing there will be losses. This is because the average marginal cost of small hydropower and wind power generation is close to zero. If generation is stopped and restarted, the restart cost would be extremely high. Therefore, to avoid frequent start-ups and shutdowns, the generation side will try to maintain a stable generation capacity. Load-side users, on the other hand, are price-sensitive and are not easily subject to losses. Therefore, it can be seen that the reward for the load side in the bilateral trading market is essentially non-zero.

4.3. The Power Balance Conditions

Figure 7 illustrates the clearing price curve throughout the bilateral bidding process and the energy balance before and after the bidding. It can be observed that, under the constraints of the market clearing price, the participants in the entire bilateral bidding process have formed a consensus to set their bids slightly below the clearing price upper limit. Although the load side would typically respond to market price changes, such responses are based on the subjective intentions of each participant. Therefore, this study does not consider the price sensitivity of the load side’s energy usage. The values slightly below the clearing price upper limit align with the profit-maximizing behavior of the agents during the market bidding process.
Furthermore, comparing the market clearing price and the energy balance before the bidding, it is evident that, when the market has excess energy, it becomes a buyer’s market. In this scenario, generation-side participants lower their energy price bids to maximize their chances of being awarded power, leading to a decrease in the market clearing price. Conversely, when there is an energy shortage in the market, it becomes a seller’s market. In this case, generation-side participants raise their bids to achieve higher revenues, resulting in an increase in the market clearing price. This negative feedback effect helps prevent abnormal bid prices from market participants, and the incompleteness of public information further stabilizes the market price, avoiding large fluctuations. The proposed method in this study alleviates the power imbalance in the system. The figure also compares the power imbalance under different algorithms. With the method proposed in this paper, the power imbalance decreased by 92.8%. The current mainstream method, MADDPG, reduces the power imbalance by 54%, while the MATD3 method, which does not introduce an incentive mechanism, reduces the imbalance by 74.2%. From the data analysis, it is evident that the incentive mechanism proposed in this paper can effectively guide the bidding behavior of various market participants, showing great application potential.

5. Conclusions

This paper proposes a multi-agent reinforcement learning-based bidding decision method for bilateral markets. First, the common market participants in a bilateral market, including PV, wind turbines, hydropower, electrical load, cooling load, and heating load, are analyzed and modeled based on their specific physical processes. Second, an incentive mechanism aimed at renewable energy absorption and power balance is established and integrated with the market clearing process, thereby guiding and incentivizing each participant in the bilateral market. Next, a multi-agent reinforcement learning algorithm is introduced. This algorithm constructs a dual-critic network to enhance the robustness of state estimation in reinforcement learning. Additionally, a TD-error-based prioritized experience replay mechanism is developed to ensure efficient exploration of reinforcement learning agents within the environment. Finally, through the construction of typical operating scenarios, the superiority of the proposed method is validated. Compared to traditional methods, the proposed approach reduces the system’s power balance deviation by 92.8% and accelerates training time by 36%, demonstrating its significant application value.
Despite the encouraging results, this study has certain limitations that point to valuable directions for future research:
(1)
A primary constraint lies in the simplified market composition adopted for this initial investigation. To fully validate the scalability and generalizability of the proposed framework, it is imperative to test its performance within a more extensive and heterogeneous market environment.
(2)
The current market model does not incorporate complex physical network constraints, such as transmission line limits and nodal pricing. Future work will, therefore, focus on enhancing the model’s fidelity by integrating these critical elements of power system operation, thereby assessing the method’s robustness under more realistic market conditions.
(3)
The inherent non-stationarity of multi-agent learning environments remains a challenge. Advancing the algorithm to improve its convergence stability and sample efficiency in large-scale applications represents a crucial avenue for further algorithmic development.

Author Contributions

Conceptualization, J.H. and Y.B.; methodology, M.Y. and Y.B.; software, L.W., J.H. and K.L.; validation, J.H., M.M. and K.L.; formal analysis, J.Y.; investigation, M.M. and Y.B.; writing—original draft preparation, J.H. and M.Y.; writing—review and editing, J.H., K.L. and Y.B.; supervision, Y.B. All authors have read and agreed to the published version of the manuscript.

Funding

This project is supported by the Science and Technology Project of State Grid Zhejiang Electric Power Co. Ltd. (5211LS240008).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Conflicts of Interest

Authors Jian Huang; Ming Yang; Li Wang; Mingxing Mei; Jianfang Ye were employed by the company State Grid Zhejiang Electric Power Company Lishui Power Supply Company. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest”. The authors declare that this study received funding from State Grid Zhejiang Electric Power Company Lishui Power Supply Company. The funder had the following involvement with the study: State Grid Zhejiang Electric Power Company Lishui Power Supply Company.

Abbreviation

DADay-ahead
HMRHeterogeneous Multi-resource
MADDPGMulti-agent Deep Deterministic Policy Gradient
MARLMulti-agent Reinforcement Learning
MATD3Multi-agent Twin Delayed Deep Deterministic Policy Gradient
MTLMulti-task Learning
SHSmall Hydropower
RLReinforcement Learning
TDTemporal Difference
PVPhotovoltaic
PWERPriority Weighted Experience Replay
P2PPeer-to-peer
MTMulti-task

References

  1. Ibrahim, M.S.; Dong, W.; Yang, Q. Machine Learning Driven Smart Electric Power Systems: Current Trends and New Perspectives. Appl. Energy 2020, 272, 115237. [Google Scholar] [CrossRef] [Scilit]
  2. Shu, Y.; Chen, G.; He, J.; Zhang, F. Building a New Electric Power System Based on New Energy Sources. Strateg. Study Chin. Acad. Eng. 2021, 23, 61–69. [Google Scholar] [CrossRef] [Scilit]
  3. Xu, Z.; Guo, Y.; Sun, H. Competitive Pricing Game of Virtual Power Plants: Models, Strategies, and Equilibria. IEEE Trans. Smart Grid 2022, 13, 4583–4595. [Google Scholar] [CrossRef] [Scilit]
  4. Wang, J.; Zhong, H.; Yang, Z.; Lai, X.; Xia, Q.; Kang, C. Incentive Mechanism for Clearing Energy and Reserve Markets in Multi-Area Power Systems. IEEE Trans. Sustain. Energy 2020, 11, 2470–2482. [Google Scholar] [CrossRef] [Scilit]
  5. Lin, H. Demand Index and Supply Index Based on Principal Component Analysis: Evidence from US Labor Market. Am. J. Econ. Sociol. 2025; early view. [CrossRef] [Scilit]
  6. Pezzutto, S.; Bottino-Leone, D.; Wilczynski, E.; Fraboni, R. Drivers and Barriers in the Adoption of Green Heating and Cooling Technologies: Policy and Market Implications for Europe. Sustainability 2024, 16, 6921. [Google Scholar] [CrossRef] [Scilit]
  7. Zhang, L.; Tian, C.; Li, Z.; Yin, S.; Xie, A.; Wang, P.; Ding, Y. The Impact of Participation Ratio and Bidding Strategies on New Energy’s Involvement in Electricity Spot Market Trading under Marketization Trends—An Empirical Analysis Based on Henan Province, China. Energies 2024, 17, 4463. [Google Scholar] [CrossRef] [Scilit]
  8. Shivaie, M.; Kiani-Moghaddam, M.; Weinsier, P.D. Bilateral Bidding Strategy in Joint Day-Ahead Energy and Reserve Electricity Markets Considering Techno-Economic-Environmental Measures. Energy Environ. 2022, 33, 696–727. [Google Scholar] [CrossRef] [Scilit]
  9. Javadi, M.; Baghramian, A. Electricity Trading of Multiple Home Microgrids through V2X Based on Game Theory. Sustain. Cities Soc. 2024, 101, 105046. [Google Scholar] [CrossRef] [Scilit]
  10. Cheng, Q.; Luo, P.; Liu, P.; Li, X.; Ming, B.; Huang, K.; Xu, W.; Gong, Y. Stochastic Short-Term Scheduling of a Wind-Solar-Hydro Complementary System Considering Both the Day-Ahead Market Bidding and Bilateral Contracts Decomposition. Int. J. Electr. Power Energy Syst. 2022, 138, 107904. [Google Scholar] [CrossRef] [Scilit]
  11. Nazari, M.E.; Ardehali, M.M. Optimal Bidding Strategy for a GENCO in Day-Ahead Energy and Spinning Reserve Markets with Considerations for Coordinated Wind-Pumped Storage-Thermal System and CO2 Emission. Energy Strategy Rev. 2019, 26, 100405. [Google Scholar] [CrossRef] [Scilit]
  12. Dolatabadi, S.H.H.; Bhuiyan, T.H.; Chen, Y.; Morales, J.L. A Stochastic Game-Theoretic Optimization Approach for Managing Local Electricity Markets with Electric Vehicles and Renewable Sources. Appl. Energy 2024, 368, 123518. [Google Scholar] [CrossRef] [Scilit]
  13. Wang, P.; Guo, Z.; Zhang, S.; Zhu, L.; Yi, L.; Song, X. Strategic Behaviors of Renewable Energy Generation Companies Participating in the Electricity and Carbon Coupled Markets Based on Non-Cooperative Game Theory. Energy 2024, 312, 133522. [Google Scholar] [CrossRef] [Scilit]
  14. Algarvio, H.; Lopes, F. Bilateral Contracting and Price-Based Demand Response in Multi-Agent Electricity Markets: A Study on Time-of-Use Tariffs. Energies 2023, 16, 645. [Google Scholar] [CrossRef] [Scilit]
  15. Ghasemian Koochaksaraei, M.H.; Toroghi Haghighat, A.; Rezvani, M.H. An Efficient Cloud Resource Exchange Model Based on the Double Auction and Evolutionary Game Theory. Clust. Comput. 2024, 27, 2291–2307. [Google Scholar] [CrossRef] [Scilit]
  16. Sakolkiatkajorn, P.; Chayakulkheeree, K. Bi-Level Optimization Algorithm for Trading Quantity and Surplus Maximization in P2P Electricity Market. ECTI Trans. Electr. Eng. Electron. Commun. 2025, 23, 255889. [Google Scholar]
  17. Sanz-Martín, L.; Rivas, G.; Clavijo-Buriticá, N.; Herrera, M.; Parra-Domínguez, J. Mapping the Frontier: A Review of Quantum and Evolutionary Game Theory for Complex Decision-Making. Quantum Inf. Process 2025, 24, 291. [Google Scholar] [CrossRef] [Scilit]
  18. Karaki, A.; Al-Fagih, L. Evolutionary Game Theory as a Catalyst in Smart Grids: From Theoretical Insights to Practical Strategies. IEEE Access 2024, 12, 186926–186940. [Google Scholar] [CrossRef] [Scilit]
  19. Jiang, Y.; Dong, J.; Huang, H. Optimal Bidding Strategy for the Price-Maker Virtual Power Plant in the Day-Ahead Market Based on Multi-Agent Twin Delayed Deep Deterministic Policy Gradient Algorithm. Energy 2024, 306, 132388. [Google Scholar] [CrossRef] [Scilit]
  20. Zhang, H.; Qiu, D.; Kok, K.; Paterakis, N.G. Reliability Assessment of Multi-Agent Reinforcement Learning Algorithms for Hybrid Local Electricity Market Simulation. Appl. Energy 2025, 389, 125789. [Google Scholar] [CrossRef] [Scilit]
  21. Du, Y.; Li, F.; Zandi, H.; Xue, Y. Approximating Nash Equilibrium in Day-Ahead Electricity Market Bidding with Multi-Agent Deep Reinforcement Learning. J. Mod. Power Syst. Clean Energy 2021, 9, 534–544. [Google Scholar] [CrossRef] [Scilit]
  22. Deng, L.; Li, Z.; Sun, H.; Guo, Q.; Xu, Y.; Chen, R.; Wang, J.; Guo, Y. Generalized Locational Marginal Pricing in a Heat-and-Electricity-Integrated Market. IEEE Trans. Smart Grid 2019, 10, 6414–6425. [Google Scholar] [CrossRef] [Scilit]
  23. Ye, Y.; Papadaskalopoulos, D.; Yuan, Q.; Tang, Y.; Strbac, G. Multi-Agent Deep Reinforcement Learning for Coordinated Energy Trading and Flexibility Services Provision in Local Electricity Markets. IEEE Trans. Smart Grid 2022, 14, 1541–1554. [Google Scholar] [CrossRef] [Scilit]
  24. Xu, H.; Hu, Q.; Wu, Q.; Wang, K.; Wu, F.; Wen, J. Deep Multi-Task Multi-Agent Reinforcement Learning Based Joint Bidding and Pricing Strategy of Price-Maker Load Serving Entity. IEEE Trans. Power Syst. 2025, 40, 505–517. [Google Scholar] [CrossRef] [Scilit]
  25. Yin, B.; Weng, H.; Hu, Y.; Xi, J.; Ding, P.; Liu, J. Multi-Agent Deep Reinforcement Learning for Simulating Centralized Double-Sided Auction Electricity Market. IEEE Trans. Power Syst. 2025, 40, 518–529. [Google Scholar] [CrossRef] [Scilit]
  26. Qiu, D.; Wang, J.; Wang, J.; Strbac, G. Multi-Agent Reinforcement Learning for Automated Peer-to-Peer Energy Trading in Double-Side Auction Market. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, Virtual, 19–27 August 2021; International Joint Conferences on Artificial Intelligence Organization: Montreal, QC, Canada, 2021; pp. 2913–2920. [Google Scholar]
  27. Zheng, J.; Liang, Z.-T.; Li, Y.; Li, Z.; Wu, Q.-H. Multi-Agent Reinforcement Learning with Privacy Preservation for Continuous Double Auction-Based P2P Energy Trading. IEEE Trans. Ind. Inform. 2024, 20, 6582–6590. [Google Scholar] [CrossRef] [Scilit]
  28. Zou, L.; Munir, M.S.; Tun, Y.K.; Kang, S.; Hong, C.S. Intelligent EV Charging for Urban Prosumer Communities: An Auction and Multi-Agent Deep Reinforcement Learning Approach. IEEE Trans. Netw. Serv. Manag. 2022, 19, 4384–4407. [Google Scholar] [CrossRef] [Scilit]
  29. Gao, J.; Li, X.; Liu, W.; Zhao, J. Prioritized Experience Replay Method Based on Experience Reward. In Proceedings of the 2021 International Conference on Machine Learning and Intelligent Systems Engineering (MLISE), Chongqing, China, 9–11 July 2021; pp. 214–219. [Google Scholar]
  30. Janjua, J.I.; Sabir, A.; Abbas, T.; Abbas, S.Q.; Saleem, M. Predictive Analytics and Machine Learning for Electricity Consumption Resilience in Wholesale Power Markets. In Proceedings of the 2024 2nd International Conference on Cyber Resilience (ICCR), Dubai, United Arab Emirates, 26–28 February 2024; pp. 1–7. [Google Scholar]
  31. Wang, Z.; Dong, L.; Shi, M.; Qiao, J.; Jia, H.; Mu, Y.; Pu, T. Market Power Modeling and Restraint of Aggregated Prosumers in Peer-to-Peer Energy Trading: A Game-Theoretic Approach. Appl. Energy 2023, 348, 121550. [Google Scholar] [CrossRef] [Scilit]
  32. Wu, Y.; Sun, M. Multi-Oligarch Dynamic Game Model for Regional Power Market with Renewable Portfolio Standard Policies. Appl. Math. Model. 2022, 107, 591–620. [Google Scholar] [CrossRef] [Scilit]
  33. Bansal, R.K.; Chen, Y.; You, P.; Mallada, E. Market Power Mitigation in Two-Stage Electricity Markets with Supply Function and Quantity Bidding. IEEE Trans. Energy Mark. Policy Regul. 2023, 1, 512–522. [Google Scholar] [CrossRef] [Scilit]
  34. Kenis, M.; Höschle, H.; Bruninx, K. Strategic Bidding of Wind Power Producers in Electricity Markets in Presence of Information Sharing. Energy Econ. 2022, 110, 106036. [Google Scholar] [CrossRef] [Scilit]
  35. Zhu, Z.; Chan, K.W.; Xia, S.; Bu, S. Optimal Bi-Level Bidding and Dispatching Strategy between Active Distribution Network and Virtual Alliances Using Distributed Robust Multi-Agent Deep Reinforcement Learning. IEEE Trans. Smart Grid 2022, 13, 2833–2843. [Google Scholar] [CrossRef] [Scilit]
  36. Wolgast, T.; Nieße, A. Approximating Energy Market Clearing and Bidding with Model-Based Reinforcement Learning. IEEE Access 2024, 12, 145106–145117. [Google Scholar] [CrossRef] [Scilit]
Figure 1. The MATD3 method construction and training process.
Figure 1. The MATD3 method construction and training process.
Sustainability 18 00141 g001
Figure 2. Heterogeneous multi-resource in bilateral bidding market.
Figure 2. Heterogeneous multi-resource in bilateral bidding market.
Sustainability 18 00141 g002
Figure 3. The planned generation and load curves for agents.
Figure 3. The planned generation and load curves for agents.
Sustainability 18 00141 g003
Figure 4. The convergence comparison for different algorithms.
Figure 4. The convergence comparison for different algorithms.
Sustainability 18 00141 g004
Figure 5. The market behaviors for agents.
Figure 5. The market behaviors for agents.
Sustainability 18 00141 g005
Figure 6. The reward curves for agents.
Figure 6. The reward curves for agents.
Sustainability 18 00141 g006
Figure 7. The power balance in double-sided auction market under different methods.
Figure 7. The power balance in double-sided auction market under different methods.
Sustainability 18 00141 g007
Table 2. The incentive mechanism in the double-sided auction market clearing process.
Table 2. The incentive mechanism in the double-sided auction market clearing process.
1initial the generation clearing quantity is P t G , c , load clearing quantity is P t L , c and storage clearing quantity is P t S , c , market clearing price is C t M .
2Calculate the normal unbalance power as P t n = i = 1 N P t G , n i = 1 M P t L , n
3For the generation sides
4  If P t n > 0
5    If P t G , c < P t G , n : R t G = P t G , c C t M Else P t G , c C t M k G
6  If P t n 0
7    If P t G , c < P t G , n : R t G = P t G , c C t M Else R t G = P t G , c C t M k G
8For the load sides
9  If P t n > 0 R t L = max ( 0 , ( P t L , c P t L , n ) C t M k L )
10  If P t n < 0 R t L = max ( 0 , ( P t L , n P t L , c ) C t M k L )
11  If P t n = 0
12    If P t L , c ! = P t L , n   R t L = 100 Else R t L = 0
13For the storage sides
14  If P t n > 0
15    If P t s > 0 R t s = 0 Else R t s = P t s , c C t M k G
16  If P t n < 0
17    If P t s > 0 R t s = P t s , c C t M k G Else R t s = 0
18  If P t n = 0
19    If P t s > 0 or P t s < 0 R t s = P t s , c C t M k G Else R t s = 0
Table 3. Priority replay experience mechanism for each agent.
Table 3. Priority replay experience mechanism for each agent.
1initial the buffer capacity L m a x , batch size B , priority radio α = 0.4 , weight radio γ = 0.4 ;
2
3If buffer not full
4  Push new experience L e G = ( s t G , A t G , R t G , s t + 1 G ) at position p o = t   %   L m a x
5else
6  Replace the old experience with low weights
7Push new experience L e G = ( s t G , A t G , R t G , s t + 1 G )
8Calculate the priority of L e t with δ t G = y G Q G ( s t G , A t G ; θ c r i t i c )
9Calculate the probability for L e t with L e t , p = δ t G i = 1 L m a x δ i G
10randomly Sample the batch L e = { L e 1 , L e 2 , , L e B } under probabilities L e t , p
11Calculate the weights of sampled experience with L e w , t = ( L m a x L e t , p ) γ max ( L e w ) (For loss update, c 1 = c 1     L e w , t )
12Calculate the priority of sampled experience with δ t G = δ t G α
Table 4. Key parameters in bilateral power market.
Table 4. Key parameters in bilateral power market.
Key ParametersUnitValue
Market clearingMaximum market price$1.4
Minimum market price$0
Time intervalmin15
LoadPower rate of electrical loadkW/min10
Power rate of heat loadkW/min10
Power rate of cooling loadkW/min10
StorageStorage capacitykWh300
Storage maximum changing powerkW30
Storage maximum discharging powerkW30
Storage charging efficiency/0.95
Storage discharging efficiency/0.95
Table 5. The neural network architectures and hyper-parameters.
Table 5. The neural network architectures and hyper-parameters.
Key ParametersValue
ActorNo. of Hidden layers2
No. of Neurons128
Activation FunctionReLU
Learning rate0.001
OptimizerAdam
Delay update 0.005
Critic_ Q 1 No. of Hidden layers2
No. of Neurons128
Activation FunctionReLU
Learning rate0.001
OptimizerAdam
Delay update 0.005
Critic_ Q 2 No. of Hidden layers2
No. of Neurons128
Activation FunctionReLU
Learning rate0.001
OptimizerAdam
Delay update0.005
Training processReplay buffer size1 × 1010
Batch size64
Discount factor0.99
Weights factor0.4
Sampling factor0.4
Clip-epsilon0.2
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Huang, J.; Yang, M.; Wang, L.; Mei, M.; Ye, J.; Liu, K.; Bo, Y. Multi-Agent Reinforcement Learning for Sustainable Integration of Heterogeneous Resources in a Double-Sided Auction Market with Power Balance Incentive Mechanism. Sustainability 2026, 18, 141. https://doi.org/10.3390/su18010141

AMA Style

Huang J, Yang M, Wang L, Mei M, Ye J, Liu K, Bo Y. Multi-Agent Reinforcement Learning for Sustainable Integration of Heterogeneous Resources in a Double-Sided Auction Market with Power Balance Incentive Mechanism. Sustainability. 2026; 18(1):141. https://doi.org/10.3390/su18010141

Chicago/Turabian Style

Huang, Jian, Ming Yang, Li Wang, Mingxing Mei, Jianfang Ye, Kejia Liu, and Yaolong Bo. 2026. "Multi-Agent Reinforcement Learning for Sustainable Integration of Heterogeneous Resources in a Double-Sided Auction Market with Power Balance Incentive Mechanism" Sustainability 18, no. 1: 141. https://doi.org/10.3390/su18010141

APA Style

Huang, J., Yang, M., Wang, L., Mei, M., Ye, J., Liu, K., & Bo, Y. (2026). Multi-Agent Reinforcement Learning for Sustainable Integration of Heterogeneous Resources in a Double-Sided Auction Market with Power Balance Incentive Mechanism. Sustainability, 18(1), 141. https://doi.org/10.3390/su18010141

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop