Multi-Agent Reinforcement Learning for Sustainable Integration of Heterogeneous Resources in a Double-Sided Auction Market with Power Balance Incentive Mechanism
Abstract
1. Introduction
- (1)
- In direct contrast to prevailing models that use separate, often static, penalties [28], this paper designs a novel market-clearing incentive mechanism to resolve the misalignment between individual profits and system-wide stability. This mechanism is genuinely new in that it dynamically integrates the system-level power balance residual directly into the profit calculation of each agent’s reward function. This creates an intrinsic, market-driven incentive where agents learn that maximizing their financial return is directly contingent on contributing to grid stability and renewable energy consumption. Hence, the profit-seeking behavior of individual agents will cooperate with the collective goal of market equilibrium and sustainability.
- (2)
- Conventional methods [26,27] rely on uniform experience replay, which treats all market outcomes equally and thus learns inefficiently from rare but critical events (e.g., moments of extreme power imbalance or highly profitable trades). This paper introduces a TD-error-based priority experience weighted replay mechanism. While prioritized experience replay is an established technique in single-agent RL [29], its adaptation and application to stabilize training and accelerate convergence in the volatile, high-dimensional, and sparse-reward context of multi-agent bilateral electricity markets represents a novel contribution. Our specific design weights experiences based not only on error but also on their relevance to system-level objectives, enables agents to focus on high-impact, information-rich experiences, thereby significantly enhancing training efficiency and accelerating convergence to robust bidding strategies in a dynamic, high-dimensional market landscape.
| a | b | c | d | e | f | g | |
|---|---|---|---|---|---|---|---|
| Refs. [13,30] | ✓ | ✕ | ✕ | ✕ | ✕ | ✕ | ✕ |
| Refs. [31,32] | ✓ | ✓ | ✓ | ✕ | ✕ | ✕ | ✕ |
| Refs. [21,26] | ✓ | ✕ | ✓ | ✕ | ✓ | ✕ | ✕ |
| Refs. [33,34] | ✓ | ✕ | ✕ | ✓ | ✕ | ✕ | ✕ |
| Ref. [25] | ✓ | ✓ | ✓ | ✕ | ✓ | ✕ | ✕ |
| This paper | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
2. Mathematical Modeling of Heterogeneous Multi-Resource and Bilateral Market Clearing
2.1. Mathematical Modeling of Heterogeneous Multi-Resource
- The generation sides
- 2.
- The load sides
- 3.
- The storage sides
2.2. The Double-Sided Auction Power Market Clearing Mechanism
- (1)
- Let there be N power generators, M load consumers, and Z energy storage systems participating in the market. Each market participant will submit fair bids and quantities based on publicly available market information.
- (2)
- Each power generator is assumed to be fully rational and risk-neutral, meaning they always strive to maximize their own profit.
- (3)
- The generation side, storage side, and load side engage in free bidding, meaning that for each trading period, each participant is required to submit a capacity–price pair.
3. Multi-Agent RL Method
3.1. Action and State Space
3.2. RL Agent Optimization Algorithm
- (1)
- Double Q-networks
- (2)
- Delayed Target Networks
- (3)
- Target policy construction
3.3. Priority Experience Weighted Replay Mechanism
4. The Simulation Cases
4.1. The Convergence Performance
4.2. The Market Clearing Conditions
4.3. The Power Balance Conditions
5. Conclusions
- (1)
- A primary constraint lies in the simplified market composition adopted for this initial investigation. To fully validate the scalability and generalizability of the proposed framework, it is imperative to test its performance within a more extensive and heterogeneous market environment.
- (2)
- The current market model does not incorporate complex physical network constraints, such as transmission line limits and nodal pricing. Future work will, therefore, focus on enhancing the model’s fidelity by integrating these critical elements of power system operation, thereby assessing the method’s robustness under more realistic market conditions.
- (3)
- The inherent non-stationarity of multi-agent learning environments remains a challenge. Advancing the algorithm to improve its convergence stability and sample efficiency in large-scale applications represents a crucial avenue for further algorithmic development.
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviation
| DA | Day-ahead |
| HMR | Heterogeneous Multi-resource |
| MADDPG | Multi-agent Deep Deterministic Policy Gradient |
| MARL | Multi-agent Reinforcement Learning |
| MATD3 | Multi-agent Twin Delayed Deep Deterministic Policy Gradient |
| MTL | Multi-task Learning |
| SH | Small Hydropower |
| RL | Reinforcement Learning |
| TD | Temporal Difference |
| PV | Photovoltaic |
| PWER | Priority Weighted Experience Replay |
| P2P | Peer-to-peer |
| MT | Multi-task |
References
- Ibrahim, M.S.; Dong, W.; Yang, Q. Machine Learning Driven Smart Electric Power Systems: Current Trends and New Perspectives. Appl. Energy 2020, 272, 115237. [Google Scholar] [CrossRef] [Scilit]
- Shu, Y.; Chen, G.; He, J.; Zhang, F. Building a New Electric Power System Based on New Energy Sources. Strateg. Study Chin. Acad. Eng. 2021, 23, 61–69. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Guo, Y.; Sun, H. Competitive Pricing Game of Virtual Power Plants: Models, Strategies, and Equilibria. IEEE Trans. Smart Grid 2022, 13, 4583–4595. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Zhong, H.; Yang, Z.; Lai, X.; Xia, Q.; Kang, C. Incentive Mechanism for Clearing Energy and Reserve Markets in Multi-Area Power Systems. IEEE Trans. Sustain. Energy 2020, 11, 2470–2482. [Google Scholar] [CrossRef] [Scilit]
- Lin, H. Demand Index and Supply Index Based on Principal Component Analysis: Evidence from US Labor Market. Am. J. Econ. Sociol. 2025; early view. [CrossRef] [Scilit]
- Pezzutto, S.; Bottino-Leone, D.; Wilczynski, E.; Fraboni, R. Drivers and Barriers in the Adoption of Green Heating and Cooling Technologies: Policy and Market Implications for Europe. Sustainability 2024, 16, 6921. [Google Scholar] [CrossRef] [Scilit]
- Zhang, L.; Tian, C.; Li, Z.; Yin, S.; Xie, A.; Wang, P.; Ding, Y. The Impact of Participation Ratio and Bidding Strategies on New Energy’s Involvement in Electricity Spot Market Trading under Marketization Trends—An Empirical Analysis Based on Henan Province, China. Energies 2024, 17, 4463. [Google Scholar] [CrossRef] [Scilit]
- Shivaie, M.; Kiani-Moghaddam, M.; Weinsier, P.D. Bilateral Bidding Strategy in Joint Day-Ahead Energy and Reserve Electricity Markets Considering Techno-Economic-Environmental Measures. Energy Environ. 2022, 33, 696–727. [Google Scholar] [CrossRef] [Scilit]
- Javadi, M.; Baghramian, A. Electricity Trading of Multiple Home Microgrids through V2X Based on Game Theory. Sustain. Cities Soc. 2024, 101, 105046. [Google Scholar] [CrossRef] [Scilit]
- Cheng, Q.; Luo, P.; Liu, P.; Li, X.; Ming, B.; Huang, K.; Xu, W.; Gong, Y. Stochastic Short-Term Scheduling of a Wind-Solar-Hydro Complementary System Considering Both the Day-Ahead Market Bidding and Bilateral Contracts Decomposition. Int. J. Electr. Power Energy Syst. 2022, 138, 107904. [Google Scholar] [CrossRef] [Scilit]
- Nazari, M.E.; Ardehali, M.M. Optimal Bidding Strategy for a GENCO in Day-Ahead Energy and Spinning Reserve Markets with Considerations for Coordinated Wind-Pumped Storage-Thermal System and CO2 Emission. Energy Strategy Rev. 2019, 26, 100405. [Google Scholar] [CrossRef] [Scilit]
- Dolatabadi, S.H.H.; Bhuiyan, T.H.; Chen, Y.; Morales, J.L. A Stochastic Game-Theoretic Optimization Approach for Managing Local Electricity Markets with Electric Vehicles and Renewable Sources. Appl. Energy 2024, 368, 123518. [Google Scholar] [CrossRef] [Scilit]
- Wang, P.; Guo, Z.; Zhang, S.; Zhu, L.; Yi, L.; Song, X. Strategic Behaviors of Renewable Energy Generation Companies Participating in the Electricity and Carbon Coupled Markets Based on Non-Cooperative Game Theory. Energy 2024, 312, 133522. [Google Scholar] [CrossRef] [Scilit]
- Algarvio, H.; Lopes, F. Bilateral Contracting and Price-Based Demand Response in Multi-Agent Electricity Markets: A Study on Time-of-Use Tariffs. Energies 2023, 16, 645. [Google Scholar] [CrossRef] [Scilit]
- Ghasemian Koochaksaraei, M.H.; Toroghi Haghighat, A.; Rezvani, M.H. An Efficient Cloud Resource Exchange Model Based on the Double Auction and Evolutionary Game Theory. Clust. Comput. 2024, 27, 2291–2307. [Google Scholar] [CrossRef] [Scilit]
- Sakolkiatkajorn, P.; Chayakulkheeree, K. Bi-Level Optimization Algorithm for Trading Quantity and Surplus Maximization in P2P Electricity Market. ECTI Trans. Electr. Eng. Electron. Commun. 2025, 23, 255889. [Google Scholar]
- Sanz-Martín, L.; Rivas, G.; Clavijo-Buriticá, N.; Herrera, M.; Parra-Domínguez, J. Mapping the Frontier: A Review of Quantum and Evolutionary Game Theory for Complex Decision-Making. Quantum Inf. Process 2025, 24, 291. [Google Scholar] [CrossRef] [Scilit]
- Karaki, A.; Al-Fagih, L. Evolutionary Game Theory as a Catalyst in Smart Grids: From Theoretical Insights to Practical Strategies. IEEE Access 2024, 12, 186926–186940. [Google Scholar] [CrossRef] [Scilit]
- Jiang, Y.; Dong, J.; Huang, H. Optimal Bidding Strategy for the Price-Maker Virtual Power Plant in the Day-Ahead Market Based on Multi-Agent Twin Delayed Deep Deterministic Policy Gradient Algorithm. Energy 2024, 306, 132388. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Qiu, D.; Kok, K.; Paterakis, N.G. Reliability Assessment of Multi-Agent Reinforcement Learning Algorithms for Hybrid Local Electricity Market Simulation. Appl. Energy 2025, 389, 125789. [Google Scholar] [CrossRef] [Scilit]
- Du, Y.; Li, F.; Zandi, H.; Xue, Y. Approximating Nash Equilibrium in Day-Ahead Electricity Market Bidding with Multi-Agent Deep Reinforcement Learning. J. Mod. Power Syst. Clean Energy 2021, 9, 534–544. [Google Scholar] [CrossRef] [Scilit]
- Deng, L.; Li, Z.; Sun, H.; Guo, Q.; Xu, Y.; Chen, R.; Wang, J.; Guo, Y. Generalized Locational Marginal Pricing in a Heat-and-Electricity-Integrated Market. IEEE Trans. Smart Grid 2019, 10, 6414–6425. [Google Scholar] [CrossRef] [Scilit]
- Ye, Y.; Papadaskalopoulos, D.; Yuan, Q.; Tang, Y.; Strbac, G. Multi-Agent Deep Reinforcement Learning for Coordinated Energy Trading and Flexibility Services Provision in Local Electricity Markets. IEEE Trans. Smart Grid 2022, 14, 1541–1554. [Google Scholar] [CrossRef] [Scilit]
- Xu, H.; Hu, Q.; Wu, Q.; Wang, K.; Wu, F.; Wen, J. Deep Multi-Task Multi-Agent Reinforcement Learning Based Joint Bidding and Pricing Strategy of Price-Maker Load Serving Entity. IEEE Trans. Power Syst. 2025, 40, 505–517. [Google Scholar] [CrossRef] [Scilit]
- Yin, B.; Weng, H.; Hu, Y.; Xi, J.; Ding, P.; Liu, J. Multi-Agent Deep Reinforcement Learning for Simulating Centralized Double-Sided Auction Electricity Market. IEEE Trans. Power Syst. 2025, 40, 518–529. [Google Scholar] [CrossRef] [Scilit]
- Qiu, D.; Wang, J.; Wang, J.; Strbac, G. Multi-Agent Reinforcement Learning for Automated Peer-to-Peer Energy Trading in Double-Side Auction Market. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, Virtual, 19–27 August 2021; International Joint Conferences on Artificial Intelligence Organization: Montreal, QC, Canada, 2021; pp. 2913–2920. [Google Scholar]
- Zheng, J.; Liang, Z.-T.; Li, Y.; Li, Z.; Wu, Q.-H. Multi-Agent Reinforcement Learning with Privacy Preservation for Continuous Double Auction-Based P2P Energy Trading. IEEE Trans. Ind. Inform. 2024, 20, 6582–6590. [Google Scholar] [CrossRef] [Scilit]
- Zou, L.; Munir, M.S.; Tun, Y.K.; Kang, S.; Hong, C.S. Intelligent EV Charging for Urban Prosumer Communities: An Auction and Multi-Agent Deep Reinforcement Learning Approach. IEEE Trans. Netw. Serv. Manag. 2022, 19, 4384–4407. [Google Scholar] [CrossRef] [Scilit]
- Gao, J.; Li, X.; Liu, W.; Zhao, J. Prioritized Experience Replay Method Based on Experience Reward. In Proceedings of the 2021 International Conference on Machine Learning and Intelligent Systems Engineering (MLISE), Chongqing, China, 9–11 July 2021; pp. 214–219. [Google Scholar]
- Janjua, J.I.; Sabir, A.; Abbas, T.; Abbas, S.Q.; Saleem, M. Predictive Analytics and Machine Learning for Electricity Consumption Resilience in Wholesale Power Markets. In Proceedings of the 2024 2nd International Conference on Cyber Resilience (ICCR), Dubai, United Arab Emirates, 26–28 February 2024; pp. 1–7. [Google Scholar]
- Wang, Z.; Dong, L.; Shi, M.; Qiao, J.; Jia, H.; Mu, Y.; Pu, T. Market Power Modeling and Restraint of Aggregated Prosumers in Peer-to-Peer Energy Trading: A Game-Theoretic Approach. Appl. Energy 2023, 348, 121550. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Sun, M. Multi-Oligarch Dynamic Game Model for Regional Power Market with Renewable Portfolio Standard Policies. Appl. Math. Model. 2022, 107, 591–620. [Google Scholar] [CrossRef] [Scilit]
- Bansal, R.K.; Chen, Y.; You, P.; Mallada, E. Market Power Mitigation in Two-Stage Electricity Markets with Supply Function and Quantity Bidding. IEEE Trans. Energy Mark. Policy Regul. 2023, 1, 512–522. [Google Scholar] [CrossRef] [Scilit]
- Kenis, M.; Höschle, H.; Bruninx, K. Strategic Bidding of Wind Power Producers in Electricity Markets in Presence of Information Sharing. Energy Econ. 2022, 110, 106036. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Z.; Chan, K.W.; Xia, S.; Bu, S. Optimal Bi-Level Bidding and Dispatching Strategy between Active Distribution Network and Virtual Alliances Using Distributed Robust Multi-Agent Deep Reinforcement Learning. IEEE Trans. Smart Grid 2022, 13, 2833–2843. [Google Scholar] [CrossRef] [Scilit]
- Wolgast, T.; Nieße, A. Approximating Energy Market Clearing and Bidding with Model-Based Reinforcement Learning. IEEE Access 2024, 12, 145106–145117. [Google Scholar] [CrossRef] [Scilit]







| 1 | initial the generation clearing quantity is , load clearing quantity is and storage clearing quantity is , market clearing price is . |
| 2 | Calculate the normal unbalance power as |
| 3 | For the generation sides |
| 4 | If |
| 5 | If : Else |
| 6 | If |
| 7 | If : Else |
| 8 | For the load sides |
| 9 | If |
| 10 | If |
| 11 | If |
| 12 | If Else |
| 13 | For the storage sides |
| 14 | If |
| 15 | If Else |
| 16 | If |
| 17 | If Else |
| 18 | If |
| 19 | If or Else |
| 1 | initial the buffer capacity , batch size , priority radio , weight radio ; |
| 2 | |
| 3 | If buffer not full |
| 4 | Push new experience at position |
| 5 | else |
| 6 | Replace the old experience with low weights |
| 7 | Push new experience |
| 8 | Calculate the priority of with |
| 9 | Calculate the probability for with |
| 10 | randomly Sample the batch under probabilities |
| 11 | Calculate the weights of sampled experience with (For loss update, ) |
| 12 | Calculate the priority of sampled experience with |
| Key Parameters | Unit | Value | |
|---|---|---|---|
| Market clearing | Maximum market price | $ | 1.4 |
| Minimum market price | $ | 0 | |
| Time interval | min | 15 | |
| Load | Power rate of electrical load | kW/min | 10 |
| Power rate of heat load | kW/min | 10 | |
| Power rate of cooling load | kW/min | 10 | |
| Storage | Storage capacity | kWh | 300 |
| Storage maximum changing power | kW | 30 | |
| Storage maximum discharging power | kW | 30 | |
| Storage charging efficiency | / | 0.95 | |
| Storage discharging efficiency | / | 0.95 | |
| Key Parameters | Value | |
|---|---|---|
| Actor | No. of Hidden layers | 2 |
| No. of Neurons | 128 | |
| Activation Function | ReLU | |
| Learning rate | 0.001 | |
| Optimizer | Adam | |
| Delay update | 0.005 | |
| Critic_ | No. of Hidden layers | 2 |
| No. of Neurons | 128 | |
| Activation Function | ReLU | |
| Learning rate | 0.001 | |
| Optimizer | Adam | |
| Delay update | 0.005 | |
| Critic_ | No. of Hidden layers | 2 |
| No. of Neurons | 128 | |
| Activation Function | ReLU | |
| Learning rate | 0.001 | |
| Optimizer | Adam | |
| Delay update | 0.005 | |
| Training process | Replay buffer size | 1 × 1010 |
| Batch size | 64 | |
| Discount factor | 0.99 | |
| Weights factor | 0.4 | |
| Sampling factor | 0.4 | |
| Clip-epsilon | 0.2 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Huang, J.; Yang, M.; Wang, L.; Mei, M.; Ye, J.; Liu, K.; Bo, Y. Multi-Agent Reinforcement Learning for Sustainable Integration of Heterogeneous Resources in a Double-Sided Auction Market with Power Balance Incentive Mechanism. Sustainability 2026, 18, 141. https://doi.org/10.3390/su18010141
Huang J, Yang M, Wang L, Mei M, Ye J, Liu K, Bo Y. Multi-Agent Reinforcement Learning for Sustainable Integration of Heterogeneous Resources in a Double-Sided Auction Market with Power Balance Incentive Mechanism. Sustainability. 2026; 18(1):141. https://doi.org/10.3390/su18010141
Chicago/Turabian StyleHuang, Jian, Ming Yang, Li Wang, Mingxing Mei, Jianfang Ye, Kejia Liu, and Yaolong Bo. 2026. "Multi-Agent Reinforcement Learning for Sustainable Integration of Heterogeneous Resources in a Double-Sided Auction Market with Power Balance Incentive Mechanism" Sustainability 18, no. 1: 141. https://doi.org/10.3390/su18010141
APA StyleHuang, J., Yang, M., Wang, L., Mei, M., Ye, J., Liu, K., & Bo, Y. (2026). Multi-Agent Reinforcement Learning for Sustainable Integration of Heterogeneous Resources in a Double-Sided Auction Market with Power Balance Incentive Mechanism. Sustainability, 18(1), 141. https://doi.org/10.3390/su18010141
