Next Article in Journal
Research on Long-Term Scheduling Optimization of Water–Wind–Solar Multi-Energy Complementary System Based on DDPG
Next Article in Special Issue
The Economic Optimization of a Grid-Connected Hybrid Renewable System with an Electromagnetic Frequency Regulator Using a Genetic Algorithm
Previous Article in Journal
Polish Farmers′ Perceptions of the Benefits and Risks of Investing in Biogas Plants and the Role of GISs in Site Selection
Previous Article in Special Issue
Coordination of Hydropower Generation and Export Considering River Flow Evolution Process of Cascade Hydropower Systems
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

ELM-Bench: A Multidimensional Methodological Framework for Large Language Model Evaluation in Electricity Markets

1
School of Economics and Management, North China Electric Power University, Beijing 100000, China
2
Beijing Power Exchange Center Co., Ltd., Beijing 100000, China
3
State Grid LiaoNing Electric Power Supply Co., Ltd., Electric Power Research Institute, Shenyang 110000, China
*
Author to whom correspondence should be addressed.
Energies 2025, 18(15), 3982; https://doi.org/10.3390/en18153982
Submission received: 19 June 2025 / Revised: 13 July 2025 / Accepted: 23 July 2025 / Published: 25 July 2025

Abstract

The large language model (LLM) has significant potential for application in the field of electricity markets, but there are shortcomings in professional evaluation methods for LLM: single task, limited dataset coverage, and lack of depth. To this end, this article proposes the ELM-Bench framework for evaluating the LLM of the Chinese electricity market, which evaluates the model from 3 dimensions of understanding, generation, and safety through 7 tasks (such as common-sense Q&A and terminology explanations) with 2841 samples. At the same time, a specialized domain model QwenGOLD was fine-tuned based on the general LLM. The evaluation results show that the top-level general model performs well in general tasks due to high-quality pre-training, while QwenGOLD performs better in tasks such as prediction and decision-making in professional fields, verifying the effectiveness of domain fine-tuning. The study also found that fine-tuning has limited improvement on LLM’s basic abilities, but its score in professional prediction tasks is second only to Deepseek-V3, indicating that some general LLMs can handle domain data well without professional training. This can provide a basis for model selection in different scenarios, balancing performance and training costs.
Keywords: electricity market; large language models; prompt instructions; evaluation framework; model fine-tuning electricity market; large language models; prompt instructions; evaluation framework; model fine-tuning

Share and Cite

MDPI and ACS Style

Fan, H.; Ji, S.; Yuan, P.; Zhao, Q.; Wang, S.; Tan, X.; Duan, Y. ELM-Bench: A Multidimensional Methodological Framework for Large Language Model Evaluation in Electricity Markets. Energies 2025, 18, 3982. https://doi.org/10.3390/en18153982

AMA Style

Fan H, Ji S, Yuan P, Zhao Q, Wang S, Tan X, Duan Y. ELM-Bench: A Multidimensional Methodological Framework for Large Language Model Evaluation in Electricity Markets. Energies. 2025; 18(15):3982. https://doi.org/10.3390/en18153982

Chicago/Turabian Style

Fan, Hang, Shijie Ji, Peng Yuan, Qingsong Zhao, Shuaikang Wang, Xiaowei Tan, and Yunjie Duan. 2025. "ELM-Bench: A Multidimensional Methodological Framework for Large Language Model Evaluation in Electricity Markets" Energies 18, no. 15: 3982. https://doi.org/10.3390/en18153982

APA Style

Fan, H., Ji, S., Yuan, P., Zhao, Q., Wang, S., Tan, X., & Duan, Y. (2025). ELM-Bench: A Multidimensional Methodological Framework for Large Language Model Evaluation in Electricity Markets. Energies, 18(15), 3982. https://doi.org/10.3390/en18153982

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop