Next Article in Journal
Dynamic Adaptive Artificial Hummingbird Algorithm-Enhanced Deep Learning Framework for Accurate Transmission Line Temperature Prediction
Next Article in Special Issue
A Comparison of Deep Learning Techniques for Pose Recognition in Up-and-Go Pole Walking Exercises Using Skeleton Images and Feature Data
Previous Article in Journal
An Evolutionary Deep Learning Framework for Accurate Remaining Capacity Prediction in Lithium-Ion Batteries
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Model Checking Using Large Language Models—Evaluation and Future Directions

by
Sotiris Batsakis
1,2,*,†,‡,
Ilias Tachmazidis
2,‡,
Matthew Mantle
2,‡,
Nikolaos Papadakis
1,‡ and
Grigoris Antoniou
3,*
1
Electrical and Computer Engineering Department, Hellenic Mediterranean University, 71004 Heraklion, Greece
2
School of Computing and Engineering, University of Huddersfield, Queensgate, Huddersfield HD1 3DH, UK
3
School of Built Environment, Engineering and Computing, Leeds Beckett University, Leeds LS1 3HE, UK
*
Authors to whom correspondence should be addressed.
Current address: Department of Electrical and Computer Engineering, Hellenic Mediterranean University, 71410 Heraklion, Greece.
These authors contributed equally to this work.
Electronics 2025, 14(2), 401; https://doi.org/10.3390/electronics14020401
Submission received: 14 December 2024 / Revised: 16 January 2025 / Accepted: 17 January 2025 / Published: 20 January 2025
(This article belongs to the Special Issue Advances in Information, Intelligence, Systems and Applications)

Abstract

Large language models (LLMs) such as ChatGPT have risen in prominence recently, leading to the need to analyze their strengths and limitations for various tasks. The objective of this work was to evaluate the performance of large language models for model checking, which is used extensively in various critical tasks such as software and hardware verification. A set of problems were proposed as a benchmark in this work and three LLMs (GPT-4, Claude, and Gemini) were evaluated with respect to their ability to solve these problems. The evaluation was conducted by comparing the responses of the three LLMs with the gold standard provided by model checking tools. The results illustrate the limitations of LLMs in these tasks, identifying directions for future research. Specifically, the best overall performance (ratio of problems solved correctly) was 60%, indicating a high probability of reasoning errors by the LLMs, especially when dealing with more complex scenarios requiring many reasoning steps, and the LLMs typically performed better when generating scripts for solving the problems rather than solving them directly.
Keywords: model checking; large language models; non-monotonic reasoning model checking; large language models; non-monotonic reasoning

Share and Cite

MDPI and ACS Style

Batsakis, S.; Tachmazidis, I.; Mantle, M.; Papadakis, N.; Antoniou, G. Model Checking Using Large Language Models—Evaluation and Future Directions. Electronics 2025, 14, 401. https://doi.org/10.3390/electronics14020401

AMA Style

Batsakis S, Tachmazidis I, Mantle M, Papadakis N, Antoniou G. Model Checking Using Large Language Models—Evaluation and Future Directions. Electronics. 2025; 14(2):401. https://doi.org/10.3390/electronics14020401

Chicago/Turabian Style

Batsakis, Sotiris, Ilias Tachmazidis, Matthew Mantle, Nikolaos Papadakis, and Grigoris Antoniou. 2025. "Model Checking Using Large Language Models—Evaluation and Future Directions" Electronics 14, no. 2: 401. https://doi.org/10.3390/electronics14020401

APA Style

Batsakis, S., Tachmazidis, I., Mantle, M., Papadakis, N., & Antoniou, G. (2025). Model Checking Using Large Language Models—Evaluation and Future Directions. Electronics, 14(2), 401. https://doi.org/10.3390/electronics14020401

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop