Reliability and Performance Stability of Large Language Models in Medical Knowledge Assessment: Evidence from the European Board of Nuclear Medicine Examination
Abstract
1. Introduction
- (1)
- Evaluate the performance of ten state-of-the-art LLMs on EBNM Fellowship Examination questions, including five proprietary and five open-source systems.
- (2)
- Assess inter-run reliability using Cohen’s kappa coefficient across five independent evaluations per model, establishing the first reliability benchmarks for AI in nuclear medicine examinations.
- (3)
- Identify statistically significant performance differences through rigorous pairwise comparisons with Bonferroni correction for multiple testing.
- (4)
- Compare proprietary versus open-source model capabilities to inform accessibility considerations for educational and clinical applications. We hypothesized that leading models would approach human first-time candidate performance levels while demonstrating high inter-run consistency, with significant performance stratification emerging between model tiers.
2. Materials and Methods
2.1. Study Design
2.2. AI Models
2.3. Examination
2.4. Prompt Engineering
- Context: Expert nuclear medicine physician with comprehensive subspecialty knowledge;
- Objective: Answer 50 EBNM examination questions selecting exactly one option per question;
- Style: Clinical expertise with definitive single-best-answer selections;
- Tone: Professional and evidence-based;
- Audience: European subspecialty certification level;
- Response: Strict table format (Question|Answer) with no explanatory text.
2.5. Data Collection
2.6. Statistical Analysis
2.7. Ethical Considerations
3. Results
3.1. Overall Performance
3.2. Top Performers
3.3. Proprietary Versus Open-Source Models
3.4. Inter-Run Reliability
3.5. Pairwise Model Comparisons
3.6. Non-Significant Clinical Differences
4. Discussion
- From an applied sciences perspective, this work should be interpreted as a methodological benchmark of model performance and reproducibility rather than an assessment of clinical competence.
- We evaluated 50 publicly available items rather than the full 140-question EBNM examination, which reduces representativeness and statistical power. Power calculations based on the observed data indicated limited ability to detect moderate differences in accuracy, and some differences in the 10 to 20 percentage point range may therefore have gone undetected.
- DeepSeek V3.2’s perfect performance raises concerns about possible test-set contamination or memorization, which cannot be excluded without access to the model’s training corpus [17]. This ambiguity necessitates replication on entirely withheld material before claims of true generalisable superiority can be made.
- Models were accessed via web interfaces using default settings, so sampling and decoding parameters were not controlled and may have introduced variability across runs.
- The evaluation focused on multiple-choice items and did not assess clinical reasoning, image interpretation, or hands-on skills that are central to nuclear medicine practice.
- Items were presented in English only and results may not generalize to other languages. Finally, LLM capabilities change rapidly and our results reflect models available in October 2025, so findings may not apply to future versions. Despite these limitations, this proof-of-concept study provides baseline benchmarks and underscores the importance of multi-run reliability assessment in medical AI evaluation.
- Establish reproducibility and inter-run stability as primary performance dimensions in the evaluation of large language models, alongside conventional accuracy metrics.
- Investigate drivers of inter-run variability such as prompt design, decoding temperature, deterministic decoding and backend model-version differences. Develop standardized multi-run benchmarking protocols for high-stakes knowledge assessments, including predefined reliability thresholds alongside accuracy metrics.
- Evaluate performance on the complete 140-question EBNM exam to increase power and representativeness.
- Replicate the evaluation on entirely withheld and newly authored exam items to investigate potential data contamination and to confirm generalisability.
- Audit or request training-data provenance where feasible to clarify the origin of unusually strong performance.
- Expand assessment to include image-based tasks using SPECT, PET and hybrid cases and to evaluate clinical reasoning beyond multiple-choice formats.
- Compare model outputs directly with expert nuclear medicine physicians to establish clinical performance benchmarks.
- Develop and test workflows for safe human–AI collaboration in education and clinical practice.
- Evaluate model behavior on multiple-choice questions in which the number of correct response options is not specified, to assess both knowledge representation and response strategies under uncertainty, including tendencies toward risk-averse or risk-seeking answer selection.
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| API | Application Programming Interface |
| CI | Confidence Interval |
| CV | Coefficient of Variation |
| EBNM | European Board of Nuclear Medicine |
| FEBNM | Fellow of the European Board of Nuclear Medicine |
| GPT | Generative Pre-trained Transformer |
| κ | Cohen’s Kappa |
| LLM | Large Language Model |
| MCQ | Multiple-Choice Question |
| PET | Positron Emission Tomography |
| SD | Standard Deviation |
| SPECT | Single Photon Emission Computed Tomography |
| UEMS | Union Européenne des Médecins Spécialistes |
References
- Lucas, H.C.; Upperman, J.S.; Robinson, J.R. A systematic review of large language models and their implications in medical education. Med. Educ. 2024, 58, 1276–1285. [Google Scholar] [CrossRef] [PubMed]
- Abd-Alrazaq, A.; AlSaad, R.; Alhuwail, D.; Ahmed, A.; Healy, P.M.; Latifi, S.; Aziz, S.; Damseh, R.; Alrazak, S.A.; Sheikh, J. Large Language Models in Medical Education: Opportunities, Challenges, and Future Directions. JMIR Med. Educ. 2023, 9, e48291. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Kung, T.H.; Cheatham, M.; Medenilla, A.; Sillos, C.; De Leon, L.; Elepaño, C.; Madriaga, M.; Aggabao, R.; Diaz-Candido, G.; Maningo, J.; et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digit. Health 2023, 2, e0000198. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Chaiban, T.; Nahle, Z.; Assi, G.; Cherfane, M. The intent of ChatGPT usage and its robustness in medical proficiency exams: A systematic review. Discov. Educ. 2024, 3, 232. [Google Scholar] [CrossRef]
- Lawal, I.O. Nuclear Medicine Training: Skills and Competencies Required for Practice in the 21st Century. World J. Nucl. Med. 2023, 22, 75–77. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Mirzaei, S.; Hustinx, R.; Prior, J.O.; Ozcan, Z.; Boubaker, A.; Farsad, M.; European Union of Medical Specialists and European Board for Nuclear Medicine. Improving Nuclear Medicine Practice with UEMS/EBNM Committees. J. Nucl. Med. 2020, 61, 18N–20N. [Google Scholar] [PubMed]
- European Union of Medical Specialists (UEMS). European Training Requirements for the Specialty of Nuclear Medicine. Available online: https://uems.eanm.org/wp-content/uploads/2021/07/UEMS_European_Training_Requirements__NUCMED_final_May17.pdf (accessed on 29 October 2025).
- Bhayana, R.; Krishna, S.; Bleakney, R.R. Performance of ChatGPT on a Radiology Board-style Examination: Insights into Current Strengths and Limitations. Radiology 2023, 307, 230582. [Google Scholar] [CrossRef] [PubMed]
- Almeida, L.C.; Farina, E.M.J.M.; Kuriki, P.E.A.; Abdala, N.; Kitamura, F.C. Performance of ChatGPT on the Brazilian Radiology and Diagnostic Imaging and Mammography Board Examinations. Radiol. Artif. Intell. 2024, 6, e230103. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Gonzalez, M.A.; Hernandez, M.B.; Peñaloza Perez, M.A.; Lopez Orozco, B.; Cruz Soto, J.T.; Malagon, S. Do Repetitions Matter? Strengthening Reliability in LLM Evaluations. arXiv 2025, arXiv:2509.24086. Available online: https://arxiv.org/abs/2509.24086 (accessed on 29 October 2025).
- Certificate of Fellowship of the European Board of Nuclear Medicine Examples of Multiple Choice Questions (MCQ). Available online: https://uems.eanm.org/wp-content/uploads/2021/07/FEBNM_2012_doc06_MCQ_examples.pdf (accessed on 29 October 2025).
- Kuerbanjiang, W.; Peng, S.; Jiamaliding, Y.; Yi, Y. Performance Evaluation of Large Language Models in Cervical Cancer Management Based on a Standardized Questionnaire: Comparative Study. J. Med. Internet Res. 2025, 27, e63626. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Landis, J.R.; Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef] [PubMed]
- McNemar, Q. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika 1947, 12, 153–157. [Google Scholar] [CrossRef] [PubMed]
- Xu, C.; Guan, S.; Greene, D.; Kechadi, M. Benchmark Data Contamination of Large Language Models: A Survey. arXiv 2024, arXiv:2406.04244. Available online: https://arxiv.org/abs/2406.04244v1 (accessed on 29 October 2025).
- Funk, P.F.; Hoch, C.C.; Knoedler, S.; Knoedler, L.; Cotofana, S.; Sofo, G.; Dezfouli, A.B.; Wollenberg, B.; Guntinas-Lichius, O.; Alfertshofer, M. ChatGPT’s Response Consistency: A Study on Repeated Queries of Medical Examination Questions. Eur. J. Investig. Health Psychol. Educ. 2024, 14, 657–668. [Google Scholar] [CrossRef] [PubMed]
- Krishna, S.; Bhambra, N.; Bleakney, R.; Bhayana, R. Evaluation of Reliability, Repeatability, Robustness, and Confidence of GPT-3.5 and GPT-4 on a Radiology Board-style Examination. Radiology 2024, 311, e232715. [Google Scholar] [CrossRef] [PubMed]
- Certificate of Fellowship of the European. Board of Nuclear Medicine 2012 Information. Eur. J. Nucl. Med. Mol. Imaging 2011, 38, 2289–2301. [Google Scholar] [CrossRef]
- Prigent, A.; Huic, D.; Costa, D.C. Syllabus for Postgraduate Specialization in Nuclear Medicine–2011/2012 Update: Nuclear medicine training in the European Union. Eur. J. Nucl. Med. 2012, 39, 739–743. [Google Scholar] [CrossRef] [PubMed]
- Pons, F.; Delaloye, A.B. The European board of nuclear medicine fellowship examination. Eur. J. Nucl. Med. 2006, 33, 109–110. [Google Scholar] [CrossRef] [PubMed]
- Ozcan, Z.; Kulakiene, I.; Vaz, S.C.; Garzon, J.R.G.; Boubaker, A. Challenges and possibilities for board exams in the COVID-19 era: Experience from the Fellowship Committee of European Board of Nuclear Medicine. Eur. J. Nucl. Med. 2022, 49, 1442–1446. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]


| Rank | Model | Mean Score | Mean Accuracy (%) | SD | 95% CI (%) | CV (%) | Cohen’s κ |
|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V3.2 | 50.0 | 100 | 0 | 100.0–100.0 | 0 | 1.000 |
| 2 | Gemini 2.5 Pro | 46.8 | 93.6 | 4.44 | 82.6–100.0 | 9.5 | 0.370 |
| 3 | Grok-4 | 43.6 | 87.2 | 0.55 | 85.8–88.6 | 1.3 | 0.676 |
| 4 | Mistral Medium 3.1 | 41.8 | 83.6 | 0.45 | 82.4–84.8 | 1.1 | 0.972 |
| 5 | Claude Sonnet 4.5 | 40.8 | 81.6 | 1.79 | 77.2–86.0 | 4.4 | 0.802 |
| 6 | Qwen3 Max | 40.4 | 80.8 | 0.55 | 79.4–82.2 | 1.4 | 0.947 |
| 7 | GPT-5 Pro | 36.8 | 73.6 | 1.64 | 69.6–77.6 | 4.5 | 0.684 |
| 8 | ERNIE 4.5 Turbo | 33.6 | 67.2 | 1.14 | 64.4–70.0 | 3.4 | 0.500 |
| 9 | Llama 3.3 70B | 32.0 | 64.0 | 0.71 | 62.2–65.8 | 2.2 | 0.670 |
| 10 | Falcon H1-34B | 26.8 | 53.6 | 2.49 | 47.4–59.8 | 9.3 | 0.543 |
| Model | Mean κ | Range | Interpretation |
|---|---|---|---|
| DeepSeek V3.2 | 1.000 | 1.000–1.000 | Almost Perfect |
| Mistral Medium 3.1 | 0.972 | 0.929–1.000 | Almost Perfect |
| Qwen3 Max | 0.947 | 0.864–1.000 | Almost Perfect |
| Claude Sonnet 4.5 | 0.802 | 0.732–0.922 | Almost Perfect |
| GPT-5 Pro | 0.684 | 0.508–0.896 | Substantial |
| Grok-4 | 0.676 | 0.558–1.000 | Substantial |
| Llama 3.3 70B | 0.670 | 0.566–0.826 | Substantial |
| Falcon H1-34B | 0.543 | 0.356–0.878 | Moderate |
| ERNIE 4.5 Turbo | 0.500 | 0.288–0.955 | Moderate |
| Gemini 2.5 Pro | 0.370 | 0.000–1.000 | Fair |
| Model A | Model B | χ2 | p-Value |
|---|---|---|---|
| DeepSeek V3.2 | Falcon H1-34B | 21.00 | p < 0.0011 |
| Falcon H1-34B | Gemini 2.5 Pro | 21.00 | p < 0.0011 |
| DeepSeek V3.2 | ERNIE 4.5 Turbo | 17.00 | p < 0.0011 |
| DeepSeek V3.2 | Llama 3.3 70B | 17.00 | p < 0.0011 |
| ERNIE 4.5 Turbo | Gemini 2.5 Pro | 17.00 | p < 0.0011 |
| Gemini 2.5 Pro | Llama 3.3 70B | 17.00 | p < 0.0011 |
| Falcon H1-34B | Grok-4 | 14.22 | p < 0.0011 |
| DeepSeek V3.2 | GPT-5 Pro | 14.00 | p < 0.0011 |
| GPT-5 Pro | Gemini 2.5 Pro | 14.00 | p < 0.0011 |
| Falcon H1-34B | Mistral Medium 3.1 | 11.27 | p < 0.0011 |
| Falcon H1-34B | Qwen3 Max | 11.00 | p < 0.0011 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Stelling, H.; Brink, I.; Grieb, G.; Kraus, A.; Güler, I. Reliability and Performance Stability of Large Language Models in Medical Knowledge Assessment: Evidence from the European Board of Nuclear Medicine Examination. AI 2026, 7, 77. https://doi.org/10.3390/ai7020077
Stelling H, Brink I, Grieb G, Kraus A, Güler I. Reliability and Performance Stability of Large Language Models in Medical Knowledge Assessment: Evidence from the European Board of Nuclear Medicine Examination. AI. 2026; 7(2):77. https://doi.org/10.3390/ai7020077
Chicago/Turabian StyleStelling, Henrik, Ingo Brink, Gerrit Grieb, Armin Kraus, and Ibrahim Güler. 2026. "Reliability and Performance Stability of Large Language Models in Medical Knowledge Assessment: Evidence from the European Board of Nuclear Medicine Examination" AI 7, no. 2: 77. https://doi.org/10.3390/ai7020077
APA StyleStelling, H., Brink, I., Grieb, G., Kraus, A., & Güler, I. (2026). Reliability and Performance Stability of Large Language Models in Medical Knowledge Assessment: Evidence from the European Board of Nuclear Medicine Examination. AI, 7(2), 77. https://doi.org/10.3390/ai7020077

