Clinical Reliability of Large Language Models in Complex Haematology: A Multidimensional Evaluation in Hemophilia–Oncology
Abstract
1. Introduction
2. Materials and Methods
2.1. Study Type and Settings
2.2. Haematology Cases
2.3. Clinical Data Collection
- Assessment referred to the model’s ability to appropriately frame the clinical case within a hemophilia oncology context, identify haemorrhagic risk based on hemophilia type and severity, recognise cancer related modifiers of bleeding or thrombosis risk, consider procedure related risk when surgery was involved, and identify missing but clinically essential information required for safe management, such as baseline factor levels, inhibitor status, bleeding phenotype, and planned invasive procedures.
- Decision and Rationale referred to the coherence and transparency of the clinical reasoning process, including the explicit balancing of oncologic effectiveness and haemostatic safety, the justification of proposed diagnostic and therapeutic choices, prioritisation of interventions, and consistency with established clinical principles such as multidisciplinary coordination, peri-procedural haemostatic planning, and monitoring strategies.
- Strategy referred to the overall management approach proposed by the model, including logical sequencing of interventions, feasibility in real-world clinical practice, coordination between haematology and oncology care, integration of surgery, chemotherapy, radiotherapy or immunotherapy when applicable, anticipation of potential complications, and the presence of a structured follow-up plan.
- Selected drugs referred to the appropriateness and safety of the pharmacological options proposed by the model in relation to both hemophilia care and the oncologic context, including haemostatic and supportive therapies, as well as avoidance of contraindicated or unsafe agents.
- Regimen referred to the extent to which the model translated drug choices into an actionable treatment plan, including dosing logic, timing, duration, peri-procedural scheduling, monitoring parameters, and criteria for therapy adjustment in response to bleeding events, laboratory findings, or treatment-related complications.
3. Results
3.1. Inter-Rater Agreement Between Clinicians
3.1.1. Agreement by Area
3.1.2. Magnitude of Disagreement by Area
3.1.3. Agreement by Clinical Case
3.2. Perceived Reliability of ChatGPT and Copilot
3.2.1. Reliability Scores by Clinical Case
3.2.2. Reliability by Area
4. Discussion
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Swift, A. The role of hematology in modern medicine: Insights into blood disorders and therapeutic strategies. Hematol. Blood Disord. 2024, 7, 185. [Google Scholar]
- Hoffman, R.; Benz, E.J.; Silberstein, L.E.; Heslop, H.E.; Weitz, J.I.; Anastasi, J.; Salama, M.E. Hematology: Basic Principles and Practice, 7th ed.; Elsevier: Philadelphia, PA, USA, 2018. [Google Scholar]
- Falanga, A.; Marchetti, M.; Vignoli, A. Coagulation and cancer: Biological and clinical aspects. J. Thromb. Haemost. 2013, 11, 223–233. [Google Scholar] [CrossRef] [PubMed]
- Rao, A.; Pang, M.; Kim, J.; Kamineni, M.; Lie, W.; Prasad, A.K.; Landman, A.; Dreyer, K.; Succi, M.D. Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow. medRxiv 2023. Update in J. Med. Internet Res. 2023, 25, e48659. https://doi.org/10.2196/48659. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Li, R.; Wang, X.; Yu, H. Two Directions for Clinical Data Generation with Large Language Models: Data-to-Label and Label-to-Data. In Findings of the Association for Computational Linguistics: EMNLP 2023; Association for Computational Linguistics: Singapore, 2023; Volume 2023, pp. 7129–7143. [Google Scholar] [CrossRef]
- Kung, T.H.; Cheatham, M.; Medenilla, A.; Sillos, C.; De Leon, L.; Elepaño, C.; Madriaga, M.; Aggabao, R.; Diaz-Candido, G.; Maningo, J.; et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLoS Digit. Health 2023, 2, e0000198. [Google Scholar] [CrossRef] [PubMed]
- Gilson, A.; Safranek, C.W.; Huang, T.; Socrates, V.; Chi, L.; Taylor, R.A.; Chartash, D. Correction: How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment. JMIR Med. Educ. 2024, 10, e57594, Erratum for JMIR Med. Educ. 2023, 9, e45312. https://doi.org/10.2196/45312. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Frieder, S.; Pinchetti, L.; Griffiths, R.-R.; Salvatori, T.; Lukasiewicz, T.; Petersen, P.; Chevalier, A.; Berner, J. Mathematical capabilities of ChatGPT. Adv. Neural Inf. Process. Syst. 2023, 36, 27699–27744. [Google Scholar] [CrossRef]
- Woesle, C.; Fischer-Brandies, L.; Buettner, R. A Systematic Literature Review of Hallucinations in Large Language Models. IEEE Access 2025, 13, 148231–148253. [Google Scholar] [CrossRef]
- Moor, M.; Banerjee, O.; Abad, Z.S.H.; Krumholz, H.M.; Leskovec, J.; Topol, E.J.; Rajpurkar, P. Foundation models for generalist medical artificial intelligence. Nature 2023, 616, 259–265. [Google Scholar] [CrossRef]
- Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. Large language models encode clinical knowledge. Nature 2023, 620, 172–180. [Google Scholar] [CrossRef]
- He, J.; Baxter, S.L.; Xu, J.; Xu, J.; Zhou, X.; Zhang, K. The practical implementation of artificial intelligence technologies in medicine. Nat. Med. 2019, 25, 30–36. [Google Scholar] [CrossRef]
- Cascella, M.; Montomoli, J.; Bellini, V.; Bignami, E. Evaluating the Feasibility of ChatGPT in Healthcare: An Analysis of Multiple Clinical and Research Scenarios. J. Med. Syst. 2023, 47, 33. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large language models in medicine. Nat. Med. 2023, 29, 1930–1940. [Google Scholar] [CrossRef]
- Cross, J.L.; Choma, M.A.; Onofrey, J.A. Bias in medical AI: Implications for clinical decision-making. PLoS Digit. Health 2024, 3, e0000651. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Alkaissi, H.; McFarlane, S.I. Artificial Hallucinations in ChatGPT: Implications in Scientific Writing. Cureus 2023, 15, e35179. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Atil, B.; Chittams, A.; Fu, L.; Ture, F.; Xu, L.; Baldwin, B. LLM Stability: A detailed analysis with some surprises (Version 1). arXiv 2024, arXiv:2408.04667. [Google Scholar]
- Balagurunathan, Y.; Mitchell, R.; El Naqa, I. Requirements and reliability of AI in the medical context. Phys. Med. 2021, 83, 72–78. [Google Scholar] [CrossRef] [PubMed]
- Srivastava, A.; Brewer, A.K.; Mauser-Bunschoten, E.P.; Key, N.S.; Kitchen, S.; Llinas, A.; Ludlam, C.A.; Mahlangu, J.N.; Mulder, K.; Poon, M.C.; et al. Guidelines for the management of hemophilia. Hemophilia 2013, 19, e1–e47. [Google Scholar] [CrossRef] [PubMed]
- Tarniceriu, C.C.; Hurjui, L.L.; Tanase, D.M.; Haisan, A.; Tepordei, R.T.; Statescu, G.; Vicoleanu, S.A.P.; Lupu, A.; Lupu, V.V.; Ursaru, M.; et al. Inherited hemophilia—A multidimensional chronic disease that requires a multidisciplinary approach. Life 2025, 15, 530. [Google Scholar] [CrossRef]
- Koc, B.; Zulfikar, B. A Challenge for Hemophilia Treatment: Hemophilia and Cancer. J. Pediatr. Hematol. Oncol. 2021, 43, e29–e32. [Google Scholar] [CrossRef]
- Tagliaferri, A.; Di Perna, C.; Santoro, C.; Schinco, P.; Santoro, R.; Rossetti, G.; Coppola, A.; Morfini, M.; Franchini, M.; Italian Association of Hemophilia Centers. Cancers in patients with hemophilia: A retrospective study from the Italian Association of Hemophilia Centers. J. Thromb. Haemost. 2012, 10, 90–95. [Google Scholar] [CrossRef] [PubMed]
- Kempton, C.L.; Makris, M.; Holme, P.A. Management of comorbidities in hemophilia. Hemophilia 2021, 27, 37–45. [Google Scholar] [CrossRef] [PubMed]
- Sallam, M. ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns. Healthcare 2023, 11, 887. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Kim, Y.; Jeong, H.; Chen, S.; Li, S.S.; Lu, M.; Alhamoud, K.; Park, C.; Mun, J.; Grau, C.; Jung, M.; et al. Medical hallucinations in foundation models and their impact on healthcare. arXiv 2025, arXiv:2503.05777. [Google Scholar]
- Gao, C.A.; Howard, F.M.; Markov, N.S.; Dyer, E.C.; Ramesh, S.; Luo, Y.; Pearson, A.T. Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. NPJ Digit. Med. 2023, 6, 75. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Zitu, M.M.; Le, T.D.; Duong, T.; Haddadan, S.; Garcia, M.; Amorrortu, R.; Zhao, Y.; Rollison, D.E.; Thieu, T. Large language models in cancer: Potentials, risks, and safeguards. BJR Artif. Intell. 2025, 2, ubae019. [Google Scholar] [CrossRef]
- Rezende, S.M.; Neumann, I.; Angchaisuksiri, P.; Awodu, O.; Boban, A.; Cuker, A.; Curtin, J.A.; Fijnvandraat, K.; Gouw, S.C.; Gualtierotti, R.; et al. International Society on Thrombosis and Haemostasis clinical practice guideline for treatment of congenital hemophilia A and B based on the Grading of Recommendations Assessment, Development, and Evaluation methodology. J. Thromb. Haemost. 2024, 22, 2629–2652. [Google Scholar] [CrossRef]
- Hermans, C.; Apte, S.; Santagostino, E. Invasive procedures in patients with hemophilia: Review of low dose protocols and experience with extended half life FVIII and FIX concentrates and non replacement therapies. Hemophilia 2020, 27, 46–52. [Google Scholar] [CrossRef] [PubMed]
- Johnson, D.; Goodman, R.; Patrinely, J.; Stone, C.; Zimmerman, E.; Donald, R.; Chang, S.; Berkowitz, S.; Finn, A.; Jahangir, E.; et al. Assessing the Accuracy and Reliability of AI-Generated Medical Responses: An Evaluation of the Chat-GPT Model. Res. Sq. 2023. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Lee, P.; Bubeck, S.; Petro, J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N. Engl. J. Med. 2023, 388, 1233–1239. [Google Scholar] [CrossRef] [PubMed]
- Mishra, H.P.; Gupta, R. Leveraging Generative AI for Drug Safety and Pharmacovigilance. Curr. Rev. Clin. Exp. Pharmacol. 2025, 20, 89–97. [Google Scholar] [CrossRef] [PubMed]
- Ali, A.M.A.; Alrobaian, M.M. Strengths and weaknesses of current and future prospects of artificial intelligence-mounted technologies applied in the development of pharmaceutical products and services. Saudi Pharm. J. 2024, 32, 102043. [Google Scholar] [CrossRef] [PubMed]
- Cabitza, F.; Rasoini, R.; Gensini, G.F. Unintended consequences of machine learning in medicine. JAMA 2017, 318, 517–518. [Google Scholar] [CrossRef]
- Ciobanu-Caraus, O.; Aicher, A.; Kernbach, J.M.; Regli, L.; Serra, C.; Staartjes, V.E. A critical moment in machine learning in medicine: On reproducible and interpretable learning. Acta Neurochir. 2024, 166, 14. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
- Jeong, C. Fine tuning and utilization methods of domain specific LLMs (Version 2). arXiv 2024, arXiv:2401.02981. [Google Scholar]
- Zhang, W.; Zhang, J. Hallucination mitigation for retrieval augmented large language models: A review. Mathematics 2025, 13, 856. [Google Scholar] [CrossRef]
- Xia, Y.; Zhou, J.; Shi, Z.; Chen, J.; Huang, H. Improving retrieval augmented language model with self reasoning (Version 3). arXiv 2024, arXiv:2407.19813. [Google Scholar]


| Case [21] | Hemophilia (Type and Severity) | Age at Cancer Diagnosis (Years) | Type of Cancer | HIV/HCV Status | Complications | Treatment | Outcome |
|---|---|---|---|---|---|---|---|
| 1 | A, Moderate | 14 | Acute Lymphoblastic Leukaemia (ALL) | Negative | None | Chemotherapy | Alive |
| 2 | A, Severe | 31 | Thyroid Carcinoma | Negative | None | Total thyroidectomy with radical cervical dissection, Radioiodine | Alive |
| 3 | A, Severe | 52 | Rectal Cancer | Negative | Prolonged bleeding during surgery | Laparoscopic low-anterior resection of the rectum, Chemotherapy, Radiotherapy | Alive |
| 4 | A, Severe | 67 | Malignant Melanoma | Negative | Disease progression, chemotherapy regimen change | Right forefinger amputation and axillary curettage, Interferon alpha 2-b | Alive |
| 5 | B, Moderate | 58 | Basal Cell Carcinoma | Negative | None | Excision of a skin lesion | Alive |
| 6 | B, Moderate | 79 | Gastric Cancer | Negative | None | Partial gastrectomy | Alive |
| Area | Number of Ratings | % of Agreement for ChatGPT | % of Agreement for Copilot |
|---|---|---|---|
| Assessment | 6 | 100 | 100 |
| Decision and Rationale | 6 | 66.7 | 100 |
| Drug | 6 | 0 | 0 |
| Regimen | 6 | 33.3 | 0 |
| Strategy | 6 | 83.3 | 100 |
| Case | Number of Areas | % Agree ChatGPT | Distance ChatGPT | % Agree Copilot | Distance Copilot |
|---|---|---|---|---|---|
| 1 | 5 | 40 | 0.6 | 60 | 0.4 |
| 2 | 5 | 20 | 1.0 | 60 | 0.4 |
| 3 | 5 | 60 | 0.4 | 60 | 0.6 |
| 4 | 5 | 80 | 0.2 | 60 | 0.4 |
| 5 | 5 | 80 | 0.2 | 60 | 0.4 |
| 6 | 5 | 60 | 0.4 | 60 | 0.4 |
| Case | Score ChatGPT Case | Score Copilot Case |
|---|---|---|
| 1 | 1.5 | 1.8 |
| 2 | 1.1 | 2 |
| 3 | 2 | 2.1 |
| 4 | 2.3 | 2 |
| 5 | 1.9 | 2 |
| 6 | 2 | 2 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Porreca, A.; Proietti, S.; Maturo, F.; Bonassi, S.; Zanon, E. Clinical Reliability of Large Language Models in Complex Haematology: A Multidimensional Evaluation in Hemophilia–Oncology. Hemato 2026, 7, 10. https://doi.org/10.3390/hemato7020010
Porreca A, Proietti S, Maturo F, Bonassi S, Zanon E. Clinical Reliability of Large Language Models in Complex Haematology: A Multidimensional Evaluation in Hemophilia–Oncology. Hemato. 2026; 7(2):10. https://doi.org/10.3390/hemato7020010
Chicago/Turabian StylePorreca, Annamaria, Stefania Proietti, Fabrizio Maturo, Stefano Bonassi, and Ezio Zanon. 2026. "Clinical Reliability of Large Language Models in Complex Haematology: A Multidimensional Evaluation in Hemophilia–Oncology" Hemato 7, no. 2: 10. https://doi.org/10.3390/hemato7020010
APA StylePorreca, A., Proietti, S., Maturo, F., Bonassi, S., & Zanon, E. (2026). Clinical Reliability of Large Language Models in Complex Haematology: A Multidimensional Evaluation in Hemophilia–Oncology. Hemato, 7(2), 10. https://doi.org/10.3390/hemato7020010

