Next Article in Journal
Pattern of Reported Infections Among Paediatric Patients with Sickle Cell Disease: A Single-Centre Cohort Study in Nigeria
Previous Article in Journal
A Pragmatic First-Line Screening Assay for PDGFR Rearrangements: A Real-World Clinical Validation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Clinical Reliability of Large Language Models in Complex Haematology: A Multidimensional Evaluation in Hemophilia–Oncology

1
Department of Human Sciences and Quality of Life Promotion, San Raffaele University, Via di Val Cannuta 247, 00166 Rome, Italy
2
Clinical and Molecular Epidemiology, Istituto di Ricovero e Cura a Carattere Scientifico (IRCCS) San Raffaele Roma, Via di Val Cannuta 247, 00166 Rome, Italy
3
AgEA Coordinating Body, Via Palestro 81, 00187 Rome, Italy
4
Department of Economics, Statistics and Business, Faculty of Technological & Innovation Sciences, Universitas Mercatorum, Piazza Mattei 10, 00186 Rome, Italy
5
Hemophilia Centre, Internal Medicine—Department of Medicine, Padua University Hospital, 35128 Padua, Italy
*
Author to whom correspondence should be addressed.
Hemato 2026, 7(2), 10; https://doi.org/10.3390/hemato7020010
Submission received: 12 February 2026 / Revised: 21 March 2026 / Accepted: 26 March 2026 / Published: 31 March 2026

Abstract

Background: The co-existence of hemophilia and cancer presents one of the most complex clinical scenarios, demanding individualised therapeutic planning to balance oncologic efficacy and hemostatic safety. This study evaluated the ability of two Large Language Models (LLMs)—ChatGPT (GPT-4) and Microsoft Copilot (GPT-4–based)—to generate clinically appropriate recommendations for real cases of hemophilia with concurrent malignancy. Methods: Six consecutive adult cases of hemophilia and cancer, managed at the Hemophilia Centre of Padua, Italy, were selected for evaluation. Identical structured prompts were submitted to both LLMs. Two independent expert clinicians rated the model outputs across five domains (Decision/Rationale, Strategy, Selected Drug, Regimen, and Assessment) using a four-level ordinal scale. Results: LLMs demonstrated uneven performances. Outputs were consistently rated as highly reliable in domains involving high-level synthesis, such as Assessment and Strategy. However, substantial limitations were observed in the clinically demanding domains of Selected Drug and Regimen. Critically, in the Selected Drug domain, there was complete agreement between the two expert raters for neither system. This severe lack of concordance signifies that clinicians assigned different adequacy ratings to the same output in every case, reflecting ambiguity, lack of specificity, and inconsistent clinical interpretability of the drug-related information provided by LLMs. Conclusions: While LLMs possess the capacity for high-level reasoning and strategic planning, their inability to translate principles into precise, consistent, and clinically interpretable therapeutic plans—particularly regarding drug selection and treatment regimens—is a significant constraint. These deficiencies, highlighted by the minimal expert concordance in critical domains, necessitate rigorous clinical validation before the responsible integration of LLMs into the management of this uniquely vulnerable patient population.

1. Introduction

Haematology plays an essential role in modern clinical practice, as the accurate diagnosis and effective management of blood disorders significantly influence patient outcomes and quality of life [1]. Conditions such as anemia, leukemia, hereditary coagulation defects, and thrombophilias require the integration of laboratory data, clinical findings, and personalised therapeutic plans [2]. In parallel, oncology has advanced toward increasingly complex diagnostic and therapeutic strategies, with precision medicine and multidisciplinary evaluation now standard practice [3]. The coexistence of hematologic disorders and malignancy represents one of the most challenging intersections in medicine, as clinicians must balance oncologic efficacy with hemostatic safety in patients who often have limited physiological reserves [3]. In recent years, artificial intelligence (AI) and natural language processing (NLP) have driven rapid advances in Large Language Models (LLMs). Systems such as ChatGPT-4 and Microsoft Copilot demonstrate notable proficiency in contextual understanding, summarising complex information, and generating clinically structured outputs, making them potentially valuable tools for decision support and medical education [4,5,6,7,8,9,10,11,12]. Several studies have shown that LLMs can achieve performance levels comparable to those of medical trainees on standardized examinations and clinical reasoning tasks [6,7,13,14,15]. Despite this promise, LLM-generated content remains variable and context dependent. Concerns include hallucinated clinical facts, inaccurate drug recommendations, lack of adherence to guidelines, and variability in reasoning quality across different clinical tasks [9,16,17,18]. These limitations are particularly critical in high-risk settings such as hematology and oncology, where decisions require precision and nuanced interpretation. The coexistence of hemophilia and cancer introduces a unique layer of complexity. Hemophilia inherently increases bleeding risk, requiring individualized hemostatic preparation, factor replacement strategies, and close perioperative monitoring [19,20]. Meanwhile, cancer therapies, including surgery, chemotherapy, radiotherapy, and immunotherapy, not only compromise the coagulation cascade but also frequently induce significant thrombocytopenia due to bone marrow suppression. This hematological toxicity becomes a major clinical decision point when managing invasive procedures, especially in hemophilia patients, as the risk of bleeding is multiplied by the simultaneous presence of a clotting factor deficiency and a low platelet count [3]. These complications require specialized interdisciplinary evaluation and nuanced clinical judgement. Existing evidence on the management of malignancies in patients with hemophilia is limited, but available reports emphasize the need for tailored perioperative planning, multidisciplinary coordination, and careful balancing of competing risks [21,22,23]. Given these challenges, evaluating whether LLMs can reliably support clinical decision making in this rare and complex setting is both timely and necessary. Prior work in AI evaluation has highlighted that domain-specific benchmarks and expert validation are essential to assess model performance in medical contexts, as general-purpose metrics fail to capture clinically relevant errors and hallucinations [24,25,26,27]. The present study addresses this gap by examining the ability of two LLMs—ChatGPT and Microsoft Copilot—to generate coherent, clinically appropriate recommendations for real cases of hemophilia with concurrent malignancy. Two clinicians independently rated each model’s output across key therapeutic domains (Decision and Rationale, Strategy, Selected Drug, Regimen, and Assessment) using a structured four-level rubric (Complete, Partially Complete, Mixed, Incomplete). This approach allowed a multidimensional assessment of completeness, reliability, internal coherence and inter-rater agreement. Through this structured evaluation, this study aims to identify both the strengths and the limitations of current LLMs in a high-risk hematologic–oncologic context, offering evidence to guide their safe, responsible and clinically informed integration into medical practice.

2. Materials and Methods

2.1. Study Type and Settings

This work was structured as a comparative expert-based evaluation study designed to assess the adequacy, completeness, and internal coherence of clinical outputs generated by two Large Language Models (LLMs) when applied to complex real cases of hemophilia associated with malignancy. This study was conducted at the Hemophilia Centre of Padua, Italy, where two expert clinicians independently evaluated model-generated outputs based on six published clinical cases of hemophilia with cancer derived from a previously published study [21] (Table 1). Each case was reformulated into a standardized clinical record summarizing the key elements required for therapeutic reasoning, including hemophilia type and severity, cancer diagnosis, HIV/HCV status, complications, treatments administered, and outcome. For every case, identical structured prompts were submitted to the two LLMs under evaluation: ChatGPT (OpenAI, GPT-4–based version available at the time of access) and Microsoft Copilot (GPT-4–based implementation). Both systems were accessed through their publicly available interfaces using default settings, without plug-ins, fine-tuning, or external retrieval tools. The models were accessed in October 2025, and each prompt was submitted once without repeated runs. The models were asked to provide clinical recommendations for five predefined components of the therapeutic process: Decision and Rationale, Strategy, Selected Drug, Regimen, and overall Assessment. All outputs were recorded in full and anonymised before clinical review. The evaluation was performed independently by two physicians from the Hemophilia Centre of Padua, each with extensive experience in the management of adults with inherited bleeding disorders and their comorbidities, including oncologic conditions. The clinicians were instructed to assess the adequacy of each model-generated response using a predefined four-level ordinal scale: complete, partially complete, mixed or incomplete. The criteria for each category were defined in advance and focused on the clinical utility of the response in real decision-making contexts. The reviewers were blinded to each other’s assessments, and no consensus procedure was applied to preserve independent judgments and allow quantification of inter-rater agreement and discrepancies in expert interpretation. This design enabled a systematic comparison of ChatGPT and Copilot across clinically heterogeneous cases, allowing the evaluation of model performance at the level of individual therapeutic domains as well as aggregated case-level reliability.

2.2. Haematology Cases

Six published clinical cases of patients with hemophilia and concomitant oncologic disease were extracted from a previously published study [21] and used as standardized input for the evaluation. All clinical assessments were conducted independently by expert clinicians at the Hemophilia Centre of Padua, who evaluated the outputs generated by large language models, including ChatGPT and Microsoft Copilot, in the context of complex hemophilia –oncology management. The patients described in the study, all monitored at the Comprehensive Hemophilia Care Centre at Istanbul University, presented various types of hemophilia (A and B) with different levels of severity and distinct cancer types, including acute lymphoblastic leukaemia, acute myeloid leukaemia, thyroid cancer, rectal cancer, malignant melanoma, basal cell carcinoma, and gastric cancer. For each case, ChatGPT and Copilot were asked to suggest the most appropriate clinical management, including the recommended haematological and oncological therapy, as well as strategies for hemostasis control in case of surgical interventions. A specialist physician then evaluated the responses generated by the models to assess the degree of agreement with existing clinical guidelines and the therapeutic decisions adopted in the actual cases. This process allowed us to explore whether such models can effectively provide decision support in clinical settings, especially in high-risk, complex scenarios such as the intersection of hemophilia and cancer. The analysis included comparing the models’ responses to the physician’s recommendations, assessing parameters such as accuracy and completeness, and considering specific complications, such as bleeding risk and potential interactions between treatments. The findings from this comparative evaluation can provide valuable insights into the future integration of artificial intelligence tools into the clinical management of patients with rare, multidisciplinary conditions, aiming to enhance care quality and support clinical decision-making.

2.3. Clinical Data Collection

Six published clinical cases of patients with hemophilia and concomitant oncologic disease were identified from the literature. For each case, the clinical management reported in the original publication was concealed. Case descriptions were provided to two large language models, ChatGPT and Copilot, which were instructed to independently propose a clinical management approach based solely on the available information, assuming responsibility for the case and without knowledge of the actual therapeutic decisions described in the source articles. The language models were prompted to structure their responses by explicitly addressing five predefined domains corresponding to key components of clinical reasoning and therapeutic planning in hemophilia oncology, namely assessment, decision and rationale, overall strategy, selected drugs, and treatment regimen. These domains represent the areas in which the models formulated and articulated their clinical opinions. All model-generated outputs were independently evaluated by two clinicians with expertise in hemophilia and oncologic care. For each domain and for each clinical case, the clinicians assessed the quality of the model responses using a four-level qualitative rubric designed to capture completeness and clinical adequacy. The evaluative categories were defined as complete, partially complete, mixed, and incomplete. A response was considered complete when it fully addressed the relevant clinical aspects within a domain in a coherent, accurate, and guideline-consistent manner, providing sufficient detail to support clinical decision-making. Responses were classified as partially complete when they addressed some key elements but omitted relevant information necessary for full clinical adequacy. Mixed responses contained a combination of appropriate and inappropriate or irrelevant elements within the same domain, resulting in heterogeneous quality. Incomplete responses failed to provide essential elements required for clinical assessment or therapeutic planning, rendering them inadequate for clinical decision making.
The five domains addressed by the language models were defined as follows:
  • Assessment referred to the model’s ability to appropriately frame the clinical case within a hemophilia oncology context, identify haemorrhagic risk based on hemophilia type and severity, recognise cancer related modifiers of bleeding or thrombosis risk, consider procedure related risk when surgery was involved, and identify missing but clinically essential information required for safe management, such as baseline factor levels, inhibitor status, bleeding phenotype, and planned invasive procedures.
  • Decision and Rationale referred to the coherence and transparency of the clinical reasoning process, including the explicit balancing of oncologic effectiveness and haemostatic safety, the justification of proposed diagnostic and therapeutic choices, prioritisation of interventions, and consistency with established clinical principles such as multidisciplinary coordination, peri-procedural haemostatic planning, and monitoring strategies.
  • Strategy referred to the overall management approach proposed by the model, including logical sequencing of interventions, feasibility in real-world clinical practice, coordination between haematology and oncology care, integration of surgery, chemotherapy, radiotherapy or immunotherapy when applicable, anticipation of potential complications, and the presence of a structured follow-up plan.
  • Selected drugs referred to the appropriateness and safety of the pharmacological options proposed by the model in relation to both hemophilia care and the oncologic context, including haemostatic and supportive therapies, as well as avoidance of contraindicated or unsafe agents.
  • Regimen referred to the extent to which the model translated drug choices into an actionable treatment plan, including dosing logic, timing, duration, peri-procedural scheduling, monitoring parameters, and criteria for therapy adjustment in response to bleeding events, laboratory findings, or treatment-related complications.
For each clinical case, both clinicians independently rated the model-generated content across all domains using the qualitative categories described above (Incomplete, Partially, Mixed, Complete).

3. Results

Six clinical cases were evaluated, covering different combinations of hemophilia type and severity, solid and hematologic malignancies, treatment pathways, and outcomes. For each case, two clinicians independently judged the information generated by ChatGPT and Copilot across several areas (Decision and Rationale, Strategy, Drug, Regimen, Assessment) using a four-level ordinal scale (Incomplete, Partially, Mixed, Complete). All results reported below refer to these ordinal judgements. AI-generated outputs are summarized from the original model responses and reported in condensed form for illustrative purposes. These tables are provided in the Supplementary Materials to support the interpretation of variability across clinical domains and cases.

3.1. Inter-Rater Agreement Between Clinicians

3.1.1. Agreement by Area

Table 2 summarises the agreement between the two clinicians in each content area, expressed as the percentage of ratings in which they assigned the same category to a given system.
In the Assessment area, the two clinicians always agreed on both ChatGPT and Copilot (100% agreement). This indicates that when they judged whether the summary assessment produced by the models was complete or not, their interpretations were fully concordant. In Decision and Rationale, which captures the justification of diagnostic and therapeutic choices, clinicians also showed perfect agreement for Copilot (100%), while agreement for ChatGPT was slightly lower (approximately 66.7%). This means that in a non-negligible proportion of decisions and rational judgements, one clinician considered the explanation more complete than the other.
The most problematic areas were Drug and Regimen. For the Drug information, agreement dropped to 0% for both systems; that is, in all cases, the clinicians assigned different categories to the same output. For Regimen, agreement was about 33% for ChatGPT and 0% for Copilot, again indicating substantial divergence in how the two clinicians perceived the adequacy of the proposed therapeutic regimens. In the Strategy area, which focuses on the overall management plan, agreement was high: about 83% for ChatGPT and 100% for Copilot. Overall, Table 2 shows that the two clinicians were highly consistent in areas such as Assessment and Strategy, while Drug and Regimen were associated with much larger variability in their opinions.

3.1.2. Magnitude of Disagreement by Area

While the percentage of agreement is intuitive, it does not convey how large the disagreements are when clinicians disagree. To capture this aspect, we converted the four ordinal categories into scores from 0 to 3 (0 = Incomplete, 1 = Partially, 2 = Mixed, 3 = Complete) and computed, for each area and system, the mean absolute difference between the two clinicians. Figure 1 reports these mean ordinal distances. A value of 0 corresponds to perfect agreement, a value of 1 corresponds to an average disagreement of one level on the scale, and higher values indicate more substantial discrepancies. In Assessment, the mean distance was 0 for both systems, confirming the perfect agreement already evident from Table 2. In Decision and Rationale, the mean distance was around 0.33 for ChatGPT and 0 for Copilot, reflecting that for ChatGPT, a minority of items differed by 1 level between the two raters, whereas for Copilot, all ratings coincided. In Drug, the mean distances were close to 1 for ChatGPT and slightly above 1 for Copilot. This shows that, on average, the clinicians differed by about one full category when judging the drug-related information produced by either system. In other words, what one clinician rated as Incomplete was often rated as Partially or Mixed by the other, and the pattern was even more pronounced for Copilot. In Regimen, the mean distance was approximately 0.67 for ChatGPT and 1.00 for Copilot. Thus, in this area, the degree of disagreement was again substantial, especially for Copilot, with many ratings differing by one level and some by more than one level. Finally, in Strategy, the mean distance was around 0.33 for ChatGPT and 0 for Copilot, consistent with high agreement in this domain and with slightly more variability in the interpretation of ChatGPT outputs than of Copilot outputs. Taken together, Figure 1 complements Table 2 by showing that disagreements are not only more frequent in Drug and Regimen but also larger in magnitude, whereas Assessment and Strategy show both high agreement and minimal differences between raters.

3.1.3. Agreement by Clinical Case

To evaluate whether agreement depended on the specific clinical scenario, we aggregated ratings within each case and computed agreement statistics for each case. Table 3 reports, for each of the six clinical cases, the percentage of areas in which the two clinicians assigned the same category and the mean ordinal distance between their ratings. For ChatGPT, the percentage of agreement ranged from 20% (case 2) to 80% (cases 4 and 5). The lowest agreement was observed in case 2, a severe hemophilia A patient with thyroid carcinoma, where the two clinicians often gave different complete judgments across areas. In contrast, cases 4 and 5, which involved melanoma and basal cell carcinoma, respectively, showed high agreement between clinicians, with only occasional one-level differences. For Copilot, the percentage of agreement was more stable, with a value of 60% across all six cases. This suggests that, overall, clinicians tend to evaluate Copilot’s outputs in a more homogeneous way across different clinical scenarios. The mean ordinal distances confirm this pattern. For ChatGPT, they ranged from 0.2 to 1.0, indicating considerable variability in disagreement size, whereas for Copilot, they were generally lower and more stable (around 0.4). This supports the impression that Copilot’s responses elicited more consistent interpretations among clinicians, although, as shown above, disagreement could still be substantial in specific areas, such as Drug and Regimen.

3.2. Perceived Reliability of ChatGPT and Copilot

3.2.1. Reliability Scores by Clinical Case

To summarise the reliability of the two models in each clinical case, we used the same 0–3 scoring scheme and averaged the two clinicians’ scores across all areas for each system and case. A higher score reflects a judgment closer to “Complete” across the different dimensions considered. Table 4 reports the mean aggregated scores. For ChatGPT, case-level scores ranged roughly between 1.2 and 2.8. The lowest scores were observed in the more complex cases, where therapeutic decisions and complications played a central role, and where the two clinicians also showed lower agreement. The highest ChatGPT score was observed in case 4, a patient with malignant melanoma, where both clinicians judged the generated information as close to complete across all areas. For Copilot, scores were more compact, ranging approximately between 2.0 and 2.7 across the six cases. This suggests that Copilot tended to produce a relatively stable level of information completeness, perceived by the clinicians as generally between “Partially” and “Complete” in most domains, with fewer extreme judgements. Comparing the two systems, ChatGPT appeared more variable, sometimes providing outputs judged as highly complete and sometimes as clearly incomplete, whereas Copilot tended to occupy a narrower band of mid to high reliability. These estimates should be interpreted with caution, given the small number of evaluated cases.

3.2.2. Reliability by Area

While case-level scores capture the overall impression per patient, it is also important to understand in which clinical areas the models were perceived as more or less reliable. Figure 2 shows the mean reliability score by area, again using the 0–3 scale averaged across clinicians and cases. In Assessment, both systems reached the highest scores, close to the “Complete” category. This indicates that when summarizing and integrating the overall clinical status, both ChatGPT and Copilot produced information that the clinicians considered largely adequate. In Strategy, which reflects the global management plan, both systems again achieved relatively high scores, with Copilot slightly higher on average than ChatGPT. This is consistent with the high agreement and low distance observed for this area. In Decision and Rationale, both systems received intermediate scores, reflecting that explanations of the reasoning behind diagnostic and treatment choices were often judged as mixed or only partially complete. ChatGPT and Copilot performed similarly in this respect. The most critical areas were Drug and Regimen. In Drug, mean scores were clearly lower for both systems, indicating that ratings were predominantly in the Incomplete or Partially range. This aligns with the very low agreement and large ordinal distances described earlier and suggests that drug-related information generated by the models was often perceived as insufficient or inconsistently complete. In Regimen, scores were again lower than in Assessment and Strategy, although slightly higher than in Drug, indicating that proposals for treatment regimens were judged as only partially adequate in many instances. Overall, Figure 2 shows that clinicians perceived both models as most reliable for global assessments and management strategies, and least reliable for specifying drug therapy and treatment regimens.

4. Discussion

The coexistence of hemophilia and malignant disease constitutes one of the most intricate scenarios in clinical haematology and oncology [21,22,23].
Managing cancer in patients with congenital coagulation disorders requires continuous recalibration of therapeutic decisions, as standard oncologic pathways must be reconciled with bleeding risk, factor replacement requirements, perioperative constraints, and the potential impact of cancer therapies on hemostatic stability. The literature strongly emphasises that surgical interventions, invasive diagnostic procedures, systemic anticancer treatments, and radiotherapy all demand individualized planning in this population, often incorporating enhanced perioperative monitoring, adjusted chemotherapy schedules, and tailored supportive care to minimize hemorrhagic complications [28,29]. Within this context, evaluating whether Large Language Models (LLMs) can meaningfully support such decision-making is both clinically relevant and methodologically necessary. It is important to clarify that the present evaluation does not directly assess clinical safety but rather the completeness, coherence, and interpretability of model-generated outputs, which may, in turn, influence their safe use in clinical contexts.
The present study shows that LLMs demonstrate uneven performance when confronted with the complexities inherent in haemophilia oncology cases. In domains such as Assessment and Strategy, clinicians rated the generated outputs as generally coherent, informative, and sufficiently aligned with broadly accepted clinical reasoning. These domains often involve the global synthesis of information, the identification of guiding principles for management, and the recognition of overarching clinical priorities. Recent evidence indicates that LLMs perform well in tasks involving abstraction, summary, and contextual organization, which likely explains their relative strength in these aspects of the clinical evaluation [11,12,17]. The high inter-rater agreement in these categories further suggests that the model-generated content was sufficiently clear for independent specialists to interpret it consistently.
However, more clinically demanding domains revealed substantial limitations. For Selected Drug and Regimen, the two areas most closely linked to patient safety in hemophilia complicated by malignancy showed markedly reduced reliability and very low clinician concordance. These domains require a precise understanding of the bleeding risks associated with antineoplastic agents, the hemostatic implications of chemotherapy-induced cytopenias, the perioperative factor-level targets recommended for patients undergoing major cancer surgery, and the potential interaction among factor concentrates, immunotherapies, and supportive medications. Previous studies of LLMs in pharmacology and oncology have already documented recurring weaknesses, including omission of key contraindications, inappropriate regimen sequencing, inaccurate toxicity profiles, and lack of attention to comorbid conditions that substantially modify treatment risk [30,31,32,33]. These findings are consistent with the present results, in which clinicians frequently diverged in their interpretation of model-generated therapeutic plans, reflecting the ambiguity or insufficient specificity of the LLM-provided text. More specifically, several recurring limitations were observed in model outputs. Although LLMs frequently proposed haemostatic strategies and, in some cases, specified factor levels or dosing schemes, these recommendations often lacked consistency and clinical integration. Factor target levels were not consistently aligned across comparable scenarios; key elements, such as inhibitor status and chemotherapy-induced thrombocytopenia, were insufficiently incorporated; and the interaction between oncologic treatments and bleeding risk was not systematically addressed. In addition, treatment regimens, while sometimes numerically detailed, were not consistently supported by a clear rationale or monitoring framework. These gaps substantially limit the clinical usability and reliability of the generated recommendations.
The challenges observed in these domains are amplified by the intrinsic complexity of hemophilia. Chemotherapy, immunotherapy, and radiotherapy may exacerbate bleeding risk, compromise mucosal integrity, or induce thrombocytopenia, all of which must be integrated with individualized factor replacement strategies and careful perioperative management [19,28,29]. The models rarely demonstrated a nuanced understanding of these interactions. They did not consistently acknowledge the need for specific factor VIII or IX correction targets for invasive procedures, nor did they reliably incorporate considerations of inhibitor status, the pharmacokinetic variability of replacement therapy, or the heightened risk of delayed bleeding following endoscopic or surgical procedures. Similarly, the models often failed to recognize scenarios in which certain antineoplastic agents might require modification, enhanced monitoring, or alternative regimens due to bleeding risk. The inability to operationalize these clinical subtleties underscores a fundamental limitation of current general-purpose models: they may articulate general principles correctly, but they struggle to translate them into domain-specific, clinically reliable, and operational therapeutic plans.
Differences between the two evaluated LLMs provide further insight into model behaviour. Copilot tended to produce outputs with more stable ratings across cases, whereas ChatGPT exhibited greater variability, at times generating responses judged as highly complete and at other times markedly insufficient. Such patterns are more plausibly explained by platform-level implementation differences, including system prompts, response filtering, and interface design, rather than differences in the underlying model architecture, as both systems are based on GPT-4. Studies comparing LLMs suggest that variability, hallucinations, and instability remain relevant issues in real-world applications [9,17,18]. In a clinical context such as hemophilia oncology, where safety is closely tied to precision, consistency, and reproducibility, this variability becomes an important limiting factor for real-world applicability.
The low inter-rater agreement observed in specific domains also warrants consideration. Disagreement between clinicians does not necessarily reflect purely subjective differences in clinical judgment; rather, it may also indicate that the material being evaluated lacks clarity or internal coherence. Several analyses of human–AI interaction underscore that variability in human ratings often signals deficiencies in the structure, factual basis, or interpretability of AI-generated content [34,35,36]. This study’s findings align with that interpretation. The clinicians converged when the model output was internally consistent and conceptually sound, as in Assessment and Strategy, but diverged sharply when the content became ambiguous or insufficiently specified, as seen in Selected Drug and Regimen. This pattern further emphasises that the quality of the output itself, rather than idiosyncrasies of clinician judgment, drove disagreement. The observed disagreement, particularly in the Drug domain, cannot be unambiguously attributed to model inadequacy. The study design did not include predefined reference standards for each case, and part of the variability may reflect differences in clinical judgment. However, the consistent pattern of high agreement in global domains and lower agreement in treatment-specific domains suggests that ambiguity and lack of specificity in LLM outputs likely contributed to the divergence between rates.
These results highlight a broader challenge in the use of LLMs for complex, high-risk medical decision-making. Although models trained on large corpora may internalize general medical knowledge, they do not reliably distinguish between authoritative and outdated sources, nor do they consistently adhere to disease-specific guidelines [10,11,12]. The nuanced management of hemophilia combined with malignancy requires mastery of perioperative coagulation targets, pharmacodynamic interactions, inhibitor management, scheduling of factor replacement around chemotherapy cycles, and anticipation of bleeding complications during radiation or invasive procedures. Without explicit fine-tuning on such specialized material, LLMs cannot be expected to produce consistently accurate or safe recommendations in this population.
Recent advances in retrieval-augmented generation and specialty fine-tuned medical models suggest pathways for future improvement. Models enhanced by structured clinical knowledge or trained on haematology–oncology datasets have demonstrated reductions in hallucinations and improvements in factual correctness [12,37,38]. However, until such systems are validated prospectively in complex multimorbidity scenarios, their deployment in real-world clinical settings should remain strictly supervised.
Methodologically, the present work demonstrates the importance of multidimensional evaluation frameworks that capture completeness, interpretability and inter-rater stability. Traditional accuracy metrics are inadequate for assessing free-text generative systems, particularly when clinical utility depends on a combination of factual integrity, contextual grounding and pragmatic applicability [24,25,26,27]. By integrating ordinal scoring, agreement metrics and case-level aggregation, this study provides a replicable template for evaluating emerging LLMs in narrow, safety-critical domains.
Notwithstanding its contributions, several limitations must be acknowledged. The rarity of hemophilia with concurrent malignancy inherently restricts sample size, limiting generalizability. The evaluation was performed by two expert clinicians, and broader panels of raters may reveal additional nuances in interpretability or reliability. The study also focused on text-only outputs, and future LLM iterations equipped with multimodal integration or real-time evidence retrieval may demonstrate different performance characteristics.
Taken together, the findings reveal that LLMs possess a meaningful capacity for high-level reasoning and synthesis in complex haematology–oncology contexts yet remain constrained by substantial deficiencies in the precise therapeutic domains that most affect patient safety. The capacity to articulate coherent strategic frameworks does not translate into reliable execution of drug selection, regimen design or perioperative hemostatic planning. These constraints are particularly consequential in patients with hemophilia undergoing cancer treatment, where small deviations from recommended therapeutic pathways may have profound clinical implications. Continued development of domain-specialised LLMs, alongside rigorous clinical validation and integration within supervised workflows, will be necessary before such systems can be considered suitable for supporting decision-making in this uniquely vulnerable population. The small sample size (six cases and two raters) limits the stability of agreement estimates and prevents strong generalization of the findings.

5. Conclusions

The present study demonstrates that while Large Language Models (LLMs) such as ChatGPT and Microsoft Copilot exhibit a significant capacity for high-level clinical reasoning and strategic synthesis in the management of complex hemophilia–oncology cases, they face critical limitations in providing precise, consistent, and clinically interpretable therapeutic details. The high reliability observed in global assessment and strategy contrasts sharply with the insufficient performance and lack of expert consensus in selecting specific drugs and treatment regimens. These findings highlight a persistent risk of ambiguity, inconsistency, and insufficient clinical specification in model-generated pharmacological advice. These findings refer to content-level reliability and should not be interpreted as a direct assessment of clinical safety.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/hemato7020010/s1, Table S1a: AI-generated clinical outputs for representative cases (ChatGPT). Condensed summaries of ChatGPT-generated recommendations across six hemophilia –oncology scenarios, reported by clinical domain (Decision and Rationale, Strategy, Selected Drug, Regimen). Outputs are presented in descriptive form to illustrate the structure and content of model responses. Table S1b: AI-generated clinical outputs for representative cases (Microsoft Copilot). Condensed summaries of Microsoft Copilot (GPT-4–based)-generated recommendations across the same six hemophilia –oncology scenarios, reported by clinical domain (Decision and Rationale, Strategy, Selected Drug, Regimen). Outputs are presented to enable qualitative comparison of model-generated clinical reasoning and therapeutic structuring.

Author Contributions

Conceptualization, A.P.; methodology, A.P. and F.M.; software, A.P.; validation, S.B. and E.Z.; formal analysis, A.P.; investigation, A.P.; resources, E.Z.; data curation, A.P. and F.M.; writing—original draft preparation, A.P.; writing—review and editing, A.P., S.P., F.M. and E.Z.; visualization, A.P.; supervision, F.M. and E.Z.; project administration, A.P.; funding acquisition, E.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding. The work of A.P. and S.B. was supported by grants funded by the Italian Ministry of Health (Ricerca Corrente).

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki. Ethical review and approval were waived for this study as it involved the retrospective evaluation of anonymized clinical cases already reported in the literature and did not involve direct intervention on human subjects.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study (summarized clinical cases and expert ratings) are available within the article text and tables.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT (OpenAI, GPT-4 architecture) and Microsoft Copilot (GPT-4-based) for the purposes of generating and evaluating clinical recommendations for the cases analyzed in the study. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Swift, A. The role of hematology in modern medicine: Insights into blood disorders and therapeutic strategies. Hematol. Blood Disord. 2024, 7, 185. [Google Scholar]
  2. Hoffman, R.; Benz, E.J.; Silberstein, L.E.; Heslop, H.E.; Weitz, J.I.; Anastasi, J.; Salama, M.E. Hematology: Basic Principles and Practice, 7th ed.; Elsevier: Philadelphia, PA, USA, 2018. [Google Scholar]
  3. Falanga, A.; Marchetti, M.; Vignoli, A. Coagulation and cancer: Biological and clinical aspects. J. Thromb. Haemost. 2013, 11, 223–233. [Google Scholar] [CrossRef] [PubMed]
  4. Rao, A.; Pang, M.; Kim, J.; Kamineni, M.; Lie, W.; Prasad, A.K.; Landman, A.; Dreyer, K.; Succi, M.D. Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow. medRxiv 2023. Update in J. Med. Internet Res. 2023, 25, e48659. https://doi.org/10.2196/48659. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  5. Li, R.; Wang, X.; Yu, H. Two Directions for Clinical Data Generation with Large Language Models: Data-to-Label and Label-to-Data. In Findings of the Association for Computational Linguistics: EMNLP 2023; Association for Computational Linguistics: Singapore, 2023; Volume 2023, pp. 7129–7143. [Google Scholar] [CrossRef]
  6. Kung, T.H.; Cheatham, M.; Medenilla, A.; Sillos, C.; De Leon, L.; Elepaño, C.; Madriaga, M.; Aggabao, R.; Diaz-Candido, G.; Maningo, J.; et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLoS Digit. Health 2023, 2, e0000198. [Google Scholar] [CrossRef] [PubMed]
  7. Gilson, A.; Safranek, C.W.; Huang, T.; Socrates, V.; Chi, L.; Taylor, R.A.; Chartash, D. Correction: How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment. JMIR Med. Educ. 2024, 10, e57594, Erratum for JMIR Med. Educ. 2023, 9, e45312. https://doi.org/10.2196/45312. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  8. Frieder, S.; Pinchetti, L.; Griffiths, R.-R.; Salvatori, T.; Lukasiewicz, T.; Petersen, P.; Chevalier, A.; Berner, J. Mathematical capabilities of ChatGPT. Adv. Neural Inf. Process. Syst. 2023, 36, 27699–27744. [Google Scholar] [CrossRef]
  9. Woesle, C.; Fischer-Brandies, L.; Buettner, R. A Systematic Literature Review of Hallucinations in Large Language Models. IEEE Access 2025, 13, 148231–148253. [Google Scholar] [CrossRef]
  10. Moor, M.; Banerjee, O.; Abad, Z.S.H.; Krumholz, H.M.; Leskovec, J.; Topol, E.J.; Rajpurkar, P. Foundation models for generalist medical artificial intelligence. Nature 2023, 616, 259–265. [Google Scholar] [CrossRef]
  11. Singhal, K.; Azizi, S.; Tu, T.; Mahdavi, S.S.; Wei, J.; Chung, H.W.; Scales, N.; Tanwani, A.; Cole-Lewis, H.; Pfohl, S.; et al. Large language models encode clinical knowledge. Nature 2023, 620, 172–180. [Google Scholar] [CrossRef]
  12. He, J.; Baxter, S.L.; Xu, J.; Xu, J.; Zhou, X.; Zhang, K. The practical implementation of artificial intelligence technologies in medicine. Nat. Med. 2019, 25, 30–36. [Google Scholar] [CrossRef]
  13. Cascella, M.; Montomoli, J.; Bellini, V.; Bignami, E. Evaluating the Feasibility of ChatGPT in Healthcare: An Analysis of Multiple Clinical and Research Scenarios. J. Med. Syst. 2023, 47, 33. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  14. Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large language models in medicine. Nat. Med. 2023, 29, 1930–1940. [Google Scholar] [CrossRef]
  15. Cross, J.L.; Choma, M.A.; Onofrey, J.A. Bias in medical AI: Implications for clinical decision-making. PLoS Digit. Health 2024, 3, e0000651. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  16. Alkaissi, H.; McFarlane, S.I. Artificial Hallucinations in ChatGPT: Implications in Scientific Writing. Cureus 2023, 15, e35179. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  17. Atil, B.; Chittams, A.; Fu, L.; Ture, F.; Xu, L.; Baldwin, B. LLM Stability: A detailed analysis with some surprises (Version 1). arXiv 2024, arXiv:2408.04667. [Google Scholar]
  18. Balagurunathan, Y.; Mitchell, R.; El Naqa, I. Requirements and reliability of AI in the medical context. Phys. Med. 2021, 83, 72–78. [Google Scholar] [CrossRef] [PubMed]
  19. Srivastava, A.; Brewer, A.K.; Mauser-Bunschoten, E.P.; Key, N.S.; Kitchen, S.; Llinas, A.; Ludlam, C.A.; Mahlangu, J.N.; Mulder, K.; Poon, M.C.; et al. Guidelines for the management of hemophilia. Hemophilia 2013, 19, e1–e47. [Google Scholar] [CrossRef] [PubMed]
  20. Tarniceriu, C.C.; Hurjui, L.L.; Tanase, D.M.; Haisan, A.; Tepordei, R.T.; Statescu, G.; Vicoleanu, S.A.P.; Lupu, A.; Lupu, V.V.; Ursaru, M.; et al. Inherited hemophilia—A multidimensional chronic disease that requires a multidisciplinary approach. Life 2025, 15, 530. [Google Scholar] [CrossRef]
  21. Koc, B.; Zulfikar, B. A Challenge for Hemophilia Treatment: Hemophilia and Cancer. J. Pediatr. Hematol. Oncol. 2021, 43, e29–e32. [Google Scholar] [CrossRef]
  22. Tagliaferri, A.; Di Perna, C.; Santoro, C.; Schinco, P.; Santoro, R.; Rossetti, G.; Coppola, A.; Morfini, M.; Franchini, M.; Italian Association of Hemophilia Centers. Cancers in patients with hemophilia: A retrospective study from the Italian Association of Hemophilia Centers. J. Thromb. Haemost. 2012, 10, 90–95. [Google Scholar] [CrossRef] [PubMed]
  23. Kempton, C.L.; Makris, M.; Holme, P.A. Management of comorbidities in hemophilia. Hemophilia 2021, 27, 37–45. [Google Scholar] [CrossRef] [PubMed]
  24. Sallam, M. ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns. Healthcare 2023, 11, 887. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  25. Kim, Y.; Jeong, H.; Chen, S.; Li, S.S.; Lu, M.; Alhamoud, K.; Park, C.; Mun, J.; Grau, C.; Jung, M.; et al. Medical hallucinations in foundation models and their impact on healthcare. arXiv 2025, arXiv:2503.05777. [Google Scholar]
  26. Gao, C.A.; Howard, F.M.; Markov, N.S.; Dyer, E.C.; Ramesh, S.; Luo, Y.; Pearson, A.T. Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. NPJ Digit. Med. 2023, 6, 75. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  27. Zitu, M.M.; Le, T.D.; Duong, T.; Haddadan, S.; Garcia, M.; Amorrortu, R.; Zhao, Y.; Rollison, D.E.; Thieu, T. Large language models in cancer: Potentials, risks, and safeguards. BJR Artif. Intell. 2025, 2, ubae019. [Google Scholar] [CrossRef]
  28. Rezende, S.M.; Neumann, I.; Angchaisuksiri, P.; Awodu, O.; Boban, A.; Cuker, A.; Curtin, J.A.; Fijnvandraat, K.; Gouw, S.C.; Gualtierotti, R.; et al. International Society on Thrombosis and Haemostasis clinical practice guideline for treatment of congenital hemophilia A and B based on the Grading of Recommendations Assessment, Development, and Evaluation methodology. J. Thromb. Haemost. 2024, 22, 2629–2652. [Google Scholar] [CrossRef]
  29. Hermans, C.; Apte, S.; Santagostino, E. Invasive procedures in patients with hemophilia: Review of low dose protocols and experience with extended half life FVIII and FIX concentrates and non replacement therapies. Hemophilia 2020, 27, 46–52. [Google Scholar] [CrossRef] [PubMed]
  30. Johnson, D.; Goodman, R.; Patrinely, J.; Stone, C.; Zimmerman, E.; Donald, R.; Chang, S.; Berkowitz, S.; Finn, A.; Jahangir, E.; et al. Assessing the Accuracy and Reliability of AI-Generated Medical Responses: An Evaluation of the Chat-GPT Model. Res. Sq. 2023. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  31. Lee, P.; Bubeck, S.; Petro, J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N. Engl. J. Med. 2023, 388, 1233–1239. [Google Scholar] [CrossRef] [PubMed]
  32. Mishra, H.P.; Gupta, R. Leveraging Generative AI for Drug Safety and Pharmacovigilance. Curr. Rev. Clin. Exp. Pharmacol. 2025, 20, 89–97. [Google Scholar] [CrossRef] [PubMed]
  33. Ali, A.M.A.; Alrobaian, M.M. Strengths and weaknesses of current and future prospects of artificial intelligence-mounted technologies applied in the development of pharmaceutical products and services. Saudi Pharm. J. 2024, 32, 102043. [Google Scholar] [CrossRef] [PubMed]
  34. Cabitza, F.; Rasoini, R.; Gensini, G.F. Unintended consequences of machine learning in medicine. JAMA 2017, 318, 517–518. [Google Scholar] [CrossRef]
  35. Ciobanu-Caraus, O.; Aicher, A.; Kernbach, J.M.; Regli, L.; Serra, C.; Staartjes, V.E. A critical moment in machine learning in medicine: On reproducible and interpretable learning. Acta Neurochir. 2024, 166, 14. [Google Scholar] [CrossRef] [PubMed] [PubMed Central]
  36. Jeong, C. Fine tuning and utilization methods of domain specific LLMs (Version 2). arXiv 2024, arXiv:2401.02981. [Google Scholar]
  37. Zhang, W.; Zhang, J. Hallucination mitigation for retrieval augmented large language models: A review. Mathematics 2025, 13, 856. [Google Scholar] [CrossRef]
  38. Xia, Y.; Zhou, J.; Shi, Z.; Chen, J.; Huang, H. Improving retrieval augmented language model with self reasoning (Version 3). arXiv 2024, arXiv:2407.19813. [Google Scholar]
Figure 1. Mean ordinal distance between the two clinicians by area. Bars represent the mean absolute difference between the two clinicians’ ordinal ratings (0 = Incomplete, 3 = Complete) for ChatGPT and Copilot. A value of 0 indicates perfect agreement, whereas a value of 1 indicates an average disagreement of one level on the four-point scale.
Figure 1. Mean ordinal distance between the two clinicians by area. Bars represent the mean absolute difference between the two clinicians’ ordinal ratings (0 = Incomplete, 3 = Complete) for ChatGPT and Copilot. A value of 0 indicates perfect agreement, whereas a value of 1 indicates an average disagreement of one level on the four-point scale.
Hemato 07 00010 g001
Figure 2. Mean perceived reliability score by area. Bars represent the mean aggregated scores of the two clinicians (0 = Incomplete, 3 = Complete) for ChatGPT and Copilot across all clinical cases in each area. Higher scores indicate that the information generated by the system was judged as more complete.
Figure 2. Mean perceived reliability score by area. Bars represent the mean aggregated scores of the two clinicians (0 = Incomplete, 3 = Complete) for ChatGPT and Copilot across all clinical cases in each area. Higher scores indicate that the information generated by the system was judged as more complete.
Hemato 07 00010 g002
Table 1. Summary table of the clinical cases reported in [21], describing patients with different hemophilia types and severities, age at cancer diagnosis, cancer type, HIV and HCV status, perioperative or treatment-related complications, therapeutic interventions and outcomes. All patients completed the diagnostic and therapeutic pathway with a favourable outcome.
Table 1. Summary table of the clinical cases reported in [21], describing patients with different hemophilia types and severities, age at cancer diagnosis, cancer type, HIV and HCV status, perioperative or treatment-related complications, therapeutic interventions and outcomes. All patients completed the diagnostic and therapeutic pathway with a favourable outcome.
Case [21]Hemophilia (Type and Severity)Age at Cancer Diagnosis (Years)Type of CancerHIV/HCV StatusComplicationsTreatmentOutcome
1A, Moderate14Acute Lymphoblastic Leukaemia
(ALL)
NegativeNoneChemotherapyAlive
2A, Severe31Thyroid CarcinomaNegativeNoneTotal thyroidectomy with radical cervical dissection, RadioiodineAlive
3A, Severe52Rectal CancerNegativeProlonged bleeding during surgeryLaparoscopic low-anterior resection of the rectum, Chemotherapy, RadiotherapyAlive
4A, Severe67Malignant MelanomaNegativeDisease progression, chemotherapy regimen changeRight forefinger amputation and axillary curettage, Interferon alpha 2-bAlive
5B, Moderate58Basal Cell CarcinomaNegativeNoneExcision of a skin lesionAlive
6B, Moderate79Gastric CancerNegativeNonePartial gastrectomyAlive
Table 2. Inter-rater agreement between the two clinicians for ChatGPT and Copilot across all content areas. Values are the number of ratings and the percentage of occasions in which the two clinicians assigned the same category for a given system.
Table 2. Inter-rater agreement between the two clinicians for ChatGPT and Copilot across all content areas. Values are the number of ratings and the percentage of occasions in which the two clinicians assigned the same category for a given system.
AreaNumber of Ratings% of Agreement for ChatGPT% of Agreement for Copilot
Assessment6100100
Decision and Rationale666.7100
Drug600
Regimen633.30
Strategy683.3100
Table 3. Inter-rater agreement between the two clinicians for the clinical case. The table reports, for each case, the number of evaluated areas, the percentage of areas in which the clinicians assigned the same category and the mean absolute difference on the 0–3 scale, separately for ChatGPT and Copilot.
Table 3. Inter-rater agreement between the two clinicians for the clinical case. The table reports, for each case, the number of evaluated areas, the percentage of areas in which the clinicians assigned the same category and the mean absolute difference on the 0–3 scale, separately for ChatGPT and Copilot.
CaseNumber of Areas% Agree ChatGPTDistance ChatGPT% Agree CopilotDistance Copilot
15400.6600.4
25201.0600.4
35600.4600.6
45800.2600.4
55800.2600.4
65600.4600.4
Table 4. Mean perceived reliability by clinical case. Values are the mean aggregated scores (0 = Incomplete, 3 = Complete) assigned by the two clinicians to ChatGPT and Copilot across all areas evaluated in each clinical case.
Table 4. Mean perceived reliability by clinical case. Values are the mean aggregated scores (0 = Incomplete, 3 = Complete) assigned by the two clinicians to ChatGPT and Copilot across all areas evaluated in each clinical case.
CaseScore ChatGPT CaseScore Copilot Case
11.51.8
21.12
322.1
42.32
51.92
622
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Porreca, A.; Proietti, S.; Maturo, F.; Bonassi, S.; Zanon, E. Clinical Reliability of Large Language Models in Complex Haematology: A Multidimensional Evaluation in Hemophilia–Oncology. Hemato 2026, 7, 10. https://doi.org/10.3390/hemato7020010

AMA Style

Porreca A, Proietti S, Maturo F, Bonassi S, Zanon E. Clinical Reliability of Large Language Models in Complex Haematology: A Multidimensional Evaluation in Hemophilia–Oncology. Hemato. 2026; 7(2):10. https://doi.org/10.3390/hemato7020010

Chicago/Turabian Style

Porreca, Annamaria, Stefania Proietti, Fabrizio Maturo, Stefano Bonassi, and Ezio Zanon. 2026. "Clinical Reliability of Large Language Models in Complex Haematology: A Multidimensional Evaluation in Hemophilia–Oncology" Hemato 7, no. 2: 10. https://doi.org/10.3390/hemato7020010

APA Style

Porreca, A., Proietti, S., Maturo, F., Bonassi, S., & Zanon, E. (2026). Clinical Reliability of Large Language Models in Complex Haematology: A Multidimensional Evaluation in Hemophilia–Oncology. Hemato, 7(2), 10. https://doi.org/10.3390/hemato7020010

Article Metrics

Back to TopTop