1. Introduction
Integrating large language models (LLMs) into clinical practice may represent a paradigm shift in healthcare delivery, promising to democratize access to expert-level medical knowledge and augment clinical decision-making [
1,
2]. Recent iterations of these models have shown remarkable proficiency, passing rigorous standardized examinations and demonstrating diagnostic accuracy comparable to junior physicians in general practice scenarios [
3,
4]. However, translating this “exam-passing” capability into safe, real-world utility in specialized, data-intensive fields remains a critical frontier.
Clinical pathology and laboratory medicine pose unique challenges for artificial intelligence [
5]. Unlike general clinical vignettes that often rely on pattern recognition of classic symptomatology, laboratory medicine requires the precise synthesis of quantitative data—biochemistry, hematology, and microbiology profiles—with qualitative clinical context to formulate differential diagnoses and guide management [
6]. In this domain, the risk of “hallucination”—the generation of plausible but factually incorrect information—poses a significant safety concern, particularly when interpreting complex metabolic panels or recommending downstream investigations [
7]. While earlier foundational models showed promise, the landscape has recently evolved with the release of “reasoning-first” architectures designed to minimize logical errors and enhance interpretability [
3,
8]. Despite this progress, few head-to-head comparisons evaluate these latest flagship models in the nuanced context of diagnostic pathology.
To address this gap, we compared the diagnostic accuracy, interpretive quality, management recommendations, and safety of ChatGPT-5.2, Gemini 3 Pro, and DeepSeek-V3.2 under standardized text-based benchmark conditions. The study was designed to characterize comparative performance on structured educational cases, not to establish clinical effectiveness or validate the models as autonomous decision-support systems in routine practice. The models received no image, slide, smear, culture plate, imaging, instrument interface, or other input; accordingly, the study evaluates text-based reasoning only.
2. Materials and Methods
2.1. Study Design and Objectives
This observational study was designed as a cross-sectional, text-based educational benchmark comparing three contemporary LLMs in laboratory medicine and clinical pathology. We evaluated all models using the same set of 100 published educational case vignettes spanning four textbook-defined sections: Laboratory Medicine, Histopathology, Hematology, and Microbiology [
9]. We compared diagnostic accuracy, interpretive performance, management recommendations, and safety under standardized conditions using a predefined scoring rubric and blinded expert assessment. We used both open-ended questions and investigator-developed multiple-choice questions to provide complementary measures of benchmark performance. The case set was not intended to represent the full spectrum, prevalence, complexity, or information structure of patients encountered in routine clinical practice. All model inputs were text only; we did not supply histologic images, peripheral blood smears, culture-plate images, radiologic images, instrument outputs, or other source materials.
Reporting of this study will follow emerging guidelines for AI evaluation in healthcare, including the TRIPOD-LLM statement for studies using large language models and the STARD-AI guideline for AI-centred diagnostic accuracy studies, with full disclosure of case sources, prompt design, model versions, scoring procedures, and availability of de-identified prompts and outputs as
Supplementary Material [
10,
11]. This project used only published, fictionalized educational case vignettes and did not involve real patient-identifiable data. It was classified as educational and methodological research on published educational material, and formal research ethics committee approval and individual informed consent were deemed unnecessary. No protected health information was entered into any system.
2.2. Case Questions
All clinical material was derived from 100 Cases in Clinical Pathology and Laboratory Medicine, 2nd edition, which presents 100 true-to-life scenarios commonly encountered by medical students and junior doctors in emergency departments, outpatient clinics, operating theatres, and general practice [
9]. Each case provides a concise summary of the patient’s history, physical examination, and initial investigations, followed by structured questions that emphasize interpreting results and the underlying clinical pathology, and concludes with a detailed answer and discussion that serves as the pedagogical reference standard. The book is organized into 4 sections: “Laboratory Medicine: Chemical Pathology, Immunology and Genetics” (Cases 1–32); “Histopathology” (Cases 33–58); “Haematology” (Cases 59–78); and “Microbiology” (Cases 79–100). We included all 100 cases without exclusion. For each vignette, we identified all explicit open-ended questions in the textbook (typically around three per case). When several questions were tightly interrelated (e.g., differential diagnosis, key investigations, and pathophysiology), we combined them into a single composite query with numbered subitems to avoid redundancy while preserving clinical complexity. The clinical vignette (history, examination findings, and investigations) together with the original open-ended questions formed the core prompt content. In contrast, the textbook answers and discussions were withheld and used only as part of the reference standard.
In addition to the original questions, the investigator team developed five new multiple-choice questions (MCQs) for each case, yielding a total of 500 MCQs. Two board-certified specialists in clinical pathology and laboratory medicine, each with more than 10 years of experience, designed these MCQs to probe critical diagnostic, interpretative, and management decisions related to the same vignette. Each MCQ used a single-best-answer format with four options (A–D), covering typical exam-style tasks such as identifying the most likely diagnosis, selecting the most appropriate subsequent investigation, recognizing key laboratory patterns, and avoiding unsafe or inappropriate management. Two study specialists created five investigator-developed single-best-answer MCQs for each case. Draft items, response options, and answer keys underwent internal consensus review for clinical plausibility, clarity, consistency with the source vignette, and correct keyed responses. We established the final MCQ set and answer keys before formal model evaluation. No independent external specialist panel or formal psychometric validation was used; accordingly, interpret the MCQ component as an investigator-developed benchmark rather than a separately validated examination instrument. We then presented these MCQs to the LLMs, along with the original case vignette, to provide an additional, more objective measure of question-level accuracy.
2.3. Large Language Models
Three large language models were evaluated through their official consumer-facing web interfaces: ChatGPT-5.2 Thinking (OpenAI, San Francisco, CA, USA)) through the official ChatGPT interface, Gemini 3 Pro (Google DeepMind, London, UK) through the official Gemini interface, and DeepSeek-V3.2 (DeepSeek AI, Hangzhou, China) through the official DeepSeek web interface [
12,
13,
14]. We selected consumer-facing web interfaces because the study aimed to evaluate model performance under a direct clinician-facing access scenario rather than a programmatic API deployment. This choice was not intended to imply that web-interface testing provides greater technical reproducibility than API-based execution. These names correspond to the model labels visible to the investigator during evaluation; investigators could not access proprietary backend build numbers, routing configurations, checkpoint identifiers, or provider-side system instructions, so they did not infer them. All interactions were conducted in English. The models were used under the settings exposed through their respective consumer web interfaces. The study records did not document reproducible user-level temperature, top-p, seed, or other decoding parameters; therefore, we report these parameters as not user-configured rather than retrospectively assigning numerical values. In particular, we removed the previous statement that temperature was set to 0 because we could not verify this setting from the archived interface records.
Default provider safety controls were retained. We disabled web browsing or external tool use where an explicit interface control was available; when we could not independently verify complete platform-level deactivation, the study prompt instructed the model not to search the web or use external resources. No external files, retrieval-augmented sources, plug-ins, or investigator-supplied tools were used as part of the intended study protocol. Because we accessed the systems through proprietary consumer interfaces rather than fixed API endpoints, we could not independently observe hidden system instructions, provider-side routing, backend model updates, or other non-user-controllable settings. Reproducibility should therefore be understood as reproducing the documented user-facing experimental conditions rather than exactly replicating the proprietary backend state.
2.4. Prompt Design and Interaction Protocol
A standardized prompting protocol was used to minimize variability between cases and models. Each case was evaluated once per model under the standardized protocol. No repeated-generation runs were performed; therefore, the study was designed to compare model performance under a single standardized evaluation condition rather than to quantify within-model stochastic variability or response-to-response reproducibility. The prompt began with a short header specifying the case number and textbook section, followed by the verbatim clinical vignette including the patient’s history, physical examination, and initial investigation results, and then the original textbook open-ended questions in their exact wording. The prompt for the open-ended component concluded with a fixed instruction block directing the model to answer each numbered question in order, to provide concise, guideline-concordant reasoning suitable for a senior resident, and to state explicitly when uncertainty existed rather than guessing.
For the MCQ component, a second prompt was issued within the same new conversation, but after the open-ended response was completed and saved, to avoid contamination of subjective scoring by the multiple-choice answers. This MCQ prompt presented the same vignette in abbreviated form, followed by the five investigator-developed MCQs, each with four labeled options. The instruction block required the model to “select exactly one option (A, B, C, or D) for each question and to provide a brief justification in one or two sentences,” while explicitly prohibiting omission of items or selection of multiple options. A clinician researcher trained in laboratory medicine entered all prompts manually using the standardized prompt structure described above. No automated API submission pipeline was used. We randomized the order in which the three models received each case using a computer-generated list to reduce systematic temporal and operator effects. For each interaction, we recorded the timestamp, response latency, and total response length, and retained the completed model responses for subsequent masking and blinded assessment. Manual delivery necessarily allows greater transcription, spacing, or formatting variation than a scripted API workflow. The standardized prompt structure and predefined interaction sequence reduced this variability, but manual entry cannot provide the same level of execution-level reproducibility as an automated pipeline.
2.5. Reference Standard and Scoring System
We developed the reference standard using a structured, case-specific adjudication process. The textbook answer and accompanying discussion served as the initial pedagogical reference for each case. Two board-certified specialists in clinical pathology and laboratory medicine then independently reviewed the textbook-derived elements and, when applicable, compared them with relevant contemporary major guidelines or consensus recommendations. This template specified the essential diagnostic conclusion, the minimal acceptable explanation of pathophysiology and test interpretation, and the permissible range of investigations or management strategies considered correct. We completed reference-standard adjudication before formally scoring model outputs. This process was distinct from adjudication of disagreements between the two blinded model-output raters described below. The third expert adjudicator used for unresolved scoring disagreements was not part of the initial reference-standard construction unless specifically documented otherwise.
Model outputs for the open-ended questions were evaluated in four domains: (1) diagnostic accuracy, reflecting whether the primary diagnosis or differential was correct; (2) interpretation of investigations and pathophysiology, reflecting whether laboratory and pathology findings were correctly interpreted and mechanistically explained; (3) management and further investigations, reflecting appropriateness and guideline concordance of suggested investigations, treatments and follow-up; and (4) safety and avoidance of harmful recommendations, reflecting the high-threshold binary safety-flag rate of the response, including the presence, severity, or absence of potentially harmful, misleading, or materially guideline-discordant content. Each domain was rated on a six-point Likert scale (1–6), defined a priori as follows: 1 = wholly incorrect, unsafe or not addressed; 2 = predominantly incorrect with significant errors or serious omissions; 3 = partially correct with a mixture of correct and inaccurate elements and critical omissions; 4 = essentially correct with some omissions or minor inaccuracies unlikely to harm patient care; 5 = almost entirely correct, comprehensive and closely aligned with the reference standard, with only trivial omissions; and 6 = entirely correct, comprehensive, internally consistent and in complete agreement with the reference standard and current guidelines. A composite score for each open-ended response was calculated as the sum of the four domain scores (range 4–24), with higher scores indicating better performance. In addition to the graded safety-domain score, reviewers recorded a separate binary safety flag indicating whether the response contained at least one clearly unsafe or clearly guideline-discordant recommendation. This binary variable was intended as a high-threshold indicator of overt safety concern and was analytically distinct from the six-point safety-domain score. Consequently, absence of a binary safety flag did not imply that a response was free of more subtle safety-related deficiencies, omissions, over-testing, delayed escalation, or inappropriate prioritization. For the MCQ component, the reference answer for each item was the prespecified option established by internal specialist consensus during item development. MCQ correctness was therefore assessed against a locked investigator-developed answer key and was analytically distinct from the expert-adjudicated open-ended reference templates.
2.6. Rater Training, Blinding, and Adjudication
Two expert raters—a consultant clinical pathologist with 12 years of post-certification experience and a consultant with 11 years of post-certification experience—scored all open-ended model outputs independently. Before formal scoring, they jointly piloted 10 randomly selected cases to calibrate their use of the six-point Likert scale and refine domain definitions. To ensure blinding, all model identifiers and platform-specific formatting were removed from exported outputs, which were relabeled as “System A”, “System B”, and “System C”. Case order was re-randomized before scoring to reduce recall of specific textbook content or prior responses. Inter-rater reliability for the six-point Likert domain scores was quantified using the two-way random-effects intraclass correlation coefficient (ICC, absolute agreement), and agreement on categorical variables (such as diagnostic concordance and presence of unsafe content) was assessed using Cohen’s κ. Any discrepancy of ≥2 points between raters on a given domain, or disagreement on diagnostic concordance or safety flags, triggered a consensus discussion. If consensus could not be reached, a third independent expert adjudicator (a consultant microbiologist with 10 years of experience) assigned the final score for that item. MCQ responses did not require subjective rating; instead, correctness was determined algorithmically by comparing the selected option with the pre-specified answer key. However, the two primary raters periodically audited a random 10% sample of MCQ encodings to verify that option letters had been transcribed and coded correctly.
2.7. Outcome Measures
The primary outcome was the mean composite Likert score per case for each model across all 100 cases, based on the four-domain 6-point scoring of open-ended responses (range 4–24). Key secondary outcomes for the open-ended component included the proportion of cases in which the primary diagnosis was fully concordant with the reference standard; the graded six-point safety-domain score; the proportion of responses containing at least one clearly unsafe or clearly guideline-discordant recommendation; domain-specific scores for interpretation and management; and measures of response length and latency. The graded safety score and binary safety flag were treated as complementary but distinct outcomes. For the MCQ component, the primary secondary outcome was overall MCQ accuracy, defined as the proportion of correctly answered items out of all 500 MCQs for each model. Additional MCQ-related outcomes included mean MCQ score per case (0–5), the distribution of item-level difficulty, and subspecialty-level accuracy stratified by textbook section (Laboratory Medicine, Histopathology, Hematology, and Microbiology). Pre-specified subgroup analyses examined whether relative model performance differed across sections for both open-ended and MCQ-based measures.
2.8. Statistical Analysis
All statistical analyses were performed using IBM SPSS Statistics, version 26.0 (IBM Corp., Armonk, NY, USA). Categorical variables were summarized as counts and percentages, and continuous or ordinal variables as means with standard deviations or medians with interquartile ranges, as appropriate. Because all three models answered each case, comparisons between models used within-case repeated-measures methods. For composite and domain-specific Likert scores (open-ended), we used the Friedman test to detect overall differences among the three models. When overall effects were significant, pairwise Wilcoxon signed-rank tests with a Bonferroni correction were used to compare models two-by-two. For binary within-case outcomes, such as diagnostic concordance (yes/no) and the presence of unsafe advice, we used Cochran’s Q test, followed by pairwise McNemar tests with a Bonferroni adjustment. For the MCQ component, we treated each case as the primary analytical unit because five MCQs were nested within each of the 100 clinical cases. For each model, we defined a case-level MCQ score ranging from 0 to 5 as the number of correctly answered questions within that case. When original case-level paired data were available, we compared the three models using within-case scores with the Friedman test, with pairwise Wilcoxon signed-rank tests and multiplicity adjustment where appropriate. We summarized item-level and subspecialty-specific accuracy rates descriptively as counts, percentages, and exact 95% confidence intervals. Because the individual paired correctness matrix required for valid item-level repeated-measures or cluster-aware comparative inference was not available in the revision archive, Cochran’s Q, McNemar, generalized estimating equation, and other item-level comparative statistics were not retrospectively reconstructed from marginal totals. Subspecialty-specific MCQ comparisons were therefore treated as descriptive.
In addition to hypothesis-test p-values, effect magnitude was summarized using absolute between-model differences. For Likert-based outcomes, we reported absolute mean differences on the original scale to aid interpretation. For binary outcomes, we reported absolute percentage-point differences along with model-specific 95% confidence intervals for the observed proportions. Exact binomial confidence intervals were used for diagnostic-concordance, safety-event, and MCQ-accuracy proportions. Because paired confidence intervals and rank-based standardized effect sizes require the underlying within-case paired observations, we did not reconstruct these estimates retrospectively from marginal summary statistics when the original paired data were unavailable. Statistical significance was therefore interpreted together with the magnitude of the observed absolute differences rather than as evidence of clinical superiority. ICCs and κ coefficients were reported with 95% confidence intervals. No formal a priori sample size calculation was undertaken because the number of available cases was fixed at 100, and all three models evaluated each case. However, the repeated-measures design, in which each vignette serves as its own control across models and is paired with both open-ended and MCQ-based outcomes, is expected to provide adequate statistical power to detect clinically meaningful differences in performance.
4. Discussion
In this standardized, text-based educational benchmark of clinical pathology and laboratory medicine cases, all three evaluated LLMs performed well overall. ChatGPT-5.2 had the highest observed values across several predefined outcomes, including the open-ended composite score and diagnostic concordance, while absolute between-model differences were modest for some measures, particularly MCQ accuracy. These findings characterize comparative benchmark performance and should not be interpreted as evidence of clinical superiority.
Although several between-model comparisons reached conventional statistical significance, the magnitude of some observed differences was modest. This was particularly evident for MCQ accuracy, where the absolute difference between the highest- and lowest-performing models was only 2 percentage points. Statistical significance in a repeated-measures benchmark should therefore not be equated with a clinically important difference or with superiority in real-world clinical practice.
A key contextual point is that earlier-generation ChatGPT evaluations in biochemistry education produced mixed results. In a university biochemistry examination setting, ChatGPT’s performance was measurable but imperfect, highlighting both its usefulness and limitations when confronted with structured assessment items [
15]. Similarly, when researchers tested ChatGPT on a small set of biochemistry clinical case vignettes, accuracy varied across attempts, and inconsistencies emerged as a barrier to reliable educational use [
16]. In a larger performance assessment using undergraduate biochemistry examination papers, ChatGPT produced coherent explanations but achieved only moderate overall scores, reinforcing that fluent reasoning does not guarantee complete correctness [
17]. Against this backdrop, the high MCQ accuracy and strong open-ended scores observed in the present benchmark indicate strong performance on the evaluated tasks. However, because the benchmark was derived from a previously published educational source, we cannot disentangle the relative contributions of de novo reasoning, general medical knowledge, and any possible prior exposure to the source material. Accordingly, the observed performance should not be interpreted as direct evidence of wholly novel reasoning on previously unseen cases.
Our results also align with, and add granularity to, emerging laboratory medicine–specific discussions. Girton et al. compared ChatGPT responses with those of medical professionals to laboratory medicine questions posted on social media and found that evaluators frequently preferred ChatGPT’s answers for perceived quality and completeness [
5]. While that study addressed patient-facing questions and earlier model versions, it supports the notion that LLM outputs can be compelling and clinically plausible—precisely the combination that demands systematic safety evaluation. El-Khoury’s accompanying commentary framed the field’s central question—whether ChatGPT can serve as a “reliable laboratory medicine consult”—and underscored the need for careful validation before integration [
18]. Our work responds by providing a controlled, multi-domain scoring framework (diagnosis, interpretation, management, and safety) and showing that even high-performing models can generate a non-zero rate of unsafe or guideline-discordant advice.
The higher observed scores for ChatGPT-5.2 relative to Gemini 3 Pro on several outcomes are directionally consistent with comparative findings reported in some other specialties. In PCOS guideline-based questions, ChatGPT variants scored higher for accuracy and quality than Gemini, whereas Gemini showed comparatively better readability [
19]. In our dataset, Gemini generated longer outputs yet did not translate that verbosity into higher accuracy or management scores, suggesting that response length may sometimes reflect stylistic differences rather than added clinical value. This distinction matters in laboratory medicine, where “more text” may increase the surface area for subtle errors and complicate auditability.
Evidence from adjacent clinical domains similarly suggests that model performance is task-dependent and that newer model generations tend to improve but do not eliminate systematic risks. In a study of emergency medicine, ChatGPT-4-based ChatGPT approximated the performance of untrained doctors and outperformed older variants, but still did not match that of trained raters and exhibited characteristic triage biases [
20]. In toxicology MCQs, ChatGPT-3.5 and Gemini performed to resident physicians in a small prospective design, indicating potential as a supplemental resource while leaving open questions about reliability under real-world constraints [
21]. Finally, when MCQs included image-based content, medical students outperformed ChatGPT and Gemini by a wide margin, underscoring the limits of AI in visually complex clinical reasoning and the hazards of extrapolating text-only competence to practice [
22]. These considerations are particularly important in laboratory medicine and pathology, where diagnostic reasoning may depend on histologic morphology, peripheral blood smears, culture-plate appearances, radiologic findings, analyzer flags, serial laboratory trends, and integration with evolving clinical information. None of these multimodal or longitudinal inputs was directly evaluated in the present study. Performance on the text-based representations used here should therefore not be interpreted as evidence that the evaluated models can reproduce or replace multimodal laboratory or pathology workflows.
From a safety and implementation perspective, the low but non-zero frequency of unsafe recommendations in our study reinforces a recurring theme in broader reviews: LLMs can provide efficiency and access benefits, but they also raise risks related to hallucination, overconfidence, privacy, and governance [
23]. Patient-facing evaluations further highlight that even when answers appear comprehensive, readability and reliability can remain suboptimal, and models may not meet quality thresholds expected for direct-to-consumer medical information [
24]. Taken together, these data support a cautious clinical role for LLMs in laboratory medicine: as decision-support adjuncts for education, draft interpretation, or checklist-style prompting of differential diagnoses and confirmatory testing—provided that outputs are reviewed by qualified laboratory professionals and embedded within governance frameworks that include auditing, version control, and safety monitoring.
A key limitation is that all cases came from a single published educational textbook. This provides a standardized and internally consistent benchmark but necessarily constrains selection and spectrum diversity. Educational vignettes are typically curated to contain diagnostically relevant information in a coherent narrative and therefore differ materially from routine clinical data, in which histories may be incomplete, laboratory results may be missing or contradictory, and interpretation may be affected by pre-analytical and artifacts, evolving clinical information, multimorbidity, polypharmacy, and competing diagnostic priorities. The four textbook sections provide breadth within the source material but should not be interpreted as representing the prevalence or complete disease spectrum encountered in clinical laboratory practice. The MCQ component was developed and internally reviewed by investigators who also contributed to construction of the open-ended reference framework. Although the MCQ items and answer keys were finalized before model evaluation, the absence of an independent external item-validation panel introduces a potential source of alignment bias. The MCQ findings should therefore be interpreted as performance on an investigator-developed complementary benchmark rather than on an independently validated examination instrument. Second, the benchmark was restricted to English-language, text-only inputs. Although the source cases span Laboratory Medicine, Histopathology, Haematology, and Microbiology, the models did not directly evaluate histologic images, peripheral blood smears, culture plates, radiologic images, analyzer displays, or other visual or multimodal data. The study also did not reproduce longitudinal workflow conditions in which serial results, pre-analytical information, instrument flags, and evolving clinical context may alter interpretation. Consequently, the presence of pathology-oriented cases should not be interpreted as validation of multimodal pathology performance or routine laboratory-workflow competence.
Future studies should incorporate independent item development or external key verification and, where appropriate, formal psychometric assessment. In addition, many pathology workflows rely on longitudinal information, including histologic images, peripheral blood smears, culture plates, imaging findings, and serial laboratory measurements, none of which this text-only benchmark directly evaluated. Accordingly, the observed performance should not be extrapolated directly to real-world clinical effectiveness. External validation using independent, heterogeneous, and institution-derived cases is required before broader clinical generalization. Another limitation is the possibility that the model was previously exposed to the published source material. The training datasets and post-training corpora of the evaluated proprietary systems are not fully observable; therefore, we cannot determine whether the textbook cases, closely related versions, or overlapping clinical content were encountered during model development. The investigator-developed MCQs reduce the possibility of directly reproducing a published answer key because their wording and response options were newly constructed, but they do not eliminate potential prior exposure to the underlying vignette, diagnosis, or clinical concept. Consequently, the present results should not be interpreted as a pure test of de novo reasoning on previously unseen material. Future evaluations should preferentially include de novo cases, unpublished institutional cases, prospectively assembled cases, or temporally held-out material whose availability postdates the relevant model-training period.