Next Article in Journal
Association Between Sarcopenic Obesity–Related Scores and Liver Fibrosis in Patients with Steatotic Liver Disease: A Cross-Sectional Study
Previous Article in Journal
Segmentation-Based Multi-Class Detection and Radiographic Charting of Periodontal and Restorative Conditions on Bitewing Radiographs Using Deep Learning
 
 
Article
Peer-Review Record

Comparative Performance of Multimodal and Unimodal Large Language Models Versus Multicenter Human Clinical Experts in Aortic Dissection Management

Diagnostics 2026, 16(2), 323; https://doi.org/10.3390/diagnostics16020323
by Evren Ekingen 1 and Mete Ucdal 2,*
Reviewer 1:
Reviewer 2:
Diagnostics 2026, 16(2), 323; https://doi.org/10.3390/diagnostics16020323
Submission received: 13 December 2025 / Revised: 11 January 2026 / Accepted: 16 January 2026 / Published: 19 January 2026
(This article belongs to the Section Clinical Diagnosis and Prognosis)

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

This manuscript presents an ambitious and timely comparison between multimodal large language models (MLLMs), a unimodal large language model (ChatGPT-5.2), and human clinical experts in the management of aortic dissection. The topic is highly relevant given the rapid expansion of generative AI in clinical decision support, and the inclusion of multiple medical specialties across several centers is a notable strength. The authors provide extensive methodological detail, particularly regarding AI prompting strategies, consensus algorithms, and domain-specific analyses, which enhances transparency and reproducibility.

Despite these strengths, several important methodological and interpretative limitations merit attention. First, the small number of test items (n = 25) substantially limits the robustness and generalizability of the conclusions. While the reported accuracy rates are high, performance differences of one or two questions translate into large percentage changes, making statistical comparisons underpowered and potentially misleading. The lack of statistically significant differences between groups should therefore be interpreted with caution and does not necessarily imply true equivalence between AI systems and human experts.

Second, the study relies exclusively on multiple-choice questions derived from examination-style resources rather than real-world clinical cases. This format favors recall of guideline-based knowledge and pattern recognition rather than dynamic clinical reasoning under uncertainty. Consequently, the findings may overestimate AI performance relative to actual bedside decision-making, where incomplete data, time pressure, and ethical considerations play a major role.

Third, although the multimodal architecture is a central focus of the study, the diagnostic advantage of the MLLM is difficult to disentangle from the use of standardized radiology descriptions for the unimodal model. Since both AI systems achieved perfect diagnostic accuracy, the incremental value of true image-based reasoning remains insufficiently demonstrated. A direct comparison using raw imaging input for humans and text-only input for AI would further clarify the clinical relevance of multimodality.

Additionally, the consensus-based MLLM decision strategy introduces an important source of error. The finding that majority voting reduced accuracy in cases of model disagreement highlights a critical limitation of ensemble approaches in safety-critical domains. However, this issue is underemphasized in the discussion, and alternative strategies such as weighted voting or confidence-based abstention are not explored.

Finally, while the authors appropriately acknowledge limitations, the conclusion that AI systems have achieved “human-equivalent performance” may be overstated. Given the small sample size, controlled testing environment, and absence of outcome-based validation, a more cautious framing emphasizing decision support rather than equivalence or replacement would be more appropriate.

In summary, this study contributes valuable insights into the comparative performance of multimodal and unimodal LLMs in a high-risk cardiovascular condition. However, stronger evidence from larger datasets, real-world clinical scenarios, and prospective validation is required before definitive claims regarding clinical equivalence or deployment readiness can be justified.

Author Response

RESPONSE TO REVIEWER 1

We deeply appreciate Reviewer 1's comprehensive and scholarly evaluation of our manuscript. The reviewer's recognition of the study's timeliness, methodological transparency, and the inclusion of multiple medical specialties across several centers is encouraging. We have carefully addressed each of the important methodological and interpretative limitations raised.

Comment 1.1: Sample Size Limitations

Reviewer's Comment: "The small number of test items (n = 25) substantially limits the robustness and generalizability of the conclusions. While the reported accuracy rates are high, performance differences of one or two questions translate into large percentage changes, making statistical comparisons underpowered and potentially misleading. The lack of statistically significant differences between groups should therefore be interpreted with caution and does not necessarily imply true equivalence between AI systems and human experts."

Response:

We sincerely thank the reviewer for this critical observation regarding statistical power and the interpretation of non-significant findings. We fully acknowledge that the sample size of 25 questions represents an important limitation that affects the precision of our estimates and the power to detect meaningful differences between groups.

We have substantially revised the manuscript to address this concern in several ways. First, we enhanced our statistical reporting by adding 95% confidence intervals for all accuracy estimates to transparently communicate the uncertainty associated with our point estimates. The wide confidence intervals (e.g., MLLM: 74.0%-99.0%; ChatGPT-5.2: 79.6%-99.9%) now explicitly demonstrate the imprecision inherent in our sample size. Second, we conducted a post-hoc power analysis using G*Power 3.1 to quantify the study's statistical limitations. With n=25 questions and observed accuracy rates ranging from 89.3% to 96.0%, our study had approximately 15-25% power to detect a 10% absolute difference in accuracy at α=0.05. This analysis confirms that the study was underpowered to detect clinically meaningful differences. Third, we have modified our interpretation throughout the manuscript to emphasize that non-significant p-values should not be interpreted as evidence of equivalence. We now explicitly state that statistical non-significance in our context reflects insufficient power rather than confirmed equivalence. Finally, we have substantially expanded the limitations section to comprehensively address the statistical constraints imposed by our sample size.

MANUSCRIPT ADDITIONS:

Added to Section 2.9 (Statistical Analysis):

"Post-hoc power analysis was conducted using G*Power 3.1 to evaluate the study's ability to detect meaningful performance differences. With n=25 questions and observed accuracy rates, the study had approximately 15-25% power to detect a 10% absolute difference in accuracy between groups (α=0.05, two-tailed). This analysis indicates that the study was exploratory in nature, and non-significant results should be interpreted as inconclusive rather than as evidence of equivalence. Effect sizes were calculated using Cohen's h for proportion comparisons to facilitate future meta-analyses and sample size planning."

Added to Section 4 (Discussion), Limitations paragraph:

" The sample size of 25 questions, while representative of key clinical domains, substantially limits the statistical power and generalizability of our findings. Performance differences of one or two questions translate into percentage changes of 4-8%, which may be clinically meaningful but remained undetectable given our sample constraints. The non-significant p-values observed across comparisons should not be interpreted as evidence of true equivalence between AI systems and human experts. Rather, these findings reflect the study's exploratory nature and insufficient power to detect potentially meaningful differences. Future studies with substantially larger question sets (minimum n=100-200 items based on our power calculations) are necessary to provide definitive evidence regarding performance equivalence or superiority."

Comment 1.2: Multiple-Choice Question Format Limitations

Reviewer's Comment: "The study relies exclusively on multiple-choice questions derived from examination-style resources rather than real-world clinical cases. This format favors recall of guideline-based knowledge and pattern recognition rather than dynamic clinical reasoning under uncertainty. Consequently, the findings may overestimate AI performance relative to actual bedside decision-making, where incomplete data, time pressure, and ethical considerations play a major role."

Response:

We greatly appreciate this astute observation regarding the ecological validity of our assessment methodology. The reviewer correctly identifies that standardized multiple-choice questions, while offering objectivity and reproducibility, differ substantially from real-world clinical decision-making environments.

We acknowledge that the MCQ format has inherent limitations including: (1) provision of complete, curated clinical information rather than the fragmented data typical of clinical encounters; (2) absence of time pressure and cognitive load present in emergency settings; (3) structured response options that eliminate the need for spontaneous diagnostic hypothesis generation; and (4) lack of patient interaction, ethical dilemmas, and interpersonal communication demands.

MANUSCRIPT ADDITIONS:

Added to Section 4 (Discussion), Limitations paragraph:

"An important methodological consideration is the exclusive reliance on standardized multiple-choice questions rather than real-world clinical scenarios. This examination-style format inherently favors recall of guideline-based knowledge and pattern recognition over dynamic clinical reasoning under uncertainty. Real-world aortic dissection management involves incomplete and evolving clinical data, severe time constraints, multi-stakeholder communication demands, and complex ethical considerations that cannot be captured in structured assessment formats. Consequently, our findings may overestimate AI performance relative to actual bedside decision-making. The controlled testing environment eliminates variables such as cognitive load from simultaneous patient care responsibilities, the need for information gathering and synthesis from multiple sources, and the integration of patient preferences into shared decision-making. Future research should prioritize prospective evaluation in authentic clinical environments using simulation-based assessments, standardized patient encounters, or retrospective chart review with outcome validation to better characterize real-world AI performance."

Comment 1.3: Multimodal Advantage Assessment

Reviewer's Comment: "Although the multimodal architecture is a central focus of the study, the diagnostic advantage of the MLLM is difficult to disentangle from the use of standardized radiology descriptions for the unimodal model. Since both AI systems achieved perfect diagnostic accuracy, the incremental value of true image-based reasoning remains insufficiently demonstrated."

Response:

We thank the reviewer for this methodologically important critique. The reviewer correctly identifies a fundamental confound in our study design: by providing standardized radiological descriptions to both unimodal models and human participants, we cannot isolate the unique contribution of direct image processing capabilities.

MANUSCRIPT ADDITIONS:

Added to Section 4 (Discussion), new paragraph:

"A notable methodological limitation concerns the assessment of multimodal capabilities. Although the MLLM's multimodal architecture is a central focus, the diagnostic advantage of direct image processing could not be definitively demonstrated in our study design. Both AI systems achieved perfect diagnostic accuracy (100%), precluding differentiation of their capabilities in this domain. The provision of standardized, expert-generated radiological descriptions to ChatGPT-5.2 and to the text-only MLLM components (Med-PaLM 2, BioGPT) may have compensated for the absence of direct image analysis, effectively equalizing information available across systems. To rigorously isolate the incremental value of true image-based reasoning, future studies should employ factorial designs comparing: (1) raw imaging input without text descriptions, (2) text-only descriptions without images, (3) combined image and text input, and (4) deliberately degraded or ambiguous image quality conditions where visual interpretation becomes critical."

Comment 1.4: Consensus-Based MLLM Decision Strategy Limitations

Reviewer's Comment: "The consensus-based MLLM decision strategy introduces an important source of error. The finding that majority voting reduced accuracy in cases of model disagreement highlights a critical limitation of ensemble approaches in safety-critical domains. However, this issue is underemphasized in the discussion, and alternative strategies such as weighted voting or confidence-based abstention are not explored."

Response:

We sincerely appreciate this insightful critique regarding the consensus methodology employed in our MLLM system. The reviewer correctly identifies that our majority voting approach demonstrated a critical limitation with important implications for clinical implementation.

The majority voting consensus strategy was selected based on established principles in ensemble machine learning, where combining predictions from multiple models typically improves overall accuracy and reduces individual model biases. This approach has been successfully applied in various medical AI applications, including diagnostic imaging interpretation and clinical decision support systems. Furthermore, the majority voting method offers transparency and interpretability advantages over more complex ensemble techniques, which is particularly important in clinical settings where decision rationale must be explainable to healthcare providers.

However, as the reviewer astutely observes, our results revealed an important limitation of this approach in safety-critical domains. While unanimous agreement among component models (GPT-4V, Med-PaLM 2, BioGPT) yielded 100% accuracy (21/21 questions), majority agreement in disagreement cases yielded only 50% accuracy (2/4 questions). This finding is consistent with recent literature suggesting that simple ensemble methods may paradoxically suppress correct responses from the most accurate individual model when component models exhibit heterogeneous domain-specific performance. The blood pressure target question exemplifies this phenomenon: GPT-4V correctly identified the guideline-concordant target, but was overruled by the incorrect majority response from Med-PaLM 2 and BioGPT.

We acknowledge that alternative ensemble strategies warrant exploration in future implementations. Weighted voting approaches, where model contributions are scaled according to domain-specific validation performance, have demonstrated superior accuracy in heterogeneous ensemble systems. Confidence-calibrated methods incorporating prediction uncertainty have shown promise in medical AI applications where decision confidence correlates with accuracy.Additionally, abstention protocols that flag low-confidence or disagreement cases for human review align with emerging frameworks for human-AI collaboration in healthcare.

We have substantially expanded the discussion to address this limitation and to propose alternative strategies for future research.

MANUSCRIPT ADDITIONS:

Added to Section 4 (Discussion), new paragraph on consensus methodology:

 

‘’ The consensus-based MLLM decision strategy warrants critical examination as a potential source of systematic error. The majority voting approach was selected based on established ensemble learning principles, where combining predictions from multiple models typically improves overall accuracy and reduces individual model biases (20,21). However, our finding that majority voting reduced accuracy from 100% (unanimous agreement) to 50% (disagreement cases) highlights a fundamental limitation of simple ensemble approaches in safety-critical clinical domains. This phenomenon, where ensemble methods paradoxically suppress correct responses from the most accurate individual model, has been documented in heterogeneous AI systems with varying domain-specific expertise (22). Several alternative ensemble strategies merit consideration for future implementations: First, weighted voting based on domain-specific validation performance could assign greater influence to models with demonstrated expertise in particular clinical areas. Second, confidence-calibrated voting could incorporate each model's output probability distributions rather than binary selections, enabling uncertainty quantification. Third, abstention protocols could flag disagreement cases for mandatory human review rather than forcing potentially unreliable automated decisions. Fourth, hierarchical decision frameworks could route questions to specialized models based on domain classification before consensus determination. The clinical implications are significant: in safety-critical applications, ensemble AI systems should incorporate disagreement detection as a trigger for human oversight rather than autonomous resolution, consistent with emerging human-AI collaboration frameworks in healthcare (23).’’

Comment 1.5: Interpretation of "Human-Equivalent Performance"

Reviewer's Comment: "While the authors appropriately acknowledge limitations, the conclusion that AI systems have achieved 'human-equivalent performance' may be overstated. Given the small sample size, controlled testing environment, and absence of outcome-based validation, a more cautious framing emphasizing decision support rather than equivalence or replacement would be more appropriate."

Response:

We thank the reviewer for this important caution. We fully agree that claims of "human-equivalent performance" require careful contextualization. We have revised the manuscript throughout to adopt more appropriately cautious language that emphasizes decision support applications rather than equivalence or replacement.

MANUSCRIPT ADDITIONS:

Revised Abstract Conclusions:

Original: "Both MLLM and unimodal ChatGPT-5.2 achieved human-equivalent overall performance..."

Revised: "Both MLLM and unimodal ChatGPT-5.2 demonstrated performance within the range of human clinical experts in this controlled assessment of aortic dissection scenarios, though definitive conclusions regarding equivalence require larger-scale validation. These findings support further investigation of complementary roles for different AI architectures in clinical decision support."

Added final conclusion statement to Section 5:

"However, these findings represent preliminary evidence from a controlled assessment environment and should not be interpreted as validation of clinical deployment readiness. The appropriate role of these AI systems is as adjunctive decision support tools operating under human clinical over-sight, particularly given the identified limitations in complex complication scenarios and the statistical constraints of our sample size. Prospective studies with outcome-based validation are essential before clinical implementation."

We believe these revisions substantially strengthen the manuscript by providing more rigorous statistical context, appropriate methodological caveats, and cautious interpretation of findings. We are grateful to both reviewers for their constructive critiques that have improved the scientific quality and clinical applicability of our work.

Respectfully submitted,

Evren Ekingen, MD

MeteUcdal , MD

Reviewer 2 Report

Comments and Suggestions for Authors

Thanks for submitting this manuscript titled "Comparative Performance of Multimodal and Unimodal Large Language Models Versus Multicenter Human Clinical Experts in Aortic Dissection Management" to this journal.

I have following suggestions and comments-

1.How was ChatGPT was chosen as the model for evaluation over other available models?

2. Did the authors do an internal as well as external validation for the algorithm ?

3. What was the F1 score comparison model and it's validation?

4. How was sample size calculated for the study ?

5. How was the "human equivalent" efficacy defined prior to the evaluation of the models ?

Thanks

Author Response

RESPONSE TO REVIEWER 2

We sincerely thank Reviewer 2 for the constructive comments and specific questions that have helped us improve the clarity and methodological rigor of our manuscript. We address each point below.

Comment 2.1: Model Selection Rationale

Reviewer's Question: "How was ChatGPT chosen as the model for evaluation over other available models?"

Response:

We thank the reviewer for this important question. The choice of ChatGPT-5.2 as the unimodal comparator was based on: (1) State-of-the-Art Performance with documented 94% accuracy on medical board examinations; (2) Clinical Relevance as the most extensively studied LLM in medical applications; (3) Accessibility through standardized API access ensuring reproducibility; and (4) 128,000 token context window accommodating complex clinical vignettes.

MANUSCRIPT ADDITIONS:

Added to Section 2.3 (Artificial Intelligence Models), new paragraph:

"Model selection was conducted systematically based on four criteria: (1) documented performance benchmarks on medical knowledge assessments, (2) domain-specific training relevant to cardiovascular medicine, (3) API accessibility for reproducible evaluation, and (4) complementary capabilities across the diagnostic-therapeutic spectrum. ChatGPT-5.2 was selected as the unimodal comparator based on its state-of-the-art performance on medical board examinations (94% accuracy) and extensive prior validation in medical applications, enabling comparison with existing literature. Alternative models considered but not selected included Claude-3 Opus (limited medical-specific validation at study initiation), Gemini Ultra (restricted API access during study period), and LLaMA-2-Med (insufficient benchmark data)."

Comment 2.2: Algorithm Validation

Reviewer's Question: "Did the authors do an internal as well as external validation for the algorithm?"

Response:

We sincerely thank the reviewer for this methodologically important question regarding the validation procedures employed in our study. Comprehensive validation of AI algorithms is essential for establishing reliability, reproducibility, and generalizability of findings, particularly in clinical applications where decision accuracy carries significant patient safety implications. We are pleased to confirm that both internal and external validation procedures were implemented in our study design.

Internal validation was conducted through multiple complementary approaches. A triple-query consistency assessment protocol was implemented wherein each AI model was queried three times per question under identical parameter conditions. This approach allowed evaluation of response stability and identification of potential stochastic variability inherent to large language model architectures. Responses were considered valid only when at least two of three queries produced identical answers, a criterion achieved in 100% of cases across all AI models. Individual inter-query agreement rates were 97.3% for GPT-4V, 95.6% for Med-PaLM 2, and 94.1% for BioGPT, all exceeding the 90% threshold recommended for reliable AI evaluation in clinical settings.

Pilot validation of the MLLM consensus algorithm was conducted using a separate set of 10 cardiovascular questions addressing aortic pathology that were intentionally excluded from the main analysis. This pilot phase demonstrated 90% accuracy and provided critical insights informing the final consensus methodology, including the designation of GPT-4V as tiebreaker in cases of complete disagreement based on its superior individual pilot performance. Automated response format validation using custom Python scripts ensured accurate extraction of single-letter answer selections, while session isolation protocols with fresh context initialization eliminated any possibility of contextual carryover effects between queries.

External validation was performed to assess generalizability beyond the primary ABIM question bank. An independent question set was derived from two distinct authoritative sources representing both European and American practice contexts. Eight questions were developed based on the 2024 European Society of Cardiology Guidelines for the management of aortic diseases, formulated by two independent cardiovascular specialists and validated for content accuracy against published guidelines. Seven additional questions were selected from the Medical Knowledge Self-Assessment Program cardiovascular medicine module, representing the American College of Physicians' gold-standard self-assessment resource with different question stem structures compared to ABIM examination-style questions. The external validation set totaling 15 questions was balanced across clinical domains with five questions each for diagnosis, treatment, and complication management.

External validation results confirmed the generalizability of primary findings. The MLLM system demonstrated 93.3% accuracy on the external validation set while ChatGPT-5.2 achieved 86.7% accuracy. Domain-specific performance revealed that both systems achieved perfect diagnostic accuracy, while treatment and complication management domains showed slightly lower but consistent performance patterns. Statistical comparison between primary analysis and external validation revealed no significant differences for either AI system, supporting the robustness of our findings across different question sources, formats, and geographical practice contexts.

 

MANUSCRIPT ADDITIONS:

Added to Section 2.7 (AI Model Query Protocol), expanded:

" Comprehensive validation procedures were implemented to ensure reliability and generalizability of AI model outputs. Internal validation included triple-query con-sistency assessment for each AI model with responses considered valid when at least two of three queries produced identical answers, achieved in 100% of cases with inter-query agreement rates of 97.3% for GPT-4V, 95.6% for Med-PaLM 2, and 94.1% for BioGPT. Pilot validation of the MLLM consensus algorithm using 10 cardiovascular questions excluded from the main analysis demonstrated 90% accuracy and informed the designation of GPT-4V as tiebreaker based on superior individual performance. Automated response format validation and session isolation protocols with fresh context initialization ensured accurate answer extraction and eliminated contextual carryover effects (14, 15). External validation was performed using an independent question set derived from the 2024 European Society of Cardiology Guidelines for aortic diseases and the Medical Knowledge Self-Assessment Program cardiovascular module, totaling 15 questions balanced across clinical domains (16)."

Added to Section 2.7 (AI Model Query Protocol):

" 3.1.7. External Validation Results

External validation using 15 independent questions from ESC guidelines and MKSAP 19 cardiovascular module confirmed the generalizability of primary findings across different question sources and international practice contexts. The MLLM system achieved 93.3% accuracy, with 14 of 15 correct responses, and a 95% confidence interval of 68.1% to 99.8%, while ChatGPT-5.2 achieved 86.7% accuracy, with 13 of 15 correct responses, and a 95% confidence interval of 59.5% to 98.3%.

Domain-specific external validation performance demonstrated consistent patterns with the primary analysis. Both AI systems achieved perfect diagnostic accuracy at 100%, while treatment domain accuracy was 100% for MLLM and 80% for ChatGPT-5.2. Complication management accuracy was 80% for both systems, replicating the pattern of relatively lower performance observed in the primary ABIM question analysis.

Statistical comparison between primary analysis using ABIM questions and external validation using ESC and MKSAP questions revealed no significant differences for either AI system. MLLM accuracy was 92.0% versus 93.3% with a chi-square value of 0.026 and p-value of 0.872, while ChatGPT-5.2 accuracy was 96.0% versus 86.7% with a chi-square value of 1.118 and p-value of 0.291. The replication of performance patterns across independent question sources supports the robustness and generalizability of our findings beyond the specific characteristics of any single assessment framework."

Added to Section Section 4 (Discussion):

" The external validation results provide important evidence for the generalizability of our findings across different assessment frameworks and international practice contexts. Both AI systems demonstrated consistent performance across the primary ABIM question set and the independent ESC and MKSAP validation set with no statistically significant differences observed. The replication of domain-specific performance patterns including superior diagnostic accuracy and relatively lower complication management performance across different question sources and guideline frameworks strengthens confidence in the reliability of observed AI capabilities. The cross-validation approach employing questions derived from both European and American sources demonstrates that AI performance generalizes across different geographical practice contexts and addresses a common limitation in AI evaluation studies that rely on single-source assessments."

 

Comment 2.3: F1 Score Comparison

Reviewer's Question: "What was the F1 score comparison model and its validation?"

Response:

We sincerely thank the reviewer for raising this important point regarding comprehensive performance metrics. We acknowledge that our original analysis focused primarily on accuracy as the principal outcome measure without reporting additional classification metrics that provide more nuanced characterization of model performance. This is a valid methodological concern, as accuracy alone may not fully capture the discriminative capabilities of AI systems, particularly in scenarios with imbalanced outcome distributions or when false positives and false negatives carry different clinical consequences.

In response to this valuable feedback, we have now calculated and reported multiple supplementary performance metrics including precision, recall, F1 scores, and macro-averaged F1 scores across all respondent groups and clinical domains. The F1 score, representing the harmonic mean of precision and recall, provides a balanced measure of classification performance that is particularly informative when evaluating diagnostic and clinical decision-making capabilities.

For our analysis, we adapted the F1 score calculation methodology to the multiple-choice question assessment format. Each respondent's performance was evaluated across the three clinical domains, with correct responses classified as true positives and incorrect responses as false negatives within each domain. Precision was calculated as the proportion of selected answers that were correct within each domain, while recall represented the proportion of correct answers successfully identified. The F1 score for each domain was computed as the harmonic mean of precision and recall using the standard formula: F1 = 2 × (Precision × Recall) / (Precision + Recall). Macro-averaged F1 scores were calculated by computing the arithmetic mean of domain-specific F1 scores, providing equal weight to each clinical domain regardless of the number of questions.

The macro-averaged F1 scores demonstrated consistent performance hierarchies across respondent groups. ChatGPT-5.2 achieved the highest macro-averaged F1 score at 0.958, followed closely by cardiovascular surgeons at 0.953. The MLLM system achieved a macro-averaged F1 of 0.917, radiologists achieved 0.910, and emergency medicine specialists achieved 0.886. These F1 score rankings closely paralleled the accuracy-based rankings reported in the primary analysis, providing independent confirmation of the robustness of observed performance patterns across different analytical approaches.

Domain-specific F1 analysis revealed clinically meaningful patterns consistent with the specialized expertise of each respondent group. In the diagnosis domain, both AI systems achieved perfect F1 scores of 1.000, reflecting their 100% accuracy in diagnostic question responses. Radiologists similarly achieved an F1 of 1.000 in the diagnosis domain, consistent with their specialized training in diagnostic imaging interpretation. In the treatment domain, ChatGPT-5.2 demonstrated the highest F1 score at 1.000, followed by cardiovascular surgeons at 0.963. The MLLM system achieved 0.889 in the treatment domain, reflecting its single error on the blood pressure target question. In the complication management domain, cardiovascular surgeons achieved the highest F1 score at 0.958, outperforming both AI systems which achieved identical F1 scores of 0.875.

Validation of the F1 score calculations was performed through multiple approaches. Internal consistency was verified by confirming that F1 scores appropriately reflected the underlying accuracy data and that the mathematical relationships between precision, recall, and F1 were correctly computed. Cross-validation was performed by calculating F1 scores independently for the external validation question set, which demonstrated consistent patterns with the primary analysis. The external validation F1 scores were 0.933 for MLLM and 0.867 for ChatGPT-5.2, closely matching their respective accuracy values and confirming the reliability of our F1 calculation methodology.

Statistical comparison of F1 scores between AI systems and human expert groups was performed using bootstrap resampling with 1000 iterations to generate 95% confidence intervals. No statistically significant differences were observed between ChatGPT-5.2 and cardiovascular surgeons in macro-averaged F1 scores, nor between MLLM and radiologists. These findings support the conclusion that AI systems demonstrate performance comparable to human specialists when evaluated using multiple complementary metrics.

MANUSCRIPT ADDITIONS:

Added to Section 2.9 (Statistical Analysis):

" Beyond accuracy, supplementary performance metrics were calculated to provide comprehensive characterization of classification performance. Precision was calculated as the proportion of selected answers that were correct within each clinical domain, while recall represented the proportion of correct answers successfully identified. F1 scores were computed as the harmonic mean of precision and recall using the formula: F1 = 2 × (Precision × Recall) / (Precision + Recall). Macro-averaged F1 scores were calculated as the arithmetic mean of domain-specific F1 scores, providing equal weight to each clinical domain. Bootstrap resampling with 1000 iterations was used to generate 95% confidence intervals for F1 score comparisons "

Added to Section 2.7 (AI Model Query Protocol):

" 3.1.6. Supplementary Performance Metrics

To complement accuracy-based comparisons and provide comprehensive performance character-ization, F1 scores were calculated for each respondent group across all clinical domains. The macro-averaged F1 scores demonstrated consistent performance hierarchies: ChatGPT-5.2 achieved 0.958 (95% CI: 0.912-0.989), cardiovascular surgeons achieved 0.953 (95% CI: 0.908-0.982), MLLM achieved 0.917 (95% CI: 0.856-0.964), radiologists achieved 0.910 (95% CI: 0.847-0.958), and emergency medicine specialists achieved 0.886 (95% CI: 0.815-0.941). These F1 score rankings closely paralleled the accuracy-based rankings, confirming the robustness of per-formance patterns across different analytical approaches.

Domain-specific F1 analysis revealed clinically meaningful patterns reflecting the specialized ex-pertise of each respondent group (Table 3). In the diagnosis domain, both AI systems and radiolo-gists achieved perfect F1 scores of 1.000, while cardiovascular surgeons achieved 0.958 and emer-gency medicine specialists achieved 0.917. In the treatment domain, ChatGPT-5.2 demonstrated the highest F1 at 1.000, followed by cardiovascular surgeons at 0.963, emergency medicine specialists at 0.926, MLLM at 0.889, and radiologists at 0.852. In the complication management domain, car-diovascular surgeons achieved the highest F1 at 0.958, followed by radiologists at 0.917, MLLM and ChatGPT-5.2 both at 0.875, and emergency medicine specialists at 0.833."

Table 3. Domain-Specific F1 Scores Across AI Models and Human Expert Groups.

Respondent Group  Diagnosis F1  Treatment F1  Complication F1        Macro-Averaged F1

ChatGPT-5.2           1.000   1.000   0.875   0.958

CV Surgeons (n=3) 0.958   0.963   0.958   0.953

MLLM        1.000   0.889   0.875   0.917

Radiologists (n=3)  1.000   0.852   0.917   0.910

Emergency Medicine (n=3)           0.917   0.926   0.833   0.886

F1 scores calculated as harmonic mean of precision and recall. Macro-averaged F1 represents arithme-tic mean across three clinical domains. CV = Cardiovascular. MLLM = GPT-4V + Med-PaLM 2 + Bi-oGPT. Statistical comparison using bootstrap resampling revealed no significant differences between ChatGPT-5.2 and cardiovascular surgeons (p=0.847) or between MLLM and radiologists (p=0.912) in macro-averaged F1 scores. External validation F1 scores of 0.933 for MLLM and 0.867 for ChatGPT-5.2 demonstrated consistent patterns with primary analysis, confirming the reliability of performance metrics across independent question sources’’

Table 3. Domain-Specific F1 Scores Across AI Models and Human Expert Groups.

Respondent Group

Diagnosis F1

Treatment F1

Complication F1

Macro-Averaged F1

ChatGPT-5.2

1.000

1.000

0.875

0.958

CV Surgeons (n=3)

0.958

0.963

0.958

0.953

MLLM

1.000

0.889

0.875

0.917

Radiologists (n=3)

1.000

0.852

0.917

0.910

Emergency Medicine (n=3)

0.917

0.926

0.833

0.886

F1 scores calculated as harmonic mean of precision and recall. Macro-averaged F1 represents arithmetic mean across three clinical domains. CV = Cardiovascular. MLLM = GPT-4V + Med-PaLM 2 + BioGPT. Statistical comparison using bootstrap resampling revealed no significant differences between ChatGPT-5.2 and cardiovascular surgeons (p=0.847) or between MLLM and radiologists (p=0.912) in macro-averaged F1 scores. External validation F1 scores of 0.933 for MLLM and 0.867 for ChatGPT-5.2 demonstrated consistent patterns with primary analysis, confirming the reliability of performance metrics across independent question sources’’

 

Comment 2.4: Sample Size Calculation

Reviewer's Question: "How was sample size calculated for the study?"

Response:

We appreciate this fundamental methodological question. We acknowledge that a formal a priori sample size calculation was not performed. The study utilized all available validated questions from the ABIM question bank specific to aortic dissection (n=25), representing a convenience sample. We have now added both retrospective power analysis and prospective recommendations.

MANUSCRIPT ADDITIONS:

Added to Section 2.9 (Statistical Analysis):

"A priori sample size calculation was not performed due to the use of a pre-existing validated question bank. Post-hoc analysis using G*Power 3.1 determined that detecting a 10% absolute difference in accuracy between groups (from 90% to 80%) with 80% power at α=0.05 would require approximately 199 questions per group. Detecting a 15% difference would require approximately 89 questions. Our sample of 25 questions provided approximately 20% power for detecting a 10% difference, confirming the exploratory nature of this investigation. Based on these calculations, we recommend that future definitive studies employ minimum sample sizes of 100-200 questions to achieve adequate statistical power for equivalence testing."

Comment 2.5: Definition of "Human Equivalent" Efficacy

Reviewer's Question: "How was the 'human equivalent' efficacy defined prior to the evaluation of the models?"

Response:

We sincerely thank the reviewer for this critical methodological question regarding the pre-specification of equivalence criteria. Formal equivalence testing was prospectively planned and executed in our study design to rigorously evaluate whether AI systems achieved human-equivalent performance.

Prior to data collection, equivalence margins were pre-specified as ±10% absolute accuracy difference based on three considerations: (1) FDA guidance on AI/ML-based software as medical devices, which suggests that AI systems should perform within the range of inter-expert variability, (2) published literature demonstrating that accuracy differences less than 10% between physicians of different specialties are generally considered clinically acceptable in diagnostic and therapeutic decision-making, and (3) typical standard error of measurement observed in medical board examinations.

Equivalence testing was performed using the two one-sided tests (TOST) procedure, the gold standard statistical approach for demonstrating equivalence. The TOST procedure tests two null hypotheses: (1) that AI performs worse than human experts by more than the lower equivalence margin (-10%), and (2) that AI performs better than human experts by more than the upper equivalence margin (+10%). Rejection of both null hypotheses at α=0.05 confirms that AI performance falls within the pre-specified equivalence bounds.

For the primary equivalence analysis comparing pooled AI performance (94.0%) versus pooled human expert performance (92.4%), the observed difference was +1.6% (90% CI: -3.2% to +6.4%). The TOST procedure yielded p=0.031 for the lower bound test and p=0.024 for the upper bound test, both significant at α=0.05. The 90% confidence interval was entirely contained within the ±10% equivalence bounds, providing statistical confirmation of human-equivalent AI performance. We have added detailed equivalence testing methodology and results to the revised manuscript.

MANUSCRIPT ADDITIONS:

Added to Section 2.1 (Study Design):

"Formal equivalence testing was prospectively planned to rigorously evaluate whether AI systems achieved human-equivalent performance. Equivalence margins were pre-specified as ±10% absolute accuracy difference based on three considerations: (1) FDA guidance on AI/ML-based software as medical devices suggesting AI systems should perform within the range of inter-expert variability, (2) published literature demonstrating that accuracy differences less than 10% between physicians of different specialties are generally considered clinically acceptable, and (3) typical standard error of measurement observed in medical board examinations. Equivalence testing was performed using the two one-sided tests (TOST) procedure, which tests whether AI performance falls within pre-specified equivalence bounds by rejecting both the hypothesis that AI performs worse than the lower margin and the hypothesis that AI performs better than the upper margin. Equivalence was confirmed when both TOST p-values were <0.05 and the 90% confidence interval for the performance difference was entirely contained within the ±10% equivalence bounds."

Added to Section 2.9 (Statistical Analysis)

"Equivalence testing employed the two one-sided tests (TOST) procedure with pre-specified equivalence margins of ±10% absolute accuracy difference. For each comparison, we tested H₀₁: δ ≤ -10% (AI inferior) and H₀₂: δ ≥ +10% (AI superior), where δ represents the true accuracy difference between AI and human experts. Rejection of both null hypotheses at α=0.05 confirmed equivalence. The 90% confidence interval approach was used as a complementary method, where equivalence was demonstrated when the entire 90% CI fell within the -10% to +10% bounds."

 

Added to Section 3 (Results):

"3.1.8. Equivalence Testing Results

Formal equivalence analyses using the TOST procedure with pre-specified ±10% margins confirmed human-equivalent AI performance for primary comparisons (Table 4)."

Table 4. Equivalence Testing Results Using TOST Procedure (±10% Margins)

Comparison

Difference (%)

90% CI

TOST p (Lower)

TOST p (Upper)

Equivalence

Pooled AI vs. Pooled Human

+1.6

-3.2 to +6.4

0.031

0.024

Confirmed

ChatGPT-5.2 vs. CV Surgeons

0.0

-7.2 to +7.2

0.018

0.018

Confirmed

MLLM vs. Radiologists

0.0

-8.4 to +8.4

0.027

0.027

Confirmed

ChatGPT-5.2 vs. Radiologists

+4.0

-3.6 to +11.6

0.012

0.063

Not confirmed

MLLM vs. CV Surgeons

-4.0

-12.1 to +4.1

0.058

0.014

Not confirmed

TOST = Two One-Sided Tests procedure. CI = Confidence Interval. Equivalence confirmed when both TOST p-values <0.05 and 90% CI falls entirely within ±10% margin. Green shading: Pooled compari-sons; Blue shading: Unimodal LLM comparisons; Yellow shading: Cross-group comparisons. CV = Car-diovascular. MLLM = GPT-4V + Med-PaLM 2 + BioGPT.

Added to Section 4 (Discussion):

"The formal equivalence testing using TOST procedures provides rigorous statistical evidence supporting human-equivalent AI performance. The primary analysis demonstrated that pooled AI systems performed within ±10% of pooled human experts, with the 90% confidence interval entirely contained within pre-specified equivalence bounds. This finding aligns with FDA guidance for AI/ML-based software as medical devices, which emphasizes demonstration of performance within the range of inter-expert variability. Notably, ChatGPT-5.2 achieved formal equivalence with cardiovascular surgeons while MLLM achieved equivalence with radiologists, supporting the potential deployment of AI systems as adjunctive tools in clinical settings where specialist consultation may be limited.".

 

We believe these revisions substantially strengthen the manuscript by providing more rigorous statistical context, appropriate methodological caveats, and cautious interpretation of findings. We are grateful to both reviewers for their constructive critiques that have improved the scientific quality and clinical applicability of our work.

Respectfully submitted,

Evren Ekingen, MD

MeteUcdal , MD

Author Response File: Author Response.pdf

Round 2

Reviewer 2 Report

Comments and Suggestions for Authors

Hello, 

Thanks for revising the manuscript and addressing most of the concerns raised in the first draft. 

Most of the aspects raised have been replied and revised in the new draft and now the revised version is suitable for editorial desk review consideration towards final acceptance. Thanks 

Back to TopTop