Discrepancies Between MDT Recommendations and AI-Generated Decisions in Gynecologic Oncology: A Retrospective Comparative Cohort Study
Round 1
Reviewer 1 Report
Comments and Suggestions for Authors1.The manuscript simultaneously labels the study as both a “prospective comparative cohort study” and a “retrospective observational study”, an overt contradiction. Please definitively clarify the study design; if retrospective, harmonise all text—title, abstract, and Methods—to “retrospective cohort study”.
2.ChatGPT input templates, fixed prompts, multi-turn dialogue status, human intervention checkpoints, and operator-blinding procedures are not supplied. Amend the Methods to provide the full prompt and the structured case-input format; deposit them as supplementary material to ensure full reproducibility.
3.MDT composition, member expertise, and meeting workflow are unreported, precluding appraisal of the panel’s internal consistency and reliability. Supplement the Methods with participants’ specialties, years of experience, and the decision rule (consensus vs. majority vote) used to reach clinical recommendations.
4.Any “clinically relevant difference” is treated as a discrepancy without defining relevance or grading severity. Introduce a discrepancy taxonomy: major (alters treatment trajectory) and minor (both options within guideline-acceptable ranges). Illustrate each category with concrete examples in the main text or appendix to enhance transparency.
5.Analyses are limited to descriptive statistics and χ² tests; inter-rater agreement (κ), confidence intervals, and multivariable modelling are absent. Add Cohen’s or Fleiss’ κ for MDT–AI concordance, report 95 % CIs for all discrepancy rates, and employ multivariable logistic regression to identify predictors of discrepancy (e.g. tumour type, stage, recurrence status).
6.Although endometrial and ovarian cases exhibit higher discrepancy rates, the informational deficits underlying AI errors are not elucidated. Provide an error typology: insufficient molecular-subtype interpretation; inability to integrate imaging data; absent surgical resectability context; information complexity in recurrent disease leading to misclassification.
Author Response
We thank the reviewer for their thorough, insightful, and constructive comments. In the present revision all of them are introduced and helped improve methodological clarity and interpretability of findings.
1.The manuscript simultaneously labels the study as both a “prospective comparative cohort study” and a “retrospective observational study”, an overt contradiction. Please definitively clarify the study design; if retrospective, harmonise all text—title, abstract, and Methods—to “retrospective cohort study”.
Authors reply: We thank the reviewer for highlighting this inconsistency. The study was conducted as a retrospective cohort analysis of consecutively discussed MDT cases. The term “prospective” was used inaccurately and has now been removed. The study design has been harmonised throughout the title, abstract, and Methods to consistently reflect a retrospective cohort study.
2.ChatGPT input templates, fixed prompts, multi-turn dialogue status, human intervention checkpoints, and operator-blinding procedures are not supplied. Amend the Methods to provide the full prompt and the structured case-input format; deposit them as supplementary material to ensure full reproducibility.
Authors reply: we thank the reviewer for this remark as we agree that reproducibility is essential. We have expanded the Methods to fully describe the AI interaction framework, including the structured case-input template, fixed system prompt, absence of multi-turn dialogue, and lack of human intervention during output generation. The complete prompt and input template have now been provided as Supplementary Appendix 1.
3.MDT composition, member expertise, and meeting workflow are unreported, precluding appraisal of the panel’s internal consistency and reliability. Supplement the Methods with participants’ specialties, years of experience, and the decision rule (consensus vs. majority vote) used to reach clinical recommendations.
Authors reply: We thank the reviewer for this important observation. We have added this information in the Appendix section of the present revision.
4.Any “clinically relevant difference” is treated as a discrepancy without defining relevance or grading severity. Introduce a discrepancy taxonomy: major (alters treatment trajectory) and minor (both options within guideline-acceptable ranges). Illustrate each category with concrete examples in the main text or appendix to enhance transparency.
Authors reply: In the present revision we included a predefined discrepancy taxonomy distinguishing major and minor discrepancies based on their potential impact on the overall treatment trajectory. These definitions have been incorporated into the Methods section (Definition of Discrepancy and Review Procedure part).
5.Analyses are limited to descriptive statistics and χ² tests; inter-rater agreement (κ), confidence intervals, and multivariable modelling are absent. Add Cohen’s or Fleiss’ κ for MDT–AI concordance, report 95 % CIs for all discrepancy rates, and employ multivariable logistic regression to identify predictors of discrepancy (e.g. tumour type, stage, recurrence status).
Authors reply: We thank the reviewer for this valuable input. We expanded the statistical analysis to include (i) inter-rater agreement between MDT and AI recommendations using Cohen’s κ with 95% confidence intervals, (ii) 95% confidence intervals for all discrepancy proportions, and (iii) multivariable logistic regression to identify predictors of discordance (tumor type, early/advanced stage category when available, recurrence status, and ECOG performance status). These analyses have been added to the Methods and Results section and Table 2 has been accordingly revised and a newly introduced Table 4 has been inserted that summarizes the results of the multivariate regression analysis.
6.Although endometrial and ovarian cases exhibit higher discrepancy rates, the informational deficits underlying AI errors are not elucidated. Provide an error typology: insufficient molecular-subtype interpretation; inability to integrate imaging data; absent surgical resectability context; information complexity in recurrent disease leading to misclassification.
Authors reply: in the present revision these remarks are introduced in the discussion section.
Reviewer 2 Report
Comments and Suggestions for AuthorsThe paper addresses a well-known problem in AI, which is the difficulty of LLM-based methods in providing sound reasoning on complex scenarios. An adequate discussion section remarks the pros and cons of the current study, while addressing some limitations that are likely to be addressed in the authors' future works. The authors also provide a discussion of the potential implications of the study. The paper could be improved with minor comments:
- Authors should consider adding a subsection in Section 2 comparing LLM to the statistical techniques being used in the present paper within IBM SPSS, while remarking on the pros and cons of these. Within this context, authors should better disclose at the beginning of Section 3 the way the IBM SPSS plays a role in MDT, and better remark which statistical approaches were used in particular (as it is not clear whether IBM SPSS was merely used to outline the correlation results plotted in the tables, or was also used to assist with the clinical decisions by summarizing and creating a model of the data).
- The authors should improve the reproducibility of the process used to compare the different algorithms: they should disclose how they provided prompting to the LLM-based model and compare these methods against the criteria used by the clinicians to derive their decisions. This would help determine whether the LLM had to "figure out" some steps, or whether it was basically spoon-fed the same decisions and outcomes as the experts, thus parroting their judgment.
- If adequate time is given to the authors to address this, they should consider an ablation study that considers both prompt styles. This would greatly help to understand in which part the LLMs were more defective (in actually parroting the reasoning steps, or extracting further knowledge while providing more general prompts).
Overall, the study is well conducted and the paper is overall well-written and easy to read.
Author Response
The paper addresses a well-known problem in AI, which is the difficulty of LLM-based methods in providing sound reasoning on complex scenarios. An adequate discussion section remarks the pros and cons of the current study, while addressing some limitations that are likely to be addressed in the authors' future works. The authors also provide a discussion of the potential implications of the study. The paper could be improved with minor comments:
- Authors should consider adding a subsection in Section 2 comparing LLM to the statistical techniques being used in the present paper within IBM SPSS, while remarking on the pros and cons of these. Within this context, authors should better disclose at the beginning of Section 3 the way the IBM SPSS plays a role in MDT, and better remark which statistical approaches were used in particular (as it is not clear whether IBM SPSS was merely used to outline the correlation results plotted in the tables, or was also used to assist with the clinical decisions by summarizing and creating a model of the data).
Authors reply: We thank the reviewer for this comment. IBM SPSS was not used in any way to assist MDT decision-making or clinical judgment. Its role was strictly limited to post-hoc statistical analysis of concordance, discrepancy rates, inter-rater agreement, and mutiple regression. Given that IBM SPSS was not used as a decision-support or modeling tool in the MDT process, we did not introduce a direct methodological comparison between LLMs and classical statistical approaches, as they serve fundamentally different purposes in this study (clinical decision generation vs. analytical evaluation). In the present revision the section entitled AI Model and Input Standardization has been reorganized to capture the functions that were used for the LMM techniques and a newly designed Appendix section was also structured to describe the AI interaction framework, including the structured case-input template, fixed system prompt, absence of multi-turn dialogue, and lack of human intervention during output generation.
- The authors should improve the reproducibility of the process used to compare the different algorithms: they should disclose how they provided prompting to the LLM-based model and compare these methods against the criteria used by the clinicians to derive their decisions. This would help determine whether the LLM had to "figure out" some steps, or whether it was basically spoon-fed the same decisions and outcomes as the experts, thus parroting their judgment.
Authors reply: We also thank the reviewer for this input. In the present revision we substantially expanded the Methods and added Supplementary Appendix 1, providing:
- the verbatim system prompt
- the structured case-input template
- explicit confirmation that:
- FIGO stage was not disclosed to the AI
- MDT conclusions were not included
- interactions were single-turn, with no iterative dialogue
- no human intervention occurred
These additions clarify that the LLM was required to independently infer staging and management based solely on clinical inputs, rather than reproducing or parroting expert decisions.
- If adequate time is given to the authors to address this, they should consider an ablation study that considers both prompt styles. This would greatly help to understand in which part the LLMs were more defective (in actually parroting the reasoning steps, or extracting further knowledge while providing more general prompts).
Authors reply: We thank the reviewer for this valuable suggestion, which we agree would provide important insight into LLM behavior. However, performing an ablation study comparing different prompt strategies would require a separate experimental design, additional AI runs, and predefined prompt variations that fall outside the scope of the present retrospective comparative cohort study. The current study was designed to evaluate AI performance under a single, standardized, guideline-driven prompting strategy, reflecting a realistic potential clinical implementation rather than an exploratory AI optimization framework. Acknowledging the importance of this question we added an explicit statement in the discussion inside the Implications for clinical practice section.
Reviewer 3 Report
Comments and Suggestions for AuthorsThe paper's topic is interesting and timely. The methodology sounds valid. However, I have some questions and comments as follows.
- What is the overarching goal of the study? Are you investigating replacing MDTs (or some of their tasks) with LLMs? If so, it should be justified why.
- On the same note, what is the necessity of doing this research for these types of cancer? Please, justify selection of these types.
- The number of cases in the methodology section differs from the abstract.
- The LLM's inputs are not clear. Are they notes/documents or tabular data? Can you provide clarification or visualisation if possible?
- Rewrite this section as I believe the use of brackets is confusing, "Surgical management (extent, timing, method), (Systemic therapy: chemotherapy or immunotherapy), Targeted or hormonal therapy, and Follow-up or surveillance strategy".
- How many of women in your study are with cervical cancer? It is missing.
- Can you elaborate more on the outcome including variables presented in Table 1? What is the decision if the recommendations differ on one area out of four? What does concordant mean? Is it just 1-Discrepancy%?
- I recommend reviewing topics on "human–Artificial Intelligence (AI) interaction" and "Expert–machine collaborative decision making" to enhance discussion section and future research.
- For future studies, I recommend considering uncertainty in concordance/discrepancy through multi-criteria group decision-making under uncertainty using interval data.
Author Response
The paper's topic is interesting and timely. The methodology sounds valid. However, I have some questions and comments as follows.
- What is the overarching goal of the study? Are you investigating replacing MDTs (or some of their tasks) with LLMs? If so, it should be justified why.
Authors reply: We thank the reviewer for raising this important conceptual point. The goal of the study was not to evaluate replacement of MDTs by LLMs, but rather to assess whether an advanced LLM can reproduce MDT-derived recommendations under real-world conditions and identify clinical contexts where discordance emerges. This is clarified in the Introduction section of the present revision.
- On the same note, what is the necessity of doing this research for these types of cancer? Please, justify selection of these types.
Authors reply: We thank the reviewer for this important question. Cervical, endometrial, ovarian, and vulvar cancers were selected because together they encompass the full clinical, biological, and decision-making spectrum of gynecologic oncology. These malignancies differ substantially in staging systems, reliance on imaging and surgery, incorporation of molecular classification, and degree of guideline linearity. Evaluating AI performance across these heterogeneous cancer types allows assessment of whether LLMs perform consistently across simple, standardized scenarios and complex, highly individualized clinical contexts. This rationale has now been explicitly clarified in the last paragraph of the Introduction of the present revision.
- The number of cases in the methodology section differs from the abstract.
Authors reply: We thank the reviewer for identifying this inconsistency. The correct number of patients included in the analysis is 599, reflecting the final cohort with complete MDT and AI data. This number has now been harmonized across the Abstract, Methods, Results, and all Tables to ensure consistency throughout the manuscript.
- The LLM's inputs are not clear. Are they notes/documents or tabular data? Can you provide clarification or visualisation if possible?
Authors reply: AI inputs were provided as structured textual summaries following a predefined template routinely used for MDT case presentation. These summaries were entered as narrative text fields and did not consist of tabular data, raw imaging files, pathology reports, or unstructured clinical notes. The structure and content of the input template were standardized across all cases to ensure consistency and reproducibility. A detailed representation of the input format and included data elements is provided in Supplementary Appendix 1.
- Rewrite this section as I believe the use of brackets is confusing, "Surgical management (extent, timing, method), (Systemic therapy: chemotherapy or immunotherapy), Targeted or hormonal therapy, and Follow-up or surveillance strategy".
Authors reply: this section was corrected in the present revision
- How many of women in your study are with cervical cancer? It is missing.
Authors reply: this was corrected in the present revision.
- Can you elaborate more on the outcome including variables presented in Table 1? What is the decision if the recommendations differ on one area out of four? What does concordant mean? Is it just 1-Discrepancy%?
Authors reply: We thank the reviewer for this important request for clarification. Each decision domain (staging, surgical management, systemic therapy, and targeted/hormonal therapy) was evaluated independently. Concordance was defined as agreement between MDT and AI within a given domain, whereas discrepancy was recorded only when a major difference was identified in that specific domain. Discordance in one domain did not imply discordance in others. Concordant proportions therefore represent the complement of the domain-specific discrepancy rate. These definitions have now been clarified in the Methods and Results sections and a comment was introduced in the caption of Table 1.
- I recommend reviewing topics on "human–Artificial Intelligence (AI) interaction" and "Expert–machine collaborative decision making" to enhance discussion section and future research.
Authors reply: We thank the reviewer for this valuable suggestion. We have strengthened the Discussion and Future Research sections to explicitly frame AI as part of a human–AI collaborative decision-making model, rather than as an autonomous decision-maker.
- For future studies, I recommend considering uncertainty in concordance/discrepancy through multi-criteria group decision-making under uncertainty using interval data.
Authors reply: we thank the reviewer for this remark that was added in the relevant section of the discussion in the present revision.
Round 2
Reviewer 1 Report
Comments and Suggestions for Authorsaccept
Author Response
We thank the reviewer for accepting our study
Reviewer 2 Report
Comments and Suggestions for AuthorsWe thank the authors for providing the prompts given to the LLMs to make the inference, as well as providing evidence on the phisicians' board formation. Notwithstanding the foregoing, the authors did not provide sufficient evidence to clearly examine the clinicians' thought processes, thereby preventing comparisons of the human way of making inferences from the data and the prompts' thought processes. If not applicable, the authors should still provide the guidelines on how humans make decisions. This should be reported, as "MDT recommendations were documented prospectively in the institutional MDT report and served as the reference standard for comparative analysis." So, it should be an easy task for the authors to reproduce a specimen of this process.
Author Response
We thank the authors for providing the prompts given to the LLMs to make the inference, as well as providing evidence on the physicians' board formation. Notwithstanding the foregoing, the authors did not provide sufficient evidence to clearly examine the clinicians' thought processes, thereby preventing comparisons of the human way of making inferences from the data and the prompts' thought processes. If not applicable, the authors should still provide the guidelines on how humans make decisions. This should be reported, as "MDT recommendations were documented prospectively in the institutional MDT report and served as the reference standard for comparative analysis." So, it should be an easy task for the authors to reproduce a specimen of this process.
Authors reply: MDT recommendations were derived from the institutional MDT report indeed and the proposed phrase was added in the manuscript methods as last phrase of the "AI Model and Input Standardization" section.
Reviewer 3 Report
Comments and Suggestions for AuthorsThe need of "Expert–machine collaborative decision making" and "group decision-making under uncertainty using interval data" is elaborated in the manuscript. However, the relevant literature is not reviewed.
Authors has responded to the first comment as "assess whether an advanced LLM can reproduce MDT-derived recommendations under real-world conditions and identify clinical contexts where discordance emerges". However, I believe there is still need for a longer term goal. The expert-machine collaborative decision-making can be that.
Author Response
The need of "Expert–machine collaborative decision making" and "group decision-making under uncertainty using interval data" is elaborated in the manuscript. However, the relevant literature is not reviewed.
Authors has responded to the first comment as "assess whether an advanced LLM can reproduce MDT-derived recommendations under real-world conditions and identify clinical contexts where discordance emerges". However, I believe there is still need for a longer term goal. The expert-machine collaborative decision-making can be that.
Authors reply: we thank the reviewer for this remark. In the present revision we extended this section to provide a statement that captures the potential long-term goal of this collaborative decision making.
Round 3
Reviewer 2 Report
Comments and Suggestions for AuthorsThe sentence "MDT recommendations were documented prospectively in the institutional
MDT report and served as the reference standard for comparative analysis." is not sufficient. What was the rationale for making the decisions? What is the evidence that each clinician takes? You are just describing the final review process, but not comparing how a single human makes a decision. Also, the process was not reported, which raised some eyebrows about whether the human process was actually occurring.
Author Response
We thank the reviewer for revisiting this point and appreciate the opportunity to further clarify it. In routine gynecologic oncology practice, MDT decisions are reached through the synthesis of predefined clinical evidence, including radiologic staging assessments, histopathological findings, molecular or prognostic markers where applicable, patient performance status, prior treatments, and current guideline recommendations. This process is standardized across institutions and is prospectively documented in MDT reports, which represent the formal clinical output of the decision-making process.
In our institution, ESGO clinical practice guidelines serve as the reference framework for selecting therapeutic options in each cancer type. All MDT recommendations were documented prospectively prior to treatment initiation and were subsequently retrieved retrospectively for comparison with AI-generated decisions.
