Next Article in Journal
Enhancing Lecture Interactivity Through Virtual Reality
Previous Article in Journal
The Effects of Advertisement Placement Configurations on Visual Attention and Recall According to Dynamic Road Traffic Conditions Using Virtual Reality and Eye Tracking
 
 
Article
Peer-Review Record

Dialogical AI for Cognitive Bias Mitigation in Medical Diagnosis

Appl. Sci. 2026, 16(2), 710; https://doi.org/10.3390/app16020710
by Leonardo Guiducci 1,*, Claudia Saulle 1, Giovanna Maria Dimitri 2, Benedetta Valli 3, Simona Alpini 3, Cristiana Tenti 3 and Antonio Rizzo 1,*
Reviewer 1:
Reviewer 2:
Reviewer 3: Anonymous
Reviewer 4: Anonymous
Appl. Sci. 2026, 16(2), 710; https://doi.org/10.3390/app16020710
Submission received: 27 November 2025 / Revised: 31 December 2025 / Accepted: 2 January 2026 / Published: 9 January 2026
(This article belongs to the Section Computing and Artificial Intelligence)

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

Summary
The manuscript proposes a Dialogic Reasoning Framework with three roles (Framework Coach, Socratic Guide, Red Team Partner) built on top of a RAG architecture to improve medical diagnostic in LLMs. The paper presents a proof of concept using GPT-4o with two qualitative case studies and a comparison with ChatGPT.

The topic is important, however, there are several issues that need to be addressed before the paper can be considered for publication:

 

Evidence of Performance

The paper presents two qualitative cases and a brief comparison with a ChatGPT interaction. While these are useful illustrative examples, they do not provide enough evidence on the performance of the framework. The analysis is qualitative and the outcomes are described narratively rather than measured quantitatively. No performance metrics (e.g., diagnostic accuracy, reduction in bias, etc.) are defined. There is also no quantitative comparison of diagnostic accuracy (or bias reduction) between the proposed framework and ChatGPT interactions. Since the implementation is framed as a proof of concept, the paper could provide at least some quantitative evidence on the performance of the proposed framework.

 

Methodological Detail

The current description of implementation lacks some details including number of clinicians involved, number of patient cases, and case-selection procedures. Also, prompt templates for the Framework Coach, Socratic Guide, and Red Team Partner are not presented.

Additionally, in comparison between ChatGPT and DiDi, some details are not clear. For example, is any resetting strategy used to remove session information in ChatGPT? Without resetting, ChatGPT may use memorized context of previous cases. It is also not mentioned that what decoding parameters are used for Gpt-4o API calls (e.g., temperature). Smaller temperatures may provide more precise and stable results for medical diagnosis. Without these details evaluation of the comparison and reproducibility is not straightforward.

 

Overstatement

The paper makes strong claims that the proposed framework is a paradigm shift, transforming LLMs from passive information providers into active reasoning partners, and suggests that a significant limitation of current LLMs is that they lack debate or critical self-reflection (line 196). However, metacognition, debate, and self-reflection frameworks like Reflexion are widely used in the LLM literature and many LLMs have already employed metacognitive and self-reflection techniques to improve reasoning. Additionally, several LLM-based debate and challenge frameworks (e.g., Microsoft multi-agent MAI-DxO: https://arxiv.org/html/2506.22405v1) are already introduced for diagnostic purposes.  

A more comprehensive literature review and a comparison with multi-agent frameworks for medical diagnosis would be helpful. The tone and novelty claims could be adjusted accordingly.

 

Minor Issues

  • Duplicated content: A similar paragraph “While RAG provides the technical foundation…” is repeated twice (line 323 and line 327).
  • Typos: Some typos like “coope within the …” need to be corrected.
  • Inaccurate wording: ChatGPT-4 is described as a “standard LLM”, but in fact ChatGPT is a chatbot application built on top of an LLM.

 

Author Response

Thank you for your valuable feedback. Please see the attached PDF for our detailed point-by-point response. All manuscript revisions are highlighted.

Author Response File: Author Response.pdf

Reviewer 2 Report

Comments and Suggestions for Authors

refer to pdf attachment

Comments for author File: Comments.pdf

Author Response

Thank you for your valuable feedback. Please see the attached PDF for our detailed point-by-point response. All manuscript revisions are highlighted.

Author Response File: Author Response.pdf

Reviewer 3 Report

Comments and Suggestions for Authors

This is a timely and important paper that provides a profound diagnosis of the fundamental failure patterns of current large language models (LLMs) applied in clinical decision support systems. Overall, this study has significant practical application value, but further improvements are needed to strengthen the manuscript.
1. In the introduction section, the inherent challenges in medical diagnosis and clinical decision-making processes should be clearly explained.
2. Carefully check the writing of words in the text to ensure accuracy. For example, the word 'coop' on line 51.
3. How does "The Epidemic Router" in Section 3.4 analyze inputs to determine necessary cognitive roles? Please provide a detailed explanation.
4. Currently, the effectiveness is only demonstrated through case studies without providing quantitative indicators (such as diagnostic accuracy, etc.). It is recommended to supplement comparative experimental data to more objectively verify the effectiveness of the framework.
5. The pilot experiment was only conducted in one surgical center, and the number of participating clinical doctors and cases was not clearly stated. Suggest expanding the sample size to enhance the generalization of the method.
6. It is recommended to further expand the comparative experimental subjects.

Author Response

Thank you for your valuable feedback. Please see the attached PDF for our detailed point-by-point response. All manuscript revisions are highlighted.

Author Response File: Author Response.pdf

Reviewer 4 Report

Comments and Suggestions for Authors This paper proposes a "Dialogic Reasoning Framework" to transform Large Language Models (LLMs) from passive, acquiescent tools into active, critical thinking partners for medical diagnosis by addressing their inherent lack of "selfhood" and "initiative". 1. "Selfhood" and "Initiative" could be introduced in the introduction. 2. RAG has potential limitations, such as the quality and completeness of retrieved documents, and the risk of retrieving outdated information. 3. A brief explanation should be provided regarding the specific types of reliability metadata envisioned, as well as how they would be generated and presented to clinicians within the DiDi prototype. 4. In section 3.4, the Epistemic Router and Reliability Metadata Layer should be described in detail. 5. Could a quantitative metric be provided for the scenarios presented in Table 2? 6. Could author further elaborate on how "selfhood" and "initiative" are engineered to achieve their intended functionalities?  

Author Response

Thank you for your valuable feedback. Please see the attached PDF for our detailed point-by-point response. All manuscript revisions are highlighted.

Author Response File: Author Response.pdf

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

Thank you for addressing the comments.

Author Response

Thank you for acknowledging that we have addressed your comments. We sincerely appreciate your thorough and constructive feedback throughout the review process, which has significantly strengthened the quality and clarity of our manuscript.

Reviewer 2 Report

Comments and Suggestions for Authors

I have no more comments.

Author Response

Thank you for confirming that you have no further comments. We deeply appreciate your extraordinarily comprehensive and systematic evaluation in the first round, which has been invaluable in strengthening our manuscript.

Reviewer 3 Report

Comments and Suggestions for Authors

I think the paper can be accepted. 

Author Response

Thank you for your positive evaluation and recommendation for acceptance. We sincerely appreciate your constructive feedback throughout the review process, which has significantly improved the quality and clarity of our manuscript.

Reviewer 4 Report

Comments and Suggestions for Authors

The authors have addressed all my issues.

The only thing I would suggest is the No. of each chapter. It will help readers find the correct location easier.

Author Response

Thank you for your positive evaluation and for confirming that we have adequately addressed all your previous concerns.

We sincerely appreciate your constructive suggestion regarding chapter numbering. We have added section numbers throughout the entire manuscript to improve navigability and help readers locate specific content more easily. All sections and subsections now include hierarchical numbering which will facilitate cross-referencing.

We are grateful for your thorough evaluation, which has significantly strengthened the quality and clarity of our manuscript.

Back to TopTop