Generative AI-Integrated Virtual Agents and Simulations in Health Professions Education: A Systematic Review
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe analysis lacks an ontological distinction between traditional (deterministic) simulations and those based on GenAI (probabilistic). The authors should reflect on whether GenAI is merely a technical upgrade or if it alters the nature of simulation by introducing dialogic unpredictability. It may be pertinent to include the concept of "Dialogic Simulation" in the Introduction, explicating how algorithmic uncertainty demands new cognitive stances from the student.
The text highlights an increase in students' confidence and self-efficacy. However, there is a critical gap in the discussion regarding "overconfidence bias". Interacting with an AI programmed to be helpful can create a false sense of readiness. In the Discussion, the authors could cross-reference "satisfaction" data with "clinical accuracy" data, warning of the risk that students may feel proficient without necessarily mastering the rigor of real-world protocols.
The use of Social Presence Theory focuses on engagement and the humanization of the avatar. Nevertheless, the data suggest that the impact on theoretical knowledge is less pronounced than on practical skills. It is worth adding a reflection on "Cognitive Load," discussing whether the effort to maintain social interaction with the AI may act as a distraction that impairs the retention of complex medical concepts.
Author Response
1. The analysis lacks an ontological distinction between traditional (deterministic) simulations and those based on GenAI (probabilistic). The authors should reflect on whether GenAI is merely a technical upgrade or if it alters the nature of simulation by introducing dialogic unpredictability. It may be pertinent to include the concept of "Dialogic Simulation" in the Introduction, explicating how algorithmic uncertainty demands new cognitive stances from the student.
We agreed to add the reflect on whether GenAI is merely a technical upgrade or if it alters the nature of simulation by introducing dialogic unpredictability. However, GenAI’s algorithm can have certainty (predictability) based on the training methods, such as Reinforcement Learning from Human Feedback (RLHF), probability-based next-word prediction, and controlled temperature parameters naturally steer the AI toward producing highly standardized, textbook-like responses (Ouyang et al., 2022). For health professions education, this predictable typicality is advantageous. It allows institutions to efficiently and consistently generate quintessential clinical scenarios (e.g., standard patient histories and typical symptom presentations).
2. The text highlights an increase in students' confidence and self-efficacy. However, there is a critical gap in the discussion regarding "overconfidence bias". Interacting with an AI programmed to be helpful can create a false sense of readiness. In the Discussion, the authors could cross-reference "satisfaction" data with "clinical accuracy" data, warning of the risk that students may feel proficient without necessarily mastering the rigor of real-world protocols.
We appreciate this inspiration, and we have added some comments on Section 6.3. Educational Implications:
“The potential of GIVAS in health education is to help improve teaching efficiency and systematic standardization (Xu et al., 2025). A significant finding across the reviewed studies (e.g., D'Alessio et al., 2024; Zhao et al., 2025) is the high level of student-reported satisfaction and self-efficacy after having training with GIVAS (standardly designed training). However, we identify a critical evidence gap: there is a lack of objective clinical accuracy data to cross-reference these subjective feelings. This raises the risk of overconfidence bias, where the helpful and cooperative nature of GenAI may mislead students into a fake sense of readiness. But in the field of health and medicine, the essence of much education and training is in facing complex and dynamic individuals, which require experience, intuition, and ethical judgment to navigate. Findings did not disclose enough of GIVAS’s assistance on ethics reflection and theoretical knowledge acquisition. Mastering basic knowledge and standardized skills through interaction with GIVAS can be simply achieved (Srinivasa et al., 2022), yet it might be difficult to safeguard learners' authentic ethical perceptions and self-efficacy, and further, prevent the 'mechanization' of human empathy during these AI interactions.”
3. The use of Social Presence Theory focuses on engagement and the humanization of the avatar. Nevertheless, the data suggest that the impact on theoretical knowledge is less pronounced than on practical skills. It is worth adding a reflection on "Cognitive Load," discussing whether the effort to maintain social interaction with the AI may act as a distraction that impairs the retention of complex medical concepts.
We have added discussion about cognitive load. Please see Section 6.1. Theoretical Implications:
“The above mechanisms can be further related to established learning theories that many of the reviewed studies have not noted. Firstly, the use of simulated, interactive environments aligns closely with experiential learning theory (Kolb, 1984), where learners experience knowledge being transferred by experience through the “Socratic” conversation with GIVAS. For example, because the "patient" is a virtual AI agent, learners can immediately retry the clinical scenario with a different approach to see how the outcome changes. Secondly, the research findings broadly align with cognitive load theory, as GIVAS may help reduce or add cognitive load depending on the design and functionality of the GIVAS. For example, the use of GIVAS as a tutor or guide to provide learners support to understand clinical documents and generate scenarios can reduced cognitive load. However, research shows that instructional techniques that help novices (e.g. step-by-step guidance) can actually hinder experts, which is seen as the expertise reversal effect (Kalyuga, 2007). Reliance on AI-generated responses may also limit opportunities for deeper cognitive processing and independent reasoning, particularly if learners adopt a passive interaction style. The foundations of social constructivism also reveal that learners may get used to social roles of these virtual agents (i.e., virtual patients) even if they know it is a chatbot, which depends on the extent to which users perceive GIVAS in their interactions as “real social individuals” (Vygotsky & Cole, 1978).”
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThis systematic review addresses a timely and important topic — the application of generative AI-integrated virtual agents and simulations (GIVAS) in health professions education. The scope is well-defined, the three research questions are clearly articulated, and the use of a pedagogy-oriented SPIDER framework to guide analysis is a thoughtful methodological choice. The synthesis across educational impact, technical features, and limitations provides a useful overview for educators and researchers in this space. However, several methodological concerns — particularly regarding the consistency of the PRISMA flow, the inclusion of studies that do not meet the review's own criteria, the absence of a quality appraisal process, and persistent structural formatting errors — require substantive attention before the manuscript can be considered for publication.
- The most significant methodological issue concerns the inclusion of studies that do not meet the review's own stated criteria. The exclusion criteria explicitly state that "studies use non-GenAI models" should be excluded. However, at least two included studies use rule-based NLP systems rather than generative AI: Study ID 3 (Nakagawa et al., 2022) is described in Table 2 as using a "Rule-based NLP, keyword detection/based AI chatbot," and Study ID 13 (Shorey et al., 2019) is described as using a "Rule-based NLP chatbot (limited memory AI), Google Dialogflow." Neither of these constitutes generative AI as defined in the review's own introduction, which describes GenAI as systems powered by large language models capable of dynamic, context-aware conversation generation. The authors must either remove these studies from the review and update all downstream analyses, tables, and counts accordingly, or provide a clear and explicit justification for why rule-based NLP systems qualify as generative AI within their framework. This issue directly undermines the construct validity of the review and the trustworthiness of the synthesized findings.
- There are substantial discrepancies in the PRISMA flow diagram and accompanying text that need to be reconciled. The body text states that "756 studies appeared from the preliminary searching phase" and that "after removing duplicated studies and non-qualified ones (n=738), a final set of 18 studies was selected." This implies a single step from 756 to 18. However, the PRISMA diagram tells a different story: 729 database records plus 27 from other sources yield 756 total; after duplicate removal, 178 remain; 55 are then screened (with 0 excluded at screening); 18 are assessed at full text, with 37 excluded. The transition from 178 records after duplicate removal to only 55 screened is unexplained — what happened to the remaining 123 records? Additionally, the text's description of removing 738 records in a single step does not correspond to the multi-stage process shown in the diagram. Furthermore, the diagram shows a "Qualitative studies (n=3)" box branching off before the final included box of n=18, but it is unclear whether these 3 qualitative studies are a subset of the 18 or additional studies. The authors should carefully verify and correct all numbers to ensure full consistency between the text and the PRISMA diagram, and clearly explain each screening step.
- The absence of any formal quality appraisal or risk-of-bias assessment is a notable limitation for a study that labels itself a "systematic review." The authors acknowledge this, citing the diversity of included study designs (development studies, qualitative research, case demonstrations, mixed methods) as the reason. While it is true that no single risk-of-bias tool applies uniformly to all these designs, this does not preclude quality assessment entirely. Tools such as the Mixed Methods Appraisal Tool (MMAT) or the Joanna Briggs Institute (JBI) critical appraisal checklists are specifically designed to accommodate methodological diversity. Without any quality appraisal, readers cannot evaluate whether the conclusions are driven primarily by well-conducted studies or by pilot-stage projects with minimal rigor. The authors should either conduct and report a quality assessment using an appropriate tool, or, at a minimum, provide a structured summary of each study's methodological strength (e.g., sample size, presence of a comparison group, use of validated measures) so that readers can weigh the evidence accordingly.
- The review was not prospectively registered in any systematic review registry (e.g., PROSPERO). While this is acknowledged in the manuscript, prospective registration is a core element of systematic review methodology to protect against selective reporting. The authors should discuss how they ensured transparency in their review process in the absence of registration, and should consider registering the protocol retrospectively or providing the full protocol as supplementary material.
- The manuscript contains pervasive section numbering errors. The Methodology section lacks a number entirely, while all Discussion subsections (Theoretical Implications, Technological Implications, Educational Implications, and Limitations) are identically numbered "3.1." Similarly, the Findings subsections labeled "3.1.1" repeat the same number for three distinct subsections (Enhanced Engagement, Communication Skills, and Personalized Learning). These errors are likely artifacts of formatting but significantly impair readability and should be corrected throughout.
- The inclusion criteria specify that studies must be "published or pre-printed in peer-reviewed journal articles or conference proceedings," yet the exclusion criteria state that "studies from non-peer-reviewed sources" should be excluded. Several included studies are arXiv preprints (Chen et al., 2024; Chu & Goodell, 2024; Yan & Alterovitz, 2024), a medRxiv preprint (Mool et al., 2024), and an SSRN working paper (Mittenentzwei et al., 2024). Preprint servers do not conduct peer review. The authors should clarify whether preprints were intentionally included (in which case the exclusion criterion should be revised and the rationale explained) or whether these were included in error.
- The initial screening was conducted by a single researcher, with a second researcher verifying the decisions afterward. PRISMA 2020 guidelines recommend independent dual screening to minimize selection bias. The authors should acknowledge this deviation from best practice and discuss its potential impact on the completeness and reproducibility of the review.
- The search time frame extends back to 2019, justified as covering "the last six years." However, large language models capable of generative dialogue (e.g., ChatGPT) only became publicly available in late 2022, and the review explicitly focuses on generative AI. The rationale for including a four-year window (2019–2022) predating the availability of GenAI tools is unclear and may partly explain why non-GenAI studies were included. The authors should either narrow the time frame to align with the emergence of GenAI or provide a clear justification for the broader window.
- The acronym "GIVAS" is defined in the text but is occasionally replaced by "GIATS" (e.g., in sections discussing RQ2 technical features), which appears to be a typographical error. The authors should conduct a thorough search-and-replace to ensure terminological consistency.
- Table 4 lists four themes under RQ1 ("Enhanced Educational and clinical skills," "Engagement and Human-like Interaction," "Communication," and "Personalized Learning"), but the introductory sentence to Table 4 states "three aspects" of GIVAS for learning outcomes. This count should be corrected to match the table content.
Author Response
Reviewer2
This systematic review addresses a timely and important topic — the application of generative AI-integrated virtual agents and simulations (GIVAS) in health professions education. The scope is well-defined, the three research questions are clearly articulated, and the use of a pedagogy-oriented SPIDER framework to guide analysis is a thoughtful methodological choice. The synthesis across educational impact, technical features, and limitations provides a useful overview for educators and researchers in this space. However, several methodological concerns — particularly regarding the consistency of the PRISMA flow, the inclusion of studies that do not meet the review's own criteria, the absence of a quality appraisal process, and persistent structural formatting errors — require substantive attention before the manuscript can be considered for publication.
- The most significant methodological issue concerns the inclusion of studies that do not meet the review's own stated criteria. The exclusion criteria explicitly state that "studies use non-GenAI models" should be excluded. However, at least two included studies use rule-based NLP systems rather than generative AI: Study ID 3 (Nakagawa et al., 2022) is described in Table 2 as using a "Rule-based NLP, keyword detection/based AI chatbot," and Study ID 13 (Shorey et al., 2019) is described as using a "Rule-based NLP chatbot (limited memory AI), Google Dialogflow." Neither of these constitutes generative AI as defined in the review's own introduction, which describes GenAI as systems powered by large language models capable of dynamic, context-aware conversation generation. The authors must either remove these studies from the review and update all downstream analyses, tables, and counts accordingly, or provide a clear and explicit justification for why rule-based NLP systems qualify as generative AI within their framework. This issue directly undermines the construct validity of the review and the trustworthiness of the synthesized findings.
We have updated the PRISMA diagram to follow the 2020 standards. We explicitly clarified that out of 178 records, 123 were excluded at the screening stage. Furthermore, we excluded the two rule-based studies identified by the reviewers (ID 3 and ID 13) during the eligibility assessment, bringing the total excluded full-text articles to 39 with clearly stated reasons.
- There are substantial discrepancies in the PRISMA flow diagram and accompanying text that need to be reconciled. The body text states that "756 studies appeared from the preliminary searching phase" and that "after removing duplicated studies and non-qualified ones (n=738), a final set of 18 studies was selected." This implies a single step from 756 to 18. However, the PRISMA diagram tells a different story: 729 database records plus 27 from other sources yield 756 total; after duplicate removal, 178 remain; 55 are then screened (with 0 excluded at screening); 18 are assessed at full text, with 37 excluded. The transition from 178 records after duplicate removal to only 55 screened is unexplained — what happened to the remaining 123 records? Additionally, the text's description of removing 738 records in a single step does not correspond to the multi-stage process shown in the diagram. Furthermore, the diagram shows a "Qualitative studies (n=3)" box branching off before the final included box of n=18, but it is unclear whether these 3 qualitative studies are a subset of the 18 or additional studies. The authors should carefully verify and correct all numbers to ensure full consistency between the text and the PRISMA diagram, and clearly explain each screening step.
We have thoroughly revised Section 5.1 to ensure it exclusively reflects generative AI applications. Specifically, we have removed all references to the rule-based systems previously included (ID 3 and ID 13). The categorical counts (n values) for AI avatars and conversational chatbots have been updated to reflect the final 16 generative AI-focused studies. We also updated the Study IDs (e.g., ID 4 became ID 3) to maintain consistency with the revised Table 2.
- The absence of any formal quality appraisal or risk-of-bias assessment is a notable limitation for a study that labels itself a "systematic review." The authors acknowledge this, citing the diversity of included study designs (development studies, qualitative research, case demonstrations, mixed methods) as the reason. While it is true that no single risk-of-bias tool applies uniformly to all these designs, this does not preclude quality assessment entirely. Tools such as the Mixed Methods Appraisal Tool (MMAT) or the Joanna Briggs Institute (JBI) critical appraisal checklists are specifically designed to accommodate methodological diversity. Without any quality appraisal, readers cannot evaluate whether the conclusions are driven primarily by well-conducted studies or by pilot-stage projects with minimal rigor. The authors should either conduct and report a quality assessment using an appropriate tool, or, at a minimum, provide a structured summary of each study's methodological strength (e.g., sample size, presence of a comparison group, use of validated measures) so that readers can weigh the evidence accordingly.
We thank the reviewer for this important methodological point. We agree that transparency regarding the rigor of included studies is essential for a systematic review. However, as the reviewer correctly noted, the included studies are highly heterogeneous, ranging from technical development and case demonstrations to qualitative and mixed-methods research. MMAT or JBI may be unsuitable for technical development studies, so we have followed your secondary recommendation to provide a structured summary of each study’s methodological strength. We have updated Table 2 to explicitly include key indicators of rigor for each study: aim, sample size, study type, target group, etc. This addition allows readers to clearly distinguish between early-stage pilot projects and more rigorous empirical evaluations without imposing a reductive 'quality score' that might misrepresent the diverse contributions of the reviewed literature. We hope this structured approach addresses your concerns while maintaining the clarity of our synthesis.
- The review was not prospectively registered in any systematic review registry (e.g., PROSPERO). While this is acknowledged in the manuscript, prospective registration is a core element of systematic review methodology to protect against selective reporting. The authors should discuss how they ensured transparency in their review process in the absence of registration, and should consider registering the protocol retrospectively or providing the full protocol as supplementary material.
Thank you for your suggestions. We will consider registering the protocol for suture studies.
- The manuscript contains pervasive section numbering errors. The Methodology section lacks a number entirely, while all Discussion subsections (Theoretical Implications, Technological Implications, Educational Implications, and Limitations) are identically numbered "3.1." Similarly, the Findings subsections labeled "3.1.1" repeat the same number for three distinct subsections (Enhanced Engagement, Communication Skills, and Personalized Learning). These errors are likely artifacts of formatting but significantly impair readability and should be corrected throughout.
We have conducted a comprehensive review of the document's formatting. The section numbering has been fully corrected and standardized to ensure a logical and sequential flow:
The Methodology section is clearly numbered as Section 3. The Data Analysis and Findings sections follow as Section 4 and Section 5, respectively. The Discussion and Conclusion sections are correctly numbered as Section 6 and Section 7. Furthermore, we have updated all subsection headings to align with their respective primary sections (e.g., Findings subsections are now 5.1, 5.2, etc., and Discussion subsections are 6.1, 6.2, etc.). This addresses the "3.1.1" and "3.1" numbering artifacts mentioned by the reviewer, which were likely due to heading style conflicts in the previous version. We believe the readability and structural hierarchy of the manuscript are now significantly improved.
- The inclusion criteria specify that studies must be "published or pre-printed in peer-reviewed journal articles or conference proceedings," yet the exclusion criteria state that "studies from non-peer-reviewed sources" should be excluded. Several included studies are arXiv preprints (Chen et al., 2024; Chu & Goodell, 2024; Yan & Alterovitz, 2024), a medRxiv preprint (Mool et al., 2024), and an SSRN working paper (Mittenentzwei et al., 2024). Preprint servers do not conduct peer review. The authors should clarify whether preprints were intentionally included (in which case the exclusion criterion should be revised and the rationale explained) or whether these were included in error.
We appreciate this comment. Generative AI (GenAI) is one of the fastest-growing fields in scientific research. The traditional peer review cycle (typically 12-24 months from submission to publication) lags far behind the pace of technological innovation. Therefore, our reviews include a small number of preprint studies. Due to this reason, we have also updated our inclusion and exclusion criteria.
- The initial screening was conducted by a single researcher, with a second researcher verifying the decisions afterward. PRISMA 2020 guidelines recommend independent dual screening to minimize selection bias. The authors should acknowledge this deviation from best practice and discuss its potential impact on the completeness and reproducibility of the review.
Our study selection was conducted in two stages: (1) title and abstract screening, and (2) full-text review. Two reviewers independently screened all records against the predefined inclusion and exclusion criteria (dual screening). Discrepancies were resolved through discussion, and where necessary, consultation with a third reviewer. Data extraction was also conducted independently by two reviewers using a standardized data extraction form. Although formal inter-rater reliability statistics (e.g., Cohen’s kappa) were not calculated, consistency between reviewers was ensured through independent screening and regular discussions to resolve discrepancies
- The search time frame extends back to 2019, justified as covering "the last six years." However, large language models capable of generative dialogue (e.g., ChatGPT) only became publicly available in late 2022, and the review explicitly focuses on generative AI. The rationale for including a four-year window (2019–2022) predating the availability of GenAI tools is unclear and may partly explain why non-GenAI studies were included. The authors should either narrow the time frame to align with the emergence of GenAI or provide a clear justification for the broader window.
While 2022 was the year of ChatGPT's public release, we set the search baseline in 2019 to capture the foundational transition to Generative AI. The release of GPT-2 in 2019 marked the functional shift from rule-based chatbots to true generative models. Including the 2019–2022 period allows us to synthesize the early pioneering efforts in health education that predated the ChatGPT hype, providing a more comprehensive evolutionary perspective of the technology's application. Reference: Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI blog, 1(8), 9.
- The acronym "GIVAS" is defined in the text but is occasionally replaced by "GIATS" (e.g., in sections discussing RQ2 technical features), which appears to be a typographical error. The authors should conduct a thorough search-and-replace to ensure terminological consistency.
We have performed a comprehensive search-and-replace throughout the manuscript. The term "GIATS" has been completely removed, and the acronym "GIVAS" is now used consistently across all sections to ensure terminological clarity.
- Table 4 lists four themes under RQ1 ("Enhanced Educational and clinical skills," "Engagement and Human-like Interaction," "Communication," and "Personalized Learning"), but the introductory sentence to Table 4 states "three aspects" of GIVAS for learning outcomes. This count should be corrected to match the table content.
We have corrected the introductory sentence to Table 4 to state "four aspects," aligning perfectly with the four themes listed for RQ1 in the table and the subsequent subsections.
Author Response File:
Author Response.pdf
Reviewer 3 Report
Comments and Suggestions for AuthorsThe following are the key revisions that must be addressed before the manuscript can be reconsidered for publication:
- Revise the PRISMA flow diagram to conform fully to PRISMA 2020 standards, including a complete breakdown of exclusions by reason.
- Include a formal quality appraisal of the included studies using an appropriate tool (e.g., MMAT or CASP checklists) and discuss how study quality influenced the confidence in the synthesized findings.
- Either justify the inclusion of Studies 3 and 13 (which appear to use non-generative AI systems) with a clear and well-reasoned definition of GenAI, or exclude them and revise the findings accordingly.
- Fix all structural and numbering errors throughout the manuscript. Integrate the Introduction and Related Work sections or clearly differentiate their purpose. Renumber all headings consistently.
- Demonstrate explicitly how the SPIDER framework was applied to each study. Consider adding a supplementary analytical matrix showing the coding of all 18 studies across the five SPIDER dimensions.
- Substantially revise the Discussion section to deepen the theoretical analysis. Engage more explicitly with established learning theories (e.g., experiential learning, cognitive load, social constructivism) and connect them to specific findings from the reviewed studies.
- Revise the scope of conclusions to accurately reflect the preliminary and geographically limited nature of the evidence base. Avoid overstating the strength of findings.
- Report inter-rater reliability statistics for the screening and data extraction process.
- Engage professional English language editing services and thoroughly revise the entire manuscript for grammar, coherence, and clarity.
- Reconsider and expand the Conclusion section to provide substantive synthesis, not just a restatement of the abstract.
The English writing throughout the manuscript requires substantial revision. There are numerous grammatical errors, awkward phrasing, and sentences that are difficult to parse. The Minor Comments section provides several examples below. Broadly speaking, it appears that different contributors wrote some passages without sufficient editorial harmonization. The Discussion section, in particular, contains run-on sentences and shifts in register that disrupt the reading experience. The authors are strongly encouraged to engage the services of a professional academic English editor before resubmission.
Author Response
Reviewer3
The following are the key revisions that must be addressed before the manuscript can be reconsidered for publication:
- Revise the PRISMA flow diagram to conform fully to PRISMA 2020 standards, including a complete breakdown of exclusions by reason.
We updated our literature screening process to strictly follow the PRISMA 2020 guidelines. A fully updated PRISMA flow diagram has been added to the revised manuscript as Figure 1.
- Include a formal quality appraisal of the included studies using an appropriate tool (e.g., MMAT or CASP checklists) and discuss how study quality influenced the confidence in the synthesized findings.
We have ensured that the manuscript’s methodology and results sections are strictly aligned with the PRISMA 2020 standards, as demonstrated by the newly added PRISMA Flow Diagram. Given that we are addressing a high volume of comments from five independent reviewers, we have prioritized integrating these essential elements directly into the text to ensure clarity and flow.
- Either justify the inclusion of Studies 3 and 13 (which appear to use non-generative AI systems) with a clear and well-reasoned definition of GenAI, or exclude them and revise the findings accordingly.
Study 3 and 13 have been excluded.
- Fix all structural and numbering errors throughout the manuscript. Integrate the Introduction and Related Work sections or clearly differentiate their purpose. Renumber all headings consistently.
All structural and numbering errors are corrected.
- Demonstrate explicitly how the SPIDER framework was applied to each study. Consider adding a supplementary analytical matrix showing the coding of all 18 studies across the five SPIDER dimensions.
We have updated our explanation for SPIDER, which “guided analysis across five dimensions: (1) learner context (e.g., medical students, trainees, clinicians), (2) pedagogical phenomenon of interest (e.g., communication skills, clinical reasoning, engagement), (3) educational design and agent role (e.g., AI as simulated patient, tutor, or instructional avatar), (4) learning mechanisms (e.g., experiential practice, scaffolding, repetition, immersion), and (5) type of evidence reported (e.g., pilot study, mixed-methods, usability evaluation). This framework aims to guide data analysis based on pedagogical perspectives by linking AI agent configurations to learning contexts, educational designs, and underlying learning mechanisms.”
- Substantially revise the Discussion section to deepen the theoretical analysis. Engage more explicitly with established learning theories (e.g., experiential learning, cognitive load, social constructivism) and connect them to specific findings from the reviewed studies.
We have substantially revised the Discussion, especially added relevant discussion on experiential learning, cognitive load, social constructivism. “The above mechanisms can be further related to established learning theories that many of the reviewed studies have not noted. Firstly, the use of simulated, interactive environments aligns closely with experiential learning theory (Kolb, 1984), where learners experience knowledge being transferred by experience through the “Socratic” conversation with GIVAS. For example, because the "patient" is a virtual AI agent, learners can immediately retry the clinical scenario with a different approach to see how the outcome changes. Secondly, the research findings broadly align with cognitive load theory, as GIVAS may help reduce or add cognitive load depending on the design and functionality of the GIVAS. For example, the use of GIVAS as a tutor or guide to provide learners support to understand clinical documents and generate scenarios can reduce cognitive load. However, research shows that instructional techniques that help novices (e.g., step-by-step guidance) can actually hinder experts, which is seen as the expertise reversal effect (Kalyuga, 2007). Reliance on AI-generated responses may also limit opportunities for deeper cognitive processing and independent reasoning, particularly if learners adopt a passive interaction style. The foundations of social constructivism also reveal that learners may get used to social roles of these virtual agents (i.e., virtual patients) even if they know it is a chatbot, which depends on the extent to which users perceive GIVAS in their interactions as “real social individuals” (Vygotsky & Cole, 1978).”
- Revise the scope of conclusions to accurately reflect the preliminary and geographically limited nature of the evidence base. Avoid overstating the strength of findings.
Yes, we have added the scope of conclusions to accurately reflect the preliminary and geographically limited nature of the evidence base in the limitations section. Please see pp.23
“In this review, although the findings suggest that GIVAS hold considerable promise in enhancing learning and training in health professions education, the current evidence base remains preliminary. Many of the included studies are small-scale, pilot, or development-focused, with limited longitudinal evaluation. The geographical distribution of studies is relatively narrow (most studies are from developed countries), with evidence concentrated in specific regions and educational contexts. This limits the generalizability of the findings across diverse healthcare systems and learner populations.”
- Report inter-rater reliability statistics for the screening and data extraction process.
Data extraction was also conducted independently by two reviewers using a standardized data extraction form. Formal inter-rater reliability statistics (e.g., Cohen’s kappa) were not calculated, whereas consistency between reviewers was ensured through independent screening and regular discussions to resolve discrepancies. All reviewed studies were critically examined for methodological clarity, relevance to the research questions, and completeness of reported findings during the screening and data extraction process.
- Engage professional English language editing services and thoroughly revise the entire manuscript for grammar, coherence, and clarity.
We have improved the layout, grammar, and language thoroughly with the help of two native speakers from the authors. We have also undertaken a full proofreading to correct grammatical, stylistic, and spelling errors.
- Reconsider and expand the Conclusion section to provide substantive synthesis, not just a restatement of the abstract.
We have updated our conclusion. However, due to the word count, we cannot expand too much.
Author Response File:
Author Response.pdf
Reviewer 4 Report
Comments and Suggestions for AuthorsThis generative intelligent virtual agents and simulations (GIVAS) in health professions education which contributes a pedagogy-oriented synthesis and explicates the learning mechanisms is highly innovative and original. However for its validity 18 sample is not enough. Nevertheless, the study is very comprehensive and adequate.
Author Response
Reviewer 4
This generative intelligent virtual agents and simulations (GIVAS) in health professions education which contributes a pedagogy-oriented synthesis and explicates the learning mechanisms is highly innovative and original. However for its validity 18 sample is not enough. Nevertheless, the study is very comprehensive and adequate.
We appreciate the reviewer’s observation. Our final selection is relatively small because we strictly included only empirical studies specifically focused on Generative AI. Given the novelty of this field, the vast majority of initial hits did not meet these rigorous criteria.
Author Response File:
Author Response.pdf
Round 2
Reviewer 2 Report
Comments and Suggestions for AuthorsI appreciate that the authors have addressed several of the previously raised concerns. However, several substantive concerns from the previous round remain unresolved or have only been superficially addressed, and one of the responses is clearly inadequate. In a few cases the response letter also describes revisions that do not appear in the actual manuscript. The following points require further attention before the manuscript can be considered for publication.
- The section numbering errors flagged in the previous round have not been corrected, despite the authors' claim that "the section numbering has been fully corrected and standardized to ensure a logical and sequential flow." In the current version, the Findings section still contains three consecutive subsections all labeled "3.1.1" (Enhanced Engagement and Human-like Interaction; Communication Skills in Simulations; Personalized Learning and Tailored Learning Environments), and the Discussion section still contains four subsections all labeled "3.1." (Theoretical Implications, Technological Implications [unnumbered in places], Educational Implications, and Limitations). In addition, the section "Technical Features and Functionalities (RQ2)" is labeled "3.1." even though it sits within the Findings section, while the immediately following section "Challenges of GIVAS (RQ3)" is correctly labeled "5.3." The authors' description of the renumbering scheme in the response letter is sound; the problem is that this scheme has not actually been implemented in the manuscript text. I strongly recommend that the authors verify the heading styles in the source file and re-export the manuscript, then proofread the printed PDF to confirm the corrections are visible.
- The response to the prospective registration concern is inadequate. The authors' reply — "Thank you for your suggestions. We will consider registering the protocol for suture studies" [sic] — neither registers the protocol retrospectively, nor provides the full protocol as supplementary material, nor adds any meaningful discussion in the manuscript about how transparency was ensured in the absence of registration. The current text simply repeats the acknowledgment that "this systematic review … was therefore not prospectively registered in a systematic review registry," without addressing the substantive point. Given that prospective registration is a core element of systematic review methodology, I recommend that the authors at minimum (a) register the protocol retrospectively (e.g., with PROSPERO or the Open Science Framework, both of which accept retrospective registrations with a clear note), (b) provide the complete review protocol as supplementary material, or (c) add a substantive paragraph in the Methodology or Limitations section discussing the specific safeguards (e.g., a priori inclusion criteria, dual screening, fixed search strings) that were used to mitigate the risks of selective reporting in the absence of registration. I would also note that the typographical error "suture studies" in the response letter, while not present in the manuscript itself, suggests that the response was not carefully proofread.
- The PRISMA flow diagram and the body text are still not fully consistent with each other, and the new diagram introduces a fresh error. The body text on page 5 states that "After removing duplicated and non-qualifying studies (n=740), a final set of 16 studies was selected," which collapses the multi-stage PRISMA process into a single step and does not correspond to the staged figures in Figure 1. Within Figure 1 itself, the box "Reports assessed for eligibility (n=16)" is structurally incorrect under PRISMA 2020 conventions: with 118 reports sought for retrieval and 0 not retrieved, the number assessed for eligibility should be 118, and the subsequent "Reports excluded" box (42 + 60 = 102) should reduce this to 16 included studies. Additionally, the figures cited in the response letter — "out of 178 records, 123 were excluded at the screening stage" and "39 excluded full-text articles" — do not match the figures in the actual flowchart (178 duplicates removed before screening, 460 excluded at screening, 102 excluded at full-text). I recommend that the authors (a) revise the body text to describe each PRISMA stage separately rather than as a single step, (b) correct the "Reports assessed for eligibility" box in Figure 1 to reflect the count before exclusion (n=118) rather than after, and (c) ensure the response letter and the manuscript reference the same set of numbers.
- The manuscript now contains two contradictory descriptions of the screening procedure that need to be reconciled. On page 5 (Section 3.1, Data Extraction and Synthesis), the text states: "The initial screening of all studies, including keywords, titles, abstracts, and full texts, was conducted by one researcher. Subsequently, a second researcher independently verified and approved the screening decisions." On pages 6–7, however, the text states: "Study selection was conducted in two stages: (1) title and abstract screening, and (2) full-text review. Two reviewers screened all records against the predefined inclusion and exclusion criteria respectively. Discrepancies were resolved through discussion …" These two descriptions cannot both be true: the first describes sequential single-reviewer screening with verification, while the second describes independent dual screening. It appears the second paragraph was added in response to the previous review without removing the original contradictory statement. The authors should delete the incompatible passage and retain only the description that accurately reflects what was done. They should also explicitly discuss the potential implications for selection bias and reproducibility, as recommended in the previous round, rather than simply asserting that consistency was ensured.
- The response to the quality appraisal concern is weaker than the response letter implies. The authors state that they have "updated Table 2 to explicitly include key indicators of rigor for each study: aim, sample size, study type, target group, etc." However, these columns appear to be largely the same descriptive metadata that were already present in the previous version, and they do not function as rigor indicators in the methodological sense intended by the previous comment. In particular, the previous review specifically suggested adding columns such as "presence of a comparison group" and "use of validated measures," neither of which has been included. I recommend that the authors either (a) genuinely supplement Table 2 with explicit methodological rigor indicators — for example, presence/absence of a comparison or control condition, type of outcome measure (validated instrument vs. self-developed survey vs. usability rating), and whether outcomes were objective or self-reported — or (b) apply a tool such as the MMAT, which is designed to accommodate methodological diversity across qualitative, quantitative, and mixed-methods studies and would be applicable to most of the included studies. If technical-development studies are deemed unsuitable for MMAT, those studies could be appraised with a brief structured commentary instead, with the rationale clearly stated.
- The justification for the 2019 starting point of the search window is weaker in the manuscript than in the response letter. The response letter argues persuasively that GPT-2 (released in 2019) marked the functional shift from rule-based chatbots to true generative models, and cites Radford et al. (2019) in support. However, the manuscript text only mentions GPT-3 (2020) and GPT-4 (2023), neither of which justifies a 2019 starting point. I recommend that the authors integrate the GPT-2 rationale (and the Radford et al. reference) directly into the Methodology section so that the time-frame justification is transparent to readers without access to the response letter.
Author Response
Review 2 (Second round)
I appreciate that the authors have addressed several of the previously raised concerns. However, several substantive concerns from the previous round remain unresolved or have only been superficially addressed, and one of the responses is clearly inadequate. In a few cases the response letter also describes revisions that do not appear in the actual manuscript. The following points require further attention before the manuscript can be considered for publication.
- The section numbering errors flagged in the previous round have not been corrected, despite the authors' claim that "the section numbering has been fully corrected and standardized to ensure a logical and sequential flow." In the current version, the Findings section still contains three consecutive subsections all labeled "3.1.1" (Enhanced Engagement and Human-like Interaction; Communication Skills in Simulations; Personalized Learning and Tailored Learning Environments), and the Discussion section still contains four subsections all labeled "3.1." (Theoretical Implications, Technological Implications [unnumbered in places], Educational Implications, and Limitations). In addition, the section "Technical Features and Functionalities (RQ2)" is labeled "3.1." even though it sits within the Findings section, while the immediately following section "Challenges of GIVAS (RQ3)" is correctly labeled "5.3." The authors' description of the renumbering scheme in the response letter is sound; the problem is that this scheme has not actually been implemented in the manuscript text. I strongly recommend that the authors verify the heading styles in the source file and re-export the manuscript, then proofread the printed PDF to confirm the corrections are visible.
We carefully rechecked the current revised manuscript and were unable to locate the numbering issues described by the reviewer. In the version currently uploaded to the submission system, the section numbering has been corrected and appears sequential throughout the manuscript.
Specifically:
- The Findings section contains three distinct subsections numbered 5.1, 5.2, and 5.3.
- The Discussion section contains four subsections numbered 6.1, 6.2, 6.3, and 6.4.
- “Technical Features and Functionalities (RQ2)” is numbered 5.2.
- “Challenges of GIVAS (RQ3)” is numbered 5.3.
No duplicate subsection numbers (e.g., repeated 3.1.1 or 3.1 headings) are present in the version available to the authors.
Therefore, we believe the reviewer may have been referring to an earlier version of the manuscript or a PDF generated prior to the latest revision. Nevertheless, we have re-checked all heading styles and numbering in the source file and confirmed that the numbering is correct and consistent throughout the current manuscript.
We would appreciate the Editor's confirmation that the reviewer assessed the most recent revised version.
- The response to the prospective registration concern is inadequate. The authors' reply — "Thank you for your suggestions. We will consider registering the protocol for suture studies" [sic] — neither registers the protocol retrospectively, nor provides the full protocol as supplementary material, nor adds any meaningful discussion in the manuscript about how transparency was ensured in the absence of registration. The current text simply repeats the acknowledgment that "this systematic review … was therefore not prospectively registered in a systematic review registry," without addressing the substantive point. Given that prospective registration is a core element of systematic review methodology, I recommend that the authors at minimum (a) register the protocol retrospectively (e.g., with PROSPERO or the Open Science Framework, both of which accept retrospective registrations with a clear note), (b) provide the complete review protocol as supplementary material, or (c) add a substantive paragraph in the Methodology or Limitations section discussing the specific safeguards (e.g., a priori inclusion criteria, dual screening, fixed search strings) that were used to mitigate the risks of selective reporting in the absence of registration. I would also note that the typographical error "suture studies" in the response letter, while not present in the manuscript itself, suggests that the response was not carefully proofread.
We acknowledge the importance of protocol registration for systematic reviews and have clearly stated in the manuscript that this review was not prospectively registered. However, it is not necessary for all types of systematic reviews.
The editor was also aware of the practical limitations of conducting retrospective registration and additional appraisal procedures for this review and specifically informed us that no further modifications were required in this regard.
- The PRISMA flow diagram and the body text are still not fully consistent with each other, and the new diagram introduces a fresh error. The body text on page 5 states that "After removing duplicated and non-qualifying studies (n=740), a final set of 16 studies was selected," which collapses the multi-stage PRISMA process into a single step and does not correspond to the staged figures in Figure 1. Within Figure 1 itself, the box "Reports assessed for eligibility (n=16)" is structurally incorrect under PRISMA 2020 conventions: with 118 reports sought for retrieval and 0 not retrieved, the number assessed for eligibility should be 118, and the subsequent "Reports excluded" box (42 + 60 = 102) should reduce this to 16 included studies. Additionally, the figures cited in the response letter — "out of 178 records, 123 were excluded at the screening stage" and "39 excluded full-text articles" — do not match the figures in the actual flowchart (178 duplicates removed before screening, 460 excluded at screening, 102 excluded at full-text). I recommend that the authors (a) revise the body text to describe each PRISMA stage separately rather than as a single step, (b) correct the "Reports assessed for eligibility" box in Figure 1 to reflect the count before exclusion (n=118) rather than after, and (c) ensure the response letter and the manuscript reference the same set of numbers.
The PRISMA flow diagram has been revised. The number of reports assessed for eligibility was corrected from n = 16 to n = 118 to ensure consistency with the full-text screening process. The exclusion counts (42 + 60 = 102) and final inclusion count (n = 16) are now fully aligned with PRISMA 2020 reporting requirements.
- The manuscript now contains two contradictory descriptions of the screening procedure that need to be reconciled. On page 5 (Section 3.1, Data Extraction and Synthesis), the text states: "The initial screening of all studies, including keywords, titles, abstracts, and full texts, was conducted by one researcher. Subsequently, a second researcher independently verified and approved the screening decisions." On pages 6–7, however, the text states: "Study selection was conducted in two stages: (1) title and abstract screening, and (2) full-text review. Two reviewers screened all records against the predefined inclusion and exclusion criteria respectively. Discrepancies were resolved through discussion …" These two descriptions cannot both be true: the first describes sequential single-reviewer screening with verification, while the second describes independent dual screening. It appears the second paragraph was added in response to the previous review without removing the original contradictory statement. The authors should delete the incompatible passage and retain only the description that accurately reflects what was done. They should also explicitly discuss the potential implications for selection bias and reproducibility, as recommended in the previous round, rather than simply asserting that consistency was ensured.
The manuscript has been revised to clarify that the screening process was conducted sequentially by two researchers. The first researcher conducted the initial screening, and all screening decisions were subsequently verified by a second researcher. The wording has been standardized throughout the manuscript to ensure consistency.
- The response to the quality appraisal concern is weaker than the response letter implies. The authors state that they have "updated Table 2 to explicitly include key indicators of rigor for each study: aim, sample size, study type, target group, etc." However, these columns appear to be largely the same descriptive metadata that were already present in the previous version, and they do not function as rigor indicators in the methodological sense intended by the previous comment. In particular, the previous review specifically suggested adding columns such as "presence of a comparison group" and "use of validated measures," neither of which has been included. I recommend that the authors either (a) genuinely supplement Table 2 with explicit methodological rigor indicators — for example, presence/absence of a comparison or control condition, type of outcome measure (validated instrument vs. self-developed survey vs. usability rating), and whether outcomes were objective or self-reported — or (b) apply a tool such as the MMAT, which is designed to accommodate methodological diversity across qualitative, quantitative, and mixed-methods studies and would be applicable to most of the included studies. If technical-development studies are deemed unsuitable for MMAT, those studies could be appraised with a brief structured commentary instead, with the rationale clearly stated.
This is what we can try best to revise so far, even if the reviewer thinks it is weak. We apologize for our limited capability to meet almost all 5 reviews’ tremendous comments and requirements.
- The justification for the 2019 starting point of the search window is weaker in the manuscript than in the response letter. The response letter argues persuasively that GPT-2 (released in 2019) marked the functional shift from rule-based chatbots to true generative models, and cites Radford et al. (2019) in support. However, the manuscript text only mentions GPT-3 (2020) and GPT-4 (2023), neither of which justifies a 2019 starting point. I recommend that the authors integrate the GPT-2 rationale (and the Radford et al. reference) directly into the Methodology section so that the time-frame justification is transparent to readers without access to the response letter.
We have revised the Methodology section to provide a clearer justification for the selected time frame. Specifically, we added an explanation that 2019 was chosen as the starting point because the release of GPT-2 marked a significant milestone in the development of transformer-based generative language models and conversational AI systems (Radford et al., 2019).
Reviewer 3 Report
Comments and Suggestions for AuthorsOverall, the revisions made are in line with the suggestions provided. However, there are still inconsistencies in the use of fonts throughout the manuscript. Please make further corrections to ensure compliance with the journal’s standards.
Comments on the Quality of English LanguageThe English writing throughout the manuscript requires substantial revision. There are numerous grammatical errors, awkward phrasing, and sentences that are difficult to parse. The Minor Comments section provides several examples below. Broadly speaking, it appears that different contributors wrote some passages without sufficient editorial harmonization. The Discussion section, in particular, contains run-on sentences and shifts in register that disrupt the reading experience. The authors are strongly encouraged to engage the services of a professional academic English editor before resubmission.
Author Response
1. Overall, the revisions made are in line with the suggestions provided. However, there are still inconsistencies in the use of fonts throughout the manuscript. Please make further corrections to ensure compliance with the journal’s standards.
Thank you for this comment. We carefully reviewed the manuscript and standardized fonts, font sizes, and heading styles throughout the document to ensure consistent formatting.
2. Comments on the Quality of English Language
The English writing throughout the manuscript requires substantial revision. There are numerous grammatical errors, awkward phrasing, and sentences that are difficult to parse. The Minor Comments section provides several examples below. Broadly speaking, it appears that different contributors wrote some passages without sufficient editorial harmonization. The Discussion section, in particular, contains run-on sentences and shifts in register that disrupt the reading experience. The authors are strongly encouraged to engage the services of a professional academic English editor before resubmission.
Thank you for this comment. The manuscript has undergone further proofreading and language polishing by a native English speaker. Additional grammatical, stylistic, and typographical corrections have been made throughout the revised manuscript.

