1. Introduction
Diagnostic communication is a core clinical competency in medical practice. The manner in which clinicians disclose a diagnosis shapes patient understanding, emotional processing, trust, adherence, and engagement with subsequent care. In chronic cardiometabolic diseases such as type 2 diabetes mellitus (T2DM) and obesity, these processes influence whether patients adopt and sustain the therapeutic behaviors required for long-term disease control. In oncology, where diagnostic disclosure often carries greater emotional weight, communication quality also affects distress, prognostic understanding, and shared decision-making [
1,
2,
3,
4].
These considerations make diagnostic communication highly relevant to both educational quality and patient safety. Communication failures have been repeatedly identified as important contributors to adverse events, handoff problems, and malpractice claims, supporting the need for structured training rather than reliance on informal experiential learning alone [
5,
6]. As competency-based medical education increasingly emphasizes observable and benchmarkable performance, diagnostic communication requires training models that are standardized, reproducible, and scalable.
However, scalable training in diagnostic disclosure remains difficult to implement. Standardized patients and other face-to-face simulation modalities offer substantial educational value, but they are resource-intensive and difficult to repeat at high frequency across large student cohorts. Virtual patient strategies have therefore gained attention as a means of expanding deliberate practice while preserving standardization and assessment opportunities [
7]. Generative artificial intelligence (AI) may complement these approaches by enabling dynamic, open-ended dialogue and individualized formative feedback, although its added educational value over established simulation modalities requires empirical evaluation.
Recent advances in large language models (LLMs) have further increased the feasibility of interactive virtual patients capable of generating dynamic, context-sensitive dialogue. Recent systematic reviews document the rapid expansion of ChatGPT (GPT-5) and related LLM applications in healthcare education and medicine, with recurring interest in tutoring, simulation, feedback, and communication training [
8,
9,
10]. At the same time, these reviews also highlight important limitations, including hallucinations, inconsistency, bias, and unresolved governance concerns [
8,
9,
10]. Accordingly, the educational value of AI-assisted simulation should be established through rigorous outcome-based evaluation rather than inferred from novelty alone.
A key educational question is whether post-training communication score changes are consistent across clinically distinct diagnostic disclosure contexts. Many studies evaluate a single scenario and may implicitly treat performance in one case type as a proxy for broader communicative competence. Yet diagnostic disclosure is not a uniform task. Communicating T2DM, counseling a patient with obesity, and disclosing a breast cancer diagnosis differ in explanatory demands, emotional intensity, stigma burden, uncertainty, perceived threat, and the balance required between biomedical information delivery and empathic support [
11,
12]. These differences provide a rationale for evaluating multiple scenarios rather than assuming that performance changes in one diagnostic context necessarily generalize to another.
A second challenge is the marked heterogeneity of learner response. Mean pre–post score changes may obscure important variability in how students perform across domains and scenarios, which is directly relevant to competency-based assessment, targeted remediation, and curricular design. In our initial DIALOGUE study, a single-group evaluation suggested that generative AI-assisted simulation was associated with higher post-training diagnostic communication performance in a T2DM disclosure scenario [
13]. A subsequent randomized controlled trial comparing conventional diagnostic communication training with conventional training supplemented by AI-supported adaptive simulation also supported the feasibility of this approach, although baseline-adjusted group-by-time interactions did not reach conventional thresholds for statistical significance and the added mean benefit beyond conventional training remained uncertain [
14]. Together, these findings support further evaluation of AI-assisted simulation while underscoring that its educational contribution should be interpreted cautiously and examined across different clinical communication contexts rather than inferred from a single diagnostic scenario.
The present study, therefore, extended this line of work from a single T2DM disclosure scenario to three clinically distinct diagnostic contexts: T2DM, obesity, and breast cancer. We aimed to examine whether completion of a structured generative AI-assisted simulation curriculum was associated with higher post-intervention diagnostic communication scores across scenarios, whether the magnitude of pre–post score change differed by clinical context, and whether exploratory gain-score analyses suggested cross-scenario associations or inter-individual heterogeneity in response patterns.
2. Materials and Methods
2.1. Study Design, Educational Context, and Ethical Approval
This was a prospective, single-arm, pre–post educational intervention study designed to evaluate the association between generative artificial intelligence (AI)-assisted simulation training and diagnostic communication performance across three clinically distinct diagnostic disclosure scenarios: type 2 diabetes mellitus (T2DM), obesity, and breast cancer. Because all participants received the intervention and no no-intervention, wait-list, active comparator, or non-AI simulation control group was included, the study was not designed to establish causal effectiveness. Observed pre–post changes were therefore interpreted as within-student associations after curriculum completion rather than as causal effects of the AI-assisted intervention. Potential alternative explanations included repeated exposure to the assessment format, maturation, concurrent clinical learning, rubric familiarization, testing effects, and regression to the mean. The study was conducted between September and October 2025 at the Facultad de Estudios Superiores Iztacala (FES Iztacala), Universidad Nacional Autónoma de México (UNAM), Mexico. The study protocol was finalized before participant enrollment, and no changes to the intervention sequence, assessment framework, or analytic strategy were introduced after study initiation.
Because the study comprised three scenario-based evaluations corresponding to distinct clinical communication contexts, separate ethics approvals and informed consent procedures were obtained before implementation. The study was approved by the Institutional Ethics Committee of FES Iztacala, UNAM, under protocol codes CE/FESI/052025/1922 (May 2025), CE/FESI/082025/1988 (August 2025), and CE/FESI/092025/2006 (September 2025). Written informed consent was obtained from all participants before enrollment and before participation in each of the three scenario-based evaluations.
The study was embedded within an educational activity of a comprehensive clinical course conducted during students’ clinical rotations. However, participation in the research component was voluntary. Students were explicitly informed that participation, non-participation, or withdrawal would have no effect on their academic standing, course performance, or relationship with faculty. No financial incentives were provided. The study was conducted in accordance with the Declaration of Helsinki.
2.2. Participants and Recruitment
Participants were undergraduate medical students enrolled in the clinical-cycle curriculum at FES Iztacala, UNAM. Students in the clinical cycles were invited to participate during a course-related educational activity conducted within their clinical rotations. Participants were recruited through voluntary convenience sampling from the eligible clinical-cycle student population.
The number of eligible students invited, enrolled students, scenario-specific completers, and three-scenario completers is reported in the participant flow diagram. Because recruitment was voluntary and convenience-based, selection bias was considered possible, particularly in relation to motivation, digital self-efficacy, and prior LLM use.
Eligibility criteria were active enrollment in the clinical-cycle curriculum and availability to complete both the in-person pre-test and post-test assessments. Exclusion criteria were failure to complete baseline assessment procedures or failure to provide written informed consent. All assessments were conducted individually under standardized on-campus conditions. When available, baseline characteristics of three-scenario completers and non-completers were compared descriptively to assess potential attrition-related bias.
2.3. Pre-Test Assessment
Baseline diagnostic communication performance was evaluated through an in-person pre-test conducted on campus under standardized conditions. Each assessment consisted of a live diagnostic disclosure encounter in one of three predefined clinical scenarios: T2DM, obesity, or breast cancer.
Across scenarios, students were required to perform the core components of a diagnostic communication encounter, including establishing rapport and agenda, eliciting the patient’s perspective, delivering diagnostic information clearly, responding empathically to emotional reactions, and closing the encounter with an appropriate management plan and follow-up orientation. Pre-test encounters were conducted with previously trained standardized human simulated patients. Simulated patients underwent prior preparation to ensure consistent portrayal of the clinical scenario, communicative demands, and expected emotional responses. All encounters were scored live. Because encounters were scored live rather than from anonymized video recordings, complete blinding to the assessment phase could not be guaranteed.
2.4. AI-Assisted Simulation Training
After completion of the pre-test, all enrolled participants completed an asynchronous simulation-based training program focused on diagnostic communication. The intervention was delivered using GPT-5 through the ChatGPT interface (OpenAI, San Francisco, CA, USA). GPT-5 had been officially released by OpenAI on 7 August 2025, before the September–October 2025 study implementation period. No custom fine-tuning was performed. The model was configured through standardized scenario-specific prompt templates developed a priori by the research team. Because the intervention was implemented through the ChatGPT interface rather than a fully controlled API environment, model-level parameters such as temperature, top-p, sampling seed, and backend routing were not controlled by the investigators; this limitation is acknowledged in relation to reproducibility.
Three scenario-specific standardized prompt templates were developed a priori by the research team, one for each clinical context: T2DM, obesity, and breast cancer. These prompts were designed to preserve scenario fidelity, standardize the communicative task, and maintain comparable virtual patient behavior across participants within each scenario. Each prompt instructed the model to act as a standardized virtual patient, maintain the assigned clinical and emotional profile, avoid providing expert-level medical guidance to the learner during the encounter, and generate structured formative feedback only after completion of the simulated diagnostic disclosure. Feedback was organized around the adapted Kalamazoo domains and included strengths, missed opportunities, and suggestions for subsequent practice.
Each participant completed 10 AI-assisted training encounters per scenario, for a total of 30 simulation sessions across the intervention. During each session, participants were instructed to communicate the diagnosis in a realistic and professionally appropriate manner, address patient concerns and emotional cues, and iteratively incorporate the feedback generated by the model into subsequent encounters. Formative feedback was generated automatically by ChatGPT after each simulated interaction. Completion of the 30 encounters was verified using exported chat records. Students were allowed to repeat sessions for additional practice, but only the required 10 encounters per scenario were counted for completion. Students were instructed not to use external resources during asynchronous practice; this was considered when interpreting the intervention’s ecological validity.
To address privacy and data protection, students were instructed not to enter personal health information, identifiable patient data, or confidential third-party information into the platform. All simulated cases were fictional and contained no real patient identifiers. Participants were informed that AI outputs could be incomplete, inconsistent, or inaccurate and that model-generated feedback was formative rather than authoritative clinical guidance. Prompt templates and representative outputs were reviewed by faculty before implementation to assess scenario fidelity, clinical appropriateness, empathic tone, and potential hallucination or bias. Individual student-level AI outputs were not exhaustively audited in real time, which was recorded as a limitation.
To reduce potential sequence effects and rubric familiarization bias, the order of scenario exposure was randomized across participants using a computer-generated simple random sequence. Thus, participants initiated training with different scenarios and subsequently completed the remaining ones. Although the intervention was asynchronous, all participants completed the same three-scenario training structure during the study period. Prompt templates, simulation materials, scoring materials, and training resources are publicly available in the Open Science Framework repository listed in the Data Availability Statement.
2.5. Post-Test Assessment
Following completion of the AI-assisted training phase, participants underwent an in-person post-test conducted on campus under standardized conditions. As in the baseline phase, each post-test consisted of a live diagnostic disclosure encounter in one of the three predefined clinical scenarios: T2DM, obesity, or breast cancer.
The communicative tasks required during post-test encounters were equivalent to those assessed at baseline, including rapport building, elicitation of the patient’s perspective, clear explanation of the diagnosis, empathic response to emotional reactions, and closure with an appropriate plan and follow-up orientation. Post-test encounters were conducted with previously trained standardized human simulated patients using preparation procedures analogous to those applied during the pre-test phase. All post-test encounters were scored live. Standardized patients were trained to portray the same clinical and emotional profile across participants within each scenario. Standardized patients were blinded to assessment timing and were not involved in scoring.
2.6. Outcome Measures and Scoring Instrument
The primary outcome was within-student change in diagnostic communication performance from pre-test to post-test, defined as Δ = post-test − pre-test in the total rubric score for each scenario. Secondary outcomes included pre–post changes in each of the eight communication domains, exploratory associations between scenario-specific gain scores, and inter-individual heterogeneity in learning trajectories assessed through domain-level change profiles and exploratory clustering procedures.
Performance was assessed using an adapted version of the Kalamazoo Essential Elements Communication Checklist, originally described as a seven-domain framework for communication assessment in medical encounters [
15]. In the present study, the rubric was adapted to the context of diagnostic disclosure and operationalized as 24 items distributed across eight domains, with three items per domain. The first seven domains were based on the original Kalamazoo framework, whereas an eighth domain—Scenario-Specific Communication—was developed by the research team to capture communicative competencies specific to the diagnostic demands of each clinical scenario. Each item was rated on a 5-point Likert scale. Domain scores were calculated as the sum of the three constituent items (range, 3–15), and the total score was calculated as the sum of all 24 items (range, 24–120), with higher scores indicating better communication performance. The eight domains were: (1) Build a Relationship, (2) Opening the Discussion, (3) Gather Information, (4) Understand the Patient’s Perspective, (5) Share Information, (6) Reach Agreement, (7) Provide Closure, and (8) Scenario-Specific Communication. The full adapted 24-item rubric, item-level scoring anchors, scenario-specific communication items, and scoring materials are publicly available in the OSF repository listed in the Data Availability Statement. The adaptation process included expert review by faculty with experience in medical education, simulation, endocrinology/metabolic disease, oncology/gynecology, and psychology. The adapted instrument was reviewed for content relevance, clarity, scenario fit, and rater usability before implementation.
2.7. Blinding, Rater Calibration, and Scoring Procedures
All live encounters were independently evaluated by two physician raters with prior teaching experience in medical education. Raters did not participate in the delivery of the intervention and were kept unaware of participants’ exposure to the AI-assisted training program. Specifically, they were not informed that students had completed a generative AI-supported simulation curriculum before the post-test assessments. However, because encounters were scored live and not from anonymized recordings presented in random order, complete blinding to the pre-test versus post-test phase could not be guaranteed. This was considered a potential source of expectation bias.
The same rubric, scoring anchors, and calibration procedures were used across scenarios and time points. The same physician raters did not score both pre-test and post-test encounters. Raters could not identify participants by name or previous encounter. Scenario order assignment was not concealed from raters. Before study initiation, evaluators underwent rubric training and calibration to standardize scoring criteria across scenarios and time points. Inter-rater reliability was assessed using weighted Cohen’s kappa at the domain level. Inter-rater agreement was >0.80, supporting acceptable scoring consistency for rubric-based live assessment. Disagreements between raters were resolved by third-rater review, and final analytic scores were computed using adjudicated scores. For continuous domain and total scores, inter-rater reliability was additionally evaluated using intraclass correlation coefficients (ICCs) with 95% confidence intervals.
2.8. Statistical Analysis
Analyses were performed at two complementary levels. First, scenario-level analyses included all participants with complete paired pre-test and post-test data for the corresponding scenario. Second, cross-scenario analyses were restricted to the complete-case subgroup of students with paired pre-test and post-test data across all three scenarios.
Continuous variables are presented as mean ± standard deviation (SD) or median and interquartile range (IQR), as appropriate. Categorical variables are presented as frequencies and percentages. For each scenario, pre-test and post-test total scores and domain scores were compared using paired t-tests. Within-student change was summarized as Δ = post-test − pre-test. Paired effect sizes were estimated using Cohen’s dz, calculated as the mean of the paired differences divided by the standard deviation of those paired differences. Ninety-five percent confidence intervals (95% CIs) were reported where applicable. To control the family-wise error rate in domain-level analyses, p-values were adjusted using the Holm procedure.
As a sensitivity analysis, a linear mixed-effects model was fitted to evaluate total rubric scores across time and scenario while accounting for repeated observations within students. The model included fixed effects for time, scenario, and the time-by-scenario interaction, with a random intercept for student. The time-by-scenario interaction was used to examine whether the magnitude of pre–post score change differed across T2DM, obesity, and breast cancer scenarios. Estimated marginal means and pairwise contrasts of scenario-specific pre–post changes were reported with multiplicity adjustment. If model assumptions were not met, results were interpreted cautiously and compared with the paired t-test findings.
For the analysis of average within-student gains across the subgroup of participants who completed all three scenarios, one-sample t-tests were used to test whether mean Δ differed from 0 for each domain and for the total score, with Holm correction applied across the nine outcomes. To explore baseline–gain relationships, associations between baseline pre-test scores and Δ values were examined using regression-based visualization. Because gain scores are sensitive to baseline scores, regression to the mean, ceiling effects, and measurement error, these analyses were interpreted cautiously. Where feasible, residualized change scores from baseline-adjusted models were also examined as an exploratory sensitivity analysis.
Because gain-score distributions were not assumed to be strictly Gaussian and cross-scenario associations were treated as exploratory rank-based constructs, associations between scenario-specific Δ values were evaluated using Spearman correlation coefficients. Multiple cross-scenario comparisons were interpreted cautiously because of the small number of pairwise tests and the exploratory nature of the analysis.
To characterize heterogeneity in learning trajectories, exploratory clustering analyses were performed on domain-level and scenario-level change profiles. For heatmap visualization, domain-level Δ scores were standardized within each scenario using z-scores to emphasize relative change patterns across domains. Students were grouped using k-means clustering with k = 3, selected as a pragmatic and interpretable solution for identifying exploratory response patterns across scenarios. Because clusters were derived from the same gain variables used to describe them, between-cluster differences were presented descriptively and were not interpreted as independent inferential evidence. The k = 3 solution was selected a priori as a pragmatic descriptive grouping to aid visualization and was not intended as a statistically optimized or validated clustering solution.
Baseline characteristics of three-scenario completers and non-completers were compared descriptively using t-tests, Mann–Whitney U tests, chi-square tests, or Fisher’s exact tests as appropriate. These analyses were exploratory and intended to assess potential attrition bias rather than to support causal inference. All tests were two-sided, and statistical significance was defined as p < 0.05. All analyses were performed in R using RStudio (version 2025.05.1+513).
4. Discussion
In this prospective single-arm pre–post educational study, completion of a structured generative AI-assisted simulation curriculum was associated with higher post-intervention diagnostic communication scores across three clinically distinct disclosure scenarios: T2DM, obesity, and breast cancer. Higher post-test scores were observed not only in the total rubric score, but also across all eight communication domains. However, the magnitude of pre–post score change was not uniform across scenarios or domains. The largest descriptive total-score change was observed in the breast cancer scenario, domain-level effects varied across clinical contexts, exploratory cross-scenario gain-score associations were limited and context-dependent, and individual pre–post change patterns were heterogeneous. Taken together, these findings extend the current literature by suggesting that AI-assisted communication simulation can be evaluated across multiple diagnostic disclosure contexts using standardized performance-based assessment, while also showing that score changes should not be assumed to generalize uniformly across diagnostic scenarios.
The most immediate implication of these findings is that AI-assisted simulation may be useful as a scalable adjunct for repeated diagnostic communication practice across more than one clinical context. Prior work in medical education has shown that communication competence can be strengthened through simulation-based training, but repeated exposure with individualized feedback is often constrained by the cost and logistics of standardized-patient programs [
5,
7,
16]. Large language model-based systems offer an attractive complement because they permit high-frequency, low-marginal-cost rehearsal with immediate formative feedback [
8,
9,
10,
17,
18]. The present findings are consistent with this rationale, as students demonstrated positive pre–post score changes across all three scenarios using a standardized rubric-based assessment. However, because the study lacked a control group, these changes cannot be attributed specifically to the AI-assisted component and may also reflect repeated exposure, rubric familiarization, concurrent clinical learning, or maturation. This distinction is important because much of the current literature on generative AI in health professions education remains descriptive, feasibility-focused, or based on self-reported outcomes rather than observed performance [
8,
9,
10,
19,
20].
At the same time, the present results argue against treating diagnostic communication as a generic, scenario-independent competency. The pattern of pre–post score change depended on the clinical context. The breast cancer scenario showed the largest descriptive changes in total score and several individual domains, whereas the cardiometabolic scenarios showed smaller changes and different internal profiles. This scenario dependence is educationally plausible. Communicating a breast cancer diagnosis differs from counseling a patient with obesity or disclosing T2DM not only in emotional valence, but also in explanatory structure, stigma burden, decisional urgency, perceived threat, and the balance between empathic containment and structured information delivery [
3,
4,
11,
12]. These are not interchangeable communicative tasks. The present findings support the use of multi-scenario assessment when evaluating diagnostic communication training, particularly when the goal is to benchmark performance across clinically and emotionally distinct encounters.
Domain-level findings also have curricular implications. Larger score changes were concentrated in domains related to encounter organization and information exchange, including opening the discussion, gathering information, sharing information, and providing closure. In contrast, changes were smaller in the domain of understanding the patient’s perspective. This pattern suggests that some communication behaviors may be more readily shaped through repeated scripted practice and immediate feedback, whereas deeper patient-centered skills may require longitudinal coaching, reflective debriefing, supervised clinical exposure, and explicit attention to emotional and contextual complexity. AI-assisted simulation may therefore be most appropriately conceptualized as one component of a broader communication curriculum rather than as a replacement for faculty-led feedback or standardized-patient debriefing [
6,
7,
20,
21].
The exploratory cross-scenario analyses further reinforce the need for caution when interpreting gain scores as evidence of transferable competence. Gain-score associations were limited and context-dependent: changes in one scenario did not consistently correspond to changes in another. However, these correlations should not be interpreted as direct evidence of transfer or lack of transfer. Gain scores are sensitive to baseline performance, regression to the mean, ceiling effects, and measurement error. Therefore, the more defensible interpretation is that diagnostic communication performance should be assessed across multiple contexts rather than inferred from a single case. Future studies should use controlled longitudinal designs, residualized change models, and follow-up standardized-patient assessments to evaluate whether skills practiced in one diagnostic context are retained and applied to other contexts [
12,
13,
17,
22].
The heterogeneity analyses are particularly important in this regard. Heatmaps and responder clustering showed that improvement was not distributed along a simple linear continuum. Some students exhibited consistently strong gains across scenarios, others showed intermediate but broad improvement, and others demonstrated attenuated or context-specific gains. This pattern should not be dismissed as statistical noise. In educational terms, heterogeneity is itself information. It suggests that AI-supported communication training may function not only as an intervention, but also as a measurement-rich environment that can expose learner-specific profiles not easily visible in aggregate means. That possibility aligns with contemporary views of competency-based education, in which variability is not a nuisance to be averaged away, but a signal that should guide targeted remediation, progression decisions, and curricular design [
6,
7,
19,
21].
The use of generative AI in communication training also raises methodological and governance considerations. Although standardized prompts, fictional cases, faculty review of representative outputs, and privacy instructions were implemented, large language models remain probabilistic systems that may generate inconsistent, incomplete, biased, or clinically inaccurate responses. In this study, model-generated feedback was used only for formative practice and not as an authoritative clinical source or summative assessment tool. Future implementations should incorporate prospective monitoring of AI outputs, predefined procedures for identifying hallucinations or biased responses, transparent documentation of model version and access route, and clear data-protection safeguards. These requirements are particularly important when AI systems are used in sensitive communication scenarios involving stigma, cancer diagnoses, emotional distress, or potentially vulnerable learners.
This study has limitations. First, the single-arm pre–post design precludes causal attribution of the observed score changes to the AI-assisted simulation curriculum. Without a no-intervention, wait-list, active comparator, or non-AI simulation control group, alternative explanations cannot be excluded, including testing effects, repeated exposure to the assessment format, rubric familiarization, maturation, concurrent clinical learning, and regression to the mean. Second, assessments were scored live rather than from anonymized recordings presented in random order; therefore, complete blinding to the pre-test versus post-test phase could not be guaranteed, creating potential expectation bias. Third, participants were recruited from a single institution within one cultural and linguistic context, which may limit external generalizability. Fourth, although the use of three scenarios strengthens ecological and educational relevance, it also introduces complexity: diagnostic disclosure tasks are not emotionally or cognitively equivalent, and some observed variation may reflect scenario characteristics rather than differences in underlying communication competence. Fifth, despite standardized prompts and publicly available training materials, GPT-5 was accessed through the ChatGPT interface rather than a fully controlled API environment; therefore, model-level parameters and backend routing were not controlled by the investigators, limiting exact reproducibility. Sixth, individual student-level AI outputs were not exhaustively audited in real time, so undetected variability, hallucination, bias, or inconsistency in model responses cannot be excluded. Seventh, although the adapted rubric was reviewed by experts and inter-rater agreement was acceptable, further psychometric evaluation of the instrument across institutions and learner levels is needed. Finally, the study assessed short-term post-intervention performance in simulated encounters and did not evaluate long-term retention, transfer to real clinical encounters, patient outcomes, or effects on faculty workload and curricular implementation. Although several learner-level characteristics may plausibly influence pre–post score change, including baseline communication performance, prior LLM use, digital self-efficacy, empathy, communication-related self-confidence, academic performance, and prior clinical exposure, the present study was not powered for reliable multivariable predictor modeling. We therefore limited inferential interpretation to baseline–gain relationships and reported other learner characteristics descriptively. Future controlled studies with larger samples should examine predictors and moderators of response to AI-assisted simulation to determine which learners benefit most and which require additional faculty-led support.
Future research should move in three directions. The first is causal clarification: multi-scenario randomized trials are needed to determine the added value of AI-assisted simulation over conventional communication training, non-AI virtual patients, standardized-patient practice, and faculty-led feedback. The second is mechanism: studies should examine whether observed score changes are driven primarily by repetition, immediate feedback, scenario variability, learner motivation, baseline performance, or interactions among these factors. The third is implementation: future work should evaluate how AI-assisted simulation performs when embedded longitudinally within curricula, how it interacts with faculty-led debriefing, whether performance changes persist over time, and whether skills demonstrated in simulated encounters translate to later standardized-patient assessments or real clinical communication. The present study supports the rationale for such work by showing that AI-assisted simulation can be evaluated using multi-scenario, rubric-based performance assessment, while also underscoring that diagnostic communication remains context-dependent and educationally heterogeneous.