1. Introduction
Modern clinical practice is increasingly burdened by documentation, and discharge letters remain among the most time-consuming tasks for hospital physicians. Despite their administrative burden, high-quality discharge letters are essential for continuity of care, communication between inpatient and outpatient providers, and patient safety. In routine practice, however, discharge documents are often inconsistent in structure, variable in completeness, and uneven in clarity, which may reduce their usefulness for subsequent clinical decision-making. These persistent challenges have driven growing interest in tools that could improve both the efficiency and standardization of discharge documentation.
Large language models (LLMs) have emerged as promising candidates for this role because they can transform structured or semi-structured clinical information into coherent narrative text. Recent studies suggest that LLM-generated discharge letters can, in some settings, approach the quality of clinician-authored documents [
1,
2,
3]. For example, one blinded comparative study found no significant difference in overall quality between physician- and LLM-generated hospital discharge letters, although LLM outputs were somewhat more error-prone despite a low overall potential for harm [
1]. Other work has shown that LLM-generated letters may reduce physician editing burden while producing narratives that are readable and clinically usable, particularly when supported by structured prompting strategies and carefully selected source inputs [
4,
5]. In addition, recent methodological work suggests that summarization accuracy can be improved when models are guided by highlighted key content rather than relying on raw source text alone [
5].
At the same time, the literature remains mixed and indicates that performance is strongly context-dependent. Several studies have reported that AI-generated discharge letters, while fluent and concise, may omit case-critical details, contain factual inaccuracies, or introduce hallucinated content [
2,
3]. In psychiatry, lower ratings and a notable prevalence of hallucinations have been reported in AI-generated letters [
2]. Comparative work in oncology has likewise shown meaningful differences across models in their ability to capture clinically important information for lung cancer discharge letters [
3]. Specialty-specific studies further suggest that strong performance often depends on task-specific optimization, structured inputs, and expert validation. In cardiology, fine-tuned open-source LLMs showed encouraging performance in transforming progress notes into discharge letters, but only with careful clinical assessment and controlled implementation [
6]. Likewise, in intensive care settings, models have differed substantially in their ability to identify and extract key clinical events from complex narrative records [
7].
More recent evidence also cautions against assuming broad equivalence between LLM- and clinician-written discharge documents across all settings. A recent pediatric study found that physician-authored discharge letters outperformed unedited LLM-generated letters across evaluated quality domains, underscoring the continued importance of human authorship and review in safety-critical documentation [
8]. Broader evidence syntheses similarly conclude that LLMs may support clinical documentation and reduce administrative burden, but are not yet ready to replace clinicians without robust safeguards, local validation, and human oversight [
9].
These considerations are especially relevant in ophthalmology, where discharge letters must communicate precise laterality, diagnoses, operative details, treatment plans, and follow-up recommendations in a concise but clinically reliable format. Even small omissions or inaccuracies may have consequences for subsequent care. However, despite rapidly expanding literature on AI-assisted documentation, evidence remains limited for ophthalmology-specific discharge documentation and for direct head-to-head comparisons between LLM-generated and resident-authored discharge letters derived from the same clinical information.
This study addresses that gap by evaluating whether LLM-generated discharge letters can achieve quality comparable to, or better than, resident-authored discharge letters in domains most relevant to continuity and safety of care: factual accuracy, completeness, clarity and structure, conciseness, and appropriate professional tone. Building on prior findings, we used standardized prompts and structured anonymized clinical data to reduce input variability, together with blinded expert assessment to minimize detection bias [
1,
4,
5,
6,
7]. By clarifying where LLMs may add value and where human authorship remains essential, the findings aim to inform the safe and pragmatic integration of AI-assisted drafting into ophthalmic clinical workflows.
Several considerations specifically motivate the use of LLMs for this task. First, they are particularly effective at producing grammatically polished, consistently structured, and professionally worded prose, which may improve the standardization of discharge documentation across authors and shifts. Second, unlike most narrow domain-specific natural-language-processing pipelines, current general-purpose LLMs require minimal task-specific training, making them an accessible candidate technology for resource-limited departments. Taken together, these properties make LLMs a pragmatic candidate for supporting, rather than replacing, clinician authorship of routine ophthalmology discharge documentation.
Despite this growing body of evidence, several specific gaps motivate the present study. First, prior comparisons have rarely paired LLM-generated and human-authored letters at the case level using identical source inputs and a blinded reviewer design. Second, structured assessment of clinically high-stakes content elements (laterality, operative details, findings, treatment, and follow-up) has been incompletely reported in specialty-specific work. The main contributions of this study are therefore: (i) a head-to-head paired comparison of resident-written and GPT-5.2-generated ophthalmology discharge letters derived from identical de-identified clinical input data; (ii) a blinded multi-reviewer assessment using a structured rating instrument covering accuracy, completeness, clarity, tone, conciseness, errors, omissions, and key content elements; (iii) a reviewer-adjusted sensitivity analysis to test the robustness of the primary findings; and (iv) practical, specialty-specific implications for the safe integration of LLM-assisted drafting into ophthalmic discharge workflows.
2. Materials and Methods
Study framework and overall workflow
The overall study framework is summarized in
Figure 1. For each eligible hospitalization, structured de-identified clinical data were retrieved from the hospital information system and used as the single source of truth for both drafting paths. Along the human-authored path, the discharge letter had already been written by an ophthalmology resident as part of routine clinical care. Along the model-authored path, the same de-identified inputs were submitted to GPT-5.2 together with a predefined, standardized prompt designed to constrain output to documented facts (numerical values, dates, laterality, and procedural details) and to enforce a uniform discharge-letter structure. The two letters were then paired at the case level, anonymized, reformatted into a uniform layout to strip source-specific cues, and submitted for blinded structured assessment by three board-certified ophthalmologists. The resulting paired ratings were analyzed using a paired statistical design (primary analysis), with a reviewer-adjusted model and inter-rater agreement on overlapping subsets serving as sensitivity analyses. This framework was designed to directly address the principal challenges of evaluating LLM-generated clinical documentation: case-mix confounding (addressed by the paired design), input variability (addressed by standardized de-identified source inputs), reviewer detection bias (addressed by uniform formatting and blinding), and reviewer-specific bias (addressed by the multi-reviewer sensitivity analysis).
Study design and setting
This retrospective observational study was conducted at the Clinic for Eye Diseases, University Hospital of Split, Croatia. The study protocol was approved by the Ethics Committee of University Hospital of Split.
Study material
The study cohort comprised 146 consecutive inpatient discharges beginning on 1 September 2025. For each hospitalization, the original discharge letter had been written by an ophthalmology resident as part of routine clinical care. Relevant clinical information for each case was retrospectively retrieved from the hospital information system.
Before further processing, all data were de-identified to remove patient identifiers and ensure confidentiality.
Patients ranged across the typical age distribution of an adult ophthalmology inpatient service, with a mixed-sex composition reflecting routine clinical practice. The case mix covered the principal categories of ophthalmic inpatient care at the center, including surgical admissions (e.g., cataract surgery, vitreoretinal procedures, anterior-segment and oculoplastic interventions) as well as conservatively managed conditions (e.g., retinal vascular disease, uveitis, infectious or inflammatory ocular disease, and ocular trauma evaluation). Each discharge letter was based on the structured information available in the hospital information system, including the principal diagnosis, ICD-10 code where applicable, hospital course, performed procedures, examination findings, ocular and systemic medications, discharge therapy, follow-up instructions, and discharge status.
Generation of GPT-5.2-based discharge letters
For each eligible case, a second discharge letter was generated using GPT-5.2. The model was provided with the same de-identified clinical information available from the hospital information system and prompted using a predefined, standardized instruction set developed for ophthalmology discharge letters. The prompt required the model to generate an official discharge letter in Croatian strictly on the basis of the available source material, without inventing missing information, while preserving documented numerical values, dates, laterality, and treatment details. The standardized GPT-5.2 prompt is provided in
Supplementary Material (File S1). The prompt was drafted by an ophthalmologist-led writing group (senior ophthalmology authors together with a clinical informatics co-author) and was iteratively refined over four rounds on a pilot set of cases that were not part of the study sample, until the model output reliably preserved laterality, numerical values, dates, and procedural details without introducing content that was absent from the source material. The final prompt fixed the output language to Croatian, imposed a fixed section structure (admission diagnosis, hospital course, findings, procedures, discharge therapy, follow-up instructions, and discharge status), and explicitly prohibited fabrication, speculation about missing data, and inclusion of patient-identifying information. The prompt was kept fixed throughout the study to avoid post hoc tuning on the evaluated cohort. Prompt optimization was therefore based on structured prompt engineering rather than on automated approaches such as reinforcement-style prompt search, fine-tuning, or retrieval-augmented generation. Domain ontologies (for example, SNOMED CT or ICD-10) and clinical knowledge graphs were not used to augment the prompt in this initial study, in order to keep the comparison focused on the realistic baseline of a general-purpose LLM operating on routine de-identified clinical inputs. Incorporating ophthalmology-specific ontologies, structured knowledge graphs, and retrieval-augmented context into the prompting pipeline is a planned next step and a promising avenue for further reducing residual content omissions, particularly for surgical and procedural details.
Thus, for each hospitalization, two discharge letters were available for comparison: the original resident-written version and the GPT-5.2-generated version. A resident-written and GPT-5.2-generated discharge-letter pair was available in 146 consecutive discharges.
Blinded assessment
All discharge letters were anonymized and evaluated in blinded fashion by three independent board-certified ophthalmologists who had not participated in drafting the letters. Reviewers were unaware whether a given discharge letter had been written by a resident or generated by GPT-5.2.
Blinding was implemented by assigning each evaluated document a unique anonymized file code that did not reveal its origin. In addition, before review, all discharge letters were reformatted into a uniform layout and stripped of source-specific formatting features in order to minimize the possibility that reviewers could infer document origin from visual presentation or stylistic formatting cues. The allocation key linking anonymized file codes to document source was maintained exclusively by the principal investigator and was not available to the reviewers during the assessment process.
A total of 146 paired resident/GPT-5.2 discharge-letter sets were available for primary reviewer assessment. One reviewer assessed all 146 available case pairs. The remaining two reviewers each assessed a subset of 45 case pairs. These additional reviewer subsets partially overlapped but were not identical, with one reviewer assessing case pairs 46–90 and the other case pairs 31–75. Consequently, not all case pairs were assessed by all three reviewers. Accordingly, the primary reviewer-assessed cohort consisted of the 146 cases in which both the resident-written and GPT-5.2-generated discharge letters were available.
Outcome measures
Each discharge letter was assessed using a structured study-specific rating form. The primary evaluation domains were factual accuracy, completeness, clarity/structure, tone/professional phrasing, and overall global quality. Accuracy was defined as concordance with the source clinical data. Completeness referred to whether the discharge letter contained the information necessary for safe continuity of care, including diagnosis, key intervention, hospital course, relevant findings, inpatient and discharge treatment, follow-up recommendations, and patient status at discharge. Clarity/structure assessed readability, logical organization, and practical usefulness for subsequent care. Tone/professional phrasing referred to whether the text was appropriately worded, professionally framed, and suitable for official medical documentation. Overall global quality reflected the overall suitability of the discharge letter for continuity and safety of care.
Secondary outcomes included conciseness, critical errors, minor errors, and critical omissions. Reviewers also recorded whether specific discharge-letter components were present, including the main diagnosis, operation, hospital course, findings, discharge therapy, follow-up instructions, and discharge status. Reviewer confidence was recorded on the original rating form, but because all evaluable entries received the maximum score, no inferential comparison was performed for this variable. The reviewer rating form is provided in
Supplementary Material (File S2).
Statistical analysis
The unit of analysis was the paired case, defined as one resident-written discharge letter and one GPT-5.2-generated discharge letter derived from the same hospitalization and based on the same underlying clinical information. All primary comparisons were therefore performed using a paired design.
The primary analysis was based on ratings from the reviewer who assessed all available paired cases, in order to preserve a consistent case-level comparison across the study sample. Among the three reviewers, this reviewer was selected as the primary grader because they consistently gave the lowest mean ratings across the rating domains, so that the primary paired comparison would be based on the most conservative grading available in the study and would not overstate the quality of either group. This means that the primary comparison is biased against showing high absolute scores for either resident-written or GPT-5.2-generated letters, and that any apparent advantage for GPT-5.2 on the primary analysis is conservative rather than optimistic. Because four GPT-5.2 evaluations were missing in the primary reviewer dataset, the main paired comparison for the principal quality outcomes included 142 complete discharge-letter pairs. Thus, the analytic flow for the primary comparison was as follows: 146 consecutive inpatient discharges were screened, 146 resident/GPT-5.2 discharge-letter pairs were available for primary reviewer assessment, and 142 pairs had complete primary reviewer ratings for both documents and were therefore included in the primary paired analysis. The remaining four pairs were excluded from the primary paired analysis because the primary reviewer assessment of the GPT-5.2-generated discharge letter was incomplete.
Variables rated on a 4-point Likert scale, including accuracy, completeness, clarity/structure, tone/professional phrasing, conciseness, and overall quality, were summarized as mean ± standard deviation (SD). Paired comparisons between resident-written and GPT-5.2-generated discharge letters were performed using the Wilcoxon signed-rank test, reflecting the ordinal nature of the scales and the paired structure of the data.
Discrete count outcomes, including the number of critical errors, minor errors, and critical omissions, were also compared using the Wilcoxon signed-rank test. For binary variables describing the presence or absence of specific discharge-letter components, analyses were restricted to applicable cases only, with non-applicable items excluded from the respective comparison. Paired comparisons for binary outcomes were performed using the McNemar test.
As supportive analyses, effect sizes were calculated for the primary paired comparisons as mean paired differences with 95% confidence intervals and standardized paired effect sizes (Cohen’s dz). In addition, a reviewer-adjusted sensitivity analysis including all available ratings from all three reviewers was performed, with document source as the predictor of interest, reviewer included as an adjustment factor, and standard errors clustered at the case level to account for within-case dependence. For the principal quality outcomes, this supportive analysis included 231 complete reviewer-case pairs across all reviewers. Inter-rater agreement on overlapping subsets was assessed using quadratic weighted Cohen’s kappa for ordinal outcomes.
All tests were two-sided, and p < 0.05 was considered statistically significant.
The choice to base the primary paired comparison on the single reviewer who assessed all 142 complete pairs reflected a deliberate trade-off: it preserved a fully paired case-level design and the statistical power associated with the full cohort, while still allowing reviewer-related variability to be examined through two pre-specified sensitivity analyses. The first sensitivity analysis used a reviewer-adjusted model on all available ratings from all three reviewers, with reviewer included as an adjustment factor and case-level clustering of standard errors. The second sensitivity analysis was restricted to the subset of case pairs that were independently rated by all three reviewers (the overlap between the two secondary reviewer assignments, corresponding to case pairs 46–75), and used the same Wilcoxon signed-rank tests, with reviewer ratings averaged within each case before paired comparison. Inter-rater agreement was deliberately calculated only on the overlapping subsets of cases that were independently rated by the relevant reviewer pair, since agreement statistics are not defined on cases evaluated by only one reviewer. Together, these analyses were intended to provide a transparent view of how strongly the primary findings depended on the assignment of cases to reviewers.
Study objective
The primary objective of the study was to compare the quality of discharge letters written by ophthalmology residents with those generated by GPT-5.2 using the same de-identified clinical input data.
3. Results
Of the 146 consecutive inpatient discharges, 146 resident/GPT-5.2 discharge-letter pairs were available for primary reviewer assessment, and 142 complete paired evaluations were available for the primary analysis. Thus, four otherwise available pairs were excluded from the primary analytic sample because the primary reviewer assessment of the GPT-5.2-generated discharge letter was incomplete. Overall, resident-written and GPT-5.2-generated discharge letters performed similarly across the main quality domains (
Table 1). As illustrated in
Figure 2, most estimated between-group differences were small and centered close to zero. No significant differences were observed for accuracy, completeness, clarity/structure, or overall quality. GPT-5.2-generated discharge letters received higher ratings for tone/professional phrasing, whereas resident-written discharge letters were rated as more concise (
Table 1,
Figure 2).
No significant between-group differences were found for critical errors, minor errors, or critical omissions (
Table 1,
Figure 2). Among the predefined content elements, the largest differences were seen for operation and findings: resident-written discharge letters more often documented the operation, whereas GPT-5.2-generated discharge letters more consistently included findings (
Table 2,
Figure 3). The presence of the main diagnosis, hospital course, discharge therapy, follow-up instructions, and discharge status did not differ significantly between groups (
Table 2,
Figure 3).
The distribution of reviewer scores across the main quality domains, including conciseness, is shown in
Figure 4. Both groups received predominantly high ratings overall, particularly for clarity/structure and completeness. However, GPT-5.2-generated discharge letters showed a relative shift toward higher scores for tone/professional phrasing, whereas resident-written discharge letters more often received the highest score for conciseness (
Figure 4), consistent with the paired comparisons shown in
Table 1.
Effect-size estimates for the primary paired analysis were generally small. The largest effect favored GPT-5.2 for tone/professional phrasing (mean difference 0.18, 95% CI 0.10 to 0.26; Cohen’s dz = 0.37), while resident-written discharge letters showed a small advantage in conciseness (mean difference −0.08, 95% CI −0.13 to −0.02; Cohen’s dz = −0.23). For the remaining outcomes, effect sizes were small and confidence intervals were centered close to zero (
Figure 1).
In the supportive reviewer-adjusted sensitivity analysis including all available ratings, GPT-5.2 remained associated with higher tone/professional phrasing scores, whereas resident-written discharge letters showed more favorable ratings for accuracy, clarity/structure, conciseness, and global quality (
Table 3). No significant differences were observed for completeness, critical errors, minor errors, or critical omissions in this analysis.
Inter-rater agreement on overlapping subsets was highest for completeness and overall quality, with weighted kappa values in the fair-to-moderate range, and lower for clarity/structure, tone/professional phrasing, and conciseness (
Table 4). This pattern suggests greater reviewer dependence for stylistic than for content-related assessments.
Taken together, the primary paired analysis indicates that GPT-5.2-generated discharge letters achieved overall performance comparable to resident-written discharge letters, with differences mainly related to tone/professional phrasing and conciseness. However, the reviewer-adjusted sensitivity analysis suggests that these findings should be interpreted with caution.
4. Discussion
In this retrospective paired comparison of ophthalmology discharge letters generated either by residents or by GPT-5.2 from the same de-identified clinical information, the primary analysis suggested similar performance across the main quality domains. No statistically significant differences were observed for accuracy, completeness, clarity/structure, critical errors, minor errors, critical omissions, or overall global quality. The clearest differences were stylistic: GPT-5.2-generated letters received higher ratings for tone/professional phrasing, whereas resident-written letters were rated as more concise. Overall, these findings suggest that, in this setting, LLM-generated discharge letters can approximate the quality of resident-authored letters, but with a somewhat different balance of strengths and weaknesses.
The finding of similar overall performance is consistent with a growing body of literature suggesting that LLMs can produce discharge letters of clinically acceptable quality when applied to bounded documentation tasks and supplied with sufficiently structured source material [
1,
2,
3,
4,
5,
10]. Our results are particularly aligned with studies showing that AI-generated discharge letters may approach clinician-written documents in overall quality while differing in style, brevity, or completeness rather than in gross usability [
1,
3,
4,
10]. In this sense, the present findings support the view that LLMs may already be useful as drafting tools in clinical documentation workflows, even if they are not yet suitable for fully autonomous use. This interpretation is further supported by recent studies showing that LLM-assisted medical record generation and clinical note summarization systems can produce documentation of usable quality under structured evaluation conditions [
11,
12].
At the same time, our results also reinforce the need for caution. Although the primary paired analysis suggested broad equivalence across most domains, the supportive reviewer-adjusted sensitivity analysis was less favorable to GPT-5.2, with resident-written letters receiving better ratings for accuracy, clarity/structure, conciseness, and global quality, while GPT-5.2 retained an advantage only in tone/professional phrasing. This discrepancy is important because it indicates that conclusions about comparative performance may depend on analytical framework and reviewer composition. Rather than demonstrating interchangeability between model- and clinician-generated documentation, our findings may more appropriately be interpreted as evidence that LLM performance can approach human performance under some conditions, but remains sensitive to how quality is assessed. This cautious interpretation is in line with broader literature emphasizing that apparent gains in fluency or readability should not be conflated with clinical reliability [
2,
6,
9]. It is also consistent with more recent real-world evidence showing that although LLM-generated letters may reduce editing burden and produce cohesive drafts, they may still contain confabulations or subtle inaccuracies that require clinician review before sign-off [
4,
10]. A similar pattern has been reported in recent evaluations of AI-generated clinical notes, in which LLM-authored documentation was often rated as comparatively thorough or well organized, yet less concise and more prone to hallucinated or clinically imprecise content than physician-authored notes [
13].
One clinically relevant finding was the uneven pattern in the inclusion of specific discharge-letter elements. Resident-written letters were substantially more likely to document the operation, whereas GPT-5.2-generated letters more consistently included findings. This suggests that the two approaches may prioritize content differently even when derived from the same source information. A plausible explanation is that resident authors, writing within a surgical specialty and routine workflow, naturally foreground procedural content, whereas the model may respond more strongly to descriptive clinical details present in the input text. In ophthalmology, however, both procedural details and examination findings may be critical for continuity of care. A letter that is well written but incompletely communicates the operative intervention may be less clinically useful than its overall quality score suggests. This observation supports the need for specialty-specific prompting and targeted review checklists focused on high-stakes content such as laterality, procedure details, treatment instructions, and follow-up recommendations.
The better performance of GPT-5.2 in tone/professional phrasing is unsurprising and likely reflects one of the more stable advantages of contemporary LLMs. These models are particularly effective at producing grammatically polished, standardized, and professionally worded prose, which may improve the surface readability and formal presentation of discharge documentation [
1,
4,
9,
14]. However, this strength should not be overinterpreted. In real clinical communication, stylistic polish is valuable only insofar as it supports accurate, efficient transfer of clinically relevant information. The corresponding finding that resident-written letters were more concise is therefore important. In busy clinical settings, concise letters may be easier for downstream clinicians to scan rapidly and use effectively. Thus, GPT-5.2 may have produced text that appeared more polished, while resident authors produced communication that was somewhat more information-dense and operationally efficient. Related recent work also suggests that LLMs may improve readability or simplify discharge communication, but that these advantages do not eliminate concerns regarding factual precision, contextual nuance, or clinical appropriateness [
14]. A recent ophthalmology study further supports this distinction, showing that patients preferred human empathy but did not uniformly prefer human wording when comparing GPT-generated and clinician-written discharge texts, suggesting that linguistic polish and perceived empathy may not fully overlap in patient-facing discharge communication [
15].
Our findings fit within an increasingly heterogeneous literature. Some recent studies have reported overall quality comparable to physician-authored discharge letters, particularly when LLMs are used with structured workflows or optimized prompting strategies [
1,
4,
5,
10]. Other studies have highlighted risks of hallucinations, omission of important details, and lower ratings in domains where contextual nuance is especially important, such as psychiatry [
2]. More recent work in pediatrics has been even more cautionary, showing physician-authored discharge letters to outperform unedited LLM-generated letters across all evaluated quality domains [
8]. Together with our own sensitivity analysis, this suggests that LLM performance in discharge documentation is unlikely to generalize uniformly across specialties, patient populations, and implementation settings. Local validation therefore remains essential. Recent work on discharge letter generation from structured clinical data and prompt-engineered note generation further suggests that documentation quality depends not only on model choice, but also on the degree of input standardization and task-specific prompt design [
16,
17].
This study has several strengths. First, the paired design allowed direct comparison of resident-written and GPT-5.2-generated letters derived from the same hospitalization and based on the same underlying clinical data, thereby reducing case-mix confounding. Second, the use of de-identified standardized source inputs and a predefined prompt limited unnecessary variability in model generation. Third, discharge letters were assessed in blinded fashion by board-certified ophthalmologists, which strengthens the clinical relevance of the evaluation and reduces overt source-related bias. These features provide a more rigorous framework than uncontrolled comparisons of unrelated clinician- and AI-generated documents.
It is important to underline the qualitative difference between the failure modes of generative models and those of human-authored discharge letters. Resident-written letters may be variable in style, occasionally terse, or affected by transcription errors, but the errors that occur are generally errors of omission or imprecision rather than fabrication. Generative LLMs, in contrast, introduce a distinct class of risks that does not have a clean analogue in human-authored letters: confabulation of plausible but unsupported clinical content, smoothing over of missing data with grammatically polished filler, surface fluency that masks subtle factual inaccuracies, and inconsistent prioritization of high-stakes elements such as laterality or operative details. Importantly, even a discharge letter generated “without errors” in a strict statistical sense may be misleading if it produces text that looks more complete or more confident than the underlying source data warrant. These failure modes are not eliminated by improvements in linguistic quality, and they argue strongly against treating LLM output as a finished discharge document. The implication for clinical practice is that any LLM-assisted workflow must include a structured clinician review step that explicitly targets these specific risks (hallucination, omission, and misprioritization), rather than relying on global “quality” impressions of the generated text.
A related question is how AI-generated discharge letters can be ensured to be clinically useful across the wide heterogeneity of ophthalmic conditions, which range from routine cataract surgery to complex vitreoretinal disease, glaucoma, uveitis, ocular trauma, and neuro-ophthalmology. The present study addressed this heterogeneity in three ways. First, the standardized prompt was condition-agnostic: it constrained the model to the structured input data rather than encoding diagnosis-specific templates, so the same prompt could be applied across the case mix encountered in routine inpatient ophthalmology. Second, the reviewer rating form explicitly assessed inclusion of high-stakes content elements (diagnosis, operation, findings, treatment, follow-up, discharge status) on a per-case applicability basis, so that performance could be evaluated separately on cases for which each element was clinically relevant. Third, the case-level paired design ensured that any condition-specific difficulty was matched between resident and model outputs. Across this heterogeneous case mix, the model performed broadly comparably to residents on most domains, but underperformed on documentation of operative details, suggesting that surgical and procedural cases are the clearest area in which prompt engineering, structured templates, or condition-specific safeguards will be needed. In practice, we would not expect a single general-purpose prompt to be optimal for every ophthalmic condition; rather, condition-aware prompt variants, retrieval-augmented context (for example, ICD-10 or surgery-code-conditioned templates), and checklist-based clinician review focused on high-stakes elements (laterality, procedure type, biometry, intraoperative findings) are the most likely routes to ensuring usefulness across the full breadth of ophthalmic presentations.
Several limitations should also be acknowledged. This was a single-center study conducted within one specialty, which limits generalizability. Ophthalmology discharge letters may be more structured than discharge documentation in specialties involving more complex systemic disease, prolonged admissions, or multimorbidity. The evaluation instrument was study-specific rather than a broadly validated discharge-letter quality scale, and some assessed domains, particularly clarity/structure, tone/professional phrasing, and conciseness, are inherently subjective. In addition, the primary analysis relied on one reviewer who assessed all available pairs, while the remaining two reviewers evaluated only partially overlapping subsets. Although this approach preserved a consistent paired comparison, it also means that the primary findings may have been influenced by individual reviewer tendencies. This concern is reinforced by the inter-rater agreement results, which were only fair to moderate for several domains and particularly low for clarity/structure, tone/professional phrasing, and conciseness. Accordingly, small observed differences in these domains should be interpreted cautiously. A more granular reading of
Table 4 reinforces this caveat. Using the Landis–Koch interpretation of weighted kappa, agreement was moderate for global quality (0.44–0.54) and largely moderate for completeness (0.37–0.53), mixed for accuracy (0.21–0.48), and only slight-to-fair for tone/professional phrasing (0.18–0.24). Agreement was poorer still for clarity/structure (−0.05 to 0.12) and was at or below chance for conciseness, with two of the three pairwise kappas negative (0.07, −0.05, and −0.07). This pattern indicates that the reviewers were not, in practice, applying a shared operational definition of conciseness, and that the apparently statistically significant primary-analysis finding for conciseness (with residents rated higher than GPT-5.2) should be interpreted as exploratory rather than confirmatory. The corresponding finding for tone/professional phrasing rests on slight-to-fair agreement and, while it was reproduced in the reviewer-adjusted sensitivity analysis, should likewise be treated as suggestive rather than definitive. By contrast, the absence of group differences on content-anchored outcomes such as accuracy, completeness, errors, and omissions is supported by relatively stronger agreement on those domains and is therefore more robust. We have retained the original stylistic results for transparency but recommend that downstream interpretation focus on the content-anchored domains and on the differential pattern of operation versus findings inclusion, which does not depend on subjective stylistic judgment.
Additional limitations relate to the scope of inference. The study evaluated document quality rather than operational outcomes, so it cannot determine whether GPT-5.2 would reduce drafting time, lower administrative burden, or improve clinician satisfaction in real-world use. Specifically, because the resident-written letters were drafted as part of routine clinical care before the study was initiated, the time required by residents to author each letter was not prospectively recorded and could not be reliably reconstructed retrospectively; the GPT-5.2 letters were produced in seconds per case once the standardized prompt and de-identified inputs were prepared, but this is not directly comparable to clinician drafting time, which includes data review, dictation, and editing. A direct head-to-head comparison of drafting time, editing burden, and downstream usability is therefore a key outcome for a planned prospective evaluation of this approach. Only one model and one prompting strategy were tested, and the results should not be generalized to other LLMs, future model versions, or alternative prompt designs. Moreover, although the model was explicitly instructed not to invent information, the risk of subtle hallucinations, misprioritization, or omission cannot be fully excluded, particularly in more complex cases. Prior work suggests that structured guidance, highlighted source information, and human verification may mitigate some of these risks, but not eliminate them entirely [
5,
6,
10]. Broader reviews similarly emphasize that implementation of LLM-supported clinical documentation depends not only on model performance, but also on governance, privacy safeguards, user trust, and robust mechanisms for identifying inaccuracies before documentation enters the medical record [
9]. Recent systematic and narrative reviews likewise suggest that documentation-support applications are among the more promising clinical use cases for LLMs, but that safe implementation still requires robust validation, standardization, and governance, particularly when generated text may influence downstream clinical decisions [
18,
19].
The practical implication of these findings is that LLMs should currently be viewed as potential drafting aids rather than replacements for clinician authorship or final review. In this role, GPT-5.2 may offer value by generating linguistically polished first drafts and helping standardize documentation format, while clinicians remain responsible for verifying factual accuracy, ensuring inclusion of specialty-relevant details, and maintaining accountability for the final record. This “AI-assisted drafting with clinician validation” model is also the interpretation most consistent with current evidence syntheses, which support cautious integration of LLMs into documentation workflows but not unsupervised deployment in safety-critical settings [
6,
9].
Future studies should evaluate this approach prospectively within real clinical workflow, ideally incorporating time savings, edit burden, and downstream usability as additional outcomes. It would also be valuable to test specialty-specific prompt refinements aimed at improving documentation of operative details and other ophthalmology-relevant elements, as well as checklist-based review systems designed to identify omissions before sign-off. This is particularly relevant in light of recent evidence showing that structured prompt engineering and constrained use of standardized clinical inputs may materially improve the quality of generated discharge documentation [
16,
17]. Multi-center studies with fully crossed reviewer designs and validated rating instruments would further strengthen the evidence base. Given the variability across published studies, implementation decisions should be based on local evidence rather than on assumptions that favorable findings from one setting will automatically transfer to another [
2,
8,
9,
10].
In conclusion, GPT-5.2-generated ophthalmology discharge letters showed similar performance to resident-written letters in several evaluated domains in the primary paired analysis, but differences in specific content elements and less favorable sensitivity analyses indicate that clinician oversight remains necessary to ensure accuracy, procedural completeness, and clinical usability. The results support LLMs as potentially useful tools for drafting ophthalmology discharge letters, but not as substitutes for clinician oversight. Safe implementation will require specialty-specific validation, careful prompt design, and robust human review focused on factual accuracy, procedural completeness, and clinical usability.