Abstract
Large language models (LLMs) are increasingly proposed for clinical summarization, yet evaluations relying on textual similarity overlook clinically consequential failure modes such as fabrication, omission, contradiction, and negation inversion. We treat medical summarization as controlled clinical compression, operationalizing a multidimensional framework of seven quality dimensions, an eight-category error taxonomy, and fourteen replicability components. The prototype runs on de-identified Portuguese-language Picture Archiving and Communication System (PACS)-derived multi-modality imaging reports, generating English impressions via local gemma4:latest under three prompt variants (R03) and temperature control. Factuality was scored claim-by-claim by an LLM judge with a two-pass protocol yielding 100% label coverage on 6190 claims, cross-checked by three external judges and two non-radiologist physicians. Factuality reached % (P1, 95% confidence interval (CI) 98.94–99.64), % (P2, 99.12–99.76), and % (P3, 91.60–93.79); critical errors were % (95% CI 0.35–0.71); critical omissions (source-to-summary) are outside the claim-level scheme. Pairwise Cohen’s among external judges ranged –. A controlled Portuguese (PT) → PT re-run () yielded significantly lower factuality than PT → EN ( to pp per variant, paired), restricting claims to PT → EN. Human annotation elicited a factuality–completeness criterion gap; BERTScore F1 correlated weakly with factuality (); temperature had no significant effect.
1. Introduction
Clinical documentation overload [1] creates a practical need for reliable summarization in radiology and related specialties [2,3]. LLMs are attractive because they can generate fluent text and follow instructions [4], but their clinical use is limited by the risk of unsupported findings, omissions, contradiction, and negation inversion. These risks cannot be controlled by fluency alone and cannot be measured adequately by textual overlap metrics.
We treat medical summarization as controlled clinical compression: producing concise clinical text that preserves source-supported findings while minimizing unsupported additions, omissions, contradictions, negation errors, and clinically unsafe reformulations. To make this notion operational, we apply a multidimensional evaluation framework defined within this article (Section 3.1): seven clinical quality dimensions (factuality, completeness, non-contradiction, clinical utility, clinical writing quality, source traceability, and privacy/safety), an eight-category clinical error taxonomy (fabrication, contradiction, critical omission, negation inversion or clinically unsafe polarity error, temporal error, attribution error, excessive generalization, dangerous ambiguity), and fourteen replicability components (corpus, sampling, model identity, prompt, parameters, decoding, language direction, annotation schema, evaluator model, human validation, statistical tests, hardware/runtime, data availability, code availability). The present article is self-contained: it defines the seven dimensions, eight error categories, and fourteen replicability components in Section 3.1 and reports the empirical prototype that instantiates them on multi-modality PACS-derived imaging reports; the taxonomy rows whose salience is paradigm-dependent are marked as such in the note below Table 2.
This article contributes three empirical elements. First, it provides a claim-level validation of the framework on real PACS-derived imaging reports from a local clinical corpus (multi-modality: CT 47.4%, ultrasound 17.6%, echocardiogram 9.2%, angiography 8.4%, MRI 6.8%, mammography 0.4%, unclassified 10.2%). Second, it quantifies robustness across prompt variants and temperature settings under appropriate paired statistical tests (Friedman omnibus with Wilcoxon signed-rank post hoc and Holm correction), showing a reproducible degradation associated with rigid four-item output structure. Third, it provides empirical evidence that automatic semantic similarity (BERTScore F1) diverges from claim-level factuality and that partial human re-annotation by non-radiologist physicians exposes a clinically important factuality–completeness criterion gap.
The remainder of this article is organized as follows. Section 2 surveys the related literature on clinical LLM-based summarization, faithfulness and factuality evaluation, clinical safety and completeness, LLM-as-judge protocols, and existing evaluation frameworks, concluding with the research gap this work addresses. Section 3 reports the technical prototype methodology, including the evaluation framework operationalized in this study (Section 3.1). Section 4 details the PACS-to-FHIR-LLM pipeline architecture. Section 5 presents corpus, factuality, prompt, metric, error, multi-judge, human-annotation, and temperature results. Section 6 discusses the findings relative to prior work. Section 7 reports limitations. Section 8 concludes with implications for prospective validation.
2. Related Work
This section surveys the five lines of prior work that directly inform the framework design and its positioning within the field: clinical LLM-based summarization (Section 2.1), faithfulness and factuality evaluation (Section 2.2), clinical safety and completeness (Section 2.3), LLM-as-judge protocols and multi-judge panels (Section 2.4), and the resulting research gap (Section 2.5).
2.1. Clinical LLM-Based Summarization
Clinical documentation generates a significant workload burden, with physicians spending substantial time on administrative tasks [1,5]. Large language models have emerged as a candidate technology to alleviate this burden, with reviews documenting their application to a broad range of clinical NLP tasks including named entity recognition, information extraction, and summarization [2]. In radiology, preliminary work has evaluated GPT-based models for automated report generation on single-modality chest radiographs [3], while Van Veen et al. demonstrated that adapted LLMs can outperform medical experts in completeness, correctness, and conciseness across four clinical summarization tasks (radiology reports, patient questions, progress notes, and doctor–patient dialogue), with the best model exhibiting a lower extent of potential harm than medical expert summaries [5].
Shared tasks have operationalized clinical summarization as a benchmarking problem. MEDIQA-Chat 2023 [6] attracted 17 teams and established corpora and baselines for automatic summarization of doctor–patient conversations; ACI-bench [7] introduced a large public dataset for AI-assisted visit note generation from clinical dialogue. These benchmarks are dominated by English-language text and cloud-based model evaluations. Multi-modality PACS-derived imaging reports, local on-premise deployment, and cross-lingual generation (source language ≠ output language) remain understudied configurations.
2.2. Faithfulness and Factuality Evaluation
Standard summarization metrics such as ROUGE [8] and BERTScore [9] measure lexical or embedding-level overlap and are inadequate for detecting factual inconsistencies [10,11]. Kryscinski et al. [11] established that neural abstractive summaries frequently violate factual consistency with the source document and proposed FactCC, a weakly supervised model that outperforms prior approaches on consistency detection. Maynez et al. [12] demonstrated through large-scale human evaluation that all evaluated summarization models produce substantial hallucinated content, distinguishing intrinsic hallucinations (contradicting the source) from extrinsic ones (not verifiable from the source), and showed that textual entailment correlates better with faithfulness than surface-overlap metrics.
Building on these foundations, SummaC [13] showed that sentence-level NLI outperforms document-level approaches for inconsistency detection (balanced accuracy 74.4%), while FRANK [14] introduced a typology of factual error categories enabling systematic metric benchmarking. SummEval [10] re-evaluated 14 automatic metrics against expert and crowdsourced human annotations across four quality dimensions, confirming low metric–human correlation throughout. In the medical domain, FactPICO [15] applied claim-level factuality evaluation to RCT-based summaries, finding that existing metrics correlate poorly with expert judgments at the instance level. Ji et al. [16] surveyed hallucination across NLG tasks, classifying failure modes and mitigation strategies; Farquhar et al. [17] proposed semantic entropy as a principled confabulation detection method applicable post hoc.
2.3. Clinical Safety, Harm-Aware Evaluation, and Completeness
The clinical consequences of summarization errors elevate evaluation from an NLP metric problem to a patient safety problem. Patient safety classification frameworks [18,19] distinguish error types by clinical consequence, a principle extended to LLM-generated clinical text by recent empirical studies. Williams et al. [20] conducted a blinded comparison of LLM-generated and physician-generated hospital discharge summaries at 100 inpatient encounters, finding comparable overall quality (3.67 vs. 3.77, ) and comparable individual-error harm potential (1.35 vs. 1.34, ), but with LLM summaries being significantly less comprehensive (3.72 vs. 4.13, ) and containing more unique errors per summary (2.91 vs. 1.82).
Asgari et al. [21] introduced CREOLA, a four-component framework (error taxonomy, experimental structure, clinical safety assessment, and annotation interface) applied to 450 consultation transcript–note pairs; across 12,999 annotated sentences, they observed a 1.47% hallucination rate and a 3.45% omission rate, with 44% of hallucinations classified as clinically major. Van Veen et al. [5] further confirmed that standard NLP metrics do not capture the completeness–correctness asymmetry relevant in clinical settings. Taken together, these studies establish that evaluation frameworks for clinical LLM summarization must account for both harm severity and the asymmetry between faithfulness and completeness.
2.4. LLM-as-Judge Protocols and Multi-Judge Panels
Zheng et al. [22] established LLM-as-a-judge as a scalable evaluation paradigm, showing that GPT-4 achieves over 80% agreement with human preferences on MT-bench and Chatbot Arena; their study identified systematic biases (position bias, verbosity bias, and self-enhancement bias) that compromise single-judge reliability. Panickssery et al. [23] demonstrated a linear correlation between self-recognition capability and self-preference bias, showing that LLMs have non-trivial accuracy at identifying their own outputs and consistently rate them higher when objective quality is equivalent. Wang et al. [24] showed that positional ordering alone can reverse quality rankings, enabling a weaker model to appear superior in 66 of 80 queries when serving as the reference position of the evaluating LLM.
Verga et al. [25] proposed PoLL (Panel of LLM Evaluators), demonstrating that a panel of diverse smaller models achieves higher Cohen’s with human judgements than a single large judge at seven times lower cost, and reduces intra-model bias through compositional diversity of judge families. In the clinical domain, CLEVER [26] extended the paradigm to blind, randomized evaluation by practicing medical doctors across three clinical rubrics: factuality, clinical relevance, and conciseness.
2.5. Existing Evaluation Frameworks and Research Gap
Prior work has advanced each of the five lines surveyed above, but existing frameworks address them in partial or isolated combinations. Table 1 summarizes the positioning of representative studies across five key methodological dimensions.
Table 1.
Positioning of representative prior studies relative to this work across five methodological dimensions. Claim-level: whether atomic factual claims are scored individually (vs. summary-level rubric or preference). Multi-judge : whether pairwise inter-judge agreement is reported explicitly.
Taken together, the studies summarized in Table 1 show that each of these methodological dimensions has been addressed in prior work, but not in combination. To the best of our knowledge, the literature lacks an integrated evaluation protocol that simultaneously covers real multi-modality PACS-derived clinical data, cross-lingual summarization with source language distinct from output language, local on-premise LLM deployment, atomic claim-level S/NS/C factuality annotation, and inter-judge agreement with explicit separation of auto-consistency from independent external evaluation. This gap motivates the evaluation framework and empirical prototype described in Section 3.
3. Methodology
3.1. Evaluation Framework Operationalized in This Study
The empirical study operationalizes the multidimensional framework summarized in Table 2. The seven clinical quality dimensions characterize what a clinically defensible summary must preserve; the eight-category error taxonomy characterizes how summarization can fail in clinically consequential ways; and the fourteen replicability components characterize what must be reported for an independent reader to reproduce or audit the result. The remaining methodology subsections (Corpus, Models and generation, Evaluation) instantiate the corresponding rows of this framework.
Table 2.
Evaluation framework operationalized in this prototype, with the components that this empirical study evaluates directly. Coverage entries reflect what this specific prototype evaluates on imaging-report summarization; several taxonomy rows are paradigm-dependent (see note immediately below the table).
The framework in Table 2 synthesizes existing evaluation instruments from three literatures and adapts them to imaging-report summarization. Rows marked “Yes” or “Partial” are the elements activated in this empirical prototype; the paradigm-dependence note below the table documents rows declared for framework completeness but not exercised on this corpus. The seven quality dimensions synthesize the factual-consistency literature [10,11,12,13,14,15] together with the clinical-rubric evidence [5,20,21] and the LLM-as-judge validity literature that motivates panel-of-judges consensus and bias mitigation [22,23,24,25]. The eight-category error taxonomy draws on hallucination and fabrication typologies [11,12,14] and on patient-safety and medication-error frameworks that operationalize clinical consequence [18,19]. The fourteen-component replicability index synthesizes established reproducibility reporting conventions from the NLP evaluation and summarization literature [10,14], the clinical NLP shared-task reporting practices [6,7,27], the expert-evaluation dimensions of CLEVER [26], and the safety-annotation protocol of CREOLA [21], extended with three components specific to local cross-lingual LLM deployment: model identity with version-pinned SHA checksum, language direction (PT → EN), and evaluator model specification.
3.2. Technical Prototype
A technical prototype for clinical imaging report summarization was implemented to apply and validate the framework of Section 3.1 on real PACS-derived clinical data.
Corpus. A total of 10,001 real PACS-derived imaging reports from a local hospital information system, written in Portuguese (PT), were used as the source population. The reports were provided to the research team by the data custodian after de-identification at source; the research team additionally applied heuristic PHI filters as a precautionary screening step under a local, non-networked workflow. These supplementary filters reduce but do not formally certify de-identification. Reports were converted to pseudo-FHIR format by extracting the DiagnosticReport.text.div field (source text, RTF → UTF-8 plain text, CP1252) and DiagnosticReport.conclusion (gold standard, not sent to the LLM). After four quality filters (scheduling-only reports with no clinical content, reports with no source text in the text.div field, source texts shorter than 30 characters, and heuristic detection of date-like patterns as a supplementary PHI precaution), 5819 eligible reports remained (58.2%). A proportionally stratified sample of was selected (seed = 42). The corpus is multi-modality, reflecting the heterogeneous case mix of the source PACS; the modality distribution of the 511 selected cases is reported in Table 3. Researchers seeking to independently replicate this protocol on public corpora may use MIMIC-CXR [27] for English-language radiology reports, MIMIC-III [28] for broader critical-care clinical text, the MEDIQA-Sum shared-task benchmark [6] for medical summarization, or ACI-bench [7] for clinical note generation; evaluation scripts and generation parameters are specified in the accompanying Supplementary Material.
Table 3.
Modality distribution of the stratified PACS-derived imaging cases (modality inferred from the source-text context field).
Models and generation. Three runs were executed: R01 (gemma3:4b, chronological, P1), R02 (gemma4:latest, chronological, P1), and R03 (gemma4:latest, stratified, P1 + P2 + P3). All models were run locally via Ollama. Generation parameters were those passed through the Ollama options object: num_ctx = 8192, temperature = 0.0 (greedy decoding), think = False; no explicit output-length cap (num_predict) was set, so the runtime default applied. Source texts are in Portuguese (PT) and outputs were generated in English (EN). This cross-lingual design was deliberate. English is the working language of the clinical NLP research community and of the benchmark corpora named above, and the FHIR/SNOMED-CT/ICD-10 interoperability layer operates in English. Contemporary LLMs are trained predominantly on English-language corpora and tend to produce higher fluency in English than in lower-resource languages. The PT → EN translation layer introduced by this design is recognized as a limitation and is formally declared in Section 7. The three prompt variants are: P1 (direct instruction with explicit rules), P2 (explicit fabrication restriction with traceability), and P3 (four-item output structure); their full text is reproduced verbatim in Section Prompt Variants (Verbatim).
Evaluation. Factuality was evaluated via LLM-as-a-judge (gemma4:latest) with S/NS/C labels (supported/not-supported/contradicted) per atomic claim. Annotation was executed in two passes with a SHA-pinned model checkpoint (blob digest and inter-pass audit trail in Supplementary Section S1). The primary pass covered 6190 claims from 511 stratified cases × three prompt variants. An automatic recovery pass then rescored 143 (case, variant) pairs (622 claims) whose primary-pass JSON output failed strict parsing. The recovery pass produced parseable labels on 143/143 pairs (100.0%) using a reinforced parser, yielding 100% label coverage on the full annotation grid; the two-pass provenance is preserved via the annotator column (gemma4:latest for the primary pass, gemma4:latest-auto-recovery-2026-08-29 for the recovery pass). Cross-model agreement was assessed independently of the primary annotator on a variant-stratified subsample of 306 claims, using three additional LLM judges from distinct architectural families: llama3.1:8b, mistral-small3.1, and qwen3.5:9b (Section 5.7). The external pairwise Cohen’s (Table 11) is the primary inter-model agreement statistic. It is reported separately from the auto-consistency statistics of any panel that includes gemma4. Partial human re-annotation was performed on the 209 claims of the P1 variant by two non-radiologist internal-medicine physicians (D1, D2), operating as a two-round criterion-elicitation exercise rather than an independent adjudication (Section 5.7). The task did not require primary image interpretation. Automatic metrics were limited to BERTScore F1 [9] with the xlm-roberta-base encoder [29] on pairs; ROUGE [8] was not executed due to cross-lingual incompatibility PT → EN.
The claim-level S/NS/C annotation is asymmetric by construction with respect to the two quality dimensions it can bound. Faithfulness (no unsupported additions) is directly observable at the claim level: every generated claim can be scored against the source text as Supported, Not Supported, or Contradicted. The resulting , , and rates are unbiased estimators of the fraction of generated content that is source-consistent. Completeness (no clinically relevant omissions), by contrast, requires the inverse operation: scanning the source text for clinically relevant findings that the summary failed to include, which the claim-level scheme cannot perform. The framework therefore treats faithfulness quantitatively via S/NS/C (Table 8) and completeness qualitatively via the LLM-rubric mean and the human criterion-elicitation exercise (Table 10, Section 5.7). The two dimensions must be reported together and cannot be collapsed into a single Supported/Not-Supported binary. This asymmetry is the methodological reason why the paradigm-dependence note below Table 2 lists Completeness as “Partial” and Critical omission as “No”.
3.3. Study Hypotheses and Decision Thresholds
Provenance of hypotheses. The four hypotheses formalized in Table 4 below were not deposited in a public pre-registration platform prior to data collection. They formalize the design intent that motivated the prompt-variant construction (P1/P2/P3), the choice of a claim-level S/NS/C evaluation scheme, and the temperature-sensitivity control (R04) at study inception in May 2026. The thresholds ( for , for ) were selected a priori as clinically defensible reference values rather than data-driven cut-offs, but were not archived externally. We therefore report these as design-time expectations formalized post hoc in revision: the tests operationalize decisions made before the R03 data were labeled and evaluated, but they are not pre-registered in the strict sense of the term. Where a hypothesis is not supported (), the finding is reported directly rather than reframed.
Table 4.
Study hypotheses and decision thresholds evaluated on the R03 corpus (post-recovery, claims/511 cases × 3 prompt variants) and on the R04 temperature-sensitivity subsample ( cases). Wilson 95% CIs are reported for proportions; paired Friedman + Wilcoxon (Holm-adjusted) for cross-condition comparisons. Detailed derivations are given in Section 5.3 and Section 5.8.
– are supported with wide margins. is not supported at the sample size and temperature range tested: gemma4:latest maintains across –, and no pairwise temperature contrast reaches statistical significance under Holm correction. Where is referenced in the results (Section 5.8) and in the discussion, the null finding is preserved rather than reframed. The four hypotheses are back-referenced in the corresponding results subsections (Section 5.3 for –; Section 5.8 for ), and their robustness is examined in the sensitivity analyses (Section 5.10).
4. Pipeline Architecture (PACS-to-FHIR-LLM)
The previous section established that medical summarization must be evaluated by paradigm. This section addresses a prior architectural question: how can a summary be audited against the source record? In radiology, the answer depends partly on interoperability. Medical images and reports are produced inside PACS/DICOM environments, while downstream clinical applications increasingly rely on FHIR resources for structured exchange [30,31]. If the summarization pipeline cannot identify the source field from which a claim was generated, factuality becomes difficult to verify even when the text appears clinically plausible.
A PACS system stores images in DICOM format, which includes study metadata but not the radiology report as structured text. The summarizable clinical content is the radiology report produced by the radiologist, which in FHIR is represented primarily in the DiagnosticReport resource (with fields .conclusion and .text), while ImagingStudy provides the technical context of the study. This distinction is critical: the summary is built on DiagnosticReport, not on ImagingStudy.
This article instantiates the pipeline PACS/DICOM → FHIR resources → input preparation → LLM → validation → summary as a four-layer architecture (Figure 1).
Figure 1.
Pipeline for traceable medical summarization from PACS-FHIR data using LLMs. Layer 1: Interoperability (PACS/DICOM → FHIR DiagnosticReport). Layer 2: LLM Input Preparation (free text/FHIR-derived/RAG). Layer 3: Controlled Generation (prompt + parameters). Layer 4: Multi-dimensional Validation (factuality/completeness/PHI/traceability).
The first layer resolves interoperability; the second prepares the input for the LLM (free text, FHIR-derived text, serialized FHIR, or RAG); the third executes generation under a controlled prompt; the fourth validates the output: factuality, completeness, non-contradiction, traceability, and PHI detection.
This pipeline extends the PACS-to-FHIR transformation described and validated in a prior publication by the same research group, which implemented an open-source Python 3.13 pipeline processing over 6000 computed tomography cases from a multidisciplinary department, producing structured FHIR R4 resources from JSON exports with RTF-embedded reports [32]. FHIR interoperability enables potential traceability (each claim in the summary can be referenced to a specific field of the source resource) but does not guarantee the quality of the generated summary. FHIR structuring facilitates verification; verification must be executed.
Prompt Variants (Verbatim)
The three prompt variants evaluated in R03 are reported verbatim below for reproducibility. Source-text placeholders are denoted {REPORT}; instantiated prompts containing real clinical text are not publicly released.
Note on modality wording. All three prompts contain the phrase “CT report” as a legacy of the initial CT-only pilot; the same prompts were then applied unchanged to the full multi-modality stratified sample (; CT , ultrasound , echocardiogram , angiography , MRI , mammography , unclassified ; Table 3). This wording was preserved deliberately so that the R03 grid remained internally comparable across variants and cases; a per-modality sensitivity analysis of (Section 5.10, Table 17) shows that the CT-labeled wording does not systematically disadvantage non-CT modalities under P1/P2, and the variant ordering () holds within each modality.

5. Results
This section reports the results of the prototype experiment described in Section 3. Results from R01 and R02 are definitive. Results from R03 are definitive: 6190/6190 claims annotated (100% label coverage after the two-pass protocol; see Section 5.7 and Supplementary Section S1).
5.1. Corpus and Generation Statistics
The corpus construction, quality-filter pipeline, and generation-run parameters are summarized in Table 5.
Table 5.
Corpus statistics and generation run data.
The four quality filters applied were as follows: (1) scheduling-only reports with no clinical content; (2) reports with no source text in the text.div field; (3) source texts shorter than 30 characters; and (4) heuristic detection of date-like patterns in the text (date_like), which identified 4121 files with probable patient-identifiable information, excluded as a PHI precaution.
5.2. Factuality by Run and Model Comparison
Table 6 reports the R01 and R02 factuality summary. The difference between gemma3:4b and gemma4:latest on the same chronological corpus was 14.8 percentage points in S% (98.0% vs. 83.2%), illustrating the importance of declaring the exact model version as a replicability criterion in the framework operationalized in this study. R02 confirms hypothesis (, Section 3.3) with a wide margin and hypothesis (critical errors ): 0 critical errors observed across 904 evaluated claims.
Table 6.
Factuality results by run (R01 and R02 definitive).
5.3. Robustness Across Prompt Variants (R03, Final Data)
Annotation coverage and missingness. R03 produced 6190 candidate claims across 1533 successful generations (511 cases × 3 prompt variants; 0 generation failures). Of these, 5568 claims (89.95%) were successfully labeled by the primary LLM-as-judge. The remaining 622 claims (10.05%), corresponding to 143 case×variant annotation calls, were not labeled because the judge response could not be parsed as valid JSON; internal logs identify structured-output parsing failures, including malformed or truncated JSON responses when the judge copied source-text support spans. These failures occurred after generation, not during summary production. Missingness was not perfectly balanced across prompt variants (Table 7); therefore, the magnitude of prompt differences should be interpreted with the sensitivity-analysis bounds reported below.
Table 7.
R03 annotation coverage and unlabeled-claim distribution by prompt variant. Source: r03_final_results.json.
Primary factuality result. Using labeled claims as denominator (Table 8):
Table 8.
Final factuality results by prompt variant (R03, post-recovery; denominator = all annotated claims per variant; , , ; total ). Wilson 95% CIs are given in the note below the table.
Full label coverage. The 622 claims (10.05% of the R03 grid) that initially returned JSON parse failures in the primary annotation pass were fully recovered via an automatic re-annotation pass with the identity-verified gemma4:latest checkpoint (Section 5.7, Supplementary Section S1); no imputation-based sensitivity analysis is required, since the merged corpus reaches 100% label coverage ( claims, zero missing).
Figure 2 visualizes the per-case distribution under each variant; the P3 distribution is substantially wider with a lower median. Variant P3, with a mandatory four-item output structure (main finding, secondary findings, negative findings, recommendation), generated essentially all NS claims in the final annotated data. Error analysis confirms that the forced structure induces fabrications when the source report does not contain information for one or more template items, for example, “secondary findings” or “follow-up recommendation” in brief normal ultrasound reports. This pattern illustrates the risk of rigid prompt structures, revisited in Section 5.10 and in the “Evaluator architecture and self-evaluation bias” paragraph of Section 7.
Figure 2.
Per-case distribution by prompt variant (R03 post-recovery). Box shows Q1–Q3 with median (black); whiskers extend to IQR; grey dots are per-case outliers. The reference line () is shown for scale. P1 and P2 concentrate near the ceiling with a small left tail; P3 exhibits a substantially wider distribution with a lower median and a heavy left tail, consistent with the structure-induced fabrication pattern discussed in Section 5.3.
Statistical significance under paired design. Because each case was generated under all three variants, the appropriate omnibus test for differences in per-case is the paired-samples Friedman test rather than Kruskal–Wallis (which assumes independent groups). Using all 511 paired cases (each case annotated under all three variants with 100% label coverage after the two-pass recovery protocol), the Friedman omnibus rejected the null hypothesis of equal per-case across the three variants (, , paired cases). Pairwise paired-sample post hoc tests (Wilcoxon signed-rank, two-sided, Holm-adjusted at the family level) show that P3 differs significantly from both P1 (, , , ) and P2 (, , , ), with medium-to-large effect sizes. P1 and P2 do not differ significantly (, , , ), indicating no statistically significant difference between the two instruction-based variants. Effect-size magnitudes are interpreted following the conventional small/medium/large thresholds for behavioral research [33]. These results formally confirm (Section 3.3): prompt structure has a statistically significant and practically meaningful effect on factuality, with the structured template variant (P3) performing substantially worse than the instruction-based prompts (P1 and P2).
5.4. Automatic Metrics and BERTScore F1 Finding
BERTScore F1 [9] with the xlm-roberta-base encoder [29] was calculated over N = 151 P1 generation pairs with available CONCLUSAO reference: F1 = 0.831 (SD = 0.014). The Pearson correlation between BERTScore F1 and the percentage of supported claims was r = 0.047 (n = 85 pairs with both metrics available), quasi-null.
This empirical result confirms that BERTScore F1 does not discriminate clinical factuality in this paradigm: generations with high semantic similarity to the reference may contain NS claims, and generations with lower similarity may be factually correct. ROUGE was not executed due to cross-lingual incompatibility: source reports are in Portuguese (PT); outputs were generated in English (EN); n-gram overlap is artificially low and lacks interpretative value. Surface-overlap metrics such as BLEU [34] and METEOR [35] were likewise not used in this study because they do not capture clinical factuality. The divergence between semantic-similarity scores and claim-level factuality is consistent with the broader literature on summarization evaluation re-evaluation [10] and with faithfulness studies showing that abstractive summaries can exhibit unsupported content despite high overlap or similarity scores [11,12,14]. NLI-based factuality detectors such as SummaC [13] and domain-specific frameworks for medical evidence such as FactPICO [15] further illustrate that factuality requires dedicated evaluation machinery beyond textual similarity.
5.5. Error Analysis
The LLM-as-a-judge assigned error categories to NS and C claims. In the merged annotated data (6190/6190 claims, 100% coverage after the two-pass protocol).
Critical-severity errors: 31 claims out of 6190 annotated (0.50%, Wilson 95% CI ), well below the threshold (, Section 3.3). The category-level counts are given in Table 9. Among the 156 primary-pass NS/C claims with taxonomy labels, Fabrication is the dominant category (130/156 = 83.3%), concentrated almost entirely in P3 (125/130 fabrication claims). The five unobserved taxonomy categories may characterize paradigms other than radiological reporting. Important methodological note: the claim-by-claim annotation scheme cannot detect critical omissions (category 3), as it would require an inverse perspective from the source; completeness is captured via LLM rubric (Section 5.6).
Table 9.
Error distribution by taxonomy category (final annotated data, n = 156 NS/C claims from the primary annotation pass; the 22 additional NS/C claims recovered by the second pass (11 NS and 11 C) are aggregated in the total but were not re-classified per category, since the qualitative dominance of the fabrication category is preserved).
5.6. Rubric Scores (Automatic Multidimensional Evaluation)
The LLM-as-a-judge assigned rubric scores 1–5 to all generations. Final R03 results are shown in Table 10 (mean ± SD, n = 1390 total evaluated cases):
Table 10.
Rubric scores by prompt variant (R03 final, mean ± SD, scale 1–5).
The completeness dimension is consistently the lowest across the three variants (3.67–3.91/5), consistent with the limitation of the claim-by-claim scheme for detecting omissions. P3 factuality is the lowest (4.70) and shows the highest variability (SD = 0.52), consistent with the NS pattern observed in Table 8.
5.7. Multi-Judge Validation and Human Annotation
Three additional judges from distinct architectural families were applied to a variant-stratified sample of 306 claims (150 S + 136 NS + 20 C, seed = 42), all running locally without access to gemma4:latest’s labels. The purpose of this analysis is twofold. First, it estimates the degree of independent agreement among externally selected LLM judges via pairwise Cohen’s (Table 11); this is the primary inter-model agreement statistic reported here, since it is uncontaminated by the primary annotator’s own vote. Second, it quantifies how the panel’s consensus label diverges from the primary annotator when gemma4 is excluded from the vote, which we report as “disagreement against the primary annotator” (Table 15) rather than as “external corroboration”, since a four-way panel that includes gemma4 constitutes an auto-consistency measurement rather than an independent adjudication. Drawing judges from distinct families mitigates known LLM-as-judge biases, in particular, self-preference for own generations [23], family-aligned scoring inconsistencies [24], and single-judge fragility addressed by panel-of-judges protocols [25]. The multi-judge protocol follows the LLM-as-judge baseline established for general-domain benchmarks such as MT-bench [22] and is adapted to clinical content along the lines of expert-review frameworks such as CLEVER [26]. Selection criteria for the three external judges were: (i) local availability via Ollama without external API quota; (ii) parameter count ≥8B for baseline capacity; (iii) architectural diversity (Meta, Mistral, Alibaba families) to reduce family-aligned bias; no selection was performed on prior agreement with gemma4:latest.
Table 11.
Pairwise Cohen’s among the three external judges on the stratified sample. This is the primary inter-model agreement statistic since it excludes the primary annotator from the panel.
The mistral–qwen pair reaches Almost Perfect agreement (0.908), higher than either model’s agreement with gemma4:latest ( and respectively, Table 12), while llama3.1:8b is the more discrepant member of the external panel. This structure was not observable in the four-way panel that anchored on gemma4:latest.
Table 12.
Inter-model agreement: independent judges vs. gemma4:latest annotations (n = 306).
Each external judge shows substantial agreement with gemma4:latest ( = 0.702–0.827); this is reported here as an auto-consistency statistic and should not be interpreted as external corroboration, since the reference annotator is the same model whose labels are being scored (Table 12 therefore complements, and does not replace, the pairwise external agreement of Table 11). No judge ever classified a Contradiction (C) as Supported (S), the highest-risk error in a clinical context. NS recall = 1.000 is achieved by two of three judges (mistral and qwen3.5). The C class remains the most difficult: the best-performing judge (qwen3.5:9b) correctly identifies 4/20 contradictions (F1 = 0.333), suggesting that full C detection requires domain-specific clinical knowledge; we return to the implications of this asymmetry for the reported critical-error rate in the Consensus Panel discussion below and in Limitations.
To assess whether annotation quality varies across prompt variants, the three judges were evaluated separately for P1, P2, and P3 using the same fixed sample (Table 13).
Table 13.
Cohen’s Kappa by prompt variant and judge model.
P3 (the variant with the highest density of real errors, NS = 129/186 claims in the sample) produces the most consistent inter-judge agreement across all three external models ( = 0.633–0.832). All three external models, running independently of gemma4:latest, reproduce the same variant ordering (), providing an independent cross-model replication of the H3 pattern from Section 5.3: prompt variant selection has a real and measurable effect on factuality that is not an artefact of the primary annotator’s decision criteria. This is a stronger claim than corroboration of gemma4’s labels, because it establishes that the ordering exists in the outputs themselves independent of the annotator applied.
Consensus panel (4-judge majority vote). To consolidate the multi-judge validation, labels from all four models (gemma4:latest, llama3.1:8b, mistral-small3.1, qwen3.5:9b) were combined through majority-vote consensus (≥3/4 agreement; gemma4:latest breaks 2-2 ties) on the same 306-claim stratified sample. Panel results are summarized in Table 14.
Table 14.
Four-judge consensus panel: label distribution and per-judge performance vs. consensus (n = 306 claims).
Two consensus computations are reported and compared in Table 15. The four-way panel (gemma4:latest as tie-breaker) reached unanimous agreement on % of claims and agreed with the gemma4:latest primary label on % (% override, ); this measures panel internal consistency, not external validation. Recomputed with only the three external judges under majority rule (≥2/3; 1-1-1 ties resolved to the most severe label, % of the sample), the non-circular external consensus differs from gemma4:latest on % of claims (). The four-way % figure was structurally attenuated by including gemma4 in the vote. Both figures remain valid under their respective definitions; the % is our revised primary estimate of independent disagreement.
The predominance of NS reclassifications ( under the four-way panel) indicates that gemma4:latest applies a broader definition of contradiction than any external majority, and this is amplified under the external-only vote: the three external judges never reach a majority for the C label on any sample claim. The Contradictory category as reported in this study is therefore anchored on the primary annotator; the reported critical-error rate (% overall, 95% CI –) is conditional on this scoping and revisited in Limitations. The per-variant external consensus preserves the H3 ordering (P1 = 77.8%, P2 = 82.5%, P3 = 20.4% supported, Table 15), consistent with gemma4:latest being a more permissive annotator than the external majority. Among external judges, qwen3.5:9b showed the highest C-sensitivity (F1 = 0.333 against the four-way consensus; Table 14) and llama3.1:8b the highest NS-sensitivity (F1 = 0.970) but no C detection.
Table 15.
Consensus panel comparison: four-way (includes gemma4) vs. external-only (3 judges), stratified sample. The external-only column excludes the primary annotator and constitutes the non-circular consensus.
Human re-annotation as criterion elicitation (two non-radiologist internal-medicine physicians). Two non-radiologist internal-medicine physicians (D1, D2) independently evaluated all 50 validation cases (209 claims, P1 variant). We frame this exercise a posteriori as a two-round criterion-elicitation protocol rather than an independent gold-standard adjudication, because the first round produced inter-annotator agreement statistically indistinguishable from chance (, asymptotic 95% CI under the Fleiss–Cohen standard-error approximation; % observed agreement, ), signaling that the two annotators had internalized materially different operational definitions of the Supported label. In round 1, D1 applied a clinical-completeness criterion (S = %, NS = %, C = %; 159/209, 48/209, 2/209) and D2 a stricter factuality-only criterion. After recalibrating D2 with an exhaustive completeness criterion (round 2), D2 yielded S = %, NS = %, C = % (188/209, 20/209, 1/209) and inter-annotator rose to (Fair, Landis & Koch, asymptotic 95% CI ; % observed agreement, 168/209). Of the 41 residual disagreements, 34 followed D1 = NS → D2 = S, confirming that the persistent divergence reflects criterion weighting rather than annotation error: the residual pp gap in S% quantifies the fraction of claims that are factually correct but clinically incomplete.
For the 200 claims matched to gemma4:latest P1 labels, D1 vs. gemma4:latest agreement yielded ( pp S% gap), attributable to the completeness–factuality criterion distinction rather than to model hallucination [16,17]. Read as a criterion-elicitation finding, this motivates the D1/D2 separation of Factuality and Completeness in the framework (Section 3.1) and demonstrates that any subsequent gold-standard human study must pre-register operational S/NS/C definitions, calibrate at least three clinically credentialed annotators against them, and blind annotators to LLM outputs. The present study does not perform that gold-standard step; its human component is a hypothesis-generating criterion-elicitation exercise, not a human gold standard.
5.8. Temperature Sensitivity (R04)
To empirically test hypothesis (Section 3.3; that generation temperature has a statistically significant effect on factuality), a controlled experiment was conducted using N = 50 cases selected by length-quintile stratification from the R03 corpus (511 cases, seed = 44). Three temperature levels (T = 0.0, T = 0.3, T = 0.7) were tested with three repetitions each (450 total generations), using gemma4:latest as generator; annotation was performed by gemma4:latest at temperature = 0.0. A within-case design controlled for case-level variance.
Mean factuality was (), (), and (). Because every case was evaluated under all three temperature levels, the appropriate omnibus test is the paired-samples Friedman test rather than Kruskal–Wallis (which assumes independent groups). Aggregating the three repetitions per (case, temperature) cell, the Friedman test on paired cases did not reject the null hypothesis of equal per-case across the three temperatures (, ). All paired post hoc Wilcoxon signed-rank tests were also non-significant under Holm correction at the family level (T = 0.0 vs. T = 0.3: ; T = 0.0 vs. T = 0.7: , ; T = 0.3 vs. T = 0.7: , ); the largest observed mean difference was vs. ( pp). is therefore not supported: gemma4:latest maintains factuality above across all tested temperatures.
This finding qualifies the replicability claim in Section 7 (“Computational reproducibility”). Temperature (component 7 of the framework in Section 3.1) remains a mandatory reporting parameter for transparency and cross-study comparability. However, its practical effect on factuality is paradigm-dependent: in the controlled clinical compression paradigm with gemma4:latest, temperature variation up to does not produce clinically or statistically meaningful factuality differences. This result should not be generalized to other models or paradigms where stochastic effects may be stronger.
5.9. Cross-Lingual Control (PT → PT vs. PT → EN)
Because the R03 experiment operates cross-lingually (Portuguese input, English output), a natural question is whether the observed factuality profile is a property of the model under this specific language direction or would replicate under a native-language (PT → PT) pipeline. To answer this, we ran a controlled re-execution on a modality-stratified subsample of cases (CT = 24, US = 9, Echo = 5, Unknown = 5, Angiography = 4, MRI = 3; seed = 42), holding the model checkpoint (gemma4:latest, identity-verified against R03), the generation parameters, and the three prompt-variant families constant, and changing only the output language: prompts and outputs were both in Portuguese (validated with clinical-radiology terminology, matching the source field DiagnosticReport.conclusion). The LLM-as-a-judge was also run in Portuguese-native mode over the resulting 1151 claims. Case-level per-variant was compared paired against the PT → EN labels of the same 50 cases from R03 using Wilcoxon signed-rank tests (Table 16).
Table 16.
Cross-lingual control: paired per-variant (PT → PT vs. PT → EN) on the same subsample. 95% Wilson CIs and Wilcoxon signed-rank paired p values (two-sided) are reported.
The PT → PT condition is systematically lower than PT → EN across all three variants, with pooled of pp (Wilcoxon paired , ). Two properties of the pre-declared PT → EN result nevertheless survive: (i) the variant ordering () is preserved in PT → PT (thus replicates under both language conditions); (ii) P3 remains the worst variant in each condition. The cross-lingual gap does not overturn the primary – conclusions but constrains their scope: the reported R03 factuality profile is specific to the cross-lingual PT → EN condition tested, and gemma4:latest does not exhibit the same profile under native-language generation. Mechanistic hypotheses (asymmetric claim decomposition between PT and EN outputs; language-dependent evaluator calibration by the same-family judge; language-specific idiomatic reformulation) are not adjudicated here; the honest interpretation is that language direction is a first-order factor that must be reported alongside model and prompt when replicating this framework, and this is now declared in Section 7 and in the framework replicability index (component 7).
5.10. Sensitivity Analyses
Three sensitivity analyses characterize the robustness of the R03 primary results to (i) the annotation-recovery pipeline, (ii) the composition of the LLM-as-a-judge panel, and (iii) the output-language direction. All three are back-referenced from the pre-specified hypotheses (Section 3.3) and inform the scope of the primary claims.
Pre-recovery vs. post-recovery. Before the automatic recovery pass, claims (89.95%) were labeled by the primary judge; the remaining 622 (10.05%) failed strict JSON parsing (Table 7). Pre-recovery per-variant (labeled-denominator convention) was P1 = 99.38%, P2 = 99.55%, P3 = 92.85%; the P2−P3 gap was pp. After the recovery pass (100% coverage, ), the figures become P1 = 99.38% [98.94, 99.64], P2 = 99.54% [99.12, 99.76], P3=92.77% [91.60, 93.79], with a gap of pp. The variant ordering () is preserved; the recovery step principally reveals critical errors previously masked by parse failures (global : , i.e., critical claims), and the gap magnitude is stable within pp. This robustness supports treating the recovery pass as a coverage-completion step rather than a substantive reanalysis. All primary figures reported in Section 5.3, Section 5.4, Section 5.5, Section 5.6 and Section 5.7 are post-recovery unless explicitly stated.
Panel composition: with vs. without gemma4:latest. A four-judge panel that includes the primary annotator constitutes an auto-consistency measurement rather than an independent adjudication. When the panel is restricted to the three external judges (llama3.1:8b, mistral-small3.1, qwen3.5:9b) under majority rule with most-severe fallback on 1-1-1 ties ( of the sample), the panel disagreement against the primary annotator rises from (four-way panel with gemma4 included as tie-breaker) to (external-only, Table 15). The pairwise Cohen’s among the three external judges remains within the Substantial-to-Almost-Perfect range (–, Table 11); the external judges therefore agree strongly among themselves while diverging more from the primary annotator than a self-inclusive panel would suggest. The primary factuality figures reported in Table 8 remain the authoritative / estimates of the primary annotator, but their inter-model support is quantified by the external-only agreement analysis rather than by the four-way panel. The Contradictory (C) label in particular is anchor-specific to gemma4:latest (Section 5.7 and Section 7).
Language direction: PT → EN vs. PT → PT. The controlled cross-lingual re-execution on modality-stratified cases (Section 5.9) shows a systematic PT → PT deficit of pp pooled (, ), consistent across all three variants (, , pp; all paired Wilcoxon ). The variant ordering is preserved under both directions ( replicates), but the absolute level does not: the R03 profile reported in Table 8 is specific to the PT → EN direction tested, and / should be read as PT → EN-conditional. This scope constraint is declared in the framework replicability index (component 7, Table 2) and revisited in Section 7.
Prompt wording generalization: CT vs. non-CT modalities. The three prompts (Section Prompt Variants (Verbatim)) contain the phrase “CT report” as a legacy of the CT-only pilot but were applied unchanged to the multi-modality R03 stratified sample. Stratifying by modality (Table 17) shows no systematic disadvantage of non-CT modalities: for P1 and P2, non-CT modalities pooled reach the same or slightly higher than CT ( pp for P1 and pp for P2; Wilson 95% CIs overlap heavily). Under the fragile variant P3, the pooled non-CT rate is pp below CT, driven principally by Echocardiogram ( [78.57, 88.83]), where the four-item template’s “negative findings” and “follow-up recommendation” slots are most likely to be filled by fabrication in the absence of a natural referent in the source report. Angiography (), MRI (), Ultrasound (), and Unknown () remain above the threshold of under P3, and above under P1/P2. The “CT report” wording does not, therefore, appear to induce a modality-general degradation of factuality; the variant ordering () holds within each modality.
Table 17.
Per-modality (R03 post-recovery, cases). Wilson 95% CIs in brackets. Mammography ( cases) is reported for completeness; its CIs are wide and non-informative. Source: sensitivity_by_modality.py + sensitivity_by_modality.json.
6. Discussion
The prototype confirms three findings central to the multidimensional framework operationalized in Section 3.1 on imaging-report summarization. First, factuality can be measured at claim level under a controlled local workflow: P1 and P2 achieved supported-claim rates above 99%, while P3 remained above 92% but with a clear structure-induced failure mode. Second, clinical-risk evaluation cannot be reduced to automatic similarity: BERTScore F1 was essentially uncorrelated with factuality. Third, human annotation reveals that factuality and completeness are separable dimensions, since a claim can be supported while the summary remains clinically incomplete. These findings apply to the paradigm and corpus tested; generalization to dialogue summarization, discharge-note generation, or compression-ratio-controlled summarization requires re-instantiation of the framework rows currently marked paradigm-dependent (paradigm-dependence note below Table 2).
The P3 degradation is the most actionable result. A rigid four-item template can pressure the model to fill slots for secondary findings, negative findings, or recommendations even when the source text does not support them. This result argues against assuming that more structure is always safer. In this corpus, explicit source-traceability constraints were safer than mandatory output slots.
Compared with prior empirical work, the prototype is narrower but more controlled. It does not claim clinical deployment readiness. Instead, it demonstrates that local LLM summarization can be evaluated through a reproducible factuality, multi-judge, and human-validation protocol. The results are compatible with recent evidence that LLMs can perform well in clinical summarization while still requiring harm-aware evaluation and expert oversight [5,20,21]. Our per-variant critical-error rate is overall (Wilson 95% CI ; P1, P2, P3). This figure is offered as a paradigm-internal reference point for prospective replications under matched corpora, languages, and annotation schemes, not as a cross-study performance claim. Hallucination and omission rates in adjacent clinical-NLP evaluations [21] operate on distinct English corpora and annotation schemes, and are not directly comparable numerically. Mapping the observed error severity onto established patient-safety taxonomies (the NCC MERP index for medication-error severity [18] and the WHO International Classification for Patient Safety [19]) is a natural next step before any prospective deployment.
Practical implications for framework adoption. Beyond its role as an empirical benchmark, the multidimensional framework operationalized in this study provides a structured pre-deployment evaluation protocol for clinical NLP teams. The fourteen replicability components (Section 3.1) function as a minimum documentation checklist: corpus provenance, SHA-pinned model identity, prompt specification, annotation schema, evaluator model, and statistical test choice together constitute an audit trail that enables independent verification of any reported factuality claim. The claim-level S/NS/C scheme and the eight-category error taxonomy (Table 2) enable teams to identify the dominant failure mode in their specific paradigm before clinical deployment: in this imaging-report corpus, the dominant mode was structure-induced fabrication under the P3 template (83.3% of all NS/C claims), a finding not detectable by summary-level rubric scores alone and one that dictated a direct prompt-design intervention rather than a model-level change. For clinical governance, the factuality–completeness criterion gap documented in Section 5.7 implies that human oversight mechanisms must specify in advance which dimension they are monitoring: reviewing LLM-generated impressions for unsupported additions (faithfulness) and reviewing for missed critical findings (completeness) are operationally distinct tasks that require different annotator profiles and different escalation protocols, and should not be collapsed into a single sign-off step. A defensible prospective deployment study would need, at minimum, to instantiate these framework components on an independent institutional corpus, run the full claim-level evaluation under at least one external judge family, and pre-register operational S/NS/C definitions with specialist radiologist input before any human annotation is treated as ground truth.
7. Limitations
The limitations of this study cluster into four domains: corpus and external validity; evaluator architecture; human-annotation coverage; and computational reproducibility. Each is stated together with the concrete mitigation already in place and the prospective action that would close the gap.
Corpus and external validity. The prototype was executed on a single-institution local PACS-derived corpus with a heterogeneous imaging case mix dominated by computed tomography (47.4%) but also including ultrasound, echocardiography, angiography, MRI, mammography, and unclassified imaging reports (Table 3), retrospective in design, with stratified cases generated and paired cases used for the prompt-variant analysis (100% pairing achieved after the two-pass annotation recovery). The reports are in Portuguese and outputs were generated in English. External validity is therefore limited not by single-modality restriction, but by single-institution provenance, retrospective design, local documentation style, uneven modality distribution, and the cross-lingual PT → EN direction. Generalization to other institutions, languages, report templates, modality distributions, and single-modality pipelines is not directly tested. The corpus cannot be publicly released due to institutional privacy restrictions, which limits external verification of the numerical results. Although the corpus was provided de-identified at source by the data custodian, the research team’s supplementary PHI screening relied on heuristic filtering rather than on an independently validated de-identification pipeline, which is acceptable for a methodological feasibility study but not for clinical deployment. Prospective extensions should replicate the protocol on at least one public corpus (MIMIC-CXR [27], ACI-Bench [7]), on a second institution, and per-modality, and should adopt a formally validated de-identification pipeline. A directly related scope constraint concerns language direction: the reported R03 numbers are for PT → EN generation. A controlled PT → PT re-execution on a stratified subsample using the same model, prompts (translated to PT), and parameters (Section 5.9, Table 16) yielded systematically lower across all three variants ( from to pp, all Wilcoxon paired). The variant ordering () survives but the absolute level does not. The empirical claim of this study is therefore explicitly restricted to the PT → EN cross-lingual condition tested; native-language deployment (in PT or in any other source language) requires separate validation with its own subsample and pre-registered thresholds.
Evaluator architecture and self-evaluation bias. The primary factuality evaluator and the generator share the same model checkpoint (gemma4:latest), which introduces self-evaluation bias. This bias was partially bounded, not eliminated, by two complementary analyses. First, three external LLM judges from distinct architectural families (llama3.1:8b, mistral-small3.1, qwen3.5:9b) rescored a variant-stratified sample; the resulting pairwise Cohen’s – among externals (Table 11) constitutes the primary bound on inter-model agreement independent of the primary annotator. Second, the partial human criterion-elicitation exercise (Section 5.7) triangulated the labels from a clinical perspective. We note two residual limitations. First, the primary label on the full corpus is still gemma4-generated; the external panel covers only % of claims and cannot be linearly extrapolated to per-variant on the full corpus. Second, the Contradictory category (C) is idiosyncratic to the primary annotator: on the external-only consensus of the sample, the three external judges never reach a majority for the C label (Table 15). The critical-error rate reported in Results (% on the merged corpus, 95% CI –) should therefore be read as an annotator-specific quantity rather than as a cross-model consensus rate. Prospective studies should pre-register a generator/evaluator family-separation policy, run the full corpus (not a sample) under at least one external judge, and pre-register operational definitions of C to enable cross-model reproducibility of the critical-error rate.
Human component is a criterion-elicitation exercise, not a gold standard. Human re-annotation was limited to claims (P1 variant) and to two non-radiologist internal-medicine physicians (D1, D2). The initial round produced (chance-level agreement), which we report as the primary elicitation finding rather than as a baseline to improve upon: the negative signals that the two annotators were scoring materially different underlying constructs (strict factuality vs. clinical completeness) and that any subsequent independent-adjudication study on this corpus requires a prior criterion-alignment step. The re-calibrated round produced (Fair, Landis & Koch, % observed agreement, ) with a directional disagreement pattern (34/41 disagreements followed ) that informs but does not substitute for a multi-rater expert-adjudication panel. The exercise is therefore reported as hypothesis-generating: it motivates the Factuality/Completeness separation in the evaluation framework (Section 3.1); the criterion divergence between annotators D1 and D2 (strict factuality vs. clinical completeness) is precisely the empirical anchor for this separation. It does not constitute a human gold standard against which the LLM labels can be scored. Prospective work should: (i) recruit at least three clinically credentialed annotators, including one specialist radiologist; (ii) pre-register operational definitions of S, NS, and C with worked examples; (iii) blind annotators to the LLM output; (iv) target human-annotated claims stratified by variant; and (v) report both raw agreement and with the marginal class prevalence.
Computational reproducibility. The Ollama model tag gemma4:latest is floating; the same tag may resolve to a different weight digest in the future. Strict bit-level reproducibility requires SHA-pinned model digests; the Ollama IDs and blob digests for all four models used in this study (generator + three-judge consensus panel) are reported in the Supplementary Material. Hardware details and runtime cost are documented in the Supplementary Material; wall-clock latency per claim was not benchmarked across hardware classes. Prospective replications should fix the hardware class explicitly and report latency.
Scope of the empirical claim. This study does not establish clinical utility, deployment safety, or downstream patient outcomes. It supports methodological feasibility of the multidimensional evaluation framework defined in Section 3.1 on a controlled local prototype, identifies concrete failure modes (structure-induced fabrication under P3, BERTScore–factuality decoupling, factuality–completeness criterion divergence), and yields a numeric reference (critical-error rate %, 95% CI –) that future prospective studies can use as a comparator. It does not authorize integration into clinical workflows; that step requires institutional review, prospective validation on independent corpora, and patient-safety risk assessment against established taxonomies [18,19].
8. Conclusions
The multidimensional evaluation framework operationalized in this article was tractable in a real local PACS-derived imaging summarization prototype. Five concrete conclusions follow.
(i) Factuality is high under instruction-style prompts but degrades under rigid output structure. P1 and P2 reached supported-claim rates of % and %, while the four-item structured-output variant P3 fell to %. The omnibus rejection under paired Friedman (, , ) and the Holm-adjusted Wilcoxon post hoc analyses confirm that rigid output templates pressure the model to fill empty slots and induce fabrication. Prompt design is therefore a first-order factuality control, not a stylistic choice.
(ii) Inter-rater agreement should be read as a pair of statistics, not a single number. The raw agreement rates between annotators and between human and LLM judge are uniformly high. Two non-radiologist physicians (D1, D2) on the P1 sample agreed on % of claims (168/209). The human annotator D1 agreed with the gemma4 judge on % of the matched subsample (154/200), with perfect recall for supported claims (RecallS ) and precision of (45/200 false positives by the LLM judge). The three independent LLM judges of distinct families agreed with gemma4 on –% of claims. The corresponding Cohen’s Kappa values are low to fair (D1–D2 ; D1–gemma4 ; external judges pairwise –; external judges vs. gemma4 auto-consistency –). Reading the raw agreement alone would understate the systematic skew, and reading the chance-corrected alone would overstate the disagreement: in a distribution where Supported claims dominate (∼90%), penalises the imbalance even when the raters substantively concur. Reporting both is therefore not redundant but mutually corrective; future studies should report both with the marginal class prevalence so the imbalance can be reasoned about explicitly.
(iii) The dominant disagreement is a criterion gap, not random noise. 93% of the false positives between human D1 and the gemma4 judge () and 83% of the D1–D2 disagreements (, pattern ) reflect strict factuality versus clinical completeness, not factual error detection. The framework should therefore separate these dimensions at the rubric level, as Table 2 does, and not collapse them into a single Supported/Not Supported binary.
(iv) Automatic semantic similarity does not substitute for claim-level factuality. BERTScore F1 was essentially uncorrelated with the supported-claim rate (, pairs). Replication packages that rely on similarity scores as a proxy for clinical correctness will under-detect the failure modes that drive harm in clinical summarization.
(v) Temperature is a reportable parameter, not a primary factuality lever. Across on paired cases, the paired Friedman omnibus did not reject equality (, ); the practical window stayed above % throughout. Temperature must still be reported for reproducibility, but in this paradigm it is not where the meaningful clinical risk lives.
Outlook. Future work should pursue five parallel extensions: (i) replicate the protocol on public clinical corpora and across multiple institutions and modalities; (ii) replicate under native-language conditions and across additional source languages to isolate cross-lingual effects (Section 5.9); (iii) pin model digests across the full judge panel; (iv) calibrate the eight-category error taxonomy with multidisciplinary expert panels; and (v) bring at least one specialist radiologist into the human re-annotation loop. The pre-deployment step that the prototype does not perform (mapping observed claim-level errors onto established patient-safety severity taxonomies) is, in our view, the next required milestone before any prospective clinical deployment can be defensibly considered.
Compression-aware evaluation. The present framework quantifies faithfulness at claim level but does not stratify by compression ratio. In paradigms where the summarization target has an explicit compression budget (e.g., discharge-note generation from full inpatient records, patient-facing rewrites with reading-level constraints, or executive summaries of multi-day encounters), the interaction of compression ratio with and is itself a research question. A compression-aware extension of the framework would (i) declare per-instance target compression ratios, (ii) stratify and by achieved compression, and (iii) evaluate whether tightly compressed outputs exhibit higher hallucination or omission rates than loosely compressed ones. On the present imaging-report corpus, this stratification is deferred to future work because the source DiagnosticReport.text.div and target DiagnosticReport.conclusion lengths follow institution-specific conventions that would need to be normalized before a compression-ratio analysis is clinically interpretable.
Supplementary Materials
The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/make8100305/s1: Figure S1: BERTScore F1 (xlm-roberta-base encoder) versus per-case claim-level on the P1 subset ( paired cases, Pearson ); Figure S2: Pairwise confusion matrices among the three external LLM judges (llama3.1:8b, mistral-small3.1, qwen3.5:9b) on the variant-stratified sample; Table S1: SHA-256 digests of the Ollama model blobs used in R03 (primary annotator and external panel); Table S2: Software stack at each annotation pass.
Author Contributions
Conceptualization, L.V., O.C., and R.M.; methodology, L.V.; software (pipeline implementation, multi-judge orchestration), L.V.; investigation (experiment execution on the local corpus), L.V.; resources (clinical data infrastructure and PACS access), M.D. and R.C.B.; data curation (corpus preparation and de-identification verification), L.V.; formal analysis (claim-by-claim factuality assessment and statistical analysis), L.V.; validation, L.V. and R.M.; writing—original draft preparation, L.V.; writing—review and editing, O.C. and R.M.; visualization, L.V.; supervision, O.C. and R.M. All authors have read and agreed to the published version of the manuscript.
Funding
This empirical validation is methodologically connected to the PACS-to-FHIR interoperability work [32], which forms part of BLOCKCHAIN.PT, Agenda “Descentralizar Portugal com Blockchain” (WP2: Saúde e Bem-estar), reference 02/C05-i01.01/2022.PC644918095-00000033, supported by the European Commission via the Portuguese Recovery and Resilience Plan (PRR/NextGenerationEU). The local Ollama-based prototype, claim-level evaluation pipeline, and manuscript preparation received no additional external funding.
Institutional Review Board Statement
Not applicable. The clinical imaging reports analyzed in this study were provided to the research team by the data custodian (BioGHP (Global Health Platform S.A.)) after full de-identification at source; the research team had no access to identifiable patient data at any point of the workflow. Under the applicable institutional policy and the Portuguese General Data Protection Regulation framework, secondary use of fully de-identified clinical data for methodological evaluation does not constitute human-subjects research and therefore does not require Institutional Review Board approval. The local processing pipeline was non-networked and did not transmit any report content to external services.
Informed Consent Statement
Not applicable. The study used fully de-identified secondary data; no patient interactions, recruitment, or clinical interventions were performed.
Data Availability Statement
The data presented in this study are available on request from the corresponding author, subject to data-custodian approval. The data are not publicly available due to privacy restrictions on patient-derived clinical content: the underlying PACS-derived multi-modality imaging reports and the instantiated prompts containing real clinical text cannot be released. The aggregate, non-sensitive contributions of this study (per-variant factuality results, multi-judge agreement matrices, sensitivity-analysis bounds, prompt templates with placeholders, evaluation schemas, claim-level annotation rubrics, model SHA-256 digests, and hyperparameters) are included in this article and in the accompanying Supplementary Material, and are publicly available in the Zenodo repository under a CC BY 4.0 licence at https://doi.org/10.5281/zenodo.22962371 (accessed on 21 September 2026).
Acknowledgments
The authors thank the University of Leiria and Oeste and the Centre for Informatics and Systems of the University of Coimbra (CISUC) for institutional support, and the clinical data custodian that provided the de-identified PACS-derived multi-modality imaging report corpus used in this empirical study. This work is partly supported by the BLOCKCHAIN.PT project, Agenda “Descentralizar Portugal com Blockchain” (WP2: Saúde e Bem-estar), reference 02/C05-i01.01/2022.PC644918095-00000033, financed by European funds through the Portuguese Recovery and Resilience Plan (PRR/NextGenerationEU). During the preparation of this manuscript, the authors used Claude (Anthropic; model versions Opus 4.6 and 4.7, accessed via the Claude Code CLI; May–June 2026) for bibliographic verification against PubMed/NCBI E-utilities and Crossref APIs, structural review of the manuscript, editing for grammar, consistency, and LaTeX formatting, and assistance with the interpretation of selected statistical and empirical results. All AI-generated output was reviewed, validated, and edited by the authors, who take full responsibility for the content of this publication, including all interpretations and conclusions. The generative AI tool played no role in the design or execution of the empirical study itself; the experimental pipeline, prompts, evaluator configuration, and human annotation protocol were defined and executed by the authors prior to any drafting or interpretive assistance.
Conflicts of Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Manuel Dias, Ricardo Correia Bezerra are employees of BioGHP, who participated in the research and development consortium, provided technical support, but did not provided funding for this work.
References
- Furlow, B. Information overload and unsustainable workloads in the era of electronic health records. Lancet Respir. Med. 2020, 8, 243–244. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nazi, Z.A.; Peng, W. Large language models in healthcare and medical domain: A review. Informatics 2024, 11, 57. [Google Scholar] [CrossRef] [Scilit]
- Nakaura, T.; Yoshida, N.; Kobayashi, N.; Shiraishi, K.; Nagayama, Y.; Uetani, H.; Kidoh, M.; Hokamura, M.; Funama, Y.; Hirai, T. Preliminary assessment of automated radiology report generation with generative pre-trained transformers: Comparing results to radiologist-generated reports. Jpn. J. Radiol. 2024, 42, 190–200. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, P.; Bubeck, S.; Petro, J. Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine. N. Engl. J. Med. 2023, 388, 1233–1239. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Van Veen, D.; Van Uden, C.; Blankemeier, L.; Delbrouck, J.B.; Bhatt, S.; Chheda, T.; Pareek, A.; Polacin, M.; Reis, E.P.; Seehofnerová, A.; et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat. Med. 2024, 30, 1134–1142. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shor, J.; Bi, R.A.; Venugopalan, S.; Ibara, S.; Goldenberg, R.; Rivlin, E. Overview of the MEDIQA-Chat 2023 shared tasks on the summarization and generation of doctor-patient conversations. In Proceedings of the 5th Clinical NLP Workshop at ACL; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 1–23. [Google Scholar] [CrossRef] [Scilit]
- Yim, W.W.; Fu, Y.; Ben Abacha, A.; Snider, N.; Lin, T.; Yetisgen, M. ACI-bench: A novel ambient clinical intelligence dataset for benchmarking automatic visit note generation. Sci. Data 2023, 10, 586. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, C.Y. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out (Workshop at ACL 2004); Association for Computational Linguistics: Barcelona, Spain, 2004; pp. 74–81. Available online: https://aclanthology.org/W04-1013/ (accessed on 31 August 2026).
- Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K.Q.; Artzi, Y. BERTScore: Evaluating text generation with BERT. In Proceedings of the 8th International Conference on Learning Representations (ICLR 2020), Online, 26–30 April 2020; Available online: https://openreview.net/forum?id=SkeHuCVFDr (accessed on 31 August 2026).
- Fabbri, A.R.; Kryściński, W.; McCann, B.; Xiong, C.; Socher, R.; Radev, D. SummEval: Re-evaluating summarization evaluation. Trans. Assoc. Comput Linguist. 2021, 9, 391–409. [Google Scholar] [CrossRef] [Scilit]
- Kryscinski, W.; McCann, B.; Xiong, C.; Socher, R. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, 16–20 November 2020; pp. 9332–9346. [Google Scholar] [CrossRef] [Scilit]
- Maynez, J.; Narayan, S.; Bohnet, B.; McDonald, R. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th Annual Meeting of the ACL, Seattle, WA, USA, 5–10 July 2020; pp. 1906–1919. [Google Scholar] [CrossRef] [Scilit]
- Laban, P.; Schnabel, T.; Bennett, P.N.; Hearst, M.A. SummaC: Re-visiting NLI-based models for inconsistency detection in summarization. Trans. Assoc. Comput. Linguist. 2022, 10, 163–177. [Google Scholar] [CrossRef] [Scilit]
- Pagnoni, A.; Balachandran, V.; Tsvetkov, Y. Understanding factual errors in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the 2021 Conference of the North American Chapter of the ACL: Human Language Technologies (NAACL-HLT), Online, 6–11 June 2021; pp. 4812–4829. [Google Scholar] [CrossRef] [Scilit]
- Joseph, S.; Chen, L.; Trienes, J.; Göke, H.; Coers, M.; Xu, W.; Wallace, B.; Li, J.J. FactPICO: Factuality evaluation for plain language summarization of medical evidence. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), Long Papers; Association for Computational Linguistics: Bangkok, Thailand, 2024; pp. 8437–8464. [Google Scholar] [CrossRef] [Scilit]
- Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.J.; Madotto, A.; Fung, P. Survey of hallucination in natural language generation. ACM Comput Surv. 2023, 55, 248. [Google Scholar] [CrossRef] [Scilit]
- Farquhar, S.; Kossen, J.; Kuhn, L.; Gal, Y. Detecting hallucinations in large language models using semantic entropy. Nature 2024, 630, 625–630. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- National Coordinating Council for Medication Error Reporting and Prevention (NCC MERP). NCC MERP Index for Categorizing Medication Errors; NCC MERP: Rockville, MD, USA, 2001; Revised 2022; Available online: https://www.nccmerp.org/types-medication-errors (accessed on 31 August 2026).
- World Health Organization. Conceptual Framework for the International Classification for Patient Safety; Version 1.1 (Final Technical Report) Document WHO/IER/PSP/2010.2; WHO: Geneva, Switzerland, 2009; Available online: https://www.who.int/publications/i/item/WHO-IER-PSP-2010.2 (accessed on 31 August 2026).
- Williams, C.Y.K.; Subramanian, C.R.; Ali, S.S.; Apolinario, M.; Askin, E.; Barish, P.; Cheng, M.; Deardorff, W.J.; Donthi, N.; Ganeshan, S.; et al. Physician- and large language model-generated hospital discharge summaries: A blinded, comparative quality and safety study. JAMA Intern. Med. 2025, 185, 818–825. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Asgari, E.; Montana-Brown, N.; Dubois, M.; Khalil, S.; Balloch, J.; Yeung, J.A.; Pimenta, D. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digit Med. 2025, 8, 274. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zheng, L.; Chiang, W.L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.P.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems 36 (NeurIPS 2023). 2023. Available online: https://proceedings.neurips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.html (accessed on 31 August 2026).
- Panickssery, A.; Bowman, S.R.; Feng, S. LLM Evaluators Recognize and Favor Their Own Generations. Advances in Neural Information Processing Systems 37 (NeurIPS 2024). Available online: https://dl.acm.org/doi/abs/10.5555/3737916.3740113 (accessed on 13 September 2026).
- Wang, P.; Li, L.; Chen, L.; Cai, Z.; Zhu, D.; Lin, B.; Cao, Y.; Kong, L.; Liu, Q.; Liu, T.; et al. Large language models are not fair evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024), Long Papers; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
- Verga, P.; Hofstatter, S.; Althammer, S.; Su, Y.; Piktus, A.; Arkhangorodsky, A.; Xu, M.; White, N.; Lewis, P. Replacing judges with juries: Evaluating LLM generations with a panel of diverse models. arXiv 2024, arXiv:2404.18796. [Google Scholar]
- Kocaman, V.; Kaya, M.A.; Feier, A.M.; Talby, D. Clinical large language model evaluation by expert review (CLEVER): Framework development and validation. JMIR AI 2025, 4, e72153. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Johnson, A.E.W.; Pollard, T.J.; Berkowitz, S.J.; Greenbaum, N.R.; Lungren, M.P.; Deng, C.-Y.; Mark, R.G.; Horng, S. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Sci. Data 2019, 6, 317. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Johnson, A.E.W.; Pollard, T.J.; Shen, L.; Lehman, L.H.; Feng, M.; Ghassemi, M.; Moody, B.; Szolovits, P.; Celi, L.A.; Mark, R.G. MIMIC-III, a freely accessible critical care database. Sci. Data 2016, 3, 160035. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzmán, F.; Grave, E.; Ott, M.; Zettlemoyer, L.; Stoyanov, V. Unsupervised cross-lingual representation learning at scale. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 8440–8451. [Google Scholar] [CrossRef] [Scilit]
- HL7 International. HL7 FHIR Release 4 (v4.0.1); HL7 International: Ann Arbor, MI, USA, 2019; Available online: https://hl7.org/fhir/R4/ (accessed on 31 August 2026).
- National Electrical Manufacturers Association (NEMA). Digital Imaging and Communications in Medicine (DICOM) Standard; Document PS 3.1-2024; NEMA: Rosslyn, VA, USA, 2024; Available online: https://www.dicomstandard.org/current/ (accessed on 31 August 2026).
- Vera, L.; Craveiro, O.; Malheiro, R.; Pereira, J.; Silva, P.; Dias, M.; Bezerra, R. PACS-to-FHIR transformation: A case of study with computed tomography. Procedia Comput. Sci. 2026, 278, 1342–1349. [Google Scholar] [CrossRef] [Scilit]
- Cohen, J. Statistical Power Analysis for the Behavioral Sciences, 2nd ed.; Lawrence Erlbaum Associates: Hillsdale, NJ, USA, 1988. [Google Scholar]
- Papineni, K.; Roukos, S.; Ward, T.; Zhu, W.J. BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (ACL ’02); Association for Computational Linguistics: Stroudsburg, PA, USA, 2002; pp. 311–318. [Google Scholar] [CrossRef] [Scilit]
- Banerjee, S.; Lavie, A. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization; Association for Computational Linguistics: Ann Arbor, MI, USA, 2005; pp. 65–72. Available online: https://aclanthology.org/W05-0909/ (accessed on 31 August 2026).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.

