1. Introduction
Brain tumor management relies heavily on magnetic resonance imaging (MRI) throughout the clinical pathway, from initial detection to lesion characterization, response assessment, and treatment planning. In neuro-oncology, standardized imaging frameworks such as the Response Assessment in Neuro-Oncology (RANO) criteria [
1] and the RANO criteria for brain metastases (RANO-BM) [
2] emphasize consistent evaluation of lesion size, enhancement pattern, multiplicity, and interval change. In parallel, structured reporting frameworks such as the Brain Tumor Reporting and Data System (BT-RADS) have been introduced to improve the consistency, completeness, and clinical usability of brain tumor MRI reports [
3,
4]. Together, these developments highlight the need for imaging workflows that are accurate, reproducible, and standardized across readers and clinical settings.
Brain tumor MRI interpretation remains labor-intensive. A prior time-motion analysis of brain MRI interpretation reported that the total interpretation workflow, measured from study opening to report signing, required approximately 10–18 min per case across reader roles, with time distributed among image viewing, report transcription, obtaining clinical data, education, and other related activities [
5]. Prolonged reporting turnaround time can affect departmental efficiency and may also influence downstream clinical operations. Assistive artificial intelligence (AI), including structured reporting tools and large language model (LLM)-based support systems, has therefore attracted interest as a means of improving reporting efficiency, standardization, and quality [
6,
7]. Although recent studies have reported favorable performance for AI-assisted reporting paradigms, their value in disease-specific clinical brain tumor workflows remains to be validated.
Artificial intelligence (AI) includes conventional machine learning models and deep-learning methods that learn imaging representations directly from data. These approaches differ in data requirements, interpretability, and validation needs. Recent systematic reviews in neurosurgical, neuroendocrine, and neuroimaging applications highlight rapid growth of clinical AI while emphasizing heterogeneity and the need for external validation [
8,
9,
10].
A similar challenge exists on the treatment side. Tumor delineation for surgery and radiotherapy planning is time-consuming and subject to interobserver variability, particularly when lesions are multiple, irregular, or adjacent to critical neuroanatomical structures [
11]. Recent reviews and clinical evaluations have shown that AI-based auto-segmentation can reduce contouring burden and improve reproducibility when used under physician supervision [
12]. In brain tumor imaging, meta-analytic evidence suggests that deep learning models have demonstrated performance in detection and segmentation across brain metastases, meningiomas, gliomas, and vestibular schwannomas, supporting their use as assistive tools in treatment planning rather than as fully autonomous systems [
13,
14,
15,
16].
However, most prior studies have evaluated AI assistance for diagnostic reporting and contouring separately. In brain tumor care, these processes are clinically connected: diagnostic MRI interpretation informs follow-up decisions, multidisciplinary treatment strategy, target delineation, and treatment delivery. AI-assisted tools may be useful not only for isolated algorithmic tasks, but also for creating shared lesion-level information that can be carried from diagnostic reporting into treatment planning. When lesions are preliminarily marked by AI during image review, physicians may verify report content, communicate lesion location and extent, and transfer relevant imaging information to downstream contouring workflows.
For physician-facing AI tools, workflow efficiency is a key clinical consideration. Even if an AI model demonstrates acceptable standalone diagnostic or segmentation performance, its clinical utility may be limited if its use increases reader workload, prolongs reporting or contouring time, or disrupts routine practice. From this perspective, time efficiency provides an appropriate primary endpoint for workflow evaluation: an AI-assisted workflow should reduce, or at least not increase, the physician burden required to complete clinical tasks. Measures of lesion detection, report consistency, segmentation overlap, and volumetric agreement were therefore considered complementary outcomes to determine whether any time savings were achieved without compromising output quality. This rationale led to the present study, which evaluates whether AI assistance is associated with changes in workflow efficiency, consistency, and reproducibility in two connected use scenarios: brain tumor MRI report preparation and treatment-planning tumor delineation, using controlled retrospective paired workflow testing (
Figure 1).
2. Materials and Methods
2.1. Study Design
This retrospective controlled study evaluated physician workflows with and without AI assistance in two linked clinical tasks: diagnostic MRI reporting and treatment-planning tumor delineation. The evaluation consisted of two stages: first, the establishment of a consensus reference standard from 30 retrospective MRI cases by an independent three-physician panel external to the study reader group; second, task-specific reader evaluations using the same cases. The reference-standard physicians did not participate in the subsequent diagnostic reporting or treatment-planning delineation reader-evaluation sessions. In both tasks, the AI-assisted session was performed first, followed by a 3-week washout period and then the physician-only session. Although the same 30 cases were used across sessions, case identifiers were de-identified and case presentation order was randomized separately for each physician and each session to reduce case recognition, sequence effects, and recall bias. The workflow condition order was fixed, with AI-assisted sessions performed before physician-only sessions. This sequence was selected under the assumption that any residual case familiarity after the washout would be more likely to benefit the later physician-only session rather than the earlier AI-assisted session, thereby reducing the likelihood of overestimating the apparent effect of AI assistance. However, the direction and magnitude of order effects cannot be determined from this design; therefore, the study should be interpreted as a fixed-sequence paired comparison rather than a randomized or counterbalanced design. Recall, learning, case familiarity, fatigue, or workflow adaptation may still have influenced the results despite the washout.
For diagnostic reporting, two readers generated brain tumor MRI reports under both AI-assisted and physician-only conditions. In the AI-assisted condition, MRI-based lesion localization and structured lesion-level information derived from the AI output were available for physician review and report preparation; in the physician-only condition, readers interpreted the same cases without AI support. Report generation time was measured from the start of image review to completion of the finalized report, thereby including image review, interpretation, and report generation. These measurements did not include worklist queue time, retrieval of clinical history outside the study materials, comparison with prior examinations outside the supplied study images, report-signing delay, or clinician notification.
For treatment-planning tumor delineation, two contouring readers generated final lesion contours under both AI-assisted and physician-only conditions. In the AI-assisted condition, readers reviewed and edited the AI system-generated preliminary tumor contours as needed before finalization; in the physician-only condition, readers manually delineated the same cases without AI support. Delineation time was recorded separately for each individual lesion. For cases with multiple lesions, the active drawing, editing, or contour-confirmation time for each lesion was measured independently rather than only as a single total case time. The recorded time reflected only active contouring or editing activity; image browsing, case review, and software operation time outside manual contouring or editing were excluded. These task-specific timing definitions reflect the different workflow endpoints of the two experiments: the reporting experiment evaluated the time required to review images and produce a finalized report, whereas the contouring experiment evaluated the active physician delineation or editing time required to generate treatment-planning contours. The overall study design and workflow are summarized in
Figure 2.
2.2. Sample Size Rationale and Power Consideration
This study was designed as a retrospective paired workflow-efficiency evaluation rather than a comprehensive validation study of lesion-level diagnostic or segmentation performance. Because each case was evaluated under both physician-only and AI-assisted conditions, the primary analysis was planned as a paired, within-case comparison. The primary workflow-efficiency variable was defined as the case-level relative workflow–time reduction:
where
Ri is the relative workflow–time reduction for case
i,
TPO,i is the time required under the physician-only condition, and
TAI,i is the time required under the AI-assisted condition.
Because reliable prior estimates of the standard deviation of paired relative workflow–time reduction were not available for this specific clinical setting, the sample size rationale was based on a standardized paired-effect framework and estimation precision rather than on a fixed absolute time-reduction threshold. With 30 paired cases, the study provides approximately 80% power to detect a medium standardized paired effect size (
dz = 0.53) using a two-sided paired-comparison framework at alpha = 0.05 [
17]. In this context,
dz represents the mean paired workflow–time difference divided by the standard deviation of the paired differences; it is a standardized measure of paired workflow change rather than a percentage time reduction. In addition, 30 paired cases allow the mean paired relative workflow–time reduction to be estimated with a 95% confidence interval half-width of approximately 0.37 standard deviations, providing a precision-based rationale for estimating the workflow–time effect and its variability. The same sample size also allowed balanced representation of the three target tumor categories. Therefore, 30 retrospective paired cases were included, with 10 cases each of vestibular schwannoma, meningioma, and brain metastasis. Lesion-level sensitivity, precision, report consistency, Dice coefficient, and volumetric agreement were treated as secondary descriptive performance measures and were not used as the basis for sample size determination.
2.3. Study Population and Image Acquisition
This retrospective study was approved by the Institutional Review Board (IRB; protocol code SE24344C; approval date: 23 July 2024), and informed consent was waived because only de-identified retrospective imaging data were analyzed. Cases were retrospectively selected to provide 10 cases in each target tumor category and required diagnostic-quality contrast-enhanced T1-weighted and T2-weighted imaging; the cohort was not intended to represent a consecutive or population-based sample. This study included MRI data from a total of 30 clinical cases, comprising 10 vestibular schwannomas, 10 meningiomas, and 10 brain metastases. The patient cohort consisted of 21 females and 9 males, with a mean age of 60.2 ± 13.1 years (range: 32–84 years).
All imaging data were acquired using either 1.5 Tesla (1.5 T; n = 18, 60.0%) or 3.0 Tesla (3.0 T; n = 12, 40.0%) MRI systems to reflect real-world clinical heterogeneity. Examinations were performed on multi-vendor platforms, including GE Healthcare scanners (Signa HDxt and Optima MR450w; n = 19), Siemens Healthineers scanners (MAGNETOM Verio and Lumina; n = 6), and Philips Healthcare scanners (Ingenia Elition X; n = 5).
The standard imaging protocol for all cases included multi-contrast sequences, specifically incorporating both contrast-enhanced T1-weighted imaging (CE-T1WI) and T2-weighted imaging (T2WI). Detailed acquisition parameters varied according to institutional protocols and scanner models. Slice thickness ranged from 0.9 mm to 3.0 mm, with 3.0 mm being the most prevalent value (66.7%).
2.4. Reference Standard and Physician Participants
A reference standard dataset was established before formal workflow evaluation by an independent three-physician panel external to the study reader group. The panel physicians did not participate in the diagnostic reporting or treatment-planning delineation experiments. Three board-certified radiology/neuroradiology attending physicians, each with more than 10 years of clinical experience, independently reviewed and annotated the 30 included cases. After independent annotation, lesion candidates were compared across the three physicians. Lesions annotated by two or more physicians were included in the reference standard, and the union of their annotated regions was used as the initial ground-truth contour. Lesions annotated by only one physician were flagged for consensus review. During the consensus meeting, the three physicians jointly reviewed these single-reader lesion candidates, determined whether each lesion should be included in the reference standard, and confirmed or adjusted the final contour boundary when inclusion was agreed upon. The final reference standard database, therefore, incorporated both multi-reader lesion agreement and consensus adjudication for discordant or single-reader findings. This database served as the ground truth against which physician-only, AI-assisted physician, and AI-generated preliminary outputs were compared. The union-based contour strategy was used to reduce exclusion of potential lesion extent, but it may influence precision and overlap metrics compared with majority-vote or probabilistic fusion methods.
Four physicians participated in the study. The diagnostic reporting task was performed by two neuroradiologists with 25 and 10 years of clinical experience in brain MRI interpretation, respectively. The treatment-planning segmentation task was performed by one neurosurgeon with 35 years of experience and one radiation oncologist with 10 years of experience in neuro-oncology treatment planning and contouring. All readers evaluated both AI-assisted and non-AI-assisted workflows within their respective tasks. For each evaluation arm, readers were selected to capture variation in clinical experience, allowing assessment of workflow effects across physicians with different levels of task-specific expertise. Because only two readers performed each task, reader-level results were interpreted descriptively.
2.5. AI System
DeepBT Detector-Plus v1.0.0 (AItewan BioMedical Technology Inc., Taipei, Taiwan) is an AI-assisted software-as-a-medical-device (SaMD) product for brain tumor image analysis in radiotherapy-related treatment-planning workflows. Regulatory identifiers include the Taiwan Food and Drug Administration (TFDA) medical device license (MOHW-MD-No. 008460) and the U.S. Food and Drug Administration (FDA) 510(k) clearance (K252190). In this study, the AI assistance primarily provided MRI-based lesion localization and preliminary tumor contours for physician review. For the diagnostic-reporting experiment, the study interface also displayed structured lesion-level information derived from the AI output, including lesion location and size-related measurements, to support physician review and report preparation (
Figure 3). These data were presented as structured information rather than as an automatically finalized radiology report; readers reviewed the MRI images and finalized all report content themselves.
To standardize the experimental environment and minimize interface-related variability, reference-standard annotation, diagnostic reporting, and tumor delineation experiments were conducted using the same reader interface in 3D Slicer (version 5.8.0). This controlled environment was used to review MRI examinations, reference-standard annotations, AI-generated outputs, and physician-generated outputs, and to record task-timing procedures across reference-standard generation, AI-assisted sessions, and physician-only sessions.
2.6. Outcome Measures
2.6.1. Diagnostic Reporting Outcomes
The diagnostic reporting evaluation included reporting efficiency, lesion-level diagnostic performance, and report consistency. Reporting efficiency was assessed using the total report generation time across the 30 cases for each reader. Diagnostic performance was assessed using lesion-level sensitivity and precision against the reference standard. Report consistency was assessed using Recall-Oriented Understudy for Gisting Evaluation-Longest Common Subsequence (ROUGE-L), BERTScore, and Sentence-BERT (SBERT) cosine similarity to quantify structural and semantic similarity between reports generated by different readers. ROUGE-L was used as a lexical and structural overlap metric based on the longest common subsequence between reports [
18]. BERTScore F1 was used to quantify token-level semantic similarity using contextual BERT embeddings [
19]. SBERT cosine similarity was used to compare sentence-level embeddings as a global semantic similarity measure between reports [
20].
2.6.2. Treatment-Planning Outcomes
The treatment-planning evaluation included contouring efficiency, volumetric agreement, lesion-level contouring performance, and segmentation overlap. Contouring efficiency was assessed using per-lesion delineation or contour-editing time across the 30 cases for each reader, with lesion-level times summed within each case for paired case-level workflow–time comparisons. Volumetric agreement was assessed by comparing tumor volumes among AI contours, manual physician contours, and AI-assisted physician contours. Lesion-level contouring performance was assessed using sensitivity and precision, and segmentation overlap was evaluated using the Dice coefficient.
2.7. Statistical Analysis and Computational Tools
Descriptive statistics were used to summarize workflow efficiency, diagnostic performance, contouring performance, and report consistency. Continuous variables were summarized as means with standard deviations and median with interquartile range (IQR), as appropriate.
For workflow–time analyses, case-level paired observations were used. Report generation time was analyzed at the case level. For tumor delineation, the raw time unit was the individual lesion; when a case contained more than one lesion, all lesion-level edit times within that case were summed to obtain case-level contouring time before paired comparison. Paired workflow–time comparisons between AI-assisted and physician-only conditions were performed using two-sided Wilcoxon signed-rank tests. Case-level relative workflow–time reductions were calculated for each paired case and summarized using median and IQR.
Cross-reader report-similarity metrics (ROUGE-L, BERTScore F1, and SBERT cosine similarity) were compared between physician-only and AI-assisted reports using paired two-sided Wilcoxon signed-rank tests across cases. Lesion-level sensitivity, precision, and Dice coefficient were summarized descriptively without formal hypothesis testing or confidence-interval-based inference. These metrics were considered secondary descriptive performance measures. Tumor volume differences across standalone AI contours, manual physician contours, and AI-assisted physician contours were assessed using the Kruskal–Wallis test.
Linear least-squares regression analysis was performed to examine the relationship between tumor volume and manual contouring or editing time, and regression results were reported using fitted equations and coefficients of determination (R2). A p-value < 0.05 was considered statistically significant. No formal adjustment for multiple comparisons was applied; therefore, p-values for non-primary analyses should be interpreted cautiously.
All statistical analyses were performed using Python (version 3.12.12). Data handling and numerical calculations were performed using pandas (version 2.3.3) and NumPy (version 2.3.5). Statistical testing and regression analyses were performed using SciPy (version 1.16.3) and statsmodels (version 0.14.6). Report-similarity metrics were calculated using rouge-score (version 0.1.2) for ROUGE-L, bert-score (version 0.3.13) for BERTScore, and sentence-transformers (version 5.1.2) with scikit-learn (version 1.7.2) for Sentence-BERT cosine similarity. Figures were generated using Matplotlib (version 3.10.8) and seaborn (version 0.13.2).
2.8. Preparation of Schematic Figures and AI-Assisted Manuscript Support
Figure 1 and
Figure 2 were prepared as schematic workflow illustrations during manuscript preparation. Initial figure concepts and layouts were generated using ChatGPT (OpenAI) from author-provided descriptions of the clinical imaging workflow and the retrospective paired evaluation design. The authors reviewed, edited, and finalized the schematic content to ensure consistency with the study design and manuscript text. ChatGPT was not used to generate, modify, analyze, or interpret patient-level data, clinical images, statistical results, or study conclusions.
4. Discussion
In this study, we evaluated the workflow effects of AI assistance across two connected tasks in brain tumor care: MRI reporting and tumor delineation for treatment planning. AI assistance was associated with shorter task-completion times in both settings, higher report-similarity metrics, and numerically higher contour-overlap measures. These findings suggest that the value of AI in this setting should be assessed not only by standalone algorithmic performance but also by its ability to support physician-supervised, standardized clinical workflows. This interpretation is consistent with prior work emphasizing structured reporting in neuro-oncology, LLM-supported radiology reporting, and AI-based segmentation as tools for workflow standardization rather than replacements for physician judgment [
3,
4,
5,
6,
7,
12].
The clinical contribution of AI differed between the diagnostic and treatment-planning settings. In the diagnostic workflow, the main observed benefit was shorter reporting time and higher inter-reader report-similarity metrics rather than increased sensitivity. This pattern is relevant because experienced physicians may already identify most clinically relevant lesions, whereas variability more commonly arises in report structure, description, and emphasis. By providing structured lesion-level information and preliminary annotations, the AI-assisted workflow appeared to reduce inter-reader variation and promote more standardized reporting. However, higher similarity does not necessarily indicate better clinical report quality and may partly reflect anchoring to AI-derived structured information. This finding is aligned with the rationale of BT-RADS and structured reporting frameworks, which aim to improve clarity, completeness, and actionability of brain tumor reports [
3,
4]. Broader structured-reporting evidence similarly suggests that reporting support may add value by improving completeness, consistency, and usability rather than by replacing expert interpretation [
21,
22].
In the treatment-planning task, AI assistance reduced contouring burden while producing numerically higher overlap metrics relative to the reference standard. Manual delineation is labor-intensive and depends on lesion complexity, imaging quality, and physician experience; interobserver variability in brain tumor gross tumor volume delineation has been documented even when MRI information is available [
11]. In this study, AI-generated preliminary contours shifted physician work from full manual segmentation toward review-and-edit behavior, which likely explains the observed reduction in annotation time. For Reader 1, the median AI-assisted edit time was 0.00 min, indicating that in many cases the physician accepted the AI-generated contour after review without additional manual modification; these zero-time observations therefore reflect contour confirmation rather than missing measurements. These findings are consistent with recent deep learning applications in MRI-guided radiotherapy and brain tumor segmentation, where AI supports segmentation while preserving physician oversight [
12,
13,
14,
15,
16,
23,
24,
25,
26,
27]. Clinical implementation studies of auto-segmentation also emphasize that physician review, local validation, editing burden, and clinical acceptability should be evaluated alongside geometric metrics such as the Dice coefficient [
28,
29,
30]. Nevertheless, geometric metrics do not fully establish clinical acceptability, and AI contours require physician review, particularly near critical neurovascular structures.
The volume–time regression provides an additional interpretation of manual contouring burden. For manual delineation, each additional cubic centimeter of tumor volume was associated with approximately 15.00 s and 26.28 s longer active edit time for Readers 1 and 2, respectively. The intercepts should not be interpreted as the expected time for a true zero-volume lesion, because zero volume is outside the clinically meaningful range of this task and the intercept is an extrapolated parameter of the fitted linear model. In this analysis, the intercept more appropriately reflects reader-specific baseline active-editing time not explained by tumor volume alone, such as contour initiation, minimum editing steps, slice-by-slice boundary refinement, editing granularity, and lesion-complexity factors not captured by volume. The larger intercept and steeper slope observed for Reader 2 suggest both a higher baseline active-editing burden and a stronger volume-dependent increase in manual contouring time.
The linked reporting-and-contouring workflow may also support more quantitative longitudinal assessment and interdisciplinary communication. In routine clinical practice, treatment response assessment is often constrained by workload and may rely on qualitative review or semi-quantitative measurements such as the longest lesion diameter. AI-assisted contouring can provide volumetric information that may make serial comparison more reproducible. In a longitudinal vestibular schwannoma radiosurgery study, Lee et al. used AI-derived volumetric analysis to track post-treatment tumor changes over serial MRI follow-ups, illustrating how automated volume information can complement conventional response assessment [
31]. In the present study, physician-reviewed AI contours may provide quantitative lesion-volume information for both treatment planning and future follow-up comparison, while lesion-level information generated during reporting may help maintain continuity across diagnostic and treatment teams [
1,
2,
3,
4].
These workflow effects may also be relevant to radiology turnaround time, although turnaround time is influenced by more than the time needed to compose a report. In routine practice, overall turnaround time can be affected by case arrival and worklist position, prioritization or triage, the interval before a physician opens the examination, image interpretation, report generation, finalization, notification of actionable findings, and clinician review of the report [
32,
33]. In this context, AI assistance could affect different components of the pathway: automated triage or notification could help identify brain tumor cases requiring earlier physician attention, whereas AI-derived structured lesion-level information may support report preparation once the case is opened. These possibilities should be examined in future studies that integrate AI-assisted reporting with worklist prioritization and notification systems.
Taken together, prior evidence indicates that AI assistance in radiology and radiotherapy is most informative when evaluated as part of a physician-supervised workflow rather than as an isolated algorithmic output. Current guidance for clinical AI evaluation emphasizes human–AI interaction, the intended setting of use, workflow integration, error analysis, external validation, and outcome-oriented assessment [
34,
35,
36,
37,
38,
39,
40]. The present work extends this perspective by measuring reporting time, report consistency, contouring time, lesion-level detection performance, volumetric agreement, and Dice overlap in the same paired evaluation framework. Prior reader-assistance, reporting, and auto-contouring studies have reported reductions in reading, reporting, segmentation, or contouring workload, but the magnitude of benefit varies across task definitions, measurement methods, and clinical settings [
41,
42,
43]. By evaluating two connected clinical nodes, this study highlights the potential role of AI assistance in linking diagnostic report preparation with treatment-planning tumor delineation.
This study has several limitations. First, the cohort was small, retrospective, and limited to vestibular schwannoma, meningioma, and brain metastasis; therefore, the findings should not be generalized to gliomas or other tumor types without further validation. The sample size was planned for a paired workflow–time evaluation with power and precision considerations, and the cohort was balanced across the three target tumor categories; nevertheless, the study was not intended to provide comprehensive validation of diagnostic accuracy, segmentation performance, or downstream clinical outcomes. Second, the workflow order was fixed, with AI-assisted sessions performed before physician-only sessions; recall, learning, fatigue, and workflow-adaptation effects cannot be excluded. Third, only two readers performed each task, so reader-level results may reflect individual practice patterns. The inclusion of one more experienced and one less experienced reader in each task provides preliminary observations across experience levels, but the study was not powered to formally test reader–experience interactions. Fourth, reporting and contouring times were controlled task-completion times rather than full clinical turnaround or treatment-planning workflow times. The diagnostic reporting experiment measured report preparation for a single MRI examination under a controlled study setting. In routine clinical practice, physicians often review prior imaging examinations, compare interval changes, consult clinical history, and incorporate information from the medical record before finalizing a brain tumor MRI report. Therefore, the absolute reporting times observed here should not be interpreted as the full clinical reporting time required in routine practice. This study focused on proximal workflow outcomes and did not directly measure clinical turnaround time for brain tumor cases, including time from examination completion to report initiation, final report availability, clinician notification, treatment-planning initiation, or treatment delivery. Fifth, sensitivity, precision, Dice coefficient, and volume agreement were secondary descriptive performance measures without formal confidence intervals or multiplicity adjustment. Sixth, the union-based reference contours, report-similarity metrics, and geometric contour metrics do not fully establish clinical report quality or contour acceptability, and AI outputs may anchor reader decisions. Finally, several authors are affiliated with the developer of the evaluated AI software; independent external validation and prospective multicenter testing are needed.