Next Article in Journal
Signal-Derived Feature Analysis for Cuffless Blood Pressure Estimation: Comparing Machine Learning and Deep Learning on ICU Physiological Waveforms
Next Article in Special Issue
Beyond Vital Signs: A Machine Learning Model Using Comprehensive Triage-Time Data to Detect Undertriage in Emergency Department Patients
Previous Article in Journal
The MiniMarket80 Dataset for Evaluation of Unique Item Segmentation in Point Clouds
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Technical Note

CTV Delineation in the Era of Artificial Intelligence: A Multicenter Assessment of a 3D U-Net Model as Predictive Peer Review for Hypofractionated Prostate Cancer Treatment

by
Luca Capone
1,*,
Giorgio H. Raza
1,*,
Chiara D’Ambrosio
1,
Francesco Tortorelli
1,
Francesco Aquilanti
2 and
Pier Carlo Gentile
1,3
1
UPMC Hillman Cancer Center San Pietro, 00189 Rome, Italy
2
Radiotherapy Unit, Marrelli Hospital, 88900 Crotone, Italy
3
UPMC Hillman Cancer Center Villa Maria, 83036 Mirabella Eclano, Italy
*
Authors to whom correspondence should be addressed.
Submission received: 19 January 2026 / Revised: 28 January 2026 / Accepted: 5 March 2026 / Published: 6 March 2026
(This article belongs to the Special Issue Applications of Artificial Intelligence in Medicine)

Abstract

Purpose: The aim is to evaluate the effectiveness of artificial intelligence (AI)-based automatic segmentation as a predictive tool for clinical peer review in prostate cancer patients treated with hypofractionated radiotherapy. Methodology: A retrospective analysis was conducted on 62 patients treated across three Italian centers between 2020 and 2025. CT images were segmented using software based on 3D U-net models. Three workflows were compared: manual segmentation (C man), automatic segmentation (C AI), and AI-based segmentation adjusted by clinicians (C adj). Quantitative metrics used for comparison included the Dice Similarity Coefficient (DSC) and Hausdorff Distance (HDmax). Statistical analysis involved Welch’s t-test and Cohen’s d for effect size. Results: The results showed a significant improvement in agreement between C AI and C adj compared to C man. Median DSC for CTV increased from 0.80 (C man) to 0.92 (C adj), while HDmax decreased from 12.33 mm to 9.22 mm. Similar improvements were observed for the bladder and anorectum. All differences were statistically significant (p < 0.0001), with large effect sizes (Cohen’s d > 0.8). Discussion: AI use demonstrated a reduction in interobserver variability and segmentation time, enhancing workflow standardization. The C adj workflow, where the physician acts as a reviewer of AI-generated contours, proved effective and potentially integrable into clinical peer review. The predictive peer review refers to a preliminary support step in the clinical review process rather than a substitute for medical decision-making.

1. Introduction

1.1. Clinical and Workflow Challenges

The use of artificial intelligence (AI) in many aspects of the modern radiotherapy workflow is increasing as commercial solutions become available [1]. AI-based software is used for decision making, segmentation, treatment planning, patient QA, adaptive treatment, model development, and other processes [2,3]. In the radiotherapy workflow, target and OAR contouring is strongly related to the clinical choices and the treatment outcome, and an inaccurate segmentation may affect the plan and subsequently the treatment quality [4]. Moreover, target and OAR segmentation is time consuming, particularly for adaptive radiotherapy in which physicians only have a little time to evaluate and revise the clinical volumes [5]. Furthermore, clinical volumes are related to the imaging interpretation of different physicians, which depends on the type of the images (CT, MR, PT) and the physicians’ experience [6,7]. This produces a kind of variability in the clinical volume segmentation observed by many papers [8]. Another aspect is related to the check of the treatment plan and, in general, of the entire treatment. Radiotherapy is a complex process that requires a wide spectrum of hardware, software, and human procedure tools to ensure safe clinical treatment. Part of the Quality Assurance procedure is a review of the patient data by the involved professional figures, as required by many national and international laws and guidelines. The review is time consuming and requires care mainly from physicians [9]. Furthermore, peer review, while important for patient safety, is seen as boring because the physicians must check patient data repeatedly, with most of the data being correct, and detect most of the anomaly data but not all [10].

1.2. Previous AI–Human Hybrid Approaches

Qiongge et al. proposed a software tool based on the AI which can support physicians in detecting erroneous prescriptions, and the management software and the TPS have automatic procedures to prevent the most common errors [11]. However, contouring anomalies are often hard to detect because the volumes are related to the patient’s clinical characteristics and depend on the clinical intent, so physician reviewers must check a great amount of data, from diagnostic data to volumes (target, OARs), for each patient [12].
Many methods have been applied to automatic segmentation, such as Atlas-based, statistical models, and AI. In this last decade, the quality of auto-segmentation has enhanced with the AI approach, and auto-contouring software for OARs are widely used in radiotherapy centers [13,14,15].
Unfortunately, AI-based segmentation software is still under investigation due to their racial/demographic bias between the Asian and European population for several organs [16], but despite this lack of training, they can have a significant time reduction compared to manual contouring when used in clinical practice [17]. The deep learning approach can be used for the automatic segmentation of organs at risks (OARs) on magnetic resonance (RM) images with a high level of conformality and accuracy [18,19,20]. In the case of prostate cancer patients, several studies have demonstrated a high level of agreement in normal tissue segmentations, variable agreement on CTV (prostate and seminal vesicles), and large variability on GTV (GTV with microscopic disease). Therefore, interobserver variability should be considered when margins are reduced and AI segmentation is in place, and quantitative (Dice similarity coefficient—DSC, Hausdorff distance—HDmax) metrics can be evaluated [21].
Plans with geometric differences in tumor and normal tissues on physicians’ segmentations and automatic contours do not overdose nearby OARs [22]. Nourzadeh et al. described significantly less accurate plans created from AI-based segmentation of the prostate on CT scans when compared to plans from manually delineated targets [23], while Sritharan K. et al. reported consistent data for the automatic contouring of CTV on MR Linacs images [24].

1.3. Positioning of the Present Study

The main challenge in involving AI-based auto segmentation in clinical practice is its validation, specifically in determining precisely what the “gold standard” or ground truth segmentation is [25]. The contours were created by several physicians with the assumption of being sufficiently accurate for treatment planning. However, while small inconsistencies, especially outside of the high-dose region, do not affect the plan quality, they can considerably decrease quantitative metrics such as DSC [26,27]. Quality standards in radiotherapy underline the relevance of a solid peer review for all patients, particularly when AI solutions are available, and some procedures are in batch mode [28]. Following this, our study wished to investigate the use of AI in playing a critical role in predictive clinical peer review in prostate cancer patients treated with hypofractionated radiotherapy.

2. Materials and Methods

2.1. Autocontouring and Software Analysis

ART-Plan v.2.3.1 (TheraPanacea—8–10 Avenue Ledru-Rollin, 75012, Paris, France) was used for segmentation of the CTV and OARs. ART-Plan is a Class IIb medical device that uses deep learning methods based on 3D U-net models and rule-based post-processing techniques for automatic segmentation [29]. Images from CT simulation are exported to the ART-Plan cloud server, and in a few minutes, contours are generated and imported to the Treatment Planning System (Eclipse v15.6—Varian Medical System, 3100 Hansen Way, Palo Alto, 94304, CA, USA; MiM v7.2.8 or Monaco v.6.2.2.0—Elekta, Hagaplan 4, 113 68, Stockholm, Sweden).

2.2. Contouring Workflow

For volume delineations, the international guidelines and scientific associations recommendations were followed [30,31,32,33,34]. In our workflow, prostate gland and seminal vesicles (defined as CTV), bowel bag, anorectum, bladder, femoral heads were segmented but only CTV, bladder, and anorectum were considered in our analysis.
In our study, we describe three different segmentation workflows:
  • C man (manual segmentation): Physicians delineate the Clinical Target Volume (CTV) and all OARs using Eclipse or Monaco tools on CT sims and MRI and/or PET images when available. Physicians can also use tools like interpolation and smoothing.
  • C AI (automatic segmentation): Contours are generated immediately before image acquisition based on CT sims and are delineated using the ART-Plan 3D U-net model.
  • C adj (automatic segmentation adjusted by clinicians): Physicians use ART-Plan segmentation (C AI) as the basis of contouring, and check and eventually modify the contours by Eclipse or Monaco and its tools.

2.3. Patient Dataset

A total of 62 patients affected by prostate cancer, treated with the hypofractionated protocol, between January 2020 and July 2025 at UPMC Hillman Cancer Center San Pietro in Rome (TrueBeam sTx, Varian Medical System, 3100 Hansen Way, Palo Alto, 94304, CA, USA), UPMC Hillman Cancer Center Villa Maria in Mirabella Eclano (TrueBeam sTx, Varian Medical System, 3100 Hansen Way, Palo Alto, 94304, CA, USA), and Marrelli hospital in Crotone (Versa HDmax, Elekta, Hagaplan 4, 113 68, Stockholm, Sweden) were retrospectively enrolled. No exclusion criteria were considered, so among the patients, there were two patients with hip implants. The patient preparation before simulation and treatments required a full bladder and empty rectum. All patients, with only knee-fix, were simulated with a slice thickness of 1.25 mm, matrix 512 × 512, and interactive filters. Segmentation was based on CT simulation and diagnostic images for C man and C adj, and only CT simulation for C AI. Treatment plans were computed using two co-planar VMAT arcs. In moderate hypofractionated prostate cancer radiotherapy (62 Gy/20 fractions), prostate and seminal vesicles are the CTV. PTV was created by adding a margin of 7 mm (5mm toward rectum). The OARs were the rectum, bladder, bowel bag, and femoral heads.
All patients were divided into two groups (Figure 1):
  • C AI vs. C man: 42 patients treated between January 2020 and December 2024 with volumes delineated by manual contouring (C man) compared to AI (C AI).
  • C AI vs. C adj: 20 patients treated between January 2025 and July 2025 with volumes of interest delineated by ART-Plan and reviewed by physicians (C adj). In our procedure, the physician uses automatic segmentation (C AI) like a reference or a “digital expert” that suggests the reasonable structures in order to check, review, and adapt the structures if needed. At the end of the process, the C adj volumes were compared to our ground truth (C AI).

2.4. Quantitative Metrics Tools

Comparison among contours of different workflows (C AI; C man; C adj) were computed by two quantitative metrics that based on the distance between the surfaces and the ratio of the amounts of overlapped and not overlapped contours. Volumes were imported into Velocity v4.0 (Varian Medical System) for geometric analysis so DSC [35] and HDmax [36] ere performed.
DSC provides an evaluation of the overlapping volume between two contours. It ranges from 0, in the case of not overlapping, to 1 in the case of totally overlapping.
D S C C 1 , C 2 = 2 v o l C 1 C 2 v o l C 1 + v o l C 2
HDmax is a measure between contour surfaces. The operator computes the distance to the nearest point of the other contour in both directions.
H D C 1 , C 2 = m a x h C 1 , C 2 , h C 2 , C 1
where h is the Euclidian distance between the point of the two contours.
h C 1 , C 2 = max a C 1   min b C 2 a b
HDmax compares the shapes of the contours with 0 value for a perfect superposition.

2.5. Statistical Analysis

To compare the performance of AI-generated contours, we performed Welch’s t-test for independent samples and calculated Cohen’s d as an effect size measure.
Welch’s t-test was chosen because it does not assume equal variances between groups and is robust for unequal sample sizes (here, n = 42 for C AI vs. C man and n = 20 for C AI vs. C adj). The test was one-sided, reflecting the directional hypotheses:
  • For DSC:
H0: μCadj ≤ μCman H0: μCadj ≤ μCman vs. H1: μCadj > μCman H1: μCadj > μCman
  • For HDmax:
H0: μCadj ≥ μCman H0: μCadj ≥ μCman vs. H1: μCadj < μCman H1: μCadj < μCman
The Welch t-statistic is computed as:
t = X ¯ 1 X ¯ 2 s 1 2 n 1 + s 2 2 n 2
where:
X ¯ 1 , X ¯ 2 = s a m p l e   m e a n s
s 1 2 ,   s 2 2 = s a m p l e   v a r i a n c e s
n 1 ,   n 2 = s a m p l e   s i z e s
Degrees of freedom are approximated using the Welch–Satterthwaite equation:
d f = s 1 2 n 1 + s 2 2 n 2 2 s 1 2 n 1 2 n 1 1 + s 2 2 n 2 2 n 2 1
The resulting p-value indicates whether the observed difference is statistically significant in the hypothesized direction.
To quantify the magnitude of the difference, we computed Cohen’s d for independent samples:
d = X ¯ C a d j X ¯ C m a n s p
where Sp is the pooled standard deviation:
s p = n 1 1 s 1 2 + n 2 1 s 2 2 n 1 + n 2 2
Interpretation thresholds for |d|:
  • 0.2 = small effect;
  • 0.5 = medium effect;
  • 0.8 or higher = large effect.
If C adj > C man, the d value will be positive; if C adj < C man, the d value will be negative.

2.6. Peer Review Workflow

Peer review in radiotherapy represents a structured process of clinical and technical evaluation conducted among professionals with comparable expertise. Its primary aim is to ensure the accuracy, safety, and appropriateness of radiotherapy treatments. Recognized as a cornerstone of quality assurance in radiation oncology, peer review typically involves an independent assessment of the treatment plan by one or more experienced colleagues prior to the initiation of therapy. A ‘peer review issue’ was defined as encompassing gross contouring inaccuracies, deviations from guideline-based standards, and discrepancies between imaging and delineated volumes.
This process encompasses the verification of several critical components including the correctness of the prescription, the delineation of target volumes, the adequacy of dose distribution, and the identification of potential technical or clinical discrepancies. By systematically reviewing these elements, peer review contributes to minimizing the risk of errors and enhancing the overall standard of care delivered to patients [28].

2.7. Ethical Approval

No interventions, modifications to treatment, or additional procedures were performed beyond standard clinical practice. All data were managed in accordance with applicable privacy regulations and institutional policies to ensure patient confidentiality. As a result, the study involved no additional risk to participants and qualified as a non-interventional observational study. As this analysis was a retrospective non-interventional observational study conducted exclusively on data generated during routine care, ethics committee approval was not required under the applicable Italian regulatory framework including the Ministry of Health DM 30 November 2021, the AIFA Guidelines for observational studies (425/2024), and the EU definition of non-interventional research (Reg. 536/2014). According to our institutional policy, this type of research does not require specific approval from the Institutional Review Board (IRB). The study was conducted following the standard internal procedure for non-interventional research. ART-Plan is a commercially available, pre-trained model, and no fine-tuning was performed using institutional data.

3. Results

Table 1 shows the geometric accuracy of AI-based segmentation compared to manual contours (C man) and adjusted contours (C adj) for three anatomical structures: clinical target volume (CTV), anorectum, and bladder. The focus was not on absolute accuracy, but on the consistency of the workflow and on reducing variability through the use of a common reference (AI). Two metrics were evaluated: Dice similarity coefficient (DSC), which measures the volumetric overlap (range 0–1, higher is better), and Hausdorff distance (HDmax, in mm), which assesses the boundary deviation (lower is better). For each structure and metric, the table reports the median, minimum, maximum, mean, and standard deviation (SD) values.

3.1. C AI vs. C Man

As shown in Table 1, median DSC values for the CTV, bladder, and anorectum were respectively 0.80 (0.49–0.90), 0.94 (0.78–0.97), and 0.86 (0.58–0.91). Median HDmax for the CTV, bladder, and anorectum were respectively 12.33 (8.07–26.67), 7.50 (3.52–18.45); and 17.50 (5.00–30.00).

3.2. C AI vs. C Adj

Median DSC values for the CTV, bladder, and anorectum were respectively 0.92 (0.82–0.97), 0.99 (0.97–1.00), and 0.99 (0.86–0.99) (Figure 1). Median HDmax for the CTV, bladder, and anorectum were respectively 9.22 (5.50–19.13), 3.75 (0.97–6.69), and 1.42 (0.94–18.00).
All comparisons were statistically significant (Welch p < 0.0001).
Effect sizes were large or very large for DSC and HDmax. DSC’s Cohen’s d range was +1.34, +2.39, and +1.96, and HDmax’s Cohen’s d range was −0.91, −2.15, and −1.43, respectively for the CTV, anorectum, and bladder (Table 1).
Figure 2 shows the relationship between the dispersion of DSC and HDmax in the C AI vs. C adj group (green) compared to the C AI vs. C man group (pink) respectively of the CTV, anorectum, and bladder.

4. Discussion

The main challenges in implementing AI-based auto segmentation in clinical practice is its validation and the determination of a reference contour, which in other papers are the “gold standard” or ground truth segmentation. In the literature, the ground truth volumes are based on physician manual segmentation without a standardized procedure, usually based on a commission of physicians with the assumption to be sufficiently accurate for treatment planning. In our study, the reference segmentation was the automatic one (C AI).
Automatic segmentation based on AI allows for the fast and efficient target delineation of the pelvic organs for the prostate moderate hypofractionation treatment. The segmentation is well-described in other papers [37], and is validated for OARs [38].
Our results, for the comparison between the C man vs. C AI and C adj vs. C AI showed that the median DSC improved from 0.80 to 0.92 for CTV and from 0.86 to 0.99 for the anorectum, while the median HDmax decreased from 12.33 mm to 9.52 mm for CTV and from 17.50 mm to 1.42 mm for the anorectum. Similar trends were observed for the bladder, indicating that AI-generated contours require minimal adjustments for organs at risk. From this point of view, the C adj workflow showed lower physician contouring variability with respect to the C man workflow, so this could be used to reduce the inter- and intra-observer variations and increase safety and standardization. The observed differences may therefore be influenced by the distinct temporal windows rather than exclusively by the workflow itself.
An interesting aspect is based on the contouring guidelines [34,35,36,37,38] The contouring procedures for the prostate, seminal vesicles, and femoral heads are well-defined and are the same in all considered guidelines; however, bladder, rectum, or bowel contouring procedures depend on the chosen guideline. ART-Plan can delineate the entire bowel bag, the bowel loops, the anorectum (from the sphincter to the initial sigma) with or without all contents in the lumen, the wall, and the full bladder. Thus ART-Plan can be configured appropriately to the guidelines used. However, some guidelines may require contouring the limits of the rectum different to those automatically contoured by ART-Plan, such as above, to include the sigma, or 10 or 1 cm from the PTV, and below to the anal verge, or to ischial tuberosities (or 2 cm below them), or 1 cm below the PTV.
Another aspect is related to the quality standards in radiotherapy, which suggests the use of peer review at various stages of the workflow for accuracy and safety reasons. Segmentation peer review requires a second physician (reviewer) to evaluate the clinical volumes. However, this activity is time consuming and complex because the reviewer must consider the patient clinical history and the same patient characteristics used by the primary physician. The study of Li et al. [10] showed that compliance regarding the contour increased in the presence of three or more reviewers. This situation is difficult to replicate in pre-treatment contexts in modern radiotherapy departments. Having a system that is able to follow the same “rules” ensures uniform contouring across different patients, allowing for a visual guideline to be represented on the simulation CT, just a few minutes after acquisition, thanks to batch AI contouring. However, the problem of boundary validation and continuous monitoring remains for it to be stable and reproducible even after software version updates or maintenance. Therefore, DSC and HDmax distance are valid tools for the continuous quality control of the deep learning system.
For HDmax, negative values of Cohen’s d in the CTV, anorectum, and bladder indicated improvement (lower distances for C adj), the same for DSC with positive values of Cohen’s d. These results confirm that the AI contours were substantially closer to the final adjusted contours than to the original manual contours, both in terms of overlap (DSC) and boundary accuracy (HDmax).
Our study highlights that in the case of using CT slices for defining the CTV, there was compliance in approving AI segmentation after peer review, with a median DSC of 0.84 in clinical cases in which the physician directly approved the contour without modifications. In fact, in these cases, the physician, even having PET and MRI available, tended to trust the definition of the prostate and seminal vesicles, without making changes. As shown in Figure 2, our procedure allows us to standardize the delineation of the CTV, anorectum, and bladder in three different radiotherapy departments using ART-Plan as the new standard of target delineation peer review. Results show that AI contours exhibited substantially higher agreement with the adjusted contours than with manual contours. In this phase, the initial contours could also be adjusted to MRI or PET images as needed, but their impact on the final target was not part of this study. This perspective aligns with existing hybrid paradigms in which AI serves as an assistive tool rather than a replacement for human expertise, and explicitly situating our findings within this framework helps clarify the novelty of our approach while avoiding any implication that the observed improvements reflect an intrinsic superiority of AI over clinical judgment.

5. Study Limitations

One of our imitations was the small sample size in which two patients with hip implants had lower DSC and higher HDmax due to the artifacts on CT images. In our study, the reference segmentation was the automatic one (C AI) and was used as the “gold standard” or ground truth segmentation as used in other papers. This means that we trust in a “black box” that is able to generate delineations based on models and algorithms but are still not able to merge this with clinical information such as anatomy variances or personalized conditions. Another limit of this kind of study is that there are no standardized values in the literature for DSC and HDmax but many papers that used several ground truth volumes with different types of deep learning and AI algorithms (commercial or open source) based on different images (CT, MRI, PET). It is important to underline that HDmax is highly sensitive to outliers and contour spurs. The 95th percentile HDmax or average surface distance is usually more stable for clinical interpretation, so follow-up studies will implement different metrics. The absence of contour-related metrics (such as surface DSC) also limits the depth of the analysis; however, their inclusion was not available in the software used. From this point of view, we should define HDmax or DSC limit values beyond those so that during peer review, the clinician can check the patient volumes with care. In any event, we did not evaluate these values but only suggest a workflow to introduce AI as a predictive tool for peer review even though no causal conclusions can be drawn regarding the workflow improvement. Our dataset reflects the natural evolution of the workflows adopted across the three centers over the years; however, we acknowledge that this introduces a potential temporal bias, and therefore the two cohorts are not directly comparable. Our study is exploratory in nature and does not allow for causal inferences

6. Conclusions

The study is intended as a testing methodology to assess the impact of integrating an artificial intelligence-based auto-contouring software into the radiotherapy contouring workflow. In this paper, we demonstrate how to use A- generated contours as reference volumes and how these segmentations could be a reference for the physician during the peer review procedure.

Author Contributions

Conceptualization, L.C. and G.H.R.; Data curation, L.C. and G.H.R.; Formal analysis, G.H.R., C.D. and F.T.; Investigation, L.C., G.H.R. and C.D.; Methodology, L.C. and G.H.R.; Project administration, P.C.G.; Resources, G.H.R. and F.A.; Software, L.C.; Supervision, P.C.G.; Validation, G.H.R., C.D. and F.T.; Visualization, L.C. and G.H.R.; Writing—original draft, L.C. and G.H.R.; Writing—review and editing, L.C. and G.H.R. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

This retrospective, non-interventional observational study only used routinely collected clinical data and, according to the applicable Italian regulatory framework (Ministry of Health Decree 30 November 2021; AIFA Determination 425/2024; EU Regulation 536/2014), did not require ethics committee approval.

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

The datasets generated and analyzed in this study are not publicly available due to privacy and ethical restrictions as they consist of retrospective clinical data that cannot be shared or deposited in public repositories.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Conroy, L.; Winter, J.; Khalifa, A.; Tsui, G.; Berlin, A.; Purdie, T.G. Artificial intelligence for radiation treatment planning: Bridging gaps from retrospective promise to clinical reality. Clin. Oncol. 2025, 37, 103630. [Google Scholar] [CrossRef]
  2. Thompson, R.F.; Valdes, G.; Fuller, C.D.; Carpenter, C.M.; Morin, O.; Aneja, S.; Lindsay, W.D.; Aerts, H.J.; Agrimson, B.; Deville, C.; et al. Artificial intelligence in radiation oncology: A specialty-wide disruptive transformation? Radiother. Oncol. 2018, 129, 421–426. [Google Scholar] [CrossRef]
  3. Boon, I.S.; Au Yong, P.T.; Boon, C.S. Application of artificial intelligence (AI) in radiotherapy workflow: Paradigm shift in precision radiotherapy using machine learning. Br. J. Radiol. 2019, 92, 20190716. [Google Scholar] [CrossRef]
  4. Wong, J.; Huang, V.; Wells, D.; Giambattista, J.; Giambattista, J.; Kolbeck, C.; Otto, K.; Saibishkumar, E.P.; Alexander, A. Implementation of deep learning-based auto-segmentation for radiotherapy planning structures: A workflow study at two cancer centers. Radiat. Oncol. 2021, 16, 101. [Google Scholar] [CrossRef]
  5. Kiser, K.J.; Fuller, C.D.; Reed, V.K. Artificial intelligence in radiation oncology treatment planning: A brief overview. J. Med. Artif. Intell. 2019, 2, 9. [Google Scholar] [CrossRef]
  6. Williams, L.H.; Drew, T. What do we know about volumetric medical image interpretation? A review of the basic science and medical image perception literatures. Cogn. Res. Princ. Implic. 2019, 4, 21. [Google Scholar] [CrossRef]
  7. Atsina, K.B.; Parker, L.; Rao, V.M.; Levin, D.C. Advanced Imaging Interpretation by Radiologists and Nonradiologist Physicians: A Training Issue. Am. J. Roentgenol. 2020, 214, W55–W61. [Google Scholar] [CrossRef] [PubMed]
  8. Liang, J.; Ma, Z.; Li, H.; Gao, F.; Yang, N.; Li, D.; Li, M.; Geng, D. Interobserver Agreement in Automatic Segmentation Annotation of prostate Magnetic resonance imaging. Bioengineering 2023, 10, 1340. [Google Scholar] [CrossRef] [PubMed]
  9. Wang, D.; Geng, H.; Gondi, V.; Lee, N.Y.; Tsien, C.I.; Xia, P.; Chenevert, T.L.; Michalski, J.M.; Gilbert, M.R.; Le, Q.-T.; et al. Radiotherapy plan quality assurance in NRG oncology trials for brain and head/neck cancers: An AI-enhanced knowledge-based approach. Cancers 2024, 16, 2007. [Google Scholar] [CrossRef]
  10. Lewis, P.J.; Court, L.E.; Lievens, Y.; Aggarwal, A. Structure and Processes of Existing Practice in Radiotherapy Peer Review: A Systematic Review of the Literature. Clin. Oncol. 2021, 33, 248–260. [Google Scholar] [CrossRef]
  11. Li, Q.; Wright, J.; Hales, R.; Voong, R.; McNutt, T. A digital physician peer to automatically detect erroneous prescriptions in radiotherapy. Digit. Med. 2022, 5, 158. [Google Scholar] [CrossRef]
  12. Mackay, K.; Banfill, K.; Bernstein, D.; Daniel, J.; Diez, P.; Gwynne, S.; Hoole, A.; Jena, R.; Marchant, T.; Nix, M.; et al. Royal College of Radiologists guidance statements on the use of auto-contouring in radiotherapy. Clin. Oncol. 2025, 50, 104004. [Google Scholar] [CrossRef]
  13. Brouwer, C.L.; Steenbakkers, R.J.; Heuvel, E.v.D.; Duppen, J.C.; Navran, A.; Bijl, H.P.; Chouvalova, O.; Burlage, F.R.; Meertens, H.; Langendijk, J.A.; et al. 3D variation in delineation of head and neck organs at risk. Radiother. Oncol. 2012, 7, 32. [Google Scholar] [CrossRef]
  14. Cardenas, C.E.; Yang, J.; Anderson, B.M.; Court, L.E.; Brock, K.B. Advances in auto-segmentation. Semin. Radiat. Oncol. 2019, 29, 185–197. [Google Scholar] [CrossRef]
  15. Lustberg, T.; van Soest, J.; Gooding, M.; Peressutti, D.; Aljabar, P.; van der Stoep, J.; van Elmpt, W.; Dekker, A. Clinical evaluation of atlas and deep learning based automatic contouring for lung cancer. Radiother. Oncol. 2018, 126, 312–317. [Google Scholar] [CrossRef] [PubMed]
  16. McQuinlan, Y.; Brouwer, C.L.; Lin, Z.; Gan, Y.; Kim, J.S.; van Elmpt, W.; Gooding, M.J. An investigation into the risk of population bias in deep learning autocontouring. Radiother. Oncol. 2023, 186, 109747. [Google Scholar] [CrossRef]
  17. Brouwer, C.L.; Dinkla, A.M.; Vandewinckele, L.; Crijns, W.; Claessens, M.; Verellen, D.; van Elmpt, W. Machine learning applications in radiation oncology: Current use and needs to support clinical implementation. Phys. Imaging Radiat. Oncol. 2020, 16, 144–148. [Google Scholar] [CrossRef] [PubMed]
  18. Wu, C.; Montagne, S.; Hamzaoui, D.; Ayache, N.; Delingette, H.; Renard-Penna, R. Automatic segmentation of prostate zonal anatomy on MRI: A systematic review of the literature. Insights Into Imaging 2022, 13, 202. [Google Scholar] [CrossRef] [PubMed]
  19. Savenije, M.H.F.; Maspero, M.; Sikkes, G.G.; van der Voort van Zyp, J.R.N.; Kotte, A.N.T.J.; Bol, G.H.; van den Berg, C.A.T. Clinical implementation of MRI-based organs-at-risk auto-segmentation with convolutional networks for prostate radiotherapy. Radiat. Oncol. 2020, 15, 104. [Google Scholar] [CrossRef]
  20. Molière, S.; Hamzaoui, D.; Granger, B.; Montagne, S.; Allera, A.; Ezziane, M.; Luzurier, A.; Quint, R.; Kalai, M.; Ayache, N.; et al. Reference standard for the evaluation of automatic segmentation algorithms: Quantification of inter observer variability of manual delineation of prostate contour on MRI. Diagn. Interv. Imaging 2023, 105, 65–73. [Google Scholar] [CrossRef]
  21. Arjmandi, N.; Mosleh-Shirazi, M.A.; Mohebbi, S.; Nasseri, S.; Mehdizadeh, A.; Pishevar, Z.; Hosseini, S.; Tehranizadeh, A.A.; Momennezhad, M. Evaluating the dosimetric impact of deep-learning-based auto-segmentation in prostate cancer radiotherapy: Insights into real-world clinical implementation and inter-observer variability. J. Appl. Clin. Med. Phys. 2025, 26, e14569. [Google Scholar] [CrossRef]
  22. Hoque, S.M.H.; Pirrone, G.; Matrone, F.; Donofrio, A.; Fanetti, G.; Caroli, A.; Rista, R.S.; Bortolus, R.; Avanzo, M.; Drigo, A.; et al. Clinical use of a commercial artificial intelligence-based software for autocontouring in radiation therapy: Geometric performance and dosimetric impact. Cancers 2023, 15, 5735. [Google Scholar] [CrossRef]
  23. Nourzadeh, H.; Watkins, T.; Ahmed, M.; Cheukkai, H.; Shlesinger, D.; Siebers, J.V. Clinical adequacy assessment of autocontours for prostate IMRT with meaningful endpoints. Med. Phys. 2017, 44, 1525–1537. [Google Scholar] [CrossRef]
  24. Sritharan, K.; Dunlop, A.; Mohajer, J.; Adair-Smith, G.; Barnes, H.; Brand, D.; Greenlay, E.; Hijab, A.; Oelfke, U.; Pathmanathan, A.; et al. Dosimetric comparison of automatically propagated prostate contours with manually drawn contours in MRI-guided radiotherapy: A step towards a contouring free workflow? Clin. Transl. Radiat. Oncol. 2022, 37, 25–32. [Google Scholar] [CrossRef] [PubMed]
  25. De Biase, A.; Sijtsema, N.M.; Janssen, T.; Hurkmans, C.; Brouwer, C.; van Ooijen, P. Clinical adoption of deep learning target auto-segmentation for radiation therapy: Challenges, clinical risks, and mitigation strategies. BJR Artif. Intell. 2024, 1, ubae015. [Google Scholar] [CrossRef]
  26. Zhang, L.; Tanno, R.; Xu, M.C.; Jin, C.; Jacob, J.; Cicarrelli, O.; Barkhof, F.; Alexander, D. Disentangling human error from ground truth in segmentation of medical images. Adv. Neural Inf. Process. Syst. 2020, 33, 15750–15762. [Google Scholar]
  27. Wang, T.; Tam, J.; Chum, T.; Tai, C.; Marshall, D.C.; Buckstein, M.; Liu, J.; Green, S.; Stewart, R.D.; Liu, T.; et al. Evaluation of AI-based auto-contouring tools in radiotherapy: A single-institution study. J. Appl. Clin. Med. Phys. 2025, 26, e14620. [Google Scholar] [CrossRef]
  28. Duggar, W.N.; Bhandari, R.; Yang, C.C.; Vijayakumar, S. Group consensus peer review in radiation oncology: Commitment to quality. Radiat. Oncol. 2018, 13, 55. [Google Scholar] [CrossRef]
  29. TheraPanacea. ART-Plan™ User Manual (Version 2.3.2). TheraPanacea. 2024. Available online: https://www.therapanacea.eu/technical-information/ (accessed on 4 March 2026).
  30. Salembier, C.; Villeirs, G.; De Bari, B.; Hoskin, P.; Pieters, B.R.; Van Vulpen, M.; Khoo, V.; Henry, A.; Bossi, A.; De Meerleer, G. ESTRO ACROP consensus guideline on CT- and MRI-based target volume delineation for primary radiation therapy of localized prostate cancer. Radiother. Oncol. 2018, 127, 49–61. [Google Scholar] [CrossRef] [PubMed]
  31. Hurkmans, C.; Bibault, J.E.; Brock, K.K.; van Elmpt, W.; Feng, M.; Fuller, C.D.; Jereczek-Fossa, B.A.; Korreman, S.; Landry, G.; Madesta, F.; et al. A joint ESTRO and AAPM guideline for development, clinical validation and reporting of artificial intelligence models in radiation therapy. Radiother. Oncol. 2024, 197, 110345. [Google Scholar] [CrossRef]
  32. Bentzen, S.M.; Constine, L.S.; Deasy, J.O.; Eisbruch, A.; Jackson, A.; Marks, L.B.; Ten Haken, R.K.; Yorke, E.D. Quantitative Analyses of Normal Tissue Effects in the Clinic (QUANTEC): An introduction to scientific issues. Int. J. Radiat. Oncol. Biol. Phys. 2010, 76, S3–S9. [Google Scholar] [CrossRef] [PubMed]
  33. Arcangeli, G.; Saracino, B.; Gomellini, S.; Petrongari, M.G.; Arcangeli, S.; Benassi, M. A prospective phase III randomized trial of hypofractionation versus conventional fractionation in patients with high-risk prostate cancer. Int. J. Radiat. Oncol. Biol. Phys. 2010, 78, 11–18. [Google Scholar] [CrossRef] [PubMed]
  34. Benedict, S.H.; Yenice, K.M.; Followill, D.; Galvin, J.M.; Hinson, W.; Kavanagh, B.; Keall, P.; Lovelock, M.; Meeks, S.; Papiez, L.; et al. Stereotactic body radiation therapy: The report of AAPM Task Group 101. Med. Phys. 2010, 37, 4078–4101. [Google Scholar] [CrossRef] [PubMed]
  35. Dice, L.R. Measures of the amount of ecologic association between species. Ecology 1945, 26, 297–302. [Google Scholar] [CrossRef]
  36. Rockafellar, R.T.; Wets, R.J.-B. Variational Analysis; Springer: Berlin/Heidelberg, Germany, 1998. [Google Scholar] [CrossRef]
  37. Jin, R.; Li, D.; Xiang, D.; Zhang, L.; Zhou, H.; Shi, F.; Zhu, W.; Cai, J.; Penguin, T.; Chen, X. Ai-based automatic segmentation of prostate on multi-modality images: A review. arXiv 2024, arXiv:2407.06612. [Google Scholar]
  38. Huang, S.; Wu, J.; Lin, X.; Wang, G.; Song, T.; Chen, L.; Jia, L.; Cao, Q.; Liu, R.; Liu, Y.; et al. Auto-Segmentation and Auto-Planning in Automated Radiotherapy for Prostate Cancer. Bioengineering 2025, 12, 620. [Google Scholar] [CrossRef]
Figure 1. Study design and quantitative metrics.
Figure 1. Study design and quantitative metrics.
Ai 07 00097 g001
Figure 2. Relationship between the DSC and HDmax values in CTV (a), anorectum (b), and bladder (c) segmentation. In pink, the C AI vs. C man group; in green C AI vs. C adj group.
Figure 2. Relationship between the DSC and HDmax values in CTV (a), anorectum (b), and bladder (c) segmentation. In pink, the C AI vs. C man group; in green C AI vs. C adj group.
Ai 07 00097 g002
Table 1. Median, min, max, mean, standard deviation (SD), Welch p value and Cohen’s d of DSC and HDmax in C AI vs. C man and C AI vs. C adj groups.
Table 1. Median, min, max, mean, standard deviation (SD), Welch p value and Cohen’s d of DSC and HDmax in C AI vs. C man and C AI vs. C adj groups.
C AI vs CmanC AI vs C adjStatistical Analysis
MedianMinMaxMeanSDMedianMinMaxMeanSDWelch p ValueCohen’s d
DSCCTV0.800.490.900.770.110.920.820.970.900.05<0.00011.34
ANORECTUM0.860.580.910.840.060.990.860.990.970.04<0.00022.39
BLADDER0.940.780.970.930.030.990.970.9980.990.01<0.00031.96
HDmax (mm)CTV12.338.0726.6713.995.119.525.519.139.783.07<0.0004−0.91
ANORECTUM17.505.0030.0017.846.961.420.9418.004.284.78<0.0005−2.15
BLADDER7.503.5218.458.013.443.750.976.693.601.89<0.0006−1.43
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Capone, L.; Raza, G.H.; D’Ambrosio, C.; Tortorelli, F.; Aquilanti, F.; Gentile, P.C. CTV Delineation in the Era of Artificial Intelligence: A Multicenter Assessment of a 3D U-Net Model as Predictive Peer Review for Hypofractionated Prostate Cancer Treatment. AI 2026, 7, 97. https://doi.org/10.3390/ai7030097

AMA Style

Capone L, Raza GH, D’Ambrosio C, Tortorelli F, Aquilanti F, Gentile PC. CTV Delineation in the Era of Artificial Intelligence: A Multicenter Assessment of a 3D U-Net Model as Predictive Peer Review for Hypofractionated Prostate Cancer Treatment. AI. 2026; 7(3):97. https://doi.org/10.3390/ai7030097

Chicago/Turabian Style

Capone, Luca, Giorgio H. Raza, Chiara D’Ambrosio, Francesco Tortorelli, Francesco Aquilanti, and Pier Carlo Gentile. 2026. "CTV Delineation in the Era of Artificial Intelligence: A Multicenter Assessment of a 3D U-Net Model as Predictive Peer Review for Hypofractionated Prostate Cancer Treatment" AI 7, no. 3: 97. https://doi.org/10.3390/ai7030097

APA Style

Capone, L., Raza, G. H., D’Ambrosio, C., Tortorelli, F., Aquilanti, F., & Gentile, P. C. (2026). CTV Delineation in the Era of Artificial Intelligence: A Multicenter Assessment of a 3D U-Net Model as Predictive Peer Review for Hypofractionated Prostate Cancer Treatment. AI, 7(3), 97. https://doi.org/10.3390/ai7030097

Article Metrics

Back to TopTop