Next Article in Journal
Virtual Reality-Assisted Rehabilitation for Adolescents with Cerebral Palsy: A Systematic Review and Meta-Analysis
Previous Article in Journal
Feasibility and Adoption of Low-Volume Ultrasound-Guided Superior Trunk Block for Shoulder Reduction in the Emergency Department
Previous Article in Special Issue
Building the Foundation for Standardized Care Metrics in Jejunoileal Atresia: A Systematic Review of Reported Baseline Characteristics, Treatment Variables and Outcomes
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

AI-Supported Prediction of Vesicoureteral Reflux in Children Based on Cystoscopic Configuration of the Ureteric Orifice

1
Department of Paediatric and Adolescent Surgery, Medical University of Graz, Auenbruggerplatz 34, 8036 Graz, Austria
2
Division of Paediatric Radiology, Department of Radiology, Medical University of Graz, 8036 Graz, Austria
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
J. Clin. Med. 2026, 15(16), 6246; https://doi.org/10.3390/jcm15166246
Submission received: 20 July 2026 / Revised: 7 August 2026 / Accepted: 10 August 2026 / Published: 12 August 2026

Abstract

Background: Vesicoureteral reflux (VUR) is a common pediatric urological disorder traditionally diagnosed by voiding cystourethrography (VCUG). Although ureteric orifice (UO) morphology during cystoscopy correlates with VUR severity, its assessment remains subjective. To our knowledge, artificial intelligence (AI) has not previously been applied to cystoscopic images for VUR prediction. We aimed to develop and evaluate the first AI model for predicting VUR from pediatric cystoscopic images and videos. Methods: In this retrospective single-center study, cystoscopic videos from children with and without VUR diagnosed by VCUG were analyzed. Individual frames were extracted, anonymized and annotated according to VUR grade. Four YOLOv12 object detection models were trained to classify UOs as no/grade I, low-grade (grades II–III), or high-grade (grades IV–V) VUR. To better understand sources of model failure, additional experiments evaluated binary image classification and class-agnostic UO localization. Finally, frame-level predictions were aggregated across complete cystoscopic sequences using majority voting to assess UO-level performance. Results: Three-class object detection demonstrated limited frame-level performance, with a maximum mAP@50 of 0.37 and mAP@50–95 of 0.17. Binary classification achieved a macro precision of 0.62, recall of 0.58, and F1 score of 0.49. Class-agnostic object detection improved localization performance (mAP@50 0.63). Aggregating predictions across complete video sequences substantially improved diagnostic performance, achieving UO-level accuracies of up to 77%, with leave-one-out cross-validation accuracies ranging from 0.69 to 0.77. Conclusions: This proof-of-concept study demonstrates the feasibility of AI-assisted VUR prediction from pediatric cystoscopic videos. While frame-level performance was limited, sequence-level aggregation markedly improved diagnostic accuracy, highlighting the importance of temporal information for future video-based AI models in pediatric endourology.

1. Introduction

Vesicoureteral reflux (VUR) describes the retrograde flow of urine from the bladder into the upper urinary tract and represents the most common urological malformation in children, affecting approximately 1–2% of the general pediatric population and up to 30–45% of children presenting with a febrile urinary tract infection (UTI) [1]. Left untreated or improperly managed, VUR can result in recurrent pyelonephritis, renal scarring, chronic kidney disease, and ultimately end-stage renal disease in children and young adults [2].
Current diagnostic pathways rely principally on voiding cystourethrography (VCUG), which remains the gold standard for grading VUR according to the International Reflux Study Committee (IRSC) classification (grades I–V) [3]. However, VCUG is invasive, requires urinary catheterization, and exposes patients to ionizing radiation. Moreover, VCUG grading is notoriously operator-dependent: inter-rater agreement among radiologists has been reported at only 59% in well-resourced multi-centre trials [4].
The anatomical appearance of the ureteric orifice (UO) during cystoscopy has long been recognized as a correlate of VUR severity. Based on the embryological Mackie–Stephens ureteral bud theory, ectopically positioned ureteral buds produce orifices that are laterally displaced, tunnel-deficient, and morphologically abnormal [5]. Lyon and colleagues classified UO configurations from normal cone-shaped through stadium, horseshoe, to golf-hole forms, with the latter two correlating with progressive degrees of reflux [6]. A recent clinical series confirmed that UO shape correlates significantly with VUR grade, with horseshoe and golf-hole configurations comprising the majority of severe VUR cases [7]. The dynamic hydrodistention (HD) classification introduced by Kirsch and colleagues provides an additional, more objective endoscopic grading tool with high interobserver concordance [8], though its predictive value for de novo contralateral VUR remains limited [7]. Despite these established morphological correlates, cystoscopic assessment of the UO performed by experienced pediatric urologists remains highly subjective, with significant variability depending on distension state, cystoscope angle, and lighting conditions. These observations suggest that diagnostically relevant information is present within cystoscopic appearance but it is difficult to assess reproducibly. These limitations make cystoscopic assessment an attractive target for AI-assisted image analysis, which may offer greater objectivity and reproducibility than conventional visual interpretation.
Artificial intelligence (AI) and machine learning (ML), particularly convolutional neural networks (CNNs), have demonstrated transformative potential in image-based medical diagnostics. In pediatric urology, AI applications have been expanded to include UTI risk stratification, prediction of VUR spontaneous resolution, VCUG-based grading, and hydronephrosis classification [9,10]. The AI-PEDURO living database of AI applications in pediatric urology identified 17 studies applying AI to VUR or UTI outcomes, with neural networks, tree-based algorithms, and support vector machines being the most commonly employed approaches [9]. Most notably, Li et al. (2024) developed Deep-VCUG, a ResNet-101 ensemble model for automated VCUG grading, achieving AUCs of 0.96 on internal and 0.94 on external multi-institutional testing sets [11], a landmark result demonstrating that AI-assisted VUR grading can match or exceed clinician performance on radiological data.
However, no prior study has applied AI to cystoscopic video data for VUR prediction in children. Translating AI from radiological to endoscopic video data introduces a distinct set of challenges: endoscopic images are susceptible to motion blur, specular reflections, variable lighting, and anatomical obscuration; datasets are smaller and the diagnostic signal is distributed across a temporal sequence rather than contained in a single frame. This study addresses precisely these challenges. We present the development, iterative optimization, and performance evaluation of the first AI model for VUR prediction based on cystoscopic UO morphology, and discuss the methodological lessons learned for future AI development in pediatric endourology.

2. Materials and Methods

2.1. Ethics and Patient Cohort

Following institutional ethical approval (36-262 ex 23/24), a total of 79 pediatric patients undergoing cystoscopy due to VUR and other urological diseases listed below were included in this retrospective study, comprising 51 boys (64.6%) and 28 girls (35.4%). The mean age at cystoscopy was 4.8 ± 5.0 years (range, 0–17.6 years). Patients with duplex collecting systems, obstructive ureteral orifices, secondary vesicoureteral reflux (e.g., associated with posterior urethral valves) or pronounced ureteral ectopia were excluded from the study.
Patients with other diseases including hypospadias (n = 4), urethral injury (n = 3), ureteropelvic junction obstruction (UPJO; n = 5), posterior urethritis (n = 2), and other less frequent indications (n = 11) were included as controls. In addition, five patients with neurogenic bladder dysfunction and negative VCUG findings were also included as controls. The remaining 49 patients underwent VCUG before cystoscopy as part of the diagnostic evaluation for VUR. None of the patients has undergone deflux or urological surgery before the cystoscopies included in the study.
VUR was diagnosed using VCUG in 41 of the 49 patients, all of whom subsequently underwent cystoscopy. Bilateral VUR was present in 23 patients (56.1%). Overall, 117 UOs were included in the analysis.
On the left side, VUR was absent in 25 ureteral orifices, while grade I VUR was present in 5, grade II in 8, grade III in 10, grade IV in 9, and grade V in 3 ureteral orifices. On the right side, VUR was absent in 34 ureteral orifices, while grade I VUR was present in 4, grade II in 4, grade III in 4, grade IV in 9, and grade V in 2 ureteral orifices.

2.2. Image Extraction and Annotation

Individual fully anonymized images were extracted from cystoscopic video recordings and labeled by two experienced pediatric urologists with 5 and 15 years of experience using the Computer Vision Annotation Tool (CVAT), an open-source annotation platform well established in surgical AI development. Annotations classified UOs as representing high-grade VUR (grades IV–V), low-grade VUR (grades II–III) and grade I/no VUR. Therefore, the reference standard was the grade of reflux using the results from the VCUG if available. The control patients who did not receive a VCUG were patients who were asymptomatic and without pathological sonographic results. Table 1 shows the final proportions of UOs by class in our dataset.
We split the dataset into 27,393 images from 91 UO for training and 9211 images from 26 UO for validation. Crucially, all training and validation data splits were conducted strictly at the UO level rather than by random image allocation. Randomly splitting continuous video frames across sets causes severe data leakage, as adjacent frames are near-duplicates.

2.3. Model Training and AI Performance Metrics

As described in detail below, three different evaluation tasks were performed on a workstation with one NVIDIA RTX 3060 GPU with 12GB of RAM.
First, we trained and evaluated models on a three-class object detection task, the most natural fit for the nature of the dataset, locating no, low-grade, or high-grade reflux on a per-image basis. Four models of the YOLOv12 family [12] with different sizes and parameter counts (YOLOv12n, YOLOv12s, YOLOv12m, YOLOv12l) were trained for 10 epochs each. The configuration of the models is shown in Table 2.
Otherwise, standard hyperparameters were used. All four YOLOv12 models were trained for a short, fixed budget of 10 epochs using randomly initialized weights. To prevent overfitting on highly correlated consecutive video frames, we implemented several strict countermeasures. First, the data split was maintained exclusively at the orifice-level to ensure that validation metrics would accurately expose any overfitting rather than masking it through adjacent-frame leakage. Second, we applied early stopping based on validation fitness, retaining checkpoints with the highest validation performance. Third, robust regularization was utilized throughout training, including weight decay and data augmentations, to limit the models’ capacity to memorize the dataset.
We report frame-based mean average precision (mAP@50, mAP@50-95) and per-frame accuracy as our primary evaluation metric. Second, we simplified the original task along two axes to identify the reason for the weak performance of the first task. To isolate classification capabilities (i.e., no requirement to localize the UO), additional experiments were performed on a simpler task for binary classification of no-/low- and high-grade reflux, based on the same dataset. For this task, an EfficientNet [13] architecture was employed and we report precision, recall and F1 score as our primary metrics. Additionally, to isolate localization capabilities (i.e., no requirement to classify the UO but only to localize the UO), we created a single-class object detection dataset. For this task, YOLOv12 models were used. We report frame-based mean average precision (mAP@50, mAP@50-95) as our primary evaluation metric.
Third, instead of interpreting accuracy on a frame-by-frame basis, we observed whether sequence-level aggregation yields correct conclusions in cross-sequence patient studies. Results are collected by re-using the YOLOv12 models from task 1. We report majority-voted accuracy on UO-level as our primary metric.
To aggregate frame-level predictions into sequence-level classifications, we implemented a temporal majority voting scheme. For each individual video frame within a sequence of a single ureteric orifice, candidate detections with a class confidence score below a specified threshold were discarded. The frame’s vote was then assigned to the class of the remaining detection that had the highest confidence score. If no detections in a frame exceeded the threshold, that frame was logged as a non-prediction.
The final sequence-level classification for the entire ureteric orifice was defined as the class that received the majority of votes across all predicting frames in the sequence. To prevent threshold-selection bias, we employed a leave-one-out (LOVO) cross-validation strategy: for each validation orifice in turn, the optimal threshold was determined using only the remaining orifices, ensuring the evaluated orifice never contributed to its own threshold selection. The accompanying 95% confidence intervals were computed via bootstrap resampling at the orifice level.

3. Results

3.1. Three-Class Object Detection

Training on three-class object detection showed weak results on the test set, as shown in Table 3, with maximum mAP@50 of 0.37 (YOLOv12s) and mAP@50-95 of 0.17 (YOLOv12s). This indicates that the model can only reliably place one out of every 3 target frames correctly, such that it has an overlap of at least 0.5 IoU with the ground-truth label. Similarly, mAP@50-95 of at most 0.17 suggests that the model struggles with highly accurate placement beyond the 0.5 IoU boundary. Examples images are shown in Figure 1.

3.2. Task Simplification

Due to these initial weak results, the task was simplified for binary classification of no-/low- and high-grade reflux, based on the same dataset. Binary classification results showed a slight increase in performance yet remained below expectations, achieving a macro average of 0.62 precision, 0.58 recall and 0.49 F1 score. Looking at individual classes, critical cases (high-grade reflux) showed a low recall of 0.26, limiting the clinical utility of this model. Given the simplifications, the model showed only slight improvements, predicting around half the images correctly.
Class-agnostic object detection improved mAP@0.5 to 0.63 (+0.32 mAP@50) from multi-class prediction for YOLOv12n, yielding even larger improvements than the binary classification. In around 2/3 of the frames, the UO could be located correctly, regardless of what class it may entail. This finding suggests that detecting a UO itself is a difficult task, but the model has more difficulty in classifying a UO than finding it.

3.3. Sequence Aggregation

Aggregating prediction results over cross-sequence studies substantially improved results. Across longer video sequences, up to 77% of videos reached the correct classification via a majority voting system given a confidence threshold. In addition, in a leave-one-out (LOVO) setting, in which a threshold is not selected optimally but averaged via cross-validation, results remained between 0.69 and 0.77 for all models (Table 4). This suggests that even models that produce subpar results on a per-image basis can reach correct conclusions when deployed over longer sequences of time. In addition, this also allows the suggestion that the reduced performance may not be a fundamental limitation of the model, but rather that individual frames may not carry sufficient information, which averages out across the sequence.

4. Discussion

This study represents, to our knowledge, the first application of AI to cystoscopic video analysis for predicting vesicoureteral reflux based on ureteric orifice morphology in children. While frame-based object detection and classification showed poor performance, aggregating predictions across complete cystoscopic sequences substantially improved diagnostic accuracy, reaching up to 77% at the ureteric-orifice level. These findings suggest that diagnostically relevant information is present within cystoscopic examinations but is distributed across multiple frames rather than consistently contained within isolated images.
AI-assisted image analysis has become increasingly integrated into clinical practice, particularly in radiology and gastrointestinal endoscopy, where it supports lesion detection and real-time decision-making [14,15,16,17,18,19]. These successful applications provide a rationale for exploring similar approaches in pediatric cystoscopy.
Recent advances suggest that AI may also support several aspects of VUR management, although current evidence remains limited [20]. Applications of AI in VUR have primarily focused on structured clinical variables or radiological imaging. Most notably, Khondker and colleagues have demonstrated excellent performance for automated grading of reflux on VCUGs, achieving accuracy exceeding 0.8 in both internal and external validation cohorts [21]. The comparatively lower frame-level performance observed in the present study should not be interpreted as evidence that cystoscopic prediction is intrinsically less suitable for AI-based assessment. Rather, it likely reflects the substantially greater complexity of endoscopic image analysis. Unlike previous studies, our model relied exclusively on endoscopic video appearance without incorporating clinical variables or radiological findings. In contrast to standardized radiographic images, cystoscopic videos are characterized by continuous camera motion, variable illumination, fluid-related artefacts, specular reflections, and considerable anatomical variability, all of which increase the difficulty of automated image interpretation and limit model robustness.
Regarding VUR, AI-based models have also been developed to predict the need for VCUG, estimate the risk of recurrent urinary tract infections and predict the endoscopic treatment results [22,23,24]. In addition, large language models have been evaluated as tools for parental counseling [25]. However, grading of VUR on pediatric cystoscopy images and videos using AI has not been performed yet and further studies building on our results need to be performed to evaluate the possibility of AI as a clinical decision-support tool for less experienced practitioners by assisting in the identification of anatomical and pathological features associated with VUR. The ultimate goal could be the implementation of AI-assisted systems that could facilitate earlier and more reliable assessment of patients during cystoscopy which may ultimately complement conventional diagnostic pathways after prospective validation. Furthermore, the integration of VUR detection into the cystoscopic workflow offers the opportunity not only to improve diagnostic efficiency but also to directly guide therapeutic decision-making during the same procedure.
Several factors likely contributed to the limited performance observed on individual frames. Unlike radiological imaging, cystoscopic videos are highly dynamic, with frequent changes in viewing angle, illumination, bladder distension, and partial occlusion of the ureteric orifice by mucosal folds, irrigation fluid, or instrument movement. Specular reflections further reduce image quality and introduce visual artefacts. Moreover, the morphological characteristics associated with reflux are often subtle and may only become apparent when the ureteric orifice is viewed from a favorable angle or during a particular phase of bladder filling. Consequently, many individual frames may contain insufficient diagnostic information for reliable classification, even for experienced pediatric urologists.
The marked improvement observed after sequence-level aggregation supports the concept that cystoscopic diagnosis is inherently a temporal task. During routine cystoscopy, surgeons rarely base their assessment on a single static image but instead integrate information obtained while continuously inspecting the ureteric orifice from multiple perspectives. Majority voting across consecutive frames appears to mimic this clinical decision-making process by reducing the influence of noisy or non-informative images while emphasizing consistently detected morphological features. This observation is consistent with the broader surgical computer vision literature, in which temporal information extracted from video sequences has been shown to improve scene understanding, workflow recognition, and intraoperative decision support compared with approaches based solely on individual frames [26,27].
We opted for the YOLOv12 architecture because it represents a transition toward attention-centric designs in real-time object detection [12]. Cystoscopic video frames present unique challenges, including highly variable lighting conditions, specular reflections and motion blur. YOLOv12 incorporates an attention-based backbone designed to capture global contextual relationships and long-range spatial dependencies more effectively than purely convolutional networks, while still maintaining the low latency required for real-time intraoperative deployment. In the future, however, a systematic benchmark against other model families (e.g., older YOLO generations, different convolutional architectures, etc.) is a logical next step to isolate architectural benefits, alongside the transition to dedicated video-based architectures that model temporal context explicitly.
Our empirical findings indicate that scaling the model capacity tenfold (from YOLOv12n to YOLOv12l) did not yield a corresponding monotonic increase in diagnostic accuracy, with frame-level mAP@50 remaining constrained between 0.31 and 0.37. This non-monotonic variation suggests that the primary bottleneck limiting accuracy is not model capacity or parameter count, but rather the diversity and size of the dataset (specifically, the number of unique ureteric orifices with 117 in total). While the frame count is high (36,604 images), the high visual correlation between consecutive frames from the same sequence means the effective sample size is dictated by the unique anatomical configurations. Future performance improvements will therefore depend heavily on expanding the patient cohort to encompass a broader spectrum of anatomical and clinical variations, rather than merely employing larger or more complex neural network architectures.
Beyond dataset size, the clinical class imbalance significantly influenced the models’ predictive behavior, which represents another key finding. The clinically most critical class (high-grade reflux) was also the least represented, constituting only 19.7% of the included ureteric orifices. Consequently, the models were highly conservative when predicting this class, achieving a high precision of 0.81 but a low recall of 0.26 for high-grade cases. This combination of high precision and low recall is a classic characteristic of deep learning models trained on imbalanced data [28,29]. Notably, this performance asymmetry persisted across both the YOLOv12 detectors and the structurally distinct EfficientNet classifier, indicating that it reflects the underlying composition and inherent difficulty of the endoscopic data rather than the property of a single training configuration. While this low recall is a clear clinical limitation, it highlights the necessity of exploring targeted imbalance remedies in future work (such as class-weighted focal loss or patient-level oversampling) to safely improve diagnostic sensitivity for severe cases.
This study has several limitations. First, it represents a retrospective single-center study with a relatively small number of patients and UOs available for training deep learning models, which also introduced a marked clinical class imbalance that constrained the models’ sensitivity for high-grade cases. Second, although data splitting was performed at the patient level to prevent information leakage, external validation was not available. Third, annotations were based on expert interpretation of cystoscopic morphology, which remains inherently subjective despite standardized labeling procedures. Fourth, only conventional convolutional object detection models were evaluated, whereas dedicated video-based architectures may be better suited for this application. Finally, the heterogeneous indications for cystoscopy among control patients may introduce additional variability in ureteric orifice appearance.
Our results suggest that in the future an intended clinical application could run on the live cystoscopic stream and display an overlay marking the detected ureteric orifice together with the predicted grade and associated confidence. Rather than presenting an isolated prediction for each individual frame, the display would show a continuously updated majority vote over the frames observed so far, which is the intraoperative equivalent of sequence aggregation. Nevertheless, more studies with larger datasets have to be performed to confirm our results and facilitate further development of our model.
In conclusion, this proof-of-concept study demonstrates that AI can extract diagnostically relevant information from pediatric cystoscopic videos for VUR prediction, although frame-level performance remains limited. The substantial improvement achieved through sequence-level aggregation suggests that temporal information is essential for automated interpretation of UO morphology. These findings establish a foundation for future multi-centre studies employing larger datasets including more patients and dedicated video-based deep learning architectures to develop clinically applicable decision-support systems for pediatric endourology.

Author Contributions

Conceptualization, V.W. and H.T.; Data curation, V.W., T.T. and S.T.; Formal analysis, V.W., T.T. and S.T.; Investigation, V.W. and T.T.; Methodology, V.W., T.T., S.T. and H.T.; Project administration, H.T.; Resources, G.S. and H.T.; Software, V.W., T.T. and A.B.; Supervision, G.S. and H.T.; Validation, V.W., T.T. and H.T.; Visualization, V.W. and T.T.; Writing—original draft, V.W., T.T. and H.T.; Writing—review and editing, S.T., G.S. and A.B. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Committee of the Medical University of Graz, approval number 36-262 ex 23/24, 5 April 2024.

Informed Consent Statement

Patient consent was waived due to the retrospective design of the study and the full anonymization of all data prior to analysis, as approved by the ethics committee.

Data Availability Statement

The data supporting the reported results are not publicly available due to institutional data governance restrictions. Requests for data access may be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
VURvesicoureteral reflux
UOureteric orifice
VCUGvoiding cystourethrography
AIartificial intelligence
UTIurinary tract infection
MLmachine learning
UPJOureteropelvic junction obstruction
CVATcomputer vision annotation tool
YOLOyou only look once
LOVOleave-one-out

References

  1. Sargent, M.A. What is the normal prevalence of vesicoureteral reflux? Pediatr. Radiol. 2000, 30, 587–593. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Brakeman, P. Vesicoureteral reflux, reflux nephropathy, and end-stage renal disease. Adv. Urol. 2008, 2008, 508949. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Lebowitz, R.L.; Olbing, H.; Parkkulainen, K.V.; Smellie, J.M.; Tamminen-Mobius, T.E. International system of radiographic grading of vesicoureteric reflux. International Reflux Study in Children. Pediatr. Radiol. 1985, 15, 105–109. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Schaeffer, A.J.; Greenfield, S.P.; Ivanova, A.; Cui, G.; Zerin, J.M.; Chow, J.S.; Hoberman, A.; Mathews, R.I.; Mattoo, T.K.; Carpenter, M.A.; et al. Reliability of grading of vesicoureteral reflux and other findings on voiding cystourethrography. J. Pediatr. Urol. 2017, 13, 192–198. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Mackie, G.G.; Awang, H.; Stephens, F.D. The ureteric orifice: The embryologic key to radiologic status of duplex kidneys. J. Pediatr. Surg. 1975, 10, 473–481. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Lyon, R.P.; Marshall, S.; Tanagho, E.A. The ureteral orifice: Its configuration and competency. J. Urol. 1969, 102, 504–509. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Zvizdic, Z.; Catic, A.; Zivojevic, S.; Jonuzi, A.; Glamoclija, U.; Vranic, S. The correlation between ureteric orifice morphology and primary vesicoureteral reflux grade and the impact on the effectiveness of endoscopic reflux correction. J. Pediatr. Urol. 2024, 20, 295–301. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Kirsch, A.J.; Kaye, J.D.; Cerwinka, W.H.; Watson, J.M.; Elmore, J.M.; Lyles, R.H.; Molitierno, J.A.; Scherz, H.C. Dynamic hydrodistention of the ureteral orifice: A novel grading system with high interobserver concordance and correlation with vesicoureteral reflux grade. J. Urol. 2009, 182, 1688–1692. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Khondker, A.; Kaushal, S.; Wu, J.; Gupta, N.; Kwong, J.C.; Li, T.; Erdman, L.; Rickard, M.; Lorenzo, A.J. Predicting vesicoureteral reflux outcomes using artificial intelligence: A critical appraisal using APPRAISE-AI. PLoS Digit. Health 2026, 5, e0001237. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Scott Wang, H.H.; Vasdev, R.; Nelson, C.P. Artificial Intelligence in Pediatric Urology. Urol. Clin. N. Am. 2024, 51, 91–103. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Li, Z.; Tan, Z.; Wang, Z.; Tang, W.; Ren, X.; Fu, J.; Wang, G.; Chu, H.; Chen, J.; Duan, Y.; et al. Development and multi-institutional validation of a deep learning model for grading of vesicoureteral reflux on voiding cystourethrogram: A retrospective multicenter study. EClinicalMedicine 2024, 69, 102466. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-Centric Real-Time Object Detectors; Neural Information Processing Systems Foundation, Inc. (NeurIPS): San Diego, CA, USA, 2025. [Google Scholar]
  13. Tan, M.; Le, Q. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. arXiv 2019, arXiv:1905.11946. [Google Scholar] [CrossRef] [Scilit]
  14. Geppert, J.; Auguste, P.; Asgharzadeh, A.; Ghiasvand, H.; Patel, M.; Brown, A.; Jayakody, S.; Helm, E.; Todkill, D.; Madan, J.; et al. Software with artificial intelligence-derived algorithms for detecting and analysing lung nodules in CT scans: Systematic review and economic evaluation. Health Technol. Assess. 2025, 29, 1–234. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Lu, J.; Shen, L.; Zhou, C.; Bi, Z.; Ye, X.; Zhao, Z.; Zeng, M.; Wang, M. Image Quality Improvement and Artificial Intelligence Performance in Pulmonary Embolism Detection at Deep Learning Reconstruction-Based Ultra-low Radiation Dose CT Pulmonary Angiography. Acad. Radiol. 2025, 32, 7562–7572. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Park, D.K.; Kim, E.J.; Im, J.P.; Lim, H.; Lim, Y.J.; Byeon, J.S.; Kim, K.O.; Chung, J.W.; Kim, Y.J. A prospective multicenter randomized controlled trial on artificial intelligence assisted colonoscopy for enhanced polyp detection. Sci. Rep. 2024, 14, 25453. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Scherkl, M.; Stranger, N.; Ciornei-Hoffman, A.; Singer, G.; Till, T.; Till, H.; Hrzic, F.; Tschauner, S. Automated AI fracture detection in initial presentation pediatric wrist X-rays: Effects and benefits of adding follow-up examinations. Radiol. Med. 2026, 131, 458–469. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Stranger, N.; Scherkl, M.; Stutz, D.; Janisch, M.; Singer, G.; Hankel, S.; Till, H.; Sackl, M.; Hrzic, F.; Herzog, S.; et al. Impact of Test Set Composition on AI Performance for Pediatric Radiograph Appendicular Skeleton Fracture Detection. Radiology 2026, 318, e250540. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Zheng, F.; Zhu, Q.; Yuan, P.; Wu, F.; Cui, Y.; Xiao, Z.; Li, H.; Li, X.; Wu, J.; Qu, Z.; et al. Application of artificial intelligence for detection of Helicobacter pylori infection by magnetically controlled capsule endoscopy. Surg. Endosc. 2026, 40, 1504–1513. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Kalfa, N.; Bonnin, Y.; Tessier, B.; Garnier, S.; Mekhenane, N.; Cazals, A.; Donmez, M.I. The future of pediatric vesicoureteral reflux management. J. Pediatr. Urol. 2026; in press. [CrossRef] [Scilit] [PubMed]
  21. Khondker, A.; Kwong, J.C.C.; Rickard, M.; Skreta, M.; Keefe, D.T.; Lorenzo, A.J.; Erdman, L. A machine learning-based approach for quantitative grading of vesicoureteral reflux from voiding cystourethrograms: Methods and proof of concept. J. Pediatr. Urol. 2022, 18, 78.e1–78.e7. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Scott Wang, H.H.; Li, M.; Cahill, D.; Panagides, J.; Logvinenko, T.; Chow, J.; Nelson, C. A machine learning algorithm predicting risk of dilating VUR among infants with hydronephrosis using UTD classification. J. Pediatr. Urol. 2024, 20, 271–278. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Bertsimas, D.; Li, M.; Estrada, C.; Nelson, C.; Scott Wang, H.H. Selecting Children with Vesicoureteral Reflux Who are Most Likely to Benefit from Antibiotic Prophylaxis: Application of Machine Learning to RIVUR. J. Urol. 2021, 205, 1170–1179. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Serrano-Durba, A.; Serrano, A.J.; Magdalena, J.R.; Martin, J.D.; Soria, E.; Dominguez, C.; Estornell, F.; Garcia-Ibarra, F. The use of neural networks for predicting the result of endoscopic treatment for vesico-ureteric reflux. BJU Int. 2004, 94, 120–122. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Akyol Onder, E.N.; Ensari, E.; Ertan, P. ChatGPT-4o’s performance on pediatric Vesicoureteral reflux. J. Pediatr. Urol. 2025, 21, 504–509. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. King, A.; Fowler, G.E.; Macefield, R.C.; Walker, H.; Thomas, C.; Markar, S.; Higgins, E.; Blazeby, J.M.; Blencowe, N.S. Use of artificial intelligence in the analysis of digital videos of invasive surgical procedures: Scoping review. BJS Open 2025, 9, zraf073. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Mascagni, P.; Alapatt, D.; Sestini, L.; Altieri, M.S.; Madani, A.; Watanabe, Y.; Alseidi, A.; Redan, J.A.; Alfieri, S.; Costamagna, G.; et al. Computer vision in surgery: From potential to clinical value. npj Digit. Med. 2022, 5, 163. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Oksuz, K.; Cam, B.C.; Kalkan, S.; Akbas, E. Imbalance problems in object detection: A review. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 43, 3388–3415. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Johnson, J.M.; Khoshgoftaar, T.M. Survey on deep learning with class imbalance. J. Big Data 2019, 6, 27. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Examples of a ground truth and prediction using YOLOv12n of a normal ostium (a), a low-grade (b) and a high-grade ostium (c).
Figure 1. Examples of a ground truth and prediction using YOLOv12n of a normal ostium (a), a low-grade (b) and a high-grade ostium (c).
Jcm 15 06246 g001
Table 1. Distribution of classes throughout the proposed dataset. Proportions show a class imbalance between no-, low-, and high-grade reflux.
Table 1. Distribution of classes throughout the proposed dataset. Proportions show a class imbalance between no-, low-, and high-grade reflux.
ClassUreteric OrificesProportion
No reflux6858.1%
Low-grade (grades II–III)2622.2%
High-grade (grades IV–V)2319.7%
Table 2. Model hyperparameters for training models in the YOLOv12 family. Identical configurations are used across all model sizes besides architecture configuration to ensure fair comparisons of architectures.
Table 2. Model hyperparameters for training models in the YOLOv12 family. Identical configurations are used across all model sizes besides architecture configuration to ensure fair comparisons of architectures.
ParameterValue
HardwareSingle NVIDIA RTX 3060, 12GB RAM
ArchitecturePer-model definition
Input image size640 × 640 pixels
Batch size16 (nominal batch size of 64, gradient accumulation over 4 steps
Epochs10 (with early stopping, patience of 4)
OptimizerAdamW
Initial learning rate1.43 × 10−3
Learning rate scheduleLinear decay to 1% of the initial value (final 1.56 × 10−4) with a 3-epoch warm-up
Momentum0.9 (warm-up momentum 0.8)
Weight decay5 × 10−4
Loss weightsBox 7.5, classification 0.5, distribution focal loss 1.5
AugmentationMosaic, horizontal flip (p = 0.5), HSV jitter (0.015, 0.7, 0.4), translation 0.1, scaling 0.5
NMSIoU threshold 0.7, maximum 300 detections per image
Numerical precisionAutomatic mixed precision
Seed0, deterministic mode enabled
Table 3. Three-class object detection results for models from the YOLOv12 family. The metrics mAP@50, mAP@50-95, and accuracy are reported.
Table 3. Three-class object detection results for models from the YOLOv12 family. The metrics mAP@50, mAP@50-95, and accuracy are reported.
ModelmAP@50mAP@50-95Per-Frame Accuracy
YOLOv12n0.310.140.40
YOLOv12s0.370.170.42
YOLOv12m0.310.140.36
YOLOv12l0.320.140.34
Table 4. Sequence aggregation results from models of the YOLOv12 family. Per-UO accuracy for an optimal threshold, per-UO accuracy for a threshold determined via leave-one-out cross-validation, as well as the LOVO 95%-confidence interval are reported.
Table 4. Sequence aggregation results from models of the YOLOv12 family. Per-UO accuracy for an optimal threshold, per-UO accuracy for a threshold determined via leave-one-out cross-validation, as well as the LOVO 95%-confidence interval are reported.
ModelPer-UO Accuracy (Static)Per-UO Accuracy (LOVO)95%-CI (LOVO)
YOLOv12n0.74 (@0.05)0.71[0.57–0.86]
YOLOv12s0.74 (@0.15)0.69[0.51–0.83]
YOLOv12m0.74 (@0.05)0.71[0.57–0.86]
YOLOv12l0.77 (@0.00)0.77[0.63–0.89]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wolfschluckner, V.; Till, T.; Tschauner, S.; Singer, G.; Basharkhah, A.; Till, H. AI-Supported Prediction of Vesicoureteral Reflux in Children Based on Cystoscopic Configuration of the Ureteric Orifice. J. Clin. Med. 2026, 15, 6246. https://doi.org/10.3390/jcm15166246

AMA Style

Wolfschluckner V, Till T, Tschauner S, Singer G, Basharkhah A, Till H. AI-Supported Prediction of Vesicoureteral Reflux in Children Based on Cystoscopic Configuration of the Ureteric Orifice. Journal of Clinical Medicine. 2026; 15(16):6246. https://doi.org/10.3390/jcm15166246

Chicago/Turabian Style

Wolfschluckner, Vanessa, Tristan Till, Sebastian Tschauner, Georg Singer, Alireza Basharkhah, and Holger Till. 2026. "AI-Supported Prediction of Vesicoureteral Reflux in Children Based on Cystoscopic Configuration of the Ureteric Orifice" Journal of Clinical Medicine 15, no. 16: 6246. https://doi.org/10.3390/jcm15166246

APA Style

Wolfschluckner, V., Till, T., Tschauner, S., Singer, G., Basharkhah, A., & Till, H. (2026). AI-Supported Prediction of Vesicoureteral Reflux in Children Based on Cystoscopic Configuration of the Ureteric Orifice. Journal of Clinical Medicine, 15(16), 6246. https://doi.org/10.3390/jcm15166246

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop