Next Article in Journal
Design and Research of a Dual-Target Drug Molecular Generation Model Based on Reinforcement Learning
Next Article in Special Issue
Histopathological Medical Image Classification Using ANN Optimized by PSO with CNN for Feature Extraction
Previous Article in Journal
Potential Recovery and Recycling of Condensate Water from Atlas Copco ZR315 FF Industrial Air Compressors
Previous Article in Special Issue
CTGAN-Augmented Ensemble Learning Models for Classifying Dementia and Heart Failure
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Enhancement Without Contrast: Stability-Aware Multicenter Machine Learning for Glioma MRI Imaging

1
Department of Medicine, Tehran University of Medical Science, Tehran 1417613151, Iran
2
Department of Radiology, School of Paramedical Sciences, Guilan University of Medical Sciences, Rasht 4477166595, Iran
3
School of Electrical and Computer Engineering, College of Engineering, University of Tehran, Tehran 515-14395, Iran
4
Department of Integrative Oncology, Breast Cancer Research Center, Motamed Cancer Institute, Academic Center for Education, Culture and Research (ACECR), Tehran 1517964311, Iran
5
Department of Computer Science, University of British Columbia, Vancouver, BC V6T 1Z4, Canada
6
Technological Virtual Collaboration Company (TECVICO CORP.), Vancouver, BC V5E 3H7, Canada
7
Department of Radiology, University of British Columbia, Vancouver, BC V5Z 1M9, Canada
8
Department of Medicine, University of British Columbia, Vancouver, BC V5Z 1M9, Canada
9
Department of Basic and Translational Research, BC Cancer Research Institute, Vancouver, BC V5Z 1L3, Canada
*
Author to whom correspondence should be addressed.
Inventions 2026, 11(1), 11; https://doi.org/10.3390/inventions11010011
Submission received: 8 December 2025 / Revised: 16 January 2026 / Accepted: 23 January 2026 / Published: 26 January 2026
(This article belongs to the Special Issue Machine Learning Applications in Healthcare and Disease Prediction)

Abstract

Gadolinium-based contrast agents (GBCAs) are vital for glioma imaging yet pose safety, cost, and accessibility issues; predicting contrast enhancement from non-contrast MRI via machine learning (ML) provides a safer, economical alternative, as enhancement indicates tumor aggressiveness and informs treatment planning. However, scanner and population variability hinder robust model selection. To overcome this, a stability-aware framework was developed to identify reproducible ML pipelines for predicting glioma contrast enhancement across multicenter cohorts. A total of 1367 glioma cases from four TCIA datasets (UCSF-PDGM, UPENN-GB, BRATS-Africa, BRATS-TCGA-LGG) were analyzed, using non-contrast T1-weighted images as input and deriving enhancement status from paired post-contrast T1-weighted images; 108 IBSI-standardized radiomics features were extracted via PyRadiomics 3.1, then systematically combined with 48 dimensionality reduction algorithms and 25 classifiers into 1200 pipelines, evaluated through rotational validation (training on three datasets, external testing on the fourth, repeated across rotations) incorporating five-fold cross-validation and a composite score penalizing instability via standard deviation. Cross-validation accuracies spanned 0.91–0.96, with external testing yielding 0.87 (UCSF-PDGM), 0.98 (UPENN-GB), and 0.95 (BRATS-Africa), averaging ~0.93; F1, precision, and recall remained stable (0.87–0.96), while ROC-AUC varied (0.50–0.82) due to cohort heterogeneity, with the MI + ETr pipeline ranking highest for balanced accuracy and stability. This framework enables reliable, generalizable prediction of contrast enhancement from non-contrast glioma MRI, minimizing GBCA dependence and offering a scalable template for reproducible ML in neuro-oncology.

1. Introduction

Magnetic resonance imaging (MRI) is a critically utilized modality in neuro-oncology, as other anatomical techniques do not match its precision in outlining glioma boundaries, peritumoral oedema, or in assessing treatment response [1]. As a result, almost all patients diagnosed with a diffuse glioma, regardless whether categorized as low-grade (I–II) or high-grade (III–IV), undergo multiple MRI scans from the point of diagnosis through to end-of-life care [2]. Gliomas are indeed a heterogeneous group of primary brain tumors arising from glial cells. The global incidence rate of 6–8 per 100,000 individuals highlights their relative rarity but significant clinical impact due to their aggressive nature and poor prognosis. The fact that about 30–40% of these gliomas are high-grade underscores the burden of malignancy within this tumor category [3,4]. The diagnostic utility of MRI in neuro-oncology has been significantly enhanced through the routine use of gadolinium-based contrast agents (GBCAs).
GBCAs exploit the paramagnetic properties of Gadolinium (Gd3+) to reduce T1-weighted MRI images (T1WI) relaxation time, thereby enhancing the visualization of blood-brain barrier (BBB) disruptions—a feature observed in approximately 80% of high-grade gliomas [5]. This enhancement on post-contrast T1WI provides critical biological and clinical insights, aiding in surgical resection planning, radiotherapy target definition, and the application of Response Assessment in Neuro-Oncology criteria (KJR Online) [6]. Bright (hyperintense) areas typically indicate BBB compromise due to leakage from neovascularization, characteristics of anaplastic astrocytoma, glioblastoma (GBM), and other high-grade lesions. In contrast, the absence of enhancement suggests an intact BBB, as seen in most grade I–II gliomas—though exceptions exist, such as “non-enhancing GBM” or focal enhancement in low-grade tumors [7]. These imaging patterns hold significant prognostic value: enhancement often correlates with aggressive histology and poorer outcomes, guides surgical and radiotherapy targeting, determines eligibility for anti-angiogenic trials, and serves as a key imaging marker for distinguishing true progression from pseudo-progression or pseudo-response during long-term monitoring [8]. Because contrast enhancement in gliomas directly reflects heterogeneous BBB disruption across tumor grades, spatial regions, and disease states, this intrinsic variability makes gliomas a uniquely challenging and clinically meaningful target for prediction from non-contrast MRI.
Nonetheless, the growing dependence on GBCAs is still under investigation. The use of linear chelates has caused an increase in nephrogenic systemic fibrosis among individuals with kidney issues, leading to black-box warnings and withdrawal of several products by the University of Maryland Medical System [5,9]. Further autopsy and biopsy studies have revealed that Gd3+ accumulates in a dose-dependent fashion in the dentate nucleus, globus pallidus, bone, and skin, even in those with normal renal function [10]. Although a direct clinical syndrome has not been conclusively linked, these discoveries have increased medicolegal scrutiny and patient concern, particularly among children and young adults who need ongoing imaging throughout their lives [9]. Beyond these biological and safety concerns, GBCAs also impose substantial economic and logistical burdens, particularly in resource-limited settings [11].
The financial and operational challenges associated with GBCAs have exacerbated accessibility issues in neuroimaging. By 2024, the global market for CT/MRI contrast media will have surpassed USD 6 billion, with projections reaching approximately USD 10 billion by 2030, driven largely by oncology demand. In low- and middle-income countries, the cost of a single GBCA vial can exceed a household’s weekly income, while the logistical burdens of intravenous cannulation, creation of testing, and prolonged scan durations further strain limited radiology resources [12]. Compounding these challenges is a critical scientific gap: conventional non-contrast MRI lacks sufficient sensitivity to detect subtle BBB disruptions, while advanced techniques such as diffusion- or perfusion-weighted imaging improve detection but remain imperfect standalone alternatives [13]. Consequently, there is a pressing need for tools that can predict or replicate contrast enhancement (CE) without Gd3+ administration.
Emerging computational approaches aim to address this unmet need [14,15,16,17]. Radiomics features (RFs), which provide quantitative descriptors of tumor texture, shape, and intensity from standard MR images, have shown promise in brain disease studies [18,19,20]. Single-center investigations report that machine learning (ML) models can distinguish enhancing from non-enhancing gliomas with accuracies exceeding 0.80, though these findings are limited by small sample sizes, heterogeneous preprocessing methods, and insufficient external validation [21,22]. Meanwhile, deep learning techniques—particularly convolutional neural networks and generative adversarial networks—have explored image-to-image translation to synthesize “virtual contrast” images from pre-contrast scans, achieving structural similarity indices of 0.80–0.90 in early trials [23]. Despite these advances, concerns persist regarding model generalizability, clinical interpretability, and regulatory acceptance.
A major challenge in ML for medical imaging therefore lies in robust model selection under multicenter heterogeneity [24]. With dozens of dimension reduction algorithms (DRA) and classifiers, the search space quickly expands to thousands of possible combinations, making it difficult to identify which models will generalize well beyond the training set [25,26]. This problem becomes even more critical in multicenter cohorts, where differences in scanners, acquisition protocols, and patient demographics can cause models that perform well internally to fail when applied to external data [27]. Traditional approaches often rely on single-center cross-validation, which risks overfitting and inflates performance estimates. Therefore, the key issue is not only achieving high accuracy but also ensuring stability and reproducibility across heterogeneous datasets, requiring systematic, fair, and scalable evaluation frameworks for model selection [28,29].
This paper presents a comprehensive ML framework for predicting glioma contrast enhancement from non-contrast MRI across multicenter datasets. Using 1367 cases from four large The Cancer Imaging Archive (TCIA) cohorts, the study applies rigorous preprocessing, expert labeling, and standardized RFs extraction to ensure reliable inputs. To provide a controlled and stringent evaluation, this study intentionally restricted the input to non-contrast T1-weighted imaging to address a clinically motivated question: whether contrast-related information can be reliably inferred from the sequence most directly affected by gadolinium administration, while minimizing inter-sequence and inter-center variability in multicenter settings. A large-scale model search was conducted by evaluating 1200 combinations of 48-dimensional reduction algorithms and 25 classifiers. To address variability and enhance generalizability, the pipeline incorporates rotational dataset partitioning, internal cross-validation, and independent external testing. Model performances are assessed with multiple metrics and integrated into a composite scoring system that balances accuracy and stability, enabling systematic ranking and selection of the best models. Importantly, the primary contribution of this work is the introduction of a stability-aware, multicenter framework for systematic model selection and evaluation. By explicitly accounting for performance variability across folds, rotations, and external cohorts, the proposed framework is designed to guide the identification of robust and generalizable ML pipelines suitable for real-world clinical deployment.

2. Materials and Methods

  • (i) Patient Data Preparation: Four publicly available large glioma datasets comprising a total of 1367 cases from TCIA were analyzed, each containing T1WI with expert-validated tumor segmentations. As illustrated in Table 1, the datasets included 495 samples from UCSF PDGM [30] (296 males, 199 females; mean age 56.87 ± 15.02 years), 671 samples from UPENN-GB [31] (405 males, 266 females; mean age 62.45 ± 12.36 years), 143 samples from BRATS Africa [32], and 58 samples from BRATS TCGA LGG [33]. Imaging datasets from TCIA exhibited variations in scanner types, MRI sequences, acquisition parameters, preprocessing methods, and data quality. For example, the UCSF-PDGM dataset includes preoperative 3T MRI scans with 3D and 2D sequences, processed with eddy current correction, DTI processing, and skull stripping using deep learning. This single-center dataset contains preoperative scans with no prior treatment (except biopsy), some variability in contrast agents, but no missing sequences. The UPENN-GBM dataset features multi-parametric MRI scans from GBM patients, acquired on 1.5 T and 3 T scanners, with preprocessing including skull-stripping, co-registration, automated tumor segmentation, and RF extraction. This single-center dataset supports radiogenomic studies and includes only preoperative scans with no missing sequences. The BRATS-Africa dataset includes multiparametric MRI scans from brain tumor patients acquired across six Nigerian centers using 1.5 T scanners, with preprocessing steps like N4 bias field correction, skull-stripping, and rigid registration. It contains preoperative scans with variability in scanner types but no missing sequences, supporting diagnostic tool development for African populations. The BRATS-TCGA-LGG dataset offers preoperative multi-parametric MRI scans from glioma patients, acquired across multiple institutions. Preprocessing included skull-stripping, co-registration, and automated tumor segmentation, followed by manual corrections. This dataset supports molecular and outcome studies and also has no missing sequences.
Table 1. Summary of TCIA Glioma Datasets: Demographics, Preprocessing, and Partitioning; abbreviations: T1: T1-weighted MRI sequence. T1-CE/T1-Gd: T1-weighted sequence with Contrast Enhancement. “Gd”: Gadolinium. T2: T2-weighted MRI sequence. T2-FLAIR: T2-weighted Fluid-Attenuated Inversion Recovery sequence. DWI: Diffusion-Weighted Imaging. DTI: Diffusion Tensor Imaging. GBM: Glioblastoma, HGG: High-Grade Glioma, LGG: Low-Grade Glioma.
Table 1. Summary of TCIA Glioma Datasets: Demographics, Preprocessing, and Partitioning; abbreviations: T1: T1-weighted MRI sequence. T1-CE/T1-Gd: T1-weighted sequence with Contrast Enhancement. “Gd”: Gadolinium. T2: T2-weighted MRI sequence. T2-FLAIR: T2-weighted Fluid-Attenuated Inversion Recovery sequence. DWI: Diffusion-Weighted Imaging. DTI: Diffusion Tensor Imaging. GBM: Glioblastoma, HGG: High-Grade Glioma, LGG: Low-Grade Glioma.
DatasetSubjectsMalesFemalesTumor GradesMRI ModalitiesSurvival DataAccess
Restrictions
BRATS-Africa143N/AN/ALGG/GBM/HGGT1, T1-CE, T2, T2-FLAIRNot specifiedPublic
(CC BY 4.0)
UCSF-PDGM495296199Grade II, III, IVT2, T2/FLAIR, DWI, T1-GdNot specifiedPublic
(CC BY 4.0)
BRATS TCGA LGG58N/AN/ALGG (Grades I–II)T1, T1-Gd, T2, T2-FLAIRNot specifiedPartial Restrictions (TCIA)
UPENN-GBM671405266GBM (Grade IV)T1, T1-CE, T2, T2/FLAIR, DWI, DSC, and DTIOverall SurvivalPublic
(CC BY 4.0)
  • (ii) Labeling: The dataset was meticulously labeled to ensure accurate representation of glioma CE status. Image labeling followed a structured workflow: (a) non-contrast T1WI served as the input for RF extraction, and (b) the corresponding contrast-enhanced T1WI was used as the ground truth (Figure 1). Radiologists identified enhancement patterns (e.g., ring-like, nodular) by comparing (a) and (b). Contrast enhancement is a key indicator in glioma imaging, as it reflects BBB disruption, a hallmark of tumor aggressiveness. This feature is critical for refining tumor grading, informing surgical and radiotherapy planning, and guiding treatment monitoring under the RANO criteria. Furthermore, enhancement patterns are crucial for differentiating true progression from pseudoprogression following chemoradiotherapy, a challenge that can otherwise result in premature therapy changes or unnecessary interventions. Given these factors, the ability to reliably predict enhancement, without gadolinium, offers a safer, cost-effective approach while retaining the diagnostic and prognostic value traditionally obtained from contrast imaging. Therefore, binary classification labels were assigned: enhanced = 1, non-enhanced = 0. All labels were independently verified by expert radiologists to preserve clinical relevance and ensure consistency across multicenter data. Moreover, all tumor segmentations were performed by an experienced radiologist and independently validated by a second expert to minimize inter-observer variability. This step ensured the reliability of the regions of interest (ROIs) used for RF extraction.
Figure 1. Labeling workflow: non-contrast T1WI (a) used for feature extraction, contrast-enhanced T1WI (b) as ground truth for generation of radiologist-verified binary labels (red arrow, enhanced = 1, non-enhanced = 0), to be predicted from non-contrast imaging.
Figure 1. Labeling workflow: non-contrast T1WI (a) used for feature extraction, contrast-enhanced T1WI (b) as ground truth for generation of radiologist-verified binary labels (red arrow, enhanced = 1, non-enhanced = 0), to be predicted from non-contrast imaging.
Inventions 11 00011 g001
  • (iii) MRI Intensity Normalization: Non-contrast T1WI data were standardized using min–max normalization to account for variations in scanner protocols and imaging conditions across multicenter cohorts. This preprocessing step enhanced the comparability of RFs extracted from different datasets.
  • (iv) RF Extraction: A comprehensive set of RFs was extracted using PyRadiomics [34], standardized in reference to the image biomarker standardization initiative (IBSI). RFs are quantitative descriptors extracted from medical images that characterize lesion intensity, texture, shape, and spatial heterogeneity for computational analysis. In total, 108 standardized RFs were considered as the reference set, comprising 19 first-order (FO) features that describe voxel intensity distributions, 15 shape-based features (SF) that quantify tumor geometry, 23 gray-level co-occurrence matrix (GLCM) features that capture pairwise spatial intensity relationships, 16 gray-level size zone matrix (GLSZM) features that characterize homogeneous intensity regions, 16 gray-level run length matrix (GLRLM) features that describe consecutive runs of similar intensities, five neighborhood gray-tone difference matrix (NGTDM) features that quantify local intensity variations, and 14 gray-level dependence matrix (GLDM) features that measure voxel dependency patterns within the tumor.
  • (v) Rotational Data Partitioning into Training and Test Sets: We performed a rotational ML analysis in which three datasets were combined and used for five-fold cross-validation, while the remaining dataset was reserved for external testing. This process was repeated three times to ensure robustness. The TCGA-LGG dataset (59 patients) was included only in the five-fold cross-validation due to its insufficiently balanced labels, and therefore excluded from the rotational external testing procedure. Internal validation was performed and optimized via stratified five-fold cross-validation and grid search.
  • (vi) Min–Max Normalization of RFs: Extracted RFs underwent Min–Max normalization to scale values between 0 and 1. This step ensured uniformity in feature ranges, preventing bias in ML models due to varying magnitudes. We used only the four training folds of the cross-validation process for data normalization.
  • (vii) Machine Learning Algorithms: A total of 48 dimensionality reduction techniques (25 feature selection algorithms (FSAs) and 23 attribute extraction algorithms (AEAs)) were evaluated for their ability to isolate the most informative and non-redundant features. FSAs identify and retain a subset of the original RFs based on relevance or statistical criteria, whereas AEAs transform the original feature space into a lower-dimensional representation. Together, these approaches reduce feature redundancy, mitigate overfitting, and improve model stability and generalizability in multicenter settings. FSAS/AEAs were configured to reduce the feature space to 10 dimensions.
The FSAs encompassed several categories. Filter-based methods, which rank features independently of classifiers, included the Chi-Square Test (CST), Correlation Coefficient (CC), Mutual Information (MI, a statistical measure of dependency between features and class labels), and Information Gain Ratio. Statistical hypothesis-based methods, such as ANOVA F-Test (AFT), ANOVA p-Value Selection (APT), Chi2 p-value selection, and Variance Thresholding (VT), evaluated feature discriminativeness based on statistical significance. Wrapper-based methods, including Recursive Feature Elimination (RFE), Univariate Feature Selection (UFS), Sequential Forward Selection (SFS), and Sequential Backward Selection (SBS), iteratively assessed feature subsets using classification performance. Embedded methods, such as Lasso, Elastic Net, Embedded Elastic Net, and Stability Selection, performed feature selection during model training. Ensemble-based feature selection approaches, including Random Forest feature importance, Extra Trees importance, and permutation importance, leveraged collections of decision trees to capture non-linear feature interactions. Additional methods addressed multiple testing and multicollinearity, including False Discovery Rate (FDR), Family-Wise Error (FWE), and Variance Inflation Factor (VIF). Dictionary-based strategies employed Principal Component Analysis (PCA) or sparse loadings to enhance stability and interpretability.
AEAs provided a complementary strategy by projecting features into compact subspaces that preserve variance, class separation, or non-linear structure. These included linear projection methods such as PCA, Truncated PCA, Sparse PCA (SPCA), and Kernel PCA; Independent Component Analysis (ICA) and FastICA for extracting statistically independent latent variables; Factor Analysis for modeling hidden structure; and Non-negative Matrix Factorization (NMF) for parts-based representations. Supervised linear techniques, such as Linear Discriminant Analysis (LDA), maximized class separability in the transformed space. Non-linear manifold learning methods, including t-SNE, UMAP, Isomap, Locally Linear Embedding (LLE), Spectral Embedding, Multidimensional Scaling (MDS), and Diffusion Maps, captured complex non-linear relationships in high-dimensional radiomics space. Deep learning approaches, such as shallow and deep autoencoders, enabled data-driven feature compression through reconstruction optimization. Additional strategies included Feature Agglomeration for hierarchical grouping, Truncated Singular Value Decomposition (TSVD) for matrix decomposition, and random projection methods (Gaussian Random Projection, Sparse Random Projection, and Feature Hashing) for scalable compression.
Each reduced feature set was evaluated using 25 classification algorithms. These included tree-based ensemble classifiers, such as Decision Trees, Random Forest, Extra Trees (ETr), AdaBoost, and HistGradient Boosting, which combine multiple decision trees to improve robustness and generalization. Meta-ensemble strategies, including stacking, voting classifiers, and bagging, aggregate multiple base learners to enhance predictive stability. Margin- and distance-based classifiers, such as Support Vector Machines (SVM) and k-Nearest Neighbors (KNN), modeled decision boundaries and sample similarity. Probabilistic classifiers, including Naive Bayes variants and Gaussian Process classifiers, modeled class probabilities with uncertainty estimation. Neural network-based classifiers, such as Multilayer Perceptron (MLP), captured complex non-linear patterns, while gradient-boosted frameworks such as Light Gradient Boosting Machine (LGBM) and Extreme Gradient Boosting (XGB) optimized performance through gradient-based learning. Additional classifiers, including LDA, Nearest Centroid, Decision Stump, Dummy Classifier, and Stochastic Gradient Descent Classifier (SGDC), ensured methodological diversity.
  • (viii) Rotational Model Selection: We propose a comprehensive model evaluation and selection pipeline for learning tasks in multicenter studies. The pipeline incorporates three-fold rotational validation and five-fold internal cross-validation within each rotation, ensuring a robust assessment of both performance and stability across combinations of DRAs and classifiers. Performance metrics are computed at multiple levels and aggregated through a composite scoring system, enabling systematic model ranking and selection.
Multi-Rotation and Cross-Validation Scheme. To account for center-related variability and enhance generalizability, the entire training and validation process is repeated across three rotation splits. Each rotation simulates a distinct configuration of data partitioning across sites, thereby mimicking real-world deployment scenarios. Within each rotation:
  • Each DRA-classifier pair is evaluated using five-fold cross-validation.
  • Performance metrics are computed in each fold and averaged to derive robust estimates.
  • Both internal validation metrics (from five-fold cross-validation) and external test metrics (from a held-out external set per rotation) are recorded.
This process resulted in three averaged cross-validation and external metric values per metric per model, along with corresponding standard deviations (SD) across the five folds. Importantly, only the five-fold cross-validation results were used for model scoring and selection. The external test performances were held out entirely during the model ranking process to avoid information leakage.
Performance Metrics: We computed the following five evaluation metrics: Accuracy, F1 Score, Precision, Recall, and Area Under the Curve (AUC). For each DRA-classifier pair:
  • The mean of each metric was calculated across the five-fold cross-validation within a rotation.
  • The SD of each metric across folds was also computed.
  • Over three rotations, this produced a total of:
    15 mean values per model: 3 rotations × 5 metrics
    15 SD values per model: 3 rotations × 5 metrics
Aggregation and Normalization. To enable fair comparison across different metrics and models, we applied min–max normalization to the metric values before scoring. Let:
  • Mij be the average metric i across folds in rotation j.
  • Sij be the SD of metric i across folds in rotation j.
We normalized the metric means (Equation (1)) and SD (Equation (2)) across all models
M i j ^ = M i j min M i max M i min M i
S i j ^ = S i j min S i max S i min S i
Then we inverted the normalized SD to compute a stability score (Equation (3)):
S t a b i l i t y i j = 1 S i j ^
Composite Scoring Formula: The final model selection score aggregates both accuracy and stability across all metrics and rotations (Equation (4)). Specifically, for each DRA-classifier pair, we computed:
Fin al   Score = 1 30 i = 1 5 j = 1 3   M i j ^ + S t a b i l i t y i j
where:
  • 5 metrics × 3 rotations = 15 normalized metric averages
  • 5 metrics × 3 rotations = 15 normalized SD (converted to stability)
  • Total = 30 terms, and the final score is divided by 30 to normalize it to the range [0, 1] while balancing performance and stability equally.
Model Ranking and Selection: Each model was then:
  • Assigned a final score as computed above
  • Ranked in descending order of score (higher is better)
  • Mapped to its cross-validation and external metrics for interpretability
The pipeline is compatible with multicenter designs and large-scale model comparisons, enabling automatic, fair, and interpretable model selection under realistic clinical settings.

3. Results

Table 2 presents the Top 10 performing DRA–CA pairs for glioma outcome prediction (enhanced vs. non-enhanced), evaluated through the proposed model evaluation pipeline. The pipeline incorporates three-fold rotational validation, five-fold cross-validation, and composite scoring, allowing for systematic and unbiased ranking of all 1200 tested combinations generated from 48 DRAs paired with 25 CAs.
Among these, the MI + ETr pair emerged as the top performer with the highest overall score (0.941), showing balanced and stable performance across all major metrics: accuracy (0.94 ± 0.02), F1 score (0.92 ± 0.02), precision (0.94 ± 0.02), and recall (0.93 ± 0.02). This consistent profile underscores its robustness in distinguishing enhanced from non-enhanced glioma cases. Feature Embedding (FEW) + ETr and Extra Trees Importance feature selection (ETIm) + Gaussian Process (GP) followed closely, both achieving accuracies of 0.94 ± 0.03 with strong F1 scores (0.93 ± 0.02 and 0.93 ± 0.01, respectively), demonstrating their ability to balance predictive sensitivity and specificity.
Other high-ranking models, such as AFT + LGBM classifier, Recursive Feature Elimination (RFE) + MLP, and FEW + LGBM, maintained similarly strong accuracies (~0.94) with narrow variability across cross-validation folds, reflecting the pipeline’s ability to capture stable performance across multiple multicenter data partitions. Interestingly, ETr emerged repeatedly as a dominant classifier in the top-performing combinations (e.g., MI + ETr, FEW + ETr, UFS + ETr, RFE + ETr, TSVD + ETr), suggesting its strong adaptability when paired with different DRAs.
All Top 10 pairs exhibited relatively small SD across accuracy, F1, precision, and recall, underscoring their reproducibility and resilience to data heterogeneity. While the area under the receiver operating characteristic curve (ROC-AUC) values showed more variability (ranging from 0.77 to 0.82), this did not compromise the overall robustness of the top-ranked models. Taken together, these results highlight that the proposed pipeline reliably identifies models with stable and generalizable performance, positioning them as strong candidates for multicenter glioma outcome prediction tasks. Supplemental File S1, Sheet 1, presents the complete list of average rotational performances, rankings, and scores for all 1200 CA + DRA combinations. Moreover, no significant differences were observed among the Top 10 models based on the paired t-test and the Benjamini–Hochberg false discovery rate correction for five-fold cross-validation results.
Consistent with the internal cross-validation findings, the MI + ETr pair remained the best-performing combination, achieving the highest composite score (0.941). Its performance across key metrics was balanced, with an accuracy of 0.93 ± 0.06, an F1 score of 0.91 ± 0.06, a precision of 0.93 ± 0.06, and a recall of 0.91 ± 0.08, demonstrating strong reproducibility of internal results in an unseen cohort. Similarly, FEW + ETr and ETIm + GP maintained competitive performance, each reaching an accuracy of 0.93 ± 0.06 but with slightly reduced F1 scores (0.89 ± 0.12 and 0.88 ± 0.11, respectively). These results suggest that while predictive accuracy remained stable, the balance between precision and recall was more sensitive to variations in external data.
Across the Top 10 models, external testing confirmed a high degree of stability in accuracy, precision, and recall, with mean values consistently around 0.93 and SD tightly bound. This reproducibility across independent datasets highlights the robustness of the pipeline’s ranking methodology. However, ROC-AUC values demonstrated greater variability (ranging from 0.50 to 0.77), reflecting the influence of dataset heterogeneity on threshold-dependent metrics. Notably, some models, such as RFE + MLP, achieved relatively high accuracy and precision but showed substantial fluctuations in ROC-AUC (0.50 ± 0.29), indicating that performance consistency across decision thresholds is less reliable in certain DRA-classifier pairings.
When comparing internal cross-validation (Table 2) and external validation (Table 3), a strong alignment in ranking patterns was observed, with MI + ETr, FEW + ETr, and ETIm + GP consistently emerging as top performers. The close match between cross-validation and external testing supports the robustness of the evaluation pipeline and provides confidence that the identified top models are generalizable across multicenter datasets. At the same time, the discrepancies in ROC-AUC highlight areas for further refinement, suggesting that while overall classification performance is stable, the discriminative ability across thresholds warrants closer investigation. Similar to cross-validation, no significant improvement was observed among the Top 10 models in external testing.
Table 4 summarizes the Top 10 performing DRA–CA pairs for glioma outcome prediction, trained on UPENN-GB, BRATS-Africa, and BRATS-TCGA-LGG, with UCSF-PDGM used exclusively for external testing. During internal cross-validation, all Top 10 models achieved consistently high performance, with mean accuracies of 0.96 and narrow error margins across F1 score, precision, and recall. The highest-ranked pair, MI + ETr, achieved an internal accuracy of 0.96 ± 0.01, F1 score of 0.93 ± 0.01, precision of 0.96 ± 0.01, and recall of 0.95 ± 0.01, with a ROC-AUC of 0.70 ± 0.02. Other top-performing models, including FEW + ETr, ETIm + GP, and AFT + LGBM, showed nearly identical accuracies of 0.96 and balanced performance across metrics, reflecting stable and reproducible results in multicenter cross-validation.
When evaluated on the external UCSF-PDGM dataset, model performance remained strong but with modest reductions compared to internal validation, reflecting the challenge of generalizing across independent cohorts. The top-ranked MI + ETr pair achieved an external accuracy of 0.87 ± 0.00, F1 score of 0.85 ± 0.01, precision of 0.87 ± 0.00, recall of 0.82 ± 0.01, and ROC-AUC of 0.81 ± 0.04. Comparable results were observed for FEW + ETr and ETIm + GP, with external accuracies of 0.86–0.87 and F1 scores of 0.75. These findings demonstrate that while predictive accuracy and precision translated well across datasets, F1 scores and ROC-AUC showed greater variability, suggesting sensitivity to label distribution and class imbalance in external data.
Overall, the strong alignment between internal cross-validation and external testing underscores the robustness and reproducibility of the proposed evaluation pipeline. The repeated appearance of Extra Trees (ETr) as a top-performing classifier in multiple high-ranking pairs highlights its adaptability and stability when paired with different DRAs. Although performance metrics declined modestly on external testing, the maintenance of accuracies near 0.87 across models suggests that the identified top-performing pairs are well-suited for multicenter glioma outcome prediction. Supplemental File S1, Sheet 2, presents the complete list of average five-fold cross-validation and external testing performances, rankings, and scores for all 1200 CA + DRA combinations.
Table 5 presents the Top 10 performing DRA–CA pairs for glioma outcome prediction, trained on UCSF-PDGM, BRATS-Africa, and BRATS-TCGA-LGG, with UPENN-GB used exclusively for external testing. The highest-ranked MI + ETr pair achieved an internal score of 0.941, with accuracy (0.91 ± 0.02), F1 score (0.90 ± 0.03), precision (0.91 ± 0.02), recall (0.90 ± 0.02), and ROC-AUC (0.84 ± 0.05). FEW + ETr and ETIm + GP followed closely, with nearly identical cross-validation metrics and narrow variability, indicating robust stability across multicenter datasets.
When tested externally on UPENN-GB, performance further improved compared to internal validation. All Top 10 models achieved very high accuracies of 0.98 ± 0.00, with F1 scores, precision, and recall also at 0.97–0.98 ± 0.00, reflecting near-perfect classification of glioma outcome status. For example, MI + ETr, the top-ranked pair, demonstrated an accuracy of 0.98 ± 0.00, an F1 score of 0.97 ± 0.00, a precision of 0.98 ± 0.00, and a recall of 0.98 ± 0.00, highlighting its exceptional reproducibility. While ROC-AUC values ranged from 0.65 to 0.75, these were somewhat lower relative to other metrics, suggesting that threshold-dependent discrimination is more variable than overall classification accuracy.
Table 6 reports the Top 10 performing DRA–CA pairs for glioma outcome prediction, trained on UCSF-PDGM, UPENN-GB, and BRATS-TCGA-LGG, with BRATS-Africa used exclusively for external testing. Internal cross-validation demonstrated consistently high performance across all Top 10 pairs, with mean accuracies of 0.94–0.95, F1 scores and precision values near 0.94, and recalls also around 0.94 ± 0.01. The top-ranked MI + ETr pair achieved an internal score of 0.941, with balanced performance across accuracy, F1 score, precision, and recall of 0.94 ± 0.00, as well as ROC-AUC (0.89 ± 0.03). Similar robustness was observed for FEW + ETr. On external testing with BRATS-Africa, model ROC-AUC performance remained lower than internal results. The MI + ETr pair achieved an accuracy of 0.95 ± 0.00, F1 score of 0.92 ± 0.00, precision of 0.95 ± 0.00, and recall of 0.94 ± 0.00, while the ROC-AUC was 0.56 ± 0.03, indicating reduced discriminatory power across decision thresholds.
FEW + ETr and ETIm + GP followed with comparable accuracy, F1 score, precision, and recall, but demonstrated higher ROC-AUC values (0.63 and 0.71, respectively). Several other pairs, including AFT + LGBM and RFE_MLP, also maintained external accuracies of 0.95, underscoring the consistency of top-ranked models across diverse datasets. Overall, the Top 10 pairs exhibited strong generalization on BRATS-Africa, with accuracies consistently around 0.95 and narrow variability across F1 score, precision, and recall. However, ROC-AUC values were notably lower (0.56–0.71) compared to internal validation, indicating that while the models achieved stable overall classification, their ability to discriminate across probability thresholds may be more sensitive to regional dataset differences. These findings reinforce the robustness of the proposed evaluation pipeline while highlighting the challenges of generalizing across geographically distinct imaging datasets.

4. Discussion

This study was motivated by the well-documented limitations of GBCAs and the growing need for safer, more accessible imaging strategies in neuro-oncology. In parallel, model selection remains a central challenge in medical imaging ML, particularly in multicenter settings where variability in scanners, acquisition protocols, and patient populations can substantially affect generalizability. To address these challenges, this study systematically combined 48 DRAs and 25 classifiers to construct and evaluate 1200 distinct learning pipelines. This extensive exploration was necessary given that no single algorithm consistently performs optimally across heterogeneous datasets, while also highlighting the risk of selecting models that appear strong in internal validation but fail to generalize externally. Our findings, therefore, emphasize that robust model selection must balance performance with stability to ensure reproducibility across diverse clinical environments.
The inclusion of a wide spectrum of DRAs and CAs enhanced the framework’s ability to identify robust pipelines. DRAs such as MI, FEW, and ETIm were effective in isolating informative (RFs), while projection-based methods (e.g., PCA, UMAP) captured complementary variance and manifold structure. On the classifier side, ensemble models such as ETr and LGBM demonstrated consistent superiority in handling non-linear feature interactions and redundancy, while other algorithms, such as GP and MLP, offered complementary perspectives. This diversity ensured that the framework did not rely on one family of methods but instead selected pipelines resilient across centers. The repeated dominance of MI + ETr and FEW + ETr underscores how combinations of strong feature selectors with ensemble classifiers can produce both accuracy and stability. To our knowledge, few prior studies have systematically tested such a large space of pipelines across multicenter datasets, making this work one of the most comprehensive evaluations of reproducible model selection in glioma imaging to date.
Traditional cross-validation within a single dataset often inflates performance estimates [35,36]. In contrast, rotational validation provided a more rigorous stress test by training on three datasets while reserving one for external testing. This strategy more closely reflects real-world deployment, where models developed in one center must generalize to entirely new populations. As expected, mean accuracies varied, ranging from ~0.91–0.96 in cross-validation (e.g., MI + ETr and FEW + ETr on internal folds, with an average rotational accuracy of ~0.94 across combinations) to ~0.87 in UCSF-PDGM external testing, ~0.98 in UPENN-GB external testing, and ~0.95 in BRATS-Africa external testing, corresponding to an average external rotational accuracy of ~0.93 across top-performing models. Importantly, the relative rankings of the top-performing models remained stable across rotations, indicating that the framework successfully filtered out unstable pipelines while preserving reproducible ones. Such stability is crucial for clinical translation, where reliability across sites is more valuable than isolated peaks of accuracy. This design parallels the conditions of regulatory validation, where reproducibility across independent sites is increasingly viewed as essential for AI approval.
A key strength of this framework was the integration of SD into the composite scoring system. This step penalized models that appeared strong on average but showed unstable behavior across folds or datasets. For example, MI + ETr consistently delivered reliable performance, with accuracy of 0.94 ± 0.02 in rotational validation and 0.87–0.98 across UCSF and BRATS-Africa external tests, alongside a balanced F1 of 0.92 ± 0.02. In contrast, RFE + MLP achieved similar mean accuracy but produced highly unstable ROC-AUC values, dropping to 0.50 ± 0.29 in rotational external testing. Without SD integration, such unstable models might have been ranked among the top despite their lack of reproducibility. By incorporating SD across cross-validation, rotations, and external tests, the framework favored models that combined high accuracy with low variability—an essential requirement for deployment in heterogeneous clinical environments. This illustrates how stability-aware scoring can act as a safeguard against misleading performance peaks, ultimately guiding more trustworthy model selection.
Another important observation was the contrast between consistent and inconsistent performance metrics. Accuracy, F1, precision, and recall remained stable across both internal and external evaluations, with the majority ranging from ~0.87–0.96 and in some cases reaching 0.98 in the UPENN-GB external test. These metrics consistently captured classification performance at fixed thresholds and showed narrow variability across folds and datasets. In contrast, ROC-AUC values were far less stable, spanning ~0.50 to 0.82 in rotational cross-validation and external testing, with the lowest values observed in BRATS-Africa and in unstable combinations such as RFE + MLP. This divergence suggests that while thresholded clinical decisions (enhanced vs. non-enhanced) are robust, threshold-independent discrimination is more sensitive to cohort shifts and calibration differences. The discrepancy underscores why reliance on a single metric is problematic. By integrating multiple metrics into a composite score, the framework ensured that top-ranked models—such as MI + ETr and FEW + ETr—were evaluated holistically rather than being selected on the basis of one potentially misleading criterion.
A plausible explanation for this behavior lies in the interaction between cohort heterogeneity and probabilistic model outputs. In multicenter settings, differences in class prevalence, label imbalance, and population composition can shift the distribution of predicted probabilities without substantially affecting performance at a predefined operating point. As a result, models may remain well calibrated around the decision threshold used for classification—yielding stable accuracy, precision, recall, and F1 score—while exhibiting reduced separation between positive and negative classes across the full range of thresholds, leading to lower or more variable ROC-AUC. In addition, variations in scanner characteristics and acquisition protocols can affect feature scaling and probability calibration across cohorts, further amplifying threshold-independent variability. This phenomenon explains why strong performance at clinically relevant operating points may coexist with diminished global discrimination and underscores the importance of interpreting ROC-AUC in the context of dataset shift rather than as a sole indicator of clinical utility. Collectively, these observations reinforce the need for stability-aware evaluation strategies that account for both fixed-threshold performance and probabilistic variability when assessing model robustness in heterogeneous clinical environments.
Overall, the repeated appearance of similar high-performing pipelines across rotations suggests a consistent and interpretable pattern in what types of information are most informative for contrast enhancement prediction. Across datasets, feature selection tended to favor radiomic descriptors capturing intensity distribution and intratumoral texture heterogeneity, particularly first-order and gray-level texture features, while shape-based features contributed less consistently. This indicates that enhancement-related information in non-contrast T1-weighted imaging is primarily encoded in voxel-level intensity and spatial texture characteristics rather than in global tumor morphology.
From a clinical perspective, the ability to predict contrast enhancement from non-contrast-enhanced MRI represents an important step toward reducing dependence on GBCAs. Although GBCAs provide critical diagnostic information, they are associated with safety concerns, economic costs, and logistical challenges, especially in patients requiring repeated imaging or in resource-limited settings. The demonstrated stability of accuracy, F1 score, precision, and recall across multicenter datasets highlights that ML models can provide reproducible and clinically reliable CE prediction. Although ROC-AUC values showed greater variability across rotations and external datasets, this primarily reflected threshold sensitivity rather than a fundamental weakness of the models. In practice, this suggests that stable accuracy, precision, and recall at clinically relevant thresholds are more important for decision-making than global threshold-free metrics. Furthermore, non-contrast prediction could be especially impactful in longitudinal monitoring, pediatric neuro-oncology, and renal-impaired populations, where GBCA exposure is problematic.
In terms of intended clinical use, the proposed framework is primarily envisioned as a decision-support tool that complements radiologist assessment rather than replacing contrast-enhanced imaging. In settings with limited access to GBCAs or when contrast administration is contraindicated, the framework may also serve as a screening or triage aid to identify patients most likely to benefit from contrast-enhanced studies. In addition, the framework is designed to support research and multicenter studies by providing a stability-aware approach for selecting reproducible ML pipelines, thereby improving consistency and reliability across heterogeneous clinical environments. In clinical workflows, the stability-aware frameworks identified in this study can be employed as an upstream decision layer to support downstream virtual post-contrast T1-weighted image generation approaches, such as image-to-image translation or generative models, thereby improving the reliability and safety of gadolinium-free imaging pipelines in multicenter clinical settings.
This study has limitations. First, it relied solely on RFs derived from non-contrast T1-weighted images, whereas in clinical practice, radiologists typically integrate information from additional sequences such as T2-weighted and FLAIR imaging when anticipating contrast enhancement. This restriction was a deliberate design choice to evaluate whether contrast-related information can be inferred from T1-weighted imaging alone; however, future work should incorporate multi-parametric MRI (e.g., T2, FLAIR, DWI, PWI) or deep RFs to capture complementary tumor characteristics and further improve prediction performance. Second, the variability of ROC-AUC across external datasets underscores the need for calibration and domain adaptation techniques to improve threshold-independent discrimination. Third, while the pipeline prioritized reproducibility, interpretability was not deeply explored. Future directions include applying explainability methods (e.g., SHAP values, radiomic signature analysis) to link selected (RFs) to biological correlates, as well as exploring federated learning strategies to further enhance generalization across centers without direct data sharing.

5. Conclusions

In summary, this study establishes a stability-aware framework for model selection in multicenter ML, combining large-scale evaluation of diverse DRA–CA combinations with rotational validation and external testing. Rather than relying on inflated single-center cross-validation performances, this approach highlights generalizable pipelines, with MI + ETr emerging as a robust example. The framework achieved consistently high accuracies and F1 scores (0.87–0.98 externally; 0.91–0.96 in cross-validation), with stable precision and recall, while AUC values showed greater variability (0.50–0.82), reflecting cohort heterogeneity. Beyond methodological advances, the framework addresses a critical clinical gap by offering a pathway to reduce dependence on Gd3+, particularly valuable for longitudinal monitoring, pediatric care, and patients with renal impairment. More broadly, it provides a transferable strategy for reproducible radiomics that can be adapted across neuro-oncology and other imaging domains. To our knowledge, this work is among the most systematic and comprehensive evaluations of model selection in glioma imaging, spanning 1200 pipelines across multicenter cohorts, and it sets a precedent for rigorous benchmarking in medical ML. Looking ahead, integration with explainability methods, federated learning, and prospective clinical trials will be essential for translating these pipelines into routine practice. Ultimately, the findings underscore that stability and multicenter validation are prerequisites for trustworthy clinical adoption of ML in neuro-oncology, marking a significant step toward safer, more accessible, and more reproducible precision imaging.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/inventions11010011/s1. Supplemental File S1, Sheet 1, presents the complete list of average rotational performances, rankings, and scores for all 1200 CA + DRA combinations, while Supplemental File S1, Sheet 2, presents the complete list of average five-fold cross-validation and external testing performances, rankings, and scores for all 1200 CA + DRA combinations. Supplemental Files S2–S4 provide detailed information on different rotations, including external tests with BRATS-Africa, UCSF-PDGM, and UPENN-GBM, respectively. Each file contains Sheets 1–5, which outline the specific features selected by each FSA, hyperparameters used for all ML algorithms, and the results of five-fold cross-validation, including average values and standard deviations, respectively. These files ensure transparency and reproducibility by documenting both performance outcomes and the configuration choices that guide model training and evaluation.

Author Contributions

Conceptualization, M.R.S. and S.A.; methodology, M.R.S., S.A. and M.O.; software, S.T., S.S.M., S.G. and S.D.; validation, M.O., I.H., A.R. and M.R.S.; formal analysis, M.R.S., S.G. and S.A.; investigation, S.A., S.D., S.T. and S.S.M.; resources, M.R.S., I.H., A.R. and M.O.; data curation, S.D., S.G. and S.A.; writing—original draft preparation, S.T., M.R.S., S.S.M. and S.A.; writing—review and editing, M.R.S., I.H., A.R. and M.O.; visualization, S.T., S.S.M., S.G., I.H., A.R., M.O. and M.R.S.; project administration, S.T.; funding acquisition, M.R.S., I.H. and A.R. All authors have read and agreed to the published version of the manuscript.

Funding

This study was supported by funding from the Canadian Foundation for Innovation—John R. Evans Leaders Fund (CFI-JELF; Award No. AWD-023869 CFI), as well as the Natural Sciences and Engineering Research Council of Canada (NSERC) through Awards AWD-024385, RGPIN-2023-0357, and Discovery Horizons Grant DH-2025-00119.

Institutional Review Board Statement

Ethical review and approval were waived for this study due to the use of fully anonymized, publicly available secondary data. The study was conducted in accordance with the Declaration of Helsinki.

Informed Consent Statement

Patient consent was waived due to the fact that informed consent had already been obtained by the original investigators.

Data Availability Statement

All codes and Supplemental Files supporting the findings of this study are publicly available at: https://github.com/MohammadRSalmanpour/Stability-Aware-Multicenter-Machine-Learning-for-Glioma-MRI (accessed on 5 Sepember 2025).

Acknowledgments

We gratefully acknowledge the Technological Virtual Collaboration Corporation (TECVICO CORP.) and the VirCollab group (https://vircollab.com/#/) for their support in this study.

Conflicts of Interest

Authors Mohammad R. Salmanpour and Mehrdad Oveisi were employed by the company TECVICO Corp. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ABAdaBoost
AEAAttribute Extraction Algorithm
AFTANOVA F-Test
APTANOVA p-Value Selection
BBBBlood–Brain Barrier
CAClassification Algorithm
CCCorrelation Coefficient
CEContrast Enhancement
CSTChi-Square Test
DCDummy Classifier
DRADimension Reduction Algorithm
DTIDiffusion Tensor Imaging
DWIDiffusion-Weighted Imaging
ETImExtra Trees Importance
ETrExtra Trees Classifier
FEWFeature Embedding
FDRFalse Discovery Rate
FOFirst-Order
FSAFeature Selection Algorithm
FWEFamily-Wise Error
GBGradient Boosting
GBCAGadolinium-Based Contrast Agent
GBMGlioblastoma
GLDMGray-Level Dependence Matrix
GLCMGray-Level Co-Occurrence Matrix
GLRLMGray-Level Run Length Matrix
GLSZMGray-Level Size Zone Matrix
GPGaussian Process
HGBHistGradient Boosting
IBSIImage Biomarker Standardization Initiative
ICAIndependent Component Analysis
KNNk-Nearest Neighbors
LDALinear Discriminant Analysis
LGBMLight Gradient Boosting Machine
LGGLow-Grade Glioma
LLELocally Linear Embedding
MDSMultidimensional Scaling
MIMutual Information
MLPMulti-Layer Perceptron
MRIMagnetic Resonance Imaging
NGTDMNeighborhood Gray-Tone Difference Matrix
NMFNon-negative Matrix Factorization
PCAPrincipal Component Analysis
PWIPerfusion-Weighted Imaging
RANOResponse Assessment in Neuro-Oncology
RFRadiomics Feature
RFERecursive Feature Elimination
ROC-AUCArea Under the Receiver Operating Characteristic Curve
SBSSequential Backward Selection
SDStandard Deviation
SFSSequential Forward Selection
SGDCStochastic Gradient Descent Classifier
SHAPSHapley Additive exPlanations
SPCASparse Principal Component Analysis
SVMSupport Vector Machine
T1WIT1-Weighted Imaging
TCIAThe Cancer Imaging Archive
TSVDTruncated Singular Value Decomposition
t-SNEt-distributed Stochastic Neighbor Embedding
UFSUnivariate Feature Selection
UMAPUniform Manifold Approximation and Projection
VIFVariance Inflation Factor
VTVariance Thresholding

References

  1. Wu, C.; Lin, G.; Lin, Z.; Zhang, J.; Chen, L.; Liu, S.; Tang, W.; Qiu, X.; Zhou, C. Peritumoral edema on magnetic resonance imaging predicts a poor clinical outcome in malignant glioma. Oncol. Lett. 2015, 10, 2769–2776. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Maynard, J.; Okuchi, S.; Wastling, S.; Al Busaidi, A.; Almossawi, O.; Mbatha, W.; Brandner, S.; Jaunmuktane, Z.; Murat Koc, A.; Mancini, L.; et al. World Health Organization Grade II/III Glioma Molecular Status: Prediction by MRI Morphologic Features and Apparent Diffusion Coefficient. Radiology 2020, 296, 111–121. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Comba, A.; Faisal, S.M.; Varela, M.L.; Hollon, T.; Al-Holou, W.N.; Umemura, Y.; Nunez, F.J.; Motsch, S.; Castro, M.G.; Lowenstein, P.R. Uncovering Spatiotemporal Heterogeneity of High-Grade Gliomas: From Disease Biology to Therapeutic Implications. Front. Oncol. 2021, 11, 703764. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Radhakrishnan, K.; Mokri, B.; Parisi, J.E.; O’Fallon, W.M.; Sunku, J.; Kurland, L.T. The trends in incidence of primary brain tumors in the population of Rochester, Minnesota. Ann. Neurol. 1995, 37, 67–73. [Google Scholar] [CrossRef] [Scilit]
  5. Do, C.; DeAguero, J.; Brearley, A.; Trejo, X.; Howard, T.; Escobar, G.P.; Wagner, B. Gadolinium-based contrast agent use, their safety, and practice evolution. Kidney 2020, 1, 561–568. [Google Scholar] [CrossRef] [Scilit]
  6. Won, S.E.; Suh, C.H.; Kim, S.; Park, H.J.; Kim, K.W. Summary of Key Points of the Response Assessment in Neuro-Oncology (RANO) 2.0. Korean J. Radiol. 2024, 25, 859. [Google Scholar] [CrossRef] [Scilit]
  7. Salehi, A.; Paturu, M.R.; Patel, B.; Cain, M.D.; Mahlokozera, T.; Yang, A.B.; Lin, T.-H.; Leuthardt, E.C.; Yano, H.; Song, S.-K. Therapeutic enhancement of blood–brain and blood–tumor barriers permeability by laser interstitial thermal therapy. Neuro-Oncol. Adv. 2020, 2, vdaa071. [Google Scholar] [CrossRef] [Scilit]
  8. Stoecklein, V.M.; Stoecklein, S.; Galiè, F.; Ren, J.; Schmutzer, M.; Unterrainer, M.; Albert, N.L.; Kreth, F.-W.; Thon, N.; Liebig, T. Resting-state fMRI detects alterations in whole brain connectivity related to tumor biology in glioma patients. Neuro-Oncol. 2020, 22, 1388–1398. [Google Scholar] [CrossRef] [Scilit]
  9. Stanescu, A.L.; Shaw, D.W.; Murata, N.; Murata, K.; Rutledge, J.C.; Maloney, E.; Maravilla, K.R. Brain tissue gadolinium retention in pediatric patients after contrast-enhanced magnetic resonance exams: Pathological confirmation. Pediatr. Radiol. 2020, 50, 388–396. [Google Scholar] [CrossRef] [Scilit]
  10. McDonald, R.J.; McDonald, J.S.; Kallmes, D.F.; Jentoft, M.E.; Murray, D.L.; Thielen, K.R.; Williamson, E.E.; Eckel, L.J. Intracranial gadolinium deposition after contrast-enhanced MR imaging. Radiology 2015, 275, 772–782. [Google Scholar] [CrossRef] [Scilit]
  11. Murali, S.; Ding, H.; Adedeji, F.; Qin, C.; Obungoloch, J.; Asllani, I.; Anazodo, U.; Ntusi, N.A.; Mammen, R.; Niendorf, T.; et al. Bringing MRI to low-and middle-income countries: Directions, challenges and potential solutions. NMR Biomed. 2024, 37, e4992. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Tadavarthi, Y.; Vey, B.; Krupinski, E.; Prater, A.; Gichoya, J.; Safdar, N.; Trivedi, H. The state of radiology AI: Considerations for purchase decisions and current market offerings. Radiol. Artif. Intell. 2020, 2, e200004. [Google Scholar] [CrossRef] [Scilit]
  13. Kamintsky, L.; Beyea, S.D.; Fisk, J.D.; Hashmi, J.A.; Omisade, A.; Calkin, C.; Bardouille, T.; Bowen, C.; Quraan, M.; Mitnitski, A. Blood-brain barrier leakage in systemic lupus erythematosus is associated with gray matter loss and cognitive impairment. Ann. Rheum. Dis. 2020, 79, 1580–1587. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Mallio, C.A.; Radbruch, A.; Deike-Hofmann, K.; van der Molen, A.J.; Dekkers, I.A.; Zaharchuk, G.; Parizel, P.M.; Zobel, B.B.; Quattrocchi, C.C. Artificial intelligence to reduce or eliminate the need for gadolinium-based contrast agents in brain and cardiac MRI: A literature review. Investig. Radiol. 2023, 58, 746–753. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Salmanpour, M.R.; Hosseinzadeh, M.; Modiri, E.; Akbari, A.; Hajianfar, G.; Askari, D.; Fatan, M.; Maghsudi, M.; Ghaffari, H.; Rezaeijo, S.M. Advanced survival prediction in head and neck cancer using hybrid machine learning systems and radiomics features. In Proceedings of the Medical Imaging 2022: Biomedical Applications in Molecular, Structural, and Functional Imaging, San Diego, CA, USA, 20–22 February 2022; SPIE: Bellingham, WA, USA, 2022; pp. 314–321. [Google Scholar]
  16. Hajianfar, G.; Kalayinia, S.; Hosseinzadeh, M.; Samanian, S.; Maleki, M.; Sossi, V.; Rahmim, A.; Salmanpour, M.R. Prediction of Parkinson’s disease pathogenic variants using hybrid Machine learning systems and radiomic features. Phys. Med. 2023, 113, 102647. [Google Scholar] [CrossRef] [Scilit]
  17. Fröling, E.; Rajaeean, N.; Hinrichsmeyer, K.S.; Domrös-Zoungrana, D.; Urban, J.N.; Lenz, C. Artificial intelligence in medical affairs: A new paradigm with novel opportunities. Pharm. Med. 2024, 38, 331–342. [Google Scholar] [CrossRef] [Scilit]
  18. Hosseinzadeh, M.; Gorji, A.; Fathi Jouzdani, A.; Rezaeijo, S.M.; Rahmim, A.; Salmanpour, M.R. Prediction of Cognitive Decline in Parkinson’s Disease Using Clinical and DAT SPECT Imaging Features, and Hybrid Machine Learning Systems. Diagnostics 2023, 13, 1691. [Google Scholar] [CrossRef] [Scilit]
  19. Salmanpour, M.R.; Shamsaei, M.; Hajianfar, G.; Soltanian-Zadeh, H.; Rahmim, A. Longitudinal clustering analysis and prediction of Parkinson’s disease progression using radiomics and hybrid machine learning. Quant. Imaging Med. Surg. 2022, 12, 906. [Google Scholar] [CrossRef] [Scilit]
  20. Salmanpour, M.R.; Shamsaei, M.; Saberi, A.; Hajianfar, G.; Soltanian-Zadeh, H.; Rahmim, A. Robust identification of Parkinson’s disease subtypes using radiomics and hybrid machine learning. Comput. Biol. Med. 2021, 129, 104142. [Google Scholar] [CrossRef] [Scilit]
  21. Granzier, R.W.; Verbakel, N.M.; Ibrahim, A.; Van Timmeren, J.; Van Nijnatten, T.; Leijenaar, R.; Lobbes, M.; Smidt, M.; Woodruff, H. MRI-based radiomics in breast cancer: Feature robustness with respect to inter-observer segmentation variability. Sci. Rep. 2020, 10, 14163. [Google Scholar] [CrossRef] [Scilit]
  22. Wang, W.; Wang, Y.; Meng, W.; Guo, E.; He, H.; Huang, G.; He, W.; Wu, Y. Prediction of Glioma enhancement pattern using a MRI radiomics-based model. Medicine 2024, 103, e39512. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Mukherkjee, D.; Saha, P.; Kaplun, D.; Sinitca, A.; Sarkar, R. Brain tumor image generation using an aggregation of GAN models with style transfer. Sci. Rep. 2022, 12, 9141. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Salmanpour, M.R.; Shamsaei, M.; Rahmim, A. Feature selection and machine learning methods for optimal identification and prediction of subtypes in Parkinson’s disease. Comput. Methods Programs Biomed. 2021, 206, 106131. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Varoquaux, G.; Cheplygina, V. Machine learning for medical imaging: Methodological failures and recommendations for the future. npj Digit. Med. 2022, 5, 48. [Google Scholar] [CrossRef] [Scilit]
  26. Salmanpour, M.R.; Hosseinzadeh, M.; Sanati, N.; Jouzdani, A.F.; Gorji, A.; Mahboubisarighieh, A.; Maghsudi, M.; Rezaeijo, S.M.; Moore, S.; Bonnie, L. Tensor deep versus radiomics features: Lung cancer outcome prediction using hybrid machine learning systems. J. Nucl. Med. 2023, 64, P1174. [Google Scholar]
  27. Sarma, K.V.; Harmon, S.; Sanford, T.; Roth, H.R.; Xu, Z.; Tetreault, J.; Flores, D.X.M.G.; Raman, A.G.; Kulkarni, R.; Wood, B.J.; et al. Federated learning improves site performance in multicenter deep learning without data sharing. J. Am. Med. Inform. Assoc. 2021, 28, 1259–1264. [Google Scholar] [CrossRef] [Scilit]
  28. Yang, J.; Soltan, A.A.S.; Clifton, D.A. Machine learning generalizability across healthcare settings: Insights from multi-site COVID-19 screening. npj Digit. Med. 2022, 5, 69. [Google Scholar] [CrossRef] [Scilit]
  29. Linardos, A.; Kushibar, K.; Walsh, S.; Gkontra, P.; Lekadir, K. Federated learning for multi-center imaging diagnostics: A simulation study in cardiovascular disease. Sci. Rep. 2022, 12, 3551. [Google Scholar] [CrossRef] [Scilit]
  30. Calabrese, E.; Villanueva-Meyer, J.; Rudie, J.; Rauschecker, A.M.; Baid, U.; Bakas, S.; Cha, S.; Mongan, J.T.; Hess, C.P. The University of California San Francisco Preoperative Diffuse Glioma MRI (UCSF-PDGM). Radiol. Artif. Intell. 2022, 4, e220058. [Google Scholar] [CrossRef] [Scilit]
  31. Bakas, S.; Sako, C.; Akbari, H.; Bilello, M.; Sotiras, A.; Shukla, G.; Rudie, J.D.; Santamaría, N.F.; Kazerooni, A.F.; Pati, S. The University of Pennsylvania glioblastoma (UPenn-GBM) cohort: Advanced MRI, clinical, genomics, & radiomics. Sci. Data 2022, 9, 453. [Google Scholar] [CrossRef] [Scilit]
  32. Adewole, M.; Rudie, J.D.; Gbadamosi, A.; Zhang, D.; Raymond, C.; Ajigbotoshso, J.; Toyobo, O.; Aguh, K.; Omidiji, O.; Akinola, R. The BraTS-Africa Dataset: Expanding the Brain Tumor Segmentation (BraTS) Data to Capture African Populations. Radiol. Artif. Intell. 2025, 7, e240528. [Google Scholar] [CrossRef] [Scilit]
  33. Bakas, S.; Akbari, H.; Sotiras, A.; Bilello, M.; Rozycki, M.; Kirby, J.; Freymann, J.; Farahani, K.; Davatzikos, C. Segmentation Labels for the Pre-Operative Scans of the TCGA-LGG Collection; The Cancer Imaging Archive: Online, 2017. [Google Scholar]
  34. van Griethuysen, J.J.M.; Fedorov, A.; Parmar, C.; Hosny, A.; Aucoin, N.; Narayan, V.; Beets-Tan, R.G.; Fillion-Robin, J.C.; Pieper, S.; Aerts, H.J. Computational Radiomics System to Decode the Radiographic Phenotype. Cancer Res. 2017, 77, e104–e107. [Google Scholar] [CrossRef] [Scilit]
  35. Wilimitis, D.; Walsh, C.G. Practical considerations and applied examples of cross-validation for model development and evaluation in health care: Tutorial. JMIR AI 2023, 2, e49023. [Google Scholar] [CrossRef] [Scilit]
  36. Gallitto, G.; Englert, R.; Kincses, B.; Kotikalapudi, R.; Li, J.; Hoffschlag, K.; Bingel, U.; Spisak, T. External validation of machine learning models—Registered models and adaptive sample splitting. GigaScience 2025, 14, giaf036. [Google Scholar] [CrossRef] [Scilit]
Table 2. Model evaluation pipeline with three-fold rotational validation, five-fold cross-validation. Metrics are aggregated into a composite score for systematic model ranking and selection. The Top 10 performing dimension reduction–classifier pairs for glioma outcome prediction are shown across different datasets. CA: Classification Algorithm, DRA: Dimension Reduction Algorithm, ROC-AUC: The area under the receiver operating characteristic curve, MI: Mutual Information, FEW: Feature Embedding, ETIm: Extra Trees Importance, AFT: ANOVA F-Test, RFE: Recursive Feature Elimination, UFS: Univariate Feature Selection, APT: ANOVA p-Value Selection, TSVD: Truncated Singular Value Decomposition, ETr: Extra Trees Classifier, GP: Gaussian Process, LGBM: Light Gradient Boosting Machine, MLP: Multilayer Perceptron, XGB: Extreme Gradient Boosting.
Table 2. Model evaluation pipeline with three-fold rotational validation, five-fold cross-validation. Metrics are aggregated into a composite score for systematic model ranking and selection. The Top 10 performing dimension reduction–classifier pairs for glioma outcome prediction are shown across different datasets. CA: Classification Algorithm, DRA: Dimension Reduction Algorithm, ROC-AUC: The area under the receiver operating characteristic curve, MI: Mutual Information, FEW: Feature Embedding, ETIm: Extra Trees Importance, AFT: ANOVA F-Test, RFE: Recursive Feature Elimination, UFS: Univariate Feature Selection, APT: ANOVA p-Value Selection, TSVD: Truncated Singular Value Decomposition, ETr: Extra Trees Classifier, GP: Gaussian Process, LGBM: Light Gradient Boosting Machine, MLP: Multilayer Perceptron, XGB: Extreme Gradient Boosting.
Five-Fold Cross Validation
DRA + CARankScoreAccuracyF1 ScorePrecisionRecallROC-AUC
MI + ETr10.9413810.94 ± 0.020.92 ± 0.020.94 ± 0.020.93 ± 0.020.81 ± 0.10
FEW + ETr20.9378600.94 ± 0.020.93 ± 0.020.94 ± 0.020.93 ± 0.020.78 ± 0.13
ETIm + GP30.9345100.94 ± 0.030.93 ± 0.010.94 ± 0.030.92 ± 0.030.82 ± 0.03
AFT + LGBM40.9340900.94 ± 0.030.93 ± 0.020.94 ± 0.030.93 ± 0.020.79 ± 0.11
RFE_MLP50.9332450.94 ± 0.030.92 ± 0.020.94 ± 0.030.93 ± 0.020.79 ± 0.14
FEW + LGBM60.9330740.94 ± 0.030.93 ± 0.020.94 ± 0.030.93 ± 0.030.79 ± 0.12
UFS + ETr70.9330290.94 ± 0.020.92 ± 0.020.94 ± 0.020.93 ± 0.020.80 ± 0.11
APT + XGB80.9329310.94 ± 0.030.92 ± 0.020.94 ± 0.030.93 ± 0.020.77 ± 0.14
RFE + ETr90.9328440.94 ± 0.030.93 ± 0.020.94 ± 0.030.93 ± 0.030.81 ± 0.09
TSVD + ETr100.9327620.94 ± 0.030.93 ± 0.020.94 ± 0.030.92 ± 0.030.77 ± 0.09
Table 3. Model evaluation pipeline with three-fold rotational validation and held-out external testing. Metrics are aggregated into a composite score for systematic model ranking and selection. The Top 10 performing dimension reduction–classifier pairs for glioma outcome prediction are shown across different datasets. CA: Classification Algorithm, DRA: Dimension Reduction Algorithm, ROC-AUC: The area under the receiver operating characteristic curve, MI: Mutual Information, FEW: Feature Embedding, ETIm: Extra Trees Importance, AFT: ANOVA F-Test, RFE: Recursive Feature Elimination, UFS: Univariate Feature Selection, APT: ANOVA p-Value Selection, TSVD: Truncated Singular Value Decomposition, ETr: Extra Trees Classifier, GP: Gaussian Process, LGBM: Light Gradient Boosting Machine, MLP: Multilayer Perceptron, XGB: Extreme Gradient Boosting.
Table 3. Model evaluation pipeline with three-fold rotational validation and held-out external testing. Metrics are aggregated into a composite score for systematic model ranking and selection. The Top 10 performing dimension reduction–classifier pairs for glioma outcome prediction are shown across different datasets. CA: Classification Algorithm, DRA: Dimension Reduction Algorithm, ROC-AUC: The area under the receiver operating characteristic curve, MI: Mutual Information, FEW: Feature Embedding, ETIm: Extra Trees Importance, AFT: ANOVA F-Test, RFE: Recursive Feature Elimination, UFS: Univariate Feature Selection, APT: ANOVA p-Value Selection, TSVD: Truncated Singular Value Decomposition, ETr: Extra Trees Classifier, GP: Gaussian Process, LGBM: Light Gradient Boosting Machine, MLP: Multilayer Perceptron, XGB: Extreme Gradient Boosting.
External Test
DRA + CARankScoreAccuracyF1 ScorePrecisionRecallROC-AUC
MI + ETr10.9413810.93 ± 0.060.91 ± 0.060.93 ± 0.060.91 ± 0.080.70 ± 0.13
FEW + ETr20.9378600.93 ± 0.060.89 ± 0.120.93 ± 0.060.91 ± 0.090.69 ± 0.06
ETIm + GP30.9345100.93 ± 0.060.88 ± 0.110.93 ± 0.060.91 ± 0.090.77 ± 0.10
AFT + LGBM40.9340900.93 ± 0.060.90 ± 0.100.93 ± 0.060.91 ± 0.090.73 ± 0.07
RFE_MLP50.9332450.94 ± 0.060.89 ± 0.120.94 ± 0.060.91 ± 0.090.50 ± 0.29
FEW + LGBM60.9330740.93 ± 0.060.90 ± 0.100.93 ± 0.060.91 ± 0.090.74 ± 0.06
UFS + ETr70.9330290.93 ± 0.060.92 ± 0.050.93 ± 0.060.91 ± 0.080.70 ± 0.12
APT + XGB80.9329310.93 ± 0.060.89 ± 0.120.93 ± 0.060.91 ± 0.090.73 ± 0.06
RFE + ETr90.9328440.93 ± 0.060.91 ± 0.060.93 ± 0.060.91 ± 0.080.70 ± 0.10
TSVD + ETr100.9327620.93 ± 0.060.89 ± 0.120.93 ± 0.060.91 ± 0.090.74 ± 0.05
Table 4. The Top 10 performing dimension reduction–classifier (DRA–CA) pairs for glioma outcome prediction, trained on UPENN-GB, BRATS Africa, BRATS TCGA LGG, with UCSF PDGM used exclusively for external testing. ROC-AUC: The area under the receiver operating characteristic curve, MI: Mutual Information, FEW: Feature Embedding, ETIm: Extra Trees Importance, AFT: ANOVA F-Test, RFE: Recursive Feature Elimination, UFS: Univariate Feature Selection, APT: ANOVA p-Value Selection, TSVD: Truncated Singular Value Decomposition, ETr: Extra Trees Classifier, GP: Gaussian Process, LGBM: Light Gradient Boosting Machine, MLP: Multilayer Perceptron, XGB: Extreme Gradient Boosting.
Table 4. The Top 10 performing dimension reduction–classifier (DRA–CA) pairs for glioma outcome prediction, trained on UPENN-GB, BRATS Africa, BRATS TCGA LGG, with UCSF PDGM used exclusively for external testing. ROC-AUC: The area under the receiver operating characteristic curve, MI: Mutual Information, FEW: Feature Embedding, ETIm: Extra Trees Importance, AFT: ANOVA F-Test, RFE: Recursive Feature Elimination, UFS: Univariate Feature Selection, APT: ANOVA p-Value Selection, TSVD: Truncated Singular Value Decomposition, ETr: Extra Trees Classifier, GP: Gaussian Process, LGBM: Light Gradient Boosting Machine, MLP: Multilayer Perceptron, XGB: Extreme Gradient Boosting.
Five-Fold Cross Validation by UCSF PDGM, Africa, BRATS TCGA LGGExternal Test by UPENN-GB
ClassifierScoreRankAccuracyF1 ScorePrecisionRecallROC-AUCAccuracyF1 ScorePrecisionRecallROC-AUC
MI + ETr0.94138110.96 ± 0.010.93 ± 0.010.96 ± 0.010.95 ± 0.010.70 ± 0.020.87 ± 0.000.85 ± 0.010.87 ± 0.000.82 ± 0.010.81 ± 0.04
FEW + ETr0.9378620.96 ± 0.010.94 ± 0.020.96 ± 0.010.95 ± 0.010.64 ± 0.020.86 ± 0.000.75 ± 0.000.86 ± 0.000.80 ± 0.000.69 ± 0.07
ETIm + GP0.9345130.96 ± 0.000.93 ± 0.010.96 ± 0.000.95 ± 0.000.79 ± 0.040.87 ± 0.000.75 ± 0.000.87 ± 0.000.80 ± 0.000.88 ± 0.01
AFT + LGBM0.9340940.96 ± 0.000.93 ± 0.010.96 ± 0.000.95 ± 0.010.66 ± 0.040.87 ± 0.000.79 ± 0.050.87 ± 0.000.81 ± 0.010.81 ± 0.07
RFE_MLP0.93324550.96 ± 0.000.93 ± 0.010.96 ± 0.000.95 ± 0.000.62 ± 0.040.87 ± 0.000.75 ± 0.000.87 ± 0.000.80 ± 0.000.16 ± 0.00
FEW + LGBM0.93307460.96 ± 0.000.94 ± 0.020.96 ± 0.000.95 ± 0.010.66 ± 0.040.87 ± 0.000.79 ± 0.050.87 ± 0.000.81 ± 0.010.81 ± 0.06
UFS + ETr0.93302970.96 ± 0.010.93 ± 0.010.96 ± 0.010.94 ± 0.010.68 ± 0.040.87 ± 0.000.87 ± 0.030.87 ± 0.000.82 ± 0.010.81 ± 0.04
APT + XGB0.93293180.96 ± 0.000.93 ± 0.010.96 ± 0.000.95 ± 0.000.61 ± 0.040.87 ± 0.000.75 ± 0.000.87 ± 0.000.80 ± 0.000.79 ± 0.05
RFE + ETr0.93284490.96 ± 0.000.94 ± 0.020.96 ± 0.000.95 ± 0.000.71 ± 0.060.87 ± 0.000.86 ± 0.010.87 ± 0.000.83 ± 0.010.79 ± 0.05
TSVD + ETr0.932762100.96 ± 0.000.93 ± 0.000.96 ± 0.000.94 ± 0.000.67 ± 0.040.87 ± 0.000.75 ± 0.000.87 ± 0.000.80 ± 0.000.79 ± 0.04
Table 5. The Top 10 performing dimension reduction–classifier (DRA–CA) pairs for glioma outcome prediction, trained on UCSF PDGM, UPENN-GB, BRATS TCGA LGG, with BRATS Africa used exclusively for external testing. ROC-AUC: The area under the receiver operating characteristic curve, MI: Mutual Information, FEW: Feature Embedding, ETIm: Extra Trees Importance, AFT: ANOVA F-Test, RFE: Recursive Feature Elimination, UFS: Univariate Feature Selection, APT: ANOVA p-Value Selection, TSVD: Truncated Singular Value Decomposition, ETr: Extra Trees Classifier, GP: Gaussian Process, LGBM: Light Gradient Boosting Machine, MLP: Multilayer Perceptron, XGB: Extreme Gradient Boosting.
Table 5. The Top 10 performing dimension reduction–classifier (DRA–CA) pairs for glioma outcome prediction, trained on UCSF PDGM, UPENN-GB, BRATS TCGA LGG, with BRATS Africa used exclusively for external testing. ROC-AUC: The area under the receiver operating characteristic curve, MI: Mutual Information, FEW: Feature Embedding, ETIm: Extra Trees Importance, AFT: ANOVA F-Test, RFE: Recursive Feature Elimination, UFS: Univariate Feature Selection, APT: ANOVA p-Value Selection, TSVD: Truncated Singular Value Decomposition, ETr: Extra Trees Classifier, GP: Gaussian Process, LGBM: Light Gradient Boosting Machine, MLP: Multilayer Perceptron, XGB: Extreme Gradient Boosting.
Five-Fold Cross Validation by UCSF PDGM, Africa, BRATS TCGA LGGExternal Test by UPENN-GB
ClassifierScoreRankAccuracyF1 ScorePrecisionRecallROC-AUCAccuracyF1 ScorePrecisionRecallROC-AUC
MI + ETr0.94138110.91 ± 0.020.90 ± 0.030.91 ± 0.020.90 ± 0.030.84 ± 0.050.98 ± 0.000.97 ± 0.000.98 ± 0.000.98 ± 0.000.73 ± 0.05
FEW + ETr0.9378620.91 ± 0.020.91 ± 0.020.91 ± 0.020.90 ± 0.020.82 ± 0.050.98 ± 0.000.97 ± 0.000.98 ± 0.000.98 ± 0.000.74 ± 0.07
ETIm + GP0.9345130.91 ± 0.010.91 ± 0.020.91 ± 0.010.90 ± 0.020.83 ± 0.070.98 ± 0.000.97 ± 0.000.98 ± 0.000.98 ± 0.000.72 ± 0.01
AFT + LGBM0.9340940.91 ± 0.010.90 ± 0.020.91 ± 0.010.90 ± 0.020.82 ± 0.050.98 ± 0.000.97 ± 0.000.98 ± 0.000.98 ± 0.000.72 ± 0.07
RFE_MLP0.93324550.91 ± 0.030.90 ± 0.040.91 ± 0.030.90 ± 0.040.86 ± 0.040.98 ± 0.000.97 ± 0.000.98 ± 0.000.98 ± 0.000.65 ± 0.05
FEW + LGBM0.93307460.91 ± 0.010.90 ± 0.020.91 ± 0.010.90 ± 0.020.82 ± 0.050.98 ± 0.000.97 ± 0.000.98 ± 0.000.98 ± 0.000.72 ± 0.07
UFS + ETr0.93302970.91 ± 0.020.90 ± 0.030.91 ± 0.020.90 ± 0.030.84 ± 0.050.98 ± 0.000.97 ± 0.000.98 ± 0.000.98 ± 0.000.73 ± 0.05
APT + XGB0.93293180.91 ± 0.020.90 ± 0.020.91 ± 0.020.90 ± 0.020.81 ± 0.050.98 ± 0.000.97 ± 0.000.98 ± 0.000.97 ± 0.000.72 ± 0.03
RFE + ETr0.93284490.91 ± 0.020.90 ± 0.030.91 ± 0.020.90 ± 0.030.84 ± 0.040.98 ± 0.000.97 ± 0.000.98 ± 0.000.98 ± 0.000.73 ± 0.04
TSVD + ETr0.932762100.91 ± 0.010.91 ± 0.020.91 ± 0.010.89 ± 0.020.80 ± 0.040.98 ± 0.000.97 ± 0.000.98 ± 0.000.98 ± 0.000.73 ± 0.05
Table 6. The Top 10 performing dimension reduction–classifier (DRA–CA) pairs for glioma outcome prediction, trained on UCSF-PDGM, UPENN-GB, and BRATS-TCGA-LGG, with BRATS-Africa used exclusively for external testing. CA: Classification Algorithm, DRA: Dimension Reduction Algorithm, ROC-AUC: The area under the receiver operating characteristic curve, MI: Mutual Information, FEW: Feature Embedding, ETIm: Extra Trees Importance, AFT: ANOVA F-Test, RFE: Recursive Feature Elimination, UFS: Univariate Feature Selection, APT: ANOVA p-Value Selection, TSVD: Truncated Singular Value Decomposition, ETr: Extra Trees Classifier, GP: Gaussian Process, LGBM: Light Gradient Boosting Machine, MLP: Multilayer Perceptron, XGB: Extreme Gradient Boosting.
Table 6. The Top 10 performing dimension reduction–classifier (DRA–CA) pairs for glioma outcome prediction, trained on UCSF-PDGM, UPENN-GB, and BRATS-TCGA-LGG, with BRATS-Africa used exclusively for external testing. CA: Classification Algorithm, DRA: Dimension Reduction Algorithm, ROC-AUC: The area under the receiver operating characteristic curve, MI: Mutual Information, FEW: Feature Embedding, ETIm: Extra Trees Importance, AFT: ANOVA F-Test, RFE: Recursive Feature Elimination, UFS: Univariate Feature Selection, APT: ANOVA p-Value Selection, TSVD: Truncated Singular Value Decomposition, ETr: Extra Trees Classifier, GP: Gaussian Process, LGBM: Light Gradient Boosting Machine, MLP: Multilayer Perceptron, XGB: Extreme Gradient Boosting.
Five-Fold Cross Validation by UCSF PDGM, UPENN-GB, BRATS TCGA LGGExternal Test by BRATS Africa
DRA + CARankScoreAccuracyF1 ScorePrecisionRecallROC-AUCAccuracyF1 ScorePrecisionRecallROC-AUC
MI + ETr10.9413810.94 ± 0.000.94 ± 0.000.94 ± 0.000.94 ± 0.000.89 ± 0.030.95 ± 0.000.92 ± 0.000.95 ± 0.000.93 ± 0.000.56 ± 0.03
FEW + ETr20.937860.94 ± 0.000.94 ± 0.010.94 ± 0.000.94 ± 0.010.89 ± 0.020.95 ± 0.010.94 ± 0.010.95 ± 0.010.94 ± 0.010.63 ± 0.03
ETIm + GP30.934510.94 ± 0.010.94 ± 0.020.94 ± 0.010.93 ± 0.010.85 ± 0.020.95 ± 0.000.92 ± 0.000.95 ± 0.000.93 ± 0.000.71 ± 0.01
AFT + LGBM40.934090.94 ± 0.010.94 ± 0.010.94 ± 0.010.94 ± 0.010.88 ± 0.020.95 ± 0.010.95 ± 0.010.95 ± 0.010.95 ± 0.010.67 ± 0.04
RFE_MLP50.9332450.95 ± 0.010.94 ± 0.010.95 ± 0.010.94 ± 0.010.88 ± 0.010.96 ± 0.010.96 ± 0.010.96 ± 0.010.95 ± 0.000.69 ± 0.01
FEW + LGBM60.9330740.94 ± 0.010.94 ± 0.010.94 ± 0.010.93 ± 0.010.88 ± 0.020.95 ± 0.020.94 ± 0.010.95 ± 0.020.94 ± 0.010.69 ± 0.07
UFS + ETr70.9330290.95 ± 0.000.94 ± 0.000.95 ± 0.000.94 ± 0.000.89 ± 0.040.95 ± 0.000.92 ± 0.000.95 ± 0.000.93 ± 0.000.56 ± 0.02
APT + XGB80.9329310.94 ± 0.000.94 ± 0.010.94 ± 0.000.94 ± 0.000.88 ± 0.020.95 ± 0.010.95 ± 0.000.95 ± 0.010.95 ± 0.010.68 ± 0.03
RFE + ETr90.9328440.95 ± 0.010.94 ± 0.010.95 ± 0.010.94 ± 0.010.88 ± 0.030.95 ± 0.000.92 ± 0.000.95 ± 0.000.93 ± 0.000.60 ± 0.03
TSVD + ETr100.9327620.95 ± 0.010.94 ± 0.010.95 ± 0.010.94 ± 0.010.83 ± 0.020.95 ± 0.000.94 ± 0.000.95 ± 0.000.94 ± 0.000.69 ± 0.05
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Amiri, S.; Taeb, S.; Gharibi, S.; Dehghanfard, S.; Mehrnia, S.S.; Oveisi, M.; Hacihaliloglu, I.; Rahmim, A.; Salmanpour, M.R. Enhancement Without Contrast: Stability-Aware Multicenter Machine Learning for Glioma MRI Imaging. Inventions 2026, 11, 11. https://doi.org/10.3390/inventions11010011

AMA Style

Amiri S, Taeb S, Gharibi S, Dehghanfard S, Mehrnia SS, Oveisi M, Hacihaliloglu I, Rahmim A, Salmanpour MR. Enhancement Without Contrast: Stability-Aware Multicenter Machine Learning for Glioma MRI Imaging. Inventions. 2026; 11(1):11. https://doi.org/10.3390/inventions11010011

Chicago/Turabian Style

Amiri, Sajad, Shahram Taeb, Sara Gharibi, Setareh Dehghanfard, Somayeh Sadat Mehrnia, Mehrdad Oveisi, Ilker Hacihaliloglu, Arman Rahmim, and Mohammad R. Salmanpour. 2026. "Enhancement Without Contrast: Stability-Aware Multicenter Machine Learning for Glioma MRI Imaging" Inventions 11, no. 1: 11. https://doi.org/10.3390/inventions11010011

APA Style

Amiri, S., Taeb, S., Gharibi, S., Dehghanfard, S., Mehrnia, S. S., Oveisi, M., Hacihaliloglu, I., Rahmim, A., & Salmanpour, M. R. (2026). Enhancement Without Contrast: Stability-Aware Multicenter Machine Learning for Glioma MRI Imaging. Inventions, 11(1), 11. https://doi.org/10.3390/inventions11010011

Article Metrics

Back to TopTop