Next Article in Journal
Diagnosis of Extensive DCIS with Progressive Microcalcifications over a Period of 11 Years
Next Article in Special Issue
Feasibility Study of Deep Learning-Driven Image Restoration for Fast, Low-Count [18F]FP-CIT Digital PET/CT
Previous Article in Journal
Retrospective Review of Head Injury: Are Pre-Injury Anticoagulant and Antiplatelet Therapy Associated with Acute Traumatic Intracranial Hemorrhage?
Previous Article in Special Issue
Decoding the Myocardium: Tracer-Aware Deep Learning for Patient-Level Classification in Stress–Rest SPECT Myocardial Perfusion Imaging
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Distinguishing Benign from Malignant Small Enhancing Breast Lesions Using Radiomics

by
Parlyn Hatch
1,
Christopher Louviere
2,
Mutlu Mete
3,
Oswaldo A. Guevara Tirado
2,
Ruben G. Ortiz Cordero
2,
Hector Diaz De Villegas
1 and
Kazim Z. Gumus
1,*
1
Department of Radiology, College of Medicine-Jacksonville, University of Florida, 655 West 8th Street, C506, Jacksonville, FL 32209, USA
2
Department of Radiology, University of Florida Health Jacksonville Physicians Inc., 655 West 8th Street, C506, Jacksonville, FL 32209, USA
3
Department of Information Science, University of North Texas, 620 Central Avenue, Denton, TX 76203, USA
*
Author to whom correspondence should be addressed.
Diagnostics 2026, 16(17), 2818; https://doi.org/10.3390/diagnostics16172818
Submission received: 30 June 2026 / Revised: 25 August 2026 / Accepted: 31 August 2026 / Published: 2 September 2026

Abstract

Background/Objectives: Small enhancing lesions (often termed foci if lesion ≤ 5 mm) observed on breast magnetic resonance imaging (MRI) pose a persistent challenge. Their clinical relevance and the most effective management strategy remain unclear, often leading to unnecessary biopsies. We sought to uncover whether quantitative radiomic features of these lesions could reliably discriminate malignant ones from benign ones. Methods: In this single-center retrospective study (2015–2024), we analyzed 31 contrast-enhancing breast lesions with a maximum diameter ≤ 10 mm measured on a single axial MRI slice from 30 patients who underwent 1.5-T or 3.0-T breast MRI followed by histologic confirmation (10 malignant, 21 benign). Lesions were manually delineated, and 105 radiomic variables were extracted from early-phase dynamic contrast-enhanced (DCE) subtraction images. Feature importance was quantified with a Random Forest model; iterative top-k pruning found the highest-performing variables. A multilayer perceptron (MLP) classifier was trained and validated using a leave-one-subject-out cross-validation (LOSO-CV) scheme, with the sensitivity, specificity, accuracy, and AUC as the primary performance metric. Results: Of the 105 extracted features, 44 carried predictive information (feature importance score ≥ 0.20). Progressive feature reduction yielded an optimal subset of nine radiomic features (six texture-based and three shape-based). Texture-derived features predominated among the most informative variables. The MLP achieved a sensitivity of 0.80, specificity of 0.89, accuracy of 0.86, and AUC of 0.82. Conclusions: In this exploratory study, a preliminary nine-feature radiomic signature extracted from early-phase DCE-subtraction images has the potential to distinguish benign from malignant small breast lesions when used with an MLP model. Prospective validation is needed before the model can influence clinical decision-making.

1. Introduction

Breast cancer is among the leading causes of cancer-related death in women, with U.S. incidence rising 1% between 2020 and 2021 according to the Centers for Disease Control and Prevention [1]. To standardize reporting and improve communication among healthcare professionals, the American College of Radiology (ACR) developed the Breast Imaging Reporting and Data System (BI-RADS) [2]. Since the implementation of BI-RADS in 1993, magnetic resonance imaging (MRI) has become a vital imaging tool for breast cancer diagnosis, providing a diagnostic accuracy of 86.9% [3,4]. However, studies report false-negative breast MRI results of approximately 4–7% in preoperative MRI-pathology correlation studies [5,6]. Although biopsy remains the gold-standard for definitive diagnosis of breast cancer, a unique advantage of MRI is that it allows for the detection of small enhancing lesions, including those previously defined by ACR BI-RADS as foci (≤5 mm) [7,8,9,10]. Although these small lesions are often associated with increased hormonal stimulation, they can also represent a developing malignancy [9]. Because malignancy rates of small lesions are variable (ranging from 0.6 to 23%), their clinical significance and optimal management remain unclear [9]. Early detection of malignant small lesions is of particular interest in high-risk women, where timely identification could lead to meaningful improvements in management and outcomes.
Radiomics, the extraction of mathematical features from medical images through a data characterization algorithm, could help address the uncertainty surrounding the interpretation of small enhancing lesions [11,12,13]. Selected features can be reduced to the most predictive and used to train a machine learning (ML) model to classify breast cancer lesions and estimate the probability of malignancy [7,13,14,15,16,17].
In this preliminary study, we explored radiomic features of breast lesions ≤ 1 cm in size from contrast-enhanced T1-weighted images. We hypothesized that an ML model could be trained to classify small enhancing breast lesions using the selected features and accurately predict malignancy.

2. Material and Methods

2.1. Patient Selection

We obtained Institutional Review Board (IRB) approval for this retrospective study (IRB approval number: IRB202401789; date of approval: 6 February 2025). The requirement for written informed consent was waived by the IRB given the retrospective design of the study. We reviewed the medical records of adult female patients who underwent multiparametric breast MRI using a 1.5-T or 3.0T scanner at our institution and subsequently had a breast biopsy from 2015 through 2024. A total of 30 female patients with 31 enhancing breast lesions were selected based on the MRI reports and pathology results. To ensure all lesions would be ≤10 mm in size, patients were included if their “impressions” stated that a breast cancer focus/foci was visualized. A flowchart summarizing the cohort selection process is presented in Figure 1. Lesion size was measured in millimeters (mm) utilizing the segmented area in the imaging data and ITK-SNAP software (version 3.8.0). A two-sample t-test was performed to compare the size measurements.

2.2. MRI Protocol

Images were acquired on a 1.5-T scanner (Siemens Espree) or 3-T scanners (Siemens TrioTim, Siemens Vida, and Siemens Skyra, Siemens, Erlangen, Germany). Patients were imaged in the prone position with a dedicated multi-channel phased-array breast coil per institutional standards. The imaging protocol included axial T1-weighted, axial T2-STIR, and an axial T1-weighted dynamic contrast-enhanced (DCE) images with six phases: one pre-contrast and five post-contrasts. Gadoteric acid (Dotarem®, Guerbet, Princeton, NJ, USA) was used as a contrast agent at a dose of 0.2 mL/kg with an injection rate of 2 mL/s. DCE scan parameters obtained for both fields are shown in Table 1.

2.3. Segmentation

DCE images were sent to DynaCAD (version 5.0) (Invivo, Gainesville, FL, USA) software, where pre-contrast images were subtracted from post-contrast images to create subtraction images. The first subtraction image (first-phase post-contrast image minus pre-contrast image) was used for radiomic analysis. The DICOM files were then de-identified, extracted from the Visage Picture Archive and Communication System (PACS) software, imported into ITK-SNAP, and converted to Neuroimaging Informatics Technology Initiative (NIfTI) format for segmentation [18]. Lesions were manually segmented slice-by-slice by two trained radiology research fellows in consensus and confirmed by an MR physicist and a breast radiologist. The resulting segmentation masks and original image datasets were then used to extract radiomic features and train a machine learning (ML) algorithm.

2.4. Experimental Design

The unit of analysis in this study was the image of the lesion. One patient contributed two lesions, one benign and one malignant, which were entered as separate samples. Because classification relied exclusively on radiomic features computed from the voxels of each individual lesion, and no patient-level or demographic variables were included in the feature set; treating these two lesions as independent observations does not transfer shared patient information across folds. In the context of lesion classification, the presence of breast cancer was treated as a positive class. Classification models were trained and validated using a leave-one-subject-out (LOSO) cross-validation strategy, in which a lesion is excluded from the dataset in each iteration and used as the test set while the model is trained on the data of all remaining lesions. LOSO was chosen due to its advantage of minimizing subject-specific bias, ensuring that the classifier’s performance generalizes well across lesions. Each experiment was repeated 30 times using different random seeds. The primary evaluation objective is to maximize the GeoMean of sensitivity (true positive rate) and specificity (true negative rate), given as S e n =   TP TP   +   FN , S p e   =   TN TN   +   FP , and G e o M e a n   =   S e n S p e . Geometric mean is useful for imbalanced datasets, as it ensures both sensitivity and specificity are high, preventing a model from simply maximizing accuracy by predicting the majority class.

2.5. Feature Selection

Using the PyRadiomics library (version 3.1.0) [19], first, we normalized the signal intensities on the DCE images to eliminate scan-to-scan variations. In-plane resampling was performed on the images (0.6 × 0.6 mm) to homogenize image processing findings. Then, we extracted 105 radiomic features from each 3D lesion per imaging sequence. Radiomics features were of the following feature classes: first order (n = 18), shape (n = 14), and texture (n = 73) including Grey Level Co-occurrence Matrix (GLCM, n = 22), Grey Level Size Zone Matrix (GLSZM, n = 16), Grey Level Run Length Matrix (GLRLM, n = 16), Grey Level Dependence Matrix (GLDM, n = 14), and Neighboring Gray Tone Difference Matrix (NGTDM, n = 5). Most of the features are in line with the Imaging Biomarker Standardization Initiative (IBSI) [20].
Feature selection and ranking are critical steps in high-dimensional radiomic feature classification. Although a comprehensive feature set can encapsulate rich spatial dynamics, it also increases the risk of overfitting, redundancy, and computational inefficiency. Therefore, identifying a subset of the most informative and discriminative features is essential for enhancing model performance and interpretability.
We employed the Feature Importance Score (FIS) of the Random Forest (RF) Algorithm to rank all 105 features. The RF is an ensemble of decision trees that aggregates their predictions and quantifies feature importance from the reduction in node impurity attributable to each feature [21]. Node impurity is measured by the Gini impurity, defined for a node t as G ( t ) = 1 i = 1 C p i 2 , where pi is the proportion of samples of class i among the C classes at that node. It equals zero for a pure node containing a single class and increases as the classes become more evenly mixed, and it can be read as the probability of misclassifying a randomly chosen sample labeled according to the node’s class distribution [22]. The importance of a feature is the total Gini impurity reduction across all splits on that feature, weighted by the sample count at each node and averaged over the trees, as implemented in Scikit-learn [23]. The FIS is the Gini-based mean decrease in impurity, averaged over the 64 trees of each Random Forest, and normalized to sum to one across the 105 features per run. The score was computed over 16 random seeds and aggregated by summation across runs, so the reported values lie on a 0 to 16 scale, with the most informative feature reaching 1.81; the corresponding mean per run values are obtained by dividing by 16. The 0.20 cutoff is a descriptive label for features carrying non-negligible information and was not used to select the classification feature set, which was determined by the top-k Geometric Mean sweep.
We implemented a top-k iterative pruning strategy, a form of recursive feature elimination in which the features with the lowest importance scores are incrementally removed and the classifier is re-evaluated at each step [24,25], using the Scikit-learn library [23].

2.6. Classification Algorithm: Multilayer Perceptron

Our classification approach used a Multilayer Perceptron (MLP) neural network trained by the backpropagation algorithm [26,27]. The network was optimized with the Adam optimizer [28] and implemented in PyTorch (Version 2.9) [29]. Data preparation involved subjecting all input features to standardization (Z-score normalization). This preprocessing step is vital for significantly enhancing the efficiency and stability of the training process of the neural network. By transforming the features to have a zero mean and unit variance, we ensured that the MLP algorithm encountered a more spherical and uniform error surface. This ultimately leads to faster convergence, as the optimizer is prevented from oscillating back and forth across uneven scales. Furthermore, scaling the inputs mitigates the significant threat of gradient vanishing or exploding problems that can arise from highly variable input magnitudes.
The MLP was structured with an input layer, output layer, and a sequence of four internal hidden layers. The specific size of these hidden layers was not fixed but dynamically configured based on the total number of input features, N. The MLP architecture was configured according to the number of input features N. For N greater than 64, four hidden layers of 128, 64, 32, and 16 neurons were used; for N less than or equal to 64, three hidden layers of 64, 32, and 16 neurons were used. As the final model was trained on nine selected features, it used the three hidden layer configuration (64, 32, 16).
The MLP was trained with the binary cross-entropy with logits loss, weighting the positive (malignant) class by 1.7 to address class imbalance, using the Adam optimizer (learning rate 0.001, batch size 2). Each fold was trained for up to 30 epochs with early stopping (patience 7 epochs, minimum loss improvement 0.0001). Hidden layers used ReLU activation, and the output layer produced a logit passed through a sigmoid, with a fixed decision threshold of 0.50. Weights were initialized using the Xavier uniform method with zero biases. No dropout or weight decay was applied; early stopping functioned as the sole form of regularization. The positive (malignant) class was assigned a fixed weight of 1.7 in the loss function to address class imbalance. This value was chosen a priori to approximate, while remaining below, the inverse class frequency of the cohort (21 benign versus 10 malignant lesions) and was held constant across all folds. Because the weight was not estimated from the data, it does not differ between folds and introduces no fold-dependent leakage.
The value of k in the Feature Selection algorithm was iterated from 105 to 2 by 1. For each set of k features, the models to be tested were trained using leave-one-subject-out cross-validation (LOSO-CV), and the geometric mean of each trained model was calculated. The set of features that yielded the highest GeoMean was identified as the best model. Although k was iterated from 105 down to 2, Figure 2 displays the GeoMean distribution across feature set sizes from k = 69 down to k = 2 for visualization clarity, as feature sets beyond k = 69 contributed negligibly to discriminative performance. The models were implemented using Python version 3.9.5, SciKitLearn version 1.4.0, and XGBoost version 1.5.1 [30]. The network architecture, optimizer, learning rate, batch size, weight initialization, and class weight were fixed a priori and were not tuned by grid search or nested cross-validation; the number of training epochs was governed by early stopping. The only empirically selected setting was the number of retained features, k, which was chosen to maximize the cross-validated GeoMean. This selection was not nested within the outer cross-validation folds, and a nested procedure that confines feature and hyperparameter selection to each training fold is required for unbiased estimation in future work.

2.7. Statistical Analysis

The confusion matrix was used to determine instances in a predicted class. The accuracy, sensitivity, and specificity values for the radiomics classifier were calculated. The AUC of the receiver operating characteristic curve for malignant-versus-benign was calculated.

2.8. Use of Generative Artificial Intelligence

During the preparation of this manuscript, the authors used gpt-oss-20b (OpenAI) and Claude Opus 5 to improve grammar, spelling, punctuation, and readability. The authors reviewed and edited the output and took full responsibility for the content of this publication.

3. Results

3.1. Subjects

This study included 31 enhancing breast lesions from 30 patients. One patient contributed two separate enhancing lesions (one benign and one malignant), which were treated as independent units. The mean lesion size was 5.9 mm ranging from 2.6 mm to 10 mm. Benign lesions (N = 21) had a mean size of 5.4 ± 1.7 mm (median 5.2 mm), whereas malignant lesions (N = 10) had a mean size of 6.9 ± 1.9 mm (median 6.7 mm). A two-sample t-test revealed no statistically significant size difference between the two groups (CI [−3, 0]) mm. The benign group comprised seven epithelial proliferative lesions without atypia, two papillary lesions without atypia, one benign fibroepithelial tumors, two non-proliferative fibrocystic/stromal changes, four trauma-related or inflammatory changes, one lesion with atypia/in situ disease, and four other benign lesions (Table 2). The malignant group comprised six invasive ductal carcinomas and four ductal carcinomas in situ (Table 3). We grouped DCIS and IDC as malignant since both show similar early-phase DCE-MRI enhancement patterns. The mean patient age was 53.6 ± 10.8 years (median 52, range 29–71) for benign cases and 59.8 ± 9.5 years (median 61, range 45–79) for malignant cases. Figure 3 illustrates focal enhancement on breast MRI images.

3.2. Radiomics

Of the 105 features, only 44 carried non-negligible information, scoring above 0.2 on the aggregated FIS scale, on which the most informative feature scored 1.81. This aggregated scale results from summing the per run Gini importances, each normalized to sum to one, across the repeated RF runs; the 0.2 value is therefore expressed on this scale and not out of 1 (Table 4). Figure 2 shows the relationship between the number of features and resulting GeoMean values. For each feature set size, the boxplots depict the distribution of GeoMean across repetitions, whereas the solid line represents the average performance trend. The results demonstrate a clear trend of improvement in GeoMean as redundant, or less informative, features are removed. With the full 105-feature set, the average GeoMean was relatively low (approximately 0.60–0.65). However, as the number of features was reduced, the performance gradually increased, indicating that the classifier benefited from the elimination of noisy or irrelevant variables. The peak performance was observed when nine features remained, where the average GeoMean reached values close to 0.84. Beyond this point, further feature removal leads to a decline in GeoMean as essential features are discarded. With fewer than seven features, the performance deteriorated significantly, approaching the levels observed with the initial full-feature set.
Table 4. Rank of features for which the Random Forest feature score is greater than 0.2 (glcm: Grey Level Co-occurrence Matrix, glszm: Grey Level Size Zone Matrix, glrlm: Grey Level Run Length Matrix, gldm: Grey Level Dependence Matrix, and ngtdm: Neighboring Gray Tone Difference Matrix).
Table 4. Rank of features for which the Random Forest feature score is greater than 0.2 (glcm: Grey Level Co-occurrence Matrix, glszm: Grey Level Size Zone Matrix, glrlm: Grey Level Run Length Matrix, gldm: Grey Level Dependence Matrix, and ngtdm: Neighboring Gray Tone Difference Matrix).
RankRadiomics FeatureFeature ImportanceRankRadiomics FeatureFeature Importance
#1gldm_DependenceNonUniformity1.81#23gldm_DependenceEntropy0.36
#2glszm_SizeZoneNonUniformity1.28#24gldm_LargeDependenceLowGrayLevelEmphasis0.36
#3glcm_JointEnergy1.12#25shape_Maximum2DDiameterRow0.34
#4shape_MeshVolume1.1#26firstorder_Energy0.34
#5shape_SurfaceArea1.08#27shape_Maximum2DDiameterSlice0.32
#6glcm_JointEntropy0.97#28gldm_SmallDependenceLowGrayLevelEmphasis0.32
#7glcm_MaximumProbability0.79#29glszm_ZoneEntropy0.3
#8glrlm_RunLengthNonUniformity0.75#30glcm_SumEntropy0.29
#9shape_VoxelVolume0.7#31shape_Elongation0.27
#10shape_SurfaceVolumeRatio0.68#32shape_Maximum3DDiameter0.27
#11glszm_LowGrayLevelZoneEmphasis0.56#33firstorder_Skewness0.27
#12shape_Sphericity0.48#34firstorder_Kurtosis0.26
#13glrlm_ShortRunLowGrayLevelEmphasis0.47#35shape_MajorAxisLength0.26
#14shape_Maximum2DDiameterColumn0.47#36ngtdm_Busyness0.25
#15ngtdm_Coarseness0.45#37glcm_Idn0.24
#16shape_Flatness0.44#38glszm_GrayLevelNonUniformity0.23
#17shape_LeastAxisLength0.42#39glrlm_LongRunLowGrayLevelEmphasis0.22
#18glrlm_LowGrayLevelRunEmphasis0.42#40glszm_ZonePercentage0.22
#19glcm_Correlation0.42#41firstorder_10Percentile0.21
#20gldm_LowGrayLevelEmphasis0.4#42glszm_SmallAreaLowGrayLevelEmphasis0.21
#21shape_MinorAxisLength0.4#43firstorder_Uniformity0.20
#22firstorder_TotalEnergy0.37#44glszm_LargeAreaLowGrayLevelEmphasis0.20
We also analyzed how the feature importance score varied across the radiomic variables. Figure 4 shows the feature importance scores for a set of radiomic features. The features are ranked in descending order of their importance, with the x-axis indicating the position in this ranking, from the 1st (most important) to the 105th (least important).
As observed in many similar classification problems, a small number of top features capture most of the relevant information, whereas subsequent features add diminishing value. This suggests a potential redundancy or lower relevance for lower-ranked features. Up to approximately the 12th feature, the scores drop relatively steeply, indicating a rapid decline in importance among the initial top-tier features. Starting with the 13th feature, however, the scores continue to decrease but at a noticeably slower rate, leveling off more gradually toward near-zero by k = 105. This “elbow” or inflection around k = 9 suggests that the top nine–twelve features are the most critical and including more beyond this point yields minimal additional benefit.
Using the top nine feature subset, the MLP classifier achieved an AUC of 0.82 (Figure 5). Sensitivity, specificity, accuracy, GeoMean, and the averaged confusion matrix across repetitions are reported in Table 5, while a representative single-repetition confusion matrix is shown in Figure 6. The low standard deviations across all metrics indicate stable classifier performance across subjects despite the subject-independent evaluation imposed by LOSO-CV. Ninety-five percent confidence intervals for sensitivity, specificity, accuracy, and AUC were estimated by a patient-level bootstrap of the out-of-fold predictions, resampling lesions with replacement over 2000 iterations, with predicted probabilities averaged across the repeated runs prior to resampling. The intervals were as follows: sensitivity 0.80 (0.50 to 1.00), specificity 0.89 (0.76 to 1.00), accuracy 0.86 (0.74 to 0.97), GeoMean 0.84 (0.61 to 0.99), and AUC 0.82 (0.53 to 0.95).

4. Discussion

In this study, we identified nine radiomic features from sub-centimeter enhancing breast lesions using early-phase DCE subtraction breast MRI images to predict malignancy of small lesions. We developed and trained an MLP algorithm with these features. It achieved a sensitivity of 0.80, specificity of 0.89, accuracy of 0.86, GeoMean of 0.84, and AUC of 0.82 through 30-times repeated LOSO cross-validation. A multilayer perceptron was chosen for its capacity to model nonlinear relationships among radiomic features. We recognize that, given the limited sample size, such a model is more susceptible to overfitting than simpler statistical classifiers. To mitigate this, the input space was reduced to a small feature subset, a compact network architecture was used, early stopping was applied, and class weighting addressed the class imbalance. The resulting model should nonetheless be regarded as exploratory, and comparison against conventional classifiers better suited to small datasets is warranted in future validation.
This repeated evaluation strategy was adopted to account for the sensitivity of MLP classifiers to random weight initialization, thereby minimizing stochastic bias and producing performance estimates that are less dependent on any single random seed. The close agreement between GeoMean and AUC indicates consistent, balanced performance across both classes, and the low standard deviations across all metrics suggest relatively stable classifier behavior despite the subject-independent evaluation imposed by LOSO-CV. Notably, a top-k iterative pruning strategy revealed that focusing on a small subset of top-ranked features substantially improved GeoMean relative to the full 105-feature pool. These findings highlight the importance of feature selection for model generalization and demonstrate that an MLP classifier can be effective when paired with a carefully curated subset of features rather than the entire feature pool.
Among the nine top-ranked features used for classification, six were texture-based, drawn from the GLDM, GLSZM, GLCM, and GLRLM classes, and three were shape-based descriptors (Table 4). The dominance of these mathematically derived texture and morphologic descriptors underscores the value of quantitative “hidden” imaging data in accurate classification of small enhancing breast lesions. The stability of the selected features was evaluated with respect to the RF ranking, which was aggregated over repeated runs with different random seeds. The nine selected features, drawn from complementary texture and shape feature classes, were consistently retained across these repetitions, indicating robustness to the stochasticity of the ranking procedure.
D’Amico et al. similarly applied a radiomics-based machine learning pipeline to distinguish malignant from benign enhancing foci (lesions ≤ 5 mm in maximum diameter) on breast MRI [7]. In their study, radiomic features were extracted from one pre-contrast and four post-contrast timepoints, selected using the TWIST framework, and evaluated using a k-nearest neighbors (kNN) model [7]. Using this workflow, they achieved an AUC of ~0.94 and demonstrated that their most discriminative features were primarily intensity- and texture-based with the highest prediction relevance observed in the early post-contrast phase [7]. When compared with our results, there is meaningful overlap at the feature-family level despite our differences in model choice and imaging timepoint selection. For example, second-order texture descriptors dominated in both studies. A key difference, however, is that D’Amico et al.’s top-performing descriptors also included intensity-based features [7], whereas our top-performing features were dominated by GLCM, GLSZM, GLDM, and GLRLM texture features together and shape descriptors. This observation may reflect differences in cohort compositions as our dataset included lesions up to 10 mm in size. Furthermore, differences in feature extraction and our choice to use phase-1 subtraction images rather than multi-timepoint kinetics may also contribute to varying results.
When expanded to include sub-centimeter enhancing breast lesions, MRI radiomic studies can provide additional context for how “small lesions” behave across different cohorts and workflows. Gibbs et al. evaluated small enhancing BI-RADS 4 using a support vector machine (SVM) model with first-order and texture-based radiomic features, achieving AUCs of 0.75–0.81 [14]. They found that features derived from initial enhancement images provided greater discriminative value than those of later timepoints [14]. Similarly, Lo Gullo et al. applied an SVM-based approach to sub-centimeter enhancing masses in BRCA mutation carriers [15]. They extracted over 100 radiomic features and demonstrated that a small subset of radiomic descriptors combined with clinical data and the first post-contrast scan effectively separated benign from malignant lesions [15]. Although the top-ranked features in both studies differed from ours, their findings collectively reinforce the utility of early post-contrast derived descriptors for radiomic classification and support our decision to focus on early-phase subtraction images for feature extraction.
From a clinical standpoint, incorporating radiomics into a breast radiologist’s workflow might offer the opportunity to leverage supplemental quantitative imaging data to improve the classification of small enhancing breast lesions. A radiomics-based classifier such as the one presented in our study could serve as diagnostic support for small enhancing lesions that are difficult to confidently characterize on morphology alone provided that it is validated in a prospective study with a large cohort. Given that breast MRI is highly sensitive but has suboptimal specificity, an additional quantitative malignancy probability could help radiologists better triage lesions toward short-interval MRI follow-up, targeted second-look via ultrasound, or image-guided biopsy [31,32]. Furthermore, early post-contrast subtraction images are routinely acquired as part of standard breast MRI protocols, posing no additional imaging burden or cost to patients. If validated with larger datasets and paired with automated segmentation, these tools might assist with improving diagnostic consistency [31,32].
This study is not without limitations. First, the small sample size, class imbalance, and single-center design restrict the generalizability of our findings. For small sample size, we utilized LOSO cross-validation strategy, which is a robust validation method for small cohorts. To address our imbalanced dataset, we utilized class weighting within the MLP classifier. Second, we only analyzed first-phase DCE subtraction images, as this is where malignant lesions are most likely to enhance. A dedicated optimization analysis is needed to determine the optimal enhancement phase. Third, this is a retrospective study, and our cohort is not consecutive. Patients were identified based on the presence of a breast MRI followed by a biopsy. Because a biopsy provides the only gold standard, we had to restrict the analysis to those subjects who had tissue sampling. As a result of this, this selection strategy might increase the pre-test probability of malignancy and may overestimate diagnostic performance. Fourth, the reported AUC was computed from the pooled out-of-fold predicted probabilities within each repetition and averaged over thirty repetitions; sensitivity, specificity, accuracy, and GeoMean were obtained at a fixed decision threshold of 0.50 and averaged over the same repetitions. The standard deviations reported across repetitions reflect random seed variability only and do not represent patient-level sampling uncertainty. Patient-level confidence intervals via bootstrap, model calibration assessed by a calibration plot, and decision curve analysis were not performed in this exploratory study and are planned for future confirmatory validation. Given a cohort of 31 lesions with 10 malignant events, the MLP used here is over-parameterized relative to the available sample, and LOO-CV does not fully protect against overfitting when feature and model selection draw on the same data. The reported performance should therefore be interpreted as exploratory. Future confirmatory work should benchmark the MLP against simpler classifiers (penalized logistic regression, linear SVM) using nested, patient-level cross-validation in a larger cohort with external validation before any diagnostic claim is made. Last, feature ranking by Random Forest importance and the selection of the optimal number of features k were performed on the complete dataset prior to the repeated LOSO-CV classification experiments. Because these two model-development steps used all available observations, held-out subjects may have indirectly informed feature ranking and the choice of k. The reported performance metrics should therefore be regarded as internal, potentially optimistic estimates.
We present this as a hypothesis-generating, proof-of-concept study. Given its methodological limitations, specifically, feature selection performed on the full dataset prior to cross-validation, a small sample size, and a lesion-level rather than patient-level analytical framework, the reported performance metrics should not be interpreted as unbiased estimates of generalizability. Future prospective studies are required to validate these findings.

5. Conclusions

In conclusion, our preliminary study shows that nine radiomic features extracted from sub-centimeter enhancing breast lesions on early-phase DCE subtraction breast MRI can improve differentiation between benign and malignant lesions when used with an MLP algorithm in a single-center retrospective cohort. These preliminary findings need to be confirmed in future prospective, multicenter studies with larger cohorts to establish clinical utility and reproducibility.

Author Contributions

Conceptualization: P.H., M.M., H.D.D.V., and K.Z.G.; methodology: P.H., M.M., O.A.G.T., R.G.O.C., H.D.D.V., and K.Z.G.; software: C.L., O.A.G.T., R.G.O.C., and K.Z.G.; validation: P.H., M.M., R.G.O.C., and K.Z.G.; formal analysis: C.L., R.G.O.C., and K.Z.G.; investigation: K.Z.G.; resources: O.A.G.T. and K.Z.G.; data curation: C.L., O.A.G.T., R.G.O.C., and K.Z.G.; writing—original draft: P.H., O.A.G.T., R.G.O.C., and K.Z.G.; writing—review and editing: P.H., C.L., M.M., R.G.O.C., H.D.D.V., and K.Z.G.; visualization: K.Z.G.; supervision: K.Z.G.; project administration: P.H. and K.Z.G. All authors have read and agreed to the published version of the manuscript.

Funding

This research did not receive funding.

Institutional Review Board Statement

This original retrospective research project was approved by the University of Florida Institutional Review Board (IRB202401789, 6 February 2025).

Informed Consent Statement

Patient consent was waived due to the retrospective study design.

Data Availability Statement

The datasets presented in this article are not readily available due to patient privacy and ethical reasons.

Acknowledgments

During the preparation of this manuscript, the authors used gpt-oss-20b (OpenAI) and Claude Opus 5 to improve grammar, spelling, punctuation, and readability. The authors reviewed and edited the output and took full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Giaquinto, A.N.; Sung, H.; Newman, L.A.; Freedman, R.A.; Smith, R.A.; Star, J.; Jemal, A.; Siegel, R.L. Breast cancer statistics 2024. CA Cancer J. Clin. 2024, 74, 477–495. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Eberl, M.M.; Fox, C.H.; Edge, S.B.; Carter, C.A.; Mahoney, M.C. BI-RADS Classification for Management of Abnormal Mammograms. J. Am. Board Fam. Med. 2006, 19, 161–164. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Alotaibi, B.S.; Alghamdi, R.; Aljaman, S.; Hariri, R.A.; Althunayyan, L.S.; AlSenan, B.F.; Alnemer, A.M. The Accuracy of Breast Cancer Diagnostic Tools. Cureus 2024, 16, e51776. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Kataoka, M.; Iima, M.; Miyake, K.K.; Matsumoto, Y. Multiparametric imaging of breast cancer: An update of current applications. Diagn. Interv. Imaging 2022, 103, 574–583. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Carnahan, M.B.; Harper, L.; Brown, P.J.; Bhatt, A.A.; Eversman, S.; Sharpe, R.E.; Patel, B.K. False-Positive and False-Negative Contrast-enhanced Mammograms: Pitfalls and Strategies to Improve Cancer Detection. Radiographics 2023, 43, e230100. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Korhonen, K.E.; Zuckerman, S.P.; Weinstein, S.P.; Tobey, J.; Birnbaum, J.A.; SMcDonald, E.; Conant, E.F. Breast MRI: False-Negative Results and Missed Opportunities. Radiographics 2021, 41, 645–664. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. D’Amico, N.C.; Grossi, E.; Valbusa, G.; Rigiroli, F.; Colombo, B.; Buscema, M.; Fazzini, D.; Ali, M.; Malasevschi, A.; Cornalba, G.; et al. A machine learning approach for differentiating malignant from benign enhancing foci on breast MRI. Eur. Radiol. Exp. 2020, 4, 5. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Truhn, D.; Schrading, S.; Haarburger, C.; Schneider, H.; Merhof, D.; Kuhl, C. Radiomic versus Convolutional Neural Networks Analysis for Classification of Contrast-enhancing Lesions at Multiparametric Breast MRI. Radiology 2018, 290, 290–297. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Clauser, P.; Cassano, E.; Nicolò, A.D.; Rotili, A.; Bonanni, B.; Bazzocchi, M.; Zuiani, C. Foci on breast magnetic resonance imaging in high-risk women: Cancer or not? Radiol. Med. 2016, 121, 611–617. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Spak, D.A.; Plaxco, J.S.; Santiago, L.; Dryden, M.J.; Dogan, B.E. BI-RADS® fifth edition: A summary of changes. Diagn. Interv. Imaging 2017, 98, 179–190. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Zhang, J.; Wu, Q.; Lei, P.; Zhu, X.; Li, B. MRI-based radiomics models for early predicting pathological response to neoadjuvant chemotherapy in triple-negative breast cancer: A systematic review and meta-analysis. J. Appl. Clin. Med. Phys. 2025, 26, e70296. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Chen, Q.; Qin, X.; Du, H.; Ma, X.; Tan, S. A delta radiomics model based on ultrasound images predicts response to neoadjuvant therapy in triple negative breast cancer. J. Appl. Clin. Med. Phys. 2025, 26, e70384. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Qi, X.; Wang, W.; Pan, S.; Liu, G.; Xia, L.; Duan, S.; He, Y. Predictive value of triple negative breast cancer based on DCE-MRI multi-phase full-volume ROI clinical radiomics model. Acta Radiol. 2024, 65, 173–184. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Gibbs, P.; Onishi, N.; Sadinski, M.; Gallagher, K.M.; Hughes, M.; Martinez, D.F.; Morris, E.A.; Sutton, E.J. Characterization of Sub-1 cm Breast Lesions Using Radiomics Analysis. J. Magn. Reson. Imaging JMRI 2019, 50, 1468–1477. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Gullo, R.L.; Daimiel, I.; Saccarelli, C.R.; Bitencourt, A.; Gibbs, P.; Fox, M.J.; Thakur, S.B.; Martinez, D.F.; Jochelson, M.S.; Morris, E.A.; et al. Improved characterization of sub-centimeter enhancing breast masses on MRI with radiomics and machine learning in BRCA mutation carriers. Eur. Radiol. 2020, 30, 6721–6731. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Lee, H.-J.; Nguyen, A.-T.; Ki, S.Y.; Lee, J.E.; Do, L.-N.; Park, M.H.; Lee, J.S.; Kim, H.J.; Park, I.; Lim, H.S. Classification of MR-Detected Additional Lesions in Patients With Breast Cancer Using a Combination of Radiomics Analysis and Machine Learning. Front. Oncol. 2021, 11, 744460. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Liu, T.; Lin, J.; Zhang, J.; Lou, J.; Zou, Q.; Wang, S.; Wang, C.; Jiang, Y. The clinical value of radiomics models based on multi-parameter MRI features in evaluating the different expression status of HER2 in breast cancer. Acta Radiol. 2025, 66, 597–607. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Yushkevich, P.A.; Piven, J.; Hazlett, H.C.; Smith, R.G.; Ho, S.; Gee, J.C.; Gerig, G. User-guided 3D active contour segmentation of anatomical structures: Significantly improved efficiency and reliability. Neuroimage 2006, 31, 1116–1128. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. van Griethuysen, J.J.M.; Fedorov, A.; Parmar, C.; Hosny, A.; Aucoin, N.; Narayan, V.; Beets-Tan, R.G.H.; Fillion-Robin, J.C.; Pieper, S.; Aerts, H. Computational Radiomics System to Decode the Radiographic Phenotype. Cancer Res. 2017, 77, e104–e107. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  20. Zwanenburg, A.; Vallières, M.; Abdalah, M.A.; Aerts, H.J.W.L.; Andrearczyk, V.; Apte, A.; Ashrafinia, S.; Bakas, S.; Beukinga, R.J.; Boellaard, R.; et al. The Image Biomarker Standardization Initiative: Standardized Quantitative Radiomics for High-Throughput Image-based Phenotyping. Radiology 2020, 295, 328–338. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef] [Scilit]
  22. Breiman, L.; Friedman, J.H.; Olshen, R.A.; Stone, C.J. Classification and Regression Trees; Wadsworth: Belmont, CA, USA, 1984. [Google Scholar]
  23. Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine Learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar]
  24. Guyon, I.; Weston, J.; Barnhill, S.; Vapnik, V. Gene selection for cancer classification using support vector machines. Mach. Learn. 2002, 46, 389–422. [Google Scholar] [CrossRef] [Scilit]
  25. Guyon, I.; Elisseeff, A. An introduction to variable and feature selection. J. Mach. Learn. Res. 2003, 3, 1157–1182. [Google Scholar] [CrossRef] [Scilit]
  26. Goodfellow, I.; Bengio, Y.; Courville, A. Deep Learning; MIT Press: Cambridge, MA, USA, 2016. [Google Scholar]
  27. Rumelhart, D.; Hinton, G.; Williams, R. Learning Representations by Back-Propagating Errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef] [Scilit]
  28. Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization (2014). arXiv 2017, arXiv:1412.6980. [Google Scholar] [CrossRef] [Scilit]
  29. Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 8–14 December 2019. [Google Scholar]
  30. Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, New York, NY, USA, 13–17 August 2016; pp. 785–794. [Google Scholar]
  31. Motanagh, S.A.; Dwan, D.; Azizgolshani, N.; Muller, K.E.; diFlorio-Alexander, R.M.; Marotti, J.D. Sixteen-Year Institutional Review of Magnetic Resonance Imaging–Guided Breast Biopsies: Trends in Histologic Diagnoses With Radiologic Correlation. Breast Cancer Basic Clin. Res. 2023, 17, 11782234231215193. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Pötsch, N.; Dietzel, M.; Kapetas, P.; Clauser, P.; Pinker, K.; Ellmann, S.; Uder, M.; Helbich, T.; Baltzer, P.A.T.; Pötsch, N.; et al. An AI classifier derived from 4D radiomics of dynamic contrast-enhanced breast MRI data: Potential to avoid unnecessary breast biopsies. Eur. Radiol. 2021, 31, 5866–5876. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Flowchart of patient selection and study cohort derivation. The diagram summarizes the identification, screening, inclusion, and exclusion of patients from the initial dataset retrieval to the final study cohort used for analysis. Reasons for exclusion at each stage are provided where applicable.
Figure 1. Flowchart of patient selection and study cohort derivation. The diagram summarizes the identification, screening, inclusion, and exclusion of patients from the initial dataset retrieval to the final study cohort used for analysis. Reasons for exclusion at each stage are provided where applicable.
Diagnostics 16 02818 g001
Figure 2. Box-and-whisker plot showing cross-validated distribution of Geometric Mean by Ranked Feature Sets (Sen: Sensitivity, Spe: Specificity).
Figure 2. Box-and-whisker plot showing cross-validated distribution of Geometric Mean by Ranked Feature Sets (Sen: Sensitivity, Spe: Specificity).
Diagnostics 16 02818 g002
Figure 3. Breast MRI demonstrates focal enhancement. (A) Axial T1-weighted image; (B) axial T2-weighted STIR sequence; (C) axial early-phase post-contrast subtraction image with red arrow indicating the small enhancing lesion; (D) corresponding subtraction image with the enhancing lesion segmented and highlighted in red.
Figure 3. Breast MRI demonstrates focal enhancement. (A) Axial T1-weighted image; (B) axial T2-weighted STIR sequence; (C) axial early-phase post-contrast subtraction image with red arrow indicating the small enhancing lesion; (D) corresponding subtraction image with the enhancing lesion segmented and highlighted in red.
Diagnostics 16 02818 g003
Figure 4. Ranked feature importance scores (FIS) derived from Random Forest feature-ranking. Features are sorted in descending order (rank 1–105).
Figure 4. Ranked feature importance scores (FIS) derived from Random Forest feature-ranking. Features are sorted in descending order (rank 1–105).
Diagnostics 16 02818 g004
Figure 5. ROC Curve (AUC = 0.82) for the classification of breast cancer subject using the top 9 features (Table 4).
Figure 5. ROC Curve (AUC = 0.82) for the classification of breast cancer subject using the top 9 features (Table 4).
Diagnostics 16 02818 g005
Figure 6. The confusion matrix, which is also observed in the majority of 30-time repeated LOSO-CV experiments.
Figure 6. The confusion matrix, which is also observed in the majority of 30-time repeated LOSO-CV experiments.
Diagnostics 16 02818 g006
Table 1. Dynamic contrast-enhanced (DCE) breast MRI acquisition parameters by field strength (1.5T vs. 3T).
Table 1. Dynamic contrast-enhanced (DCE) breast MRI acquisition parameters by field strength (1.5T vs. 3T).
Dynamic Contrast-Enhanced MRI Parameters
MRI StrengthRepetition Time (TR) (ms)Echo Time (TE) (ms)FOV (mm2)Matrix (-)Slice Thickness (mm)Flip Angle (°)
3.0T3.36–5.541.24–2.46220 × 346–400 × 395224 × 354–416 × 4160.9–1.311–15
1.5T3.67–4.61.42–1.55341 × 338–422 × 420384 × 346–480 × 4321.2–1.511–12
Table 2. Distribution of benign breast lesion pathology by category and final pathology diagnosis.
Table 2. Distribution of benign breast lesion pathology by category and final pathology diagnosis.
Benign CategoryFinal Pathology Diagnosisn
Benignepithelial proliferative lesions without atypiaAdenosis with focal usual ductal hyperplasia (UDH)1
UDH alone1
Columnar cell lesions/change (incl. columnar cell hyperplasia; microcysts/microcalcifications; minute foci of columnar cell change)3
Sclerosing adenosis with focal columnar change and microcalcifications1
UDH with columnar cell change, adenosis, microcysts, focal perilobular chronic inflammation, and focal pseudoangiomatous stromal hyperplasia (PASH)1
Subtotal7
Papillary lesions without atypiaSclerosed intraductal papilloma (no atypia)1
Sclerosing papilloma with associated UDH and chronic inflammation1
Subtotal2
Non-specific benign/limited-sampling findingsFragments of benign/normal mammary tissue3
No invasive or in situ carcinoma/Negative for metastatic carcinoma1
Subtotal4
Trauma-related and inflammatory lesionsFat necrosis2
Organized fat necrosis with chronic inflammation and coarse calcification1
Foreign body-type multinucleated giant cell reaction1
Subtotal4
Benign fibroepithelial tumorsFibroadenoma1
Subtotal1
Non-proliferative fibrocystic and stromal changesMild stromal sclerosis1
Apocrine metaplasia without a dominant mass lesion1
Subtotal2
Lesions with atypia/in situ diseaseLobular carcinoma in situ1
Subtotal1
Total21
Table 3. Distribution of malignant breast lesion pathology by category and final pathology diagnosis.
Table 3. Distribution of malignant breast lesion pathology by category and final pathology diagnosis.
Malignant Category n
Invasive Ductal Carcinoma (IDC) SubtypesInvasive ductal carcinoma4
Infiltrating ductal carcinoma1
Invasive ductal carcinoma, poorly differentiated with focal micropapillary features1
Subtotal6
Ductal Carcinoma in Situ (DCIS)Ductal carcinoma in situ2
Apocrine-type ductal carcinoma in situ1
Intraductal papilloma, ductal carcinoma in situ1
Subtotal4
Total10
Table 5. Confusion matrix and classification accuracy obtained with the MLP classifier using a 30-time repeated LOSO-CV.
Table 5. Confusion matrix and classification accuracy obtained with the MLP classifier using a 30-time repeated LOSO-CV.
TPTNFPFNSenSpeAccuracyGeoMean
Average8.018.82.22.00.800.890.860.84
StdDev0.1830.4300.4300.1830.0180.0210.0150.013
95% CI 0.50–1.000.76–1.000.74–0.970.61–0.99
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Hatch, P.; Louviere, C.; Mete, M.; Tirado, O.A.G.; Cordero, R.G.O.; De Villegas, H.D.; Gumus, K.Z. Distinguishing Benign from Malignant Small Enhancing Breast Lesions Using Radiomics. Diagnostics 2026, 16, 2818. https://doi.org/10.3390/diagnostics16172818

AMA Style

Hatch P, Louviere C, Mete M, Tirado OAG, Cordero RGO, De Villegas HD, Gumus KZ. Distinguishing Benign from Malignant Small Enhancing Breast Lesions Using Radiomics. Diagnostics. 2026; 16(17):2818. https://doi.org/10.3390/diagnostics16172818

Chicago/Turabian Style

Hatch, Parlyn, Christopher Louviere, Mutlu Mete, Oswaldo A. Guevara Tirado, Ruben G. Ortiz Cordero, Hector Diaz De Villegas, and Kazim Z. Gumus. 2026. "Distinguishing Benign from Malignant Small Enhancing Breast Lesions Using Radiomics" Diagnostics 16, no. 17: 2818. https://doi.org/10.3390/diagnostics16172818

APA Style

Hatch, P., Louviere, C., Mete, M., Tirado, O. A. G., Cordero, R. G. O., De Villegas, H. D., & Gumus, K. Z. (2026). Distinguishing Benign from Malignant Small Enhancing Breast Lesions Using Radiomics. Diagnostics, 16(17), 2818. https://doi.org/10.3390/diagnostics16172818

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop