Next Article in Journal
A Problem Landscape Visualisation Method for Multi-Objective Optimisation
Next Article in Special Issue
A Comparison Between Heuristic and Automatic Design in Variational Quantum Circuits for the MaxCut Problem Under Noise Effects
Previous Article in Journal / Special Issue
Neuroevolution of Liquid State Machine Based on Neural Configurations and Positions
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

AI-Driven Biopsychosocial Screening for Breast Cancer: Enhancing Risk Prediction via Differential Evolutionary Linear Discriminant Analysis for Feature Extraction

by
José Luis Llaguno-Roque
1,
Adriana Laura López-Lobato
2,
Juan Carlos Pérez-Arriaga
3,
Héctor Gabriel Acosta-Mesa
2,
Ángel J. Sánchez-García
3,
Gabriel Gutiérrez-Ospina
4,5,
Antonia Barranca-Enríquez
1 and
Tania Romo-González
1,*
1
Laboratorio de Biología y Salud Integral, Instituto de Investigaciones Biológicas, Universidad Veracruzana, Dr. Luis Castelazo Ayala s/n, Col. Industrial, Ánimas km 2.5 Carretera Xalapa-Veracruz, Xalapa 91190, Veracruz, Mexico
2
Instituto de Investigaciones en Inteligencia Artificial, Universidad Veracruzana, Campus Sur, Calle Paseo Lote II, Sección Segunda N° 112, Nuevo Xalapa, Xalapa 91097, Veracruz, Mexico
3
Facultad de Estadística e Informática, Universidad Veracruzana, Av. Xalapa s/n, Col. Obrero Campesina, Xalapa 91020, Veracruz, Mexico
4
Laboratorio de Biología de Sistemas, Instituto de Investigaciones Biomédicas, Universidad Nacional Autónoma de México, Coyoacán 04510, CDMX, Mexico
5
Coordinación de Psicobiología y Neurociencias, Facultad de Psicología, Universidad Nacional Autónoma de México, Coyoacán 04510, CDMX, Mexico
*
Author to whom correspondence should be addressed.
Math. Comput. Appl. 2026, 31(3), 66; https://doi.org/10.3390/mca31030066
Submission received: 25 February 2026 / Revised: 13 April 2026 / Accepted: 21 April 2026 / Published: 24 April 2026
(This article belongs to the Special Issue New Trends in Computational Intelligence and Applications 2025)

Abstract

In Mexico, the high prevalence and mortality rates associated with breast cancer (BC) constitute a critical public health challenge that demands context-specific preventive measures. This study proposes an integrative framework for predicting BC risk based on a biopsychosocial model. We hypothesize that emotional suppression and repression act as key neuroendocrine disruptors and predisposing factors within the Mexican female population. To test this, we systematically compared the predictive performance of various machine learning classification models using the clinical, psychological, and combined profiles of 110 women. These models were evaluated with and without the application of a robust evolutionary algorithm: Differential Evolutionary Linear Discriminant Analysis for Feature Extraction ( D E L D A F E ). The results demonstrated that integrating clinical and psychological data into a combined latent space significantly improved the performance of the classification algorithms. The Artificial Neural Network achieved the highest metrics (0.9975 Precision; 0.9976 F1-score). However, due to the inherent “black-box” nature of these models (limited clinical interpretability), the Decision Tree emerged as the optimal practical alternative, providing highly competitive (0.8874 Precision; 0.8853 F1-score) and interpretable results. These findings provide empirical evidence that psychological factors, rather than being mere incidental comorbidities, could be associated with the etiology of breast cancer and be used as risk factors in predicting the disease. Ultimately, this AI-driven biopsychosocial screening model offers a scalable, low-cost, and context-adapted risk assessment tool for early BC diagnosis in Mexican women.

1. Introduction

Breast cancer (BC) is the most prevalent neoplasm in women worldwide. Nationally, it has been the most frequent and deadliest cancer affecting Mexican women since 2006 (Table A1 in Appendix A) [1,2,3,4]. With projections estimating over 50,000 new annual cases by 2050, BC represents an urgent and expanding public health challenge. There is broad consensus that early detection is essential to reduce morbidity and mortality. While mammography remains the gold standard for screening programs worldwide, it presents limitations. Widespread mammographic screening is often restricted by age guidelines and carries concerns regarding radiation exposure, particularly for younger populations [5,6,7]. This is especially relevant in Mexico, where advanced-stage BC is frequently diagnosed in women aged 20–41 [8,9]. Since 2013, our group has developed non-radiographic resources to identify young Mexican women at elevated risk, aiming to optimize early detection while minimizing unnecessary radiation exposure.
Although the Official Mexican Standard [10] acknowledges that biological, iatrogenic, reproductive, and lifestyle factors interplay to determine BC relative risk, current national screening programs rely almost exclusively on biological and reproductive history. This approach has not substantially reduced incidence or mortality rates, likely because these factors explain, at best, 50% of BC prevalence. We have recently suggested that the accuracy of risk identification would increase by incorporating psycho-affective factors into screening designs [11,12]. Evidence suggests that specific personality traits—such as low emotional containment, low defensiveness, and high stress—are consistently observed in Mexican women with BC prior to diagnosis [11]. Physiologically, emotional suppression and repression are known to alter diurnal cortisol secretion, a dysregulation associated with early mortality in metastatic BC patients [13,14]. Furthermore, principal component analyses have demonstrated that healthy women and those with benign or malignant breast lesions cluster differentially when emotional suppression, repression, and stress symptoms are included in the phenotyping [12]. Notably, when clinical and psychological variables are analyzed together, emotional suppression variables significantly define differential group clustering. These observations support the hypothesis that screening for psychological factors is crucial for prevention, early diagnosis, and understanding BC pathogenesis [15,16,17,18,19].
To improve risk identification in young Mexican women, García-Camacho et al. [20] previously designed a decision tree-based model assessing the Stress Symptom Inventory (ISE), the Courtauld Emotional Control Suppression Scale (CECS), and the Weinberger Adjustment Inventory (WAI). Their results showed that while these tools achieved low accuracy individually (34–40%), their combination improved prediction to 42.7%. This suggests that while the emotional dimension has discriminatory value, the modeling approach required refinement to achieve clinical utility.
On the other hand, recent advances in biopsychosocial risk prediction and explainable AI highlight the importance of integrating multi-domain data for oncology applications. Hussain et al. [21] reviewed machine learning approaches for breast cancer risk prediction, emphasizing the added value of psychosocial variables in enhancing model accuracy. Similarly, radiomics-based studies have demonstrated that multimodal feature fusion, such as parenchymal enhancement and breast symmetry in MRI, improves discrimination power in early detection [22]. In parallel, evolutionary feature extraction methods have been increasingly applied in oncology, as shown by Wang et al. [23], who underscored the role of metaheuristic optimization in refining tumor subtype classification. Explainable AI (XAI) frameworks, such as those discussed by Carriero et al. [24], further stress the necessity of transparency in clinical models, particularly when integrating heterogeneous data sources.
Beyond breast cancer, comparable strategies have been successfully deployed in other medical domains. The NEDL-GCP model [25] introduced a nested ensemble deep learning framework for gynecological cancer risk prediction, demonstrating how ensemble architectures can robustly manage complex clinical datasets. Doumari et al. [26] proposed a novel high-accuracy model for the early diagnosis of Parkinson’s disease, illustrating the potential of feature-fusion and classifier optimization in neurodegenerative contexts.
Building on this foundation, we present a new risk assessment model that employs Differential Evolutionary Linear Discriminant Analysis for Feature Extraction ( D E L D A F E ). This approach integrates biological, reproductive, and psychological variables to manage the complexity of the dataset. Unlike previous models, D E L D A F E simplifies the observation of risk classes and reveals the synergistic behavior of the data. Furthermore, this work is distinguished by its focus on the Mexican population and its key innovation: the use of psychological data alongside clinical data—an approach that, unlike our group [11,12,27], no other previous study in the country has integrated into breast cancer risk assessment, a gap that this work seeks to fill. Although the analyzed cohort consists of 110 individuals, it accurately represents the psychological and clinical information of Mexican women, a group in which identifying risk factors for breast cancer is a priority.
The D E L D A F E method was used to improve classification using machine learning algorithms, prioritizing those that offer explainability in their results. This approach lays the groundwork for future applications, such as the development of a web module that estimates breast cancer risk in an accessible, non-invasive, highly accurate, and timely manner.

2. Methodology

2.1. Participants and Study Design

This study utilized clinical and psychological data from 150 women attending routine gynecological consultations at the General Hospital of Mexico “Dr. Eduardo Liceaga.” All participants provided written informed consent, ensuring confidentiality and anonymity in accordance with the Declaration of Helsinki (printed in the British Medical Journal, 18 July 1964).
The sample was selected to represent a broad age range (see Table A2 in Appendix A), reflecting the epidemiological reality of BC in Mexican women; it is frequently diagnosed approximately a decade earlier than global averages [5]. Participants were recruited prior to histopathological diagnosis and were subsequently stratified into three groups based on their clinical pathology reports. The Healthy Group (H, n = 50) clustered women with no palpable or radiological breast pathology. The Benign Breast Pathology Group (BBP, n = 50) was formed by women diagnosed with fibrocystic changes, fibroadenomas, or mastitis. The Breast Cancer Group (BC, n = 50) included women diagnosed with infiltrating ductal carcinoma (naïve to treatment).

2.2. Data Collection Instruments

Physical, hereditary, and lifestyle data were collected using a 57-item General Data Questionnaire. Psychological profiles—specifically emotional repression, suppression, and stress symptomatology—were assessed using three validated instruments:

2.2.1. Inventory of Stress Symptoms (ISE)

Designed by Benevides-Pereira et al. [28], the ISE evaluates the frequency of stress manifestations in daily life across three dimensions: psychological, physical, and social symptoms. It consists of 30 items rated on a 5-point Likert scale (0 = never to 4 = always). Global symptomatology scores are categorized as low (0–22), medium (22–33), or high (>33).

2.2.2. Courtauld Emotional Control Scale (CECS)

Developed by Watson and Greer [29] and validated for Spanish speakers by Durá et al. [30]. The CECS assesses the extent to which individuals suppress reactions to negative emotions. It comprises 21 items across three subscales: Anger, Anxiety, and Depression. Higher scores indicate greater emotional suppression. The instrument demonstrates high internal consistency (α = 0.95). For this study, a median split (score = 32) was used to categorize participants into “Emotional Suppression” (score ≥ 32) or “Emotional Expression” (score < 32) groups.

2.2.3. The Weinberger Adjustment Inventory (WAI)

The WAI [31,32,33] evaluates adjustment styles through three primary scales: Subjective Distress, Restraint, and Defensiveness. The interactions between these scales allow the identification of six typologies of adjustment, including the Repressive (Type C) coping style. The Spanish version validated for the Mexican population [34] was used (Cronbach’s α ranges from 0.69 to 0.89 across subscales).

2.3. Feature Encoding and Binarization

Given the binary nature of the clinical data, and in order to compare them with the results of the psychological instruments, the latter were transformed into dichotomous variables to assess the presence or absence of certain psychological traits in the groups of women. The psychological data, originally expressed as continuous scores or subscales, were categorized into low, medium, and high levels according to their values. Subsequently, these categories were recoded into binary variables (1 for the presence of the characteristic; 0 for its absence), using cut-off points previously defined and described in Romo-González et al. [27]. In most cases, the original cut-off points for each instrument were retained; however, in the case of the CECS, these were established based on 95% confidence intervals. The specific binarization criteria are detailed in Table A3 of Appendix A. Appendix C examines the potential loss or gain of information resulting from the binarization of psychological data.

2.4. Machine Learning Identification Model

To manage dataset complexity and achieve optimal separation among the three groups (H, BBP, and BC), a supervised learning approach, Differential Evolutionary Linear Discriminant Analysis for Feature Extraction ( D E L D A F E ), was employed. This method performs dimensionality reduction while enhancing class separability by combining linear transformations of the original variables with evolutionary optimization to obtain a projection matrix that maximizes class separation in a two-dimensional space. The following sections describe the main components of this method.

2.4.1. Projection-Based Dimensionality Reduction

Let X R m × n be a dataset with m samples and n features. The goal is to obtain a reduced representation Y R m × d , with d = 2 , defined as:
Y = X W
where W R n × d   is the projection matrix whose columns are orthonormal vectors. This transformation produces new features that are linear combinations of the original variables, while orthonormality ensures that the columns form a valid basis for the projected space.

2.4.2. Differential Evolution Optimization

Differential Evolution (DE) is a population-based optimization method that iteratively refines a set of candidate solutions using mutation, crossover, and selection operators.
At iteration (generation) g , the population consists of N candidate solutions (individuals) randomly initialized within the domain of a fitness function f . Each individual is represented by a vector x i R d , where d denotes the dimensionality of the problem. In the DE-LDA_{FE} framework, each individual corresponds to a vector w i R n × d that encodes a candidate projection matrix W i R n × d , with d = 2 , as presented in Equation (2). The particularity of these matrixes is that they are orthonormal by columns, to define the projection directions.
W i = w 11 w 12 w 21 w 22 w n 1 w n 2   w i = w 11   w 21     w n 1   w 12   w 22     w n 2
The DE/best/1/bin strategy is used. For each target vector w i in the population, a mutant vector is generated as:
v i = w b e s t + F ( w r 1 w r 2 )
where w b e s t is the current best solution, w r 1 , w r 2 are random individuals, and F is the mutation factor.
A trial vector is then created via binomial crossover:
u i , j = v i , j if   r a n d j C R   or   j = j r a n d w i , j otherwise
where C R is the crossover rate, and j r a n d corresponds to a randomly selected position. The selection step retains the solution with the highest fitness between the trial vector u i and the target vector w i .
These steps are repeated until a stopping criterion is met, such as reaching a maximum number of generations or population convergence. Thus, through this iterative evolutionary process, the population of candidate solutions is progressively refined by combining and selecting the best individuals at each generation. As a result, the algorithm converges toward projection matrices that increasingly improve class separability, enabling the discovery of a near-optimal solution in the lower-dimensional space.

2.4.3. Fitness Function

For the D E L D A F E method, the quality of each projection is evaluated using Fisher’s discriminant criterion, which seeks to maximize the separation between class means while minimizing the dispersion within each class. This is achieved by maximizing the ratio between the between-class scatter and the within-class scatter in the projected space, as presented in Equation (5).
J ( W ) = W T S B W W T S W W
where S B and S W represent the between-class and within-class scatter matrices, respectively, as presented in Equations (6) and (7), where G denotes the number of classes, μ g the mean of class g , and μ the global mean.
S B = g = 1 G m g ( μ g μ ) ( μ g μ ) T
S w = g = 1 G x C g x μ g x μ g T
This criterion promotes projections in which samples from the same class are tightly clustered and those from different classes are well separated in the reduced space.

2.5. Evaluation of Classifiers with/Without the D E L D A F E Algorithm

In this research, to analyze the prediction of breast cancer risk by integrating clinical and psychological variables, a comparison of six classification models is performed: K-nearest neighbors (K-nn; k = 3, 5, 7, 9, and 11); Linear Discriminant Analysis (Linear); Gaussian classifier (Gaussian); Decision Tree (DT); Support Vector Machine (SVM); and Artificial Neural Network (ANN). This comparison is made by considering the raw data of the clinical variables, the psychological variables, and finally, both. In addition, the feature extraction method, Feature Extraction ( D E L D A F E ), is used to graphically analyze classification performance and improve the classifier’s accuracy., as presented in Figure 1.
To determine the most suitable Artificial Neural Network (ANN) configuration, a systematic experimental evaluation was conducted using feedforward fully connected neural networks for multiclass classification. The tested architectures consisted of 1 to 5 hidden layers, each with the same number of neurons. The number of neurons per layer was varied across the set {10, 20, 30, 40, 50, 100, 200, 300, 400, 500}, generating a broad range of shallow and deep network configurations. In all cases, each network received the dataset as input and produced class predictions at the output layer. This structured exploration enabled comparison of architectures with different depths and capacities to identify the configuration that achieved the best classification performance according to the selected evaluation metrics.
For the D E L D A F E method, a grid search strategy was employed by systematically exploring combinations of the Differential Evolution (DE) parameters, including population size N { 100 , 300 , 500 , 1000 , 2000 } , number of generations N G { 300 , 500 , 700 , 900 } , crossover rate C R { 0.3 , 0.5 , 0.7 } , and scale factor F { 0.3 , 0.5 , 0.7 } . The best classification values were obtained with the configurations shown in Table 1 for each data type.

3. Results

This section presents the comparative performance of the selected classification models in predicting BC risk using clinical, psychological, and integrated datasets. Following the exclusion of records with missing data, a final cohort of 110 patients (out of the original 150) was included in the analysis.
Model performance was assessed using both the raw data and data obtained after using the   D E L D A F E feature extraction method. To rigorously evaluate the classifiers, standard performance metrics were utilized, including True Positive (TP) Rate, False Positive (FP) Rate, Accuracy, Recall, F1-score, and the Area Under the Receiver Operating Characteristic Curve (AUC-ROC).
For the experiments, the classification performance was evaluated on the training set. This choice was made because the primary goal of the proposed method is to compare the discriminative quality of different feature representations rather than to estimate generalization performance. Although this approach may yield optimistic accuracy values, it ensures a fair and controlled comparison across methods, since all models are evaluated under identical conditions. Additionally, the dataset’s limited size supports the use of this evaluation strategy.
For the raw variables, the results presented consider only one execution, since the dataset employed does not vary; however, for the projections of the variables with the D E L D A F E method, since it is a stochastic, population-based metaheuristic algorithm, the projections are not unique, so the results presented are the mean of 10 runs, and, to perform a fair comparison, the best solution obtained on these 10 runs is employed for comparison with the results of the raw data.

3.1. Classification Based on Clinical Variables

3.1.1. Analysis of Raw Clinical Data

In this initial experimental phase, the classification models were trained using the clinical variables in their original, unprocessed state. This baseline analysis serves to evaluate the inherent discriminatory power of standard biological and reproductive risk factors prior to the application of any feature extraction or dimensionality reduction techniques. The comparative performance metrics for each classifier in this scenario are detailed in Table 2.
As shown in Table 1, the performance of classifiers trained on raw clinical data varied significantly depending on the algorithm’s architecture. The ANN with 30 neurons demonstrated superior predictive capability, achieving the highest scores across all metrics, with an F1-score of 0.9034 and an AUC-ROC of 0.9892. The Gaussian classifier also performed robustly, securing the second-best performance (F1-score: 0.8064).
Conversely, the K-nn algorithm exhibited an inverse relationship between the number of neighbors (K) and model performance; as K increased from 3 to 11, the F1-score progressively declined from 0.6690 to 0.5070. Linear and SVM models yielded moderate results (F1-scores ≈ 0.60–0.62), suggesting that the boundaries between risk groups in the raw clinical space are non-linear and complex, requiring more sophisticated modeling approaches like ANNs to be effectively decoded.

3.1.2. Classification Using Clinical Data Projections via D E L D A F E Method

As established in previous sections, the   D E L D A F E method was implemented to improve class separation by projecting the original variables into an optimized, lower-dimensional latent space. This projection aimed to enhance the discriminative performance of the classification models. The average performance metrics obtained using these clinical projections, calculated over 10 independent runs, are presented in Table 3.
As indicated in Table 3, the models demonstrating the highest predictive capability when utilizing the projected clinical variables were the ANN architecture with 400 neurons (Precision: 0.9111; F1-score: 0.8990) and the DT model (Precision: 0.8043; F1-score: 0.7831). The remaining models exhibited moderate, comparable performances. Consistent with the raw data findings, the K-nn model displayed a decline in both Precision and F1-score as the number of neighbors (K) increased, confirming that K = 3 remains the optimal configuration for this spatial mapping.
While Table 3 reflects the average robustness across 10 iterations, it is also highly relevant to identify the peak potential of each classifier. Table 4 details the single best-performing run for each model based on the clinical projections.
The results in Table 4 mirror the trends observed in the average metrics, reaffirming the superiority of the ANN model. Notably, the peak performance was achieved during Experiment 3 using an expanded architecture of 500 neurons, which yielded an exceptional Precision of 0.9149 and an F1-score of 0.9024. The DT model similarly achieved its peak during Experiment 8, solidifying its position as the second most effective classifier for these specific clinical projections.

3.1.3. Performance Comparison: Raw vs. Projected Clinical Data

Figure 2 illustrates the comparative performance of the classification models when trained on raw clinical variables versus the optimized latent projections generated by the D E L D A F E method. As observed, for K-nn, DT, SVM, and Linear, the application of the D E L D A F E method successfully enhanced class separability, thereby contributing to an overall improvement in classification metrics.
However, two notable exceptions emerged: Gaussian and ANN classifiers achieved their peak performance using the raw, not projected clinical data. This outcome suggests that these specific non-linear algorithms possess an inherent capacity to map the complex, multi-dimensional relationships of biological risk factors in their original state, without requiring prior dimensionality reduction or feature extraction.

3.2. Classification Based on Psychological Variables

3.2.1. Analysis of Raw Psychological Variables

In this phase of the analysis, the classification models were trained using the raw psychological variables exclusively (without any preprocessing or projection). This experiment establishes a baseline to evaluate the inherent predictive power of psychological factors, such as emotional repression, suppression, and stress, when analyzed in isolation. The comparative performance metrics for each classifier in this raw psychological space are presented in Table 5.
As shown in Table 5, the classification models exhibited varying degrees of success when processing the psychological variables. Most notably, the ANN, configured with an optimized hidden layer of 20 neurons, achieved extraordinary, perfect predictive metrics across the board (Precision: 1.0000; F1-score: 1.0000; AUC-ROC: 1.0000). This indicates an absolute separation of the risk classes by the network within this specific training configuration. The DT model emerged as the second most effective classifier for raw psychological data, recording a Precision of 0.7402 and an F1-score of 0.7298. The remaining probabilistic and linear models demonstrated moderate performance. Consistent with the trends observed in the clinical data analysis, the K-nn algorithm displayed an inverse relationship between the number of neighbors (K) and model efficacy; as K increased, both Precision and the F1-score systematically declined, confirming that K = 3 remains the optimal parameter within this configuration.

3.2.2. Classification Using Psychological Data Projections via D E L D A F E Method

Following the same procedure applied to the clinical data, the psychological variables were projected into an optimized latent space using the D E L D A F E method. This transformation aims to maximize class separability and enhance the predictive performance of the classification algorithms. The average performance metrics obtained from these psychological projections, calculated over 10 independent runs, are presented in Table 6.
As detailed in Table 6, the ANN, dynamically configured with an expanded hidden layer of 500 neurons for this projected space, maintained its perfect predictive capability across the key metrics (Precision: 1.0000; F1-score: 1.0000). The DT model demonstrated notable improvement compared to its raw data baseline, securing the second-highest average performance (Precision: 0.8224; F1-score: 0.8149). The remaining classifiers exhibited similar, moderate performances. Regarding the K-nn model, the inverse relationship between the number of neighbors (K) and model accuracy persisted, confirming K = 3 as the optimal configuration for this projected space.
While Table 6 illustrates the average stability of the classifiers, it is equally important to evaluate their peak predictive potential. Table 7 presents the single best-performing iteration out of the 10 total executions for each model using the projected psychological variables.
The data in Table 7 corroborate the findings from the average metrics. The ANN model achieved a flawless classification outcome during Experiment 1 (Precision: 1.0000; F1-score: 1.0000). The DT similarly reached its peak performance during Experiment 3, yielding a Precision of 0.8668 and an F1-score of 0.8682, buttressing its status as a highly effective secondary classifier for this dataset. Lastly, for the KNN model, as the number of neighbors (K) increases, both the Precision and the F1-score tend to decrease. Therefore, the value of K = 3 corresponds to the best performance within this configuration.

3.2.3. Performance Comparison: Raw vs. Projected Psychological Data

Figure 3 illustrates the comparative performance of the classification models when trained on raw psychological variables versus the optimized latent projections generated by the D E L D A F E method. Consistent with the findings from the clinical dataset, the application of the D E L D A F E method successfully enhanced class separability, contributing to a systematic improvement in classification metrics across K-nn, Linear, DT, and SVM.
The sole exception in this experimental phase was the ANN model. Because the dynamically configured ANN had already achieved perfect predictive accuracy (1.0000) using the raw psychological data, its performance metrics remained identical when processing the projected data, demonstrating a ceiling effect in its classification capability for this specific model.

3.3. Classification Based on Combined Clinical and Psychological Variables

3.3.1. Analysis of Raw Combined Data

In this comprehensive phase of the study, the classification models were evaluated using an integrated dataset comprising both the raw clinical and psychological variables. This experiment was designed to determine whether combining these two distinct domains, the biological risk factors and the psychological states, without any prior preprocessing or projection, enhances the overall predictive accuracy for breast cancer risk. The comparative performance metrics are presented in Table 8.
As detailed in Table 8, the ANN, configured with 20 neurons in its hidden layer, once again demonstrated exceptional predictive capacity on the raw data, achieving perfect scores in several key metrics (Precision: 1.0000; F1-score: 1.0000). The integration of both variable sets significantly improved the performance of the SVM model, which emerged as the second most robust classifier for this combined dataset (Precision: 0.8929; F1-score: 0.8693). Consistent with previous experimental phases, the K-nn model’s performance degraded progressively as the number of neighbors (K) increased, reaffirming that K = 3 remains the optimal parameter setting for this algorithm.
Intriguingly, the Gaussian classifier failed to converge valid predictions in this combined raw space (yielding 0.0000 across all metrics). This was further analyzed and is primarily attributed to the high dimensionality of the feature space relative to the number of samples. In this setting, the covariance matrix becomes ill-conditioned or singular, preventing its reliable inversion and leading to degenerate probability estimates, which explains the zero values observed across all evaluation metrics. This issue was verified by inspecting the rank and condition number of the covariance matrix. These findings further justify the use of dimensionality reduction methods, such as D E L D A F E , which project the data into a lower-dimensional space where covariance estimation is stable, and class separability is significantly improved.

3.3.2. Classification Using Combined Data Projections via the D E L D A F E Method

To address the complexities of the combined dataset, the integrated clinical and psychological variables were projected into an optimized latent space using the D E L D A F E method. Table 8 presents the average performance metrics of the classification models, calculated over 10 independent runs, when trained on these combined, optimized projections.
As shown in Table 9, the models demonstrating the highest predictive capability on the projected combined variables were the ANN, utilizing an expanded 500-neuron architecture, which achieved near-perfect average metrics (Precision: 0.9975; F1-score: 0.9976), followed by the DT model (Precision: 0.8874; F1-score: 0.8853).
Notably, the application of the   D E L D A F E method resolved the prior convergence failure of the Gaussian model on the raw combined data. By optimizing the dimensional space, the Gaussian algorithm successfully generated predictions, achieving a highly competitive Precision of 0.7956. Lastly, among the K-nn configurations, an increase in neighbors (K) again correlated with a decrease in Precision and F1-score, establishing again that K = 3 is the optimal baseline for this model.
Table 10 details the single best-performing run out of the 10 total executions for each model, highlighting the peak potential of these classifiers when utilizing the combined data projections.
The results in Table 10 mirror the trends observed in the average metrics, reinforcing the superiority of the ANN model, which achieved flawless classification (Precision: 1.0000; F1-score: 1.0000) during Experiment 3. The DT model similarly peaked during experiment 5, yielding an outstanding Precision of 0.9148 and an F1-score of 0.9077. The remaining algorithms exhibited consistent, high-tier performance, validating the efficacy of the D E L D A F E transformation on combined, heterogeneous datasets. Once again, in the case of the K-nn model, as the number of neighbors (K) increases, both the Precision and the F1-score tend to decrease. Therefore, the value of K = 3 corresponds to the best performance within this configuration.

3.3.3. Performance Comparison: Raw vs. Projected Combined Data

Figure 4 illustrates the comparative performance of the classification models when trained on the raw, combined clinical and psychological dataset versus the optimized latent projections generated by the D E L D A F E method. Consistent with the findings from the isolated datasets, the application of the D E L D A F E method successfully enhanced class separability, contributing to improved classification metrics for K-nn, Linear, and DT models.
Two notable exceptions to this trend were the SVM and the ANN models, both of which achieved their peak performance utilizing the raw, unprojected data. As established in Section 3.3.1, the dynamically configured ANN had already reached perfect predictive accuracy (1.0000) on the raw combined data, while the SVM demonstrated a remarkably high natural affinity for delineating the raw integrated feature space.
In stark contrast, the Gaussian model benefited the most profoundly from the D E L D A F E projections. While the algorithm completely failed to converge on the raw, heterogeneous data (yielding metrics of 0.0000), the optimized dimensionality reduction provided by the D E L D A F E method resolved these mathematical singularities. This transformation allowed the Gaussian model to successfully compute probabilities and achieve highly competitive performance, underscoring the vital role of D E L D A F E in stabilizing sensitive algorithms when integrating complex, multi-domain datasets.

3.4. Comprehensive Analysis of Variable Integration

Figure 5 illustrates a comparative analysis of the precision achieved by the various classification models across the raw clinical, psychological, and integrated datasets. Across all unprojected data types, the ANN consistently demonstrated superior performance, achieving the highest precision and proving to be the best-performing architecture for this study’s baseline data.
Regarding the K-nn model, the analysis confirms that precision inversely correlates with the number of neighbors (K). The peak performance within this family of models was consistently observed at K = 3, particularly within the clinical and combined datasets, whereas higher K values systematically reduced the algorithm’s discriminative capacity. The Gaussian model performed adequately with isolated clinical and psychological data; however, it experienced a total convergence failure (precision dropping to zero) when attempting to process the raw combined variables, suggesting profound mathematical difficulty in handling this unprojected, heterogeneous data integration. The DT model exhibited robust and stable performance across all three data domains, standing out slightly on the clinical data with accuracy metrics superior to most baseline models, though not reaching the absolute ceiling of the ANN. Finally, the SVM and Linear maintained a generally intermediate performance profile, yielding acceptable precision without significantly outperforming the top-tier algorithms.
Figure 6 presents a parallel comparison of the models’ predictive accuracy when utilizing the optimized data projections (clinical, psychological, and combined) generated by the D E L D A F E method. Significantly, the data indicates that the integration of both clinical and psychological variables into a combined, projected latent space yields the highest precision values across the majority of the evaluated models. This outcome strongly suggests that the complementary information derived from pairing biological risk factors with psychological states, when properly optimized and mathematically smoothed via the D E L D A F E method, significantly enhances the overall discriminatory power and predictive reliability of the classifiers.
While the main objective of this work is to use clinical, psychological, and combined data to construct a breast cancer risk profile, it is also possible to identify, using the DE-LDA_FE method, isolated characteristics of each data type that may contribute to risk prediction. This is achieved through the use of a linear model, whose transformation matrix provides coefficients that reflect the relevance of the variables, ordered in descending order. The variables considered relevant are presented in Appendix B (Table A4, Table A5 and Table A6). However, as noted by Montes-Nogueira et al. (2018) and Romo-Gonzalez et al. (2018) [11,12], the breast cancer risk profile is more robust when clinical and psychological data are integrated together, rather than using isolated characteristics.

4. Discussion

4.1. The Need for Context-Specific Risk Models

Breast cancer remains one of the most critical public health issues globally, disproportionately affecting women in developing countries such as Mexico. Despite advancements in diagnosis and treatment, mortality rates remain high, largely due to delayed detection and the limited reach of effective screening programs. In Mexico, breast cancer has become the leading cause of cancer-related deaths among women, and its clinical presentation occurs approximately a decade earlier than in high-income countries. These circumstances highlight the urgent need for early detection tools adapted to the specific characteristics of the Mexican population.
Several models have been designed to evaluate the risk of developing breast cancer throughout a woman’s life. The most widely used—such as the Breast Cancer Risk Assessment Tool (Gail Model), the International Breast Cancer Intervention Study (Tyrer-Cuzick Model), and the Breast and Ovarian Analysis of Disease Incidence and Carrier Estimation Algorithm (BOADICEA)—rely on biological and reproductive risk data to provide estimates that guide physician decisions regarding future screening, preventive interventions, or chemoprophylaxis [35,36]. While these tools have proven their efficacy and are widely used in developed countries [37], they rely primarily on biological metrics, failing to account for the complex sociocultural or psychological dimensions that are critical in other demographic groups.
Currently, Mexico lacks risk-assessment tools calibrated to the national population that consider lifestyle or psychological factors. Consequently, these standardized models may lack predictive power when applied to Mexican women, who experience different environmental stressors, cultural pressures, and emotional coping strategies than the populations for which the models were originally developed. Our baseline results support this limitation: the clinical-only models failed to distinguish accurately between benign pathology and cancer, suggesting that biological risk factors overlap significantly between these groups. It was only the addition of psychological profiling that successfully “unlocked” the models’ discriminative power.

4.2. Model Performance and the “Black Box” Problem

Our results demonstrate that integrating psychological variables—specifically emotional repression, suppression, and stress symptoms—with standard clinical history significantly enhances breast cancer risk prediction. The classification models achieved their greatest accuracy when utilizing the combined projections of clinical and psychological variables optimized through the D E L D A F E method (Precision: 0.8874–0.9975; F1-score: 0.8853–0.9976), far outperforming the individual, not projected datasets.
Across the evaluated algorithms, the ANN clearly stood out as the most accurate model, reaching its peak when clinical and psychological variables were combined. However, while the ANN achieved the highest accuracy metrics, its complex architecture renders it a “black box.” This limits the interpretability of the results, which is a critical issue in the medical field where understanding how a given prediction is reached is essential for clinical trust and patient communication. This issue of explainability has been highlighted in the recent literature [38], emphasizing the need for using Explainable Artificial Intelligence (XAI) in healthcare. Consequently, as our work points out, the DT model, which ranked second in overall performance, appears to be a more suitable and practical alternative for medical applications. It allows clinicians to visually track the classification process and transparently explain the specific criteria used to arrive at a risk assessment.

4.3. Psychological Variables and Physiological Pathways

The empirical evidence generated by our models supports the hypothesis that psychological factors are not merely incidental comorbidities, but rather associated risk factors in the multidimensional etiology of breast cancer. The predictive value of psychological variables aligns closely with Type C Personality Theory [15,16,17]. Our data suggest that high levels of emotional suppression (measured by the CECS) and defensiveness (measured by the WAI) could be related as distinctive markers in the breast cancer group.
Physiologically, the relationship of these traits within the multicausal etiology of the disease could be explained by established psychoneuroimmunological mechanisms. That is, chronic emotional inhibition is known to dysregulate the hypothalamic–pituitary–adrenal axis, altering cortisol secretion patterns. This dysregulation could, in turn, affect immune surveillance and tumor suppression mechanisms [11,13,14]. Therefore, the greater accuracy achieved by combining psychological and clinical data indicates that emotional suppression and distress symptoms are not only consequences of the disease but could also serve as early indicators of physiological vulnerability.

4.4. Clinical Implications and Future Applications

The proposed D E L D A F E algorithmic pipeline offers a scalable, non-invasive alternative for early screening. Unlike standard mammography, which is resource-intensive and carries radiation risks, this methodology could be deployed as a web-based or mobile pre-screening module. Such a tool would be particularly valuable in low-resource settings, offering the following systemic benefits:
(a) Empowering Patients: Providing women with access to validated self-assessment tools in contexts where direct medical services may be limited or geographically inaccessible.
(b) Optimizing Triage: enabling primary care providers to prioritize referrals for women who exhibit a high-risk biopsychosocial profile, even if they are young or asymptomatic.
This module could be readily adapted for broader use across Latin America, significantly enhancing early detection frameworks in resource-constrained areas.

4.5. Limitations/Future Work

While the results are very promising, this study has limitations that should guide future research. First, the sample size (n = 110), although statistically sufficient for this exploratory validation, requires further expansion. Future studies should incorporate a larger multicenter cohort to ensure the model’s generalizability across the diverse geographic and socioeconomic regions of Mexico. They should also incorporate techniques for handling missing data and unbalanced cohorts. Although the proposed algorithm successfully separates risk groups, the cross-sectional nature of the data limits definitive causal inference. Longitudinal studies are needed to confirm whether these specific psychological traits precede cellular malignancy or evolve concurrently with it. Reference methods such as entropy-regulated two-view NMF and semi-supervised adaptive symmetric NMF could also be applied, as they offer comparative frameworks and could reveal how the observed benefits behave under more diverse age distributions, patterns of missing data, or larger-scale multicenter cohorts. In-depth sensitivity and robustness analyses should also be performed on the experiments.

5. Conclusions

This study provides strong empirical support of the biopsychosocial model in oncology. It demonstrates that combining engineered clinical and psychological features via the D E L D A F E   projection method significantly enhances the discriminative power of machine learning classifiers. Ultimately, this research illustrates that Artificial Intelligence can effectively bridge the gap between psychology and biological oncology, offering a powerful, context-specific, and non-invasive tool to address the rising burden of BC in Mexico.

Author Contributions

Conceptualization, T.R.-G., H.G.A.-M. and Á.J.S.-G.; methodology, A.L.L.-L. and J.L.L.-R.; software, A.L.L.-L. and J.L.L.-R.; validation, H.G.A.-M. and A.L.L.-L.; formal analysis, A.L.L.-L. and J.L.L.-R.; investigation, T.R.-G., A.B.-E., H.G.A.-M. and Á.J.S.-G.; resources, T.R.-G., A.B.-E. and G.G.-O.; data curation, T.R.-G. and A.B.-E.; writing—original draft preparation, J.L.L.-R. and T.R.-G.; writing—review and editing, A.B.-E., G.G.-O., H.G.A.-M., A.L.L.-L., Á.J.S.-G. and J.C.P.-A. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

The study was conducted following the Declaration of Helsinki at the Hospital General de Mexico “Dr. Eduardo Liceaga” and the Instituto de Investigaciones Biomédicas (Universidad Nacional Autónoma de México). All protocols were revised and approved by the Ethics and Research Committees of the Hospital General de México “Dr. Eduardo Liceaga” (Approval No: DI/12/111/03/064).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study. Copies of the written consent forms are on file and can be made available upon reasonable request.

Data Availability Statement

All data are available at: https://github.com/jllagunoroque/Risk_Predicction_via_DE-LDA_FE (accessed on 10 April 2026).

Acknowledgments

We are grateful to all the patients who agreed to participate in the study. We also want to acknowledge Gabriela Baltazar Rosario, Carlos Lara, and the administrative staff of the Hospital General de México “Eduardo Liceaga” for their assistance in recruiting female patients. Claudia Vázquez proofread the manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Table A1. Epidemiological profile of Breast Cancer in Mexico (2022). Data presents the distribution of incidence (new cases) and mortality stratified by age group, along with the federal entities reporting the highest rates.
Table A1. Epidemiological profile of Breast Cancer in Mexico (2022). Data presents the distribution of incidence (new cases) and mortality stratified by age group, along with the federal entities reporting the highest rates.
VariableIncidence (New Cases)Mortality (Deaths)
Reported Total31,0438195
By Age Group
20–44 years7839 (23.0%)1162 (13.2%)
40–59 years14,846 (43.5%)3555 (40.5%)
>60 years11,430 (33.5%)4061 (46.3%)
States with Highest Rates1. Colima1. Mexico City
(Top 3)2. Durango2. Nuevo León
3. Jalisco3. Chihuahua
Table A2. Clinical and social characteristics of the participants.
Table A2. Clinical and social characteristics of the participants.
BC (n = 50)
%
BBP (n = 50)
%
H (n = 50)
%
Age group
16–19060
20–2961434
30–39142814
40–49302628
50–59262418
60–691006
>701420
Marital status
Never married163658
Married443432
Free union18164
Divorced or Separated686
Widowed1660
Years of education
0600
1–6522610
7–12304616
13–18121450
>1901424
Family history of cancer
Yes405646
No604454
Menarche
Early 8–10 years122
Normal 11–12 years232728
Late ≥ 13 years262120
Menopause
Yes24710
No264340
Childbirths
Yes372324
No132726
Hormone replacement therapy
Yes25104
No254046
IMC
Normal 18.5–24.9121522
Overweight 25.0–29.9262213
Obesity ≥ 30121215
Table A3. Classification criteria for clinical and psychological study variables.
Table A3. Classification criteria for clinical and psychological study variables.
Clinical VariableCriterion
Age40–69 years = 1Under 40 or over 70 = 0
Family history of cancer (FamHis)Family members with history of cancer = 1Family members with no history of cancer = 0
Overweight (OW)BMI between 25 and 29.9 = 1Higher or lower than that = 0
Obesity (Obe)BMI between 30 and 39.9 = 1Higher or lower than that = 0
Smoking (Taba)Addiction present = 1No addiction = 0
Alcoholism (Alco)Addiction present = 1No addiction = 0
Drug use (Toxi)Addiction present = 1No addiction = 0
Early menarche (EarlyMen)Younger than 12 years = 113 years or older = 0
Late menopause (LateMenop)50 years or older = 149 years or younger = 0
Pregnancies (Preg)No = 0Yes = 1
Vaginal delivery (Da)Presence = 1Absence = 0
Abortion (Abo)Presence = 1Absence = 0
Cesarean section (Cesa)Presence = 1Absence = 0
Age at first delivery (AFD)Younger than 30 =031 years or older = 1
Use of exogenous hormones (EH)Yes = 1No = 0
Breastfeeding (Lac)Did not breastfeed = 1Did breastfeed = 0
Hormone replacement therapy (HRT)Uses HRT = 1Does not use HRT = 0
CECS Evaluation Criteria
VariableCriterion
Low Anger (AL)0–15 = 1Greater than 15 = 0
Middle Anger (AM)16–18 = 1Below 16 or above 18 = 0
High Anger (AH)Greater than or equal to 19 = 1Less than 19 = 0
Low Depression (DL)0–15 = 1Greater than 15 = 0
Middle Depression (DM)16–18 = 1Below 16 or above 18 = 0
High Depression (DH)Greater than or equal to 19 = 1Less than 19 = 0
Low Anxiety (ANL)0–15 = 1Greater than 15 = 0
Middle Anxiety (ANM)16–18 = 1Below 16 or above 18 = 0
High Anxiety (ANH)Greater than or equal to 19 = 1Less than 19 = 0
Low CECS (SL)0–50 = 0Greater than 50 = 0
Middle CECS (SM)51–55 = 1Below 51 or above 55 = 0
High CECS (SH)Greater than or equal to 56= 0Less than 56 = 0
WAI Evaluation Criteria
VariableCriterion
DSL (Low)Less than 47 = 1Above 47 = 0
DSH (High)Less than 47 = 0Above 47 = 1
RL (Low)Less than or equal to 94 = 1Above 94 = 0
RM (Middle)95–107 = 1Below 95 or above 107 = 0
RH (High)Greater than or equal to 108 = 1Less than 108 = 0
RDL (Low)Less than 58 = 1Greater than 58 = 0
RDH (High)Less than 58 = 0Greater than 58 = 1
ISE Evaluation Criteria
Variable Criterion
Low Physical (SFL)0–5 = 1Greater than 5 = 0
Middle Physical (SFM)6–8 = 1Below 6 or above 8 = 0
High Physical (SFH)Greater than or equal to 9 =1Less than 9 = 0
Low Psychological (SPL)0–11 = 1Greater than 11 = 0
Middle Psychological (SPM)12–19 = 1Below 12 or above 19 = 0
High Psychological (SPH)Greater than or equal to 20 = 1Less than 20 = 0
Low Social (SSL)0–5 = 1Greater than 5 = 0
Middle Social (SSM)6–8 = 1Below 6 or above 8 = 0
High Social (SSH)Greater than or equal to 9 = 1Less than 9 = 0
Low Global (SGL)0–22 = 1Less than 22 = 0
Middle Global (SGM)23–32 = 1Below 23 or above 32 = 0
High Global (SGH)Greater than or equal to 33 = 1Less than 33 = 0

Appendix B. Isolated Variables of Each Data Type That May Contribute to Predicting the Risk of Developing Breast Cancer

Since the D E L D A F E   method is based on a linear projection, the magnitude of each coefficient directly reflects the contribution of the corresponding variable to class discrimination, allowing variables to be ordered in descending relevance. As shown in the following tables, the ten most relevant variables overall were identified, along with separate rankings for the five most relevant clinical and psychological variables.
In general, the 10 most relevant variables are:
EPPToxigestaTRHDHSGLANLRDLAHAbo
The 5 most relevant clinical variables are:
EPPToxigestaTRHAbo
The 5 most relevant psychological variables are:
DHSGLANLRDLAH
Table A4. Clinical and psychological data.
Table A4. Clinical and psychological data.
VariableCoefficient
EPP0.17075
Toxi0.14542
gesta0.13314
TR0.12534
DH0.098024
SG0.092286
ANL0.072743
RDL0.066021
AH0.065688
abo0.062628
Table A5. Clinical Data.
Table A5. Clinical Data.
VariableCoefficient
EPP0.17075
Toxi0.14542
gesta0.13314
TRH0.12534
abo0.062628
Table A6. Psychological data.
Table A6. Psychological data.
VariableCoefficient
DH0.098024
SGL0.092286
ANL0.072743
RDL0.066021
AH0.065688
Among the clinical factors, EPP, Toxi, gesta, TRH, and abortion history showed the highest discriminative contribution, while DH, SGL, ANL, RDL, and AH were the most influential psychological variables. These variables consistently appeared among the highest-ranked coefficients across repeated runs, indicating stable relevance in the classification process. This ranked analysis improves interpretability by explicitly linking the predictive model to the factors that contribute most strongly to breast cancer risk discrimination.

Appendix C

Table A7 shows the descriptive statistics for the quantitative psychological variables of the risk prediction instruments (CECS/WAI/ISE). These variables are compared with the five most relevant binarized psychological variables obtained using DE-LDA_FE (DH, SGL, ANL, RDL, AH), which contribute to class discrimination, as described in Appendix B. This allows us to verify whether there was a loss or gain of information when performing the binarization. In this regard, several of the significant variables in Table A7, such as A and An, were also presented using D E L D A F E . While the D E L D A F E model highlighted the importance of Global (SG) and Restraint, even though these quantitative variables were not significant in their descriptive statistics, the results did not coincide only in the case of Distress and Psychological (SP).
Table A7. Descriptive statistics of the CECS/WAI/ISE quantitative variables by pathological group.
Table A7. Descriptive statistics of the CECS/WAI/ISE quantitative variables by pathological group.
VariableBreast Cancer Median (IQR)Benign Breast Pathology Median (IQR)Healthy Median (IQR)p a
Anger (A)18 (7)16 (5)14.5 (5.75)0.008 *
Depression (D)17 (7)19 (7)16.5 (6.75)0.069
Anxiety (An)18 (6)17 (6)15.5 (6.25)0.017 *
Global CECS52 (12)52 (16)47 (17.25)0.022
Distress50.26 (14)44.05 (16.16)33.08 (20.08)<0.001 *
Restraint103.15 (19.02)105.04 (14.04)102.61 (15.1)0.769
Defensive Attitude55.85 (10.98)56.35 (11.41)58.7 (8.55)0.938
Physics (SF)5 (6)6 (6)8 (6)0.104
Psychological (SP)10 (10)16 (12)16 (9)0.047 *
Social (SS)4 (8)6 (9)5.5 (9.75)0.375
Global (SG)20 (28)28 (27)30 (18.75)0.148
Note: a significance of the Kruskal–Wallis Test; * significant variables (p < 0.05); significant variables obtained using the D E L D A F E   model are presented in bold. IQR stands for Interquartile Range.

References

  1. World Health Organization. Data Visualization Tools for Exploring the Global Cancer Burden in 2022; Cancer Today; World Health Organization: Geneva, Switzerland, 2024. [Google Scholar]
  2. Maffuz-Aziz, A.; Labastida-Almendaro, S.; Espejo-Fonseca, A.; Rodríguez-Cuevas, S. Características clinicopatológicas del cáncer de mama en una población de mujeres en México. Cirugía Cir. 2017, 85, 201–207. [Google Scholar] [CrossRef]
  3. Cárdenas-Sánchez, J. Consenso mexicano sobre diagnóstico y tratamiento del cáncer mamario. Mex. J. Oncol. 2022, 20, 6923. [Google Scholar] [CrossRef]
  4. Díaz Godinez, M.M. Estadísticas a Propósito del Día Internacional de la Lucha Contra el Cáncer de Mama (19 de Octubre); INEGI: Aguascalientes, Mexico, 2023. [Google Scholar]
  5. Seimenis, I.; Chouchos, K.; Prassopoulos, P. Radiation Risk Associated with X-Ray Mammography Screening: Communication and Exchange of Information via Tweets. J. Am. Coll. Radiol. 2018, 15, 1033–1039. [Google Scholar] [CrossRef] [PubMed]
  6. Di Maria, S.; Van Nijnatten, T.J.A.; Jeukens, C.R.L.P.N.; Vedantham, S.; Dietzel, M.; Vaz, P. Understanding the Risk of Ionizing Radiation in Breast Imaging: Concepts and Quantities, Clinical Importance, and Future Directions. Eur. J. Radiol. 2024, 181, 111784. [Google Scholar] [CrossRef]
  7. Chaudhary, L.; Knapp, S.; Wen, S.; Xiao, J.; Marano, G.; Kurian, S.; Layne, G.; Jacobson, G.; Abraham, J. Radiation Exposure from Diagnostic Procedures in Patients with Newly Diagnosed Breast Cancer. J. Community Support. Oncol. 2015, 13, 27–29. [Google Scholar] [CrossRef][Green Version]
  8. Chávarri-Guerra, Y.; Villarreal-Garza, C.; Liedke, P.E.; Knaul, F.; Mohar, A.; Finkelstein, D.M.; Goss, P.E. Breast Cancer in Mexico: A Growing Challenge to Health and the Health System. Lancet Oncol. 2012, 13, e335–e343. [Google Scholar] [CrossRef] [PubMed]
  9. Villarreal-Garza, C.; Platas, A.; Miaja, M.; Fonseca, A.; Mesa-Chavez, F.; Garcia-Garcia, M.; Chapman, J.-A.; Lopez-Martinez, E.A.; Pineda, C.; Mohar, A.; et al. Young Women with Breast Cancer in Mexico: Results of the Pilot Phase of the Joven & Fuerte Prospective Cohort. JCO Glob. Oncol. 2020, 6, 395–406. [Google Scholar] [CrossRef] [PubMed]
  10. Avila, M.H. NORMA Oficial Mexicana NOM-041-SSA2-2011, Para la Prevención, Diagnóstico, Tratamiento, Control y Vigilancia Epidemiológica del Cáncer de Mama; Diario Oficial de la Federación: CDMX, México, 2011. [Google Scholar]
  11. Romo-González, T.; Martínez, A.J.; Hernández-Pozo, M.D.R.; Gutiérrez-Ospina, G.; Larralde, C. Psychological Features of Breast Cancer in Mexican Women I: Personality Traits and Stress Symptoms. Adv. Neuroimmune Biol. 2018, 7, 3–15. [Google Scholar] [CrossRef]
  12. Montes-Nogueira, I.; Campos-Uscanga, Y.; Gutiérrez-Ospina, G.; Hernández-Pozo, M.D.R.; Larralde, C.; Romo-González, T. Psychological Features of Breast Cancer in Mexican Women II: The Psychological Network. Adv. Neuroimmune Biol. 2018, 7, 91–105. [Google Scholar] [CrossRef]
  13. Giese-Davis, J.; Wilhelm, F.H.; Conrad, A.; Abercrombie, H.C.; Sephton, S.; Yutsis, M.; Neri, E.; Taylor, C.B.; Kraemer, H.C.; Spiegel, D. Depression and Stress Reactivity in Metastatic Breast Cancer. Psychosom. Med. 2006, 68, 675–683. [Google Scholar] [CrossRef]
  14. Giese-Davis, J.; Koopman, C.; Butler, L.D.; Classen, C.; Cordova, M.; Fobair, P.; Benson, J.; Kraemer, H.C.; Spiegel, D. Change in Emotion-Regulation Strategy for Women with Metastatic Breast Cancer Following Supportive-Expressive Group Therapy. J. Consult. Clin. Psychol. 2002, 70, 916–925. [Google Scholar] [CrossRef]
  15. Greer, S.; Morris, T. Psychological Attributes of Women Who Develop Breast Cancer: A Controlled Study. J. Psychosom. Res. 1975, 19, 147–153. [Google Scholar] [CrossRef]
  16. Temoshok, L. Personality, Coping Style, Emotion and Cancer: Towards an Integrative Model. Cancer Surv. 1987, 6, 545–567. [Google Scholar]
  17. Iwamitsu, Y.; Shimoda, K.; Abe, H.; Tani, T.; Kodama, M.; Okawa, M. Differences in Emotional Distress between Breast Tumor Patients with Emotional Inhibition and Those with Emotional Expression. Psychiatry Clin. Neurosci. 2003, 57, 289–294. [Google Scholar] [CrossRef] [PubMed]
  18. Myers, L.B. The Importance of the Repressive Coping Style: Findings from 30 Years of Research. Anxiety Stress Coping 2010, 23, 3–17. [Google Scholar] [CrossRef] [PubMed]
  19. McGregor, B.A.; Antoni, M.H. Psychological Intervention and Health Outcomes among Women Treated for Breast Cancer: A Review of Stress Pathways and Biological Mediators. Brain Behav. Immun. 2009, 23, 159–166. [Google Scholar] [CrossRef] [PubMed]
  20. García Camacho, A. Módulo Web Para la Obtencion de Riesgo de Padecer Cáncer de Mama Mediante Datos Psicológicos y Clínicos. Bachelor’s Thesis, Universidad Veracruzana, Xalapa, Ver., Mexico, 2017. [Google Scholar]
  21. Hussain, S.; Ali, M.; Naseem, U.; Nezhadmoghadam, F.; Jatoi, M.A.; Gulliver, T.A.; Tamez-Peña, J.G. Breast Cancer Risk Prediction Using Machine Learning: A Systematic Review. Front. Oncol. 2024, 14, 1343627. [Google Scholar] [CrossRef]
  22. Geißler, K.; Koller, T.L.; Ambroladze, A.; Fallenberg, E.M.; Ingrisch, M.; Hahn, H.K. Breast Cancer Risk Prediction Using Background Parenchymal Enhancement, Radiomics, and Symmetry Features on MRI. In Proceedings of the Medical Imaging 2025: Computer-Aided Diagnosis; SPIE: Cergy-Pontoise, France, 2025; Volume 13407, pp. 564–568. [Google Scholar]
  23. Wang, J.; Zhang, Z.; Wang, Y. Utilizing Feature Selection Techniques for AI-Driven Tumor Subtype Classification: Enhancing Precision in Cancer Diagnostics. Biomolecules 2025, 15, 81. [Google Scholar] [CrossRef]
  24. Carriero, A.; De Hond, A.; Cappers, B.; Paulovich, F.; Abeln, S.; Moons, K.G.; Van Smeden, M. Explainable AI in Healthcare: To Explain, to Predict, or to Describe? Diagn. Progn. Res. 2025, 9, 29. [Google Scholar] [CrossRef]
  25. Berahmand, K.; Zhou, X.; Li, Y.; Gururajan, R.; Barua, P.D.; Acharya, U.R.; Chennakesavan, S.K. NEDL-GCP: A Nested Ensemble Deep Learning Model for Gynecological Cancer Risk Prediction. Array 2025, 27, 100468. [Google Scholar] [CrossRef]
  26. Doumari, S.A.; Berahmand, K.; Ebadi, M.J. Early and High-Accuracy Diagnosis of Parkinson’s Disease: Outcomes of a New Model. Comput. Math. Methods Med. 2023, 2023, 1493676. [Google Scholar] [CrossRef]
  27. Romo-González, T.; Barranca-Enríquez, A.; León-Díaz, R.; Del Callejo-Canal, E.; Gutiérrez-Ospina, G.; Jimenez Urrego, A.M.; Bolaños, C.; Botero Carvajal, A. Psychological Suppressive Profile and Autoantibodies Variability in Women Living with Breast Cancer: A Prospective Cross-Sectional Study. Heliyon 2022, 8, e10883. [Google Scholar] [CrossRef]
  28. Benevides-Pereira, A.M.T.; Moreno-Jiménez, B.; Garrosa Hernández, E.; González Gutiérrez, J.L. La evaluación específica del síndrome de burnout en psicólogos: El Inventario de burnout de psicólogos. Clínica Salud 2002, 13, 257–283. [Google Scholar]
  29. Watson, M.; Greer, S. Development of a Questionnaire Measure of Emotional Control. J. Psychosom. Res. 1983, 27, 299–305. [Google Scholar] [CrossRef] [PubMed]
  30. Durá, E.; Andreu, Y.; Galdón, M.J.; Ibáñez, E.; Pérez, S.; Ferrando, M.; Murgui, S.; Martínez, P. Emotional Suppression and Breast Cancer: Validation Research on the Spanish Adaptation of the Courtauld Emotional Control Scale (CECS). Span. J. Psychol. 2010, 13, 406–417. [Google Scholar] [CrossRef]
  31. Weinberger, D.A. The Construct Validity of the Repressive Coping Style. In Repression and Dissociation: Implications for Personality Theory, Psychopathology and Health; Singer, J.L., Ed.; University of Chicago Press: Chicago, IL, USA, 1990; pp. 337–386. [Google Scholar]
  32. Weinberger, D.A.; Schwartz, G.E. Distress and Restraint as Superordinate Dimensions of Self-Reported Adjustment: A Typological Perspective. J. Personal. 1990, 58, 381–417. [Google Scholar] [CrossRef]
  33. Weinberger, D.A. Distress and Self-Restraint as Measures of Adjustment across the Life Span: Confirmatory Factor Analyses in Clinical and Nonclinical Samples. Psychol. Assess. 1997, 9, 132–135. [Google Scholar] [CrossRef]
  34. Romo González, T.; Enríquez-Hernández, C.B.; Hernández Pozo, M.D.R.; Ruiz Montalvo, M.E.; Castillo, R.L.; Ehrenzweig Sánchez, Y.; Marván, M.L.; Larralde, C. Validación en México del inventario de ajuste de Weinberger (WAI). Salud Ment. 2014, 37, 247. [Google Scholar] [CrossRef]
  35. Amir, N.; Beard, C.; Burns, M.; Bomyea, J. Attention Modification Program in Individuals with Generalized Anxiety Disorder. J. Abnorm. Psychol. 2009, 118, 28–33. [Google Scholar] [CrossRef]
  36. Brentnall, A.R.; Harkness, E.F.; Astley, S.M.; Donnelly, L.S.; Stavrinos, P.; Sampson, S.; Fox, L.; Sergeant, J.C.; Harvie, M.N.; Wilson, M.; et al. Mammographic Density Adds Accuracy to Both the Tyrer-Cuzick and Gail Breast Cancer Risk Models in a Prospective UK Screening Cohort. Breast Cancer Res. 2015, 17, 147. [Google Scholar] [CrossRef] [PubMed]
  37. Terry, M.B.; Liao, Y.; Whittemore, A.S.; Leoce, N.; Buchsbaum, R.; Zeinomar, N.; Dite, G.S.; Chung, W.K.; Knight, J.A.; Southey, M.C.; et al. 10-Year Performance of Four Models of Breast Cancer Risk: A Validation Study. Lancet Oncol. 2019, 20, 504–517. [Google Scholar] [CrossRef] [PubMed]
  38. Hassija, V.; Chamola, V.; Mahapatra, A.; Singal, A.; Goel, D.; Huang, K.; Scardapane, S.; Spinelli, I.; Mahmud, M.; Hussain, A. Interpreting Black-Box Models: A Review on Explainable Artificial Intelligence. Cogn. Comput. 2024, 16, 45–74. [Google Scholar] [CrossRef]
Figure 1. Schematic representation of the Biopsychosocial Risk Prediction Model. This flow diagram illustrates the four computational stages of the study: (A) Data Acquisition: Collection of biological, reproductive, and psychological variables from the study population (n = 150). (B) Preprocessing: Transformation of raw data into binary feature vectors (0, 1) to standardize input. (C) Feature Extraction: Application of the Differential Evolutionary Linear Discriminant Analysis ( D E L D A F E ) to identify the optimal projection matrix (W) that maximizes class separability. (D) Classification: Probabilistic assignment of new cases into Healthy, Benign Pathology, or Breast Cancer categories based on the classification models.
Figure 1. Schematic representation of the Biopsychosocial Risk Prediction Model. This flow diagram illustrates the four computational stages of the study: (A) Data Acquisition: Collection of biological, reproductive, and psychological variables from the study population (n = 150). (B) Preprocessing: Transformation of raw data into binary feature vectors (0, 1) to standardize input. (C) Feature Extraction: Application of the Differential Evolutionary Linear Discriminant Analysis ( D E L D A F E ) to identify the optimal projection matrix (W) that maximizes class separability. (D) Classification: Probabilistic assignment of new cases into Healthy, Benign Pathology, or Breast Cancer categories based on the classification models.
Mca 31 00066 g001
Figure 2. Performance comparison of classification models based on clinical data. The graph contrasts the predictive performance of each algorithm when utilizing raw clinical variables (baseline) versus the optimized projections derived from the D E L D A F E method. Abbreviations: K-nn: K-nearest neighbor; Linear: Linear Discriminant Classifier; DT: Decision Tree; SVM: Support Vector Machine; ANN: Artificial Neural Network.
Figure 2. Performance comparison of classification models based on clinical data. The graph contrasts the predictive performance of each algorithm when utilizing raw clinical variables (baseline) versus the optimized projections derived from the D E L D A F E method. Abbreviations: K-nn: K-nearest neighbor; Linear: Linear Discriminant Classifier; DT: Decision Tree; SVM: Support Vector Machine; ANN: Artificial Neural Network.
Mca 31 00066 g002
Figure 3. Comparison of classification models based on psychological data. The graph contrasts the predictive performance of each algorithm when utilizing raw psychological variables (baseline) versus the optimized projections derived from the D E L D A F E method. Abbreviations: K-nn: K-nearest neighbor; Linear: Linear Discriminant Classifier; DT: Decision Tree; SVM: Support Vector Machine; ANN: Artificial Neural Network.
Figure 3. Comparison of classification models based on psychological data. The graph contrasts the predictive performance of each algorithm when utilizing raw psychological variables (baseline) versus the optimized projections derived from the D E L D A F E method. Abbreviations: K-nn: K-nearest neighbor; Linear: Linear Discriminant Classifier; DT: Decision Tree; SVM: Support Vector Machine; ANN: Artificial Neural Network.
Mca 31 00066 g003
Figure 4. Comparison of classification models based on combined clinical and psychological data. The graph contrasts the predictive performance of each algorithm when utilizing raw integrated variables (baseline) versus the optimized projections derived from the D E L D A F E method. Abbreviations: K-nn: K-nearest neighbor; Linear: Linear Discriminant Classifier; DT: Decision Tree; SVM: Support Vector Machine; ANN: Artificial Neural Network.
Figure 4. Comparison of classification models based on combined clinical and psychological data. The graph contrasts the predictive performance of each algorithm when utilizing raw integrated variables (baseline) versus the optimized projections derived from the D E L D A F E method. Abbreviations: K-nn: K-nearest neighbor; Linear: Linear Discriminant Classifier; DT: Decision Tree; SVM: Support Vector Machine; ANN: Artificial Neural Network.
Mca 31 00066 g004
Figure 5. Performance comparison of classification models using raw datasets. The graph contrasts the precision of each classifier across the raw clinical variables, raw psychological variables, and the unprojected combination of both domains. Abbreviations: K-nn: K-nearest neighbor; Linear: Linear Discriminant Classifier; DT: Decision Tree; SVM: Support Vector Machine; ANN: Artificial Neural Network.
Figure 5. Performance comparison of classification models using raw datasets. The graph contrasts the precision of each classifier across the raw clinical variables, raw psychological variables, and the unprojected combination of both domains. Abbreviations: K-nn: K-nearest neighbor; Linear: Linear Discriminant Classifier; DT: Decision Tree; SVM: Support Vector Machine; ANN: Artificial Neural Network.
Mca 31 00066 g005
Figure 6. Performance comparison of classification models using D E L D A F E data projections. The graph contrasts the classification accuracy across the projected clinical space, projected psychological space, and the optimized projection of the combined variables. Abbreviations: K-nn: K-nearest neighbor; Linear: Linear Discriminant Classifier; DT: Decision Tree; SVM: Support Vector Machines; ANN: Artificial Neural Network.
Figure 6. Performance comparison of classification models using D E L D A F E data projections. The graph contrasts the classification accuracy across the projected clinical space, projected psychological space, and the optimized projection of the combined variables. Abbreviations: K-nn: K-nearest neighbor; Linear: Linear Discriminant Classifier; DT: Decision Tree; SVM: Support Vector Machines; ANN: Artificial Neural Network.
Mca 31 00066 g006
Table 1. Parameter settings used for the D E L D A F E method across the experiments for each dataset.
Table 1. Parameter settings used for the D E L D A F E method across the experiments for each dataset.
ParameterClinicalPsychologicalBoth
Population size (N)20002000300
Generations (NG)300300700
Crossover rate (CR)0.50.50.5
Mutation rate (F)0.30.30.7
Table 2. Performance metrics of classification models using raw clinical data.
Table 2. Performance metrics of classification models using raw clinical data.
Classification ModelsTP RATEFP RATEPrecisionRecallF1-ScoreAUC-ROC
K-nn (3)0.65480.36730.74660.65180.66900.8676
K-nn (5)0.63260.37420.69470.59720.60560.8300
K-nn (7)0.66260.43090.65470.58110.58570.7790
K-nn (9)0.61160.39900.61670.55980.55750.7313
K-nn (11)0.61640.40990.57200.51670.50700.7263
Linear0.69440.39780.64710.59510.60760.7978
Gaussian0.78940.37570.86280.79370.80640.9528
DT0.72850.32720.79390.69690.71650.8966
SVM0.68270.40070.68920.61440.62550.7856
ANN (30)0.85190.25200.91440.89640.90340.9892
Classification models: K-Nearest Neighbors (K-nn), Linear Discriminant Classifier (Linear), Gaussian Classifier (Gaussian), Decision Tree (DT), Support Vector Machine (SVM), Artificial Neural Network (ANN). The optimal ANN architecture consists of two hidden layers with 30 neurons each.
Table 3. Average performance metrics of classification models using clinical variable projections via D E L D A F E   method.
Table 3. Average performance metrics of classification models using clinical variable projections via D E L D A F E   method.
Classification ModelsTP RATEFP RATEPrecisionRecallF1-ScoreAUC-ROC
K-nn (3)0.68350.34160.78100.70960.72730.8937
K-nn (5)0.65490.34780.71410.65530.66730.8586
K-nn (7)0.68230.39840.69260.62620.63700.8228
K-nn (9)0.67050.38120.67950.60100.61370.8039
K-nn (11)0.66010.38850.66840.58790.59970.7961
Linear0.69140.39930.65040.59860.61010.7940
Gaussian0.67950.39930.67450.60310.61590.7808
DT0.73600.32270.80430.77330.78310.9089
SVM0.69020.39530.66450.58220.59520.7993
ANN (400)0.82380.32260.91110.89110.89900.9875
Classification models: K-Nearest Neighbors (K-nn), Linear Discriminant Classifier (Linear), Gaussian Classifier (Gaussian), Decision Tree (DT), Support Vector Machine (SVM), Artificial Neural Network (ANN). The optimal ANN architecture comprises two hidden layers with 400 neurons per layer.
Table 4. Best individual run performance for each classification model using clinical variable projections.
Table 4. Best individual run performance for each classification model using clinical variable projections.
Classification ModelsExperimentTP RATEFP RATEPrecisionRecallF1-ScoreAUC-ROC
K-nn (3)100.68950.33960.80080.71730.73800.8966
K-nn (5)20.64620.34320.75870.69220.70800.8520
K-nn (7)100.67870.39330.74010.65040.66440.8235
K-nn (9)30.67650.39100.70120.63430.64920.8059
K-nn (11)80.67520.38290.69390.62920.64550.8148
Linear10.69350.39950.67030.60410.62040.7967
Gaussian80.67840.40040.68650.60920.62500.7797
DT80.72680.32460.82340.78420.79350.9125
SVM70.68900.39720.67790.59070.60490.7981
ANN (400)30.83790.28290.91490.89470.90240.9894
Classification models: K-Nearest Neighbors (K-nn), Linear Discriminant Classifier (Linear), Gaussian Classifier (Gaussian), Decision Tree (DT), Support Vector Machine (SVM), Artificial Neural Network (ANN).
Table 5. Performance metrics of classification models using raw psychological data.
Table 5. Performance metrics of classification models using raw psychological data.
Classification ModelTP RATEFP RATEPrecisionRecallF1-ScoreAUC-ROC
K-nn (3)0.60910.42520.56890.51870.52570.7865
K-nn (5)0.59960.44160.53220.47720.45750.7487
K-nn (7)0.66010.50530.49420.46680.45020.7129
K-nn (9)0.60490.47930.47350.44550.43300.6717
K-nn (11)0.62430.51840.47810.45260.43850.6456
Linear0.70050.40110.59990.56560.57170.8036
Gaussian0.68700.39360.47800.62550.54190.7983
DT0.71280.37350.74020.72940.72980.8796
SVM0.72230.39120.69740.61580.62630.8362
ANN (20)0.89330.17791.00001.00001.00001.0000
Classification models: K-Nearest Neighbors (K-nn), Linear Discriminant Classifier (Linear), Gaussian Classifier (Gaussian), Decision Tree (DT), Support Vector Machine (SVM), Artificial Neural Network (ANN). The optimal ANN architecture comprises one hidden layer with 20 neurons per layer.
Table 6. Average performance metrics of classification models using psychological variable projections via D E L D A F E method.
Table 6. Average performance metrics of classification models using psychological variable projections via D E L D A F E method.
Classification ModelTP RATEFP RATEPrecisionRecallF1-ScoreAUC-ROC
K-nn (3)0.69140.33900.77080.72380.73670.9129
K-nn (5)0.66980.33930.69810.65810.66670.8714
K-nn (7)0.70040.38440.66700.64010.64500.8492
K-nn (9)0.68790.38760.65350.61940.62480.8350
K-nn (11)0.68920.41290.64220.60730.61250.8251
Linear0.69350.40410.59410.57340.57740.7934
Gaussian0.69450.40310.61470.58800.59380.7954
DT0.74990.30080.82240.81690.81490.9339
SVM0.70000.40100.65810.59180.59830.8031
ANN (500)0.89300.18761.00001.00001.00001.0000
Classification models: K-Nearest Neighbors (K-nn), Linear Discriminant Classifier (Linear), Gaussian Classifier (Gaussian), Decision Tree (DT), Support Vector Machine (SVM), Artificial Neural Network (ANN). The optimal ANN architecture comprises two hidden layers with 500 neurons per layer.
Table 7. Best individual run performance for each classification model using psychological variable projections.
Table 7. Best individual run performance for each classification model using psychological variable projections.
Classification ModelExperimentTP RATEFP RATEPrecisionRecallF1-ScoreAUC-ROC
K-nn (3)70.69320.35220.82120.73530.75650.9150
K-nn (5)30.69480.33670.76050.70970.72060.8953
K-nn (7)70.71260.38670.73130.69420.70440.8520
K-nn (9)80.69460.37610.70350.68950.69460.8593
K-nn (11)70.70360.41660.66610.64010.64760.8370
Linear60.69980.40140.63090.59640.60450.8023
Gaussian30.69750.40150.64270.61960.62460.7999
DT30.77800.25620.86680.87780.86820.9640
SVM70.70030.39980.70590.60630.61790.8046
ANN (500)10.90820.14981.00001.00001.00001.0000
Classification models: K-Nearest Neighbors (K-nn), Linear Discriminant Classifier (Linear), Gaussian Classifier (Gaussian), Decision Tree (DT), Support Vector Machine (SVM), Artificial Neural Network (ANN).
Table 8. Performance metrics of classification models using raw combined clinical and psychological data.
Table 8. Performance metrics of classification models using raw combined clinical and psychological data.
Classification ModelTP RATEFP RATEPrecisionRecallF1-ScoreAUC-ROC
K-nn (3)0.63690.39980.72590.64040.65130.8510
K-nn (5)0.60260.41450.63400.54320.54680.7579
K-nn (7)0.61670.45620.56320.50150.49530.6959
K-nn (9)0.57430.45560.48330.42370.40300.6510
K-nn (11)0.59480.48820.49590.43270.40360.6378
Linear0.78120.36230.79310.77620.78340.9227
Gaussian0.00000.00000.00000.00000.00000.0000
DT0.72150.32930.77700.76290.76790.9053
SVM0.79580.35500.89290.85540.86930.9448
ANN (20)0.87650.22491.00001.00001.00001.0000
Classification models: K-Nearest Neighbors (K-nn), Linear Discriminant Classifier (Linear), Gaussian Classifier (Gaussian), Decision Tree (DT), Support Vector Machine (SVM), Artificial Neural Network (ANN). The optimal ANN architecture comprises one hidden layer with 20 neurons per layer.
Table 9. Projection results of the clinical and psychological variables obtained with the D E L D A F E method.
Table 9. Projection results of the clinical and psychological variables obtained with the D E L D A F E method.
Classification ModelTP RATEFP RATEPrecisionRecallF1-ScoreAUC-ROC
K-nn (3)0.74540.27360.84350.83840.84000.9586
K-nn (5)0.74390.25780.79830.77000.78020.9375
K-nn (7)0.76770.27510.78740.77240.77840.9236
K-nn (9)0.76100.27980.77900.77010.77380.9208
K-nn (11)0.75470.29240.76260.75440.75750.9188
Linear0.78150.36210.79000.77160.77920.9232
Gaussian0.78490.36040.79560.78170.78760.9283
DT0.76200.30640.88740.88670.88530.9593
SVM0.78300.36130.78800.76320.77230.9254
ANN (500)0.90730.14510.99750.99770.99761.0000
Classification models: K-Nearest Neighbors (K-nn), Linear Discriminant Classifier (Linear), Gaussian Classifier (Gaussian), Decision Tree (DT), Support Vector Machine (SVM), Artificial Neural Network (ANN). The optimal ANN architecture comprises two hidden layers with 500 neurons per layer.
Table 10. Best individual run performance for each classification model using combined data projections via D E L D A F E .
Table 10. Best individual run performance for each classification model using combined data projections via D E L D A F E .
Classification ModelExperimentTP RATEFP RATEPrecisionRecallF1-ScoreAUC-ROC
K-nn (3)10.74780.27480.86630.86300.86430.9636
K-nn (5)40.75420.25280.82390.79560.80620.9459
K-nn (7)20.77160.27230.81600.78470.79510.9276
K-nn (9)90.76330.27590.81600.80520.81000.9277
K-nn (11)20.75740.29410.78680.77050.77720.9220
Linear20.78230.36160.80080.78530.79180.9246
Gaussian60.78620.35980.80370.79240.79730.9303
DT50.77830.25780.91480.90990.90770.9730
SVM100.78420.36050.79910.77760.78600.9275
ANN (500)30.87390.22051.00001.00001.00001.0000
Classification models: K-Nearest Neighbors (K-nn), Linear Discriminant Classifier (Linear), Gaussian Classifier (Gaussian), Decision Tree (DT), Support Vector Machine (SVM), Artificial Neural Network (ANN).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Llaguno-Roque, J.L.; López-Lobato, A.L.; Pérez-Arriaga, J.C.; Acosta-Mesa, H.G.; Sánchez-García, Á.J.; Gutiérrez-Ospina, G.; Barranca-Enríquez, A.; Romo-González, T. AI-Driven Biopsychosocial Screening for Breast Cancer: Enhancing Risk Prediction via Differential Evolutionary Linear Discriminant Analysis for Feature Extraction. Math. Comput. Appl. 2026, 31, 66. https://doi.org/10.3390/mca31030066

AMA Style

Llaguno-Roque JL, López-Lobato AL, Pérez-Arriaga JC, Acosta-Mesa HG, Sánchez-García ÁJ, Gutiérrez-Ospina G, Barranca-Enríquez A, Romo-González T. AI-Driven Biopsychosocial Screening for Breast Cancer: Enhancing Risk Prediction via Differential Evolutionary Linear Discriminant Analysis for Feature Extraction. Mathematical and Computational Applications. 2026; 31(3):66. https://doi.org/10.3390/mca31030066

Chicago/Turabian Style

Llaguno-Roque, José Luis, Adriana Laura López-Lobato, Juan Carlos Pérez-Arriaga, Héctor Gabriel Acosta-Mesa, Ángel J. Sánchez-García, Gabriel Gutiérrez-Ospina, Antonia Barranca-Enríquez, and Tania Romo-González. 2026. "AI-Driven Biopsychosocial Screening for Breast Cancer: Enhancing Risk Prediction via Differential Evolutionary Linear Discriminant Analysis for Feature Extraction" Mathematical and Computational Applications 31, no. 3: 66. https://doi.org/10.3390/mca31030066

APA Style

Llaguno-Roque, J. L., López-Lobato, A. L., Pérez-Arriaga, J. C., Acosta-Mesa, H. G., Sánchez-García, Á. J., Gutiérrez-Ospina, G., Barranca-Enríquez, A., & Romo-González, T. (2026). AI-Driven Biopsychosocial Screening for Breast Cancer: Enhancing Risk Prediction via Differential Evolutionary Linear Discriminant Analysis for Feature Extraction. Mathematical and Computational Applications, 31(3), 66. https://doi.org/10.3390/mca31030066

Article Metrics

Back to TopTop