1. Introduction
Breast cancer (BC) is the most prevalent neoplasm in women worldwide. Nationally, it has been the most frequent and deadliest cancer affecting Mexican women since 2006 (
Table A1 in
Appendix A) [
1,
2,
3,
4]. With projections estimating over 50,000 new annual cases by 2050, BC represents an urgent and expanding public health challenge. There is broad consensus that early detection is essential to reduce morbidity and mortality. While mammography remains the gold standard for screening programs worldwide, it presents limitations. Widespread mammographic screening is often restricted by age guidelines and carries concerns regarding radiation exposure, particularly for younger populations [
5,
6,
7]. This is especially relevant in Mexico, where advanced-stage BC is frequently diagnosed in women aged 20–41 [
8,
9]. Since 2013, our group has developed non-radiographic resources to identify young Mexican women at elevated risk, aiming to optimize early detection while minimizing unnecessary radiation exposure.
Although the Official Mexican Standard [
10] acknowledges that biological, iatrogenic, reproductive, and lifestyle factors interplay to determine BC relative risk, current national screening programs rely almost exclusively on biological and reproductive history. This approach has not substantially reduced incidence or mortality rates, likely because these factors explain, at best, 50% of BC prevalence. We have recently suggested that the accuracy of risk identification would increase by incorporating psycho-affective factors into screening designs [
11,
12]. Evidence suggests that specific personality traits—such as low emotional containment, low defensiveness, and high stress—are consistently observed in Mexican women with BC prior to diagnosis [
11]. Physiologically, emotional suppression and repression are known to alter diurnal cortisol secretion, a dysregulation associated with early mortality in metastatic BC patients [
13,
14]. Furthermore, principal component analyses have demonstrated that healthy women and those with benign or malignant breast lesions cluster differentially when emotional suppression, repression, and stress symptoms are included in the phenotyping [
12]. Notably, when clinical and psychological variables are analyzed together, emotional suppression variables significantly define differential group clustering. These observations support the hypothesis that screening for psychological factors is crucial for prevention, early diagnosis, and understanding BC pathogenesis [
15,
16,
17,
18,
19].
To improve risk identification in young Mexican women, García-Camacho et al. [
20] previously designed a decision tree-based model assessing the Stress Symptom Inventory (ISE), the Courtauld Emotional Control Suppression Scale (CECS), and the Weinberger Adjustment Inventory (WAI). Their results showed that while these tools achieved low accuracy individually (34–40%), their combination improved prediction to 42.7%. This suggests that while the emotional dimension has discriminatory value, the modeling approach required refinement to achieve clinical utility.
On the other hand, recent advances in biopsychosocial risk prediction and explainable AI highlight the importance of integrating multi-domain data for oncology applications. Hussain et al. [
21] reviewed machine learning approaches for breast cancer risk prediction, emphasizing the added value of psychosocial variables in enhancing model accuracy. Similarly, radiomics-based studies have demonstrated that multimodal feature fusion, such as parenchymal enhancement and breast symmetry in MRI, improves discrimination power in early detection [
22]. In parallel, evolutionary feature extraction methods have been increasingly applied in oncology, as shown by Wang et al. [
23], who underscored the role of metaheuristic optimization in refining tumor subtype classification. Explainable AI (XAI) frameworks, such as those discussed by Carriero et al. [
24], further stress the necessity of transparency in clinical models, particularly when integrating heterogeneous data sources.
Beyond breast cancer, comparable strategies have been successfully deployed in other medical domains. The NEDL-GCP model [
25] introduced a nested ensemble deep learning framework for gynecological cancer risk prediction, demonstrating how ensemble architectures can robustly manage complex clinical datasets. Doumari et al. [
26] proposed a novel high-accuracy model for the early diagnosis of Parkinson’s disease, illustrating the potential of feature-fusion and classifier optimization in neurodegenerative contexts.
Building on this foundation, we present a new risk assessment model that employs Differential Evolutionary Linear Discriminant Analysis for Feature Extraction (
). This approach integrates biological, reproductive, and psychological variables to manage the complexity of the dataset. Unlike previous models,
simplifies the observation of risk classes and reveals the synergistic behavior of the data. Furthermore, this work is distinguished by its focus on the Mexican population and its key innovation: the use of psychological data alongside clinical data—an approach that, unlike our group [
11,
12,
27], no other previous study in the country has integrated into breast cancer risk assessment, a gap that this work seeks to fill. Although the analyzed cohort consists of 110 individuals, it accurately represents the psychological and clinical information of Mexican women, a group in which identifying risk factors for breast cancer is a priority.
The method was used to improve classification using machine learning algorithms, prioritizing those that offer explainability in their results. This approach lays the groundwork for future applications, such as the development of a web module that estimates breast cancer risk in an accessible, non-invasive, highly accurate, and timely manner.
2. Methodology
2.1. Participants and Study Design
This study utilized clinical and psychological data from 150 women attending routine gynecological consultations at the General Hospital of Mexico “Dr. Eduardo Liceaga.” All participants provided written informed consent, ensuring confidentiality and anonymity in accordance with the Declaration of Helsinki (printed in the British Medical Journal, 18 July 1964).
The sample was selected to represent a broad age range (see
Table A2 in
Appendix A), reflecting the epidemiological reality of BC in Mexican women; it is frequently diagnosed approximately a decade earlier than global averages [
5]. Participants were recruited prior to histopathological diagnosis and were subsequently stratified into three groups based on their clinical pathology reports. The Healthy Group (H,
n = 50) clustered women with no palpable or radiological breast pathology. The Benign Breast Pathology Group (BBP,
n = 50) was formed by women diagnosed with fibrocystic changes, fibroadenomas, or mastitis. The Breast Cancer Group (BC,
n = 50) included women diagnosed with infiltrating ductal carcinoma (naïve to treatment).
2.2. Data Collection Instruments
Physical, hereditary, and lifestyle data were collected using a 57-item General Data Questionnaire. Psychological profiles—specifically emotional repression, suppression, and stress symptomatology—were assessed using three validated instruments:
2.2.1. Inventory of Stress Symptoms (ISE)
Designed by Benevides-Pereira et al. [
28], the ISE evaluates the frequency of stress manifestations in daily life across three dimensions: psychological, physical, and social symptoms. It consists of 30 items rated on a 5-point Likert scale (0 = never to 4 = always). Global symptomatology scores are categorized as low (0–22), medium (22–33), or high (>33).
2.2.2. Courtauld Emotional Control Scale (CECS)
Developed by Watson and Greer [
29] and validated for Spanish speakers by Durá et al. [
30]. The CECS assesses the extent to which individuals suppress reactions to negative emotions. It comprises 21 items across three subscales: Anger, Anxiety, and Depression. Higher scores indicate greater emotional suppression. The instrument demonstrates high internal consistency (α = 0.95). For this study, a median split (score = 32) was used to categorize participants into “Emotional Suppression” (score ≥ 32) or “Emotional Expression” (score < 32) groups.
2.2.3. The Weinberger Adjustment Inventory (WAI)
The WAI [
31,
32,
33] evaluates adjustment styles through three primary scales: Subjective Distress, Restraint, and Defensiveness. The interactions between these scales allow the identification of six typologies of adjustment, including the Repressive (Type C) coping style. The Spanish version validated for the Mexican population [
34] was used (Cronbach’s α ranges from 0.69 to 0.89 across subscales).
2.3. Feature Encoding and Binarization
Given the binary nature of the clinical data, and in order to compare them with the results of the psychological instruments, the latter were transformed into dichotomous variables to assess the presence or absence of certain psychological traits in the groups of women. The psychological data, originally expressed as continuous scores or subscales, were categorized into low, medium, and high levels according to their values. Subsequently, these categories were recoded into binary variables (1 for the presence of the characteristic; 0 for its absence), using cut-off points previously defined and described in Romo-González et al. [
27]. In most cases, the original cut-off points for each instrument were retained; however, in the case of the CECS, these were established based on 95% confidence intervals. The specific binarization criteria are detailed in
Table A3 of
Appendix A.
Appendix C examines the potential loss or gain of information resulting from the binarization of psychological data.
2.4. Machine Learning Identification Model
To manage dataset complexity and achieve optimal separation among the three groups (H, BBP, and BC), a supervised learning approach, Differential Evolutionary Linear Discriminant Analysis for Feature Extraction (), was employed. This method performs dimensionality reduction while enhancing class separability by combining linear transformations of the original variables with evolutionary optimization to obtain a projection matrix that maximizes class separation in a two-dimensional space. The following sections describe the main components of this method.
2.4.1. Projection-Based Dimensionality Reduction
Let
be a dataset with
samples and
features. The goal is to obtain a reduced representation
, with
, defined as:
where
is the projection matrix whose columns are orthonormal vectors. This transformation produces new features that are linear combinations of the original variables, while orthonormality ensures that the columns form a valid basis for the projected space.
2.4.2. Differential Evolution Optimization
Differential Evolution (DE) is a population-based optimization method that iteratively refines a set of candidate solutions using mutation, crossover, and selection operators.
At iteration (generation)
, the population consists of
candidate solutions (individuals) randomly initialized within the domain of a fitness function
. Each individual is represented by a vector
, where
denotes the dimensionality of the problem. In the DE-LDA_{FE} framework, each individual corresponds to a vector
that encodes a candidate projection matrix
, with
, as presented in Equation (2). The particularity of these matrixes is that they are orthonormal by columns, to define the projection directions.
The DE/best/1/bin strategy is used. For each target vector
in the population, a mutant vector is generated as:
where
is the current best solution,
are random individuals, and
is the mutation factor.
A trial vector is then created via binomial crossover:
where
is the crossover rate, and
corresponds to a randomly selected position. The selection step retains the solution with the highest fitness between the trial vector
and the target vector
.
These steps are repeated until a stopping criterion is met, such as reaching a maximum number of generations or population convergence. Thus, through this iterative evolutionary process, the population of candidate solutions is progressively refined by combining and selecting the best individuals at each generation. As a result, the algorithm converges toward projection matrices that increasingly improve class separability, enabling the discovery of a near-optimal solution in the lower-dimensional space.
2.4.3. Fitness Function
For the
method, the quality of each projection is evaluated using Fisher’s discriminant criterion, which seeks to maximize the separation between class means while minimizing the dispersion within each class. This is achieved by maximizing the ratio between the between-class scatter and the within-class scatter in the projected space, as presented in Equation (5).
where
and
represent the between-class and within-class scatter matrices, respectively, as presented in Equations (6) and (7), where
denotes the number of classes,
the mean of class
, and
the global mean.
This criterion promotes projections in which samples from the same class are tightly clustered and those from different classes are well separated in the reduced space.
2.5. Evaluation of Classifiers with/Without the Algorithm
In this research, to analyze the prediction of breast cancer risk by integrating clinical and psychological variables, a comparison of six classification models is performed: K-nearest neighbors (K-nn; k = 3, 5, 7, 9, and 11); Linear Discriminant Analysis (Linear); Gaussian classifier (Gaussian); Decision Tree (DT); Support Vector Machine (SVM); and Artificial Neural Network (ANN). This comparison is made by considering the raw data of the clinical variables, the psychological variables, and finally, both. In addition, the feature extraction method, Feature Extraction (
), is used to graphically analyze classification performance and improve the classifier’s accuracy., as presented in
Figure 1.
To determine the most suitable Artificial Neural Network (ANN) configuration, a systematic experimental evaluation was conducted using feedforward fully connected neural networks for multiclass classification. The tested architectures consisted of 1 to 5 hidden layers, each with the same number of neurons. The number of neurons per layer was varied across the set {10, 20, 30, 40, 50, 100, 200, 300, 400, 500}, generating a broad range of shallow and deep network configurations. In all cases, each network received the dataset as input and produced class predictions at the output layer. This structured exploration enabled comparison of architectures with different depths and capacities to identify the configuration that achieved the best classification performance according to the selected evaluation metrics.
For the
method, a grid search strategy was employed by systematically exploring combinations of the Differential Evolution (DE) parameters, including population size
, number of generations
, crossover rate
, and scale factor
. The best classification values were obtained with the configurations shown in
Table 1 for each data type.
3. Results
This section presents the comparative performance of the selected classification models in predicting BC risk using clinical, psychological, and integrated datasets. Following the exclusion of records with missing data, a final cohort of 110 patients (out of the original 150) was included in the analysis.
Model performance was assessed using both the raw data and data obtained after using the feature extraction method. To rigorously evaluate the classifiers, standard performance metrics were utilized, including True Positive (TP) Rate, False Positive (FP) Rate, Accuracy, Recall, F1-score, and the Area Under the Receiver Operating Characteristic Curve (AUC-ROC).
For the experiments, the classification performance was evaluated on the training set. This choice was made because the primary goal of the proposed method is to compare the discriminative quality of different feature representations rather than to estimate generalization performance. Although this approach may yield optimistic accuracy values, it ensures a fair and controlled comparison across methods, since all models are evaluated under identical conditions. Additionally, the dataset’s limited size supports the use of this evaluation strategy.
For the raw variables, the results presented consider only one execution, since the dataset employed does not vary; however, for the projections of the variables with the method, since it is a stochastic, population-based metaheuristic algorithm, the projections are not unique, so the results presented are the mean of 10 runs, and, to perform a fair comparison, the best solution obtained on these 10 runs is employed for comparison with the results of the raw data.
3.1. Classification Based on Clinical Variables
3.1.1. Analysis of Raw Clinical Data
In this initial experimental phase, the classification models were trained using the clinical variables in their original, unprocessed state. This baseline analysis serves to evaluate the inherent discriminatory power of standard biological and reproductive risk factors prior to the application of any feature extraction or dimensionality reduction techniques. The comparative performance metrics for each classifier in this scenario are detailed in
Table 2.
As shown in
Table 1, the performance of classifiers trained on raw clinical data varied significantly depending on the algorithm’s architecture. The ANN with 30 neurons demonstrated superior predictive capability, achieving the highest scores across all metrics, with an F1-score of 0.9034 and an AUC-ROC of 0.9892. The Gaussian classifier also performed robustly, securing the second-best performance (F1-score: 0.8064).
Conversely, the K-nn algorithm exhibited an inverse relationship between the number of neighbors (K) and model performance; as K increased from 3 to 11, the F1-score progressively declined from 0.6690 to 0.5070. Linear and SVM models yielded moderate results (F1-scores ≈ 0.60–0.62), suggesting that the boundaries between risk groups in the raw clinical space are non-linear and complex, requiring more sophisticated modeling approaches like ANNs to be effectively decoded.
3.1.2. Classification Using Clinical Data Projections via Method
As established in previous sections, the
method was implemented to improve class separation by projecting the original variables into an optimized, lower-dimensional latent space. This projection aimed to enhance the discriminative performance of the classification models. The average performance metrics obtained using these clinical projections, calculated over 10 independent runs, are presented in
Table 3.
As indicated in
Table 3, the models demonstrating the highest predictive capability when utilizing the projected clinical variables were the ANN architecture with 400 neurons (Precision: 0.9111; F1-score: 0.8990) and the DT model (Precision: 0.8043; F1-score: 0.7831). The remaining models exhibited moderate, comparable performances. Consistent with the raw data findings, the K-nn model displayed a decline in both Precision and F1-score as the number of neighbors (K) increased, confirming that K = 3 remains the optimal configuration for this spatial mapping.
While
Table 3 reflects the average robustness across 10 iterations, it is also highly relevant to identify the peak potential of each classifier.
Table 4 details the single best-performing run for each model based on the clinical projections.
The results in
Table 4 mirror the trends observed in the average metrics, reaffirming the superiority of the ANN model. Notably, the peak performance was achieved during Experiment 3 using an expanded architecture of 500 neurons, which yielded an exceptional Precision of 0.9149 and an F1-score of 0.9024. The DT model similarly achieved its peak during Experiment 8, solidifying its position as the second most effective classifier for these specific clinical projections.
3.1.3. Performance Comparison: Raw vs. Projected Clinical Data
Figure 2 illustrates the comparative performance of the classification models when trained on raw clinical variables versus the optimized latent projections generated by the
method. As observed, for K-nn, DT, SVM, and Linear, the application of the
method successfully enhanced class separability, thereby contributing to an overall improvement in classification metrics.
However, two notable exceptions emerged: Gaussian and ANN classifiers achieved their peak performance using the raw, not projected clinical data. This outcome suggests that these specific non-linear algorithms possess an inherent capacity to map the complex, multi-dimensional relationships of biological risk factors in their original state, without requiring prior dimensionality reduction or feature extraction.
3.2. Classification Based on Psychological Variables
3.2.1. Analysis of Raw Psychological Variables
In this phase of the analysis, the classification models were trained using the raw psychological variables exclusively (without any preprocessing or projection). This experiment establishes a baseline to evaluate the inherent predictive power of psychological factors, such as emotional repression, suppression, and stress, when analyzed in isolation. The comparative performance metrics for each classifier in this raw psychological space are presented in
Table 5.
As shown in
Table 5, the classification models exhibited varying degrees of success when processing the psychological variables. Most notably, the ANN, configured with an optimized hidden layer of 20 neurons, achieved extraordinary, perfect predictive metrics across the board (Precision: 1.0000; F1-score: 1.0000; AUC-ROC: 1.0000). This indicates an absolute separation of the risk classes by the network within this specific training configuration. The DT model emerged as the second most effective classifier for raw psychological data, recording a Precision of 0.7402 and an F1-score of 0.7298. The remaining probabilistic and linear models demonstrated moderate performance. Consistent with the trends observed in the clinical data analysis, the K-nn algorithm displayed an inverse relationship between the number of neighbors (K) and model efficacy; as K increased, both Precision and the F1-score systematically declined, confirming that K = 3 remains the optimal parameter within this configuration.
3.2.2. Classification Using Psychological Data Projections via Method
Following the same procedure applied to the clinical data, the psychological variables were projected into an optimized latent space using the
method. This transformation aims to maximize class separability and enhance the predictive performance of the classification algorithms. The average performance metrics obtained from these psychological projections, calculated over 10 independent runs, are presented in
Table 6.
As detailed in
Table 6, the ANN, dynamically configured with an expanded hidden layer of 500 neurons for this projected space, maintained its perfect predictive capability across the key metrics (Precision: 1.0000; F1-score: 1.0000). The DT model demonstrated notable improvement compared to its raw data baseline, securing the second-highest average performance (Precision: 0.8224; F1-score: 0.8149). The remaining classifiers exhibited similar, moderate performances. Regarding the K-nn model, the inverse relationship between the number of neighbors (K) and model accuracy persisted, confirming K = 3 as the optimal configuration for this projected space.
While
Table 6 illustrates the average stability of the classifiers, it is equally important to evaluate their peak predictive potential.
Table 7 presents the single best-performing iteration out of the 10 total executions for each model using the projected psychological variables.
The data in
Table 7 corroborate the findings from the average metrics. The ANN model achieved a flawless classification outcome during Experiment 1 (Precision: 1.0000; F1-score: 1.0000). The DT similarly reached its peak performance during Experiment 3, yielding a Precision of 0.8668 and an F1-score of 0.8682, buttressing its status as a highly effective secondary classifier for this dataset. Lastly, for the KNN model, as the number of neighbors (K) increases, both the Precision and the F1-score tend to decrease. Therefore, the value of K = 3 corresponds to the best performance within this configuration.
3.2.3. Performance Comparison: Raw vs. Projected Psychological Data
Figure 3 illustrates the comparative performance of the classification models when trained on raw psychological variables versus the optimized latent projections generated by the
method. Consistent with the findings from the clinical dataset, the application of the
method successfully enhanced class separability, contributing to a systematic improvement in classification metrics across K-nn, Linear, DT, and SVM.
The sole exception in this experimental phase was the ANN model. Because the dynamically configured ANN had already achieved perfect predictive accuracy (1.0000) using the raw psychological data, its performance metrics remained identical when processing the projected data, demonstrating a ceiling effect in its classification capability for this specific model.
3.3. Classification Based on Combined Clinical and Psychological Variables
3.3.1. Analysis of Raw Combined Data
In this comprehensive phase of the study, the classification models were evaluated using an integrated dataset comprising both the raw clinical and psychological variables. This experiment was designed to determine whether combining these two distinct domains, the biological risk factors and the psychological states, without any prior preprocessing or projection, enhances the overall predictive accuracy for breast cancer risk. The comparative performance metrics are presented in
Table 8.
As detailed in
Table 8, the ANN, configured with 20 neurons in its hidden layer, once again demonstrated exceptional predictive capacity on the raw data, achieving perfect scores in several key metrics (Precision: 1.0000; F1-score: 1.0000). The integration of both variable sets significantly improved the performance of the SVM model, which emerged as the second most robust classifier for this combined dataset (Precision: 0.8929; F1-score: 0.8693). Consistent with previous experimental phases, the K-nn model’s performance degraded progressively as the number of neighbors (K) increased, reaffirming that K = 3 remains the optimal parameter setting for this algorithm.
Intriguingly, the Gaussian classifier failed to converge valid predictions in this combined raw space (yielding 0.0000 across all metrics). This was further analyzed and is primarily attributed to the high dimensionality of the feature space relative to the number of samples. In this setting, the covariance matrix becomes ill-conditioned or singular, preventing its reliable inversion and leading to degenerate probability estimates, which explains the zero values observed across all evaluation metrics. This issue was verified by inspecting the rank and condition number of the covariance matrix. These findings further justify the use of dimensionality reduction methods, such as , which project the data into a lower-dimensional space where covariance estimation is stable, and class separability is significantly improved.
3.3.2. Classification Using Combined Data Projections via the Method
To address the complexities of the combined dataset, the integrated clinical and psychological variables were projected into an optimized latent space using the
method.
Table 8 presents the average performance metrics of the classification models, calculated over 10 independent runs, when trained on these combined, optimized projections.
As shown in
Table 9, the models demonstrating the highest predictive capability on the projected combined variables were the ANN, utilizing an expanded 500-neuron architecture, which achieved near-perfect average metrics (Precision: 0.9975; F1-score: 0.9976), followed by the DT model (Precision: 0.8874; F1-score: 0.8853).
Notably, the application of the method resolved the prior convergence failure of the Gaussian model on the raw combined data. By optimizing the dimensional space, the Gaussian algorithm successfully generated predictions, achieving a highly competitive Precision of 0.7956. Lastly, among the K-nn configurations, an increase in neighbors (K) again correlated with a decrease in Precision and F1-score, establishing again that K = 3 is the optimal baseline for this model.
Table 10 details the single best-performing run out of the 10 total executions for each model, highlighting the peak potential of these classifiers when utilizing the combined data projections.
The results in
Table 10 mirror the trends observed in the average metrics, reinforcing the superiority of the ANN model, which achieved flawless classification (Precision: 1.0000; F1-score: 1.0000) during Experiment 3. The DT model similarly peaked during experiment 5, yielding an outstanding Precision of 0.9148 and an F1-score of 0.9077. The remaining algorithms exhibited consistent, high-tier performance, validating the efficacy of the
transformation on combined, heterogeneous datasets. Once again, in the case of the K-nn model, as the number of neighbors (K) increases, both the Precision and the F1-score tend to decrease. Therefore, the value of K = 3 corresponds to the best performance within this configuration.
3.3.3. Performance Comparison: Raw vs. Projected Combined Data
Figure 4 illustrates the comparative performance of the classification models when trained on the raw, combined clinical and psychological dataset versus the optimized latent projections generated by the
method. Consistent with the findings from the isolated datasets, the application of the
method successfully enhanced class separability, contributing to improved classification metrics for K-nn, Linear, and DT models.
Two notable exceptions to this trend were the SVM and the ANN models, both of which achieved their peak performance utilizing the raw, unprojected data. As established in
Section 3.3.1, the dynamically configured ANN had already reached perfect predictive accuracy (1.0000) on the raw combined data, while the SVM demonstrated a remarkably high natural affinity for delineating the raw integrated feature space.
In stark contrast, the Gaussian model benefited the most profoundly from the projections. While the algorithm completely failed to converge on the raw, heterogeneous data (yielding metrics of 0.0000), the optimized dimensionality reduction provided by the method resolved these mathematical singularities. This transformation allowed the Gaussian model to successfully compute probabilities and achieve highly competitive performance, underscoring the vital role of in stabilizing sensitive algorithms when integrating complex, multi-domain datasets.
3.4. Comprehensive Analysis of Variable Integration
Figure 5 illustrates a comparative analysis of the precision achieved by the various classification models across the raw clinical, psychological, and integrated datasets. Across all unprojected data types, the ANN consistently demonstrated superior performance, achieving the highest precision and proving to be the best-performing architecture for this study’s baseline data.
Regarding the K-nn model, the analysis confirms that precision inversely correlates with the number of neighbors (K). The peak performance within this family of models was consistently observed at K = 3, particularly within the clinical and combined datasets, whereas higher K values systematically reduced the algorithm’s discriminative capacity. The Gaussian model performed adequately with isolated clinical and psychological data; however, it experienced a total convergence failure (precision dropping to zero) when attempting to process the raw combined variables, suggesting profound mathematical difficulty in handling this unprojected, heterogeneous data integration. The DT model exhibited robust and stable performance across all three data domains, standing out slightly on the clinical data with accuracy metrics superior to most baseline models, though not reaching the absolute ceiling of the ANN. Finally, the SVM and Linear maintained a generally intermediate performance profile, yielding acceptable precision without significantly outperforming the top-tier algorithms.
Figure 6 presents a parallel comparison of the models’ predictive accuracy when utilizing the optimized data projections (clinical, psychological, and combined) generated by the
method. Significantly, the data indicates that the integration of both clinical and psychological variables into a combined, projected latent space yields the highest precision values across the majority of the evaluated models. This outcome strongly suggests that the complementary information derived from pairing biological risk factors with psychological states, when properly optimized and mathematically smoothed via the
method, significantly enhances the overall discriminatory power and predictive reliability of the classifiers.
While the main objective of this work is to use clinical, psychological, and combined data to construct a breast cancer risk profile, it is also possible to identify, using the DE-LDA_FE method, isolated characteristics of each data type that may contribute to risk prediction. This is achieved through the use of a linear model, whose transformation matrix provides coefficients that reflect the relevance of the variables, ordered in descending order. The variables considered relevant are presented in
Appendix B (
Table A4,
Table A5 and
Table A6). However, as noted by Montes-Nogueira et al. (2018) and Romo-Gonzalez et al. (2018) [
11,
12], the breast cancer risk profile is more robust when clinical and psychological data are integrated together, rather than using isolated characteristics.
4. Discussion
4.1. The Need for Context-Specific Risk Models
Breast cancer remains one of the most critical public health issues globally, disproportionately affecting women in developing countries such as Mexico. Despite advancements in diagnosis and treatment, mortality rates remain high, largely due to delayed detection and the limited reach of effective screening programs. In Mexico, breast cancer has become the leading cause of cancer-related deaths among women, and its clinical presentation occurs approximately a decade earlier than in high-income countries. These circumstances highlight the urgent need for early detection tools adapted to the specific characteristics of the Mexican population.
Several models have been designed to evaluate the risk of developing breast cancer throughout a woman’s life. The most widely used—such as the Breast Cancer Risk Assessment Tool (Gail Model), the International Breast Cancer Intervention Study (Tyrer-Cuzick Model), and the Breast and Ovarian Analysis of Disease Incidence and Carrier Estimation Algorithm (BOADICEA)—rely on biological and reproductive risk data to provide estimates that guide physician decisions regarding future screening, preventive interventions, or chemoprophylaxis [
35,
36]. While these tools have proven their efficacy and are widely used in developed countries [
37], they rely primarily on biological metrics, failing to account for the complex sociocultural or psychological dimensions that are critical in other demographic groups.
Currently, Mexico lacks risk-assessment tools calibrated to the national population that consider lifestyle or psychological factors. Consequently, these standardized models may lack predictive power when applied to Mexican women, who experience different environmental stressors, cultural pressures, and emotional coping strategies than the populations for which the models were originally developed. Our baseline results support this limitation: the clinical-only models failed to distinguish accurately between benign pathology and cancer, suggesting that biological risk factors overlap significantly between these groups. It was only the addition of psychological profiling that successfully “unlocked” the models’ discriminative power.
4.2. Model Performance and the “Black Box” Problem
Our results demonstrate that integrating psychological variables—specifically emotional repression, suppression, and stress symptoms—with standard clinical history significantly enhances breast cancer risk prediction. The classification models achieved their greatest accuracy when utilizing the combined projections of clinical and psychological variables optimized through the method (Precision: 0.8874–0.9975; F1-score: 0.8853–0.9976), far outperforming the individual, not projected datasets.
Across the evaluated algorithms, the ANN clearly stood out as the most accurate model, reaching its peak when clinical and psychological variables were combined. However, while the ANN achieved the highest accuracy metrics, its complex architecture renders it a “black box.” This limits the interpretability of the results, which is a critical issue in the medical field where understanding how a given prediction is reached is essential for clinical trust and patient communication. This issue of explainability has been highlighted in the recent literature [
38], emphasizing the need for using Explainable Artificial Intelligence (XAI) in healthcare. Consequently, as our work points out, the DT model, which ranked second in overall performance, appears to be a more suitable and practical alternative for medical applications. It allows clinicians to visually track the classification process and transparently explain the specific criteria used to arrive at a risk assessment.
4.3. Psychological Variables and Physiological Pathways
The empirical evidence generated by our models supports the hypothesis that psychological factors are not merely incidental comorbidities, but rather associated risk factors in the multidimensional etiology of breast cancer. The predictive value of psychological variables aligns closely with Type C Personality Theory [
15,
16,
17]. Our data suggest that high levels of emotional suppression (measured by the CECS) and defensiveness (measured by the WAI) could be related as distinctive markers in the breast cancer group.
Physiologically, the relationship of these traits within the multicausal etiology of the disease could be explained by established psychoneuroimmunological mechanisms. That is, chronic emotional inhibition is known to dysregulate the hypothalamic–pituitary–adrenal axis, altering cortisol secretion patterns. This dysregulation could, in turn, affect immune surveillance and tumor suppression mechanisms [
11,
13,
14]. Therefore, the greater accuracy achieved by combining psychological and clinical data indicates that emotional suppression and distress symptoms are not only consequences of the disease but could also serve as early indicators of physiological vulnerability.
4.4. Clinical Implications and Future Applications
The proposed algorithmic pipeline offers a scalable, non-invasive alternative for early screening. Unlike standard mammography, which is resource-intensive and carries radiation risks, this methodology could be deployed as a web-based or mobile pre-screening module. Such a tool would be particularly valuable in low-resource settings, offering the following systemic benefits:
(a) Empowering Patients: Providing women with access to validated self-assessment tools in contexts where direct medical services may be limited or geographically inaccessible.
(b) Optimizing Triage: enabling primary care providers to prioritize referrals for women who exhibit a high-risk biopsychosocial profile, even if they are young or asymptomatic.
This module could be readily adapted for broader use across Latin America, significantly enhancing early detection frameworks in resource-constrained areas.
4.5. Limitations/Future Work
While the results are very promising, this study has limitations that should guide future research. First, the sample size (n = 110), although statistically sufficient for this exploratory validation, requires further expansion. Future studies should incorporate a larger multicenter cohort to ensure the model’s generalizability across the diverse geographic and socioeconomic regions of Mexico. They should also incorporate techniques for handling missing data and unbalanced cohorts. Although the proposed algorithm successfully separates risk groups, the cross-sectional nature of the data limits definitive causal inference. Longitudinal studies are needed to confirm whether these specific psychological traits precede cellular malignancy or evolve concurrently with it. Reference methods such as entropy-regulated two-view NMF and semi-supervised adaptive symmetric NMF could also be applied, as they offer comparative frameworks and could reveal how the observed benefits behave under more diverse age distributions, patterns of missing data, or larger-scale multicenter cohorts. In-depth sensitivity and robustness analyses should also be performed on the experiments.