Next Article in Journal
Deep Learning Model for Volume Measurement of the Remnant Pancreas After Pancreaticoduodenectomy and Distal Pancreatectomy
Previous Article in Journal
Clinical Relevance and Follow-Up of Incidental CT Imaging Findings for COVID-19 Diagnosis: A Retrospective Analysis
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Hybrid MCDM and Machine Learning Framework for Thalassemia Risk Assessment in Pregnant Women

by
Shefayatuj Johara Chowdhury
1,
Tanjim Mahmud
2,*,
Farzana Tasnim
1,
Sanjida Sharmin
1,
Saida Nawal
1,
Umme Habiba Papri
1,
Samia Afreen Dolon
1,
Md. Eftekhar Alam
3,
Mohammad Shahadat Hossain
4,5 and
Karl Andersson
5
1
Department of Computer Science and Engineering, International Islamic University Chittagong, Chittagong 4318, Bangladesh
2
Department of Computer Science and Engineering, Rangamati Science and Technology University, Rangamati 4500, Bangladesh
3
Department of Electrical and Electronic Engineering, International Islamic University Chittagong, Chittagong 4318, Bangladesh
4
Department of Computer Science and Engineering, University of Chittagong, Chittagong 4331, Bangladesh
5
Cybersecurity Laboratory, Luleå University of Technology, S-931 87 Skellefteå, Sweden
*
Author to whom correspondence should be addressed.
Diagnostics 2025, 15(22), 2833; https://doi.org/10.3390/diagnostics15222833
Submission received: 21 September 2025 / Revised: 28 October 2025 / Accepted: 5 November 2025 / Published: 8 November 2025
(This article belongs to the Section Machine Learning and Artificial Intelligence in Diagnostics)

Abstract

Background: Thalassemia has been recognized as a critical public health issue in Bangladesh, especially among pregnant women, due to its hereditary nature and the lack of early screening infrastructure. Early identification of at-risk individuals is essential to prevent the transmission of this genetic disorder to future generations and to reduce the burden on an already strained healthcare system. Methods: In this study, an innovative framework for thalassemia risk assessment has been developed by integrating Multi-Criteria Decision-Making (MCDM) methods—specifically AHP-TOPSIS—with machine learning algorithms including Random Forest, XGBoost, and CatBoost. Explainable Artificial Intelligence (XAI) techniques such as SHAP and LIME have also been incorporated to improve model transparency and trustworthiness. Real-world clinical and demographic data, consisting of 16 features and 1200 samples, have been collected through a structured survey and processed using rigorous feature selection and ranking methods. Risk stratification has been performed to classify patients into high, medium, and low categories, enabling targeted intervention. Results: Among all models, the XGBoost classifier trained on AHP–TOPSIS–prioritized features achieved a consistent accuracy of 99.28% under stratified 20-fold cross-validation, demonstrating robust diagnostic classification performance. The model predominantly captures hematologic patterns characteristic of thalassemia manifestations, functioning as an assistive diagnostic framework rather than a causal risk predictor. The explainability of predictions, ensured through comprehensive visual and statistical analyses, further enhances the model’s clinical transparency and reliability. Conclusions: The proposed MCDM–machine learning framework demonstrates strong potential for improving thalassemia risk assessment, enabling early detection and informed decision-making in maternal healthcare. The proposed framework should be regarded as a preliminary proof-of-concept system that demonstrates the feasibility of integrating Multi-Criteria Decision-Making (AHP–TOPSIS) with advanced machine learning and explainable-AI techniques for thalassemia assessment. Although the model achieved strong diagnostic performance under nested cross-validation, additional external validation and inclusion of causal predictors are required before clinical deployment.

1. Introduction

Thalassemia is not just a medical condition; it represents a lifelong challenge for countless families in Bangladesh. This inherited blood disorder impairs the body’s ability to produce hemoglobin, leading to chronic anemia, profound fatigue, and a range of severe health complications [1]. The two main types of thalassemia, α -thalassemia and β -thalassemia, with β -thalassemia being more widespread and severe, contribute to the high burden of this disorder in the country [2]. While thalassemia is common in regions like Southeast Asia, the Mediterranean, and India, Bangladesh faces distinct challenges due to limited awareness, inadequate healthcare infrastructure, and the absence of regular screening programs [3].
For pregnant women, the risks associated with thalassemia are particularly concerning [4]. Many expectant mothers are unaware that they carry the thalassemia gene, unknowingly passing it to their children. This lack of awareness can lead to devastating outcomes, including miscarriage, stillbirth, low birth weight, and severe anemia in newborns [5]. In rural areas, where healthcare access is already limited, the absence of early diagnosis and treatment compounds these challenges, leaving many families struggling with preventable complications. Without timely intervention, thalassemia significantly impacts maternal and child health, putting further strain on an already overburdened healthcare system [6].
Addressing this issue requires a transformative approach. This study proposes an innovative framework that integrates multi-criteria decision-making (MCDM) techniques [7,8], such as AHP-TOPSIS [9], with machine learning models to improve diagnostic accuracy and risk assessment. The incorporation of Explainable Artificial Intelligence (XAI) techniques [10], such as SHAP and LIME, enhances transparency, enabling healthcare professionals to better understand and trust the predictions of the model. By utilizing real-world clinical and demographic data, this research aims to develop a practical, scalable solution tailored to the specific needs of Bangladesh’s healthcare system.
By bridging the gap between technology and healthcare, this study aspires to provide a reliable and cost-effective solution for improving thalassemia screening, especially for pregnant women who are most at risk. Early diagnosis and accurate risk assessment can save lives, prevent complications, and significantly enhance maternal and child health outcomes across the nation.

1.1. Research Questions

This study aims to address the existing gap in thalassemia diagnosis and risk assessment by focusing on improving the screening process for pregnant women in Bangladesh. The following research questions will guide the investigation:
  • RQ1: What are the key factors affecting the diagnostic accuracy and risk assessment of thalassemia in pregnant women in Bangladesh?
  • RQ2: How can a combination of multi-criteria decision-making (MCDM) techniques (AHP-TOPSIS) and machine learning models enhance thalassemia screening in Bangladesh’s healthcare system?
  • RQ3: In what ways can the integration of Explainable Artificial Intelligence (XAI) methods like SHAP and LIME improve the interpretability and trustworthiness of thalassemia predictions for healthcare providers?

1.2. Motivations

The primary motivation for this research arises from the pressing need to address the challenges of thalassemia screening in Bangladesh, particularly for pregnant women who are most vulnerable. Despite the importance of early diagnosis, current practices are hindered by limited awareness and a lack of technological integration in healthcare systems. The aim of this study is to bridge these gaps by providing a data-driven, machine learning-enhanced framework for thalassemia risk assessment.

1.3. Novelty and Contributions

The present study has been designed to advance existing AHP–TOPSIS–ML frameworks through the following contributions:
  • All preprocessing, SelectKBest filtering, model-based importance (mean |SHAP|), AHP weighting, and TOPSIS ranking have been executed within the training folds of a repeated stratified or nested cross-validation, and performance has been evaluated exclusively on the held-out folds.
  • Causal or antecedent predictors (e.g., family history, socioeconomic context) have been distinguished from concurrent hematologic manifestations (e.g., MCV, RDW, Hb), and separate Etiology and Diagnostic models have been assessed to prevent label leakage and clarify their intended applications.
  • Feature-influence differences have been validated beyond SHAP and LIME visualizations by applying the Kruskal–Wallis and pairwise Mann–Whitney U tests, and statistical significance has been confirmed through a SHAP p-value heatmap.
  • A Combined Importance Index (normalized SHAP ⊕ SelectKBest) has been formulated and integrated with AHP (ensuring C R < 0.1 ) and TOPSIS so that final feature weights have been derived in an auditable and transparent manner.
The paper is organized as follows: Section 2 provides a comprehensive related works, discussing the current state of research on thalassemia screening, diagnostic techniques, and the application of machine learning and decision-making frameworks in healthcare, with a particular focus on maternal health. Section 3 outlines the methodology adopted in this study, including details on data collection, the integration of AHP-TOPSIS for ranking critical success factors, the application of machine learning models for predictive analysis, and the use of Explainable Artificial Intelligence (XAI) techniques such as SHAP and LIME to improve model interpretability. In Section 4, the experimental results are presented, including the evaluation of the proposed framework’s performance, accuracy, and the impact of XAI methods on trustworthiness and decision-making. Finally, Section 5 concludes the paper by summarizing the key findings, discussing their practical implications for improving thalassemia screening in Bangladesh, and suggesting possible directions for future research.

2. Related Works

Thalassemia remains a pressing public health concern, especially in regions with a high prevalence of inherited genetic disorders. A wide range of studies has focused on improving early diagnosis [11], risk prediction [8], and disease management strategies [12]. Central to these efforts is the integration of clinical, genetic, and demographic data to strengthen risk assessment models—particularly for vulnerable populations such as pregnant women. In parallel, significant attention has been given to the development of predictive models that address critical challenges, including data imbalance, feature selection, and the need for hybrid analytical approaches.
Recent advancements in machine learning (ML) and multi-criteria decision-making (MCDM) techniques have significantly enhanced diagnostic accuracy and decision support systems in medical informatics. Researchers have increasingly integrated ML algorithms with decision-analysis frameworks to improve the interpretability, scalability, and cost-effectiveness of healthcare solutions across a variety of clinical domains.
Machine learning has also been applied to study anemia in pregnant women. In Ethiopia, a predictive model was developed to assess the prevalence of anemia, identifying rural residency and lower education levels as significant risk factors. The use of decision trees and random forests improved prediction accuracy, highlighting the importance of targeted interventions to enhance maternal healthcare and support evidence-based management strategies for anemia prevention [13]. Table 1 presents a comparative summary of recent studies in disease prediction and screening, highlighting their domains, methodologies, datasets, performance metrics, and identified research gaps. Overall, existing research continues to face challenges related to achieving high accuracy, generalizability, and resource efficiency in medical AI applications. This study builds on previous work by proposing a comprehensive methodology tailored to the unique healthcare and socioeconomic landscape of Bangladesh. Through the integration of machine learning, AHP-TOPSIS, and Explainable AI (XAI), this research aims to bridge critical gaps in thalassemia risk prediction and management, contributing to more effective and accessible screening solutions [14,15]. The current study outlines significant advancements in thalassemia risk assessment and prenatal care, particularly in resource-constrained settings. In study  [14] proposed a supervised machine learning model to identify β -thalassemia carriers using complete blood count (CBC) data. By applying oversampling techniques such as SMOTE and ADASYN to address class imbalance, the authors achieved a promising accuracy of 96%. Feature reduction through PCA and SVD further improved performance, contributing to early-stage diagnosis.
Dejene et al. [13] applied ensemble ML techniques to predict anemia levels among pregnant women in Ethiopia. Using demographic data from over 11,000 participants, the CatBoost algorithm yielded the highest accuracy (97.6%). The study emphasized socioeconomic factors like education and residence, demonstrating the significance of integrating contextual features into predictive modeling.
In the context of mass screening for β -thalassemia traits, Jain et al. [16] evaluated the effectiveness of 42 RBC-based formulas using multi-criteria decision-making (MCDM) approaches. The newly developed SCS BTT formula, validated using TOPSIS and COPRAS, showed 100% sensitivity when the mean corpuscular volume (MCV) was below 80 fL. The authors advocate its use in low-resource settings due to affordability and accuracy.
The CWBCM method, introduced by Parishani and Rasti-Barzoki [17], enhances classifier selection by integrating confusion matrix metrics (accuracy, sensitivity, and specificity) within MCDM frameworks like AHP and Shannon Entropy. Applied to datasets on COVID-19, diabetes, and thyroid disease, CWBCM demonstrated improved decision quality compared to traditional weighting methods.
For breast cancer detection, Mustapha et al. [18] combined AHP and TOPSIS with supervised learning models using the Wisconsin Diagnostic Breast Cancer dataset. Their method incorporated not only prediction performance but also deployment feasibility, offering a more holistic evaluation framework for model selection.
Stenwig et al. [19] addressed interpretability concerns in predictive healthcare models. Using SHAP values, they analyzed model behavior across random forests, logistic regression, naive Bayes, and AdaBoost on the eICU dataset. Although the models had similar predictive capacities, the study revealed disparities in how each algorithm processed input features, underlining the importance of explainable AI in clinical practice.

3. Methods

The goal of this study is to build an intelligent and transparent decision-support framework for thalassemia risk assessment by integrating AHP, TOPSIS, machine learning, and explainable AI. AHP is used to assign weights to key risk factors based on expert judgment, while TOPSIS ranks these factors to identify high-risk individuals. The ranked dataset is then used to train various machine learning models, improving prediction accuracy and consistency. To ensure interpretability and trust in the results, explainable AI techniques like SHAP and LIME are applied, offering clear insights into how each model arrives at its predictions. This integrated approach provides a reliable, data-driven solution tailored for clinical decision-making.

3.1. Workflow Overview

The proposed framework establishes a comprehensive and intelligent pipeline for thalassemia diagnostic classification by integrating conventional machine learning techniques, multi-criteria decision-making (AHP–TOPSIS), and explainable AI (XAI). The overall process, illustrated in Figure 1, has been structured into several sequential phases as follows.
Step 1: Data Collection
  • Clinical and hematological data, including complete blood count (CBC) indices and demographic features, were collected from verified medical centers to ensure diagnostic reliability and diversity of samples.
Step 2: Data Preprocessing
  • Preprocessing ensured data quality and suitability for machine learning. Missing values were handled through appropriate imputation, numerical features were normalized, and categorical attributes were encoded into numerical form.
Step 3: Feature Selection
  • A hybrid feature-selection strategy was employed using Random Forest [20] and SelectKBest [21]. The importance scores from both techniques were min–max normalized and averaged to produce a unified ranking. The top ten features were selected for further prioritization through the Analytic Hierarchy Process (AHP).
Step 4: AHP–TOPSIS Based Prioritization
  • To integrate expert judgment, the AHP method [22] was used to derive feature weights, which were subsequently applied in the TOPSIS framework to compute ranked risk scores for each patient. This combination provided both data-driven and expert-informed weighting for model input preparation.
Step 5: Risk Stratification and Model Training
  • Based on TOPSIS scores, patients were stratified into three diagnostic categories: High Risk (≥0.66), Medium Risk ( 0.33 0.66 ), and Low Risk (<0.33). Machine learning models, including Random Forest, XGBoost, and CatBoost, were trained on the ranked dataset. To ensure methodological rigor, a stratified 20-fold cross-validation scheme was adopted, with feature selection and AHP–TOPSIS weighting performed within each training fold and evaluation restricted to the held-out fold.
Step 6: Explainable AI Integration
  • Interpretability was achieved using SHAP and LIME, which provided global and local explanations of model behavior, respectively. These visual and statistical interpretations enabled clinicians to understand the influence of hematologic parameters on diagnostic outcomes, enhancing transparency and clinical trust.

3.2. Survey Design for Data Collection

To facilitate the early identification and assessment of thalassemia risk, we developed a comprehensive, structured survey that captures a broad range of demographic, clinical, genetic, and socioeconomic information. The questionnaire was designed with guidance from specialists in hematology, genetics, and public health to ensure that all relevant diagnostic indicators are covered. The survey has consisted of sixteen carefully selected features with potential associations to thalassemia diagnosis or carrier status. These include age, Body Mass Index (BMI), hemoglobin level, and major hematologic indices such as Hematocrit (Hct), Mean Corpuscular Volume (MCV), Mean Corpuscular Hemoglobin (MCH), Mean Corpuscular Hemoglobin Concentration (MCHC), Red Cell Distribution Width (RDW), and RBC count. In addition, hereditary risk factors such as family history of thalassemia and the presence of genetic markers were included (see Table 2).
Recognizing the influence of social and environmental factors on health outcomes, the questionnaire has also collected information on socioeconomic status, education level, and residential location (urban or rural). Parity—the number of times a woman has given birth—has been included as a reproductive health indicator. Each question has been clearly structured using a mix of multiple-choice, yes/no, and percentage-based response formats. For instance, participants have been asked to estimate the prevalence of thalassemia within their families using a percentage input. The survey was pilot-tested on a small representative sample to refine question clarity and ensure usability before large-scale deployment. In total, responses from 1200 pregnant women have been collected for analysis.
All participants were informed about the study objectives and voluntarily provided consent prior to completing the questionnaire. No personally identifying information was recorded, and all data were anonymized to maintain confidentiality. Although data were gathered from several diagnostic centers to ensure clinical and population diversity, regional bias may remain since most respondents were from the Chattogram Division. This limitation has been acknowledged in the discussion section. Nevertheless, the structured survey provides a reliable and representative foundation for machine learning–based diagnostic modeling of thalassemia risk (see Table 3).

3.3. Feature Engineering

Feature engineering plays a crucial role in enhancing the efficiency and accuracy of machine learning models by eliminating irrelevant or redundant attributes. In this study, two well-established techniques—Random Forest Feature Importance and SelectKBest—were employed to identify the most significant features contributing to the prediction of thalassemia. These methods were chosen to provide both a model-based and a statistical perspective on feature relevance. While Random Forest evaluates the contribution of each feature based on its impact on predictive performance, SelectKBest ranks features using statistical tests that measure class separability. The integration of both methods ensured a robust and comprehensive feature selection process.

3.3.1. Random Forest Feature Importance Method

Random Forest, an ensemble learning algorithm composed of multiple decision trees, was utilized to estimate the importance of each feature. During the training phase, the algorithm constructs numerous trees, where each node is split based on a feature that minimizes an impurity criterion. The decrease in impurity—measured either by Gini impurity or entropy—was aggregated to determine the contribution of each feature.
The impurity measures are defined as follows:
Gini Impurity : G i n i = 1 k P k 2
Entropy : E n t r o p y = k p k log 2 p k
where P k represents the proportion of samples belonging to class k.
The importance score for a feature j was computed by summing the impurity reductions attributed to that feature across all trees in the forest:
F I j = t = 1 T I t j · Impurity t j
where T denotes the total number of trees, and I t j indicates the number of times feature j was used for splitting in tree t. To facilitate comparison, the scores were normalized as follows:
F I j = F I j k = 1 m F I k
with m representing the total number of features. Features with higher normalized importance scores ( F I j ) were deemed more influential. In this study, features such as Mean Corpuscular Volume (MCV) and hemoglobin level were found to have the highest importance scores, indicating their strong predictive power in identifying thalassemia risk.

3.3.2. SelectKBest Feature Selection Method

In parallel with the Random Forest approach, the SelectKBest method was applied to assess feature relevance using univariate statistical analysis. This method evaluates each feature individually with respect to the target variable and selects the top K features based on the highest test scores. For classification tasks, the ANOVA F-value was utilized as the scoring function.
The selection process followed these steps:
  • Dataset Definition: The dataset comprising input features X and target labels y was prepared and stratified by class.
  • Global Mean Calculation: The mean of each feature j across all samples was computed:
    x ¯ j = 1 n i = 1 n x i j
  • Class-wise Mean Calculation: The mean of feature j within each class k was determined:
    x ¯ j ( k ) = 1 n k i class k x i j
    where n k is the number of instances in class k.
  • Between-Class Variance:
    S B ( x j ) = k = 1 K n k ( x ¯ j ( k ) x ¯ j ) 2
  • Within-Class Variance:
    S W ( x j ) = k = 1 K i class k ( x i j x ¯ j ( k ) ) 2
  • F-statistic Computation: The F-value for each feature x j was calculated as follows:
    F j = S B ( x j ) S W ( x j )
    A higher F-value suggests a greater capacity of the feature to distinguish between classes.
  • Feature Ranking and Selection: Features were ranked according to their F-values, and the top K features were selected for further modeling.
  • Integration with Classifier: The selected features were subsequently used to train the classification model. This process helped reduce dimensionality and improve the model’s interpretability and generalization performance.
By employing both Random Forest and SelectKBest, the strengths of model-driven and statistical approaches were harnessed. This hybrid strategy ensured that the selected features were not only statistically relevant but also practically effective in improving model accuracy for thalassemia prediction (see Figure 2).

3.4. AHP-TOPSIS-Based Risk Evaluation Framework

To prioritize the influence of key indicators and rank patients by thalassemia risk, a combined AHP-TOPSIS framework was applied. This approach enabled a structured integration of expert knowledge with data-driven analysis, ensuring that each selected feature contributed proportionally to the overall assessment.
Step 1: Selection of Relevant Features
A total of nine indicators were selected based on domain knowledge and clinical relevance. These included blood-related biomarkers (Hemoglobin, Hematocrit, MCV, MCH, MCHC, and RBC count), lifestyle (BMI), genetic predisposition (family history of thalassemia), and environmental influence (socioeconomic status). The inclusion of these diverse parameters allowed a comprehensive evaluation of risk factors.
Step 2: Construction of the Pairwise Comparison Matrix
The Analytic Hierarchy Process (AHP) began with building a pairwise comparison matrix, where each feature was compared against others using Saaty’s scale (ranging from 1 to 9) [23] (see Table 4). The matrix structure is as follows:
A = 1 a 12 a 1 n 1 a 12 1 a 2 n 1 a 1 n 1 a 2 n 1
Step 3: Normalization of the Matrix
Each element of the matrix was normalized by dividing it by the sum of its corresponding column:
a i j = a i j i = 1 n a i j
Step 4: Deriving Feature Weights
The average of each row in the normalized matrix was calculated to derive the weight w i for each feature:
w i = 1 n j = 1 n a i j
These weights represent the relative importance of each feature in the overall decision-making process.
Step 5: Consistency Check
To validate the consistency of expert judgments, several calculations were carried out:
Weighted Sum Vector:
A W = A · W
Maximum Eigenvalue:
λ max = 1 n i = 1 n ( A W ) i w i
Consistency Index (CI):
C I = λ max n n 1
Consistency Ratio (CR):
C R = C I R I
where R I refers to the Random Index based on the size n. A CR value below 0.1 indicates that the level of consistency is acceptable [24].
Figure 3 shows the AHP for deriving feature weights.
Step 6: Data Normalization (TOPSIS)
Following the derivation of feature weights, the dataset was normalized using min-max scaling to ensure comparability:
x = x x min x max x min
Step 7: Weighted Normalized Decision Matrix
Each normalized feature value was then multiplied by its corresponding AHP-derived weight:
Weighted Feature i = x i × w i
Step 8: Determination of Positive-Ideal and Negative-Ideal Solutions
For each feature, the best (Positive-ideal) and worst (negative-ideal) values were identified across all samples:
Positive-ideal i = max ( Weighted Feature i )
Negative-ideal i = min ( Weighted Feature i )
Step 9: Distance Measurement
The Euclidean distances of each sample from the Positive-ideal and negative-ideal solutions were calculated as follows:
D Positive-ideal = i = 1 n ( Weighted Feature i Positive-Ideal i ) 2
D negative-ideal = i = 1 n ( Weighted Feature i Negative-ideal i ) 2
Step 10: Computation of Relative Closeness
The closeness of each patient to the ideal solution was computed using the following:
Closeness i = D negative-ideal D Positive-ideal + D negative-ideal
Step 11: Final Risk Ranking
Based on the closeness scores, each patient was ranked accordingly (see Figure 4). A higher closeness value implied a higher degree of risk, thereby aiding in targeted intervention.

3.5. Methodological Validation and Improvements

To enhance methodological rigor, three major improvements have been implemented in the present study. First, proper data partitioning and cross-validation were carried out through a repeated stratified 20-fold protocol. All preprocessing, feature selection, and AHP–TOPSIS weighting steps were executed exclusively within each training fold, while model evaluation was conducted only on the corresponding held-out fold, thereby eliminating the possibility of information leakage.
Second, causal predictors were explicitly separated from diagnostic manifestations to avoid circular reasoning. Variables representing antecedent or contextual factors, such as Family History of Thalassemia and Socioeconomic Status, were treated as causal predictors, whereas hematologic parameters (MCV, MCH, MCHC, Hct, RDW, and Hemoglobin Level) were categorized as disease manifestations. Separate Etiology and Diagnostic models were analyzed to examine their individual predictive contributions.
Third, the computation of feature importance and AHP–TOPSIS weights was explicitly defined. Feature importance scores obtained from Random Forest and SelectKBest were min–max normalized to a 0–1 scale and averaged to form a Combined Importance Index. This index was subsequently integrated with expert-assigned AHP weights using the TOPSIS algorithm to generate the final, hybrid feature-weight ranking. Through these refinements, the analytical framework achieved greater transparency, reproducibility, and statistical reliability.

3.6. Machine Learning Models for Thalassemia Diagnostic Classification

To predict thalassemia risk with high accuracy, multiple machine learning models have been developed and evaluated. These models utilize selected features to classify patients into distinct risk categories.

3.6.1. Random Forest

Random Forest is an ensemble learning method that constructs multiple decision trees during training and outputs the average prediction of the individual trees [25]. It enhances accuracy and reduces overfitting by leveraging feature randomness.
P ( Y ) = 1 m i = 1 m T i ( X )
where T i ( X ) is the prediction from the i th tree.

3.6.2. XGBoost

XGBoost is a powerful boosting algorithm that builds trees sequentially to correct errors from previous iterations [26]. It incorporates regularization to prevent overfitting and optimize model performance.
ŷ ( t ) = ŷ ( t 1 ) + η f t ( X )
It minimizes the loss:
L = i = 1 n l ( y i , ŷ i ) + t = 1 T Ω ( f t )

3.6.3. CatBoost

CatBoost is a gradient boosting framework designed to natively handle categorical features with minimal preprocessing [27]. It improves model stability and generalization by using ordered boosting and permutation techniques.
F m ( x ) = F m 1 ( x ) + ν · h m ( x )

3.7. Dataset Categorization

To facilitate targeted analysis and model development, the dataset of 1200 samples has been segmented according to the Pareto Principle (commonly known as the 80/20 rule)  [28], which holds that a small portion of the data typically accounts for the majority of significant outcomes. Following this principle, the dataset was categorized into three risk levels:
  • High Risk: Comprising the top 21.25% of risk-ranked instances (255 samples), this group has been identified as contributing most critically to thalassemia diagnosis.
  • Medium Risk: Representing the next 63.5% (762 samples), these cases have indicated a moderate level of diagnostic concern.
  • Low Risk: The remaining 15.25% (183 samples) have been considered to carry minimal or negligible risk.
This stratification has aligned with Pareto-based reasoning, suggesting that a relatively small subset of patients accounts for the most significant diagnostic focus. As a result, it has enhanced model interpretability and supported more efficient allocation of clinical resources.

3.8. Explainability Using SHAP

SHAP computes feature contributions using Shapley values:
ϕ i ( f ) = S N { i } | S | ! ( | N | | S | 1 ) ! | N | ! f ( S { i } ) f ( S )
f ( x ) = ϕ 0 + i = 1 N ϕ i

3.9. Explainability Using LIME

LIME provides local explanations by fitting a surrogate model:
(1)
Select instance x.
(2)
Generate perturbed samples.
(3)
Predict outputs using the complex model.
(4)
Compute weights:
π x ( z i ) = exp d ( x , z i ) 2 σ 2
(5)
Train a surrogate model g.
(6)
Minimize:
L ( f , g , π x ) = z i Z π x ( z i ) f ( z i ) g ( z i ) 2 + Ω ( g )

3.10. Model Training and Evaluation

The models were trained using 80% of the dataset after comprehensive preprocessing, which included normalization and the incorporation of AHP–TOPSIS–derived risk scores. To ensure methodological rigor and eliminate any possibility of data leakage, a stratified 20-fold cross-validation approach was implemented, wherein feature selection and AHP–TOPSIS weighting were conducted within each training fold, and performance evaluation was carried out exclusively on the corresponding held-out test fold. Because the input variables—such as Hct, MCV, MCH, MCHC, RDW, and RBC count—reflect the hematologic manifestations of thalassemia rather than its underlying genetic or causal determinants, the proposed model serves as a diagnostic classification system. Its objective is to identify characteristic blood-profile patterns associated with the disease rather than to predict its future onset. In subsequent work, the integration of genetic and biochemical features is planned to enable more comprehensive and causally grounded risk modeling.

3.11. Performance Metrics

Standard metrics were used:
Accuracy = T P + T N T P + T N + F P + F N ,
Precision = T P T P + F P ,
Recall = T P T P + F N
F1-Score = 2 · Precision · Recall Precision + Recall
MCC = T P × T N F P × F N ( T P + F P ) ( T P + F N ) ( T N + F P ) ( T N + F N )
where
    T P = True Positives;
    T N = True Negatives;
    F P = False Positives;
    F N = False Negatives.

3.12. Implementation Details and Libraries Used

The implementation of the proposed framework was carried out using Python, leveraging its robust ecosystem for data analysis, machine learning, and explainable AI. The workflow included data preprocessing, AHP-TOPSIS-based risk scoring, classification model training, and explainability analysis.

3.13. Environment

All experiments were conducted on a system running Windows 11 (64-bit), equipped with an Intel® Core™ i7 processor at 2.40 GHz and 16 GB of RAM (Intel Corporation, Santa Clara, CA, USA). The implementation and analysis were carried out using Python version 3.9.

3.14. Libraries and Tools

The implementation of the thalassemia diagnostic classification framework relied on several Python libraries to ensure efficiency, accuracy, and clarity throughout the process. NumPy and Pandas were used for handling data and performing numerical operations, particularly during preprocessing. To build and evaluate machine learning models like Random Forest, Scikit-learn played a central role, also supporting tasks such as splitting the data, scaling features, and tuning hyperparameters. For more advanced modeling, XGBoost and CatBoost were chosen due to their strong performance with tabular data. Visualization was performed using Matplotlib (version 3.9.4) and Seaborn (v0.13.2), which helped in generating correlation heatmaps and other insightful plots. SciPy contributed to the statistical analysis and computation of AHP weights. In addition, custom Python scripts were developed to implement AHP and TOPSIS methods for feature ranking and decision analysis. To make the models more interpretable, explainable AI tools like SHAP and LIME were used, providing insights into how each feature influences the predictions. Finally, GridSearchCV from Scikit-learn was utilized to fine-tune the models for optimal performance. Together, these tools formed a cohesive and transparent workflow that supports the credibility of the results.

3.15. Hyperparameter Optimization

GridSearchCV was used to tune parameters such as learning rate, tree depth, and the number of estimators for ensemble models.

4. Results and Discussion

This section presents the results obtained through a carefully designed and executed pipeline for thalassemia diagnostic classification. A structured process of data preprocessing, feature selection, and model training ensured reliability and robustness in the predictive outcomes. Key features were effectively identified, contributing to enhanced model accuracy and interpretability.
Machine learning models were assessed in two scenarios—with and without the integration of a multi-criteria decision-making framework. The incorporation of AHP-TOPSIS significantly improved performance metrics. Among all models, XGBoost delivered the best results, achieving an accuracy of 99.28%, an F1-Score of 99.28%, and an MCC of 98.55%. CatBoost and Random Forest have also shown marked improvements, with accuracies of 98.73% and 94.58%, respectively.
Furthermore, risk stratification enabled the classification of individuals into high, medium, and low-risk categories, supporting targeted clinical intervention. Explainable AI techniques, including SHAP and LIME, were employed to interpret model predictions, enhancing transparency and clinical trust. Overall, the proposed framework demonstrated high predictive performance, interpretability, and practical utility in thalassemia risk screening.

4.1. Data Preprocessing

Important libraries were imported, including LabelEncoder for categorical feature encoding, NumPy for numerical calculations, and pandas for data processing.

4.1.1. Handling Missing Values

No imputation techniques were applied, as the dataset was found to contain no missing values. This was confirmed using the df.isnull().sum() function, which verified the completeness of the dataset (see Table 5).

4.1.2. Categorical Data Encoding

Categorical variables such as Family History of Thalassemia, Socioeconomic Status, Education Level, Residence, and Carrier Status were transformed into a numerical format using suitable encoding techniques, including Label Encoding and one-hot encoding, based on the nature of each attribute.

4.2. Feature Selection

Feature selection was carried out using two complementary approaches—Random Forest and SelectKBest—to identify the most influential predictors of thalassemia risk. The use of both filter-based and model-based techniques ensured that statistically significant and clinically meaningful features were jointly captured. Across both methods, Hct, RBC count, and Hemoglobin Level consistently appeared among the top-ranked variables, highlighting their diagnostic relevance in thalassemia classification Section 4.2. A combined analysis of their importance scores provided a more robust and interpretable foundation for subsequent hybrid weighting through AHP–TOPSIS integration.

4.2.1. Random Forest Feature Importance

Random Forest feature importances were computed using mean Gini impurity reduction values. The most critical predictors were BMI (0.090), RBC count (0.092), and Hct (0.095), as shown in Table 6 and Figure 5. Features with relatively lower relevance, such as Carrier Status (0.009) and Genetic Marker Presence (0.012), contributed minimally to the classification outcome. This ranking has facilitated the prioritization of variables with substantial influence on thalassemia diagnostic assessment.

4.2.2. SelectKBest Feature Selection

The SelectKBest method was applied using mutual information scores to quantify the statistical relationship between each feature and the target class. According to Table 6, Hct (4.308) received the highest feature score, demonstrating its strong diagnostic association. Hemoglobin Level (2.164) and Education Level (2.256) also achieved high scores, whereas Parity (0.089) and Genetic Marker Presence (0.045) exhibited comparatively lower influence (see Figure 6). This ranking guided the selection of key variables for model optimization and interpretability.

4.2.3. Combination of Random Forest Classifier and SelectKBest Feature Selection

The combined feature-weight distribution for thalassemia analysis is presented in Figure 7. The top ten features have been ranked based on a unified importance score derived from the integration of the Random Forest and SelectKBest feature-selection methods. To merge both statistical and expert-driven perspectives, a combined importance index was developed. Feature-importance scores from both methods were normalized to a 0–1 scale and then averaged to produce a consolidated relevance vector. This unified importance ranking was subsequently used as the quantitative input for the AHP–TOPSIS procedure to ensure consistency with expert judgment.
As shown in Figure 7, Socioeconomic Status, Family History of Thalassemia, and MCV achieved the highest combined importance scores, emphasizing their contextual and diagnostic relevance in thalassemia risk assessment. Hemoglobin Level, RBC count, and RDW also showed notable contributions, aligning with hematological parameters typically affected in thalassemia carriers and major cases. In contrast, features such as Hct and MCHC received relatively lower combined weights, reflecting their more specific rather than global influence on risk stratification.
This integrated weighting approach ensures that both clinical interpretability and statistical reliability are considered when prioritizing features, producing a balanced synthesis of data-driven evidence and expert evaluation.

4.3. AHP-TOPSIS Models

AHP was applied to derive feature weights based on expert judgment through pairwise comparisons. These weights guided the normalization and preparation of data for TOPSIS-based ranking. Subsequently, relative closeness scores were calculated to categorize individuals into high, medium, and low-risk thalassemia groups.
Table 7 presents the pairwise comparison matrix used in the Analytic Hierarchy Process (AHP) to evaluate the relative importance of selected clinical and socioeconomic features. Each value indicates the degree to which one criterion was preferred over another in the decision-making process.
We normalized the pairwise comparison matrix using AHP to derive relative importance weights for each feature. Each cell represents the normalized value of a feature’s priority relative to another, based on expert judgments (see Table 8).
The AHP weights were calculated to determine the relative importance of each feature (see Table 9). It was observed that Hct was assigned the highest weight (0.2989), while MCV was given the lowest weight (0.0428) (see Figure 8).
Table 10 presents the normalized values (scaled between 0 and 1) of the top ten selected features based on Analytic Hierarchy Process (AHP) weights. Each value was computed using min-max normalization to prepare the dataset for the TOPSIS multi-criteria decision-making process.
Table 11 presents the computed relative closeness values for the first 10 records, as well as their corresponding risk ranks and final risk categories. A higher closeness value indicates a better position relative to the ideal solution.

4.4. Experimental Analysis and Comparison

4.4.1. Performance Comparison of Machine Learning Models Without AHP-TOPSIS

Figure 9 and Table 12 present a comparative analysis of three machine learning classifiers—Random Forest, XGBoost, and CatBoost—evaluated without the integration of AHP-TOPSIS feature weighting. Among the models, XGBoost achieved the highest overall performance, with an accuracy of 92.31%, an F1-Score of 92.21%, and a Matthews Correlation Coefficient (MCC) of 84.66%. The highest AUC of 97.91% was also recorded by XGBoost in the ROC curve, indicating superior discriminative capability. CatBoost followed with an accuracy of 88.97% and an AUC of 95.82%, while Random Forest attained an accuracy of 86.44% and an AUC of 94.59%. These results demonstrate that even without AHP-TOPSIS, ensemble models—particularly XGBoost—are capable of delivering strong predictive performance in thalassemia risk classification tasks.

4.4.2. Performance Comparison of Machine Learning Models After AHP-TOPSIS

Figure 10 and Table 13 present the performance of Random Forest, XGBoost, and CatBoost classifiers after the integration of AHP-TOPSIS-based feature weighting. Significant improvements were observed in classification performance across all models due to the inclusion of AHP-TOPSIS. The best results were achieved by XGBoost, with an accuracy of 99.28%, an F1-Score of 99.28%, and MCC of 98.55%, indicating excellent predictive capability and balanced classification. CatBoost followed closely, with a precision of 99.45% and an AUC of 99.96%, demonstrating strong discriminative power. Random Forest also improved significantly, attaining an AUC of 98.62% and an accuracy of 94.58%. These findings were validated by the ROC curves [29], where all models demonstrated near-perfect true positive rates—particularly XGBoost and CatBoost, whose curves nearly align with the top-left boundary. Under the 20-fold cross-validation protocol, XGBoost achieved a consistent mean accuracy of 99.28% with negligible variance across folds, confirming that the model’s performance was not due to overfitting or data leakage. Overall, the AHP-TOPSIS framework has been confirmed as a valuable enhancement for improving the accuracy and robustness of thalassemia diagnostic classification models.
Among them, the tuned XGBoost model achieved the highest overall accuracy of 99.28%. To further illustrate its predictive reliability, the corresponding confusion matrix is presented in Figure 11.

4.4.3. Comparative Analysis Both with and Without AHP-TOPSIS-Based Feature Weighting

Table 14 provides a comparative overview of machine learning classifiers—Random Forest, XGBoost, and CatBoost—evaluated both with and without AHP-TOPSIS-based feature weighting. Significant enhancements in classification performance were observed following the integration of AHP-TOPSIS.
XGBoost (see Figure 12) achieved the most notable improvement, with its accuracy increasing from 92.31% to 99.28%, F1-Score from 92.21% to 99.28%, and MCC from 84.66% to 98.55%. Its AUC also increased from 97.91% to nearly perfect. CatBoost (see Figure 13) showed similar gains, with accuracy improving from 88.97% to 98.73%, and precision reaching 99.45%. Random Forest (see Figure 14) benefited as well, with its accuracy rising from 86.44% to 94.58% and AUC from 94.59% to 98.62%. These enhancements were validated by ROC curves, where all models—especially XGBoost and CatBoost—demonstrated near-ideal true positive rates.
Overall, the integration of AHP-TOPSIS [30] has been confirmed as an effective strategy for boosting model robustness and predictive accuracy in thalassemia risk classification.

4.5. Comparative Analysis with Related Works

To emphasize the quantitative significance of the proposed framework, a comparative evaluation was performed against recent studies that combined multi-criteria decision-making (MCDM) methods and machine learning algorithms for hematologic disease diagnosis. Representative benchmark models from Rustam et al. [14] and Endalamaw et al. [13] were selected based on their highest reported performance. The summarized metrics are presented in Table 15.
The results clearly demonstrate that while prior MCDM–ML frameworks achieved commendable performance (97–97.6% accuracy), the proposed hybrid approach surpassed these benchmarks by reaching a mean accuracy of 99.28% under stratified 20-fold cross-validation (see Figure 15). This improvement of approximately 2% validates the effectiveness of integrating AHP–TOPSIS-derived feature weighting with the XGBoost ensemble model.
Beyond quantitative gains, the proposed system offers methodological advances:
  • A rigorously nested 20-fold validation pipeline eliminates data leakage and ensures generalizability.
  • Expert-informed weighting through AHP enhances clinical interpretability and transparency.
  • Integration of Explainable AI techniques (SHAP and LIME) provides local and global insights absent in previous studies.
Overall, this comparative analysis confirms that the proposed AHP–TOPSIS–XGBoost model achieves state-of-the-art diagnostic accuracy while maintaining interpretability, thereby strengthening its contribution to explainable and reliable clinical decision support.

4.6. XAI as SHAP and LIME

To enhance interpretability, both global (SHAP) and local (LIME) explainable AI methods were employed. These analyses were re-executed after implementing the stratified 20-fold cross-validation pipeline to ensure the consistency and reliability of interpretive outcomes. The resulting explanations remained qualitatively identical to earlier runs, confirming model stability under different data partitions.

4.6.1. High_Risk_Thalassemia

The explainability analysis showed that a few hematological features strongly influence the prediction of high-risk thalassemia. As seen in the SHAP feature importance plot (Figure 16), RDW, MCV, and BMI were the most dominant factors. A high RDW indicates greater variation in red blood cell size, which is a known sign of thalassemia major. A low MCV value also contributed to the high-risk prediction, consistent with the smaller red blood cells found in thalassemic patients. BMI showed a moderate effect, suggesting an indirect relation to blood parameters.
Hemoglobin Level, MCHC, and MCH had a moderate impact, while Socioeconomic Status and Family History of Thalassemia contributed slightly to model confidence.
The LIME explanation (Figure 17) supports these results for a specific case. Low MCV, MCH, and Hct values, together with abnormal RDW and BMI, pushed the model toward a high-risk prediction, whereas higher Socioeconomic Status and RBC count slightly opposed it. Overall, both SHAP and LIME confirm that hematological factors such as MCV, RDW, and MCHC play the most important roles in identifying high-risk thalassemia.

4.6.2. Medium_Risk_Thalassemia

The explainability analysis for the medium-risk thalassemia group showed that a few hematological parameters strongly affected the model predictions. As presented in the SHAP feature importance plot (Figure 18), Hct, RBC count, and RDW were the most influential variables. Higher Hct and RBC count values were associated with medium-risk classification, while moderate RDW levels indicated red cell variability typical of intermediate thalassemia conditions. BMI, MCHC, and MCV had additional but smaller effects, reflecting their secondary role in defining borderline hematological patterns. The lowest impact was observed for Socioeconomic Status and Family History of Thalassemia.
The LIME explanation (Figure 19) confirmed these relationships for a representative case. Low MCH and high Hct values supported the model’s medium-risk prediction, whereas slightly increased MCHC and BMI reduced the probability of this outcome. Together, the SHAP and LIME interpretations indicated that Hct, RBC count, and RDW play the most important roles in distinguishing medium-risk patients from both normal and high-risk groups.

4.6.3. Low_Risk_Thalassemia

The explainability analysis for the low-risk thalassemia group showed that several hematological parameters influenced the model’s classification. According to the SHAP feature importance plot (Figure 20), MCV, BMI, and RBC count were the strongest predictors of low-risk outcomes. Higher MCV and RBC count values were associated with normal or mild thalassemia cases, while an optimal BMI further supported this classification. Hemoglobin Level, MCHC, and MCH showed moderate effects, indicating their secondary role in identifying low-risk profiles. RDW, Socioeconomic Status, and Family History of Thalassemia contributed less to the prediction.
The LIME local explanation (Figure 21) confirmed these results for an individual case. Higher MCV and RBC count, combined with normal RDW and BMI values, increased the probability of a low-risk outcome, while lower MCHC and higher BMI slightly reduced it. Overall, both SHAP and LIME interpretations indicated that MCV, BMI, and RBC count play key roles in distinguishing low-risk thalassemia from moderate and severe categories.

4.7. SHAP Statistical Validation and Pairwise Significance Visualization

To further substantiate the statistical authentication of the explainability results, the outcomes of the pairwise Mann–Whitney U tests were visualized as a heatmap (Figure 22). This visualization provides a comparative overview of the p-values obtained from all feature pairs, thereby highlighting statistically significant relationships among the SHAP value distributions.
Darker regions in the heatmap correspond to lower p-values ( p < 0.05 ), indicating that the corresponding pairs of features differ significantly in their SHAP contributions. The resulting pattern confirms that the influence of critical hematological predictors—particularly MCV, BMI, and RBC count—is statistically distinct from other variables, supporting their authentic contribution to diagnostic decision-making. This statistical visualization complements the Kruskal–Wallis test results and reinforces the robustness and reliability of the proposed model’s interpretability analysis.

5. Conclusions

This study presents a comprehensive and interpretable framework for the diagnostic classification of thalassemia among pregnant women in Bangladesh by integrating Multi-Criteria Decision-Making (AHP–TOPSIS), advanced machine learning algorithms, and explainable AI techniques. The framework identifies characteristic hematologic patterns associated with thalassemia manifestations rather than predicting future disease onset. It has been designed to prioritize clinically relevant features, allowing effective stratification of patients into high-, medium-, and low-risk groups. Using a real-world dataset collected through a structured survey, the framework addresses key challenges such as feature redundancy, data imbalance, and the limited interpretability of traditional diagnostic models. Among the evaluated models, XGBoost trained on AHP–TOPSIS-ranked features achieved the best diagnostic performance, with a consistent mean accuracy of 99.28% under stratified 20-fold cross-validation. The LIME and SHAP analyses were re-run under this protocol, confirming stable interpretability results. The integration of SHAP and LIME further enhanced interpretability by providing visual and statistical explanations that align with clinical reasoning. Such transparency is essential in medical applications, as it fosters trust and supports evidence-based decision-making by healthcare professionals. Given that several top-ranked variables represent hematologic consequences of thalassemia rather than causal predictors, the proposed framework should be regarded as a proof-of-concept diagnostic model rather than a deployable screening system. Its results demonstrate methodological feasibility and clinical relevance but require validation on larger, independent datasets before real-world implementation. Future extensions will incorporate genetic and biochemical variables to transition from diagnostic assessment toward causal diagnostic classification. Looking ahead, several research directions have been identified to strengthen this work. The inclusion of genetic sequencing data, lifestyle and environmental factors, and longitudinal health records could substantially improve the robustness and accuracy of the model. The development of user-friendly mobile or web-based applications is envisioned to facilitate real-time diagnostic support, particularly in low-resource and rural settings. To ensure privacy and scalability, future research will explore federated learning and edge computing for secure, distributed health data processing. Collectively, these advancements aim to evolve the proposed proof-of-concept framework into a reliable, intelligent, and inclusive diagnostic system capable of supporting large-scale maternal and child health initiatives.

Author Contributions

Conceptualization, S.J.C., T.M., F.T., S.S., S.N., U.H.P. and S.A.D.; methodology, S.J.C., T.M., F.T., S.S., S.N., U.H.P. and S.A.D.; software, S.J.C., T.M., F.T., S.S., S.N., U.H.P. and S.A.D.; validation, S.J.C., T.M., F.T., S.S., S.N., U.H.P., S.A.D. and M.E.A.; formal analysis, S.J.C., T.M. and F.T.; data curation, S.J.C., T.M. and F.T.; writing—original draft preparation, S.J.C., T.M., F.T., S.S., S.N., U.H.P. and S.A.D.; writing—review and editing, S.J.C., T.M., F.T., S.S., S.N., U.H.P., S.A.D., M.E.A., M.S.H. and K.A.; visualization, S.J.C., T.M. and F.T.; supervision, T.M. and F.T. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

All necessary informed consent was obtained from all subjects involved in the study.

Data Availability Statement

The data underlying the findings of this study are available from the corresponding authors upon reasonable request and with proper justification.

Acknowledgments

We would like to express our sincere gratitude to Iffat Ara Tondra, Medical Officer (Maternal Health & Family Planning), Satkania Upazila Family Planning Office, Chattogram, for her generous support and cooperation in facilitating the collection of the dataset used in this study.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Shafique, F.; Ali, S.; Almansouri, T.; Van Eeden, F.; Shafi, N.; Khalid, M.; Khawaja, S.; Andleeb, S.; ul Hassan, M. Thalassemia, a human blood disorder. Braz. J. Biol. 2021, 83, e246062. [Google Scholar] [CrossRef] [Scilit]
  2. Singh, P.; Shaikh, S.; Parmar, S.; Gupta, R. Current Status of β-Thalassemic Burden in India. Hemoglobin 2023, 47, 181–190. [Google Scholar] [CrossRef] [Scilit]
  3. Islam, M.; Kamruzzaman, M.; Sarker, M.; Riaaz, R.; Ilhan, N. The Parental Perspective of Thalassemia in Bangladesh: Challenges for Prevention and Management of Thalassemia. Sch. J. App. Med. Sci. 2024, 5, 519–527. [Google Scholar]
  4. Ruangvutilert, P.; Phatihattakorn, C.; Yaiyiam, C.; Panchalee, T. Pregnancy outcomes among women affected with thalassemia traits. Arch. Gynecol. Obstet. 2023, 307, 431–438. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Hossain, M.S.; Mahbub Hasan, M.; Petrou, M.; Telfer, P.; Mosabbir, A.A. The parental perspective of thalassaemia in Bangladesh: Lack of knowledge, regret, and barriers. Orphanet J. Rare Dis. 2021, 16, 315. [Google Scholar] [CrossRef] [Scilit]
  6. Ghafor, F.; Ali, T. Compulsory Pre-marital Thalassemia Screening to Mitigate the Burden of Thalassemia Major on Society and Healthcare System of Pakistan. Med. Sci. J. Adv. Res. 2024, 5, 238. [Google Scholar] [CrossRef] [Scilit]
  7. Azhar, N.A.; Radzi, N.A.M.; wan Ahmad, W.S.H.M. Multi-Criteria Decision Making: A Systematic Review. Rec. Adv. Electr. Electron. Eng. (Formerly Recent Patents Electr. Electron. Eng.) 2021, 14, 779–801. [Google Scholar] [CrossRef] [Scilit]
  8. Karim, R.; Mahmud, T.; Hossain, S.; Hossain, M.S.; Zobaier, A.; Ahmed, M.; Sharmen, N.; Hossain, M.S.; Andersson, K. A Belief Rule Based Decision Support System to Assess Multiple Disease Suspicion from Signs and Symptoms Under Uncertainty. In Proceedings of the International Conference on Intelligent Computing & Optimization, Phnom Penh, Cambodia, 26–27 October 2023; Springer: Berlin/Heidelberg, Germany, 2023; pp. 239–250. [Google Scholar]
  9. Sharma, D.; Sridhar, S.; Claudio, D. Comparison of AHP-TOPSIS and AHP-AHP methods in multi-criteria decision-making problems. Int. J. Ind. Syst. Eng. 2020, 34, 203–223. [Google Scholar] [CrossRef] [Scilit]
  10. Tiwari, R. Explainable AI (xai) and its applications in building trust and understanding in ai decision making. Int. J. Sci. Res. Eng. Manag. 2023, 7, 1–13. [Google Scholar] [CrossRef] [Scilit]
  11. Dey, P.; Mahmud, T.; Hossain, M.S.; Andersson, K. Improving Pneumonia Detection with Deep Learning Models: Insights from Chest X-Rays. In Proceedings of the International Conference on Intelligent Computing & Optimization, Phnom Penh, Cambodia, 26–27 October 2023; Springer: Berlin/Heidelberg, Germany, 2023; pp. 164–173. [Google Scholar]
  12. Chowdhury, T.; Mahmud, T.; Mallik, A.; Mujdhalifa, N.A.; Barua, K.; Sharmen, N.; Hossain, M.S.; Andersson, K. Development of an Android-Based Pneumonia Detection App: Bridging Healthcare Gaps. In Proceedings of the International Conference on Intelligent Computing & Optimization, Phnom Penh, Cambodia, 26–27 October 2023; Springer: Berlin/Heidelberg, Germany, 2023; pp. 229–238. [Google Scholar]
  13. Endalamaw, B.; Abuhay, T.M.; Shibabaw, D. Predicting the Level of Anemia among Ethiopian Pregnant Women using Homogeneous Ensemble Machine Learning Algorithm. In Proceedings of the 2nd Deep Learning Indaba-X Ethiopia Conference 2021, Adama, Ethiopia, 16–18 December 2021. [Google Scholar]
  14. Rustam, F.; Ashraf, I.; Jabbar, S.; Tutusaus, K.; Mazas, C.; Barrera, A.E.P.; de la Torre Diez, I. Prediction of β-thalassemia carriers using complete blood count features. Sci. Rep. 2022, 12, 19999. [Google Scholar] [CrossRef] [Scilit]
  15. Göl, M.; Aktürk, C.; Talan, T.; Vural, M.S.; Türkbeyler, İ.H. Predicting malnutrition-based anemia in geriatric patients using machine learning methods. J. Eval. Clin. Pract. 2025, 31, e14142. [Google Scholar] [CrossRef] [Scilit]
  16. Jain, A.K.; Sharma, P.; Saleh, S.; Dolai, T.K.; Saha, S.C.; Bagga, R.; Khadwal, A.R.; Trehan, A.; Nielsen, I.; Kaviraj, A.; et al. Multi-criteria decision making to validate performance of RBC-based formulae to screen β-thalassemia trait in heterogeneous haemoglobinopathies. BMC Med. Inform. Decis. Mak. 2024, 24, 5. [Google Scholar] [CrossRef] [Scilit]
  17. Parishani, M.; Rasti-Barzoki, M. CWBCM method to determine the importance of classification performance evaluation criteria in machine learning: Case studies of COVID-19, Diabetes, and Thyroid Disease. Omega 2024, 127, 103096. [Google Scholar] [CrossRef] [Scilit]
  18. Mustapha, M.T.; Ozsahin, D.U.; Ozsahin, I.; Uzun, B. Breast cancer screening based on supervised learning and multi-criteria decision-making. Diagnostics 2022, 12, 1326. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Stenwig, E.; Salvi, G.; Rossi, P.S.; Skjærvold, N.K. Comparative analysis of explainable machine learning prediction models for hospital mortality. BMC Med. Res. Methodol. 2022, 22, 53. [Google Scholar] [CrossRef] [Scilit]
  20. Akhiat, Y.; Manzali, Y.; Chahhou, M.; Zinedine, A. A new noisy random forest based method for feature selection. Cybern. Inf. Technol. 2021, 21, 10–28. [Google Scholar] [CrossRef] [Scilit]
  21. Tislenko, M.; Gaidel, A.; Kupriyanov, A. Comparison of feature selection algorithms for Data classification problems. In Proceedings of the 2022 VIII International Conference on Information Technology and Nanotechnology (ITNT), Samara, Russia, 23–27 May 2022; pp. 1–5. [Google Scholar]
  22. Gyani, J.; Ahmed, A.; Haq, M.A. MCDM and various prioritization methods in AHP for CSS: A comprehensive review. IEEE Access 2022, 10, 33492–33511. [Google Scholar] [CrossRef] [Scilit]
  23. Yu, D.; Kou, G.; Xu, Z.; Shi, S. Analysis of Collaboration Evolution in AHP Research: 1982–2018. Int. J. Inf. Technol. Decis. Mak. 2021, 20, 7–36. [Google Scholar] [CrossRef] [Scilit]
  24. Handayani, A.; Farikhin, F.; Surarso, B. Statisticam approaches for consistency index in analytical hierarchy process. Aksioma: J. Mat. Pendidik. Mat. 2023, 14, 462–468. [Google Scholar] [CrossRef] [Scilit]
  25. Talekar, B. A Detailed Review on Decision Tree and Random Forest. BIoscience Biotechnol. Res. Commun. 2020, 13, 245–248. [Google Scholar] [CrossRef] [Scilit]
  26. Ali, Z.A.; Abduljabbar, Z.H.; Tahir, H.A.; Sallow, A.B.; Almufti, S.M. eXtreme gradient boosting algorithm with machine learning: A review. Acad. J. Nawroz Univ. 2023, 12, 320–334. [Google Scholar] [CrossRef] [Scilit]
  27. Kulkarni, C.S. Advancing Gradient Boosting: A Comprehensive Evaluation of the CatBoost Algorithm for Predictive Modeling. J. Artif. Intell. Mach. Learn. Data Sci. 2022, 1, 54–57. [Google Scholar] [CrossRef] [Scilit]
  28. Abyad, A. The pareto principle: Applying the 80/20 rule to your business. Middle East J. Bus. 2020, 15, 6–9. [Google Scholar]
  29. Çorbacıoğlu, Ş.K.; Aksel, G. Receiver operating characteristic curve analysis in diagnostic accuracy studies: A guide to interpreting the area under the curve value. Turk. J. Emerg. Med. 2023, 23, 195–198. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Pinho, M.; Moura, A. A decision support system to solve the problem of health care priority-setting. J. Sci. Technol. Policy Manag. 2021, 13, 610–624. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Proposed Method.
Figure 1. Proposed Method.
Diagnostics 15 02833 g001
Figure 2. Feature Selection Process Using integrated ranking from Random Forest and SelectKBest.
Figure 2. Feature Selection Process Using integrated ranking from Random Forest and SelectKBest.
Diagnostics 15 02833 g002
Figure 3. Analytical Hierarchy Process (AHP) for Deriving Feature Weights.
Figure 3. Analytical Hierarchy Process (AHP) for Deriving Feature Weights.
Diagnostics 15 02833 g003
Figure 4. TOPSIS Method for Risk-Based Feature Ranking.
Figure 4. TOPSIS Method for Risk-Based Feature Ranking.
Diagnostics 15 02833 g004
Figure 5. Feature importances obtained from the Random Forest Classifier.
Figure 5. Feature importances obtained from the Random Forest Classifier.
Diagnostics 15 02833 g005
Figure 6. Feature scores obtained from SelectKBest feature selection.
Figure 6. Feature scores obtained from SelectKBest feature selection.
Diagnostics 15 02833 g006
Figure 7. Top ten features ranked according to combined importance derived from Random Forest and SelectKBest feature-selection integration.
Figure 7. Top ten features ranked according to combined importance derived from Random Forest and SelectKBest feature-selection integration.
Diagnostics 15 02833 g007
Figure 8. AHP Feature Weight Distribution.
Figure 8. AHP Feature Weight Distribution.
Diagnostics 15 02833 g008
Figure 9. Comparing algorithms’ accuracy without AHP TOPSIS. The dashed line in an ROC curve indicates the performance of a random guess (baseline).
Figure 9. Comparing algorithms’ accuracy without AHP TOPSIS. The dashed line in an ROC curve indicates the performance of a random guess (baseline).
Diagnostics 15 02833 g009
Figure 10. ROC curve of the proposed hybrid model. The dashed line indicates the performance of a baseline.
Figure 10. ROC curve of the proposed hybrid model. The dashed line indicates the performance of a baseline.
Diagnostics 15 02833 g010
Figure 11. Confusion matrix of the tuned XGBoost model.
Figure 11. Confusion matrix of the tuned XGBoost model.
Diagnostics 15 02833 g011
Figure 12. Performance Improvement of XGBoost Model with AHP-TOPSIS Integration.
Figure 12. Performance Improvement of XGBoost Model with AHP-TOPSIS Integration.
Diagnostics 15 02833 g012
Figure 13. Performance Improvement of CatBoost Model with AHP-TOPSIS Integration.
Figure 13. Performance Improvement of CatBoost Model with AHP-TOPSIS Integration.
Diagnostics 15 02833 g013
Figure 14. Performance Improvement of Random Forest Model with AHP-TOPSIS Integration.
Figure 14. Performance Improvement of Random Forest Model with AHP-TOPSIS Integration.
Diagnostics 15 02833 g014
Figure 15. Comparative Accuracy of Existing Studies and the Proposed Method in Disease Prediction [13,14].
Figure 15. Comparative Accuracy of Existing Studies and the Proposed Method in Disease Prediction [13,14].
Diagnostics 15 02833 g015
Figure 16. SHAP feature importance for the high-risk thalassemia dataset.
Figure 16. SHAP feature importance for the high-risk thalassemia dataset.
Diagnostics 15 02833 g016
Figure 17. LIME local explanation for a high-risk thalassemia case.
Figure 17. LIME local explanation for a high-risk thalassemia case.
Diagnostics 15 02833 g017
Figure 18. SHAP feature importance for the medium-risk thalassemia dataset.
Figure 18. SHAP feature importance for the medium-risk thalassemia dataset.
Diagnostics 15 02833 g018
Figure 19. LIME local explanation for a medium-risk thalassemia case.
Figure 19. LIME local explanation for a medium-risk thalassemia case.
Diagnostics 15 02833 g019
Figure 20. SHAP feature importance for the low-risk thalassemia dataset. MCV, BMI, and RBC count have been the most influential predictors.
Figure 20. SHAP feature importance for the low-risk thalassemia dataset. MCV, BMI, and RBC count have been the most influential predictors.
Diagnostics 15 02833 g020
Figure 21. LIME local explanation for a low-risk thalassemia case. Green bars support and red bars oppose the model’s prediction.
Figure 21. LIME local explanation for a low-risk thalassemia case. Green bars support and red bars oppose the model’s prediction.
Diagnostics 15 02833 g021
Figure 22. Heatmap of Mann–Whitney U test p-values indicating significant SHAP feature differences ( p < 0.05 ).
Figure 22. Heatmap of Mann–Whitney U test p-values indicating significant SHAP feature differences ( p < 0.05 ).
Diagnostics 15 02833 g022
Table 1. Summary of Related Works in Disease Prediction and Screening.
Table 1. Summary of Related Works in Disease Prediction and Screening.
ReferenceDomainMethodologyDataset DescriptionMethods UsedAccuracy/
Performance
Research Gap
F. Rustam et al. [14]Medical GeneticsSupervised ML, PCA, SVD5066 samples (2015 carriers)SMOTE, ML models96% accuracyImbalanced dataset; need for generalizable models
B. Endalamaw et al. [13]Maternal HealthML for risk predictionEthiopian DHS (11,174 samples)CatBoost, RF, XGBoost97.08%Identification of socioeconomic risk factors
Jain et al. [16]HemoglobinopathiesMCDM for formula selection6388 (PGIMER), 939 (Kolkata)TOPSIS, COPRAS, SECASCS BTT: 100% sensitivity (MCV < 80)Inexpensive yet accurate tools for mass screening
Parishani & Rasti-Barzoki [17]Classifier SelectionCWBCM with Confusion Matrix6 datasets (COVID-19, diabetes, thyroid)CWBCM, AHP, EntropyCWBCM outperformed othersAHP and Entropy lack performance weighting nuance
Mustapha et al. [18]Breast Cancer ScreeningSupervised ML + MCDMWBCD dataset (569 records)AHP, TOPSIS, 10-fold CVEvaluated holisticallyFew models assess usability alongside accuracy
Stenwig et al. [19]ICU Mortality PredictionExplainable ML (SHAP)eICU DB (200K+ patients)RF, NB, LR, AdaBoostSimilar accuracy, different interpretabilityNeed for trustworthy and interpretable ML in healthcare
Table 2. Summary of Dataset Variables and Survey Questions.
Table 2. Summary of Dataset Variables and Survey Questions.
#Field NameQuestionAnswer Format
1AgeWhat is your age?Enter your age in years
2BMIWhat is your Body Mass Index (BMI)?Enter BMI value
3Hemoglobin_Level (Hb)What is your hemoglobin level?Enter in g/dL
4Family_History_ThalassemiaDo you have a family history of thalassemia?Yes/No
5Genetic_Marker_PresenceIs a genetic marker for thalassemia present?Yes/No
6Socioeconomic_StatusWhat is your socioeconomic status?Low/Middle/High
7Education_LevelWhat is your highest level of education completed?No formal
education/Primary/Secondary/Higher Secondary/Graduate
8ResidenceDo you live in an urban or rural area?Urban/Rural
9ParityHow many children have you given birth to?Enter number
10Carrier_StatusAre you a known carrier of thalassemia?Yes/No
11HctWhat is your hematocrit (Hct) level?Enter in %
12MCVWhat is your Mean Corpuscular Volume (MCV)?Enter in fL
13MCHWhat is your Mean Corpuscular Hemoglobin (MCH)?Enter in pg
14MCHCWhat is your Mean Corpuscular Hemoglobin Concentration (MCHC)?Enter in g/dL
15RDWWhat is your Red Cell Distribution Width (RDW)?Enter in %
16RBC countWhat is your Red Blood Cell (RBC) count?Enter in million cells/µL
Table 3. Patient Information and Blood Test Parameters.
Table 3. Patient Information and Blood Test Parameters.
IDAgeBMIHbFam. HistGen. MarkerSE StatusEducation LevelResidenceParityCarrierHctMCVMCHMCHCRDWRBC
13624.011.8No1MiddleHigher SecondaryRural0Non-Car.36.072.322.131.314.25.25
23922.610.4No0LowSecondaryUrban3Non-Car.36.869.923.733.312.35.34
32221.99.3Yes1MiddleSecondaryRural0Non-Car.30.862.724.631.011.64.99
41819.612.2No0MiddleSecondaryUrban1Non-Car.29.063.722.333.614.04.55
52722.59.1No0LowSecondaryRural1Non-Car.29.276.919.933.312.14.91
62729.312.3Yes1MiddleSecondaryUrban3Non-Car.31.362.019.230.314.35.22
73719.310.5No0LowPrimaryUrban3Non-Car.33.370.719.031.112.75.18
83927.410.5No0HighPrimaryRural2Carrier32.862.920.334.011.25.87
93122.011.5No0LowPrimaryUrban3Non-Car.32.467.423.031.012.54.95
103622.610.5No0LowPrimaryUrban2Carrier29.275.523.131.712.14.96
Table 4. Saaty’s 1–9 Fundamental Scale of Relative Importance.
Table 4. Saaty’s 1–9 Fundamental Scale of Relative Importance.
Intensity of ImportanceDefinition
1Equal importance
2Between equal and moderate importance
3Moderate importance
4Between moderate and strong importance
5Strong importance
6Between strong and very strong importance
7Very strong importance
8Between very strong and extreme importance
9Extreme importance
Table 5. Missing Value Summary for All Features.
Table 5. Missing Value Summary for All Features.
FeatureMissingFeatureMissing
Patient_ID0Hct0
Age0MCV0
BMI0MCH0
Hemoglobin_Level0MCHC0
Family_History_Thalassemia0RDW0
Genetic_Marker_Presence0RBC count0
Socioeconomic_Status0Diagnosis0
Education_Level0Residence0
Parity0Carrier_Status0
Table 6. Comparison of feature importance (Random Forest) and feature scores (SelectKBest).
Table 6. Comparison of feature importance (Random Forest) and feature scores (SelectKBest).
FeatureRandom Forest ImportanceSelectKBest Score
Hct0.1034.309
RBC count0.1030.222
BMI0.1010.420
MCV0.0940.290
MCH0.0920.630
Hemoglobin Level0.0910.575
MCHC0.0891.429
RDW0.0881.336
Age0.0760.201
Education Level0.0382.256
Parity0.0370.090
Socioeconomic Status0.0241.064
Residence0.0170.129
Family History of Thalassemia0.0171.555
Genetic Marker Presence0.0160.045
Carrier Status0.0111.688
Table 7. Pairwise Comparison Matrix for AHP.
Table 7. Pairwise Comparison Matrix for AHP.
HctMCHCBMIMCHRDWRBC CountHemoglobinFamily_HistSocioecoMCV
Hct1354357645
MCHC0.333121345323
BMI0.20.511234212
MCH0.25111234222
RDW0.3330.3330.50.5112111
RBC count0.20.250.3330.333113222
Hemoglobin0.1430.20.250.250.50.3331433
Family_Hist0.1670.3330.50.510.50.25122
Socioeco0.250.510.510.50.3330.511
MCV0.20.3330.50.510.50.3330.511
Table 8. Normalized AHP Pairwise Comparison Matrix.
Table 8. Normalized AHP Pairwise Comparison Matrix.
HctMCHCBMIMCHRDWRBC CountHemoglobin
_Level
Family_
History_
Thalassemia
Socio-
Economic _Status
MCV
Hct0.32510.40270.41380.41740.19350.26550.26010.27270.21050.2273
MCHC0.10840.13420.16550.10430.19350.21240.18580.13640.10530.1364
BMI0.06500.06710.08280.10430.12900.15930.14860.09090.05260.0909
MCH0.08130.13420.08280.10430.12900.15930.14860.09090.10530.0909
RDW0.10840.04470.04140.05220.06450.05310.07430.04550.05260.0455
RBC count0.06500.03360.02760.03480.06450.05310.11150.09090.10530.0909
Hemoglobin Level0.04640.02680.02070.02610.03230.01770.03720.18180.15790.1364
Family History Thalassemia0.05420.04470.04140.05220.06450.02650.00930.04550.10530.0909
Socioeconomic Status0.08130.06710.08280.05220.06450.02650.01240.02270.05260.0455
MCV0.06500.04470.04140.05220.06450.02650.01240.02270.05260.0455
Table 9. AHP-Derived Feature Weights.
Table 9. AHP-Derived Feature Weights.
FeatureAHP Weight
Hct0.2989
MCHC0.1482
MCH0.1127
BMI0.0991
Hemoglobin_Level0.0683
RBC count0.0677
RDW0.0582
Family_History_Thalassemia0.0534
Socioeconomic_Status0.0508
MCV0.0428
Table 10. Min-Max Normalized Values of Selected Features (First 10 Entries).
Table 10. Min-Max Normalized Values of Selected Features (First 10 Entries).
EntryHctMCHCMCHBMIHemoglobinRBC CountRDWFamily Hist.Socioeco. StatusMCV
10.88890.21670.58570.47830.84440.50000.91430.01.00.6150
20.97780.55000.81430.35650.53330.56000.37140.00.50.4950
30.31110.16670.94290.29570.28890.32670.17141.01.00.1350
40.11110.60000.61430.09570.93330.03330.85710.01.00.1850
50.13330.55000.27140.34780.24440.27330.31430.00.50.8450
60.26670.43330.30000.41300.24440.33330.08571.00.50.5350
70.17780.58330.27140.46960.20000.28670.62860.00.00.8650
80.93330.65000.78570.38690.51110.40670.68570.01.00.4250
90.55560.40000.50000.38260.53330.36670.74291.00.00.5050
100.22220.51670.61430.45650.64440.22670.74290.00.50.8850
Table 11. Relative Closeness and Final Risk Ranking (First 10 Entries).
Table 11. Relative Closeness and Final Risk Ranking (First 10 Entries).
IDRelative ClosenessRisk RankRisk Category
10.7064141High Risk
20.743373High Risk
30.3414992Medium Risk
40.24141162Low Risk
50.29061094Low Risk
60.24711157Low Risk
70.3601961Medium Risk
80.6082410Medium Risk
90.4325835Medium Risk
100.27441115Low Risk
Table 12. Performance Comparison of Machine Learning Models without AHP-TOPSIS.
Table 12. Performance Comparison of Machine Learning Models without AHP-TOPSIS.
ModelAccuracy (%)Precision (%)Recall (%)F1-Score (%)MCC (%)
Random Forest86.4485.7987.3486.5672.89
XGBoost92.3193.4990.9692.2184.66
CatBoost88.9788.6989.3389.0177.94
Table 13. Performance Comparison of Machine Learning Models After AHP-TOPSIS (in %).
Table 13. Performance Comparison of Machine Learning Models After AHP-TOPSIS (in %).
ModelAccuracyPrecisionRecallF1-ScoreMCC
Random Forest94.58%93.78%95.48%94.62%89.16%
XGBoost99.28%99.10%99.46%99.28%98.55%
CatBoost98.73%99.45%98.01%98.72%97.48%
Table 14. Side-by-Side Performance Comparison of Machine Learning Models With and Without AHP-TOPSIS (in %).
Table 14. Side-by-Side Performance Comparison of Machine Learning Models With and Without AHP-TOPSIS (in %).
ModelAHP-TOPSISAccuracyPrecisionRecallF1-ScoreMCC
Random
Forest
Without86.4485.7987.3486.5672.89
With94.5893.7895.4894.6289.16
XGBoostWithout92.3193.4990.9692.2184.66
With99.2899.1099.4699.2898.55
CatBoostWithout88.9788.6989.3389.0177.94
With98.7399.4598.0198.7297.48
Table 15. Quantitative comparison of the proposed AHP–TOPSIS–XGBoost framework with related MCDM–ML studies.
Table 15. Quantitative comparison of the proposed AHP–TOPSIS–XGBoost framework with related MCDM–ML studies.
StudyModel/
Framework
Accuracy (%)Precision (%)Recall (%)F1-Score (%)ROC/AUC (%)
Rustam et al. [14]CatBoost97.0897.0997.0597.0699.9
Endalamaw et al. [13]CatBoost + Ensemble (Anemia)97.6097.6097.4097.5099.9
Proposed StudyAHP + TOPSIS + XGBoost99.2899.1099.4699.2899.96
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Johara Chowdhury, S.; Mahmud, T.; Tasnim, F.; Sharmin, S.; Nawal, S.; Papri, U.H.; Dolon, S.A.; Alam, M.E.; Hossain, M.S.; Andersson, K. A Hybrid MCDM and Machine Learning Framework for Thalassemia Risk Assessment in Pregnant Women. Diagnostics 2025, 15, 2833. https://doi.org/10.3390/diagnostics15222833

AMA Style

Johara Chowdhury S, Mahmud T, Tasnim F, Sharmin S, Nawal S, Papri UH, Dolon SA, Alam ME, Hossain MS, Andersson K. A Hybrid MCDM and Machine Learning Framework for Thalassemia Risk Assessment in Pregnant Women. Diagnostics. 2025; 15(22):2833. https://doi.org/10.3390/diagnostics15222833

Chicago/Turabian Style

Johara Chowdhury, Shefayatuj, Tanjim Mahmud, Farzana Tasnim, Sanjida Sharmin, Saida Nawal, Umme Habiba Papri, Samia Afreen Dolon, Md. Eftekhar Alam, Mohammad Shahadat Hossain, and Karl Andersson. 2025. "A Hybrid MCDM and Machine Learning Framework for Thalassemia Risk Assessment in Pregnant Women" Diagnostics 15, no. 22: 2833. https://doi.org/10.3390/diagnostics15222833

APA Style

Johara Chowdhury, S., Mahmud, T., Tasnim, F., Sharmin, S., Nawal, S., Papri, U. H., Dolon, S. A., Alam, M. E., Hossain, M. S., & Andersson, K. (2025). A Hybrid MCDM and Machine Learning Framework for Thalassemia Risk Assessment in Pregnant Women. Diagnostics, 15(22), 2833. https://doi.org/10.3390/diagnostics15222833

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop