Next Article in Journal
Women’s Cardiovascular Disease and Stroke Risk Stratification Using a Precision and Personalized Framework Embedded with an Explainable Artificial Intelligence Paradigm: A Narrative Review
Next Article in Special Issue
Optimized Machine Learning Pipeline for Lung Cancer Classification: Feature Reduction and Hyperparameter Tuning
Previous Article in Journal
Rapid Molecular Diagnostics for Bloodstream Infection in Patients with Chronic Kidney Disease
Previous Article in Special Issue
A Hybrid Multimodal Cancer Diagnostic Framework Integrating Deep Learning of Histopathology and Whispering Gallery Mode Optical Sensors
 
 
Article
Peer-Review Record

A Risk-Oriented and Explainable Hierarchical AI Framework for Chronic Kidney Disease Classification

Diagnostics 2026, 16(8), 1157; https://doi.org/10.3390/diagnostics16081157
by Sara Alhaifi 1,*, Fatmah M. A. Naemi 2 and Nahed Alowidi 1
Reviewer 1: Anonymous
Reviewer 2:
Diagnostics 2026, 16(8), 1157; https://doi.org/10.3390/diagnostics16081157
Submission received: 10 March 2026 / Revised: 5 April 2026 / Accepted: 10 April 2026 / Published: 14 April 2026

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

My comments on the study are as follows:

1. It is stated in the abstract that a hierarchical machine learning framework is proposed. What kind of model is this framework? Please clearly describe it in the abstract. Which models are included hierarchically? Please specify.

2. Could you explain the novelty of the Hierarchical CKD classification framework? What kind of contribution does it make to the literature? Please elaborate.

3. Would the classification framework you developed not be effective when applied to general datasets, rather than only region-specific datasets? In the contributions section, you mention that you collected region-specific data and addressed the lack of region-specific evidence.

4. To address the issue in item 3, you should include a public dataset in your study. You need to test your proposed model on a public dataset.

5. The figures appear to have been generated using GenAI. Figure 2 is incomplete; the number of CKD patients is not visible.

6. You state that you performed binary classification (Non-CKD and CKD). However, when the target class in Table 1 is examined, there are different stages. Would it not be more appropriate to perform multiclass classification instead of binary classification? Alternatively, after performing binary classification, how does your model determine the stages?

7. Three different data splits were used, and the 70/15/15 split yielded the best result. Instead, you should consider using 5-fold cross-validation.

8. In the experimental studies, was the data split performed on a patient basis?

9. Figure 3 is not readable. Please adjust the font size according to the journal format to improve readability.

10. What are the hyperparameters of the study? Was a programming language used? On what kind of computer were the training and testing conducted? Please clarify these points.

11. Why were macro values used in the evaluation metrics?

12. Why are there no RFECV results for the MLP?

13. In Table 2, does the reported runtime refer to training time or testing time? Please state this clearly. If it is only training time, please also include the testing time.

Author Response

Please see the attachment.

Author Response File: Author Response.pdf

Reviewer 2 Report

Comments and Suggestions for Authors

This manuscript presents an innovative risk-oriented and explainable hierarchical AI framework for chronic kidney disease (CKD) classification and risk assessment. The framework is tightly aligned with real-world clinical laboratory workflows: it first performs binary classification (CKD vs. non-CKD), conducts staging (stages 3–5) only for confirmed cases, and generates a continuous risk score with highlighted contributing features for non-CKD individuals. Using a real-world Saudi kidney clinic dataset (746 patients, 60 routine laboratory parameters), combined with RFE/RFECV/SelectKBest feature selection, XGBoost/MLP models, and SHAP explainability, the system achieves excellent performance (binary classification accuracy up to 0.97, AUC approaching 0.99; staging accuracy up to 0.85, AUC 0.96–0.97). The design explicitly addresses deployment readiness (MICE imputation, MinMax scaling, threshold optimization, normalized risk scoring) and offers strong clinical translational potential. I recommend acceptance after minor revision.

Main Concerns:

  1. The sample is modest (n=746) and single-center (Saudi Arabia). The limitations regarding generalizability should be more explicitly acknowledged in the Discussion, and the Conclusion should include a clear plan for future multi-center external validation.
  2. Although MICE was used, the manuscript does not report per-feature missing rates or sensitivity analyses. Please add a supplementary table (e.g., Table S1) listing missing percentages for all features or briefly describe them in the Preprocessing section.
  3. The 0–100 risk score is well-motivated by training-set quantiles, but no actionable clinical cut-off is provided (e.g., “score >30 warrants UACR testing”). Please add an example decision rule in Section 6.5 or Figure 7 and briefly discuss sensitivity analyses.
  4. The Related Work section focuses on 2023–2025 papers and overlooks important mechanistic and epidemiological evidence. Authors are strongly recommended inserting the following high-impact studies at the indicated locations to enhance background depth and international relevance. Introduction section should discuss CKD public-health burden and early risk factors (doi: 10.1002/med4.8) at lines 34–41. Discussion section should propose CKD Risk Assessment System when discussing clinical impact and long-term outcomes (PMID: 33209661) at lines 454–465 and discussing the clinical validity of selected features and CKD complications mechanisms (PMID: 35088922) at lines 424–433.
  5. Only four classifiers (RF, XGBoost, AdaBoost, MLP) were evaluated. A short statement comparing against recent tabular deep-learning approaches (or justifying their exclusion on grounds of interpretability/computational cost) would strengthen the Methods/Discussion.
  6. Improve resolution/legend font size in Figures 2, 4, 5, and 7 for final publication.
  7. Standardize spelling should revise (“hierarchal” → “hierarchical”).

Author Response

Please see the attachment.

Author Response File: Author Response.pdf

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

Thank you to the authors for the changes they have made.

Back to TopTop