A New Era in Diagnosis: From Biomarkers to Artificial Intelligence, 2nd Edition

A Special Issue of Diagnostics (ISSN 2075-4418) belonging to the section "Machine Learning and Artificial Intelligence in Diagnostics".

Deadline for manuscript submissions: 31 December 2026 | Viewed by 2121

Editors


E-Mail Website
Guest Editor
Department of Medical Informatics and Biostatistics, Iuliu Hațieganu University of Medicine and Pharmacy Cluj-Napoca, Louis Pasteur Str., No. 6, 400349 Cluj-Napoca, Romania
Interests: medical research methodology; biostatistics; bioinformatics
Special Issues, Collections and Topics in MDPI journals

E-Mail Website
Guest Editor
Department of Medical Informatics and Biostatistics, Iuliu Hațieganu University of Medicine and Pharmacy Cluj-Napoca, Louis Pasteur Str., No. 6, 400349 Cluj-Napoca, Romania
Interests: medical research methodology; biostatistics
Special Issues, Collections and Topics in MDPI journals

Special Issue Information

Dear Colleagues,

This Special Issue reflects the experience of the editorial team over the past two years, which successfully concluded its first edition with no fewer than 19 published articles. We are now returning with a second edition because the field has a huge potential for development.

In recent years, it has become more important to look into new ways to diagnose diseases using biomarkers and artificial intelligence (AI). In fact, technologies like AI and machine learning might be able to change how cancer is diagnosed and treated, as well as find signs that can predict the future.

The Special Issue covers building multimodal machine learning, a branch of machine learning that focuses on creating and training models that can use different types of data, like genomic, proteomic, and imaging data, to make their performance more predictable. One of its benefits is that it can combine different types of info. For instance, imaging data can be turned into sound data so that some of the older systems could better tell the difference between cancerous and benign tumours. Putting together different types of clinical data into a single AI model is a big step toward making more complete images of clinical data.

AI helps combine different types of data, and there is also a big shift toward using digital data with numbers in clinical research, making use of the huge amounts of biomarker data created around the world. It has been found that AI, especially machine learning and high-performance computers, is the only way to combine very large datasets from genomics, proteomics, and other "omics" technologies. This change makes it possible to create new treatments and models that can predict how a drug will work. Instead of using just one biomarker, scientists can now use combinations of biomarkers to make better decisions about diagnosis and treatment.

In addition, most biomarker discovery for diseases has relied on using methodologies in AI toward obtaining predictive biomarkers or scores to accelerate the development of diagnosis and treatment. Using supervised and non-supervised machine learning algorithms while analyzing vast datasets without any bias is evident in the identification of novel biomarker candidates. Medicines for different illnesses have an unmet need for biomarker discovery, highlighting that AI will have a high impact in predictive diagnostics and therapeutic strategies.

These breakthroughs reinforce the critical role of AI and machine learning in enhancing understanding and capacity toward the proper diagnosis and treatment of diseases. These are necessary in an era of personalised medicine, in which data-driven insights steer health solutions into more accurate and effective directions.

Prof. Dr. Tudor Drugan
Dr. Daniel Leucuta
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Diagnostics is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2600 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • biomarkers
  • diagnosis
  • personalized medicine
  • therapeutic strategies
  • artificial intelligence
  • machine learning

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (4 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

22 pages, 1691 KB  
Article
Hematological Versus Clinical Predictors of Tumor Grade in Endometrial Cancer: A Comparative Machine Learning Study with SHAP-Based Explainability
by Esra Akaydın Gültürk and Şerife Özlem Genç
Diagnostics 2026, 16(18), 2987; https://doi.org/10.3390/diagnostics16182987 - 15 Sep 2026
Abstract
Background: Systemic inflammatory indices derived from routine complete blood counts have been proposed as inexpensive biomarkers in gynecological malignancies. Their ability to discriminate tumor grade in endometrial cancer, however, has rarely been evaluated against routinely available clinical variables under rigorous validation. Methods [...] Read more.
Background: Systemic inflammatory indices derived from routine complete blood counts have been proposed as inexpensive biomarkers in gynecological malignancies. Their ability to discriminate tumor grade in endometrial cancer, however, has rarely been evaluated against routinely available clinical variables under rigorous validation. Methods: We retrospectively analyzed 225 women with endometrial cancer (166 low-grade, G1–G2; 59 high-grade, G3). Nine inflammatory indices were computed. Four feature sets were compared using a prespecified elastic-net logistic regression under repeated nested cross-validation (5 outer × 5 inner folds, 20 repetitions) with 5000 stratified bootstrap confidence intervals: a collinearity-reduced inflammatory panel, the full inflammatory panel, preoperative clinical variables (age, CA-125 and preoperative albumin), and their combination. Selection among nine classifiers was retained as an exploratory analysis. Because 66% of high-grade events were non-endometrioid, every analysis was repeated in three cohorts: the full cohort, endometrioid tumors only, and the subgroup with assessed hormone receptor status. Histologic subtype was deliberately excluded because non-endometrioid carcinomas are high-grade by definition. Model behavior was interpreted with TreeSHAP and cross-checked against permutation importance. Results: Of nine inflammatory indices, only the eosinophil-to-lymphocyte ratio remained statistically associated with grade after false discovery rate correction (AUC 0.626, 95% CI 0.543–0.705; q = 0.036), with all others showing negligible effect sizes (|r| < 0.08); its discrimination was nonetheless modest and insufficient for individual-level prediction. The inflammatory panel achieved AUC 0.533 (95% CI 0.443–0.621), within the range attainable by chance in this design (permutation p = 0.254). Preoperative clinical variables reached AUC 0.687 (95% CI 0.590–0.772) in the full cohort but 0.390 (0.257–0.524) in endometrioid tumors alone and 0.627 (0.485–0.765) in the receptor-assessed subgroup. Adding inflammatory indices to clinical variables did not improve discrimination (ΔAUC = −0.025, 95% CI −0.062 to +0.012; one-sided upper bound +0.007 against a prespecified margin of +0.05). The results were unchanged under an alternative grade dichotomization and after excluding patients with inconsistent blood-count entries. Conclusions: Most inflammatory indices did not demonstrate clinically useful discrimination between G3 and G1–G2 disease in this cohort, and did not add incremental discriminative value beyond the prespecified margin over routinely available preoperative clinical variables. The apparent discrimination of the clinical variables in the full cohort did not survive restriction to endometrioid histology, indicating that it reflected histologic composition rather than grade. Decision curve analysis on recalibrated probabilities showed that no model exceeded the treat-all strategy at thresholds below the prevalence of high-grade disease, and that, above it, any advantage over treating no one was small in the full cohort and absent within endometrioid tumors. Full article
22 pages, 1683 KB  
Article
Machine Learning-Based Prediction of Masaoka–Koga Stage and WHO Histological Risk Group in Thymic Epithelial Tumors Using Biomarker Combinations
by Konstantinos Kitrou, Georgios Mandrakis, Georgios Tsirogiannis, Stamatios Theocharis, Constantinos Halkiopoulos and Yannis Stamatiou
Diagnostics 2026, 16(13), 2118; https://doi.org/10.3390/diagnostics16132118 - 7 Jul 2026
Viewed by 658
Abstract
Background: Thymic epithelial tumors (TETs) are the most common primary neoplasms of the anterior mediastinum and present a dual classification challenge, namely anatomical staging according to the Masaoka–Koga system and histological risk stratification according to the World Health Organization (WHO) classification. Both tasks [...] Read more.
Background: Thymic epithelial tumors (TETs) are the most common primary neoplasms of the anterior mediastinum and present a dual classification challenge, namely anatomical staging according to the Masaoka–Koga system and histological risk stratification according to the World Health Organization (WHO) classification. Both tasks rely on expert pathological assessment and may be affected by interobserver variability. This study applied supervised machine learning (ML) to quantitative immunohistochemical (IHC) H-score profiles to predict Masaoka–Koga stage and WHO risk group in TETs. Methods: Logistic regression (LR) and XGBoost were applied to 19 biomarkers, including cellular localization, across two parallel analyses. Masaoka–Koga stage prediction was performed in 81 patients, including 59 early-stage and 22 advanced-stage cases, using the Synthetic Minority Oversampling Technique (SMOTE) across 100 train/test splits. WHO risk group prediction was performed in 89 patients, including 45 low-risk and 44 high-risk tumors, without oversampling. A cross-endpoint analysis applied the optimal Masaoka–Koga model to the WHO endpoint. Results: LR consistently outperformed XGBoost. The optimal Masaoka–Koga model combined Eph receptor A6 (EphA6) membranous, Yes-associated protein (YAP) nuclear, and histone deacetylase 4 (HDAC4) cytoplasmic H-scores, achieving an area under the curve (AUC) of 0.756. The optimal WHO model combined transcriptional coactivator with PDZ-binding motif (TAZ) cytoplasmic, EphA6 membranous, and YAP nuclear H-scores, achieving an AUC of 0.936. The Masaoka–Koga triad predicted WHO risk group with an AUC of 0.901. No tetrad improved trivariate performance. Conclusions: IHC H-score profiling combined with supervised ML identifies biologically interpretable candidate signatures for TET classification, although prospective external validation is required before clinical application. Full article
Show Figures

Figure 1

33 pages, 4009 KB  
Article
Machine Learning Integration of Clinical and Molecular Biomarkers to Predict Vascular Complications in Type 2 Diabetes
by Gerardo García-Gil, Víctor Manuel Medina-Pérez, Joaquín Becerra-Contreras, José Alfonso Cruz-Ramos, Esteban González-Díaz, Héctor Raúl Pérez-Gómez, Kevin Javier Arellano-Arteaga, Arailym Yessenbekova, Botagoz Ussipbek, Nurzhanyat Ablaikhanova, Iryna Rusanova and Gabriela del C. López-Armas
Diagnostics 2026, 16(13), 2040; https://doi.org/10.3390/diagnostics16132040 - 30 Jun 2026
Viewed by 564
Abstract
Background/Objectives: Type 2 diabetes mellitus (T2DM) is a major global health challenge due to its high prevalence and association with chronic complications, highlighting the need for reliable predictive tools to support clinical decision-making. Methods: This study proposes a two-stage hierarchical prediction system based [...] Read more.
Background/Objectives: Type 2 diabetes mellitus (T2DM) is a major global health challenge due to its high prevalence and association with chronic complications, highlighting the need for reliable predictive tools to support clinical decision-making. Methods: This study proposes a two-stage hierarchical prediction system based on a Random Forest (RF) classifier. In Stage 1, the model performs multiclass classification into healthy (H), T2DM without complications (D), and T2DM with complications (C). In Stage 2, patients classified as C are further stratified into microvascular or macrovascular complications. The dataset included 31 biochemical, molecular, inflammatory, and oxidative stress variables from Mexican and Spanish cohorts. Feature selection was performed using Pearson correlation, and feature relevance was further assessed using RF importance measures. Model training used stratified cross-validation, with additional evaluation on a hold-out set to approximate real-world performance. Results: The optimized RF achieved an accuracy of 92% and a macro F1-score of 0.92, outperforming baseline models, with an AUC-ROC of 0.89 for complication prediction. Key predictive features included IL-18, miR-126, duration of T2DM, HbA1c, and IL-10. Conclusions: The novelty of this study lies in integrating heterogeneous biomarkers within a hierarchical predictive framework, rather than in the machine learning algorithm itself. This multimodal approach, combined with interpretable machine learning techniques, is designed to deliver clinically meaningful insights for patient stratification and personalized management in T2DM. Full article
Show Figures

Figure 1

17 pages, 2369 KB  
Article
Development and Validation of an Interpretable Machine Learning Model Based on Routine Blood Biomarkers: For Predicting Age-Related Hearing Loss
by Dan He, Yiting Liu, Jing Ke, Xu Jiang, Haiyu Ma, Ya Shi and Wei Yuan
Diagnostics 2026, 16(13), 2025; https://doi.org/10.3390/diagnostics16132025 - 29 Jun 2026
Viewed by 499
Abstract
Background/Objectives: Age-related hearing loss (ARHL) is a common sensory impairment in the elderly, and its early prediction and intervention are crucial for improving the quality of life in older adults. This study aims to develop and validate an interpretable machine learning model based [...] Read more.
Background/Objectives: Age-related hearing loss (ARHL) is a common sensory impairment in the elderly, and its early prediction and intervention are crucial for improving the quality of life in older adults. This study aims to develop and validate an interpretable machine learning model based on routine blood biomarkers to predict the risk of ARHL occurrence. Methods: A total of 542 participants were selected from the National Health and Nutrition Examination Survey (NHANES) database, including 271 ARHL patients and 271 healthy controls. The samples were randomly divided into a training set (50%) and two independent internal validation sets (25% each). Through systematic comparison of 113 machine learning algorithm combinations, the optimal predictive model (glmBoost+Stepglm[forward]) was constructed, and the SHAP method was employed for feature interpretation. To evaluate the model’s generalization ability, external validation was further performed using a cohort of 92 cases from Chongqing People’s Hospital. Additionally, an openly accessible interactive prediction web page was developed based on the R Shiny framework, supporting real-time clinical risk assessment and visual interpretation. Results: The model achieved an AUC of 0.948 in the training set, with AUCs of 0.893 and 0.945 in two internal validation sets, respectively, and an overall accuracy rate of 86.3%. In the external validation cohort (albeit with a limited sample size of 92 from a single center), the model maintained good performance with an AUC of 0.839 (95% CI: 0.750–0.918) and an accuracy of 77.2%. The model identified nine key predictive features, with the top three being glycated hemoglobin (HbA1c), mean corpuscular volume (MCV), and blood glucose according to SHAP interpretability analysis. Conclusions: This study successfully developed and validated an interpretable machine learning model based on routine blood biomarkers for community-based risk stratification of age-related hearing loss. The model demonstrated robust performance in internal and external validations, including an age-matched elderly subgroup. An interactive web tool was developed to facilitate real-time risk assessment. While the model is intended as a prescreening tool for large-scale populations rather than a diagnostic test for age-matched individuals, it provides a novel approach for early identification of individuals at higher risk of ARHL and offers insights into its systemic pathogenesis. Full article
Show Figures

Figure 1

Back to TopTop