Next Article in Journal
Association of Depressive and Negative Symptoms with Quality of Life in Schizophrenia: A Cross-Sectional Inpatient Study
Next Article in Special Issue
Multimodal Assessment of Consciousness with Brain-Computer Interfaces and Artificial Intelligence: From Acquired Brain Injury to Neurodegenerative Disease
Previous Article in Journal
Convergent Hybrid Ablation and Concomitant Left Atrial Appendage Exclusion for Stroke Prevention and Rhythm Control in Persistent Atrial Fibrillation
Previous Article in Special Issue
AI Based Clinical Decision-Making Tool for Neurologists in the Emergency Department
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Are AI Neuroimaging Models Ready for Clinical Use? A Systematic Methodological Review

1
Department of Neurological Surgery, School of Medicine and Public Health, University of Wisconsin-Madison, Madison, WI 53792, USA
2
School of Medicine, Acibadem University, Istanbul 34752, Turkey
3
BerbeeWalsh Department of Emergency Medicine, School of Medicine and Public Health, University of Wisconsin-Madison, Madison, WI 53792, USA
*
Author to whom correspondence should be addressed.
J. Clin. Med. 2026, 15(9), 3441; https://doi.org/10.3390/jcm15093441
Submission received: 6 April 2026 / Revised: 24 April 2026 / Accepted: 28 April 2026 / Published: 30 April 2026

Abstract

Background/Objectives: Artificial intelligence (AI) has rapidly expanded across medical imaging with proposed applications in diagnosis, prognostication, and surgical planning. Concerns remain regarding methodological robustness and clinical readiness for many published models. This systematic review aimed to conduct a methodological audit of AI imaging studies relevant to contemporary neurosurgical practice—including intracranial, cerebrovascular, spinal, and connectomics-based applications—published in 2025. Methods: Following PRISMA guidelines and PROSPERO registration (CRD420261284068), PubMed was searched for studies published in 2025 evaluating machine learning or deep learning applications in MRI- or CT-based imaging. Three reviewers independently extracted data on validation strategy, data leakage risk, human comparator use, calibration reporting, and CLAIM/TRIPOD-AI adherence. Risk of bias was assessed using PROBAST+AI. Results: Of 1776 screened records, 91 studies met the inclusion criteria. China led contributions (54.9%), oncology was the most common domain (37.4%), and MRI was the predominant modality (67.0%). External validation was reported in 75.8% of studies, and 66.0% used multicenter cohorts. Data leakage risk was low in 93.4%. However, only 18.7% included human comparators, calibration was reported in 30.8%, and none achieved full CLAIM/TRIPOD-AI compliance. Conclusions: AI imaging studies published in 2025 demonstrate encouraging progress in multicenter design and external validation. However, persistent gaps in human benchmarking, calibration, and reporting suggest further methodological development is needed.

1. Introduction

Artificial intelligence (AI) technologies have gained considerable traction in medical imaging research. Machine learning (ML) and deep learning (DL) models are increasingly proposed for predictive purposes, including prognostication, diagnosis, and surgical planning. AI applications have become especially widespread in oncologic imaging but also demonstrated their relevance in neurologic and neuro-radiologic diagnostics, surgery, and other fields. The performance of AI systems in various clinically valuable tasks is quite promising, including tumor characterization, hemorrhage detection, vascular risk estimation, and outcome modeling. Thus, an increased number of AI-based studies are now being published in radiology and associated disciplines.
Alongside this proliferation, there appear to be numerous problems related to methodological robustness and generalizability of many models being used. The adoption of models not validated well enough may result in serious issues, as seen by the example of clinically used AI systems, which revealed themselves to be much less efficient than initially claimed [1,2]. Previously published systematic reviews outlined major problems associated with insufficient external validation, poor generalizability, methodological biases, potential data leakage, and insufficiently transparent reporting [3,4,5]. In addition, many studies still do not fully conform to relevant guidelines for reporting (particularly Checklist for Artificial Intelligence in Medical Imaging (CLAIM) and Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis using Artificial Intelligence (TRIPOD-AI)) despite the increasing awareness of the problem. Finally, the role of Consolidated Standards of Reporting Trials for Artificial Intelligence (CONSORT-AI) as a supplement to clinical trials evaluating AI should not be underestimated [6,7,8]. Furthermore, strong discriminatory performance alone does not ensure clinical utility. Adequate calibration, seamless workflow integration, and robust inter-institutional transferability are often lacking.
To address some of these problems, several checklists such as CLAIM, TRIPOD-AI, and CONSORT-AI were developed in order to promote more transparent and robust reporting of AI studies [6,7,8]. Despite the increasing recognition of these frameworks, however, many recently published evaluations concentrated on certain aspects or specific checklists instead of conducting a more thorough analysis of validation strategies, data leakage risk, calibration, clinical benchmarking, and reporting compliance.
In light of all this, the present study sought to conduct a comprehensive methodological assessment of studies employing AI in surgical practice published in 2025. To our knowledge, no prior systematic review has provided a contemporaneous, year-specific methodological audit of validation methods, risk of data leakage, calibration, benchmarking against human evaluators, and reporting quality within this field. Instead of making comparisons in terms of discriminatory power (AUC, etc.), we conducted a systematic assessment of datasets, validation procedures, AI models employed, comparison to human experts, calibration, and compliance with reporting standards.

2. Materials and Methods

The research protocol for this systematic review was registered in the International Prospective Register of Systematic Reviews (PROSPERO ID: CRD420261284068).

2.1. Study Identification

The PubMed database was systematically searched in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines [9]. Eligible studies were those presenting primary data on AI applications in medical imaging involving human subjects.
The following Boolean search strategy was employed: (“Artificial Intelligence” OR “Machine Learning” OR “Neural Networks, Computer” OR “artificial intelligence” OR “machine learning” OR “deep learning” OR “neural network*” OR “radiomics”) AND (“Diagnostic Imaging” OR “medical imaging” OR “radiology” OR “MRI” OR “CT”) AND (“bias” OR “external validation” OR “generalizability” OR “reproducibility” OR “data leakage”) AND (“Humans”) AND (“2025/01/01” [Date–Publication]: “2025/12/31” [Date–Publication]).

2.2. Eligibility Criteria

Articles were considered eligible for inclusion if they discussed the use of AI techniques in medical imaging relevant to neurosurgical practice and fulfilled the following criteria: (1) publication in 2025; (2) original research article; (3) application of AI methods on MRI and CT imaging modalities and/or angiographic imaging (i.e., computed tomography angiography (CTA) and magnetic resonance angiography (MRA)); (4) human subjects’ studies; (5) development, validation, and/or evaluation of models; and (6) availability of the full texts. For the purpose of this review, “relevant to neurosurgical practice” was defined to include: (a) intracranial pathology and neurological disease (e.g., tumors, stroke, hemorrhage, epilepsy, hydrocephalus, aneurysms); (b) spinal pathology routinely managed by neurosurgeons (e.g., degenerative spine disease, vertebral fractures, surgical planning and screw placement); (c) extracranial cerebrovascular disease contributing to stroke or requiring surgical or endovascular intervention; and (d) connectomics and functional imaging studies with direct applicability to neurosurgical planning workflows (e.g., tractography, eloquent-area mapping, DBS targeting). Studies originating from related clinical specialties (orthopedics, otolaryngology, radiology, neuroscience) were considered eligible when their investigated imaging conditions fell within one of these categories.
Papers that did not meet the following criteria were excluded: (1) review papers, systematic reviews, and/or meta-analysis; (2) conference abstracts, editorials, letter to editors, or comments; (3) studies involving only animals, phantoms, or simulations; (4) purely technical articles in computer sciences with no medical imaging applications; (5) AI articles irrelevant to medical imaging; (6) papers not applying machine learning/deep learning algorithms; and (7) papers published before 2025.

2.3. Study Selection Process

Duplicated entries were eliminated via Rayyan. Titles and abstracts were then independently reviewed by three reviewers (U.S., N.S., and B.D.) based on their eligibility for further consideration. The full texts of articles identified as eligible by the aforementioned three independent reviewers were subsequently evaluated again for the purpose of finalizing the article selections for this systematic review. Any disagreements between the independent reviewers were settled through discussion and consensus among all reviewers. At each step of the selection procedure, reviewers reached agreement via consensus building. Reference lists of the selected papers were manually checked for any additional eligible citations. The study selection process followed the PRISMA guidelines, and the selection flowchart can be found in Figure 1.

2.4. Data Extraction

Data were independently collected by three reviewers using a standardized extraction form, with any disagreements resolved through discussion and consensus. For each included study, information was gathered across several domains: (1) study characteristics (author, publication year, journal, country, and medical field); (2) imaging and dataset features (imaging modality, dataset type, sample size, and ground truth definition); (3) AI model details (type of AI, model architecture, and task). Hybrid models were defined as those combining deep learning and traditional machine learning within a single pipeline. In multi-component pipelines, classification was based on the primary modeling approach described. Additional domains were (4) validation and transparency (validation method, use of external datasets, clarity of data splitting, and risk of data leakage); (5) methodological features (study design, inclusion of human comparators, reporting of calibration metrics, clarity of performance metrics, and adherence to CLAIM/TRIPOD-AI guidelines); and (6) reported outcomes (main performance metric and claims of clinical applicability). Studies involving multiple clinical domains were categorized as mixed.
Variables were recorded as either categorical or numerical, as appropriate (e.g., sample size as numerical; split clarity, external validation, human comparator, calibration reporting, and clinical applicability as yes/no; and data leakage risk and CLAIM/TRIPOD-AI adherence as predefined categories). The extracted data were summarized using descriptive statistics.
Data leakage risk was evaluated based on reported data handling and validation procedures, guided by CLAIM and TRIPOD-AI principles. Risk was classified as low when training, validation, and test sets were clearly separated at the patient or center level without overlap. Moderate risk existed when the data division was vague or described in an equivocal manner. High or unclear risk was characterized by clear overlap between datasets, preprocessing steps undertaken prior to data splitting, or inadequate reporting to eliminate leakage. The clarity of split descriptions was marked as “yes” when studies clearly detailed how datasets were divided (including whether splitting was done at the patient or center level) and “no” when this information was incomplete or unclear.
Adherence to CLAIM and TRIPOD-AI was evaluated at the domain level rather than through formal item-level scoring. Six predefined reporting domains derived from both frameworks were assessed: (1) data partitioning transparency; (2) model development and architecture clarity; (3) validation methodology; (4) completeness of performance reporting; (5) calibration reporting; and (6) human comparator analysis, when applicable. Three reviewers (U.S., N.S., B.D.) assessed each study independently and reconciled by consensus. Studies were classified as showing full adherence (all applicable domains adequately reported), partial adherence (one or more domains incompletely reported), or no adherence (key domains absent). Full item-level scoring (42 CLAIM items, 27+ TRIPOD-AI items) across 91 heterogeneous studies was beyond the feasible scope of this field-level audit. Direct comparisons with expert performance by radiologists or neuroradiologists on the same data and task were recorded as “yes” for human comparators; research lacking such comparisons was categorized as “no”. Calibration reporting was considered present if any standard metric (e.g., calibration curves, Brier score, Hosmer–Lemeshow test, or equivalent) was reported.

2.5. Data Analysis

The analysis of the data was conducted as a descriptive systematic review. Counts and proportions were used to summarize extracted variables. The multinational studies were those studies that were performed in two or more countries. In the case of multiple imaging modalities, the primary modality was used to classify them. The variables that were categorical (e.g., the type of dataset, the validation strategy, external validation, data leakage risk, calibration reporting, and the use of human comparators) were reported as frequencies and percentages.
Descriptive measures (means or medians, depending on the type of continuous variables) were used to summarize continuous variables (sample size, performance metrics, e.g., AUC, accuracy). Because of the high level of heterogeneity in study designs, protocols of imaging, patient groups, measures of outcomes, and reporting of outcomes, a meta-analysis was not performed.
Subgroup analyses were done where they were necessary, especially studies with or without external validation. Synthesized findings were presented in narrative form, and the quality of the methods, practices of validation, and possible biases in relation to clinical use were considered.

2.6. Risk of Bias Assessment

The risk of bias was assessed by three reviewers (U.S., N.S., and B.D.) with the help of Prediction Model Risk of Bias Assessment Tool of AI (PROBAST+AI) [10]. All studies were evaluated on the following areas: participants, predictors, outcomes, and analysis, where the risk of bias was evaluated, in general, according to the guidelines of PROBAST+AI.

3. Results

3.1. Study Selection

The initial search in the database revealed 1776 records. After removing 1 duplicate, 1775 records were screened by title and abstract. Subsequently, 737 full-text articles were assessed for the final screen. Finally, 91 studies were eligible for the inclusion criteria and were used in the qualitative synthesis. Figure 1 (PRISMA flow diagram) presents a description of the selection procedure.

3.2. General Study Characteristics

The 91 studies that were included were very widely internationally represented and represented 14 countries and 8 multinationals. China contributed the largest share (54.9%, n = 50), followed by the United States (11.0%, n = 10) and multinational studies (8.8%, n = 8) (Figure 2).
The most represented field was oncology (34 studies, 37.4%). This was followed by neurology (n = 19, 20.9%), radiology (n = 12, 13.2%), and neuroradiology (n = 5, 5.5%). It was also found that multidisciplinary studies were done, such as oncology/radiology (n = 3, 3.3%) and neurology/radiology (n = 1, 1.1%). The proportion of all other specialties was below 5% each (Figure 3).
MRI was the most common imaging modality (61 of 91 studies, 67.0%). The remaining 30 studies (33.0 percent) were founded on CT imaging, such as non-contrast CT, CT angiography, and similar methods.

3.3. Artificial Intelligence Methodological Characteristics

The most widespread method was the use of deep learning (DL), which was mentioned in 47 studies (51.6%) [11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47,48,49,50,51,52,53,54,55,56,57]. Machine learning (ML) methods were used in 34 studies (37.4%) [58,59,60,61,62,63,64,65,66,67,68,69,70,71,72,73,74,75,76,77,78,79,80,81,82,83,84,85,86,87,88,89,90,91], while hybrid DL/ML approaches were reported in 10 studies (11.0%) [92,93,94,95,96,97,98,99,100,101] (Figure 4). In general, DL was observed, but traditional ML was still used extensively. China had the most studies of any category of methodology (Figure 4).
The most common architecture was the neural network-based models, with 44 studies (48.4%). The next most common models were hybrid or ensemble models (n = 22, 24.2%), followed by regression-based models in 15 studies (16.5%). In 5 studies (5.5%), kernel/distance-based methods and gradient boosting were reported. These results demonstrate the supremacy of neural networks, with the ongoing application of traditional and hybrid approaches (Figure 5).
Prediction was the most frequent task (n = 35, 38.5%), followed by classification (n = 26, 28.6%). Detection and segmentation were reported in 10 (11.0%) and 8 (8.8%) studies, respectively. Multi-task approaches were used in 7 studies (7.7%). Less common tasks included diagnosis (n = 2, 2.2%) and synthesis, identification, and grading (each n = 1, 1.1%). Together, prediction and classification accounted for approximately 67.0% of all studies (Figure 6). The detailed characteristics of all included studies are presented in Table 1.

3.4. Validation and Dataset Characteristics

External validation was performed in 69 studies (75.8%), whereas 22 studies (24.2%) relied solely on internal validation (Figure 7). Cross-tabulation analysis indicated that most studies with external validation did not include comparisons with human experts.
External datasets were used in 66 studies (72.5%), while 25 studies (27.5%) did not use them. This pattern closely aligns with the proportion of studies reporting external validation, with minor differences likely reflecting variations in how external datasets and validation approaches were defined.
Multicenter datasets were the most commonly used, appearing in 60 studies (66.0%). Single-center studies accounted for 22 (24.2%), while 9 studies (9.8%) relied on public datasets. In general, this distribution indicates a significant move towards the use of multi-institutional data sources, and it implies that heterogeneity in the development and validation of AI models should be prioritized.

3.5. Methodological Transparency and Bias Indicators

The majority of the studies (85; 93.4%) were categorized as having a low risk of data leakage based on their reported data-handling and validation procedures. Moderate risk was identified in 2 studies (2.2%), unclear risk in 3 studies (3.3%), and high or potentially unclear risk in 1 study (1.1%). Almost all studies described their data-splitting plans in a transparent manner. Cross-tabulation showed that low leakage risk was most commonly associated with studies using external datasets (Figure 8). As the assessment was based on reported methodology rather than independent verification of practice, subtle forms of leakage not captured in written descriptions cannot be excluded.
Direct comparison with expert human performance was reported in 17 of 91 studies (18.7%); the remaining 74 studies (81.3%) did not include a human benchmark (Figure 7). However, this aggregate rate masks substantial task-dependent heterogeneity (Table 2): 60.0% in detection (6/10), 57.1% in multi-task (4/7), 14.3% in prediction (5/35), 3.8% in classification (1/26), and 0% in segmentation (0/8). The absence of human comparators in segmentation reflects the appropriate use of expert-annotated ground-truth labels, whereas the low rates in classification and prediction tasks that directly inform clinical decision-making represent a more concerning gap than the aggregate figure suggests. Even among externally validated studies, the proportion including a human comparator remained low (Figure 7).
Calibration was reported in 28 of 91 studies (30.8%), with 62 studies (68.1%) not reporting any calibration metric. Upon re-extraction, the most commonly used metric was the calibration curve (n = 25, 89.3%), followed by decision curve analysis (n = 24, 85.7%), Hosmer–Lemeshow test (n = 5, 17.9%), Brier score (n = 4, 14.3%), and expected calibration error (n = 2, 7.1%). Among studies reporting calibration, 19 (67.9%) assessed calibration on both internal and external validation cohorts, 7 (25.0%) on internal cohorts only, and 2 (7.1%) on external cohorts only. Calibration reporting was ambiguous in one study (1.1%).
The studies showed partial compliance with either CLAIM and/or TRIPOD-AI guidelines, yet none of them were fully compliant, which means that gaps in complete reporting persist (Supplementary Table S2).
The most prevalent reported measure was the area under the receiver operating characteristic curve (AUC), especially when predicting and classifying. The values of AUC were also diverse, with a number of studies showing more than one result in different validation cohorts. Other measures were accuracy, sensitivity, specificity and concordance index (C-index), and Dice similarity coefficient (segmentation). There was a lot of heterogeneity in the definition of outcomes, measures of evaluation, and practices of reporting among studies.

4. Discussion

4.1. Scope of Included Studies

The present review was deliberately designed with a broad definition of ”neurosurgical relevance” to capture the full methodological landscape of AI imaging studies informing contemporary neurosurgical practice. Beyond intracranial pathology, neurosurgery encompasses spinal surgery, cerebrovascular disease, and increasingly, connectomics-based surgical planning. Accordingly, of the 91 included studies, 86.8% address core intracranial or neurological applications (tumors, stroke, hemorrhage, epilepsy, aneurysms, hydrocephalus), while the remaining 13.2% consist of spinal imaging (8.8%), extracranial cerebrovascular imaging (2.2%), and connectomics or functional imaging applicable to surgical mapping (2.2%). This breadth reflects the multidisciplinary reality of neurosurgical practice and ensures that methodological conclusions regarding validation, calibration, and reporting adherence are drawn from the full spectrum of AI imaging work relevant to the specialty, rather than from a narrower intracranial subset.
The amount of research in AI medical imaging has been increasing significantly over the past few years, and the number of healthcare-related publications per year has risen from 2113 in 2021 to 4587 in 2023, which is more than doubled in two years [102]. Predictive modeling is commonly utilized in various fields of medicine, with oncologic imaging being one of them. In that regard, we performed a dedicated systematic review of AI studies applicable to neurosurgical practice published in 2025 to evaluate the current methodological situation. Instead of subjecting the performance of models to comparison, we intended to assess the methodological rigor and applicability of methods critically and in a real-world setting. We considered the characteristics of datasets (single- vs. multicenter), methods of validation (internal vs. external), transparency of data-splitting processes, and the risk of data leakage. We also evaluated major transparency and quality indicators, such as study design, the use of human comparators, calibration reporting, and compliance with the established frameworks, such as CLAIM and TRIPOD-AI [6,7].
The review of 91 studies suggests that there are significant advances; however, limitations that may influence clinical translation remain [3,103,104]. Deep learning methods predominated, and multicenter datasets were commonly employed, indicating increased attention to generalizability. External validation was reported in three-quarters of studies (n = 69/91), a marked increase over previous reports, in which external validation was found in only 6–10% of AI medical imaging studies [105,106]. A more recent methodological audit by Spaanderman et al. evaluating AI imaging studies against the CLAIM and FUTURE-AI frameworks through July 2024 similarly identified persistent gaps in external validation reporting; however, direct comparison is limited by differences in clinical scope [107]. This trend reflects growing awareness of the need to test models on independent data.
Regardless of these developments, there are a number of substantial drawbacks. There were few studies (n = 17/91) that had direct comparisons with human experts. Though these comparisons might not necessarily be valid, especially with some segmentation or new predictive tasks, the lack of them in research that suggests diagnostic or decision-support uses can restrict the evaluation of added clinical value [103]. Although automated processes may be used, expert monitoring is frequently still required, which highlights the significance of reporting ground-truth quality and inter-rater reliability. Also, almost two-thirds of studies (n = 62/91) did not include calibration assessment, which casts doubt on the predictability of the predicted probabilities in clinical decision-making [104].
Even though the majority of the studies were considered to be of low risk of data leakage, discrepancies in the reporting imply that certain methodology problems might still go unnoticed. Moreover, all studies reported partial compliance with CLAIM/TRIPOD-AI guidelines; however, none of them fully complied, and this remains an issue of standardized reporting. Altogether, these results suggest that despite the increase in methodological rigor, there remain major concerns regarding the transparency, validation, and clinical relevance of AI models that prevent their use in a standard clinical environment.

4.2. External Validation and Generalizability

The proportion of studies reporting external validation (75.8%) represents a clear improvement over prior literature, in which such validation was typically absent [3,103]. The shift toward external data divides to multicenter clinical cohorts or international competition datasets is a sign that internal validation is an incomplete approach to clinical generalizability [39,59,74].
Nevertheless, external validation does not always ensure strong real-life performance [12]. Numerous investigations used datasets of similar geographically located institutions or processed by similar pipelines [11,15,17,23,31,75,77,80,81,82,83,84,86,87,88,89,90,91,95,96,97,98]. Although this can increase internal consistency, it might not respond well to variation as experienced in normal clinical practice, including variation in scanners, protocols, and patients [39]. The relative lack of experimental studies on true cross-institutional transportability, including between different vendors and workflows, is relatively under-investigated in the literature, even though a few multicenter studies have demonstrated it, including that by Dai M. et al. [40]. Moreover, whereas dataset-level generalizability is gaining more importance, the same cannot be said of human benchmarking or calibration assessment. Consequently, models can be seen to be more generalizable, yet without adequate evaluation to be used in decision support.
A further methodological consideration concerns database coverage in the present review. PubMed was selected as the sole search database to align with the clinical scope of this review, as databases such as EMBASE, Web of Science, and Scopus index a larger proportion of engineering, computer science, and informatics publications that fall outside our predefined clinical focus and were a priori excluded through our eligibility criteria. This choice is consistent with prevailing practice in the field: a recent umbrella review of 158 AI imaging systematic reviews demonstrated that PubMed remains the most frequently used primary search source, appearing in 71.5% of reviews in this area [108]. Nevertheless, single-database searches carry an inherent coverage limitation. Empirical analyses of systematic review database yields have shown that approximately 16% of eligible references may be retrievable from only one database, with EMBASE frequently contributing the largest share of unique references not indexed in MEDLINE/PubMed [109]. Based on these estimates, the expected coverage gap for our PubMed-only search may approximate 5–15% of potentially eligible clinically oriented studies. While this may modestly affect field-level generalizability, the excluded literature would largely consist of engineering-focused work outside the intended scope of this review, making it unlikely that this gap systematically biased our conclusions regarding validation practices, calibration reporting, or guideline adherence.
This restriction is especially critical in the neurosurgical practice, where decisions are made in a multidisciplinary team that involves neurosurgeons, neuroradiologists, and oncologists. Models that are only validated based on performance measures and not on comparison with expert judgment are unable to completely indicate their additional clinical value. In the absence of such benchmarking, claims of clinical applicability are incomplete. In addition to performance measures, the other important but less investigated part of validation is workflow-level evaluation. Even though a recent systematic review identified a significant number of studies that showed an implementation of AI led to a decrease in the time of task completion, workflow impact evaluations are not common [110]. An interesting exception is the clinical trial conducted by Kang et al. [41], which demonstrated better reader performance under AI assistance. On the whole, external validation can be regarded as an indispensable step, but it cannot be considered as being enough. As the field moves towards explainable and clinically integrated AI systems, comprehensive evaluation across multiple dimensions—including transportability, human comparison, calibration, and workflow impact—will be required to establish genuine clinical readiness [42].

4.3. Data Leakage and Methodological Transparency

One of the most important but least realized risks to the validity of AI model performance in medical imaging is data leakage. The majority of the studies in this review were rated to be at a low risk of data leakage (n = 85/91, 93.4) according to the data handling and partitioning mechanisms reported [43,59] (Supplementary Table S2). Almost all studies presented their training, validation, and test splits explicitly, signifying a higher level of transparency on methodological transparency than the previous literature.
Nevertheless, the completeness and clarity of reporting are critical to assessing the risk of leakage in systematic reviews. Even though the vast majority of studies had separated patient- or center-level datasets, the differences in preprocessing, feature extraction procedures and cross-validation methods can make one less confident about the ability to eliminate less obvious types of leakage. As an example, the less transparent automated pipelines described by Sina et al. were categorized as high/unclear risk because the description lacked adequate detail on internal data partitioning [55]. Similarly, intricate model structures, ambiguous preprocessing procedures, and some feature extraction approaches occasionally blurred dataset boundaries, resulting in moderate or ambiguous risk classifications [42,44,45,92,99].
It has been demonstrated in prior studies that even small amounts of preprocessing that occur before splitting datasets can artificially boost model performance in an artificial manner [3,104]. Notably, the large percentage of low-risk studies should be taken with a grain of salt. Data partitioning is not reported clearly enough to rule out the chance of leakage, especially in deep learning pipelines that may include augmentation, normalization, or transfer learning across overlapping data domains. These results emphasize the importance of better reporting of data handling procedures, which are more standardized and detailed, according to the CLAIM and TRIPOD-AI guidelines, to make those reproducible and facilitate the translation of such evidence to a reliable clinical application [7,111].

4.4. Clinical Benchmarking and Calibration

Despite increasing emphasis on external validation, direct comparison between AI models and human experts remains limited. Only 18.7% of studies (17/91) included a human comparator. However, this aggregate rate masks substantial task-dependent variation (Table 2): 60.0% of detection studies and 57.1% of multi-task studies benchmarked against radiologist performance for clinically actionable tasks, whereas only 14.3% of prediction studies and 3.8% of classification studies incorporated human benchmarking. The 0% rate in segmentation is methodologically appropriate, as segmentation is typically benchmarked against expert-annotated ground-truth labels using metrics such as the Dice coefficient [18,25].
This pattern refocuses the clinical significance of our finding: the scarcity of human benchmarking is not uniform across the field but concentrated in classification and prediction—precisely the task categories where demonstrating clinical added value depends on comparison against expert care. In neurosurgical contexts involving AI-derived probabilistic outputs (e.g., glioma grading, stroke outcome, aneurysm rupture risk), the absence of benchmarking in 96.2% of classification and 85.7% of prediction studies raises serious concerns about clinical translational readiness. Kang et al. provide a notable exception, using a prospective design to directly quantify AI-assisted improvement in clinician performance [41].
Interestingly, in studies that conducted external validation, there was a minimal percentage of studies that proceeded to conduct external validation of their results [38,48,51,52,53,59,61,65]. It implies that, though technical generalizability is becoming more important, meaningful clinical benchmarking has yet to be established. Consequently, assertions of clinical utility can be exaggerated in regard to the actual preparedness of these models for actual application. Whereas measurement of discrimination (e.g., AUC) was consistently reported, calibration assessment was less frequent. Only 30.8% of studies (n = 28/91) measured calibration [15,65,66,67,77,83], and almost two-thirds did not measure the alignment of predicted probabilities and observed results [60,92].
Such an imbalance demonstrates a larger pattern in medical AI research, where discrimination is given more importance than probabilistic reliability. The problem is of special interest in the neurosurgical setting, e.g., the glioblastoma prognosis, stroke recovery, or recurrence risk prediction, where probability estimates are themselves directly used by the clinician to make a decision [11,12,14,16,20,21,28,31,34,76,77,78,82,83,86,87,88,89,91,93,95,97]. Such settings have a high chance of poor calibration, resulting in misleading predictions not only in clinical judgment but also in patient counseling.
Notably, despite the high-level of discriminative performance of the models, it is also possible that the outputs of such models are poorly calibrated, which leads to overconfidence or excessive risk stratification. Both discrimination and calibration focus on methodological frameworks like TRIPOD-AI and CONSORT-AI, which underline that a strong model evaluation requires both [7,103]. Enhancement of calibration assessment and reporting is thus most important to achieve safe and effective clinical implementation.
Among the 28 studies reporting calibration (30.8%), calibration curves were the most frequently used metric (89.3%), most often accompanied by decision curve analysis (85.7%); formal calibration statistics such as the Hosmer–Lemeshow test (17.9%) and Brier score (14.3%) were less common, and 67.9% of reporters assessed calibration on both internal and external cohorts. The under-reporting of standard calibration measures, including a calibration curve, Brier scores, or Hosmer-Lemeshow tests, indicates that a lot of modern AI research focuses on comparisons of relative performance, but not on real clinical reliability. In clinical environments, the interpretability and safe application of probability-based predictions can be affected without a regular assessment of calibration. Enhanced reporting of calibration should thus be one of the priorities in order to facilitate sound clinical integration.

4.5. Incomplete Adherence to Standardized Reporting Frameworks

Although there has been an improvement in methodologies, there is still little adherence to established reporting standards. All the studies in this review (n = 91) partially adhered to major elements of the CLAIM and /or TRIPOD-AI frameworks and none of them fulfilled them fully (Supplementary Table S2). Although the majority of the studies reported data partitioning and key performance measures unambiguously, other vital aspects, including a description of model development procedures, preprocessing transparency, calibration studies, and implications of clinical implementation, were reported inconsistently.
Such frameworks as CLAIM and TRIPOD-AI have been designed to improve the transparency, reproducibility and interpretability of AI-based prediction research [7,111]. The lack of their full compliance may interfere with the possibility of readers, reviewers, and clinicians to assess the quality of the methodology accurately and determine possible sources of bias. In addition, inadequate reporting can hide such key problems as data leakage, overfitting, or ineffective validation plans.
The lack of complete compliance with all studies reviewed points to the imbalance between accelerated technological development and the slowness of reporting practices development. Enhancing compliance with standardized structures must, therefore, not be considered as a formal requirement, but rather as an essential measure on the way toward better reproducibility, building trust, and making AI models responsible to clinical application.

4.6. Translational Maturity and Future Directions

Altogether, these results suggest that AI studies in the field of medical imaging are moving in the direction of more rigorous methods, but the existing evidence does not show that it is ready to be implemented in clinical practice. The increased use of multicenter datasets and external validation is a sign of increased awareness of the necessity of generalizability [40,48,51,59]. Simultaneously, persistent gaps in clinical benchmarking, calibration evaluation, and comprehensive reporting also reveal the major areas that need to be addressed to facilitate the translation between technical development and real-life implementation. Interestingly, this review failed to include research in which prospective clinical implementation or formal workflow impact analysis was conducted, a gap that highlights how the field currently focuses on the development of models rather than implementation science. These findings reflect a contemporaneous snapshot of AI medical imaging publications from 2025, following the 2024 CLAIM and TRIPOD-AI updates, rather than a longitudinal assessment of practice over time.
The fact that human comparator analyses are used sparingly and reporting of calibration metrics is low indicates that most models are tuned to be highly statistical but not necessarily reliable in supporting a decision. The presence of strong discriminatory performance itself is not sufficient to be able to safely integrate into complex neurosurgical and radiological workflows, in which probability estimates directly affect treatment decisions, surgical planning, and prognostic evaluation [72,73]. Even in the absence of strong comparison to expert performance and careful calibration evaluation, the danger of overstating clinical utility is significant.
Also, the lack of complete compliance with standardized reporting frameworks restricts reproducibility and independent validation, which would be necessary to receive regulatory approval and clinical confidence [13,19,20,26,30,32,33,36,37,79,85,90,93]. To be able to make meaningful translational impact, AI models should be evaluated not just in terms of accuracy but also in terms of transparency, reliability, and performance in multidisciplinary clinical settings.
Combined, these findings indicate that the discipline is shifting its focus towards early exploratory development to early phases of translational maturity. The way to attain the desired clinical readiness will be to place an increased emphasis on holistic external validation, strict benchmarking of the end product with human experts, open calibration practice, and the maintenance of accepted reporting standards.

4.7. Strengths and Limitations

A number of strengths can be identified with this study. First, it provides a dedicated and current methodological audit of AI imaging studies published in the same year, which allows accurate assessment of the existing validation practices and translational preparation of various medical areas. It is also a method that gives a point of reference in examining how neuroimaging AI has changed as time progresses. Second, a pre-existing, formalized data extraction framework was uniformly used across various areas, such as dataset properties, validation procedure, calibration reporting and compliance with standard guidelines. Third, the risk of data leakage and methodological transparency was systematically defined and evaluated, which made it possible to use cross-study comparisons in a structured manner and not based only on story interpretation.
Nevertheless, a few limitations are to be noted. First, although PubMed was deliberately chosen to align with the clinical focus of this review, the use of a single database may have excluded clinically oriented AI imaging studies published in journals indexed exclusively by EMBASE, Web of Science, or Scopus. This may have introduced a modest selection bias favoring MEDLINE-indexed journals, and our conclusions should therefore be interpreted as characterizing the predominantly clinical AI imaging literature rather than the full multidisciplinary scope of the field.
The review was also restricted to 2025 publications. This year-specific design was chosen to capture the first full publication cycle following the 2024 CLAIM and TRIPOD-AI updates and to enable a contemporaneous audit against current reference standards. Our findings therefore characterize the methodological state of the field at a defined point in time rather than its longitudinal evolution.
Second, adherence to CLAIM and TRIPOD-AI was assessed at the domain level rather than through formal item-level scoring of the 42 CLAIM and 27+ TRIPOD-AI items. Our finding that no study achieved full adherence therefore reflects gaps across major reporting domains rather than precise per-item compliance. Future reviews with a narrower clinical scope could extend this work through formal item-level scoring.
Third, published descriptions were used to classify types of datasets, validation strategies, and risk of data leakage. In situations where reporting was not done fully or clearly, this can cause misclassification, especially in measuring complex or hybrid model architectures.
Lastly, like any other literature-based review, one cannot eliminate the potential of publication bias. The publications of studies with highly favorable findings or high external validation have a higher chance of publication, and this can result in an overestimation of the methodological quality of the field and its clinical preparedness.
Nevertheless, these restrictions do not invalidate the fact that this review offers a detailed and systematic evaluation of the current validation practices, which have significant gaps that need to be reduced to further advance the clinical translation of AI in medical imaging research.

4.8. Future Directions for AI Methodology in Medical Imaging

Further advancements in AI-based medical imaging need to be oriented at going beyond the continued enhancement of model architecture to the enhancement of methodological rigor and clinical integration. External validation must not only progress beyond the replication of datasets but also actual cross-institutional transportability testing, with a variety of scanners, heterogeneous acquisition protocols, and multinational populations of patients. Furthermore, future validation experiments and pragmatic clinical trials will also be necessary to understand whether AI systems have any meaningful contribution to real-world clinical decision-making, but not to improve any statistical measures of performance.
It is also critical to have a regular program of structured benchmarking with human specialists and a regular check of calibration. AI models to be used in decision-support systems should not only exhibit excellent discriminatory results but also robust probability forecasting and evident value addition in multidisciplinary clinical processes. The combination of decision-curve analysis (DCA), clinical impact analysis, and workflow-based analysis will play a significant role in closing the gap between algorithm development and actual clinical utility.
Moreover, compliance with standardized reporting models like CLAIM and TRIPOD-AI should be considered as the key to reproducibility, regulatory approval, and clinical trust. The current stage of AI development in medical imaging must subsequently focus on transparency, methodological maturity, and evidence-based patient-centered outcomes instead of technical innovation.

5. Conclusions

This is a systematic review of AI-based medical imaging studies published in 2025, which shows growing signs of progress in methodological rigor, especially the increased use of multicenter datasets and external validation. Nevertheless, there are continuing weaknesses in clinical benchmarking, calibration evaluation, and overall reporting, which suggest that not all models have been adequately assessed to be used in routine clinical practice. Continued focus on openness, excellent validation, and standard reporting will play a pivotal role in facilitating the safe and successful integration of AI into neurosurgical and radiological practice.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/jcm15093441/s1, Supplementary Data S1: Detailed search strategy; Supplementary Table S1: Summary of included studies; Supplementary Table S2: Full data extraction of included studies; Supplementary Table S3: Risk of bias assessment of included studies using PROBAST+AI; PRISMA 2020 checklist.

Author Contributions

M.K.B.: conceptualization, supervision, project administration, review and editing. U.S., N.S., A.M., B.D., Y.S., A.R.B., M.S.A.-J., I.U., O.O., M.N., E.Ö., S.G.A., A.K., U.E.: review and editing. U.S., N.S., U.E. and Y.S.: writing—original draft preparation, review and editing. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial intelligence
AUCArea under the receiver operating characteristic curve
C-indexConcordance index
CLAIMChecklist for Artificial Intelligence in Medical Imaging
CONSORT-AIConsolidated Standards of Reporting Trials for Artificial Intelligence
CTComputed tomography
CTAComputed tomography angiography
DCADecision-curve analysis
DLDeep learning
MLMachine learning
MRAMagnetic resonance angiography
MRIMagnetic resonance imaging
PRISMAPreferred Reporting Items for Systematic Reviews and Meta-Analyses
PROBAST+AIPrediction model Risk of Bias Assessment Tool for Artificial Intelligence
PROSPEROInternational Prospective Register of Systematic Reviews
TRIPOD-AITransparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis using Artificial Intelligence

References

  1. Wong, A.; Otles, E.; Donnelly, J.P.; Krumm, A.; McCullough, J.; DeTroyer-Cooley, O.; Pestrue, J.; Phillips, M.; Konye, J.; Penoza, C.; et al. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Intern. Med. 2021, 181, 1065–1070. [Google Scholar] [CrossRef]
  2. Ötleş, E.; Denton, B.T.; Qu, B.; Murali, A.; Merdan, S.; Auffenberg, G.B.; Hiller, S.C.; Lane, B.R.; George, A.K.; Singh, K. Development and Validation of Models to Predict Pathological Outcomes of Radical Prostatectomy in Regional and National Cohorts. J. Urol. 2022, 207, 358–366. [Google Scholar] [CrossRef]
  3. Nagendran, M.; Chen, Y.; Lovejoy, C.A.; Gordon, A.C.; Komorowski, M.; Harvey, H.; Topol, E.J.; Ioannidis, J.P.A.; Collins, G.S.; Maruthappu, M. Artificial intelligence versus clinicians: Systematic review of design, reporting standards, and claims of deep learning studies. BMJ 2020, 368, m689. [Google Scholar] [CrossRef] [PubMed]
  4. Koçak, B.; Köse, F.; Keleş, A.; Şendur, A.; Meşe, İ.; Karagülle, M. Adherence to the Checklist for Artificial Intelligence in Medical Imaging (CLAIM): An umbrella review with a comprehensive two-level analysis. Diagn. Interv. Radiol. 2025, 31, 440–455. [Google Scholar] [CrossRef]
  5. Chen, Z.; Liu, X.; Yang, Q.; Wang, Y.J.; Miao, K.; Gong, Z.; Yu, Y.; Leonov, A.; Liu, C.; Feng, Z.; et al. Evaluation of Risk of Bias in Neuroimaging-Based Artificial Intelligence Models for Psychiatric Diagnosis: A Systematic Review. JAMA Netw. Open 2023, 6, e231671. [Google Scholar] [CrossRef]
  6. Tejani, A.S.; Klontzas, M.E.; Gatti, A.A.; Mongan, J.T.; Moy, L.; Park, S.H.; Kahn, C.E., Jr.; Panel, C.U. Checklist for Artificial Intelligence in Medical Imaging (CLAIM): 2024 Update. Radiol. Artif. Intell. 2024, 6, e240300. [Google Scholar] [CrossRef]
  7. Collins, G.S.; Moons, K.G.M.; Dhiman, P.; Riley, R.D.; Beam, A.L.; Van Calster, B.; Ghassemi, M.; Liu, X.; Reitsma, J.B.; van Smeden, M.; et al. TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024, 385, e078378. [Google Scholar] [CrossRef] [PubMed]
  8. Martindale, A.P.L.; Llewellyn, C.D.; de Visser, R.O.; Ng, B.; Ngai, V.; Kale, A.U.; di Ruffano, L.F.; Golub, R.M.; Collins, G.S.; Moher, D.; et al. Concordance of randomised controlled trials for artificial intelligence interventions with the CONSORT-AI reporting guidelines. Nat. Commun. 2024, 15, 1619. [Google Scholar] [CrossRef]
  9. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef]
  10. Moons, K.G.M.; Damen, J.A.A.; Kaul, T.; Hooft, L.; Andaur Navarro, C.; Dhiman, P.; Beam, A.L.; Van Calster, B.; Celi, L.A.; Denaxas, S.; et al. PROBAST+AI: An updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ 2025, 388, e082505. [Google Scholar] [CrossRef] [PubMed]
  11. Chen, J.; Wang, Z.; Yang, B. mpMRI-based MGMT methylation status prediction for glioblastoma through off-the-shelf deep features: A multi-dataset feasibility study. J. Appl. Clin. Med. Phys. 2025, 26, e70373. [Google Scholar] [CrossRef]
  12. Chen, Y.; Rivier, C.A.; Mora, S.A.; Torres Lopez, V.; Payabvash, S.; Sheth, K.N.; Harloff, A.; Falcone, G.J.; Rosand, J.; Mayerhofer, E.; et al. Deep learning survival model predicts outcome after intracerebral hemorrhage from initial CT scan. Eur. Stroke J. 2025, 10, 225–235. [Google Scholar] [CrossRef]
  13. Dai, Y.; Zhong, Z.; Qin, Y.; Wang, Y.; Yu, G.; Kobets, A.; Swenson, D.W.; Boxerman, J.L.; Li, G.; Robinson, S.; et al. AI Model Integrating Imaging and Clinical Data for Predicting CSF Diversion in Neonatal Hydrocephalus: A Preliminary Study. Hum. Brain Mapp. 2025, 46, e70363. [Google Scholar] [CrossRef]
  14. Dong, Y.; Pachade, S.; Roberts, K.; Jiang, X.; Sheth, S.A.; Giancardo, L. Generalizable self-supervised learning for brain CTA in acute stroke. Comput. Biol. Med. 2025, 184, 109337. [Google Scholar] [CrossRef]
  15. Gui, Y.; Hu, W.; Ren, J.; Tang, F.; Wang, L.; Zhang, F.; Zhang, J. Preoperative diagnosis of meningioma sinus invasion based on MRI radiomics and deep learning: A multicenter study. Cancer Imaging 2025, 25, 20. [Google Scholar] [CrossRef]
  16. Hamon, G.; Legrand, L.; Hmeydia, G.; Turc, G.; Hassen, W.B.; Charron, S.; Debacker, C.; Naggara, O.; Thirion, B.; Chen, B.; et al. Multicenter validation of synthetic FLAIR as a substitute for FLAIR sequence in acute ischemic stroke. Eur. Stroke J. 2025, 10, 161–171. [Google Scholar] [CrossRef] [PubMed]
  17. Jeon, E.T.; Kim, S.M.; Jung, J.M. Automated rating of Fazekas scale in fluid-attenuated inversion recovery MRI for ischemic stroke or transient ischemic attack using machine learning. Sci. Rep. 2025, 15, 32219. [Google Scholar] [CrossRef] [PubMed]
  18. Kamel, P.; Kanhere, A.; Kulkarni, P.; Khalid, M.; Steger, R.; Bodanapally, U.; Gandhi, D.; Parekh, V.; Yi, P.H. Optimizing Acute Stroke Segmentation on MRI Using Deep Learning: Self-Configuring Neural Networks Provide High Performance Using Only DWI Sequences. J. Imaging Inform. Med. 2025, 38, 717–726. [Google Scholar] [CrossRef] [PubMed]
  19. Kesari, A.; Maurya, S.; Sheikh, M.T.; Gupta, R.K.; Singh, A. Large blood vessel segmentation in quantitative DCE-MRI of brain tumors: A Swin UNETR approach. Magn. Reson. Imaging 2025, 118, 110342. [Google Scholar] [CrossRef]
  20. Ketabi, S.; Wagner, M.W.; Hawkins, C.; Tabori, U.; Ertl-Wagner, B.B.; Khalvati, F. Multimodal contrastive learning for enhanced explainability in pediatric brain tumor molecular diagnosis. Sci. Rep. 2025, 15, 10943. [Google Scholar] [CrossRef]
  21. Krag, C.H.; Muller, F.C.; Gandrup, K.L.; Plesner, L.L.; Sagar, M.V.; Andersen, M.B.; Nielsen, M.; Kruuse, C.; Boesen, M. Impact of spectrum bias on deep learning-based stroke MRI analysis. Eur. J. Radiol. 2025, 188, 112161. [Google Scholar] [CrossRef]
  22. Kulathilake, C.D.; Udupihille, J.; Abeysundara, S.P.; Senoo, A. Deep learning-driven multi-class classification of brain strokes using computed tomography: A step towards enhanced diagnostic precision. Eur. J. Radiol. 2025, 187, 112109. [Google Scholar] [CrossRef]
  23. Liang, X.; Ke, X.; Hu, W.; Jiang, J.; Li, S.; Xue, C.; Liu, X.; Dend, J.; Yan, C.; Gao, M.; et al. Deep learning radiomic nomogram outperforms the clinical model in distinguishing intracranial solitary fibrous tumors from angiomatous meningiomas and can predict patient prognosis. Eur. Radiol. 2025, 35, 2670–2680. [Google Scholar] [CrossRef]
  24. Lilhore, U.K.; Sunder, R.; Simaiya, S.; Alsafyani, M.; Monish Khan, M.D.; Alroobaea, R.; Alsufyani, H.; Baqasah, A.M. AG-MS3D-CNN multiscale attention guided 3D convolutional neural network for robust brain tumor segmentation across MRI protocols. Sci. Rep. 2025, 15, 24306. [Google Scholar] [CrossRef]
  25. Lin, X.; Zou, E.; Chen, W.; Chen, X.; Lin, L. Advanced multi-label brain hemorrhage segmentation using an attention-based residual U-Net model. BMC Med. Inform. Decis. Mak. 2025, 25, 286. [Google Scholar] [CrossRef]
  26. Liu, J.; Xing, F.; Elkhill, C.; Linguraru, M.G.; Miles, R.C.; Cruz-Guerrero, I.A.; Porras, A.R. Population-Driven Synthesis of Personalized Cranial Development From Cross-Sectional Pediatric CT Images. IEEE Trans. Biomed. Eng. 2025, 72, 2732–2741. [Google Scholar] [CrossRef]
  27. Lv, C.; Shu, X.J.; Qiu, J.; Xiong, Z.C.; Bo Ye, J.; Bo Li, S.; Chen, S.B.; Rao, H. AI-enabled precise brain tumor segmentation by integrating Refinenet and contour-constrained features in MRI images. Med. Phys. 2025, 52, e17958. [Google Scholar] [CrossRef] [PubMed]
  28. Pelcat, A.; Le Berre, A.; Ben Hassen, W.; Debacker, C.; Charron, S.; Thirion, B.; Legrand, L.; Turc, G.; Oppenheim, C.; Benzakoun, J. Generative T2*-weighted images as a substitute for true T2*-weighted images on brain MRI in patients with acute stroke. Diagn. Interv. Imaging 2025, 106, 264–271. [Google Scholar] [CrossRef]
  29. Rastogi, D.; Johri, P.; Donelli, M.; Kadry, S.; Khan, A.A.; Espa, G.; Feraco, P.; Kim, J. Deep learning-integrated MRI brain tumor analysis: Feature extraction, segmentation, and Survival Prediction using Replicator and volumetric networks. Sci. Rep. 2025, 15, 1437. [Google Scholar] [CrossRef] [PubMed]
  30. Ruhling, S.; Petzsche, M.R.H.; Loffler, M.T.; Sollmann, N.; Baum, T.; Bodden, J.; Schwarting, J.; Lange, N.; Aftahy, K.; Wostrack, M.; et al. Opportunistic osteoporosis screening in intraoperative CT can accurately identify patients with low volumetric bone mineral density and osteoporosis during spine surgery. Eur. Spine J. 2025, 34, 1461–1469. [Google Scholar] [CrossRef] [PubMed]
  31. Tu, J.; Shen, C.; Liu, J.; Hu, B.; Chen, Z.; Yan, Y.; Li, C.; Xiong, J.; Daoud, A.M.; Wang, X.; et al. Detection of Microscopic Glioblastoma Infiltration in Peritumoral Edema Using Interactive Deep Learning With DTI Biomarkers: Testing via Stereotactic Biopsy. J. Magn. Reson. Imaging 2025, 62, 1802–1811. [Google Scholar] [CrossRef]
  32. Wang, B.; Wei, J.; Wang, Z.; Niu, P.; Yang, L.; Hu, Y.; Shao, D.; Zhao, W. Development of a deep learning-based MRI diagnostic model for human Brucella spondylitis. BioMed. Eng. Online 2025, 24, 87. [Google Scholar] [CrossRef]
  33. Wang, K.; Lin, F.; Liao, Z.; Wang, Y.; Zhang, T.; Wang, R. Development of a Dual-Plane MRI-Based Deep Learning Model to Assess the 1-Year Postoperative Outcomes in Lumbar Disc Herniation After Tubular Microdiscectomy. J. Magn. Reson. Imaging 2025, 61, 2294–2307. [Google Scholar] [CrossRef]
  34. Wang, T.; Jiang, C.; Ding, W.; Chen, Q.; Shen, D.; Ding, Z. Deep-Learning Generated Synthetic Material Decomposition Images Based on Single-Energy CT to Differentiate Intracranial Hemorrhage and Contrast Staining Within 24 Hours After Endovascular Thrombectomy. CNS Neurosci. Ther. 2025, 31, e70235. [Google Scholar] [CrossRef]
  35. Wang, Y.; Wen, Z.; Bao, S.; Huang, D.; Wang, Y.; Yang, B.; Li, Y.; Zhou, P.; Zhang, H.; Pang, H. Diffusion-CSPAM U-Net: A U-Net model integrated hybrid attention mechanism and diffusion model for segmentation of computed tomography images of brain metastases. Radiat. Oncol. 2025, 20, 50. [Google Scholar] [CrossRef]
  36. Xing, L.P.; Liu, G.; Zhang, H.C.; Wang, L.; Zhu, S.; Bao, M.D.H.; Wang, Y.N.; Chen, C.; Wang, Z.; Liu, X.Y.; et al. Evaluating CNN Architectures for the Automated Detection and Grading of Modic Changes in MRI: A Comparative Study. Orthop. Surg. 2025, 17, 233–243. [Google Scholar] [CrossRef] [PubMed]
  37. Zahoora, U.; Shahid, A.R.; Gondal, F.F. A bias-resilient client selection analysis for federated brain tumor segmentation. Sci. Rep. 2025, 15, 37670. [Google Scholar] [CrossRef] [PubMed]
  38. Jia, X.; Chen, Y.; Zheng, K.; Chen, C.; Liu, J. Deep Learning-Driven Multimodal Fusion Model for Prediction of Middle Cerebral Artery Aneurysm Rupture Risk. Acad. Radiol. 2025, 32, 6114–6124. [Google Scholar] [CrossRef]
  39. Harper, J.P.; Lee, G.R.; Pan, I.; Nguyen, X.V.; Quails, N.; Prevedello, L.M. External Validation of a Winning Artificial Intelligence Algorithm from the RSNA 2022 Cervical Spine Fracture Detection Challenge. AJNR Am. J. Neuroradiol. 2025, 46, 1852–1858. [Google Scholar] [CrossRef] [PubMed]
  40. Dai, M.; Tiu, B.C.; Schlossman, J.; Ayobi, A.; Castineira, C.; Kiewsky, J.; Avare, C.; Chaibi, Y.; Chang, P.; Chow, D.; et al. Validation of a Deep Learning Tool for Detection of Incidental Vertebral Compression Fractures. J. Comput. Assist. Tomogr. 2025, 49, 669–674. [Google Scholar] [CrossRef]
  41. Kang, D.W.; Kim, M.; Park, G.H.; Kim, Y.S.; Han, M.K.; Lee, M.; Kim, D.; Ryu, W.S.; Jeong, H.G. Deep learning-assisted detection of intracranial hemorrhage: Validation and impact on reader performance. Neuroradiology 2025, 67, 1511–1519. [Google Scholar] [CrossRef] [PubMed]
  42. Hossain, M.M.; Ahmed, M.M.; Nafi, A.A.N.; Islam, M.R.; Ali, M.S.; Haque, J.; Miah, M.S.; Rahman, M.M.; Islam, M.K. A novel hybrid ViT-LSTM model with explainable AI for brain stroke detection and classification in CT images: A case study of Rajshahi region. Comput. Biol. Med. 2025, 186, 109711. [Google Scholar] [CrossRef]
  43. Chen, R.; Lu, Y.; Tian, Z.; Chen, J.; Li, W.; Wang, C.; Zhang, Z.; Huang, X.; Ding, C.; Liu, X.; et al. DWI-based deep learning radiomics nomogram for predicting the impaired quality of life in patients with unruptured intracranial aneurysm developing new iatrogenic cerebral infarcts following stent placement: A multicenter cohort study. Neurosurg. Rev. 2025, 48, 508. [Google Scholar] [CrossRef]
  44. Felefly, T.; Francis, Z.; Roukoz, C.; Fares, G.; Achkar, S.; Yazbeck, S.; Nasr, A.; Kordahi, M.; Azoury, F.; Nasr, D.N.; et al. A 3D Convolutional Neural Network Based on Non-enhanced Brain CT to Identify Patients with Brain Metastases. J. Imaging Inform. Med. 2025, 38, 858–864. [Google Scholar] [CrossRef] [PubMed]
  45. Nalentzi, K.; Gerogiannis, K.; Bougias, H.; Stogiannos, N.; Papavasileiou, P. Comparative analysis of transformer-based deep learning models for glioma and meningioma classification. J. Med. Imaging Radiat. Sci. 2025, 56, 102008. [Google Scholar] [CrossRef]
  46. Liao, L.; Puel, U.; Sabardu, O.; Harsan, O.; Medeiros, L.L.; Loukoul, W.A.; Anxionnat, R.; Kerrien, E. AI-assisted detection of cerebral aneurysms on 3D time-of-flight MR angiography: User variability and clinical implications. J. Neuroradiol. 2025, 52, 101388. [Google Scholar] [CrossRef]
  47. Pettersson, S.D.; Filo, J.; Liaw, P.; Skrzypkowska, P.; Klepinowski, T.; Szmuda, T.; Fodor, T.B.; Ramirez-Velandia, F.; Zielinski, P.; Chang, Y.M.; et al. Addressing Limited Generalizability in Artificial Intelligence-Based Brain Aneurysm Detection for Computed Tomography Angiography: Development of an Externally Validated Artificial Intelligence Screening Platform. Neurosurgery 2025, 97, 1388–1396. [Google Scholar] [CrossRef] [PubMed]
  48. Ryu, W.S.; Schellingerhout, D.; Park, J.; Chung, J.; Jeong, S.W.; Gwak, D.S.; Kim, B.J.; Kim, J.T.; Hong, K.S.; Lee, K.B.; et al. Deep learning-based automatic segmentation of cerebral infarcts on diffusion MRI. Sci. Rep. 2025, 15, 13214. [Google Scholar] [CrossRef]
  49. Yang, H.; Wang, W.; Zhao, X.; Xuan, Q.; Jiang, C.; Zhao, B. Deep learning models based on DWI-MRI for prognosis prediction in acute ischemic stroke receiving intravenous thrombolysis: Development and validation. J. Neuroradiol. 2025, 52, 101391. [Google Scholar] [CrossRef]
  50. Kong, C.; Yan, D.; Liu, K.; Yin, Y.; Ma, C. Multiple deep learning models based on MRI images in discriminating glioblastoma from solitary brain metastases: A multicentre study. BMC Med. Imaging 2025, 25, 171. [Google Scholar] [CrossRef]
  51. Topff, L.; Petrychenko, L.; Jain, N.; Lingier, S.; Bertels, J.; Astudillo, P.; Prosec, M.; Menendez Fernandez-Miranda, P.; Gevaert, O.; Smits, M.; et al. A Data-Centric Approach to Deep Learning for Brain Metastasis Analysis at MRI. Radiology 2025, 315, e242416. [Google Scholar] [CrossRef]
  52. Mahootiha, M.; Tak, D.; Ye, Z.; Zapaishchykova, A.; Likitlersuang, J.; Climent Pardo, J.C.; Boyd, A.; Vajapeyam, S.; Chopra, R.; Prabhu, S.P.; et al. Multimodal deep learning improves recurrence risk prediction in pediatric low-grade gliomas. Neuro-Oncology 2025, 27, 277–290. [Google Scholar] [CrossRef] [PubMed]
  53. Zeng, L.; Wen, L.; Jing, Y.; Xu, J.X.; Huang, C.C.; Zhang, D.; Wang, G.X. Assessment of the stability of intracranial aneurysms using a deep learning model based on computed tomography angiography. Radiol. Medica 2025, 130, 248–257. [Google Scholar] [CrossRef]
  54. Tuxunjiang, P.; Huang, C.; Zhou, Z.; Zhao, W.; Han, B.; Tan, W.; Wang, J.; Kukun, H.; Zhao, W.; Xu, R.; et al. Prediction of NIHSS Scores and Acute Ischemic Stroke Severity Using a Cross-attention Vision Transformer Model with Multimodal MRI. Acad. Radiol. 2025, 32, 5453–5467. [Google Scholar] [CrossRef]
  55. Sina, E.M.; Limage, K.; Anisman, E.; Pudik, N.; Tam, E.; Kahn, C.; Daggumati, S.; Evans, J.J.; Rabinowitz, M.R.; Rosen, M.R.; et al. Automated Machine Learning Differentiation of Pituitary Macroadenomas and Parasellar Meningiomas Using Preoperative Magnetic Resonance Imaging. Otolaryngol.–Head Neck Surg. 2025, 173, 1376–1384. [Google Scholar] [CrossRef] [PubMed]
  56. Wang, H.; Schwirtlich, T.; Houskamp, E.; Hutch, M.; Murphy, J.; Nascimento, J.; Zini, A.; Brancaleoni, L.; Giacomozzi, S.; Luo, Y.; et al. Automated Detection of the Black Hole Sign for Patients with Intracerebral Hemorrhage Using Self-Supervised Learning. AJNR Am. J. Neuroradiol. 2025, 46, 2300–2309. [Google Scholar] [CrossRef]
  57. Nada, A.; Sayed, A.A.; Hamouda, M.; Tantawi, M.; Khan, A.; Alt, A.; Hassanein, H.; Sevim, B.C.; Altes, T.; Gaballah, A. External validation and performance analysis of a deep learning-based model for the detection of intracranial hemorrhage. Neuroradiol. J. 2025, 38, 312–321. [Google Scholar] [CrossRef]
  58. Hao, M.; Yan, J.; Wang, X.; Tan, Y.; Zhang, H.; Yang, G. Survival prediction in gliomas based on MRI radiomics combined with clinical factors and molecular biomarkers. PeerJ 2025, 13, e19906. [Google Scholar] [CrossRef]
  59. Akbari, H.; Bakas, S.; Sako, C.; Fathi Kazerooni, A.; Villanueva-Meyer, J.; Garcia, J.A.; Mamourian, E.; Liu, F.; Cao, Q.; Shinohara, R.T.; et al. Machine learning-based prognostic subgrouping of glioblastoma: A multicenter study. Neuro-Oncol. 2025, 27, 1102–1115. [Google Scholar] [CrossRef] [PubMed]
  60. Belke, M.; Zahnert, F.; Steinbrenner, M.; Halimeh, M.; Miron, G.; Tsalouchidou, P.E.; Linka, L.; Keil, B.; Jansen, A.; Möschl, V.; et al. Automatic detection of hippocampal sclerosis in patients with epilepsy. Epilepsia 2025, 66, 3852–3864. [Google Scholar] [CrossRef]
  61. Demirel, E.; Dilek, O. Utilizing Radiomics of Peri-Lesional Edema in T2-FLAIR Subtraction Digital Images to Distinguish High-Grade Glial Tumors From Brain Metastasis. J. Magn. Reson. Imaging 2025, 61, 1728–1737. [Google Scholar] [CrossRef] [PubMed]
  62. Fatania, K.; Frood, R.; Mistry, H.; Short, S.C.; O’Connor, J.; Scarsbrook, A.F.; Currie, S. Impact of intensity standardisation and ComBat batch size on clinical-radiomic prognostic models performance in a multi-centre study of patients with glioblastoma. Eur. Radiol. 2025, 35, 3354–3366. [Google Scholar] [CrossRef]
  63. Foltyn-Dumitru, M.; Mahmutoglu, M.A.; Brugnara, G.; Kessler, T.; Sahm, F.; Wick, W.; Heiland, S.; Bendszus, M.; Vollmuth, P.; Schell, M. Shape matters: Unsupervised exploration of IDH-wildtype glioma imaging survival predictors. Eur. Radiol. 2025, 35, 1351–1360. [Google Scholar] [CrossRef] [PubMed]
  64. Roh, Y.H.; Cheong, E.N.; Jung, S.C.; Yun, J.; Ko, J.S.; Cho, S.J.; Choi, K.M.; Park, S.I.; Jeong, S.Y.; Lee, D.H.; et al. Prediction of hemorrhagic transformation in acute ischemic stroke patients using clinico-radiomics models. Sci. Rep. 2025, 15, 38628. [Google Scholar] [CrossRef] [PubMed]
  65. Yang, Q.; Wang, Y.; Wu, J.; Hu, H.; He, Y.; Wang, Y.; Yang, B. Preoperative prediction of pituitary neuroendocrine tumor consistency based on multiparametric MRI radiomics: A multicenter study. BMC Cancer 2025, 25, 1501. [Google Scholar] [CrossRef]
  66. Fan, Y.; Zhang, W.; Mou, A.; Fang, H.; Guo, S.; Feng, M. Development and Validation of Multiparametric MRI-based Clini-radiomic Model for Preoperative Prediction of Somatotroph Adenomas Subtypes: A Multicenter Study. Acad. Radiol. 2025, 32, 7471–7485. [Google Scholar] [CrossRef] [PubMed]
  67. Huang, L.; Xu, X.; Tian, B.; Liao, A.; Wang, L.; Shen, X.; Cao, Z.; Liu, X.; Lu, S.; Li, J.; et al. Development and validation of a radiomics model based on the ASPECTS framework using CT imaging for predicting malignant cerebral edema. Eur. J. Radiol. 2025, 192, 112410. [Google Scholar] [CrossRef]
  68. Liang, Q.; Duan, X.; Yan, H.; Li, X.; Li, Z.; Niu, W.; Liu, X.; Tan, Y.; Wang, X.; Yang, G.; et al. Development and validation of radiopathomics models for predicting molecular subtypes and WHO grades in adult-type diffuse gliomas: A multicenter study. J. Transl. Med. 2025, 23, 1120. [Google Scholar] [CrossRef]
  69. Xia, X.; Liu, J.; Cui, J.; You, Y.; Huang, C.; Li, H.; Zhang, D.; Ren, Q.; Jiang, Q.; Meng, X. A nomogram incorporating CT-based peri-hematoma radiomics features to predict functional outcome in patients with intracerebral hemorrhage. Eur. J. Radiol. 2025, 183, 111871. [Google Scholar] [CrossRef]
  70. Liu, J.; Tu, J.; Hu, B.; Li, C.; Piao, S.; Lu, Y.; Li, A.; Ding, T.; Xiong, J.; Zhu, F.; et al. Prognostic Assessment in Patients With Primary Diffuse Large B-Cell Lymphoma of the Central Nervous System Using MRI-Based Radiomics. J. Magn. Reson. Imaging 2025, 61, 1142–1152. [Google Scholar] [CrossRef]
  71. Song, D.; Wei, Q.; Zhao, S.; Lou, Y.; Zhang, K.; Duan, C.; Wang, F.; Gao, Q.; Yan, J.; Yan, D.; et al. Exploring a recurrence model for atypical meningioma based on multiparametric MRI radiomic and clinical characteristics: A multicenter retrospective cohort study. Radiat. Oncol. 2025, 20, 30. [Google Scholar] [CrossRef] [PubMed]
  72. Patel, B.K.; Zohdy, Y.M.; Lohana, S.; Tariciotti, L.; Rodas, A.; Alawieh, A.; Jahangiri, A.; Faraj, R.R.; Maldonado, J.; Uribe-Pacheco, R.; et al. Predictive Modeling of Nonfunctioning Giant Pituitary Neuroendocrine Tumor Resection: A Multi-Planar Perspective. World Neurosurg. 2025, 195, 123653. [Google Scholar] [CrossRef]
  73. Zheng, B.; Zhu, Z.; Ma, K.; Liang, Y.; Liu, H. Three-Dimensional Radiomics and Machine Learning for Predicting Postoperative Outcomes in Laminoplasty for Cervical Spondylotic Myelopathy: A Clinical-Radiomics Model. World Neurosurg. 2025, 203, 124464. [Google Scholar] [CrossRef]
  74. Hu, W.; Lin, G.; Chen, W.; Wu, J.; Zhao, T.; Xu, L.; Qian, X.; Shen, L.; Yan, Z.; Chen, M.; et al. Radiomics based on dual-energy CT virtual monoenergetic images to identify symptomatic carotid plaques: A multicenter study. Sci. Rep. 2025, 15, 10415. [Google Scholar] [CrossRef]
  75. Cai, Z.Y.; Hu, K.; Linli, Z.Q. Sexual dimorphism of white-matter functional connectome in healthy young adults. Prog. Neuro-Psychopharmacol. Biol. Psychiatry 2025, 142, 111486. [Google Scholar] [CrossRef]
  76. Feng, L.; Han, H.; Mo, J.; Huang, Y.; Huang, K.; Zhou, C.; Wang, X.; Zhang, J.; Yang, Z.; Liu, D.; et al. Individualized structural network deviations predict surgical outcome in mesial temporal lobe epilepsy: A multicenter validation study. Int. Surg. J. 2025, 111, 7594–7605. [Google Scholar] [CrossRef]
  77. Ma, Z.; Zhang, C.; Guo, Y.; Li, Y.; Wang, Z.; Hu, Y.; Zhang, X.; Duan, M.; Wang, W.; Yan, D.; et al. Improving radiomics-based differentiation of supratentorial malignant brain tumors preoperatively with diffusion-weighted imaging: A three-class machine learning algorithm. Eur. J. Surg. Oncol. 2025, 51, 110533. [Google Scholar] [CrossRef]
  78. Sun, K.; Shi, R.; Yu, X.; Wang, Y.; Zhang, W.; Yang, X.; Zhang, M.; Wang, J.; Jiang, S.; Li, H.; et al. Noninvasive imaging biomarker reveals invisible microscopic variation in acute ischaemic stroke (</= 24 h): A multicentre retrospective study. Sci. Rep. 2025, 15, 3743. [Google Scholar] [CrossRef] [PubMed]
  79. Sunavsky, A.; Hashmi, M.A.; Robertson, J.W.; Veinot, J.; Hashmi, J.A. The nucleus accumbens-prefrontal connectivity as a predictor of chronic low back pain. Pain 2025, 166, e363–e377. [Google Scholar] [CrossRef]
  80. Wang, G.; Zhang, Y.; Xu, L.; Ni, J.; Shen, Y.; Jin, Q. Radiomics-based diagnosis of carotid artery stenosis using non-contrast CT: Model development and validation. Eur. J. Med. Res. 2025, 30, 1237. [Google Scholar] [CrossRef] [PubMed]
  81. Wang, H.; Kong, J.F.; Wen, L.; Wang, X.J.; Zhang, W.T.; Wang, Z.Q.; Zeng, L.; Huang, Y.T.; Yang, S.H.; Li, M.; et al. Development of predictive models to identify the intracranial aneurysm responsible for subarachnoid hemorrhage in patients with multiple saccular aneurysms. Eur. J. Radiol. 2025, 193, 112466. [Google Scholar] [CrossRef]
  82. Xia, X.; Wu, W.; Tan, Q.; Gou, Q. Interpretable Machine Learning Models for Differentiating Glioblastoma From Solitary Brain Metastasis Using Radiomics. Acad. Radiol. 2025, 32, 5388–5400. [Google Scholar] [CrossRef] [PubMed]
  83. Xu, W.; Li, Y.; Zhang, J.; Zhang, Z.; Shen, P.; Wang, X.; Yang, G.; Du, J.; Zhang, H.; Tan, Y. Predicting the molecular subtypes of 2021 WHO grade 4 glioma by a multiparametric MRI-based machine learning model. BMC Cancer 2025, 25, 1171. [Google Scholar] [CrossRef]
  84. Xu, X.; Zhou, Y.; Sun, S.; Cui, L.; Chen, Z.; Guo, Y.; Jiang, J.; Wang, X.; Sun, T.; Yang, Q.; et al. Risk prediction for elderly cognitive impairment by radiomic and morphological quantification analysis based on a cerebral MRA imaging cohort. Eur. Radiol. 2025, 35, 4300–4314. [Google Scholar] [CrossRef]
  85. Ye, B.; Sun, Y.; Chen, G.; Wang, B.; Meng, H.; Shan, L. Development and validation of machine learning models to predict vertebral artery injury by C2 pedicle screws. Eur. Spine J. 2025, 34, 3950–3961. [Google Scholar] [CrossRef]
  86. Zeng, Q.; Jia, F.; Tang, S.; He, H.; Fu, Y.; Wang, X.; Zhang, J.; Tan, Z.; Tang, H.; Wang, J.; et al. Ensemble learning-based radiomics model for discriminating brain metastasis from glioblastoma. Eur. J. Radiol. 2025, 183, 111900. [Google Scholar] [CrossRef] [PubMed]
  87. Zhai, D.; Wu, Y.; Cui, M.; Liu, Y.; Zhou, X.; Hu, D.; Wang, Y.; Ju, S.; Fan, G.; Cai, W. Combinations of Clinical Factors, CT Signs, and Radiomics for Differentiating High-Density Areas after Mechanical Thrombectomy in Patients with Acute Ischemic Stroke. AJNR Am. J. Neuroradiol. 2025, 46, 66–74. [Google Scholar] [CrossRef]
  88. Zhao, K.; Chen, C.; Zhang, Y.; Huang, Z.; Zhao, Y.; Yue, Q.; Xu, J. Preoperative Assessment of Ki-67 Labeling Index in Pituitary Adenomas Using Delta-Radiomics Based on Dynamic Contrast-Enhanced MRI. J. Magn. Reson. Imaging 2025, 62, 508–518. [Google Scholar] [CrossRef]
  89. Zhao, K.; Deng, Y.; Su, X.; Hu, W.; Yin, T.; Yang, X.; Zhang, D.; Sun, J.; Li, Y.; Xu, J.; et al. Differential Diagnosis of Early-Stage Atypical Primary Central Nervous System Lymphoma and Low-Grade Glioma Using Magnetic Resonance Imaging-Based Radiomics. World Neurosurg. 2025, 196, 123740. [Google Scholar] [CrossRef]
  90. Zhuang, X.; Wang, J.; Kang, J.; Lin, Z. Diagnosis of Acute Versus Chronic Thoracolumbar Vertebral Compression Fractures Using CT Radiomics Based on Machine Learning: A Preliminary Study. J. Imaging Inform. Med. 2025, 38, 2183–2193. [Google Scholar] [CrossRef] [PubMed]
  91. Sun, Y.; Wang, Y.; Jiang, M.; Jia, W.; Chen, H.; Wang, H.; Ding, Y.; Wang, X.; Yang, C.; Sun, B.; et al. Habitat-based MRI radiomics to predict the origin of brain metastasis. Med. Phys. 2025, 52, 3075–3087. [Google Scholar] [CrossRef]
  92. Albadr, R.J.; Sur, D.; Yadav, A.; Rekha, M.M.; Jain, B.; Jayabalan, K.; Kubaev, A.; Taher, W.M.; Alwan, M.; Jawad, M.J.; et al. Optimizing meningioma grading with radiomics and deep features integration, attention mechanisms, and reproducibility analysis. Eur. J. Med. Res. 2025, 30, 808. [Google Scholar] [CrossRef]
  93. Chen, Y.; Hu, X.; Fan, T.; Zhou, Y.; Yu, C.; Yu, J.; Zhou, X.; Wang, B. Predicting Postoperative Prognosis in Pediatric Malignant Tumor With MRI Radiomics and Deep Learning Models: A Retrospective Study. J. Craniofacial Surg. 2025, 36, 1929–1935. [Google Scholar] [CrossRef] [PubMed]
  94. Choi, J.H.; Sobisch, J.; Kim, M.; Park, J.C.; Ahn, J.S.; Kwun, B.D.; Spiclin, Z.; Bizjak, Z.; Park, W. Prediction of intracranial aneurysm rupture from computed tomography angiography using an automated artificial intelligence framework. Comput. Biol. Med. 2025, 197, 110965. [Google Scholar] [CrossRef]
  95. Li, D.; Hu, W.; Ma, L.; Yang, W.; Liu, Y.; Zou, J.; Ge, X.; Han, Y.; Gan, T.; Cheng, D.; et al. Deep learning radiomics nomograms predict Isocitrate dehydrogenase (IDH) genotypes in brain glioma: A multicenter study. Magn. Reson. Imaging 2025, 117, 110314. [Google Scholar] [CrossRef]
  96. Li, Z.; Gao, Y.; Zhang, K.; Wang, J.; Han, L.; Xie, B.; Sun, Y.; Yan, R.; Li, Y.; Cui, H. Predicting the prognosis of symptomatic intracranial atherosclerotic stenosis (sICAS) patients using deep learning models: A multicenter study based on high-resolution magnetic resonance vessel wall imaging. Clin. Radiol. 2025, 91, 107092. [Google Scholar] [CrossRef] [PubMed]
  97. Liu, J.; Jiang, S.; Wu, Y.; Zou, R.; Bao, Y.; Wang, N.; Tu, J.; Xiong, J.; Liu, Y.; Li, Y. Deep learning-based radiomics and machine learning for prognostic assessment in IDH-wildtype glioblastoma after maximal safe surgical resection: A multicenter study. Int. Surg. J. 2025, 111, 4576–4585. [Google Scholar] [CrossRef] [PubMed]
  98. Saadh, M.J.; Albadr, R.J.; Sur, D.; Yadav, A.; Roopashree, R.; Sangwan, G.; Krithiga, T.; Aminov, Z.; Taher, W.M.; Alwan, M.; et al. Reproducible meningioma grading across multi-center MRI protocols via hybrid radiomic and deep learning features. Neuroradiology 2025, 67, 2741–2761. [Google Scholar] [CrossRef]
  99. Yin, L.; Wang, J. Enhancing brain tumor classification by integrating radiomics and deep learning features: A comprehensive study utilizing ensemble methods on MRI scans. J. X-Ray Sci. Technol. 2025, 33, 47–57. [Google Scholar] [CrossRef]
  100. Yin, S.; Ming, J.; Chen, H.; Sun, Y.; Jiang, C. Integrating deep learning and radiomics for preoperative glioma grading using multi-center MRI data. Sci. Rep. 2025, 15, 36756. [Google Scholar] [CrossRef]
  101. Yonar, A. A swarm intelligence-driven hybrid framework for brain tumor classification with enhanced deep features. Sci. Rep. 2025, 15, 37543. [Google Scholar] [CrossRef]
  102. Senthil, R.; Anand, T.; Somala, C.S.; Saravanan, K.M. Bibliometric analysis of artificial intelligence in healthcare research: Trends and future directions. Future Healthc. J. 2024, 11, 100182. [Google Scholar] [CrossRef]
  103. Liu, X.; Cruz Rivera, S.; Moher, D.; Calvert, M.J.; Denniston, A.K. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Nat. Med. 2020, 26, 1364–1374. [Google Scholar] [CrossRef]
  104. Haibe-Kains, B.; Adam, G.A.; Hosny, A.; Khodakarami, F.; Waldron, L.; Wang, B.; McIntosh, C.; Goldenberg, A.; Kundaje, A.; Greene, C.S.; et al. Transparency and reproducibility in artificial intelligence. Nature 2020, 586, E14–E16. [Google Scholar] [CrossRef]
  105. Kim, D.W.; Jang, H.Y.; Kim, K.W.; Shin, Y.; Park, S.H. Design Characteristics of Studies Reporting the Performance of Artificial Intelligence Algorithms for Diagnostic Analysis of Medical Images: Results from Recently Published Papers. Korean J. Radiol. 2019, 20, 405–410. [Google Scholar] [CrossRef]
  106. Yu, A.C.; Mohajer, B.; Eng, J. External Validation of Deep Learning Algorithms for Radiologic Diagnosis: A Systematic Review. Radiol. Artif. Intell. 2022, 4, e210064. [Google Scholar] [CrossRef] [PubMed]
  107. Spaanderman, D.J.; Marzetti, M.; Wan, X.; Scarsbrook, A.F.; Robinson, P.; Oei, E.H.G.; Visser, J.J.; Hemke, R.; van Langevelde, K.; Hanff, D.F.; et al. AI in radiological imaging of soft-tissue and bone tumours: A systematic review evaluating against CLAIM and FUTURE-AI guidelines. EBioMedicine 2025, 114, 105642. [Google Scholar] [CrossRef]
  108. Xu, H.L.; Gong, T.T.; Song, X.J.; Chen, Q.; Bao, Q.; Yao, W.; Xie, M.M.; Li, C.; Grzegorzek, M.; Shi, Y.; et al. Artificial Intelligence Performance in Image-Based Cancer Identification: Umbrella Review of Systematic Reviews. J. Med. Internet Res. 2025, 27, e53567. [Google Scholar] [CrossRef] [PubMed]
  109. Bramer, W.M.; Rethlefsen, M.L.; Kleijnen, J.; Franco, O.H. Optimal database combinations for literature searches in systematic reviews: A prospective exploratory study. Syst. Rev. 2017, 6, 245. [Google Scholar] [CrossRef]
  110. Wenderott, K.; Krups, J.; Zaruchas, F.; Weigl, M. Effects of artificial intelligence implementation on efficiency in medical imaging—A systematic literature review and meta-analysis. npj Digit. Med. 2024, 7, 265. [Google Scholar] [CrossRef] [PubMed]
  111. Mongan, J.; Moy, L.; Kahn, C.E., Jr. Checklist for Artificial Intelligence in Medical Imaging (CLAIM): A Guide for Authors and Reviewers. Radiol. Artif. Intell. 2020, 2, e200029. [Google Scholar] [CrossRef]
Figure 1. PRISMA flow diagram illustrating the study selection process.
Figure 1. PRISMA flow diagram illustrating the study selection process.
Jcm 15 03441 g001
Figure 2. Geographic distribution of included AI neuroimaging studies (n = 91). The bar chart (left) displays the frequency of studies by country of origin across 15 geographic categories. The donut chart (right) illustrates the proportional country share. China was the predominant contributor (n = 50, 54.9%), followed by the United States (n = 10, 11.0%) and multinational collaborations (n = 8, 8.8%).
Figure 2. Geographic distribution of included AI neuroimaging studies (n = 91). The bar chart (left) displays the frequency of studies by country of origin across 15 geographic categories. The donut chart (right) illustrates the proportional country share. China was the predominant contributor (n = 50, 54.9%), followed by the United States (n = 10, 11.0%) and multinational collaborations (n = 8, 8.8%).
Jcm 15 03441 g002
Figure 3. Distribution of included studies across medical specialties (n = 91, 16 specialties). The bar chart (left) shows the number of studies per specialty. The donut chart (right) illustrates the proportional specialty share. Oncology represented the largest group (n = 34, 37.4%), followed by neurology (n = 19, 20.9%) and radiology (n = 12, 13.2%).
Figure 3. Distribution of included studies across medical specialties (n = 91, 16 specialties). The bar chart (left) shows the number of studies per specialty. The donut chart (right) illustrates the proportional specialty share. Oncology represented the largest group (n = 34, 37.4%), followed by neurology (n = 19, 20.9%) and radiology (n = 12, 13.2%).
Jcm 15 03441 g003
Figure 4. AI methodology types across included studies: deep learning versus machine learning (n = 91). The donut chart (left) shows the overall distribution of AI types: deep learning (DL, 51.6%), machine learning (ML, 37.4%), and hybrid DL/ML (11.0%). The stacked bar charts illustrate AI type distribution by country (center) and by clinical task (right). DL, deep learning; ML, machine learning.
Figure 4. AI methodology types across included studies: deep learning versus machine learning (n = 91). The donut chart (left) shows the overall distribution of AI types: deep learning (DL, 51.6%), machine learning (ML, 37.4%), and hybrid DL/ML (11.0%). The stacked bar charts illustrate AI type distribution by country (center) and by clinical task (right). DL, deep learning; ML, machine learning.
Jcm 15 03441 g004
Figure 5. AI model architecture classification in neuroimaging studies (n = 91). Ground-truth labels were mapped to five subgroups. The donut chart (left) shows the proportional distribution. The lollipop chart (center) displays absolute counts and percentages per category. The legend (right) provides category definitions and clinical examples. NNM, neural network models; REG, regression-based models; KDM, kernel/distance-based models; GBM, gradient boosting machines.
Figure 5. AI model architecture classification in neuroimaging studies (n = 91). Ground-truth labels were mapped to five subgroups. The donut chart (left) shows the proportional distribution. The lollipop chart (center) displays absolute counts and percentages per category. The legend (right) provides category definitions and clinical examples. NNM, neural network models; REG, regression-based models; KDM, kernel/distance-based models; GBM, gradient boosting machines.
Jcm 15 03441 g005
Figure 6. AI task distribution across included neuroimaging studies (n = 91, 9 task categories). The donut chart (left) illustrates the proportional task distribution. The bar chart (right) displays absolute counts and percentages for each task category. Prediction (n = 35, 38.5%) and classification (n = 26, 28.6%) were the most frequently reported tasks, together accounting for approximately 67% of all included studies.
Figure 6. AI task distribution across included neuroimaging studies (n = 91, 9 task categories). The donut chart (left) illustrates the proportional task distribution. The bar chart (right) displays absolute counts and percentages for each task category. Prediction (n = 35, 38.5%) and classification (n = 26, 28.6%) were the most frequently reported tasks, together accounting for approximately 67% of all included studies.
Jcm 15 03441 g006
Figure 7. Validation strategy and human comparator analysis (n = 91). The donut chart (left) illustrates the proportion of studies using external versus internal validation. The bar chart (center) shows the frequency of human comparator inclusion. The heatmap (right) displays cross-tabulation of validation strategy by human comparator use. Only 19% of studies included a direct human comparator, and this proportion remained low even among externally validated studies.
Figure 7. Validation strategy and human comparator analysis (n = 91). The donut chart (left) illustrates the proportion of studies using external versus internal validation. The bar chart (center) shows the frequency of human comparator inclusion. The heatmap (right) displays cross-tabulation of validation strategy by human comparator use. Only 19% of studies included a direct human comparator, and this proportion remained low even among externally validated studies.
Jcm 15 03441 g007
Figure 8. Data leakage risk and methodological transparency analysis (n = 91). The bar charts display data leakage risk classification (left), split description clarity (center-left), and use of external datasets (center-right). The heatmap (right) presents cross-tabulation of leakage risk by external dataset use. The majority of studies were classified as low risk (n = 85, 93%), and all studies provided clear descriptions of data-splitting procedures.
Figure 8. Data leakage risk and methodological transparency analysis (n = 91). The bar charts display data leakage risk classification (left), split description clarity (center-left), and use of external datasets (center-right). The heatmap (right) presents cross-tabulation of leakage risk by external dataset use. The majority of studies were classified as low risk (n = 85, 93%), and all studies provided clear descriptions of data-splitting procedures.
Jcm 15 03441 g008
Table 1. Summary of included studies: author, country of origin, medical field, specific clinical domain, AI type, AI task (n = 91).
Table 1. Summary of included studies: author, country of origin, medical field, specific clinical domain, AI type, AI task (n = 91).
AuthorCountry of OriginMedical FieldSpecific FieldAI TypeAI Task During Research
Akbari H. et al. [59]MultinationalOncologyGlioblastoma prognostic subgroupingMLSurvival prediction and prognostic subgrouping
Albadr R.J. et al. [92]MultinationalOncologyMeningioma gradingHybrid (DL/ML)Preoperative classification/grading
Belke M. et al. [60]GermanyNeurologyEpilepsy imaging/hippocampal sclerosis detectionMLDetection/diagnosis
Cai Z.Y. et al. [75]ChinaNeuroscienceWhite-matter functional connectomics/sex classificationMLClassification
Chen J. et al. [11]ChinaOncologyGlioblastoma molecular marker prediction (MGMT)DLClassification/biomarker prediction
Chen R. et al. [43]ChinaNeurosurgeryIntracranial aneurysm outcome predictionDLPrediction/risk modeling
Chen Y. et al. [93]ChinaPediatricsPediatric brain tumor prognosisHybrid (DL/ML)Prognosis prediction
Chen Y. et al. [12]USANeurologyIntracerebral hemorrhage outcome predictionDLFunctional outcome prediction
Choi J.H. et al. [94]S. KoreaNeurosurgeryIntracranial aneurysm rupture predictionHybrid (DL/ML)Classification/rupture risk prediction
Dai M. et al. [40]MultinationalRadiologyVertebral compression fracture detectionDLDetection
Dai Y. et al. [13]MultinationalPediatricsNeonatal hydrocephalus/CSF diversion predictionDLPrediction
Demirel E. et al. [61]TurkeyOncologyBrain tumor differential diagnosisMLClassification
Dong Y. et al. [14]USANeurology/RadiologyGeneralizable CTA representation learning for acute stroke tasksDLDetection/classification/prediction
Fan Y. et al. [66]ChinaRadiologyPituitary adenoma subtype predictionMLPreoperative classification
Fatania K. et al. [62]UKRadiologyGlioblastoma radiomics survival modelingMLPrognosis modeling
Felefly T. et al. [44]MultinationalOncologyBrain metastasis detection on CTDLDetection/classification
Feng L. et al. [76]ChinaNeurologyEpilepsy surgery outcome predictionMLPrediction
Foltyn-Dumitru M. et al. [63]GermanyNeuroradiologyGlioma imaging phenotyping/survival predictionMLUnsupervised clustering/prognosis
Gui Y. et al. [15]ChinaOncologyMeningioma sinus invasion diagnosisDLPreoperative classification
Hamon G. et al. [16]FranceNeurologySynthetic MRI/DWI-FLAIR mismatch assessmentDLImage synthesis/diagnostic support
Hao M. et al. [58]ChinaOncologyMGMT promoter methylation prediction in glioblastomaMLSurvival prediction/risk stratification
Harper J.P. et al. [39]USARadiologyCervical spine fracture detectionDLDetection
Hossain M.M. et al. [42]BangladeshNeurologyBrain stroke classification on CTDLClassification
Hu W. et al. [74]ChinaRadiologyCarotid plaque symptom classificationMLIdentification
Huang L. et al. [67]ChinaNeurologyMalignant cerebral edema predictionMLPrediction
Jeon E.T. et al. [17]S. KoreaNeurologyWhite matter hyperintensity/Fazekas gradingDLSegmentation/grading
Jia X. et al. [38]ChinaNeuroradiologyMiddle cerebral artery aneurysm rupture risk predictionDLPrediction
Kamel P. et al. [18]USANeurologyIschemic stroke infarct segmentation on MRIDLSegmentation
Kang D.W. et al. [41]S. KoreaRadiologyIntracranial hemorrhage detectionDLDetection
Kesari A. et al. [19]IndiaOncologyBrain tumor blood-vessel segmentationDLSegmentation
Ketabi S. et al. [20]CanadaOncologyPediatric low-grade glioma genetic marker classificationDLClassification
Kong C. et al. [50]ChinaRadiation OncologyGlioblastoma versus solitary brain metastasis differentiationDLClassification
Krag C.H. et al. [21]DenmarkNeurologyAcute ischemic stroke lesion detection on MRIDLClassification
Kulathilake C.D. et al. [22]MultinationalNeurologyBrain stroke CT classificationDLClassification
Li D. et al. [95]ChinaOncologyIDH mutation prediction from MRIHybrid (DL/ML)Prediction
Li Z. et al. [96]ChinaRadiologyPrediction of stroke recurrence in symptomatic intracranial atherosclerotic stenosisHybrid (DL/ML)Prediction
Liang Q. et al. [68]ChinaOncologyAdult diffuse glioma grading/molecular subtypingMLPrediction
Liang X. et al. [23]ChinaOncologyIntracranial solitary fibrous tumor (ISFT) versus angiomatous meningioma differentiationDLClassification
Liao L. et al. [46]FranceNeuroradiologyCerebral aneurysm detection on TOF-MRADLDetection
Lilhore U.K. et al. [24]IndiaOncologyBrain tumor segmentation on multimodal MRIDLSegmentation
Lin X. et al. [25]ChinaRadiologyIntracranial hemorrhage segmentation on CTDLSegmentation
Liu J. et al. [26]USAPediatricsPrediction of normative pediatric brain development from MRIDLPrediction
Liu J. et al. [70]ChinaOncologyMRI-based survival prediction in primary CNS lymphomaMLSurvival prediction
Liu J. et al. [97]ChinaOncologyGlioblastoma prognostic stratificationHybrid (DL/ML)Survival prediction/risk stratification
Lv C. et al. [27]ChinaOncology/RadiologyBrain tumor MRI segmentationDLSegmentation
Ma Z. et al. [77]ChinaOncologyMRI radiomics-based classification of malignant brain tumorsMLClassification
Mahootiha M. et al. [52]USAOncologyPediatric low-grade glioma recurrence predictionDLPrediction/risk modeling
Nada A. et al. [57]USARadiologyIntracranial hemorrhage detectionDLDetection
Nalentzi K. et al. [45]GreeceOncologyBrain tumor MRI classification (glioma versus meningioma)DLClassification
Patel B.K. et al. [72]USAOncologyPrediction of extent of resection in giant pituitary neuroendocrine tumorsMLPrediction
Pelcat A. et al. [28]FranceNeurologyMRI hemorrhage detection in acute strokeDLSynthesis
Petterson S. et al. [47]USANeuroradiologyBrain aneurysm detection on CTADLDetection/screening
Rastogi D. et al. [29]MultinationalOncology/RadiologyBrain tumor segmentation and survival prediction from MRIDLSegmentation/prediction
Roh Y.H. et al. [64]S. KoreaNeurologyHemorrhagic transformation prediction in acute ischemic strokeMLPrediction/risk modeling
Rühling S. et al. [30]GermanyRadiologyOsteoporosis screening/bone mineral density analysisDLDetection
Ryu W.S. et al. [48]S. KoreaNeurologyAcute infarct segmentation on MRIDLSegmentation
Saadh M.J. et al. [98]MultinationalOncologyMeningioma gradingHybrid (DL/ML)Classification/grading
Sina E.M. et al. [55]USAOtolaryngologyPituitary macroadenoma vs. parasellar meningioma MRI differentiationDLClassification
Song D. et al. [71]ChinaOncologyAtypical meningioma recurrence predictionMLPrediction
Sun K. et al. [78]ChinaNeurologyAcute ischemic stroke CT radiomicsMLRadiomics-based detection/classification of MRI-occult ischemic stroke lesions on non-contrast CT
Sun Y. et al. [91]ChinaOncologyBrain metastasis primary tumor origin predictionMLPrediction
Sunavsky A. et al. [79]CanadaNeurologyChronic low back pain classification using fMRI connectivityMLClassification
Topff L. et al. [51]NetherlandsOncologyDetection, segmentation, and longitudinal tracking of brain metastases on MRIDLDetection, segmentation, and longitudinal tracking
Tu J. et al. [31]ChinaOncologyGlioblastoma infiltration detection in peritumoral edemaDLDetection/segmentation
Tuxunjiang P. et al. [54]ChinaNeurologyStroke severity prediction using multimodal MRIDLPrediction/severity estimation
Wang B. et al. [32]ChinaInfectious DiseaseMRI differentiation of Brucella and tuberculosis spondylitisDLClassification/diagnosis
Wang G. et al. [80]ChinaRadiologyCarotid artery stenosis detection on non-contrast CTMLClassification/diagnosis
Wang H. et al. [81]ChinaNeuroradiology/StrokeResponsible aneurysm identification in SAH patients with multiple aneurysmsMLPrediction
Wang H et al. [56]ChinaRadiologyICH black hole sign identification on CTMLPrediction
Wang K. et al. [33]ChinaOrthopedicsPostoperative outcome prediction after tubular microdiscectomy for lumbar disc herniationDLPrediction/outcome classification
Wang T. et al. [34]ChinaRadiology/StrokePost-thrombectomy intracranial hemorrhage CT differentiationDLImage generation/diagnostic classification
Wang Y. et al. [35]ChinaOncologyBrain metastasis segmentationDLSegmentation
Xia X. et al. [82]ChinaOncologyGlioblastoma versus solitary brain metastasis differentiationMLClassification/diagnosis
Xia X. et al. [69]ChinaNeurologyFunctional outcome prediction after ICHMLPrediction/prognosis
Xing L. et al. [36]ChinaOrthopedics/SpineModic changes detection and grading on lumbar spine MRIDLDetection/grading
Xu W. et al. [83]ChinaOncologyGrade 4 glioma molecular subtyping with MRI radiomicsMLPreoperative molecular subtype classification and prognostic stratification
Xu X. et al. [84]ChinaNeurologyPrediction of cerebrovascular disease related cognitive impairmentMLPrediction/risk stratification
Yang H. et al. [49]ChinaNeurologyPrognostic prediction in acute ischemic stroke after thrombolysisDLPrediction/prognosis
Yang Q. et al. [65]ChinaOncologyPituitary neuroendocrine tumor consistency prediction using mpMRI radiomicsMLClassification/prediction
Ye B. et al. [85]ChinaOrthopedics/SpinePrediction of vertebral artery injury during C2 pedicle screw placementMLRisk prediction/classification
Yin L. et al. [99]ChinaOncology/RadiologyPreoperative glioma grading using MRIHybrid (DL/ML)Classification/diagnosis
Yin S. et al. [100]ChinaOncologyPreoperative glioma gradingHybrid (DL/ML)Classification/grading
Yonar A. et al. [101]TurkeyOncologyBrain tumor type classification using MRIHybrid (DL/ML)Classification/diagnosis
Zahoora U. et al. [37]PakistanOncologyBrain tumor segmentation on MRIDLSegmentation
Zeng L. et al. [53]ChinaNeuroradiologyIntracranial aneurysm stability prediction on CTADLClassification/risk prediction
Zeng Q. et al. [86]ChinaOncologyGlioblastoma versus solitary brain metastasis differentiationMLClassification/diagnosis
Zhai D. et al. [87]ChinaNeuroradiology/StrokeHemorrhagic transformation versus contrast extravasation differentiation after mechanical thrombectomyMLClassification/diagnosis
Zhao K. et al. [89]ChinaOncologyDifferential analysis between PCNSL versus low grade gliomaMLClassification
Zhao K. et al. [88]ChinaOncologyPituitary adenoma Ki-67 predictionMLPrediction
Zheng B. et al. [73]ChinaNeurosurgeryCervical spondylotic myelopathy prognosis predictionMLPrediction
Zhuang X. et al. [90]ChinaOrthopedicsSpine fracture imaging analysisMLClassification
Abbreviations: AI, artificial intelligence; ML, machine learning; DL, deep learning; Hybrid (DL/ML), hybrid deep learning and machine learning. Full data extraction, including study aims, main conclusions, detailed methodological variables, and additional abbreviations, is provided in Supplementary Table S1.
Table 2. Human comparator inclusion stratified by AI task category (n = 91).
Table 2. Human comparator inclusion stratified by AI task category (n = 91).
AI Task CategoryTotal (n)With Human Comparator (n)%Clinical Interpretation
Classification2613.8%Critical gap—benchmarking against clinicians essential for diagnostic AI
Prediction35514.3%Substantial gap—benchmarking against clinical scores/experts needed
Detection10660.0%Adequate—most detection systems benchmarked against radiologists
Segmentation800.0%Methodologically appropriate—Dice coefficient vs. expert ground-truth
Multi-task7457.1%Adequate—mostly driven by detection sub-components
Other *5120.0%-
Total911718.7%
* Other includes Diagnosis (n = 2), Identification (n = 1), Grading (n = 1), Synthesis (n = 1).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Sulaimanov, U.; Sanlier, N.; Moniri, A.; Demir, B.; Serikkanov, Y.; Bayramoglu, A.R.; Al-Jebur, M.S.; Uslu, I.; Ozturk, O.; Nizzola, M.; et al. Are AI Neuroimaging Models Ready for Clinical Use? A Systematic Methodological Review. J. Clin. Med. 2026, 15, 3441. https://doi.org/10.3390/jcm15093441

AMA Style

Sulaimanov U, Sanlier N, Moniri A, Demir B, Serikkanov Y, Bayramoglu AR, Al-Jebur MS, Uslu I, Ozturk O, Nizzola M, et al. Are AI Neuroimaging Models Ready for Clinical Use? A Systematic Methodological Review. Journal of Clinical Medicine. 2026; 15(9):3441. https://doi.org/10.3390/jcm15093441

Chicago/Turabian Style

Sulaimanov, Umid, Nafiye Sanlier, Ariorad Moniri, Behman Demir, Yerkebulan Serikkanov, Ahmed Rasim Bayramoglu, Maryam Sabah Al-Jebur, Irem Uslu, Oyku Ozturk, Mariagrazia Nizzola, and et al. 2026. "Are AI Neuroimaging Models Ready for Clinical Use? A Systematic Methodological Review" Journal of Clinical Medicine 15, no. 9: 3441. https://doi.org/10.3390/jcm15093441

APA Style

Sulaimanov, U., Sanlier, N., Moniri, A., Demir, B., Serikkanov, Y., Bayramoglu, A. R., Al-Jebur, M. S., Uslu, I., Ozturk, O., Nizzola, M., Ötleş, E., Ammanuel, S. G., Keles, A., Erginoglu, U., & Baskaya, M. K. (2026). Are AI Neuroimaging Models Ready for Clinical Use? A Systematic Methodological Review. Journal of Clinical Medicine, 15(9), 3441. https://doi.org/10.3390/jcm15093441

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop