Next Article in Journal
Evidence of Validity for the Artificial Intelligence Competence and Literacy Test (CAIA) in Spanish University Students
Previous Article in Journal
A Multi-Criteria Decision Model for Evaluating WPAN Network Security Testing Methods in Educational Institutions
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Methodological Quality and Clinical Translation of Deep Learning in Traditional Chinese Medicine Disease Diagnosis: A Systematic Review and Validation Gap Analysis

1
The First Clinical Medical College, Yunnan University of Chinese Medicine, Kunming 650000, China
2
The Jockey Club School of Public Health and Primary Care, The Chinese University of Hong Kong, Hong Kong SAR 999077, China
3
School of Chinese Medicine, Hong Kong Baptist University, Hong Kong SAR 999077, China
4
Harris School of Public Policy, The University of Chicago, Chicago, IL 60615, USA
5
Centre for Smart Health, School of Nursing, The Hong Kong Polytechnic University, Hong Kong SAR 999077, China
6
School of Nursing, The Hong Kong Polytechnic University, Hong Kong SAR 999077, China
*
Authors to whom correspondence should be addressed.
These authors contributed equally to this work.
Information 2026, 17(6), 554; https://doi.org/10.3390/info17060554
Submission received: 4 April 2026 / Revised: 25 May 2026 / Accepted: 26 May 2026 / Published: 3 June 2026
(This article belongs to the Section Biomedical Information and Health)

Abstract

Traditional Chinese Medicine (TCM) diagnosis is characterized by substantial subjectivity and inter-practitioner variability. Although deep learning (DL) offers potential for standardization and intelligent diagnostic support, its methodological rigor and clinical applicability remain unclear. This systematic review evaluated the applications, methodological quality, and translational readiness of DL in TCM disease diagnosis. Eight databases were searched from January 2010 to May 2026, and 21 studies were included. Quality assessment was conducted using TRIPOD-AI and QUADAS-2. The review identified a technological evolution from conventional classification models toward knowledge graph-enhanced frameworks, large language models, and multimodal integration. However, major methodological limitations were observed. The mean TRIPOD-AI reporting rate was only 50.5%, with deficiencies in transparency and validation reporting. QUADAS-2 assessment showed high risk of bias in the Index Test and Reference Standard domains in 47.6% of studies, and none performed independent external validation. In conclusion, despite rapid technological advances, DL-based TCM diagnostic research remains limited by insufficient validation, poor reporting transparency, and reference standard bias. Future research should prioritize multi-center external validation, prospective real-world testing, standardized reference standards, and theory-driven explainable AI approaches.

Graphical Abstract

1. Introduction

1.1. The Diagnostic Framework and Modernization Challenges of TCM

Traditional Chinese Medicine (TCM) diagnostics, a comprehensive medical system with a history spanning thousands of years, is centered around the Four Diagnostic Methods: observation, listening and smelling, inquiry, and palpation [1]. These methods systematically gather clinical symptoms and signs from patients, followed by an overall syndrome differentiation based on TCM theory to determine the appropriate diagnosis and guide treatment [2]. In some clinical trials of TCM treatments, such syndrome differentiation has been included to better resemble clinical practice and optimize treatment effectiveness [3,4,5]. However, the traditional diagnostic process heavily relies on the subjective experience and individualized judgment of practitioners [6,7]. It inevitably leads to inherent limitations such as non-uniform diagnostic standards, low reproducibility of results, and difficulties in passing down experience [8,9,10]. In the context of modern evidence-based medicine, advancing the objectivity and standardization of TCM diagnostics has become a key scientific issue, as reliance on subjective methods hinders its clinical credibility and widespread application [11,12,13].

1.2. The Technological Opportunity: Deep Learning in Medicine

Artificial Intelligence (AI) technologies, represented by Deep Learning (DL), are sparking a global technological revolution [14]. DL builds neural networks with multiple hidden layers, or deep structures, which can automatically learn from vast amounts of raw data and progressively extract abstract feature representations from low-level to high-level [15,16]. In recent years, the application DL in the medical field has rapidly expanded from its initial use in medical image-assisted diagnosis to include electronic medical record mining [17], genomics analysis [18], physiological signal interpretation [19], and even drug discovery [20]. This provides a robust technological toolkit and methodological insights to address the challenge of objectifying TCM diagnostics, transforming the practitioner’s sensory observations, such as inspection, auscultation, and palpation, and subjective inquiries, such as questioning, into quantifiable and computable data, which can significantly enhance diagnostic accuracy and reproducibility. From a bioengineering perspective, TCM diagnosis presents unique challenges for automated systems: the need to transform subjective, multi-modal observations (inspection, auscultation, palpation, inquiry) into quantifiable engineering signals, while maintaining fidelity to complex theoretical frameworks.

1.3. Evolution of Intelligent TCM Diagnostic Research

The application of DL in TCM diagnosis has evolved through three distinct developmental stages, each representing a significant advancement in technical capability and clinical utility. Early research focused on automating individual diagnostic methods. Convolutional neural networks (CNNs) were applied to classify and segment tongue images [21,22], while 1-D CNNs analyzed pulse waveforms collected through pressure sensors [23]. These initial efforts established the feasibility of using DL for objective feature extraction in TCM diagnosis. The field then progressed toward multi-diagnostic integration. Representative works include TongueNet [24], which combined tongue image analysis with clinical text data for multi-label disease classification, and constitution recognition models [25] that integrated traditional and DL-derived features. This integration phase demonstrated the superiority of multi-modal approaches over single-feature analysis in diagnostic accuracy. Recent developments have shifted toward knowledge-enhanced cognitive assistance. This advancement is characterized by two key innovations: (1) the integration of TCM domain knowledge through knowledge graphs, as exemplified by Xiao et al.’s [26] work on diabetic retinopathy, which enabled structured reasoning about syndrome-symptom-treatment relationships; and (2) the application of large language models (LLMs) for intelligent consultation, demonstrated by systems like OpenTCM [27] that combine GraphRAG models with LLMs to facilitate natural language understanding in TCM diagnosis. These developments mark a crucial transition from perceptual automation to cognitive assistance in TCM diagnostic practice.

1.4. Persistent Challenges and Knowledge Gaps

Despite this progress, significant challenges impede the translation of research into clinical evidence. There is a lack of systematic evaluations specifically focusing on the application of DL in TCM disease diagnosis, leading to a gap in the comprehensive assessment of the overall progress, technical pathways, and evidence quality in this field. For example, Lim et al. [28] noted that research on machine learning in TCM, many studies still focus on specific technologies or diseases, lacking cross-modal integration. Wang et al. [29] reviewed the application of DL in tongue diagnosis, their research focuses solely on a single diagnostic method without integrating cross-disease analysis. Regarding methodological quality assessment and clinical translation potential, this field exhibits notable gaps in evaluation tools and risks within the validation framework. One prominent issue is the current lack of systematic assessment using specialized tools for diagnostic accuracy studies, such as QUADAS-2, to evaluate the methodological bias and clinical applicability of primary research on DL in TCM diagnosis. The other involves the rigorous assessment of reporting completeness, which is largely absent despite the availability of established guidelines such as the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis-Artificial Intelligence (TRIPOD-AI) statement. These two interconnected gaps together hinder a credible evaluation of progress in this field.
To address these gaps, we conducted this systematic review to examine the application of DL in TCM disease diagnosis, with specific objectives: (1) To analyze the current landscape of DL applications in TCM diagnosis, focusing on diagnostic tasks, data modalities, and model architectures; (2) To assess the methodological quality and reporting completeness of included studies using the TRIPOD-AI checklist; (3) To evaluate risk of bias and validation rigor using QUADAS-2 framework, particularly addressing external validation issues; (4) To identify gaps between AI capabilities and TCM diagnostic principles, and propose future research directions for advancing evidence-based intelligent TCM diagnostics.

2. Methods

This is a systematic review of DL in TCM diagnosis. Reporting of this study follows the PRISMA guideline [30] and the PRISMA-AI reporting guideline [31]. The study has been registered in PROSPERO (CRD42024569225).

2.1. Eligibility Criteria

The studies included in this review are original and fully accessible, focusing on the use of DL in TCM diagnosis. The inclusion criteria were: (1) provision of detailed information on datasets and data processing methodologies; (2) clear description of methods, including DL model architecture; (3) application of DL models for TCM diagnostic tasks; and (4) evaluation of model performance using specific metrics with precise results. Studies were excluded if they met any of the following criteria: (1) insufficient methodological details; (2) unclear model architecture description; (3) lack of clinical diagnostic application; (4) absence of disease-specific focus; or (5) focus limited to isolated feature recognition without diagnostic context. This review specifically targeted studies on TCM disease diagnosis for defined diseases to ensure clinical relevance and comparability of findings.

2.2. Information Sources and Search

An electronic search was conducted in the following electronic databases: Medline (via PubMed), Scopus, Embase, Web of Science, CNKI, Wanfang, CQVIP, and SinoMed. Search results were limited to publications from 1 January 2010, to 1 May 2026, accounting for the emergence of modern DL applications [32]. There was no limitation in the language. The keywords “deep learning” was combined with the “diagnose” term. The specific query searched was: (“deep learning” OR “deep neural network” OR “convolution neural network” OR “recurrent neural network” OR “deep belief network” OR “computer vision” OR “natural language processing” OR “large language model” OR “transformer”) AND (diagnosis OR inspection OR listening OR smelling OR inquiry OR palpation OR auscultation-olfaction OR pulse OR tongue OR auscultation-olfaction OR syndrome OR pattern OR symptom OR sign OR different OR classification OR identification) AND (“Chinese medicine” OR TCM OR CM). The search terms were expanded to include emerging AI technologies such as “natural language processing,” “large language model,” and “Transformer”, considering their increasing application in TCM clinical practice for medical record interpretation, syndrome reasoning, and knowledge extraction. After the initial search, duplicates were removed, and bibliographies of included papers were manually cross-referenced. Detailed search strategies for each database are provided in Section S1 of the Supplementary Materials. Only peer-reviewed journal articles were included.

2.3. Study Selection

Citation management was conducted using Endnote X9 (Clarivate Analytics, Philadelphia, PA, USA). Two independent reviewers (LJL and HL) screened the titles and abstracts after duplicate records were removed. Subsequently, full texts of potentially eligible studies were assessed according to the inclusion and exclusion criteria. Any disagreements were resolved through consensus involving a third reviewer (SCC).

2.4. Data Collection and Extraction

Two reviewers (HL and YML) independently extracted data from the included studies, with a third reviewer (SCC) resolving any discrepancies. Data extraction was conducted in two complementary parts to systematically capture both foundational study characteristics and technical modeling details.
The first part focused on fundamental study attributes and design elements: (1) basic publication information including first author, publication year, and language; (2) disease focus and corresponding TCM patterns investigated; (3) diagnostic criteria used for both Western medicine diseases and TCM syndromes; (4) participant characteristics including age and gender distributions; (5) data sources with specification of single-center or multi-center design; (6) dataset specifications including size, type, and train-test splits; (7) modality types such as tongue images, facial images, text data, or clinical measurements; and (8) labeling sources detailing how ground truth labels were established.
The second part concentrated on modeling methodology and performance evaluation, encompassing: (1) tasks: specific application scenarios and objectives; (2) models: architectural details including DL layers, configurations, and ensemble techniques, with only the best-performing model reported when multiple were employed; (3) process: strategies for model development, training, and validation; (4) performance metrics: key evaluation indicators including accuracy, F1-score, and AUC; (5) external validation: whether testing was conducted on independent external datasets; (6) control type: baseline methods or comparison groups used; (7) main findings: principal conclusions derived from the analysis; and (8) limitations: identified study constraints. Studies were indexed by first author’s surname and publication year.

2.5. Quality Assessment

Two authors (LJL and HL) independently assessed compliance with reporting guidelines using the TRIPOD-AI checklist [33] for DL models. TRIPOD-AI provides standardized guidance for transparent reporting of prediction model development and validation, comprising 27 sections with a total of 52 items. Each item was rated as “Y” (reported), “N” (not reported), or “NA” (not applicable).
Additionally, the same two authors independently reviewed the risk of bias in the selected studies using the QUADAS-2 tool, which is most used quality assessment tool for systematic reviews of diagnostic accuracy studies [34] and encouraged by current PRISMA 2020 guidance. The assessment covered four domains (patient selection, index test, reference standard, and flow/timing), evaluating both risk of bias and applicability concerns. Each domain was rated as “high,” “uncertain,” or “low” risk.
Given that QUADAS-2 was originally developed for conventional diagnostic accuracy studies, predefined adaptations were applied to the “Index Test” and “Reference Standard” domains to improve its applicability and consistency in the evaluation of AI-based diagnostic models.
For the Index Test domain, assessments focused on: (1) data separation, including whether test datasets were appropriately isolated prior to feature selection and hyperparameter tuning; (2) validation strategy, including the use of independent test datasets or cross-validation; and (3) reproducibility, including the reporting of model architecture, training procedures, and preprocessing methods. Studies with potential data leakage or unclear validation procedures were considered at increased risk of bias.
For the Reference Standard domain, assessments focused on: (1) validity of the reference standard, including whether diagnoses were based on recognized TCM diagnostic criteria such as national guidelines or expert consensus; (2) blinding, including whether reference standard assessments were conducted without knowledge of index test results; and (3) annotation transparency, including reporting of annotator qualifications and inter-rater agreement measures (e.g., kappa statistics). Studies relying on expert annotations without clear diagnostic criteria, inter-rater agreement, or annotation transparency were considered at increased risk of bias.
Disagreements between the reviewers were resolved through discussion, and if they could not reach an agreement, a third author (SCC) was consulted to make the final determination.

3. Results

3.1. Literature Search

Our systematic search across eight electronic databases yielded 1743 initial records. Following duplicate removal (n = 752), 991 unique publications underwent title and abstract screening. After excluding 872 records that failed to meet inclusion criteria, 119 articles qualified for full-text assessment. Further exclusions (n = 98) were made based on: absence of DL applications (n = 32), non-TCM disease diagnosis focus (n = 44), lack of specific disease applications (n = 17), missing performance metrics (n = 3), and inaccessible full texts (n = 2). The final review included 21 studies. The study selection process is illustrated in the PRISMA flow diagram (Figure 1).

3.2. Description of Included Studies

As shown in Table 1, the review analyzed 21 studies published between 2019 and 2026. The majority of studies were published in Chinese (n = 15, 71.4%). Regarding disease categories, cardiovascular and cerebrovascular diseases were the most common (n = 6, 28.6%), followed by digestive system diseases (n = 5, 23.8%) and endocrine and metabolic diseases (n = 4, 19.0%). Most studies reported the TCM patterns under investigation (n = 20, 95.2%). For disease diagnostic criteria, guidelines or consensus were most frequently used (n = 10, 47.6%). For TCM pattern diagnostic criteria, industry or national standards were the most common (n = 9, 42.9%).
Regarding data and methodological characteristics, single-center studies predominated (n = 12, 57.1%), and none of the included studies performed external validation (n = 21, 100%). Text-only data were the most commonly used modality (n = 12, 57.1%), while expert or clinician annotation served as the primary label source (n = 12, 57.1%). Classification tasks were the most common task category (n = 13, 61.9%), and accuracy-related metrics were reported in nearly all studies (n = 20, 95.2%; as detailed in Section S2 of the Supplementary Materials for the formulas of the outcome measures). Comparisons with other computational models were the most common evaluation approach (n = 12, 57.1%). Detailed methodology and model specifications are presented in Table 2 and Sections S3 and S4 of the Supplementary Materials.

3.3. Reporting Quality and Risk of Bias

3.3.1. Reporting Quality Assessed Using the TRIPOD-AI Checklist

The mean reporting rate across all items was 46.8%, and the mean reporting rate across the 21 included studies was 50.5% (range: 39.6% to 62.5%), as presented in Table 3 and Section S5 of the Supplementary Materials.
Eighteen items (34.6%) were reported in all 21 studies: Title (1), Abstract (2), Background (3a, 3b), Objectives (4), Data sources (5a), Participant eligibility criteria (6a, 6b), Data preparation (7), Outcome definition (8a), Predictor definition (9a, 9b), Analytical methods (12a–c, 12e), Model development (21), and Limitations (26).
The following 15 items (28.8%) were not reported in any study: one item in the Section 3 (health inequalities, 3c); five items in the Section 4 (cross-cluster heterogeneity, 12d; model prediction calculation, 12g; fairness, 14; model output, 15; training vs. evaluation comparison, 16); three items in Section 5 (study protocol, 18c; registration, 18d; code sharing, 18f); one item in Section 6 (patient involvement, 19); three items in the Section 7 (participant characteristics by data source, 20b; predictor distribution comparison, 20c; model specification, 22); and two items in the Section 8 (poor quality input data handling, 27a; user interaction requirements, 27b). Overall, reporting deficiencies were mainly concentrated in model transparency, validation reporting, open science practices, and clinical usability-related domains.

3.3.2. Risk of Bias and Applicability Concerns Assessed Using QUADAS-2

Risk of bias and applicability concerns of all 21 included studies were evaluated using QUADAS-2. The results are summarized in Figure 2 and Section S6 of the Supplementary Materials. In the Patient Selection domain, 15 studies (71.4%) were rated as low risk, 2 (9.5%) as high risk, and 4 (19.0%) as unclear risk. For the Index Test domain, 11 studies (52.4%) were rated as low risk and 10 (47.6%) as high risk. For the Reference Standard domain, 11 studies (52.4%) were rated as low risk and 10 (47.6%) as high risk. For the Flow and Timing domain, all 21 studies (100%) were rated as low risk. Applicability concerns were consistently low across all domains, indicating strong alignment with the review objectives. In summary, the risk of bias in the included studies was predominantly concentrated in the domains of the index test and the reference standard.

3.4. Deep Learning Development Trends in TCM Diagnosis

3.4.1. Evolution of Diagnostic Data Modalities

Among the 21 included studies, single-modal data predominated (n = 16, 76.2%). Text-based modalities were the most common (n = 12, 57.1%) [35,36,38,39,43,46,48,49,50,51,52,55], primarily comprising symptom descriptions, physical sign records, and electronic medical record texts for symptom analysis, medical record modeling, and TCM syndrome classification tasks. Image-based modalities were used in four studies (19.0%), including tongue images [36,53] and facial images [44,47], mainly for disease recognition and TCM syndrome classification.
Multimodal data were applied in five studies (23.8%). Among these, one study integrated text, numerical laboratory indicators, and tongue/facial images [37]; two studies combined tongue images with structured symptom text [40,45]; and two studies integrated pulse signals, sublingual collateral vessel images, speech features, and clinical questionnaire texts across four modalities [41,42].
Regarding temporal distribution, all five multimodal studies were published in or after 2023, whereas the five studies published between 2019 and 2022 exclusively employed single-modal data.

3.4.2. Evolution of Model Architectures

The model architectures adopted in the included studies demonstrated a clear technological evolution, progressing from conventional deep learning models toward knowledge-enhanced architectures, pretrained language models (PLMs), and LLMs.
Conventional DL Architectures
Early studies primarily employed conventional DL models for TCM syndrome classification and disease diagnosis based on structured symptoms, electronic medical record texts, or single-image modalities. Common architectures included deep neural networks (DNNs), artificial neural networks (ANNs), deep belief networks (DBNs), CNNs, and recurrent neural networks (RNNs/GRUs) [35,38,48,51]. These studies mainly relied on single-modal data and structured feature inputs. A modified Transformer architecture was subsequently introduced for coronary artery disease syndrome element diagnosis [49]. In addition, a stacking ensemble learning strategy integrating decision trees, support vector machines, random forests, and XGBoost with a backpropagation neural network as the meta-learner was applied to liver cirrhosis syndrome classification [39].
CNN and Visual Diagnostic Models
With the development of computer vision techniques, CNN-based architectures and their variants gradually became the dominant frameworks for tongue and facial image analysis. A fine-tuned DenseNet201 combined with Cubic Support Vector Machine (Cubic SVM) was used for tongue image classification in stroke [36]. Multibranch ResNet-18 was applied to facial image analysis in chronic renal failure [47]. A HybridModel integrating EfficientNet-B3 and Swin-Tiny with cross-modal multi-head attention was developed for tongue segmentation and TCM syndrome classification [53]. In addition, EfficientNet, MobileNet V3, and ResNet18 were compared for facial image-based depression diagnosis, with EfficientNet achieving the highest accuracy (98.6%) [44].
Knowledge-Enhanced Architectures and Pretrained Language Models
In recent years, knowledge graph-enhanced methods and PLMs have increasingly been applied to TCM text-based diagnostic tasks. The integration of knowledge graphs and recurrent neural networks was initially explored through the development of a Knowledge-Based RNN (KBRNN) for cerebral palsy diagnosis [50]. Subsequently, models integrating knowledge graphs, attention mechanisms, and ComplEx embeddings were applied to syndrome classification [46]. A Knowledge Graph Pre-trained Language Model (KG-PLM) framework was further proposed, integrating relational graph convolutional networks (RGCN) or graph attention networks (GAT) with PLMs, including BERT, Longformer, and ERNIE-Health, for multi-label TCM syndrome classification in diabetic kidney disease. This framework concatenated knowledge graph subgraph embeddings with text embeddings, with the GAT-BERT configuration achieving the best performance (Macro-F1 = 0.821) [43].
Large Language Models and Diagnostic Reasoning Frameworks
LLMs have recently been introduced for intelligent TCM diagnosis, syndrome reasoning, and clinical decision support. AcupunctureGPT, fine-tuned from GPT-4-0613, was developed for acupuncture diagnosis and treatment generation by incorporating semantic similarity evaluation models (SSEM) and knowledge-driven key feature prompting (GKFP) [54]. Continual pretraining and two-stage LoRA fine-tuning were applied to Qwen2.5-7B, combined with chain-of-thought reasoning, for diarrhea diagnosis, syndrome differentiation, and prescription generation [52]. DeepSeek-r1:32b with a three-stage chain-of-thought prompting framework (3STCoT) was further used for coronary artery disease syndrome element identification and syndrome classification through zero-shot inference [55]. Compared with traditional classification models, these studies extended DL applications from pattern recognition to diagnostic reasoning and knowledge generation.
Multimodal Fusion Frameworks
Multimodal fusion gradually emerged as an important direction in intelligent TCM diagnosis. A multimodal fusion model (TS-Model) integrating U2-Net, ResNet34, and fully connected networks was developed to combine tongue images and structured symptom texts [40]. A CNN-RNN hybrid model was adopted to integrate tongue images and textual data for spleen deficiency syndrome classification in gastrointestinal tumors [45]. Pulse signals, sublingual collateral vessel images, speech features, and clinical questionnaire texts were integrated for blood stasis syndrome classification in coronary artery disease and type 2 diabetes mellitus [41,42]. In addition, TCM symptom features, laboratory indicators, tongue images, and facial images were fused for heat syndrome prediction in acute ischemic stroke [37]. All multimodal studies adopted feature-level concatenation as the primary fusion strategy.

4. Discussion

This study systematically summarized the current applications of DL in TCM disease diagnosis and, for the first time, combined the TRIPOD-AI and QUADAS-2 tools to evaluate methodological quality and risk of bias in this field. Overall, the included studies demonstrated clear technological evolution, including a transition from single-modal to multimodal approaches and from conventional DL models toward knowledge-enhanced architectures and large language models. However, this review also identified substantial methodological and validation limitations in the current evidence base, particularly the absence of external validation, insufficient transparency of reference standards, and low adherence to TRIPOD-AI reporting recommendations. These limitations fundamentally compromise the strength of evidence and impede reliable clinical translation. The following sections systematically examine these findings and their implications for advancing the field.

4.1. Technological Evolution and Paradigm Shift

This review demonstrated that the development of DL in TCM diagnosis reflects an important paradigm shift from basic perceptual processing toward advanced cognitive enhancement.
In terms of model architectures, research has evolved from conventional classification models, such as DNNs and CNNs [48,51], toward knowledge graph-enhanced frameworks [43,49], and more recently to large language models for syndrome reasoning and clinical decision support [52,55]. This evolution was driven by the growing recognition that the core of TCM syndrome differentiation lies in knowledge reasoning rather than simple feature recognition, while advances in general artificial intelligence technologies have also been gradually introduced into TCM diagnostic research. This transition indicates that the research focus has shifted from pattern recognition toward knowledge-driven diagnostic reasoning.
Regarding data modalities, research has evolved from predominantly single-modal text-based approaches toward multimodal fusion integrating tongue images, pulse signals, speech features, and symptom descriptions [37,41]. This trend represents a technological response to the holistic characteristics of TCM “four diagnostic methods.” Previous research has suggested that the integration of the four diagnostic methods fundamentally involves multimodal cognitive integration, and that Transformer architectures based on self-attention mechanisms may better model this reasoning process [56]. However, current multimodal studies mainly rely on feature-level concatenation strategies [40,45], and theory-driven interactive fusion approaches remain lacking.
Overall, TCM artificial intelligence research is gradually moving toward an open-domain cognitive interaction stage, in which AI systems are evolving from “assisted recognition” toward “assisted reasoning,” potentially providing greater clinical decision-support value for primary care and junior practitioners.

4.2. Methodological Limitations and Barriers to Clinical Translation

Despite the rapid development of DL technologies in TCM diagnosis, methodological rigor and clinical validation frameworks remain substantially underdeveloped, representing a major barrier to clinical translation. QUADAS-2 assessment indicated that the primary sources of bias were concentrated in the Index Test and Reference Standard domains, while TRIPOD-AI evaluation further revealed widespread deficiencies in reporting transparency and reproducibility.
First, external validation was entirely absent. None of the 21 included studies conducted independent external validation, and all reported high-performance metrics were derived solely from internal validation methods, such as cross-validation or hold-out test sets. Consequently, reported model performance primarily reflects the ability to fit specific data distributions rather than true generalizability across different clinical settings, patient populations, or data acquisition conditions. Reliance on validation within a single dataset is insufficient to demonstrate model robustness, and this gap between internal validity and external generalizability has become a key obstacle to clinical translation [57,58].
Second, reporting transparency was inadequate. TRIPOD-AI assessment showed a mean reporting rate of only 50.5%, with key items including model specification, comparison between training and evaluation datasets, code sharing, and study registration entirely unreported. These deficiencies directly compromise reproducibility, making it difficult for readers to determine whether data leakage occurred during model development and preventing independent verification or replication of reported findings [33].
In addition, QUADAS-2 assessment suggested substantial limitations in both validation procedures and reference standard construction. In the Index Test domain, 10 studies (47.6%) were rated as high risk of bias, mainly due to the lack of blinding, independent validation cohorts, and rigorous data separation procedures. These deficiencies increase the risks of data leakage and overfitting, thereby inflating model performance estimates. In TCM diagnostic studies characterized by high-dimensional features and relatively small sample sizes, such optimistic bias may be further amplified [59]. Meanwhile, 10 studies (47.6%) were also rated as high risk in the Reference Standard domain. None of the 12 studies using expert annotations as reference standards reported inter-rater reliability metrics. Because TCM syndrome differentiation relies heavily on expert subjective consensus, and model input features and reference labels were often derived from the same expert judgments, models may learn specific diagnostic preferences rather than stable and generalizable disease characteristics. Previous studies have shown that even in relatively standardized fields such as pathological diagnosis, inter-observer variability can substantially affect AI performance evaluation results [60]. This fundamentally challenges the validity of the “gold standard” in current TCM AI research.
Notably, most included studies were retrospective single-center studies with limited data sources and population diversity, increasing the likelihood of overfitting to specific data environments [61]. Moreover, existing studies primarily emphasized performance metrics, while relatively limited attention was given to clinical interpretability, real-world applicability, and deployment feasibility. These findings suggest that the field remains largely at the stage of “technical feasibility validation” and has not yet progressed to rigorous clinical validation.

4.3. The Gap Between Theoretical Framework and Technical Implementation

Current DL approaches to TCM diagnosis exhibit three fundamental limitations in bridging theoretical principles with technical implementation. First, there remains an inherent mismatch between statistical modeling paradigms and the dynamic nature of TCM syndromes [62]. Among the reviewed studies, only one study attempted to analyze the temporal evolution of syndrome elements during the progression of coronary artery disease [49], while all other studies treated syndromes as fixed labels. Second, current information fusion approaches fail to embody the holistic diagnostic principles of TCM. Although recent studies have begun integrating multimodal information such as tongue images, pulse signals, speech features, and symptom descriptions, all included multimodal studies primarily adopted feature-level concatenation strategies [40,41,42,45]. This “input-feature-output” fusion paradigm [63] fails to reflect the essence of TCM holistic diagnosis, which emphasizes theory-driven cross-validation and the organic integration of multiple diagnostic methods [64,65,66]. Third, a substantial disconnect persists between model interpretability and clinical reasoning requirements. Some studies attempted to improve interpretability through Class Activation Mapping (CAM) or graph attention visualization methods [43,44]; however, these approaches mainly highlight regions of model attention and remain insufficient to establish clear semantic links with TCM pathogenesis theories.

4.4. Clinical Implications and Future Directions

4.4.1. Clinical Implications

Despite the notable methodological limitations of current studies, DL still demonstrates certain clinical application potential in TCM diagnosis. Existing models can be used for syndrome classification, information integration, and preliminary pattern differentiation support, particularly in primary care, telemedicine, and settings with limited TCM resources [67,68,69]. In primary care settings, artificial intelligence-based clinical decision support systems have been preliminarily applied to diagnostic assistance, treatment recommendations, and complication prediction, showing potential for improving clinical management and healthcare accessibility [67]. Furthermore, LLMs pre-trained and fine-tuned on large-scale TCM textual knowledge and question-answering datasets have demonstrated promising capabilities in knowledge retrieval, medical case diagnosis, and herbal or formula recommendation tasks [70], suggesting their potential value in supporting TCM clinical education and decision-making.

4.4.2. Future Directions

Future research should prioritize the establishment of more rigorous clinical validation frameworks, including multi-center dataset construction, prospective study designs, and independent external validation, in order to improve model generalizability and real-world applicability. Particularly in TCM diagnostic research, adequate validation should extend beyond internal data splitting and include evaluations across different institutions, geographic regions, patient populations, and clinical environments [71]. In addition, real-world clinical testing is needed to assess model stability, clinical acceptability, and deployment feasibility in practical healthcare settings [72].
Future studies should also strengthen methodological transparency and reproducibility through protocol pre-registration, explicit reporting of data partitioning strategies, and adherence to established AI reporting guidelines such as TRIPOD-AI [33] and PROBAST-AI [73]. Pre-registration represents an important measure to improve research transparency, prevent selective reporting, and ensuring reproducibility [74,75]. These measures may help reduce methodological bias and improve consistency across future studies.
In addition, further efforts are required to standardize TCM reference standards, including the development of unified annotation protocols, reporting of inter-rater reliability, and improved transparency of expert labeling procedures [76,77].
From a technical perspective, future TCM AI research may need to move beyond purely data-driven approaches toward theory-driven AI frameworks. This includes integrating TCM theoretical systems, syndrome differentiation logic, and interpretable reasoning mechanisms into model design [78]. In parallel, the development of explainable AI approaches may help improve the alignment between model outputs and TCM clinical reasoning processes, thereby enhancing clinical interpretability and physician trust [68].

4.5. Advantages and Limitations of the Study

This systematic review offers four key contributions: First, a comprehensive analysis of DL applications for TCM disease diagnosis, spanning research across diseases, data types, and methodological approaches. Second, it represents the first systematic application of the TRIPOD-AI checklist to assess reporting completeness in this field, identifying critical deficiencies in model transparency, validation reporting, open science practices, and clinical usability domains. Third, application of the QUADAS-2 framework for assessing diagnostic accuracy enabling systematic evaluation of methodological quality and identification of key bias and clinical applicability issues. Fourth, implementation of a standardized data extraction methodology, systematically integrating clinical and technical parameters to establish a multidimensional foundation for assessing the field. Fifth, identification of key trends and bottlenecks, particularly the evolution from unimodal to multimodal approaches, challenges of in capturing dynamic and holistic of TCM principles, and lack of external validation, providing clear direction to future development.
Several limitations warrant consideration: First, ensuring a focused evaluation of model performance and validation in well-defined clinical contexts, necessitated restricting inclusion to disease-specific studies. While this excluded broader syndrome classification studies, it ensured greater comparability in terms of clinical objectives and evaluation criteria. Second, despite comprehensive searches across major Chinese and English databases, unpublished or recent studies may have been missed, introducing potential publication bias. Third, substantial heterogeneity in diseases, data types, models, and evaluation metrics prevented quantitative meta-analysis; therefore, limiting synthesis to qualitative analysis. Finally, QUADAS-2 quality assessment was dependence on reporting completeness, particularly regarding patient selection, blinding, and reference standards documentation.

5. Conclusions

This systematic review represents the first combined application of the TRIPOD-AI and QUADAS-2 frameworks to evaluate deep learning in TCM disease diagnosis. From a technological perspective, DL in TCM diagnosis is evolving from perceptual automation toward cognitive enhancement, progressing from conventional classification models to knowledge graph-enhanced frameworks and large language models, while data modalities are expanding from unimodal to multimodal integration. However, substantial methodological limitations remain. None of the 21 included studies performed external validation, the mean TRIPOD-AI reporting rate was only 50.5%, and QUADAS-2 assessments identified high risks of bias predominantly in the Index Test and Reference Standard domains. In particular, the systematic absence of external validation substantially limits the generalizability and real-world applicability of current models. Collectively, these deficiencies undermine the reliability, reproducibility, and clinical translation of existing evidence. Future research should prioritize rigorous validation frameworks incorporating multi-center datasets, prospective study designs, independent external validation, and real-world clinical testing. Improvements in reporting transparency, standardization of TCM reference standards, and the development of theory-driven and explainable AI approaches will be essential for establishing a more robust evidence base for the modernization of TCM diagnosis.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/info17060554/s1, Section S1: Search Strategies; Section S2: Formulas for the outcome measures in the included studies; Section S3: Summary of detailed basic characteristics of included studies; Section S4: Summary of other model and methodological characteristics for included studies: comparison types, key findings, and limitations; Section S5: Assessment of reporting quality using the TRIPOD-AI checklist for included studies; Section S6: Risk of bias and applicability concerns assessed using QUADAS-2 for the included studies.

Author Contributions

Conceptualization, J.Q. and S.-C.C.; methodology, J.Q., W.-F.Y., C.C.Z., L.X. and S.-C.C.; software, H.-X.D.; validation, S.-C.C., H.L. and Y.-M.L.; formal analysis, H.L., Y.-M.L. and H.-X.D.; investigation, L.-C.L. and H.L.; data curation, L.-C.L. and H.L.; writing—original draft preparation, L.-C.L. and H.L.; writing—review and editing, J.Q., W.-F.Y., C.C.Z., L.X. and S.-C.C.; visualization, Y.-M.L.; supervision, L.X. and S.-C.C.; project administration, L.-C.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No datasets were generated or analysed during the current study.

Acknowledgments

We would like to express our sincere appreciation to all the researchers who contributed to this study for their invaluable support and efforts.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

3STCoTThree-Stage Chain-of-Thought
AIArtificial Intelligence
ANNArtificial Neural Network
AUCArea Under the Curve
BERTBidirectional Encoder Representations from Transformers
BiLSTMBidirectional Long Short-Term Memory
BLEUBilingual Evaluation Understudy
BP-NNBackpropagation Neural Network
CAMClass Activation Mapping
CHAIDChi-squared Automatic Interaction Detector
CHDCoronary Heart Disease
CHD-SEDDCoronary Heart Disease Syndrome Element Diagnostic Device
CNNConvolutional Neural Network
CoTChain of Thought
CRFChronic Renal Failure
Cubic SVMCubic Support Vector Machine
DBNDeep Belief Network
DenseNetDensely Connected Convolutional Network
DKDDiabetic Kidney Disease
DLDeep Learning
DNNDeep Neural Network
EMRElectronic Medical Record
ERNIEEnhanced Representation through kNowledge IntEgration
FCNFully Connected Network
FFTFast Fourier Transform
GATGraph Attention Network
GKFPKnowledge-Driven Key Feature Prompting
GPTGenerative Pre-trained Transformer
GraphRAGGraph Retrieval-Augmented Generation
GRUGated Recurrent Unit
KBRNNKnowledge-Based Recurrent Neural Network
KG-PLMKnowledge Graph Pre-trained Language Model
LassoLeast Absolute Shrinkage and Selection Operator
LightGBMLight Gradient Boosting Machine
LLMsLarge Language Models
LoRALow-Rank Adaptation
LSLeast Squares
LSTMLong Short-Term Memory
MFCCMel-Frequency Cepstral Coefficients
MLPMultilayer Perceptron
MLRMultiple Linear Regression
NANot Applicable
NLPNatural Language Processing
PCAPrincipal Component Analysis
PCOSPolycystic Ovary Syndrome
PDADPattern Diagnosis and Acupuncture Dataset
PLMPre-trained Language Model
PRISMAPreferred Reporting Items for Systematic Reviews and Meta-Analyses
PRISMA-AIPreferred Reporting Items for Systematic Reviews and Meta-Analyses for Artificial Intelligence
PROBAST-AIPrediction model Risk Of Bias ASsessment Tool for Artificial Intelligence
PROSPEROInternational Prospective Register of Systematic Reviews
QUADAS-2Quality Assessment of Diagnostic Accuracy Studies 2
RAGRetrieval-Augmented Generation
RCNNRecurrent Convolutional Neural Network
ResNetResidual Network
RGCNRelational Graph Convolutional Network
RNNRecurrent Neural Network
ROUGERecall-Oriented Understudy for Gisting Evaluation
SKQDSpleen-Kidney Qi Deficiency
SSEMSemantic Similarity Evaluation Model
SVMSupport Vector Machine
T2DMType 2 Diabetes Mellitus
TCMTraditional Chinese Medicine
TRIPOD-AITransparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis—Artificial Intelligence
TS-ModelTongue-Syndrome Multimodal Fusion Model
XGBoostExtreme Gradient Boosting

References

  1. Sui, D.; Zhang, L.; Yang, F. Data-driven based four examinations in TCM: A survey. Digit. Chin. Med. 2022, 5, 377–385. [Google Scholar] [CrossRef]
  2. Jiang, M.; Lu, C.; Zhang, C.; Yang, J.; Tan, Y.; Lu, A.; Chan, K. Syndrome differentiation in modern research of traditional Chinese medicine. J. Ethnopharmacol. 2012, 140, 634–642. [Google Scholar] [CrossRef]
  3. Bensoussan, A.; Talley, N.J.; Hing, M.; Menzies, R.; Guo, A.; Ngu, M. Treatment of irritable bowel syndrome with Chinese herbal medicine: A randomized controlled trial. JAMA 1998, 280, 1585–1589. [Google Scholar] [CrossRef] [PubMed]
  4. Chen, S.C.; Yu, J.; Wang, H.S.; Wang, D.D.; Sun, Y.; Cheng, H.L.; Suen, L.K.; Yeung, W.F. Parent-administered pediatric Tuina for attention deficit/hyperactivity disorder symptoms in preschool children: A pilot randomized controlled trial embedded with a process evaluation. Phytomedicine 2022, 102, 154191. [Google Scholar] [CrossRef] [PubMed]
  5. Yeung, W.F.; Yu, B.Y.; Yuen, J.W.; Ho, J.Y.S.; Chung, K.F.; Zhang, Z.J.; Mak, D.S.Y.; Suen, L.K.; Ho, L.M. Semi-Individualized Acupuncture for Insomnia Disorder and Oxidative Stress: A Randomized, Double-Blind, Sham-Controlled Trial. Nat. Sci. Sleep 2021, 13, 1195–1207. [Google Scholar] [CrossRef]
  6. Tian, D.; Chen, W.; Xu, D.; Xu, L.; Xu, G.; Guo, Y.; Yao, Y. A review of traditional Chinese medicine diagnosis using machine learning: Inspection, auscultation-olfaction, inquiry, and palpation. Comput. Biol. Med. 2024, 170, 108074. [Google Scholar] [CrossRef]
  7. Poon, M.M.; Chung, K.F.; Yeung, W.F.; Yau, V.H.; Zhang, S.P. Classification of insomnia using the traditional chinese medicine system: A systematic review. Evid. Based Complement. Altern. Med. 2012, 2012, 735078. [Google Scholar] [CrossRef]
  8. Liu, J.Y.; Li, X. Standardization, objectification, and essence research of traditional Chinese medicine syndrome: A 15-year bibliometric and content analysis from 2006 to 2020 in Web of Science database. Anat. Rec. 2023, 306, 2974–2983. [Google Scholar] [CrossRef] [PubMed]
  9. Lu, A.; Jiang, M.; Zhang, C.; Chan, K. An integrative approach of linking traditional Chinese medicine pattern classification and biomedicine diagnosis. J. Ethnopharmacol. 2012, 141, 549–556. [Google Scholar] [CrossRef]
  10. Matos, L.C.; Machado, J.P.; Monteiro, F.J.; Greten, H.J. Can Traditional Chinese Medicine Diagnosis Be Parameterized and Standardized? A Narrative Review. Healthcare 2021, 9, 177. [Google Scholar] [CrossRef]
  11. Hu, Y.; Wang, Z.; Ni, K.; Yang, J. Challenges in Traditional Chinese Medicine Clinical Trials: How to Balance Personalized Treatment and Standardized Research? Ther. Clin. Risk Manag. 2025, 21, 1085–1094. [Google Scholar] [CrossRef]
  12. Wang, J.; Guo, Y.; Li, G.L. Current Status of Standardization of Traditional Chinese Medicine in China. Evid. Based Complement. Altern. Med. 2016, 2016, 9123103. [Google Scholar] [CrossRef]
  13. Chen, S.C.; Chen, Y.; Yeung, W.F.; Ren, G.; Chen, S.; Du, H.X.; Qin, J. Deep learning in acupuncture: A systematic review. Artif. Intell. Med. 2026, 171, 103300. [Google Scholar] [CrossRef]
  14. Faiyazuddin, M.; Rahman, S.J.Q.; Anand, G.; Siddiqui, R.K.; Mehta, R.; Khatib, M.N.; Gaidhane, S.; Zahiruddin, Q.S.; Hussain, A.; Sah, R. The Impact of Artificial Intelligence on Healthcare: A Comprehensive Review of Advancements in Diagnostics, Treatment, and Operational Efficiency. Health Sci. Rep. 2025, 8, e70312. [Google Scholar] [CrossRef] [PubMed]
  15. Mienye, I.D.; Swart, T.G. A Comprehensive Review of Deep Learning: Architectures, Recent Advances, and Applications. Information 2024, 15, 755. [Google Scholar] [CrossRef]
  16. Sarker, I.H. Deep Learning: A Comprehensive Overview on Techniques, Taxonomy, Applications and Research Directions. SN Comput. Sci. 2021, 2, 420. [Google Scholar] [CrossRef] [PubMed]
  17. Shickel, B.; Tighe, P.J.; Bihorac, A.; Rashidi, P. Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis. IEEE J. Biomed. Health Inform. 2018, 22, 1589–1604. [Google Scholar] [CrossRef]
  18. Chen, R.J.; Lu, M.Y.; Williamson, D.F.K.; Chen, T.Y.; Lipkova, J.; Noor, Z.; Shaban, M.; Shady, M.; Williams, M.; Joo, B.; et al. Pan-cancer integrative histology-genomic analysis via multimodal deep learning. Cancer Cell 2022, 40, 865–878.e866. [Google Scholar] [CrossRef] [PubMed]
  19. Muzammil, M.A.; Javid, S.; Afridi, A.K.; Siddineni, R.; Shahabi, M.; Haseeb, M.; Fariha, F.N.U.; Kumar, S.; Zaveri, S.; Nashwan, A.J. Artificial intelligence-enhanced electrocardiography for accurate diagnosis and management of cardiovascular diseases. J. Electrocardiol. 2024, 83, 30–40. [Google Scholar] [CrossRef]
  20. Askr, H.; Elgeldawi, E.; Aboul Ella, H.; Elshaier, Y.; Gomaa, M.M.; Hassanien, A.E. Deep learning in drug discovery: An integrative review and future challenges. Artif. Intell. Rev. 2023, 56, 5975–6037. [Google Scholar] [CrossRef]
  21. Tian, Z.; Wang, D.; Sun, X.; Fan, Y.; Guan, Y.; Zhang, N.; Zhou, M.; Zeng, X.; Yuan, Y.; Bu, H.; et al. Current status and trends of artificial intelligence research on the four traditional Chinese medicine diagnostic methods: A scientometric study. Ann. Transl. Med. 2023, 11, 145. [Google Scholar] [CrossRef] [PubMed]
  22. Meng, D.; Cao, G.; Duan, Y.; Zhu, M.; Tu, L.; Xu, D.; Xu, J. Tongue Images Classification Based on Constrained High Dispersal Network. Evid. Based Complement. Altern. Med. 2017, 2017, 7452427. [Google Scholar] [CrossRef]
  23. Quanyu, E. Pulse Signal Analysis Based on Deep Learning Network. BioMed Res. Int. 2022, 2022, 6256126. [Google Scholar] [CrossRef]
  24. Yang, L.; Dong, Q.; Lin, D.; Lü, X. TongueNet: A multi-modal fusion and multi-label classification model for traditional Chinese Medicine tongue diagnosis. Front. Physiol. 2025, 16, 1527751. [Google Scholar] [CrossRef]
  25. Liu, Y.; Fan, L.; Zhao, M.; Wei, D.; Zhao, M.; Dong, Y.; Zhang, X. Study on a Traditional Chinese Medicine constitution recognition model using tongue image characteristics and deep learning: A prospective dual-center investigation. Chin. Med. 2025, 20, 84. [Google Scholar] [CrossRef]
  26. Xiao, L.; Wang, J.W.; Wang, C.W.; Wang, Y.; Yan, J.F.; Peng, Q.H. Knowledge graph for traditional Chinese medicine diagnosis and treatment of diabetic retinopathy: Design, construction, and applications. Int. J. Ophthalmol. 2025, 18, 2011–2021. [Google Scholar] [CrossRef]
  27. He, J.; Guo, Y.; Lam, L.; Leung, W.; He, L.; Jiang, Y.; Wang, C.; Xing, G.; Chen, H. OpenTCM: A GraphRAG-Empowered LLM-based System for Traditional Chinese Medicine Knowledge Retrieval and Diagnosis. arXiv 2025, arXiv:2504.20118. [Google Scholar] [CrossRef]
  28. Lim, J.; Li, J.; Zhou, M.; Xiao, X.; Xu, Z. Machine Learning Research Trends in Traditional Chinese Medicine: A Bibliometric Review. Int. J. Gen. Med. 2024, 17, 5397–5414. [Google Scholar] [CrossRef]
  29. Wang, S.; Zhao, X. Deep Learning in Tongue Diagnosis for Traditional Chinese Medicine: A Review. Front. Comput. Intell. Syst. 2025, 12, 139–143. [Google Scholar] [CrossRef]
  30. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [PubMed]
  31. Cacciamani, G.E.; Chu, T.N.; Sanford, D.I.; Abreu, A.; Duddalwar, V.; Oberai, A.; Kuo, C.J.; Liu, X.; Denniston, A.K.; Vasey, B.; et al. PRISMA AI reporting guidelines for systematic reviews and meta-analyses on AI in healthcare. Nat. Med. 2023, 29, 14–15. [Google Scholar] [CrossRef]
  32. Krizhevsky, A.; Sutskever, I.; Hinton, G.E. ImageNet classification with deep convolutional neural networks. Commun. ACM 2017, 60, 84–90. [Google Scholar] [CrossRef]
  33. Collins, G.S.; Moons, K.G.M.; Dhiman, P.; Riley, R.D.; Beam, A.L.; Van Calster, B.; Ghassemi, M.; Liu, X.; Reitsma, J.B.; van Smeden, M.; et al. TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024, 385, e078378. [Google Scholar] [CrossRef] [PubMed]
  34. Whiting, P.F.; Rutjes, A.W.; Westwood, M.E.; Mallett, S.; Deeks, J.J.; Reitsma, J.B.; Leeflang, M.M.; Sterne, J.A.; Bossuyt, P.M. QUADAS-2: A revised tool for the quality assessment of diagnostic accuracy studies. Ann. Intern. Med. 2011, 155, 529–536. [Google Scholar] [CrossRef] [PubMed]
  35. Liu, J.; Dan, W.; Liu, X.; Zhong, X.; Chen, C.; He, Q.; Wang, J. Development and validation of predictive model based on deep learning method for classification of dyslipidemia in Chinese medicine. Health Inf. Sci. Syst. 2023, 11, 21. [Google Scholar] [CrossRef] [PubMed]
  36. Wang, C.Y.; Huang, K.L.; Dai, G.W.; Qiang, M.; Wang, Q. Research on Tongue Image Classification of Traditional Chinese Medicine Syndrome Differentiation in Stroke Based on Convolutional Neural Networks. J. Hunan Univ. Chin. Med. 2023, 43, 1460–1467. (In Chinese) [Google Scholar] [CrossRef]
  37. Yu, X.; He, L.; Wang, Q.; Zhang, Z.; Zhu, H.; Song, J. Heat syndrome types prediction of traditional Chinese medicine in acute ischemic stroke through deep learning: A pilot study. Front. Pharmacol. 2025, 16, 1601601. [Google Scholar] [CrossRef]
  38. Zhang, Y.; Ling, N.; Zhang, G.H.; Chen, X.; Ren, C.L. Construction and Implementation of Syndrome Differentiation Classification for Polycystic Ovary Syndrome Based on Deep Learning. J. Liaoning Univ. Tradit. Chin. Med. 2019, 21, 13–16. (In Chinese) [Google Scholar] [CrossRef]
  39. Zhang, M.Q.; Deng, X. Exploration of Constructing an Intelligent Syndrome Differentiation Model for Compensated Cirrhosis Based on Integrated Learning and Neural Network Algorithm. J. Guangzhou Univ. Chin. Med. 2023, 40, 2650–2660. (In Chinese) [Google Scholar] [CrossRef]
  40. Zhao, Z.H.; Zhou, Y.; Li, W.H.; Tang, C.H.; Guo, Q.; Chen, R.G. Construction of a Traditional Chinese Medicine Syndrome Differentiation Model for Type 2 Diabetes Based on Deep Learning Multimodal Fusion. World Sci. Technol. Mod. Chin. Med. 2024, 26, 908–918. (In Chinese) [Google Scholar] [CrossRef]
  41. Ding, Y.; Wang, Y.; Liu, J.; Wang, N.Y. Construction of a Diagnostic Model for Coronary Heart Disease with Blood Stasis Pattern Based on Objective Multimodal Data Fusion. J. Tradit. Chin. Med. 2025, 66, 2239–2248. (In Chinese) [Google Scholar] [CrossRef]
  42. Ding, Y.; Wang, Y.; Liu, J.; Wang, N.Y. Diagnostic Model Construction for Blood Stasis Syndrome in T2DM Based on Objective Multimodal Data Fusion. J. Basic. Chin. Med. 2025, 31, 1592–1598. (In Chinese) [Google Scholar] [CrossRef]
  43. Zhang, Q.Q.; Yang, H.; Wu, Y.X.; Li, S.J.; Peng, X. Diagnostic and Treatment Decision Support System for TCM Diabetic Nephropathy Based on Artificial Intelligence. Chin. J. Health Inform. Manag. 2025, 22, 998–1006, 1021. (In Chinese) [Google Scholar] [CrossRef]
  44. Li, H.P.; Han, Z.Y.; Hu, W.Y.; Wang, H.Y.; Li, Y.L. Deep learning-assisted diagnosis model of depression based on TCM facial inspection. Mod. Chin. Clin. Med. 2026, 33, 27–32. (In Chinese) [Google Scholar] [CrossRef]
  45. Cao, M.R.; Wang, Z.C. Accuracy Study of AI-Assisted Syndrome Differentiation in Determining Traditional Chinese Medicine Syndrome Types (Using Spleen Deficiency Pattern as an Example) in Patients with Gastrointestinal Tumors. Mystery 2023, 6, 49–51. Available online: https://qikan.cqvip.com/Qikan/Article/Detail?id=7201982295 (accessed on 10 May 2026). (In Chinese)
  46. Liu, Z.F.; Wu, J.H. Research on TCM Syndrome Diagnosis Based on Knowledge Graph and Attention Mechanism. J. Sci. Res. Appl. 2025, 1, 89–93. Available online: https://qikan.cqvip.com/Qikan/Article/Detail?id=7202913824 (accessed on 10 May 2026). (In Chinese)
  47. Yang, L.; Tong, T.; Lu, Y.; Wang, W.; Wei, C.Y.; Wang, R.T.; Ouyang, Y.L.; Moossavi, M.; Liu, H.X.; Ma, X.L. Deep Learning Based on the Facial Color Images for Assisted Diagnosis of Chronic Renal Failure with Spleen-Kidney Qi Deficiency Syndrome. World J. Tradit. Chin. Med. 2026, 12, 55–67. [Google Scholar] [CrossRef] [PubMed]
  48. Ding, L.; Zhang, X.Y.; Liu, L.P.; Niu, X.L.; Guo, Y.K. A Primary Liver Cancer Syndrome Diagnosis Classification Prediction Model Based on Deep Neural Networks. World Sci. Technol. Mod. Chin. Med. 2020, 22, 4185–4192. (In Chinese) [Google Scholar] [CrossRef]
  49. Li, H.Z.; Wang, J.; Zhang, Z.P.; Guo, Y.C.; Du, Q.; Gao, J.L.; Dong, Y.; Li, J.N.; Li, Q.Y. Research on the Distribution and Combination Patterns of Syndrome Elements Throughout the Course of Coronary Heart Disease Based on an Improved Deep Learning Algorithm. World Sci. Technol. Mod. Chin. Med. 2021, 23, 3086–3094. (In Chinese) [Google Scholar] [CrossRef]
  50. Li, D.; Qu, J.; Tian, Z.; Mou, Z.; Zhang, L.; Zhang, X. Knowledge-Based Recurrent Neural Network for TCM Cerebral Palsy Diagnosis. Evid. Based Complement. Altern. Med. 2022, 2022, 7708376. [Google Scholar] [CrossRef]
  51. Zhu, L.; Zheng, W.T.; Zhang, Z.L.; Tian, S.L.; Li, J.H. Research on the Syndrome Prediction Model of Ulcerative Colitis Based on Convolutional Neural Networks. China Digit. Med. 2022, 17, 49–55. (In Chinese) [Google Scholar] [CrossRef]
  52. Jiaze, W.; Hao, L.; Haoran, D.; Hongliang, R.; Baoli, L. Clinical decision and prescription generation for diarrhea in traditional Chinese medicine based on large language model. Digit. Chin. Med. 2026, 9, 13–30. [Google Scholar] [CrossRef]
  53. Cai, S.X.; Zhu, D.N.; Xu, H.W.; Huang, J.Q.; Li, X.Y.; Li, Z.H.; Liu, X.F.; Yan, Y.H. Application of a deep learning-based tongue image classification system in TCM syndrome identification of psoriasis. J. Pract. Med. 2026, 42, 395–405. (In Chinese) [Google Scholar] [CrossRef]
  54. Li, S.; Tan, W.; Zhang, C.; Li, J.; Ren, H.; Guo, Y.; Jia, J.; Liu, Y.; Pan, X.; Guo, J.; et al. Taming large language models to implement diagnosis and evaluating the generation of LLMs at the semantic similarity level in acupuncture and moxibustion. Expert. Syst. Appl. 2025, 264, 125920. [Google Scholar] [CrossRef]
  55. Wang, J.; Song, Y.J.; Hui, X.S.; Zhang, Z.P.; Zhang, X.C. Chain Thinking-driven Large Language Model for Traditional Chinese Medicine Syndrome Element Differentiation of Coronary Heart Disease. Chin. J. Exp. Tradit. Med. Formulae 2025, 31, 1–10. (In Chinese) [Google Scholar] [CrossRef]
  56. Lin, S.Y.; Huang, H.W.; Liu, C.; Liu, W.T.; Li, J.M.; Qu, Y.Q.; Cao, L.Y. Cognitive mechanisms and multimodal research methodologies of traditional Chinese medicine diagnostics under the large language model perspective. CJTCMP 2025, 40, 97–102. (In Chinese) [Google Scholar]
  57. Aggarwal, R.; Sounderajah, V.; Martin, G.; Ting, D.S.W.; Karthikesalingam, A.; King, D.; Ashrafian, H.; Darzi, A. Diagnostic accuracy of deep learning in medical imaging: A systematic review and meta-analysis. npj Digit. Med. 2021, 4, 65. [Google Scholar] [CrossRef] [PubMed]
  58. Nwanosike, E.M.; Conway, B.R.; Merchant, H.A.; Hasan, S.S. Potential applications and performance of machine learning techniques and algorithms in clinical practice: A systematic review. Int. J. Med. Inform. 2022, 159, 104679. [Google Scholar] [CrossRef]
  59. Vecchi, E.; Pospíšil, L.; Albrecht, S.; Kane, T.J.O.; Horenko, I. eSPA+: Scalable Entropy-Optimal Machine Learning Classification for Small Data Problems. Neural Comput. 2022, 34, 1220–1255. [Google Scholar] [CrossRef]
  60. Chen, P.-H.C.; Mermel, C.H.; Liu, Y. Evaluation of artificial intelligence on a reference standard based on subjective interpretation. Lancet Digit. Health 2021, 3, e693–e695. [Google Scholar] [CrossRef]
  61. Wu, D.; Smith, D.; VanBerlo, B.; Roshankar, A.; Lee, H.; Li, B.; Ali, F.; Rahman, M.; Basmaji, J.; Tschirhart, J.; et al. Improving the Generalizability and Performance of an Ultrasound Deep Learning Model Using Limited Multicenter Data for Lung Sliding Artifact Identification. Diagnostics 2024, 14, 1081. [Google Scholar] [CrossRef]
  62. Cascarano, A.; Mur-Petit, J.; Hernández-González, J.; Camacho, M.; de Toro Eadie, N.; Gkontra, P.; Chadeau-Hyam, M.; Vitrià, J.; Lekadir, K. Machine and deep learning for longitudinal biomedical data: A review of methods and applications. Artif. Intell. Rev. 2023, 56, 1711–1771. [Google Scholar] [CrossRef]
  63. Li, Y.; El Habib Daho, M.; Conze, P.H.; Zeghlache, R.; Le Boité, H.; Tadayoni, R.; Cochener, B.; Lamard, M.; Quellec, G. A review of deep learning-based information fusion techniques for multimodal medical image classification. Comput. Biol. Med. 2024, 177, 108635. [Google Scholar] [CrossRef]
  64. Wang, J.; Liu, Y.M.; Li, J.; He, H.Q.; Liu, C.; Song, Y.J.; Ma, S.Y. Artificial Intelligence in Traditional Chinese Medicine: Multimodal Fusion and Machine Learning for Enhanced Diagnosis and Treatment Efficacy. Curr. Med. Sci. 2025, 45, 1013–1022. [Google Scholar] [CrossRef]
  65. Song, Y.J.; Ma, S.Y.; Dai, Y.S.; Lu, J. AI-Assisted TCM Syndrome Differentiation: Key Issues and Technical Challenges. Chin. Eng. Sci. 2024, 26, 234–244. [Google Scholar] [CrossRef]
  66. Wenderoth, L. Exploring Multi-Modality Dynamics: Insights and Challenges in Multimodal Fusion for Biomedical Tasks. arXiv 2022, arXiv:2411.00725. [Google Scholar] [CrossRef]
  67. Gomez-Cabello, C.A.; Borna, S.; Pressman, S.; Haider, S.A.; Haider, C.R.; Forte, A.J. Artificial-Intelligence-Based Clinical Decision Support Systems in Primary Care: A Scoping Review of Current Clinical Implementations. Eur. J. Investig. Health Psychol. Educ. 2024, 14, 685–698. [Google Scholar] [CrossRef]
  68. Yanhong, W.; Xin, Y.; Yun, Y.; Ji, C.; Ge, Z.; Jianhui, T. Research progress and challenges in artificial intelligence-driven intelligent traditional Chinese medicine diagnosis and treatment. Shanghai J. Tradit. Chin. Med. 2026, 60, 1–11. [Google Scholar] [CrossRef]
  69. Gao, Z.; Chen, T.; Ha, Y.; Shi, Y.; Xu, X.; Li, B.; Liu, Q. Staged identification of CAP in fever patients across epidemic environments: Modeling & validation. Sci. Rep. 2025, 16, 258. [Google Scholar] [CrossRef]
  70. Dai, Y.; Shao, X.; Zhang, J.; Chen, Y.; Chen, Q.; Liao, J.; Chi, F.; Zhang, J.; Fan, X. TCMChat: A generative large language model for traditional Chinese medicine. Pharmacol. Res. 2024, 210, 107530. [Google Scholar] [CrossRef]
  71. Deng, R.Y.; Ding, K.; Zhu, Y.X.; Li, M.; Zhu, H.T.; Du, L. Distribution Characteristics and Risk Factors of Traditional Chinese Medicine Syndrome Types of Mixed Hemorrhoids Based on Large Language Models and Latent Class Models: A Multicenter Real-World Study. J. Tradit. Chin. Med. 2026, 67, 755–763. (In Chinese) [Google Scholar] [CrossRef]
  72. Syed, M.; Hamidi, M.; Bikkanuri, M.; Dierschke, N.A.; Katragadda, H.V.; Zozus, M.; Teixeira, A.L. Translating evidence into practice: Adapting TrialGPT for real-world clinical trial eligibility screening. J. Am. Med. Inform. Assoc. 2026, 33, 909–913. [Google Scholar] [CrossRef]
  73. Moons, K.G.M.; Damen, J.A.A.; Kaul, T.; Hooft, L.; Andaur Navarro, C.; Dhiman, P.; Beam, A.L.; Van Calster, B.; Celi, L.A.; Denaxas, S.; et al. PROBAST+AI: An updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ 2025, 388, e082505. [Google Scholar] [CrossRef]
  74. Berlin, J.A.; Fihn, S.D. Encouraging the Registration of Observational Studies. JAMA Netw. Open 2025, 8, e2524181. [Google Scholar] [CrossRef]
  75. Binney, R.J.; Smith, L.J.; Rossit, S.; Demeyere, N.; Learmonth, G.; Olgiati, E.; Halai, A.D.; Rounis, E.; Evans, J.; Edelstyn, N.M.J.; et al. Practical routes to preregistration: A guide to enhanced transparency and rigour in neuropsychological research. Brain Commun. 2025, 7, fcaf162. [Google Scholar] [CrossRef]
  76. Gou, X.; Yao, J.; Lai, W.; Gao, Y.; Wang, S.; Zhou, C.; Ye, H.; Tian, J.; Yi, J.; Cao, D. A framework for normalized extraction of fine-grained traditional Chinese medicine symptom entities and relations. BMC Med. Inform. Decis. Mak. 2025, 25, 441. [Google Scholar] [CrossRef]
  77. Jacobson, E.; Conboy, L.; Tsering, D.; Shields, M.; McKnight, P.; Wayne, P.M.; Schnyer, R. Experimental Studies of Inter-Rater Agreement in Traditional Chinese Medicine: A Systematic Review. J. Altern. Complement. Med. 2019, 25, 1085–1096. [Google Scholar] [CrossRef]
  78. Weikang, K.; Chuanbiao, W.; Yue, L. Knowledge graph-enhanced long-tail learning approach for traditional Chinese medicine syndrome differentiation. Digit. Chin. Med. 2026, 9, 57–67. [Google Scholar] [CrossRef]
Figure 1. PRISMA flow diagram.
Figure 1. PRISMA flow diagram.
Information 17 00554 g001
Figure 2. Summary of Risk of Bias and Applicability Concerns Assessed by QUADAS-2.
Figure 2. Summary of Risk of Bias and Applicability Concerns Assessed by QUADAS-2.
Information 17 00554 g002
Table 1. Basic characteristics of included studies.
Table 1. Basic characteristics of included studies.
Variablen (%)
Year of Publication
 2019–20213 (14.3)
 2022–20247 (33.3)
 2025–202611 (52.4)
Language
 Chinese15 (71.4)
 English6 (28.6)
Disease category
 Cardiovascular and cerebrovascular diseases6 (28.6)
 Digestive system diseases5 (23.8)
 Endocrine and metabolic diseases4 (19.0)
 Others (renal, neurological, psychiatric, dermatological, reproductive, multisystem, etc.)6 (28.6)
TCM patterns reported
 Yes20 (95.2)
 No1 (4.8)
Disease diagnostic criteria *
 Guidelines/Consensus10 (47.6)
 Industry/National standard3 (14.3)
 Textbook1 (4.8)
 Study-specific criteria1 (4.8)
 Not reported7 (33.3)
TCM pattern diagnostic criteria *
 Guidelines/Consensus8 (38.1)
 Industry/National standard9 (42.9)
 Textbook3 (14.3)
 Study-specific criteria2 (9.5)
 Not reported3 (14.3)
Data sources
 Self-built database17 (81.0)
 Public dataset1 (4.8)
 Mixed sources3 (14.3)
Data source center type
 Single-center12 (57.1)
 Multi-center8 (38.1)
 Not applicable1 (4.8)
Modality type
 Text12 (57.1)
 Image4 (19.0)
 Multimodal5 (23.8)
Label source
 Expert/Clinician annotation12 (57.1)
 Guideline/Standard/EMR extraction3 (14.3)
 LLM-assisted + manual review1 (4.8)
 Not reported5 (23.8)
Task category
 Classification Tasks13 (61.9)
 Identification/Analysis Tasks6 (28.6)
 Generation/Decision Tasks2 (9.5)
Performance metrics category *
 Accuracy metrics20 (95.2)
 Error/Loss metrics2 (9.5)
 Generation quality metrics2 (9.5)
 Efficiency/Computational metrics3 (14.3)
 Others (segmentation metrics, etc.)1 (4.8)
External validation
 No21 (100)
Type of comparison *
 Comparison with other computational models12 (57.1)
 Comparison with clinical/human experts2 (9.5)
 Ablation study/Internal control7 (33.3)
 Other comparisons (healthy/disease controls)1 (4.8)
 Not applicable1 (4.8)
Remarks: * Percentages sum to more than 100% because multiple responses were permitte.
Table 2. Model and methodological characteristics.
Table 2. Model and methodological characteristics.
Study IDTasksModelsModality TypeProcessPerformance
Metrics
Classification Tasks
Liu
(2023)
[35]
Predict dyslipidemia occurrenceANNsText (structured clinical data: symptoms, tongue, pulse)Data Preparation: Clinical data cleaned and features selected.
Model Training: ANN trained with class balancing and early stopping.
Model Validation: Evaluated on independent test set.
Model-11 (test): TP = 51, FP = 15, TN = 129, FN = 9;
Loss 0.3241;
Accuracy 0.8672;
Precision 0.7138;
Recall 0.8286;
AUC 0.9268.
Wang (2023)
[36]
Classify stroke tongue images into eight TCM patternsDenseNet201Tongue imagesData Preparation: Tongue images processed and deep features extracted.
Model Training: DenseNet + SVM used for classification.
Model Validation: Evaluated by cross-validation.
Accuracy 95.74–98.22%;
F1 96.49–98.31%;
Precision & Sensitivity >95%.
Yu
(2025)
[37]
Classify heat vs. non-heat patterns in acute ischemic strokeCNNImage/Text (TCM pattern characteristics, lab indicators, tongue/facial)Data Preparation: Data preprocessed and features selected.
Model Training: CNN trained with leave-one-out validation.
Model Validation: Tested on independent dataset.
Accuracy 0.95;
F1 0.95;
AUC 0.91 (test).
Zhang (2019)
[38]
Predict TCM patterns for PCOSDBNText (structured clinical indicators)Data Preparation: Data normalized.
Model Training: DBN built and optimized.
Model Validation: Evaluated model accuracy.
Total accuracy 87.07%;
Liver stagnation 81.58%;
Kidney deficiency 82.5%; Phlegm-dampness 92.42%;
Blood stasis 88.24%.
Zhang (2023)
[39]
Predict TCM patterns for compensated liver cirrhosisBP-NNText (structured symptom, sign, tongue, pulse data)Data Preparation: Multiple ML models built.
Model Training: BP neural network used for stacking fusion.
Model Validation: Evaluated hybrid model performance.
Fusion model (BP-NN): Accuracy 0.94;
Precision 0.92;
Recall 0.90;
F1 0.95;
AUC 0.99.
Zhao
(2024)
[40]
Predict 14 Zheng elements (location and nature)U2-Net;
ResNet34;
FCN;
TS-Model
Tongue images + Text (structured symptom data)Data Preparation: Tongue and symptom data preprocessed.
Model Training: T-Model and S-Model fused for multimodal learning.
Model Validation: Compared with unimodal baselines.
TS-Model F1 range: T 0–86.73%, S 0–97.83%, TS 55.56–99.07%;
TS model more stable and superior to unimodal models.
Ding
(2025) a
[41]
Classify blood stasis vs. non-blood stasis in CHD patientsCNNText + Image/Video (pulse: 193 features; tongue: sublingual vessel images; voice: MFCC; inquiry: clinical questionnaire)Data Preparation: Multi-center collection, cleaning, single-modal analysis (pulse/tongue/inquiry);
Model Training: Multimodal fusion with DL;
Model Validation: Test set evaluation.
Single-modal tongue: 83.33%;
Four-modal fusion: 86.11% (precision 86.17%, recall 86.35%, F1 86.11%)
Ding
(2025) b
[42]
Classify blood stasis vs. non-blood stasis in T2DM patientsAttention networkText + Image/Video (pulse: time/frequency domain; tongue: sublingual vessel images; voice: MFCC; inquiry: symptom features)Data Preparation: National multi-center collection, cleaning, single-modal analysis (pulse/tongue/voice/inquiry);
Model Training: Multimodal fusion with deep learning;
Model Validation: Test set evaluation.
Single-modal inquiry: 78.38%;
Four-modal fusion: 86.12% (precision 86.45%, recall 86.09%, F1 86.27%)
Zhang (2025)
[43]
Classify DKD patients into 11 TCM patterns (multi-label)GAT-BERTText (EMR: chief complaint, tongue/pulse, present illness, specialty exam, past history, physical exam, personal history)Data Preparation: DKD knowledge graph construction, subgraph generation;
Model Training: Graph + text embedding with multi-label classification;
Model Validation: 5-fold cross-validation.
Micro-F1 0.901,
Macro-F1 0.821
Li
(2026)
[44]
Classify depression with TCM pattern vs. non-depressed same pattern vs. healthy controlEfficientNet, MobileNet V3, ResNet18Image (facial images, RGB)Data Preparation: Facial image QC, preprocessing, augmentation;
Model Training: EfficientNet/MobileNet V3/ResNet18;
Model Validation: Validation, test, CAM visualization.
EfficientNet: 98.6% (4.01 M);
MobileNet V3: 92.7% (1.52 M);
ResNet18: 92.2% (11.18 M)
Cao
(2023)
[45]
Classify spleen deficiency vs. non-spleen deficiencyCNN + RNNText + Image (tongue images via CNN; symptoms/signs/pulse via RNN)Data Preparation: Expert consensus establishment, clinical data collection;
Model Training: AI system differentiation;
Model Validation: Comparison with expert consensus, subgroup analysis.
Accuracy 91.46%,
Kappa 0.808 (gastric 93.55%,
colorectal 89.66%, esophageal 90.91%)
Liu
(2025)
[46]
Classify 5 TCM syndromesKnowledge Graph + AttentionText (symptoms only)Data Preparation: Knowledge graph construction, embedding;
Model Training: Attention mechanism with DNN;
Model Validation: Testing on held-out data.
Macro F1 0.8998,
Macro AUC 0.9833
Yang
(2026)
[47]
Classify SKQD vs. non-SKQD in CRF patientsResNet-18Image (facial images, 3 angles: frontal/left/right; RGB + Lab)Data Preparation: Facial image color correction, ROI extraction, augmentation;
Model Training: Multibranch ResNet-18, CHAID, logistic regression;
Model Validation: Test set + 10-fold cross-validation.
DL model: accuracy 73.77%, AUC 0.74;
DL + clinical: accuracy 75.41%, AUC 0.75
Identification/Analysis Tasks
Ding
(2020)
[48]
Diagnose and classify TCM patterns in primary liver cancerDNNText (TCM symptoms, signs, tongue, pulse)Data Preparation: Medical records collected and syndrome factors quantified.
Model Training: DNN constructed to predict TCM syndromes.
Model Validation: Tested and validated with association rule consistency.
Pattern prediction accuracy 82.86–92.76%;
Rule validation consistency 75–100%.
Li
(2021)
[49]
Diagnose TCM pattern elements in coronary artery diseaseTransformerText (symptoms, tongue, pulse descriptions)Data Preparation: Standardized symptom and syndrome data.
Model Training: Transformer model applied for syndrome element diagnosis.
Model Validation: Compared with physician diagnosis for accuracy.
Diagnostic accuracy 96.46 ± 8.96%.
Li
(2022)
[50]
Identify TCM patterns using electronic medical recordsKBRNNText (symptom descriptions)Data Preparation: Extracted and standardized EMR and knowledge graph data.
Model Training: KBRNN fine-tuned with knowledge injection.
Model Validation: Evaluated on test set against baseline models.
KBRNN accuracy: untrained 79.31%, trained 83.12%.
Zhu
(2022)
[51]
Predict TCM patterns of ulcerative colitis from clinical recordsCNN-GRUText (clinical manifestations, tongue, pulse, questionnaires)Data Preparation: Symptoms extracted and labels digitized.
Model Training: CNN–GRU trained for pattern classification.
Model Validation: Tested for accuracy and generalization.
CNN: Accuracy 88%, Recall 88%, F1 0.88;
GRU: Accuracy 86%, Recall 86%, F1 0.86.
Wu
(2026)
[52]
Generate diagnosis (binary), differentiate 6-class syndromes, and recommend prescriptionQwen2.5-7BText (clinical records, symptoms, tongue/pulse descriptions)Data Preparation: Filtering, deduplication, instruction/CoT annotation;
Model Training: Two-stage LoRA fine-tuning on Qwen2.5-7B;
Model Validation: 10-fold cross-validation.
Disease diagnosis: accuracy 97.05%, F1 91.48%;
Syndrome differentiation: accuracy 74.54%, F1 74.21%
Cai
(2026)
[53]
Tongue segmentation and identify TCM syndromesImproved U-Net, EfficientNet-B3, Swin-TinyImage (tongue images, RGB)Data Preparation: Tongue image collection, manual annotation, preprocessing;
Model Training: Improved U-Net (segmentation) + EfficientNet-B3/Swin-Tiny (classification);
Model Validation: 5-fold cross-validation.
Segmentation: Dice 0.98; Classification: HybridModel accuracy 98.16%, AUC 99.93%
Generation/Decision Tasks
Li
(2025)
[54]
Generate acupuncture diagnosis and treatment plansAcupunctureGPTText (patient description, clinical records, diagnosis and treatment plan texts extracted from hospital case database, electronic textbooks, and acupuncture guides)Data Preparation: Built PDAD dataset for acupuncture diagnosis.
Model Training: Fine-tuned GPT model (AcupunctureGPT).
Model Validation: Improved reasoning and evaluation with SSEM and GKFP.
BLEU-1 F1 0.2012;
ROUGE-1 F1 0.3268;
Higher semantic similarity vs. other LLMs.
Wang (2025)
[55]
Identify 8 syndrome elements and 6 target syndromesDeepSeek-r1:32bText (symptoms from medical records)Data Preparation: Symptom standardization (7973 → 218 terms), prompt template;
Model Training: DeepSeek-r1:32b with zero-shot inference;
Model Validation: Output interpretation.
CHD-SEDD: Macro-F1 85.0%;
No Prompt baseline: Macro-F1 61.2%
Table 3. Item-specific reporting quality of the TRIPOD-AI checklist for included studies.
Table 3. Item-specific reporting quality of the TRIPOD-AI checklist for included studies.
Section/TopicItemReportedNot ReportedNot ApplicableItem-Specific Reporting Rates
Title
Title1210 0100.0%
Abstract
Abstract2100 0100.0%
Introduction
Background3a2100100.0%
3b2100100.0%
3c02100.0%
Objectives42100100.0%
Methods
Data5a2100100.0%
5b174081.0%
Participants6a2100100.0%
6b2100100.0%
6c00210.0%
Data preparation72100100.0%
Outcome8a2100100.0%
8b1011047.6%
8c12004.8%
Predictors9a2100100.0%
9b2100100.0%
9c318014.3%
Sample size1021909.5%
Missing data11183085.7%
Analytical12a2100100.0%
12b2100100.0%
12c2100100.0%
12d02100.0%
12e2100100.0%
12f00210.0%
12g02100.0%
Class imbalance13516023.8%
Fairness1402100.0%
Model output1502100.0%
Training vs. evaluation1602100.0%
Ethical approval17138061.9%
Open Science
Funding18a714033.3%
Conflicts18b912042.9%
Protocol18c02100.0%
Registration18d02100.0%
Data sharing18e516023.8%
Code sharing18f02100.0%
Patient & public involvement
Patient involvement1902100.0%
Results
Participants20a174081.0%
20b02100.0%
20c02100.0%
Model development212100100.0%
Model specification2202100.0%
Model performance23a138061.9%
23b00210.0%
Model updating2400210.0%
Discussion
Interpretation25129057.1%
Limitations262100100.0%
Usability27a02100.0%
27b02100.0%
27c12004.8%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Lu, L.-C.; Li, H.; Liu, Y.-M.; Du, H.-X.; Qin, J.; Yeung, W.-F.; Zhong, C.C.; Xiong, L.; Chen, S.-C. Methodological Quality and Clinical Translation of Deep Learning in Traditional Chinese Medicine Disease Diagnosis: A Systematic Review and Validation Gap Analysis. Information 2026, 17, 554. https://doi.org/10.3390/info17060554

AMA Style

Lu L-C, Li H, Liu Y-M, Du H-X, Qin J, Yeung W-F, Zhong CC, Xiong L, Chen S-C. Methodological Quality and Clinical Translation of Deep Learning in Traditional Chinese Medicine Disease Diagnosis: A Systematic Review and Validation Gap Analysis. Information. 2026; 17(6):554. https://doi.org/10.3390/info17060554

Chicago/Turabian Style

Lu, Li-Chun, Han Li, Yu-Meng Liu, Hao-Xun Du, Jing Qin, Wing-Fai Yeung, Claire Chenwen Zhong, Lei Xiong, and Shu-Cheng Chen. 2026. "Methodological Quality and Clinical Translation of Deep Learning in Traditional Chinese Medicine Disease Diagnosis: A Systematic Review and Validation Gap Analysis" Information 17, no. 6: 554. https://doi.org/10.3390/info17060554

APA Style

Lu, L.-C., Li, H., Liu, Y.-M., Du, H.-X., Qin, J., Yeung, W.-F., Zhong, C. C., Xiong, L., & Chen, S.-C. (2026). Methodological Quality and Clinical Translation of Deep Learning in Traditional Chinese Medicine Disease Diagnosis: A Systematic Review and Validation Gap Analysis. Information, 17(6), 554. https://doi.org/10.3390/info17060554

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop