1. Introduction
Thyroid nodules represent one of the most common endocrine findings in clinical practice, with a prevalence ranging from 19% to 68% when detected by high-resolution ultrasound (US), although only a small proportion prove to be malignant. Ultrasound has therefore become the first-line imaging modality for thyroid nodule assessment owing to its wide availability, non-invasiveness, and cost-effectiveness. However, despite the use of standardized reporting systems, ultrasound evaluation remains highly operator-dependent and is subject to considerable interobserver variability, particularly for features such as echogenicity, margins, and echogenic foci [
1,
2]. To address this variability and improve risk stratification, several ultrasound-based classification systems have been developed, including ACR TI-RADS, EU-TIRADS, ATA guidelines, and other national frameworks. These systems rely on visual interpretation of predefined sonographic features—composition, echogenicity, shape, margins, and calcifications—that are weighted to estimate malignancy risk and guide indications for fine-needle aspiration (FNA). Although these systems enhance standardization, multiple studies have demonstrated only fair-to-moderate agreement for several individual features, even among experienced endocrinologists, highlighting intrinsic limitations of human visual assessment [
2,
3]. In recent years, artificial intelligence (AI)-based solutions, particularly computer-aided diagnosis (CAD) and deep-learning systems, have emerged as promising tools in thyroid ultrasound. AI systems can automatically extract and analyze image features, offering objective and reproducible assessments that may reduce interobserver variability and support clinical decision-making. Multiple studies have reported high agreement and discrimination analysis of AI-assisted ultrasound for differentiating benign and malignant thyroid nodules. Moreover, AI systems have demonstrated substantial time-efficiency advantages, suggesting a potential role in high-volume clinical workflows [
4,
5]. Despite these encouraging results, important limitations remain. Many AI studies focus primarily on exploratory performance metrics, such as sensitivity, specificity, or the area under the receiver operating characteristic curve (AUC), while comparatively less attention has been paid to feature-level agreement between AI systems and clinicians. Furthermore, several high-performance AI models function as “black boxes,” which may limit clinical trust and hinder implementation in routine endocrine practice. Recent investigations have therefore emphasized the need for transparent validation studies that directly compare AI-based feature assessment with clinician interpretation using established agreement measures, such as Cohen’s kappa coefficient [
6,
7]. Additionally, while AI has shown promising accuracy for malignancy prediction, evidence regarding its ability to reliably predict cytological outcomes—particularly indeterminate categories—based on individual ultrasound features remains inconsistent. This underscores the need for cautious evaluation of AI as a supportive rather than standalone diagnostic tool, especially in clinical settings where cytology remains the standard of care [
8,
9].
Against this background, the present study aimed to compare an AI-based ultrasound system with experienced endocrinologists assessment in the evaluation of thyroid nodules, focusing on: (1) agreement for individual sonographic features, (2) concordance of size measurements, and (3) exploratory diagnostic discrimination using cytology as the reference standard. By prioritizing agreement analyses alongside clinically interpretable outcomes, this study aims to delineate the practical role of AI as a supportive tool in routine thyroid ultrasound rather than a replacement for human assessment.
Objectives
The aim of this study was to evaluate the performance of an AI-based ultrasound system in the assessment of thyroid nodules in comparison with experienced endocrinologists interpretation. Specifically, the study aimed to:
- (1)
assess the agreement between the AI system and experienced endocrinologists for individual sonographic features,
- (2)
evaluate the concordance of nodule size measurements, and
- (3)
explore the agreement and discrimination analysis of the AI system in differentiating benign and malignant nodules using cytology as the reference standard.
3. Methods
The AI analysis was performed using a commercially available software system Mindray Resona I9 (Shenzhen Mindray Bio-Medical Electronics Co., Ltd., Shenzhen, China). The system is designed for automated assessment of thyroid nodules based on ultrasound images, including feature extraction and classification according to a standardized ultrasound lexicon. The software operates as a proprietary system, and detailed information regarding model architecture and training data is not publicly available. During the study period, the system was used in a locked mode, without adaptive updates. AI analysis was performed without manual region-of-interest selection, based on a single representative image per nodule. All included cases successfully underwent AI analysis; examinations with insufficient image quality were excluded prior to inclusion in the study. All statistical analyses were performed using Python (version 3.12) with standard scientific libraries. Categorical features were harmonized to a common ultrasound lexicon based on predefined mapping rules aligning AI outputs with ACR TI-RADS terminology.
Fine-needle aspiration cytology (FNAC) served as the reference standard for exploratory diagnostic analyses. However, the primary aim of the study was to assess agreement between AI-based assessment and clinician evaluation.
Agreement between AI-based and experienced endocrinologist-based measurements of thyroid nodule dimensions was assessed using Bland–Altman analysis. For each transverse, anteroposterior, and longitudinal diameter, the mean of the paired measurements and their differences were calculated. The mean difference (bias) and the 95% limits of agreement were estimated and visualized using Bland–Altman plots. Measurement calibration was further evaluated using scatter plots comparing AI-derived and experienced endocrinologists measurements across the observed size range, with the line of identity (y = x) included to assess proportional bias and consistency. Interpretation focused on the clinical acceptability of agreement and visual assessment of concordance rather than formal hypothesis testing.
Cytological data were available for a subset of nodules and were used for exploratory analyses only. All ultrasound examinations and reference assessments were performed by four endocrinologists with more than 5 years of experience in thyroid ultrasound. Examinations were conducted as part of routine clinical practice using a standardized acquisition protocol. The term “manual reference” refers to ultrasound assessment performed by an experienced endocrinologist during routine clinical examination. Human assessments were performed independently and were not influenced by AI outputs or cytological results at the time of ultrasound examination. AI analysis was conducted retrospectively on archived images. Due to the retrospective design and reliance on routine clinical assessments, formal intra-reader and inter-reader variability were not evaluated.
AI measurements were treated as continuous scores for receiver operating characteristic (ROC) analysis, with area under the ROC curve (AUROC) estimated using trapezoidal integration and Youden-optimal operating points identified. For composition, echogenicity, shape, margin, and echogenic foci, variables were binary or recoded to binary, with solid composition, hypoechogenicity, and the presence of echogenic foci (any nonzero entry) defined as positive. AI-based size measurements proved clinically informative, with the longitudinal diameter demonstrating the strongest discriminatory performance among the evaluated size pairs while maintaining high sensitivity and specificity at the Youden-optimal threshold. Hypoechogenicity was robustly identified by the AI system, showing a well-balanced and high sensitivity and specificity profile. Echogenic foci were detected with high sensitivity and good overall discriminatory ability. In contrast, composition and shape tended to prioritize sensitivity over specificity, whereas margin assessment showed a more balanced performance. These findings suggest that greater feature subtyping—such as differentiation between punctate and macro- or rim-type calcifications—and human adjudication may help reduce false-positive classifications. Methodological considerations included the use of Wilson 95% confidence intervals for proportions, nonparametric bootstrap 95% confidence intervals for AUROC estimates, and DeLong tests for pairwise comparison of AUROCs across size measurements drawn from a shared case set. Sensitivity analyses demonstrated that the principal clinical conclusions were robust across a range of ground-truth thresholds (5–15 mm), with the longitudinal diameter consistently emerging as the best-performing axis.
Archived ultrasound datasets included both static images and cine loops. However, Doppler imaging and elastography were not systematically available and were not included in the analysis. For each nodule, the AI system analyzed a single representative image selected from the archived examination.
5. Discussion
Reported AI performance varies widely across datasets, acquisition protocols, and disease prevalence. Many published studies rely on retrospective, single-center cohorts, while head-to-head comparisons with human readers often lack standardized endpoints or rigorous prospective designs.
To address these limitations, international reporting frameworks—CONSORT-AI and SPIRIT-AI for clinical trials of AI interventions, and the recently introduced STARD-AI for exploratory performance studies involving AI—provide structured guidance to improve transparency in dataset curation, human–AI interaction, error analysis, algorithm versioning, and bias and fairness considerations. Adherence to these frameworks is increasingly recognized as essential for generating interpretable and generalizable evidence capable of informing regulatory evaluation and clinical adoption [
10,
11,
12].
This study has several limitations. First, it is based on a relatively small, single-center dataset consisting of 74 thyroid nodules. Such a limited sample size may increase the risk of overestimating agreement measures and reduce the robustness of the findings. Additionally, the single-center design may limit the generalizability of the results to other populations and clinical settings. Therefore, the findings should be interpreted with caution and considered as exploratory. Future multicenter studies with larger cohorts are needed to validate these results.
The relatively small sample size and limited availability of comprehensive cytological or histopathological reference data further restrict the ability to draw definitive conclusions regarding exploratory performance. Cytological data were used for exploratory purposes and not as a definitive diagnostic reference for all cases.
Limited transparency regarding the internal functioning of the proprietary AI system, including training data and algorithm architecture, may affect interpretability and external validity.
Against this background, the present study demonstrates high agreement between AI-based and manual endocrinologist measurements of thyroid nodule dimensions, with ICC (2,1) values exceeding 0.80 across all axes. A small but statistically significant underestimation of the anteroposterior (AP) diameter by AI—on the order of approximately 1 mm—was consistently identified using Bland–Altman and calibration analyses. In contrast, no significant differences were observed for transverse or longitudinal measurements after correction for multiple testing. These results support the use of AI as a reliable assistant for quantitative nodule sizing, while underscoring that even modest, axis-specific systematic deviations merit attention when clinical decisions hinge on size-based thresholds, such as TI-RADS biopsy cut-offs. This interpretation is concordant with prior work showing that deep-learning systems can match experienced endocrinologists performance in thyroid nodule risk stratification and may enhance measurement standardization when terminology is harmonized and models are appropriately calibrated [
13,
14].
Calibration analyses further revealed an approximately linear relationship between AI-derived and manual measurements across the observed size range, closely tracking the line of perfect agreement. Importantly, this calibration pattern remained qualitatively stable in stratified analyses by sex, thyroid lobe, and nodule size category, suggesting consistent algorithm behavior across clinically relevant subgroups. Nonetheless, limited sample sizes in the smallest and largest nodules constrain the precision of these subgroup estimates and preclude definitive conclusions at the extremes of the size spectrum.
For categorical ultrasound features, inter-method agreement was limited overall, with Cohen’s kappa values indicating only slight-to-fair to fair agreement depending on the feature. Although percent agreement was relatively high for some features (e.g., shape), the corresponding kappa values were low, likely reflecting a prevalence effect and limited agreement beyond chance. These findings underscore that percent agreement alone may overestimate true concordance and that categorical interpretation remains challenging for AI systems. In particular, agreement for clinically important TI-RADS features such as echogenicity, margins, and shape was suboptimal, which may limit direct clinical applicability. While normalization to a standardized ultrasound lexicon improved comparability between AI and human assessments, the overall level of agreement remained modest. This suggests that standardization alone is insufficient and that further methodological refinement is required before categorical AI outputs can be reliably integrated into risk stratification frameworks.
These results align with the broader literature. A recent meta-analysis reported pooled AI sensitivity of 0.86 and specificity of 0.78 (AUC 0.89), not inferior to r experienced endocrinologists (AUC 0.91) [
13]. In a prospective real-world study, AI achieved agreement and discrimination analysis comparable to that of senior experienced endocrinologists and significantly improved junior readers’ accuracy, particularly for nodules ≤ 1.5 cm, while reducing potentially unnecessary biopsies in larger nodules—an effect with direct implications for clinical workflow and patient management [
14]. Independent validation studies across ultrasound vendors further demonstrate that device-specific differences influence absolute performance for both AI systems and human readers, highlighting the importance of external validation and ongoing vigilance for domain shift [
15].
Technical reviews describe a maturing AI pipeline evolving from detection and segmentation to comprehensive classification, with segmentation-based approaches—such as U-Net, TransUNet, or CNN–ViT hybrids—supporting more reproducible dimensional measurements that are directly relevant to TI-RADS decision thresholds [
16,
17]. Implementation-focused analyses further emphasize that FDA-cleared products vary in how they integrate with TI-RADS workflows and stress the necessity of careful deployment, continuous performance monitoring, and strict adherence to ACR lexicon standards [
18]. Global epidemiological analyses from GLOBOCAN highlight a high incidence but low mortality of thyroid cancer, with this divergence largely attributed to overdiagnosis [
19,
20]. Within this context, AI systems should aim not merely to maximize sensitivity, but to help rebalance sensitivity and specificity in a manner that reduces unnecessary procedures without compromising oncological safety.
Analysis of measurement discrepancies suggests that larger differences between AI and human assessments may arise in more challenging imaging conditions. Based on visual inspection of outlier cases in Bland–Altman plots, discrepancies were most frequently associated with heterogeneous nodules, partially cystic composition, indistinct or irregular margins, and reduced image quality. Additional contributing factors may include posterior acoustic shadowing and difficulties in defining precise lesion boundaries, particularly in nodules with complex internal architecture. These observations highlight typical scenarios in which AI-based segmentation and measurement may be less reliable and reinforce the importance of maintaining clinician oversight, especially in technically challenging cases. Future studies incorporating systematic segmentation-level evaluation and representative imaging examples would be valuable for a more detailed characterization of AI error patterns.
Clinical Implications
- (1)
Quantitative measurements: The high ICC values support incorporating AI into routine nodule sizing, with potential benefits in examination efficiency and reproducibility—particularly where maximum diameter determines TI-RADS-based recommendations for FNA or follow-up. The modest AP bias (~1 mm) suggests that simple calibration offsets or context-aware alerts at decision thresholds could mitigate clinically relevant effects [
21].
- (2)
Qualitative features: After normalization to the ACR TI-RADS lexicon, agreement improves for features such as echogenic foci, supporting the implementation of explicit terminology mapping (e.g., punctate echogenic foci to microcalcifications) within AI user interfaces and structured reporting templates [
21].
- (3)
Deployment and oversight: Real-world evidence indicates susceptibility to vendor- and device-related domain shifts, underscoring the need for local validation before full clinical deployment and for continuous quality assurance of both measurements and categorical classifications [
15,
18].
Importantly, thyroid nodule management is based on integrated risk stratification rather than isolated ultrasound features. Therefore, discrepancies observed between AI and human assessment at the feature level—particularly for echogenicity and margins—may have direct clinical consequences through their influence on TI-RADS scoring. Even small differences in feature classification can lead to changes in TI-RADS category assignment, potentially affecting recommendations for fine-needle aspiration (FNA), surveillance intervals, and patient counseling.
In this context, the fair-to-moderate agreement observed for selected categorical features should be interpreted cautiously. Misclassification by AI may theoretically result in both over-biopsy, due to overestimation of suspicious characteristics, and under-biopsy, due to failure to detect high-risk features. These findings reinforce that AI systems should currently be regarded as supportive tools rather than standalone decision-making systems, particularly in cases where management decisions depend on subtle qualitative features.
Our results further underscore the importance of maintaining clinician oversight and integrating AI outputs within a broader clinical framework that includes patient risk factors, longitudinal follow-up, and cytological assessment where indicated. While the present study was not powered for definitive analysis of downstream clinical decisions, these considerations highlight the need for future research evaluating the impact of AI-assisted ultrasound on biopsy rates, follow-up strategies, and patient outcomes, ideally with stratification according to Bethesda categories.
Key strengths of this study include paired AI–human comparisons on identical cases and the combined use of agreement, calibration, and discriminatory analyses, extending beyond classification metrics alone. Limitations include potential domain shift related to probe type, acquisition settings, and disease prevalence, which may affect generalizability. In addition, the lack of external validation limits the generalizability of the findings across different clinical settings and ultrasound systems.
Additionally, small sample sizes in extreme nodule-size strata limit estimate precision and regression stability, as noted in prior prospective and implementation studies [
14,
15,
18].
Given the persistent heterogeneity of study designs and under-specified human–AI interactions in this field, rigorous application of CONSORT-AI and SPIRIT-AI for interventional studies and STARD-AI for exploratory performance research remains essential. Explicit reporting of algorithm versioning, input data handling, external validation procedures, error and bias analyses, and fairness considerations will be critical to enable assessment of quality, transferability, and clinical readiness by both readers and regulators [
10,
11,
12].
Future research should prioritize prospective, multicentre investigations with external validation across ultrasound platforms, reported in accordance with STARD-AI. Such studies should quantify downstream clinical impact—including changes in biopsy rates, reporting time, cost-effectiveness, and equity-aware performance—and incorporate TI-RADS-aligned deployment strategies with adaptive calibration and active performance surveillance [
12,
18].
Additionally, the absence of formal intra- and inter-reader variability assessment limits conclusions regarding reproducibility across different observers.