Review Reports
- Komei Iwai 1,
- Tetsuji Azuma 1 and
- Takaaki Tomofuji 1,*
- et al.
Reviewer 1: Anonymous Reviewer 2: Anonymous Reviewer 3: Jianchen Yang
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe manuscript tackles a clinically relevant and timely problem by examining whether routinely collected administrative health records, when augmented with dental checkup data, can improve the prediction of incident dementia among older adults. The cohort size and record linkage strategy are clear strengths, and the observed increases in sensitivity and specificity after incorporating dental variables indicate that oral health–related information may add nontrivial signals for dementia risk stratification. This premise is potentially valuable and merits careful evaluation.
That said, the current presentation raises substantial concerns regarding how model performance is defined and interpreted. The manuscript describes an “automatic” accuracy calculation based on agreement between two independently trained models, rather than on concordance between predictions and observed clinical outcomes. This approach does not reflect predictive accuracy in the conventional sense and may overstate performance by capturing model stability or internal consistency rather than true discrimination of dementia onset. Given that several conclusions in the Discussion are anchored to these so-called accuracy benchmarks, the reliance on this internal agreement metric is problematic. At minimum, it should be clearly distinguished from outcome-based performance measures, and preferably replaced by standard evaluation metrics derived directly from labeled outcomes.
Related issues arise in the reporting of sensitivity and specificity. The validation cohort exhibits a low event rate, with incident dementia occurring in only 6% of participants over two years. In this context, incremental gains in sensitivity are difficult to interpret without additional information. Metrics such as positive and negative predictive values, precision–recall curves, and justification of decision thresholds are notably absent. Confidence intervals for all reported performance measures are also needed to convey uncertainty. Without these elements, it remains unclear whether the model would be clinically useful for screening or prioritization, particularly since sensitivity remains below 0.80 even after dental variables are included.
The manuscript’s treatment of feature contribution further complicates interpretation. Feature importance is described in terms of performance degradation after removing individual variables, yet the underlying performance metric again relies on inter-model agreement rather than outcome-based accuracy. This disconnect undermines the interpretability claims. A more widely accepted explanation framework should be adopted, and the reported contributions should be explicitly tied to changes in clinically meaningful performance measures. Moreover, several highly ranked predictors—such as care-need certification status and pneumonia—appear to reflect global frailty or preclinical decline rather than oral health per se. This raises the possibility that the model is capturing downstream vulnerability signals, a point that warrants explicit discussion to avoid overstating causal implications.
From a reproducibility perspective, the dependence on a proprietary automated modeling platform introduces additional opacity. Key details remain insufficiently specified, including the precise model configuration, software versioning, preprocessing steps applied to both NDB and dental datasets, handling of missing values, and any hyperparameter optimization procedures. While the Methods section outlines data splits and basic statistical tests, it does not provide enough information for an independent group to replicate the modeling pipeline. This limitation is particularly relevant given that the reported performance gains are relatively modest in absolute terms.
Finally, the manuscript would benefit from a clearer and more quantitative demonstration of the incremental value of dental checkup data beyond administrative records alone. Comparative analyses explicitly testing added predictive value, along with a discussion of plausible real-world deployment scenarios, would strengthen the contribution. Clarifying whether the intended use case is early rule-out screening, targeted preventive intervention, or resource allocation would help contextualize the reported operating characteristics. As currently presented, the findings are intriguing but remain preliminary, and stronger methodological grounding is required to support the level of interpretation offered in the Discussion.
Author Response
"Please see the attachment."
Author Response File:
Author Response.docx
Reviewer 2 Report
Comments and Suggestions for AuthorsThe paper presents a retrospective cohort study investigating the accuracy of an artificial intelligence model in predicting the development of dementia over a two-year period in adults aged ≥75 years. The model integrates data from the Japanese National Health Insurance Database with information on dental visits and is implemented using a commercial automated AI platform. The authors demonstrate that the integration of dental health variables moderately improves the accuracy, sensitivity, and specificity of the prediction compared to models trained exclusively on medical data. Although the dataset is large, the manuscript remains predictive rather than mechanistic. Therefore, several issues need to be addressed to strengthen scientific rigour and generalisability.
1. Why did the authors not perform any analysis of sample size or statistical power, and how might this omission affect the reliability of the reported performance metrics?
2. The improvement in accuracy from 77.3% to 79.0% is modest. How do the authors assess the clinical significance, and not just the statistical improvement, of this increase?
3. How do the authors justify the clinical interpretability of the model without explainable AI techniques? For this reason, the authors should supplement the paper with the study doi: 10.3390/computers14090344 needed to strengthen the discussion on the interpretability and reliability of the model, particularly in healthcare decision-making.
4. Sensitivity remains below 80%, limiting its usefulness for early diagnosis. How do the authors envision the use of this model in real clinical or public health workflows?
5. The contribution analysis is proprietary and lacks quantification of uncertainty. How can clinicians assess the robustness of the reported predictor importance rankings?
6. The study lacks external validation on an independent cohort. How do the authors justify the generalisability of the model beyond the dataset of a single prefecture? Recent studies on FEM-based modelling and AI-enhanced monitoring systems demonstrate the existence of structured validation pipelines that could provide useful insights for improving the current methodology and should be discussed in the Methods or Limitations sections.
7. The manuscript frames its contribution narrowly around prediction accuracy. How do the authors position their study in the broader context of AI-integrated biomedical systems, where optimisation, system-level modelling, and hybrid numerical-AI approaches are increasingly emphasised?
8. Why were confidence intervals for sensitivity, specificity, and accuracy not reported, and how does this omission affect reproducibility?
9. Given the great influence of care-need certification, how do the authors address data leakage or circularity between predictors and outcomes?
10. Dental variables appear among the main predictors. How do the authors distinguish correlation from causation, and how should these results be interpreted clinically?
Author Response
"Please see the attachment."
Author Response File:
Author Response.docx
Reviewer 3 Report
Comments and Suggestions for AuthorsPlease find more details in the comments attached.
Comments for author File:
Comments.pdf
Author Response
"Please see the attachment."
Author Response File:
Author Response.docx
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe revised manuscript is improved in that it more clearly separates the software-reported internal metric from outcome-based evaluation on the validation cohort, and it now includes clinically interpretable quantities such as PPV/NPV in addition to sensitivity and specificity. This clarification reduces a major source of confusion from the previous version and moves the paper in a more appropriate direction.
However, several major issues remain that limit both interpretability and the credibility of the conclusions. The paper still does not define a concrete operating point for decision-making under a low event rate setting. With incident dementia occurring in only about 6% of the validation cohort, small changes in sensitivity and specificity are not meaningful without a clearly stated decision threshold (or an explicit threshold selection procedure) and complementary evaluation that is standard in imbalanced clinical prediction, such as precision–recall analysis and calibration assessment. While the manuscript mentions that thresholds can be adjusted, it does not specify which threshold produced the reported headline results, nor does it justify that choice in relation to a realistic use case (e.g., rule-out screening versus risk prioritization). As a result, the reported gains remain difficult to translate into clinical utility.
The contribution and novelty of the work also need to be articulated more explicitly. The manuscript lists many authors and appears to leverage a proprietary automated modeling platform, yet it is not clear what the unique methodological contribution is beyond applying an existing tool to linked administrative and dental records. If the main contribution is the dataset linkage and the demonstration of incremental predictive value from dental variables, that should be stated plainly and supported with a rigorous added-value analysis. If the contribution is methodological, then the modeling and validation strategy must be described in enough detail to demonstrate that the approach is not a straightforward application of standard pipelines. In its current form, the work reads as relatively simple and descriptive, and it is difficult to see what is new for the community beyond the setting itself.
Relatedly, the “feature contribution” analysis remains hard to interpret. Although the manuscript now warns against causal interpretation and clarifies that contribution is measured via changes in an internal software-reported metric, this weakens the interpretability narrative rather than strengthening it. The explanation section would be substantially more convincing if it used a standard, outcome-linked interpretability framework on the validation set and reported how feature removal/perturbation affects outcome-based metrics directly. Moreover, several highly ranked predictors appear to reflect broader frailty or preclinical decline rather than oral health mechanisms, which should be discussed more explicitly to avoid over-interpretation.
Finally, the presentation and communication of the study are not yet at a publishable standard. The current manuscript contains no figures, which is a serious limitation for a paper that claims clinical relevance and aims to communicate model performance, operating characteristics, and variable effects. At minimum, the authors should include clear visual summaries of the study design and evaluation, such as a cohort flow diagram, a schematic of the modeling pipeline, and one or more figures that convey threshold-dependent performance (e.g., precision–recall behavior), calibration, and the stability of key findings. Without such visual evidence, the paper reads as incomplete and makes it harder for readers to assess the robustness and practical implications of the results.
Overall, the premise remains potentially valuable, and the revision addresses part of the earlier confusion, but additional methodological grounding, clearer articulation of the core contribution, stronger outcome-linked evaluation under class imbalance, and a more complete presentation are required before the conclusions can be considered adequately supported.
Author Response
Please see the attachment.
Author Response File:
Author Response.docx
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors have given good answers to the review comments and have revised the paper well.
Author Response
We would like to express our sincere appreciation for your very useful advice. We look forward to working with you in the future.
Round 3
Reviewer 1 Report
Comments and Suggestions for AuthorsThe revised manuscript has adequately addressed the previous comments… I recommend acceptance in present form.