Review Reports
- Swathi Ariyapadi 1,2,3,
- Emily Garavatti 2,3,4 and
- Terrie Inder 1,2,3,*
Reviewer 1: Anonymous Reviewer 2: Li Shu
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThis manuscript provides a timely and comprehensive overview of neonatal brain MRI, covering an impressive breadth of modalities from conventional T1/T2 imaging through advanced diffusion techniques and emerging artificial intelligence applications. The authors have assembled a substantial body of contemporary literature and appropriately structured the review around the principal imaging techniques and validated scoring systems. The inclusion of population-specific considerations; distinguishing between preterm and term infants, as well as special populations such as those with congenital heart disease, represents a valuable organizational framework.
However, the manuscript would benefit from significant revision to address several substantive concerns regarding its scholarly contribution, critical synthesis, and practical utility. While the review is encyclopedic in scope, it often reads as a descriptive catalog rather than a critical synthesis of the evidence. The authors have summarized what various modalities show but have not adequately grappled with why certain techniques have achieved clinical adoption while others remain investigational, nor have they meaningfully addressed the evidentiary gaps that limit clinical translation.
1. The manuscript's primary weakness lies in its failure to articulate a clear, novel scholarly contribution beyond what existing reviews already provide. Several high-quality reviews on neonatal brain MRI have been published in recent years, including Dubois et al. (2021) in Journal of Magnetic Resonance Imaging and comprehensive clinical guidelines from pediatric neurology societies. The authors cite many of these works but do not clearly distinguish what their review adds to this existing literature.
Recommendation: The Introduction should explicitly state the unique contribution of this review. Is the value in the comprehensive coverage of scoring systems? The synthesis of preterm versus term differences? The inclusion of very recent AI developments? The authors should frame their review as filling a specific gap; perhaps the integration of technical principles with clinical prognostic applications in a single resource for clinicians who are not neuroimaging specialists.
2. Throughout the manuscript, the authors present findings from individual studies with limited critical appraisal. For example, when discussing MRI scoring systems, the authors note that "all MRI scoring systems have similar predictive accuracy" (Finnegan systematic review) but do not interrogate why this might be the case or what it implies about the underlying biology. Do all systems essentially capture the same injury phenotypes? Does this suggest a ceiling effect in prediction? These are the kinds of analytical questions that would elevate the review from descriptive to scholarly.
The discussion of pseudonormalization on DWI following therapeutic hypothermia is appropriately detailed, but the clinical implications are not fully developed. The authors note that pseudonormalization is delayed to days 8-12 in cooled infants (this is clinically important!) but they do not address how this should change clinical practice regarding timing of prognostication or whether delayed imaging protocols should be adopted.
Recommendation: For each major section, include a paragraph that critically evaluates the evidence quality, identifies unresolved questions, and offers practical guidance. The "Clinical Applications" subsections should be more directive and grounded in evidence strength.
3. The review appropriately acknowledges that "mild and moderate injury" presents significant prognostic challenges, noting that some studies found similar developmental scores between infants with mild/moderate injury and those with no injury. However, this observation is not meaningfully integrated into the review's structure or recommendations. The manuscript would benefit from a dedicated section exploring why mild-moderate injury is so difficult to prognosticate; is it a measurement issue (insufficient sensitivity of current scoring systems)? A biological issue (plasticity and compensation)? A statistical issue (small effect sizes requiring larger samples)?
The authors touch on this when discussing maternal education as a modulating factor, but the insight is not developed. This is arguably one of the most clinically important challenges in the field, and it deserves more sustained attention.
4. The review covers scoring systems extensively but does not adequately address their practical implementation. How are these scoring systems actually used in clinical settings? Are there barriers to adoption (time, training, inter-rater reliability)? Do different centers use different systems, creating problems for multicenter research and clinical communication? The authors mention that MRI "remains superior to cranial ultrasonography for detecting WMI, cerebellar hemorrhage, and other abnormalities," but this statement is overly general and does not address the real-world trade-offs between the two modalities.
The discussion of TEA MRI in preterm infants appropriately acknowledges controversy but does not offer a balanced assessment of the evidence. The authors note that "some studies have reported that adding MRI to early and late cranial ultrasonography does not improve the prediction of severe intellectual disability"; this is a significant finding that warrants more than a single sentence. Why do some studies show prognostic value while others do not? Is this a methodologic issue, a population issue, or does it suggest that MRI findings are not as determinative as often assumed?
Recommendation: Add a section explicitly addressing the barriers to clinical implementation for each major modality and scoring system. Discuss evidence for clinical utility versus clinical efficacy; what works in research settings versus what is feasible and beneficial in routine practice.
5. Section 2 provides a clear and accurate exposition of T1/T2 physical principles and the characteristic appearance of the neonatal brain. The PLIC discussion is appropriately highlighted. However, the section on clinical applications is somewhat superficial. The authors mention that T1 is superior for detecting basal ganglia-thalamus injury patterns and T2 for watershed patterns, but they do not provide specific guidance on when to prioritize one sequence over another or how to interpret discordant findings.
Minor issue: The figure legend for Figure 1 describes PLIC hyperintensity as "indicating mature myelination," but this is somewhat imprecise; PLIC myelination is ongoing at term and signal intensity reflects myelin content, but complete maturity is not achieved. The wording should be more careful.
6. The technical explanation of b-values and ADC calculation is appropriate. The discussion of pseudonormalization timing with and without therapeutic hypothermia is valuable and one of the stronger clinical sections. However, the authors do not adequately address the limitations of DWI in the neonatal brain; particularly the challenges of motion artifact and the difficulty of obtaining reliable ADC measurements in small structures. The mention of "optimal timing" for preterm DWI is acknowledged as "less well known," but the section would benefit from more specific guidance based on available evidence.
7. Section 4 is the most technically sophisticated section, but it suffers from trying to cover too much ground too quickly. NODDI, CSD, and FBA are each introduced with one sentence. For a review aimed at clinicians, these advanced techniques require more explanation of what they actually measure (not just what they are called) and why they matter clinically. The authors note that these techniques "can detect aspects of brain maturation not visible by DTI alone"—this is the key point, but it needs expansion. What specific clinical questions can NODDI answer that DTI cannot?
8. The MRS section is appropriately focused on the lactate/NAA ratio and the MARBLE study findings. This is one of the few sections that offers clear quantitative prognostic information. However, the statement that "MRS can detect injury earlier than conventional MRI" should be nuanced; MRS has a specific temporal window of sensitivity, and the type of injury detectable (metabolic) differs from structural changes. The authors should clarify that MRS and DWI provide complementary information and that MRS is not necessarily "earlier" for all injury types.
9. Section 10 is the most extensive section and arguably the manuscript's strongest contribution. The coverage of term and preterm scoring systems is comprehensive, and the authors appropriately highlight the shift from simple systems (Woodward) to more comprehensive approaches (Kidokoro, Goreal). The inclusion of the Miller classification for punctate WMI and the Martinez-Biarge system is valuable, as these are less widely known than the general preterm scoring systems.
However, the section suffers from several problems:
Redundancy and organization: The Woodward and Kidokoro systems are described in separate subsections, but much of the content overlaps. The authors should consider reorganizing around themes rather than individual systems; for example, "White Matter Injury Scoring Systems" and "Comprehensive Global Scoring Systems," with the individual systems discussed within these categories.
Limited critical synthesis: The authors describe the components of each system but do not systematically compare their strengths, weaknesses, and evidence base. For example, the Kidokoro system is described as "the most widely used and comprehensive", but no evidence is provided to support this claim. When would a clinician choose one system over another? Are there situations where simpler systems are preferable?
Outcome prediction findings: The finding that "environmental factors become increasingly important as children grow older" (Jansen study) is critically important but buried. This has profound implications for how we think about the prognostic value of neonatal imaging; it suggests that early imaging can identify children at risk, but outcomes are not fixed. This should be highlighted more prominently and discussed in the context of early intervention.
The Goreal system: This is appropriately highlighted as an advance over conventional IVH grading. However, the authors do not address whether this system has been validated in independent cohorts or whether it has been adopted clinically. The distinction between research validation and clinical adoption is important.
10. Section 11 addresses the practical questions that clinicians care about: when to image, how to interpret, and how to counsel families. The content is appropriate but should be expanded. The "feed and wrap" technique is mentioned, but the section does not address how successful this approach is in practice or what the failure rate is. The discussion of MRI during therapeutic hypothermia is appropriately cautious but does not provide specific guidance on when such imaging is clinically indicated versus when it should be deferred.
The controversy regarding TEA MRI in preterm infants is acknowledged but not adequately resolved. The authors note that "obtaining routine MRI appeared to slightly reduce maternal anxiety, although it may increase the cost of care"; this is an important trade-off that deserves more attention. The authors should engage with the ethical and economic dimensions of routine screening imaging in this population.
The authors state that "specificity of MRI findings can also direct the quality and focus of in-patient and out-patient therapies"; this is an important claim!, but is there evidence to support it? This would be an ideal place to discuss whether MRI findings actually change clinical management in ways that improve outcomes, versus simply providing prognostic information.
This manuscript covers an important and rapidly evolving topic and provides a comprehensive survey of the available literature. With significant revision to strengthen critical analysis, clarify clinical translation, and address the key evidence gaps, it could make a valuable contribution to the literature. In its current form, however, the review is overly descriptive and does not sufficiently engage with the substantive challenges facing the field.
Author Response
Dear Reviewers,
Reviewer 1:
We are grateful to the extensive review that Reviewer 1 undertook and many of the helpful comments. We did wish to clarify that this review was intended for a naiver reader than the reviewer's skill set and to be an introductory review for the standard MRI sequences acquired during clinical acquisition with some additional information relevant to research applications. It was not feasible to be completely comprehensive given the space requirements and many of the thoughtful comments of the reviewer would require a more extensive review related to each of the techniques. Thus, we have attempted to respond in each domain as feasible, and all changes related to these comments are highlighted in red in the revised manuscript that is uploaded.
This manuscript provides a timely and comprehensive overview of neonatal brain MRI, covering an impressive breadth of modalities from conventional T1/T2 imaging through advanced diffusion techniques and emerging artificial intelligence applications. The authors have assembled a substantial body of contemporary literature and appropriately structured the review around the principal imaging techniques and validated scoring systems. The inclusion of population-specific considerations; distinguishing between preterm and term infants, as well as special populations such as those with congenital heart disease, represents a valuable organizational framework.
We are grateful for the recognition of our review by this Reviewer.
However, the manuscript would benefit from significant revision to address several substantive concerns regarding its scholarly contribution, critical synthesis, and practical utility. While the review is encyclopedic in scope, it often reads as a descriptive catalog rather than a critical synthesis of the evidence. The authors have summarized what various modalities show but have not adequately grappled with why certain techniques have achieved clinical adoption while others remain investigational, nor have they meaningfully addressed the evidentiary gaps that limit clinical translation.
This is a very thoughtful comment and we have addressed this with a summary statement at the end of each section to characterize the current state of the application of each technique. Its adoption - research versus clinical - can vary by site and country and thus it is important to summarize in general rather than specifics or mandate the utility.
1. The manuscript's primary weakness lies in its failure to articulate a clear, novel scholarly contribution beyond what existing reviews already provide. Several high-quality reviews on neonatal brain MRI have been published in recent years, including Dubois et al. (2021) in Journal of Magnetic Resonance Imaging and comprehensive clinical guidelines from pediatric neurology societies. The authors cite many of these works but do not clearly distinguish what their review adds to this existing literature.
Recommendation: The Introduction should explicitly state the unique contribution of this review. Is the value in the comprehensive coverage of scoring systems? The synthesis of preterm versus term differences? The inclusion of very recent AI developments? The authors should frame their review as filling a specific gap; perhaps the integration of technical principles with clinical prognostic applications in a single resource for clinicians who are not neuroimaging specialists.
This Review does differ from other reviews in that it is more recent and comprehensive in its inclusion of scoring systems and summaries. It is not typical to distinguish how it differs from other reviews as "better or more comprehensive" but rather aim to summarize knowledge that may be relevant to the reader in this domain. However, given this Reviewer's comments we have updated the Abstract with:
Abstract: Magnetic resonance imaging (MRI) has become the gold standard for evaluating brain injury and development in the newborn infant, providing structural, metabolic, and functional information without ionizing radiation. This review comprehensively examines the principal MRI modalities used in current neonatal neuroimaging for a novice/intermediate level reader —including conventional T1- and T2-weighted imaging, diffusion-weighted imaging (DWI) with apparent diffusion coefficient (ADC) mapping, diffusion tensor imaging (DTI), magnetic resonance spectroscopy (MRS), susceptibility-weighted imaging (SWI), volumetric analysis, arterial spin labeling (ASL), and functional connectivity MRI, with particular attention to their applications in preterm and term populations. Additionally, validated MRI scoring systems for quantifying brain injury severity and predicting neurodevelopmental outcomes are reviewed and summarized. Understanding the technical principles, clinical applications, and limitations of these modalities is essential for optimal interpretation of neonatal brain MRI and for advancing prognostication and therapeutic decision-making in this vulnerable population
- Throughout the manuscript, the authors present findings from individual studies with limited critical appraisal. For example, when discussing MRI scoring systems, the authors note that "all MRI scoring systems have similar predictive accuracy" (Finnegan systematic review) but do not interrogate why this might be the case or what it implies about the underlying biology. Do all systems essentially capture the same injury phenotypes? Does this suggest a ceiling effect in prediction? These are the kinds of analytical questions that would elevate the review from descriptive to scholarly.
The discussion of pseudonormalization on DWI following therapeutic hypothermia is appropriately detailed, but the clinical implications are not fully developed. The authors note that pseudonormalization is delayed to days 8-12 in cooled infants (this is clinically important!) but they do not address how this should change clinical practice regarding timing of prognostication or whether delayed imaging protocols should be adopted.
Recommendation: For each major section, include a paragraph that critically evaluates the evidence quality, identifies unresolved questions, and offers practical guidance. The "Clinical Applications" subsections should be more directive and grounded in evidence strength.
We agree with the reviewer and have both:
- a) Included a new Table with summaries of the scoring systems in preterm and term infants
- b) Included a summary of the application - clinical or research - at the end of each major section. We have not fully defined utility but rather summarized the current state of the field.
- The review appropriately acknowledges that "mild and moderate injury" presents significant prognostic challenges, noting that some studies found similar developmental scores between infants with mild/moderate injury and those with no injury. However, this observation is not meaningfully integrated into the review's structure or recommendations. The manuscript would benefit from a dedicated section exploring why mild-moderate injury is so difficult to prognosticate; is it a measurement issue (insufficient sensitivity of current scoring systems)? A biological issue (plasticity and compensation)? A statistical issue (small effect sizes requiring larger samples)?
The authors touch on this when discussing maternal education as a modulating factor, but the insight is not developed. This is arguably one of the most clinically important challenges in the field, and it deserves more sustained attention.
This is of interest but could honestly form the entire basis for a new paper of the value of mild or moderate injury in the preterm infant at term equivalent, the importance of the mediators of the relationships between structural brain injury and outcomes and indeed the detection of the nature of structural and functional alterations in the preterm brain. Although this is very worthy of thoughtful discussion, this was not the purpose of this review paper on MRI methods and their current use in neonatal populations and was beyond further expansive discussion in the scope and size of this manuscript's focus.
- The review covers scoring systems extensively but does not adequately address their practical implementation. How are these scoring systems actually used in clinical settings? Are there barriers to adoption (time, training, inter-rater reliability)? Do different centers use different systems, creating problems for multicenter research and clinical communication? The authors mention that MRI "remains superior to cranial ultrasonography for detecting WMI, cerebellar hemorrhage, and other abnormalities," but this statement is overly general and does not address the real-world trade-offs between the two modalities.
The discussion of TEA MRI in preterm infants appropriately acknowledges controversy but does not offer a balanced assessment of the evidence. The authors note that "some studies have reported that adding MRI to early and late cranial ultrasonography does not improve the prediction of severe intellectual disability"; this is a significant finding that warrants more than a single sentence. Why do some studies show prognostic value while others do not? Is this a methodologic issue, a population issue, or does it suggest that MRI findings are not as determinative as often assumed?
Recommendation: Add a section explicitly addressing the barriers to clinical implementation for each major modality and scoring system. Discuss evidence for clinical utility versus clinical efficacy; what works in research settings versus what is feasible and beneficial in routine practice.
We have included a summary of the application - clinical or research - at the end of each major section. We have not fully defined the recommendation of the utility or efficacy as that is typically determined by a body or society but rather summarized the current state of the field.
- Section 2 provides a clear and accurate exposition of T1/T2 physical principles and the characteristic appearance of the neonatal brain. The PLIC discussion is appropriately highlighted. However, the section on clinical applications is somewhat superficial. The authors mention that T1 is superior for detecting basal ganglia-thalamus injury patterns and T2 for watershed patterns, but they do not provide specific guidance on when to prioritize one sequence over another or how to interpret discordant findings.
This is very specific specialized request for radiological expertise beyond the scope of this manuscript.
Minor issue: The figure legend for Figure 1 describes PLIC hyperintensity as "indicating mature myelination," but this is somewhat imprecise; PLIC myelination is ongoing at term and signal intensity reflects myelin content, but complete maturity is not achieved. The wording should be more careful.
We have corrected this as recognized by the Reviewer with:
Figure 1. Demonstration of the posterior limb of the internal capsule (PLIC) on T1 imaging (A) with brightness indicating the presence of early myelination; conversely, hypointensity (B) represents a lack or loss of myelination. Images are original and obtained from Rady Children’s Hospital deidentified patient cases
- The technical explanation of b-values and ADC calculation is appropriate. The discussion of pseudonormalization timing with and without therapeutic hypothermia is valuable and one of the stronger clinical sections. However, the authors do not adequately address the limitations of DWI in the neonatal brain; particularly the challenges of motion artifact and the difficulty of obtaining reliable ADC measurements in small structures. The mention of "optimal timing" for preterm DWI is acknowledged as "less well known," but the section would benefit from more specific guidance based on available evidence.
We have included a summary of the application of DWI including the reference to the timing and considerations of pseudonormalization.
- Section 4 is the most technically sophisticated section, but it suffers from trying to cover too much ground too quickly. NODDI, CSD, and FBA are each introduced with one sentence. For a review aimed at clinicians, these advanced techniques require more explanation of what they actually measure (not just what they are called) and why they matter clinically. The authors note that these techniques "can detect aspects of brain maturation not visible by DTI alone"—this is the key point, but it needs expansion. What specific clinical questions can NODDI answer that DTI cannot?
This is again beyond the scope of the level of this comprehensive review as the strengths and limitations of DTI methods can be addressed by an entire review of this topic. As such it remains as summarized with a reader requiring to seek additional reference material to focus on this question.
- The MRS section is appropriately focused on the lactate/NAA ratio and the MARBLE study findings. This is one of the few sections that offers clear quantitative prognostic information. However, the statement that "MRS can detect injury earlier than conventional MRI" should be nuanced; MRS has a specific temporal window of sensitivity, and the type of injury detectable (metabolic) differs from structural changes. The authors should clarify that MRS and DWI provide complementary information and that MRS is not necessarily "earlier" for all injury types.
This has been summarized in the following:
MRS is recommended as a routine clinical component of MRI protocols for infants with neonatal encephalopathy, adding only 6–7 minutes to scan time while substantially improving prognostic accuracy.5,8-9 The lactate/NAA ratio from the basal ganglia and thalamus provides the most robust prognostic biomarker, with 88% sensitivity and 90% specificity for predicting adverse 2-year neurodevelopmental outcomes, and the combination of MRI scoring with MRS metrics outperforms MRI scoring alone.9 Clinically, MRS is particularly valuable for detecting injury alongside diffusion imaging earlier than conventional sequences to assist in early goals-of-care discussions. It can also provide additional prognostic information at later imaging timepoints. From a research perspective, MRS is being investigated as a surrogate endpoint in neuroprotection trials, and additional metabolites (glutamate/glutamine, myo-inositol) are under study for their prognostic value in both term and preterm populations.9
- Section 10 is the most extensive section and arguably the manuscript's strongest contribution. The coverage of term and preterm scoring systems is comprehensive, and the authors appropriately highlight the shift from simple systems (Woodward) to more comprehensive approaches (Kidokoro, Goreal). The inclusion of the Miller classification for punctate WMI and the Martinez-Biarge system is valuable, as these are less widely known than the general preterm scoring systems.
However, the section suffers from several problems:
Redundancy and organization: The Woodward and Kidokoro systems are described in separate subsections, but much of the content overlaps. The authors should consider reorganizing around themes rather than individual systems; for example, "White Matter Injury Scoring Systems" and "Comprehensive Global Scoring Systems," with the individual systems discussed within these categories.
We prefer the current presentation as this is the manner in whcih these scoring systems are described and reported in the literature.
Limited critical synthesis: The authors describe the components of each system but do not systematically compare their strengths, weaknesses, and evidence base. For example, the Kidokoro system is described as "the most widely used and comprehensive", but no evidence is provided to support this claim. When would a clinician choose one system over another? Are there situations where simpler systems are preferable?
Again, the systematic critique of each of these scoring systems requires its own paper rather than inclusion in a more comprehensive review such as this and is beyond the scope to expand the current section.
Outcome prediction findings: The finding that "environmental factors become increasingly important as children grow older" (Jansen study) is critically important but buried. This has profound implications for how we think about the prognostic value of neonatal imaging; it suggests that early imaging can identify children at risk, but outcomes are not fixed. This should be highlighted more prominently and discussed in the context of early intervention.
The Goreal system: This is appropriately highlighted as an advance over conventional IVH grading. However, the authors do not address whether this system has been validated in independent cohorts or whether it has been adopted clinically. The distinction between research validation and clinical adoption is important.
Again, these critique of these scoring systems would benefit from a deep and meaningful review in its own paper rather than inclusion in a more comprehensive review, such as this one. Such expansive discussion of the role of the demographic and environmental factors and the lack of clinical use of any of these scoring systems and the possible reasons for this is beyond the scope to expand the current section.
- Section 11 addresses the practical questions that clinicians care about: when to image, how to interpret, and how to counsel families. The content is appropriate but should be expanded. The "feed and wrap" technique is mentioned, but the section does not address how successful this approach is in practice or what the failure rate is. The discussion of MRI during therapeutic hypothermia is appropriately cautious but does not provide specific guidance on when such imaging is clinically indicated versus when it should be deferred.
Again these points are described for context but further expansion is beyond the scope of this introductory methodological review.
The controversy regarding TEA MRI in preterm infants is acknowledged but not adequately resolved. The authors note that "obtaining routine MRI appeared to slightly reduce maternal anxiety, although it may increase the cost of care"; this is an important trade-off that deserves more attention. The authors should engage with the ethical and economic dimensions of routine screening imaging in this population.
Again these points are described for context but further expansion is beyond the scope of this introductory methodological review.
The authors state that "specificity of MRI findings can also direct the quality and focus of in-patient and out-patient therapies"; this is an important claim, but is there evidence to support it? This would be an ideal place to discuss whether MRI findings actually change clinical management in ways that improve outcomes, versus simply providing prognostic information.
Again these points are described for context but further expansion is beyond the scope of this introductory methodological review.
This manuscript covers an important and rapidly evolving topic and provides a comprehensive survey of the available literature. With significant revision to strengthen critical analysis, clarify clinical translation, and address the key evidence gaps, it could make a valuable contribution to the literature. In its current form, however, the review is overly descriptive and does not sufficiently engage with the substantive challenges facing
We feel that this is a comprehensive and introductory review for readers on the methods and applications of MRI in neonatal populations that is robust and adds some of the areas that controversy exist making it worthy of publication as an addition to the field. We have added additional summaries as thoughtfully suggested by the reviewer. We are grateful to the reviewer for their comments and have attempted to address all of those that can be addressed in the space and scope constraints of the intentions of this review. We hope that the reviewer will appreciate this and accept that their additional requirements, although important, are beyond the scope of this format of the review.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThis review addresses an important topic in neonatal neuroimaging and covers a wide range of MRI modalities, scoring systems, clinical applications, and emerging AI/ML approaches. However, the manuscript would benefit from some major reviews as follows.
-
The manuscript is broad in scope and remains largely descriptive in several sections. It covers conventional MRI, DWI/ADC, DTI, MRS, SWI, volumetric MRI, ASL, functional connectivity MRI, scoring systems, clinical outcomes, and AI/ML applications. Much of the text explains the principles and potential uses of each modality, but the manuscript does not consistently compare which modalities have the strongest evidence for specific clinical scenarios and which remain primarily research tools. The authors should revise the manuscript to provide clearer analytical synthesis, including comparative strengths, limitations, clinical readiness, and evidence gaps across modalities.
-
The AI/ML section requires more detailed critical appraisal. The manuscript reports high performance metrics for AI/ML-based neonatal MRI interpretation and outcome prediction, including up to 83% accuracy and high correlations/R² values for Bayley score prediction. These results should be interpreted in relation to the underlying study design, sample size, training/testing strategy, external validation, multicenter generalizability, bias, interpretability, and implementation barriers. Although some limitations are briefly mentioned, the manuscript should more clearly distinguish promising research tools from clinically validated approaches and should temper conclusions.
-
The scoring systems are central to the manuscript title but are not compared in a sufficiently systematic format. The manuscript discusses scoring systems for term infants, including Barkovich, Rutherford, NICHD NRN, Weeke, and Trivedi, as well as preterm scoring systems such as Woodward, Kidokoro, Miller, Martinez-Biarge, and Goeral. However, these systems are presented mainly in separate narrative sections rather than through a unified comparative framework. A dedicated table comparing population, MRI timing, scoring domains, outcome measures, validation status, strengths, and limitations would substantially improve the clarity and usefulness of the review.
-
The source and status of the MRI images are unclear. The manuscript includes several illustrative MRI figures, but the figure legends do not state whether these are original clinical images from the authors’ institution, institutional teaching-file images, or reproduced/adapted images from previously published sources. The authors should clearly specify the origin of each image and provide appropriate permission, patient consent, or ethics/IRB clarification where applicable.
Author Response
Reviewer 2:
This review addresses an important topic in neonatal neuroimaging and covers a wide range of MRI modalities, scoring systems, clinical applications, and emerging AI/ML approaches. However, the manuscript would benefit from some major reviews as follows.
- The manuscript is broad in scope and remains largely descriptive in several sections. It covers conventional MRI, DWI/ADC, DTI, MRS, SWI, volumetric MRI, ASL, functional connectivity MRI, scoring systems, clinical outcomes, and AI/ML applications. Much of the text explains the principles and potential uses of each modality, but the manuscript does not consistently compare which modalities have the strongest evidence for specific clinical scenarios and which remain primarily research tools. The authors should revise the manuscript to provide clearer analytical synthesis, including comparative strengths, limitations, clinical readiness, and evidence gaps across modalities.
We have added summary paragraphs to each area (as shown in red) and hope that this has addressed this requirement.
- The AI/ML section requires more detailed critical appraisal. The manuscript reports high performance metrics for AI/ML-based neonatal MRI interpretation and outcome prediction, including up to 83% accuracy and high correlations/R² values for Bayley score prediction. These results should be interpreted in relation to the underlying study design, sample size, training/testing strategy, external validation, multicenter generalizability, bias, interpretability, and implementation barriers. Although some limitations are briefly mentioned, the manuscript should more clearly distinguish promising research tools from clinically validated approaches and should temper conclusions.
We have addressed this with edits in this section with:
Automated neuroprognostication using ML models incorporating MRI-based measures has shown promise in one study where it was shown to predict 18-month Bayley scores in neonates with HIE, with correlations between predicted and observed outcomes reaching 0.94 and predictive R² of 0.87 across cognitive, language, and motor domains. 53 These models may help predict outcomes across the full spectrum of injury severity, not just severe cases, but require more extensive investigation and validation in many research and clinical settings. 54
and
However, several challenges remain, including insufficient standardization of monitoring parameters, limited multicenter data sharing, and the "black-box" nature and poor clinical interpretability of AI algorithms, which hinder clinical translation. 59 The lack of transparency and regulatory neonatal frameworks in AI algorithms can be problematic when applying conclusions and interpretations to clinical care.60 Furthermore, these algorithms are particularly constrained by adequate sample size and data; given neonatal datasets are particularly limited, this can further impact accuracy and generalizability of output.61 It is unclear what role algorithmic bias may also have, given many data sets involve primarily White, North American and European populations, which can exacerbate existing health inequities.61 Implementation costs can also offer further challenges in widespread use.62
- The scoring systems are central to the manuscript title but are not compared in a sufficiently systematic format. The manuscript discusses scoring systems for term infants, including Barkovich, Rutherford, NICHD NRN, Weeke, and Trivedi, as well as preterm scoring systems such as Woodward, Kidokoro, Miller, Martinez-Biarge, and Goeral. However, these systems are presented mainly in separate narrative sections rather than through a unified comparative framework. A dedicated table comparing population, MRI timing, scoring domains, outcome measures, validation status, strengths, and limitations would substantially improve the clarity and usefulness of the review.
We have added Tables summarizing the scoring systems, their scores, the time requirement and current common uses.
- The source and status of the MRI images are unclear. The manuscript includes several illustrative MRI figures, but the figure legends do not state whether these are original clinical images from the authors’ institution, institutional teaching-file images, or reproduced/adapted images from previously published sources. The authors should clearly specify the origin of each image and provide appropriate permission, patient consent, or ethics/IRB clarification where applicable.
We have corrected the Figure legends to appropriately acknowledge all sources.
Author Response File:
Author Response.pdf
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe authors have made meaningful improvements to the manuscript, particularly through the addition of summary statements and the new tables. The manuscript now provides a more balanced view of what is clinical versus research-oriented.
Author Response
We are grateful to the reviewer for their acknowledgment of our edits and the revised manuscript now being acceptable.
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors have clarified that the illustrative MRI figures are original de-identified clinical images from Rady Children’s Hospital. However, I could not identify a corresponding ethics/IRB, consent-waiver, or institutional-permission statement in the manuscript. Please ensure that an appropriate statement is included, if not already provided elsewhere in the submission system, in accordance with the journal’s requirements.
Author Response
We are grateful to the reviewer comments and acceptance of the manuscript. With regard to the images, it is standard for all de-identified images to be able to be used in publication without direct patient consent as per standard publishing and institutional radiological guidelines. Such images are often shared in national and international databases in addition for analysis and inclusion. Thus, we specifically included "de-identified images" to allow this clarification. We are not used to any additional documentation being required in any publication. Please let us know if this journal differs from this standard radiological publishing guidelines.