A Systematic Review of the Psychometric Quality of Instruments for Assessing Adverse Childhood Experiences
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThis manuscript addresses an important and highly relevant topic by systematically reviewing the psychometric quality of instruments used to assess adverse childhood experiences (ACEs). The application of the PRISMA framework for study identification and selection, together with the COSMIN guidelines for evaluating psychometric properties, represents a major strength of the study. The review covers the principal ACE instruments currently available and provides a useful synthesis of their psychometric characteristics, making the manuscript potentially valuable for researchers, clinicians, and public health professionals.
Despite these strengths, several aspects should be improved before the manuscript is suitable for publication.
First, the Introduction would benefit from a stronger justification of the review. Although the importance of ACEs and their impact on health is well established, the authors should more explicitly identify the existing gap in the literature that this review intends to address. In particular, it would be helpful to compare the present review with previous systematic reviews or psychometric overviews and explain more clearly why a new review based on the updated COSMIN methodology is needed.
The Methods section requires greater methodological transparency. While the overall review design is appropriate, important details necessary for reproducibility are missing. The complete electronic search strategies for each database should be reported, either within the manuscript or as supplementary material. It is also recommended that the authors specify whether the review protocol was preregistered (e.g., PROSPERO). If no protocol was registered, this should be acknowledged as a limitation.
Furthermore, the manuscript should provide a more detailed description of the data extraction process, including how disagreements between reviewers were resolved and whether standardized extraction forms were used. Although the screening process is described, the procedures for extracting psychometric data and assigning COSMIN ratings remain insufficiently detailed.
A major methodological limitation is the absence of a formal assessment of the risk of bias of the included studies. Although COSMIN evaluates methodological quality related to measurement properties, it does not fully replace an assessment of study-level bias. The authors should clarify how potential sources of bias within the included validation studies may have influenced their conclusions or justify why a separate risk-of-bias assessment was not performed.
The Results section is comprehensive, but several tables could be presented more clearly. Tables 3, 4, and especially Table 5 contain numerous abbreviations, symbols, and percentages that make interpretation difficult. The calculation of the "strength of evidence" percentages should be explained explicitly in the Methods, and the presentation of these tables could be simplified to improve readability.
The Discussion accurately summarizes the findings but remains largely descriptive. The manuscript would benefit from a more critical interpretation of the observed differences between instruments. For example, the authors could discuss why some psychometric properties, such as responsiveness and measurement error, have received little attention in validation studies and what implications this has for clinical practice and future research. Similarly, greater emphasis could be placed on the practical implications of selecting one ACE instrument over another in different research or clinical settings.
The Conclusions are generally supported by the findings but should be expressed more cautiously. Statements suggesting that ACE instruments are generally "useful" or "appropriate" should acknowledge that the quality of psychometric evidence varies substantially among instruments and across psychometric domains. Instruments supported by stronger evidence should be distinguished from those requiring additional validation.
The manuscript would also benefit from a more comprehensive discussion of cross-cultural validation and measurement invariance. Since ACE instruments are increasingly used internationally, greater attention should be paid to cultural adaptation procedures and the comparability of scores across different populations.
Although the overall quality of English is acceptable, the manuscript requires careful language editing. Several sentences are unnecessarily long, certain expressions (e.g., "With regard to," "As regards," "Furthermore") are overused, and there are minor grammatical and typographical errors (e.g., "blind-ed", "interceded") that should be corrected to improve readability and stylistic consistency.
Finally, the authors may wish to consider including a brief practical summary or recommendation table highlighting the main strengths, limitations, and recommended applications of each ACE instrument. Such a synthesis would considerably increase the practical value of the review for researchers and clinicians seeking guidance in instrument selection
Comments on the Quality of English LanguageThe manuscript is generally written in clear and understandable English, and the scientific terminology is appropriate for the subject matter. The overall organization of the text facilitates comprehension, and the manuscript communicates its objectives and findings effectively.
However, the language would benefit from careful professional editing before publication. Several sentences are unnecessarily long and could be divided into shorter, more concise statements to improve readability. There is also frequent repetition of introductory phrases such as "With regard to," "As regards," "Furthermore," and "Overall," which affects the fluency and stylistic variety of the text.
In addition, a number of minor grammatical, typographical, and formatting errors should be corrected. Examples include expressions such as "blind-ed" and "in-terceded", inconsistent punctuation and spacing, and occasional awkward sentence constructions. Some paragraphs, particularly in the Results and Discussion sections, are overly descriptive and repetitive, and could be streamlined without compromising scientific content.
Author Response
The authors of the manuscript entitled ‘A Systematic Review of the Psychometric Quality of Instruments for Assessing Adverse Childhood Experiences’ would like to thank the reviewers for their comments. The changes made in response to these comments are outlined below.
Reviewer 2
Comments and Suggestions for Authors
This manuscript addresses an important and highly relevant topic by systematically reviewing the psychometric quality of instruments used to assess adverse childhood experiences (ACEs). The application of the PRISMA framework for study identification and selection, together with the COSMIN guidelines for evaluating psychometric properties, represents a major strength of the study. The review covers the principal ACE instruments currently available and provides a useful synthesis of their psychometric characteristics, making the manuscript potentially valuable for researchers, clinicians, and public health professionals.
Despite these strengths, several aspects should be improved before the manuscript is suitable for publication.
COMMENT: First, the Introduction would benefit from a stronger justification of the review. Although the importance of ACEs and their impact on health is well established, the authors should more explicitly identify the existing gap in the literature that this review intends to address. In particular, it would be helpful to compare the present review with previous systematic reviews or psychometric overviews and explain more clearly why a new review based on the updated COSMIN methodology is needed.
ANSWER: Whilst there are various review studies on adverse childhood experiences, none of them focus on the psychometric properties of the instruments used to assess them. This was confirmed because, prior to this review, a search was carried out in databases and on Prospero to check for the existence of a similar review. As none was found, it was decided to undertake this study. Furthermore, another reason for carrying out this study was the research focus of two of the authors in this area, seeking to identify instruments that are more widely used internationally and possess better psychometric properties. However, the following paragraph has been included in the introduction:
Since Fellitti and Kaiser introduced the term ‘Adverse Childhood Experiences’ in 1998, research in this field has grown steadily over the years, focusing primarily on the impact these experiences have on people’s mental health. Over the years, the concept has evolved to include other experiences that were not initially considered, leading to the development of new instruments. In this context, and following a review of the literature in various databases and in Prospero, a lack of studies was identified that comprehensively examine the psychometric properties of these instruments; this is why it was decided to carry out the present review.
COMMENT: The Methods section requires greater methodological transparency. While the overall review design is appropriate, important details necessary for reproducibility are missing. The complete electronic search strategies for each database should be reported, either within the manuscript or as supplementary material. It is also recommended that the authors specify whether the review protocol was preregistered (e.g., PROSPERO). If no protocol was registered, this should be acknowledged as a limitation.
ANSWER: The following paragraph has been included in the discussion section: Another limitation relates to the fact that the review was not registered with PROSPERO, although prior to its commencement, a check was carried out to ensure that there were no similar reviews in the included databases or in PROSPERO itself.
The search query was similar across all databases. The only difference was the field used. For example, in WoS the ‘topic’ field was used, whilst in ProQuest the terms were enclosed in quotation marks, selecting ‘advanced search’, ‘scientific journals’ and ‘peer-reviewed articles’.
COMMENT: Furthermore, the manuscript should provide a more detailed description of the data extraction process, including how disagreements between reviewers were resolved and whether standardized extraction forms were used. Although the screening process is described, the procedures for extracting psychometric data and assigning COSMIN ratings remain insufficiently detailed.
ANSWER: In response to the reviewer’s comment, the following paragraph has been included: The data relating to the psychometric properties reported in the analysed articles were collated by two researchers (FG-S and LC); where there was disagreement, a third reviewer intervened (MI-C). This information was transferred to the results tables, which summarised the characteristics of the samples (see Table 2) and the psychometric properties listed in the COSMIN checklist (see Table 3), namely: structural validity, internal consistency, cross-cultural validity/measurement invariance, reliability, measurement error, criterion validity, hypothesis testing for construct validity and responsiveness. All items comprising the checklist were rated on a 5 point scale (excellent, good, fair, poor and unknown/NA). The overall score for each box was based on the lowest score (1–5) for an item within that box [21,22]. When assessing the quality of each instrument, the GRADE criteria were followed, with each property rated on a three-point scale: sufficient (+), uncertain (?) and insufficient (-) (see Table 4), as specified in the COSMIN manual (Mokkink) and in the guidelines provided by Prinsen et al. [23]and Terwee et al. [24]. Finally, the quality of evidence for each instrument was summarised in a table (see Table 5), as suggested by Prinsen et al. [22], distinguishing between high-quality evidence when several articles demonstrated good methodology or one article was of excellent quality; moderate-quality evidence when there were several articles with acceptable methodology or one of good quality; limited-quality evidence when the quality was acceptable; and conflicting evidence when the articles were of low quality. In cases where studies lacked information on psychometric properties, the evidence was assessed as conflicting.
COMMENT: A major methodological limitation is the absence of a formal assessment of the risk of bias of the included studies. Although COSMIN evaluates methodological quality related to measurement properties, it does not fully replace an assessment of study-level bias. The authors should clarify how potential sources of bias within the included validation studies may have influenced their conclusions or justify why a separate risk-of-bias assessment was not performed.
ANSWER: The following paragraph has been included amongst the limitations of this study: Another limitation of this review is that content validity has not been assessed. When assessing which instruments exhibit the best psychometric properties, content validity is one of the most important measurement properties in Patient-Reported Outcome Measures (PROMs) [49]. This has an impact on clinical practice [50]. In the case of ACEs, it is important to bear in mind the evolution of this dimension, taking into account not only the now-classic events but also other contextual and cultural factors, as well as protective factors [4,15,16], which may affect content validity. In this regard, amongst the instruments analysed in this review, it should be noted that the ACE-IQ, the SC-ACE-IQ and the ACE-THL are the ones that include the greatest number of items relating, in some cases, to stressors and, in others, to protective factors. Furthermore, not all the studies analysed provided details on how the instrument had been developed. Consequently, content validity could not be assessed
COMMENT: The Results section is comprehensive, but several tables could be presented more clearly. Tables 3, 4, and especially Table 5 contain numerous abbreviations, symbols, and percentages that make interpretation difficult. The calculation of the "strength of evidence" percentages should be explained explicitly in the Methods, and the presentation of these tables could be simplified to improve readability.
ANSWER: As the reviewer suggests, these points have been addressed in the methodology section. As regards the tables, some of them have been simplified.
Quality of the PROM was assessed with the updated criteria for good measurement properties shown in the COSMIN manual [21] and based on Prinsen et al. [23] and Terwee et al. [24] indications. These criteria were evaluated in a threepointscale: sufficient (+), insufficient (-), and indeterminate (?). Strength of evidence assessment was performed based on the modified GRADE (Grades of Recommendation, Assessment, Development and Evaluation) approach, which determines an overall assessment of each property ranging from strong evidence, to moderate, limited, conflicting, and unknown; based on the methodological quality and consistency of results for each study [22]. Strong evidence was reported when there were several methodologically good articles or one of excellent quality, moderate evidence was indicated when there are several methodologically fair articles or one of good quality, limited evidence was considered with articles with fair quality, and conflicting evidence was assigned to low quality articles. Unknown evidence was assigned to studies in which information regarding psychometric properties was missing.
In the tables, it has been decided to retain the abbreviations with their corresponding key, following the example set by other articles and as set out in the COSMIN application manual.
COMMENT: The Discussion accurately summarizes the findings but remains largely descriptive. The manuscript would benefit from a more critical interpretation of the observed differences between instruments. For example, the authors could discuss why some psychometric properties, such as responsiveness and measurement error, have received little attention in validation studies and what implications this has for clinical practice and future research. Similarly, greater emphasis could be placed on the practical implications of selecting one ACE instrument over another in different research or clinical settings.
The manuscript would also benefit from a more comprehensive discussion of cross-cultural validation and measurement invariance. Since ACE instruments are increasingly used internationally, greater attention should be paid to cultural adaptation procedures and the comparability of scores across different populations
ANSWER: As suggested by the reviewer, the discussion has been revised to include the points raised by the reviewer. Specifically, the following paragraphs have been included:
Measurement error has been another of the properties that has been assessed to a lesser extent, specifically in 13 of the 20 studies analysed. This may be linked to memory biases, given that instruments used to assess ACEs do so retrospectively; the presence or absence of an event may therefore be affected if the individual denies or downplays it. Furthermore, other biases associated with self-reports, gender, and social or cultural groups may also contribute to this lack of measurement; this would require assessing whether each item within an instrument behaves in the same way across the different groups being assessed. In this regard, it is important to bear in mind the normalisation of certain behaviours that may occur in different groups, such as the use of violence or phys-ical punishment in the upbringing of children, or abandonment or neglect in care that may occur simply because someone is a man or a woman. All these factors also complicate the assessment of criterion validity, in addition to other issues associated with this type of in-strument, which provide a total score based on the number of incidents the individual has experienced, without weighting the score according to the severity of the incident, the age at which it occurred, or whether the incident or incidents became chronic. However, some retrospective epidemiological studies [44,45] highlight the usefulness of this type of in-strument despite the difficulties associated with memory biases.
As mentioned above, cultural biases can, to a certain extent, influence cross-cultural validity. The use of these instruments across different cultures involves more than simply translating the items, given that terms such as ‘abuse’, ‘mistreatment’ or ‘neglect’ can vary from one culture to another; in some cases, certain practices may even be normalised, whereas in other contexts they would be regarded as forms of mistreatment. Existing val-ues and prejudices towards certain communities, minorities or specific social groups, as well as the laws and customs specific to each community, make it difficult to draw com-parisons between cultures and, consequently, to study cross-cultural validity.
COMMENT: The Conclusions are generally supported by the findings but should be expressed more cautiously. Statements suggesting that ACE instruments are generally "useful" or "appropriate" should acknowledge that the quality of psychometric evidence varies substantially among instruments and across psychometric domains. Instruments supported by stronger evidence should be distinguished from those requiring additional validation.
ANSWER: The conclusion has been amended and the cultural adaptation of the instruments has been discussed in greater depth in the discussion section.
From a statistical and psychometric perspective, those instruments that demonstrate greater validity and reliability may be superior to those that score lower on these con-structs. Furthermore, it is important to ensure that these psychometric properties are maintained across different cultures, contexts and samples. It can be concluded that, amongst the instruments analysed in this review, the one with the greatest number of studies evaluating its psychometric properties most of which demonstrate good structural validity and internal consistency, whilst also having been evaluated in a wider range of cultural contexts with good cross-cultural validity is the World Health Organisation’s ACE-IQ [1].
In light of the results obtained, the ACE-IQ may be the most suitable tool for conduct-ing epidemiological studies; however, given the dynamic nature of adverse experiences in childhood, it would be advisable to review the instruments with a view to improving their content validity and ensuring that they are genuinely adapted to the norms, customs and traditions of each culture.
COMMENT: Although the overall quality of English is acceptable, the manuscript requires careful language editing. Several sentences are unnecessarily long, certain expressions (e.g., "With regard to," "As regards," "Furthermore") are overused, and there are minor grammatical and typographical errors (e.g., "blind-ed", "interceded") that should be corrected to improve readability and stylistic consistency.
Finally, the authors may wish to consider including a brief practical summary or recommendation table highlighting the main strengths, limitations, and recommended applications of each ACE instrument. Such a synthesis would considerably increase the practical value of the review for researchers and clinicians seeking guidance in instrument selection.
ANSWER: The English has been revised as suggested by the reviewer. Finally, in the discussion section, we have chosen to specify which instrument may be the most suitable, taking into account the psychometric properties assessed, their psychometric quality and the empirical evidence.
Reviewer 2 Report
Comments and Suggestions for AuthorsThe manuscript presents a systematic review of the psychometric quality of instruments used to assess adverse childhood experiences (ACEs), applying the COSMIN framework. It reviews 20 studies covering instruments such as ACE-10, ACE-Q, ACE-SQ, ACE-IQ, ACE-IQ-10, SC-ACE-IQ and ACE-THL, and examines properties including structural validity, internal consistency, cross-cultural validity, reliability, measurement error, criterion validity, construct validity and responsiveness. Its central conclusion is that ACE instruments are widely used and often show acceptable internal consistency and structural validity, but the evidence remains uneven, especially for responsiveness, measurement error, criterion validity and cross-cultural validation.
The topic is useful, but the manuscript remains largely descriptive. It does not yet fully exploit the methodological framework offered by COSMIN, and several important conceptual and methodological issues require clarification before the review can be considered a reliable guide for researchers or clinicians.
(1) The manuscript would benefit from a clearer conceptual distinction between adverse childhood experiences (ACEs) as a public health construct and child maltreatment as a clinical or legal construct. At present these concepts are introduced almost interchangeably, despite representing overlapping but non-equivalent frameworks. Since one of the review’s objectives is to evaluate measurement instruments, the conceptual target being measured should first be explicitly defined. Otherwise, readers cannot determine whether differences between instruments reflect psychometric superiority or differences in construct coverage.
(2) The review repeatedly concludes that certain instruments have “better psychometric quality”, yet the comparison is made across studies with markedly different samples, designs, statistical methods and validation objectives. COSMIN evaluates the quality of evidence within a validation study rather than providing a direct ranking between instruments. Consequently, statements implying superiority of one questionnaire over another should either be substantially qualified or supported by a formal evidence synthesis strategy specifically designed for between-instrument comparison.
(3) The review focuses almost exclusively on classical psychometric properties while giving little attention to the underlying measurement model adopted by each instrument. Several ACE questionnaires were not developed under the assumption that adversity represents a reflective latent construct. Rather, cumulative ACE scores are frequently treated as formative indices. This distinction has important implications because conventional psychometric indices such as Cronbach’s alpha or factor analysis may not always be theoretically appropriate. The manuscript should discuss this issue explicitly instead of evaluating every instrument through the same psychometric lens.
(4) Although COSMIN is appropriately cited, the manuscript does not clearly describe how COSMIN ratings were operationalised. It remains unclear whether each psychometric property was independently rated by two reviewers, how disagreements were resolved, whether the “worst score counts” principle was applied, and how methodological ratings were translated into the summary tables. Greater transparency is needed to ensure reproducibility.
(5) The search strategy appears relatively narrow for a COSMIN review. Instrument validation studies are frequently difficult to retrieve because they may not explicitly use terms such as “validation” or “psychometric properties.” The authors should justify why no validated COSMIN search filter for measurement-property studies was employed and discuss the potential impact on retrieval sensitivity.
(6) The inclusion criteria require clarification. The review states that only adult samples were included because ACE questionnaires retrospectively assess childhood experiences. However, several included studies involve adolescents. This creates an inconsistency between the stated eligibility criteria and the final sample. The eligibility section should be rewritten to accurately reflect the population actually reviewed.
(7) The manuscript summarises which psychometric properties were evaluated but does not sufficiently distinguish between absence of evidence and evidence of poor quality. For example, responsiveness is described as weak across instruments, whereas most studies simply never evaluated responsiveness because these questionnaires are designed primarily for exposure assessment rather than monitoring longitudinal change. This distinction should be emphasised throughout the Results and Discussion.
(8) The Results section is dominated by narrative descriptions of individual studies. Much of this information duplicates Tables 2–5. Instead, the Results should synthesise patterns across instruments. For example, which psychometric properties consistently demonstrate strong evidence across questionnaires? Which properties systematically remain unevaluated? Which methodological weaknesses recur across validation studies? Such synthesis would considerably strengthen the review.
(9) The Discussion tends to restate the descriptive findings instead of interpreting their implications. For example, why has responsiveness rarely been evaluated? Is this because ACE instruments are intended for epidemiological exposure measurement rather than repeated clinical assessment? Similarly, why do measurement error and criterion validity remain poorly studied? A stronger discussion should move beyond reporting frequencies toward explaining the methodological reasons behind these patterns.
(10) The manuscript repeatedly refers to percentages of “strong evidence” (e.g., 37.5%) without sufficiently explaining how these percentages were calculated or what they actually represent. Readers may mistakenly interpret these values as quantitative quality scores. A detailed methodological explanation should accompany these calculations or, alternatively, the percentages should be omitted if they oversimplify the COSMIN framework.
(11) More attention should be paid to content validity. COSMIN considers content validity the most important measurement property. Nevertheless, relatively little discussion is devoted to whether existing ACE instruments adequately capture contemporary conceptualisations of childhood adversity, including community violence, discrimination, digital victimisation, chronic poverty, or protective experiences. This omission limits the practical significance of the review.
(12) The discussion of cross-cultural validity remains relatively superficial. Cross-cultural adaptation extends well beyond linguistic translation. Measurement invariance, conceptual equivalence, cultural relevance of adversity domains, and differential interpretation of abuse-related items deserve substantially deeper consideration, particularly given the international application of the ACE-IQ.
(13) The manuscript would benefit from a more explicit discussion of retrospective recall bias. Beyond acknowledging recall limitations, the authors should discuss how retrospective reporting may influence estimates of reliability, construct validity and criterion validity, and whether this limitation differentially affects the interpretation of psychometric findings across instruments.
(14) The conclusion currently suggests that available instruments are generally reliable for assessing ACEs. This statement is somewhat stronger than the evidence presented. The review actually demonstrates that evidence is uneven across psychometric domains, with substantial gaps in responsiveness, measurement error, criterion validity and cross-cultural evaluation. The conclusion should more accurately reflect this heterogeneity.
(15) Finally, the manuscript would have greater practical value if it concluded with explicit recommendations for researchers selecting an ACE instrument. Rather than simply stating that further validation studies are needed, the review could indicate which questionnaire is currently preferable for epidemiological surveys, which is most suitable for international comparisons, which has the strongest evidence for structural validity, and where the major evidence gaps remain. Such guidance would considerably increase the translational value of the review.
Author Response
The authors of the manuscript entitled ‘A Systematic Review of the Psychometric Quality of Instruments for Assessing Adverse Childhood Experiences’ would like to thank the reviewers for their comments. The changes made in response to these comments are outlined below.
Reviewer 1.
COMMENT: The manuscript would benefit from a clearer conceptual distinction between adverse childhood experiences (ACEs) as a public health construct and child maltreatment as a clinical or legal construct. At present these concepts are introduced almost interchangeably, despite representing overlapping but non-equivalent frameworks. Since one of the review’s objectives is to evaluate measurement instruments, the conceptual target being measured should first be explicitly defined. Otherwise, readers cannot determine whether differences between instruments reflect psychometric superiority or differences in construct coverage.
ANSWER: In response to the reviewer’s comment, the following paragraph has been included: ACEs are an epidemiological concept used in healthcare to predict the risk of physical and psychological illnesses that may develop in a person who has experienced these situations, which constitute one of the main sources of stress in childhood [1]. These experiences include having suffered abuse or neglect; whilst these have an impact on health, they are also addressed from a legal perspective, insofar as they seek to establish the criminal liability of carers.
COMMENT: The review repeatedly concludes that certain instruments have “better psychometric quality”, yet the comparison is made across studies with markedly different samples, designs, statistical methods and validation objectives. COSMIN evaluates the quality of evidence within a validation study rather than providing a direct ranking between instruments. Consequently, statements implying superiority of one questionnaire over another should either be substantially qualified or supported by a formal evidence synthesis strategy specifically designed for between-instrument comparison.
ANSWER: In response to the reviewer’s comment, the following paragraph has been included: From a statistical and psychometric perspective, those instruments that demonstrate greater validity and reliability may be superior to those that score lower on these constructs. Furthermore, it is important to ensure that these psychometric properties are maintained across different cultures, contexts and samples. It can be concluded that, amongst the instruments analysed in this review, the one with the greatest number of studies evaluating its psychometric properties—most of which demonstrate good structural validity and internal consistency, whilst also having been evaluated in a wider range of cultural contexts with good cross-cultural validity is the World Health Organisation’s ACE-IQ [1].
COMMENT: The review focuses almost exclusively on classical psychometric properties while giving little attention to the underlying measurement model adopted by each instrument. Several ACE questionnaires were not developed under the assumption that adversity represents a reflective latent construct. Rather, cumulative ACE scores are frequently treated as formative indices. This distinction has important implications because conventional psychometric indices such as Cronbach’s alpha or factor analysis may not always be theoretically appropriate. The manuscript should discuss this issue explicitly instead of evaluating every instrument through the same psychometric lens.
ANSWER: In response to the reviewer’s comment, the following paragraph has been included: In the case of adverse childhood experiences, the use of Cronbach’s alpha or factor analysis may be inappropriate if one assumes that it is the indicators that define the construct (formative model), as opposed to models in which adversity is the cause of the indicators captured in the instrument’s items (reflective model), thereby generating erroneous estimates, since the indicators that make up the construct do not necessarily have to be interrelated [45]. For example, having experienced school-related difficulties does not necessarily have to be linked to having suffered abuse or neglect at the hands of primary carers. In such cases, as Cruz-Avelar et al. [46] point out, it is important to focus on content validity through scientific evidence when selecting the items for the instrument, taking into account the specific characteristics of the population to be assessed or the purpose of the instrument, amongst other aspects.
COMMENT: Although COSMIN is appropriately cited, the manuscript does not clearly describe how COSMIN ratings were operationalised. It remains unclear whether each psychometric property was independently rated by two reviewers, how disagreements were resolved, whether the “worst score counts” principle was applied, and how methodological ratings were translated into the summary tables. Greater transparency is needed to ensure reproducibility.
ANSWER: In response to the reviewer’s comment, the following paragraph has been included: The data relating to the psychometric properties reported in the analysed articles were collated by two researchers (FG-S and LC); where there was disagreement, a third reviewer intervened (MI-C). This information was transferred to the results tables, which summarised the characteristics of the samples (see Table 2) and the psychometric properties listed in the COSMIN checklist (see Table 3), namely: structural validity, internal consistency, cross-cultural validity/measurement invariance, reliability, measurement error, criterion validity, hypothesis testing for construct validity and responsiveness. All items comprising the checklist were rated on a 5 point scale (excellent, good, fair, poor and unknown/NA). The overall score for each box was based on the lowest score (1–5) for an item within that box [21,22]. When assessing the quality of each instrument, the GRADE criteria were followed, with each property rated on a three-point scale: sufficient (+), uncertain (?) and insufficient (-) (see Table 4), as specified in the COSMIN manual [21] and in the guidelines provided by Prinsen et al. [23] and Terwee et al. [24]. Finally, the quality of evidence for each instrument was summarised in a table (see Table 5), as suggested by Prinsen et al. [22], distinguishing between high-quality evidence when several articles demonstrated good methodology or one article was of excellent quality; moderate-quality evidence when there were several articles with acceptable methodology or one of good quality; limited-quality evidence when the quality was acceptable; and conflicting evidence when the articles were of low quality. In cases where studies lacked information on psychometric properties, the evidence was assessed as conflicting.
COMMENT: The search strategy appears relatively narrow for a COSMIN review. Instrument validation studies are frequently difficult to retrieve because they may not explicitly use terms such as “validation” or “psychometric properties.” The authors should justify why no validated COSMIN search filter for measurement-property studies was employed and discuss the potential impact on retrieval sensitivity.
ANSWER: The search terms used combine the names of the instruments that assess ACE which were identified in an initial phase – with measurement properties; terms relating to samples were not included as the aim was to obtain the maximum number of articles. However, taking the reviewer’s comment into account, the search strategy itself has been included amongst the study’s limitations, as it may have restricted the number of articles identified and assessed in this review. In light of this, the following paragraph has been included: Among the limitations of this review, it is worth noting the limited number of studies assessing the psychometric properties of some of the instruments included in this review. This may be related to the limited number of databases used, as well as to the search query itself, which may have excluded other studies that did not include terms related to psychometric properties in their title, abstract or keywords.
COMMENT: The inclusion criteria require clarification. The review states that only adult samples were included because ACE questionnaires retrospectively assess childhood experiences. However, several included studies involve adolescents. This creates an inconsistency between the stated eligibility criteria and the final sample. The eligibility section should be rewritten to accurately reflect the population actually reviewed.
ANSWER: The omission has been rectified by including adolescents.
COMMENT: The manuscript summarises which psychometric properties were evaluated but does not sufficiently distinguish between absence of evidence and evidence of poor quality. For example, responsiveness is described as weak across instruments, whereas most studies simply never evaluated responsiveness because these questionnaires are designed primarily for exposure assessment rather than monitoring longitudinal change. This distinction should be emphasised throughout the Results and Discussion.
ANSWER: In response to the reviewer’s comment, the following paragraph has been included in the discussion section. Although this may be explained by the fact that the instruments used to assess ACEs were designed to measure exposure to such experiences rather than to track longitudinal changes.
COMMENT: The Results section is dominated by narrative descriptions of individual studies. Much of this information duplicates Tables 2–5. Instead, the Results should synthesise patterns across instruments. For example, which psychometric properties consistently demonstrate strong evidence across questionnaires? Which properties systematically remain unevaluated? Which methodological weaknesses recur across validation studies? Such synthesis would considerably strengthen the review.
ANSWER: This text has been included in relation to Table 4.
With respect to cross-cultural validity, this has been assessed primarily for the ACE-IQ instrument [29,31,32,34–39]. Furthermore, it should be noted that the least frequently assessed psychometric properties have been reliability, which was assessed in only six studies [34,35,40,42–44]; no study was found to have assessed it for the ACE-10 instrument, whilst responsiveness was the least frequently assessed, except in the study by Schauss et al. [40].
Among the studies that have assessed a greater number of psychometric properties, those relating to the ACE-IQ [32,34,35], the study by Hietamäki et al. [44] for the ACE-THL instrument, and the study by Chen et al. [43] for the SC-ACE-IQ instrument are particularly noteworthy.
COMMENT: The Discussion tends to restate the descriptive findings instead of interpreting their implications. For example, why has responsiveness rarely been evaluated? Is this because ACE instruments are intended for epidemiological exposure measurement rather than repeated clinical assessment? Similarly, why do measurement error and criterion validity remain poorly studied? A stronger discussion should move beyond reporting frequencies toward explaining the methodological reasons behind these patterns.
ANSWER: in response to the reviewer’s comment, this text has been included in the discussion section.
Measurement error has been another of the properties that has been assessed to a lesser extent, specifically in 13 of the 20 studies analysed. This may be linked to memory biases, given that instruments used to assess ACEs do so retrospectively; the presence or absence of an event may therefore be affected if the individual denies or downplays it. Furthermore, other biases associated with self-reports, gender, and social or cultural groups may also contribute to this lack of measurement; this would require assessing whether each item within an instrument behaves in the same way across the different groups being assessed. In this regard, it is important to bear in mind the normalisation of certain behaviours that may occur in different groups, such as the use of violence or physical punishment in the upbringing of children, or abandonment or neglect in care that may occur simply because someone is a man or a woman. All these factors also complicate the assessment of criterion validity, in addition to other issues associated with this type of instrument, which provide a total score based on the number of incidents the individual has experienced, without weighting the score according to the severity of the incident, the age at which it occurred, or whether the incident or incidents became chronic. However, some retrospective epidemiological studies [47,48] highlight the usefulness of this type of instrument despite the difficulties associated with memory biases.
As mentioned above, cultural biases can, to a certain extent, influence cross-cultural validity. The use of these instruments across different cultures involves more than simply translating the items, given that terms such as ‘abuse’, ‘mistreatment’ or ‘neglect’ can vary from one culture to another; in some cases, certain practices may even be normalised, whereas in other contexts they would be regarded as forms of mistreatment. Existing values and prejudices towards certain communities, minorities or specific social groups, as well as the laws and customs specific to each community, make it difficult to draw comparisons between cultures and, consequently, to study cross-cultural validity.
COMMENT: The manuscript repeatedly refers to percentages of “strong evidence” (e.g., 37.5%) without sufficiently explaining how these percentages were calculated or what they actually represent. Readers may mistakenly interpret these values as quantitative quality scores. A detailed methodological explanation should accompany these calculations or, alternatively, the percentages should be omitted if they oversimplify the COSMIN framework.
ANSWER: As the reviewer points out, these percentages have been omitted from the tables and the text.
COMMENT: More attention should be paid to content validity. COSMIN considers content validity the most important measurement property. Nevertheless, relatively little discussion is devoted to whether existing ACE instruments adequately capture contemporary conceptualisations of childhood adversity, including community violence, discrimination, digital victimisation, chronic poverty, or protective experiences. This omission limits the practical significance of the review.
ANSWER: The following paragraph has been included amongst the limitations of this study: Another limitation of this review is that content validity has not been assessed. When assessing which instruments exhibit the best psychometric properties, content validity is one of the most important measurement properties in Patient-Reported Outcome Measures (PROMs) [49]. This has an impact on clinical practice [50]. In the case of ACEs, it is important to bear in mind the evolution of this dimension, taking into account not only the now-classic events but also other contextual and cultural factors, as well as protective factors [4,15,16], which may affect content validity. In this regard, amongst the instruments analysed in this review, it should be noted that the ACE-IQ, the SC-ACE-IQ and the ACE-THL are the ones that include the greatest number of items relating, in some cases, to stressors and, in others, to protective factors. Furthermore, not all the studies analysed provided details on how the instrument had been developed. Consequently, content validity could not be assessed.
COMMENT: The discussion of cross-cultural validity remains relatively superficial. Cross-cultural adaptation extends well beyond linguistic translation. Measurement invariance, conceptual equivalence, cultural relevance of adversity domains, and differential interpretation of abuse-related items deserve substantially deeper consideration, particularly given the international application of the ACE-IQ.
ANSWER: The following paragraph has been included in the discussion: This implies not only linguistic adaptation but also taking into account specific cultural and contextual factors that may affect the validity of the questionnaire and the epidemiological studies themselves, thereby skewing reporting rates. Furthermore, a review of administration procedures and data collection strategies is required, particularly with regard to vulnerable groups.
COMMENT: The manuscript would benefit from a more explicit discussion of retrospective recall bias. Beyond acknowledging recall limitations, the authors should discuss how retrospective reporting may influence estimates of reliability, construct validity and criterion validity, and whether this limitation differentially affects the interpretation of psychometric findings across instruments.
ANSWER: This point has already been addressed in the discussion section in response to another previous comment from the reviewer.
COMMENT: The conclusion currently suggests that available instruments are generally reliable for assessing ACEs. This statement is somewhat stronger than the evidence presented. The review actually demonstrates that evidence is uneven across psychometric domains, with substantial gaps in responsiveness, measurement error, criterion validity and cross-cultural evaluation. The conclusion should more accurately reflect this heterogeneity.
ANSWER: In response to the reviewer’s comment, the conclusion has been amended, as previously noted.
COMMENT: Finally, the manuscript would have greater practical value if it concluded with explicit recommendations for researchers selecting an ACE instrument. Rather than simply stating that further validation studies are needed, the review could indicate which questionnaire is currently preferable for epidemiological surveys, which is most suitable for international comparisons, which has the strongest evidence for structural validity, and where the major evidence gaps remain. Such guidance would considerably increase the translational value of the review.
ANSWER: In response to the reviewer’s comment, the following paragraph has been included in the final paragraph of the article. In light of the results obtained, this instrument may be the most suitable for conducting epidemiological studies; however, given the dynamic nature of adverse experiences in childhood, the instruments should be reviewed with a view to improving their content validity.
Reviewer 3 Report
Comments and Suggestions for AuthorsGeneral Comments
This well-structured systematic review addresses a crucial clinical need by utilizing the PRISMA guidelines and COSMIN checklist to evaluate the psychometric quality of retrospective Adverse Childhood Experiences (ACEs) assessment tools across 20 empirical studies The manuscript successfully exposes a major structural vulnerability in the existing literature: while internal consistency and structural validity are robustly supported, vital parameters such as measurement error and clinical responsiveness (sensitivity to change over time) remain heavily under-evaluated
To enhance the depth and clinical utility of the manuscript, a minor revision is recommended to expand your Discussion and Future Lines of Research sections
Key Points for Revision
Demographic Sample Biases: Please elaborate on how severe demographic sample imbalances (the over-representation of young females, adolescents, and university students) constrain wide clinical generalizability to male or older adult cohorts
Methodological Adaptations for Vulnerable Cohorts: Explicitly suggest how these retrospective self-reports can be methodologically adapted (e.g., simplifying complex item phrasing or providing communication accommodations) to safely and reliably capture trauma data in highly vulnerable, historically omitted cohorts like individuals with disabilities
Contextualizing Evidence Scarcity: Clearly clarify in the text that the high evidence percentages achieved by certain newer versions (e.g., ACE-THL and ACE-IQ-10) are mathematically derived from a single study evaluating that specific tool, rather than a broad, multi-study global evidence base like the ACE-IQ
Cultural Sensitivity vs. Translation: Briefly emphasize that linguistic translation does not inherently equate to full cross-cultural validation, noting how varying global social stigmas and baseline definitions of "abuse" can skew reporting rates
Minor Technical Alignment
Please double-check that the index numbers used for the studies in Table 3 correspond exactly with the study citations and numbering sequences distributed throughout the main text to prevent reading confusion
Author Response
The authors of the manuscript entitled ‘A Systematic Review of the Psychometric Quality of Instruments for Assessing Adverse Childhood Experiences’ would like to thank the reviewers for their comments. The changes made in response to these comments are outlined below.
Review 3
General Comments
This well-structured systematic review addresses a crucial clinical need by utilizing the PRISMA guidelines and COSMIN checklist to evaluate the psychometric quality of retrospective Adverse Childhood Experiences (ACEs) assessment tools across 20 empirical studies The manuscript successfully exposes a major structural vulnerability in the existing literature: while internal consistency and structural validity are robustly supported, vital parameters such as measurement error and clinical responsiveness (sensitivity to change over time) remain heavily under evaluated
To enhance the depth and clinical utility of the manuscript, a minor revision is recommended to expand your Discussion and Future Lines of Research sections
COMMENT: Demographic Sample Biases: Please elaborate on how severe demographic sample imbalances (the over-representation of young females, adolescents, and university students) constrain wide clinical generalizability to male or older adult cohorts. Methodological Adaptations for Vulnerable Cohorts: Explicitly suggest how these retrospective self-reports can be methodologically adapted (e.g., simplifying complex item phrasing or providing communication accommodations) to safely and reliably capture trauma data in highly vulnerable, historically omitted cohorts like individuals with disabilities
ANSWER: The following paragraph has been included in the discussion: This selection bias restricts generalization to other population groups, such as men, older people or groups in more vulnerable situations, given that the magnitude of the effect observed in these studies does not necessarily reflect what would occur in the general or clinical population, thereby failing to control for other variables that may influence the results obtained. The lack of studies that have adapted ACE questionnaires for vulnerable groups highlights a real need, particularly given that these groups (children with disabilities, neurodevelopmental disorders, children undergoing long-term hospitalization…) may be at greater risk of experiencing adverse events. In this regard, adaptations are required that use clear and comprehensible language tailored to the difficulties faced by these individuals, as well as adaptations and interviews conducted by experts, accompanied by pilot studies to validate the adaptation carried out.
COMMENT: Contextualizing Evidence Scarcity: Clearly clarify in the text that the high evidence percentages achieved by certain newer versions (e.g., ACE-THL and ACE-IQ-10) are mathematically derived from a single study evaluating that specific tool, rather than a broad, multi-study global evidence base like the ACE-IQ.
ANSWER: The following sentence has been included in the discussion: It should be noted that the ACE-IQ-10, the SC-ACE-IQ and the ACE-THL have each been evaluated in only one study [36,40,41].
COMMENT: Cultural Sensitivity vs. Translation: Briefly emphasize that linguistic translation does not inherently equate to full cross-cultural validation, noting how varying global social stigmas and baseline definitions of "abuse" can skew reporting rates
ANSWER: The following paragraph has been included in the discussion: This implies not only linguistic adaptation but also taking into account specific cultural and contextual factors that may affect the validity of the questionnaire and the epidemiological studies themselves, thereby skewing reporting rates. Furthermore, a review of administration procedures and data collection strategies is required, particularly with regard to vulnerable groups.
Minor Technical Alignment
COMMENT: Please double-check that the index numbers used for the studies in Table 3 correspond exactly with the study citations and numbering sequences distributed throughout the main text to prevent reading confusion
ANSWER: As suggested by the reviewer, the numbers have been checked and it has been decided to include the reference number for each article in Table 3.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe manuscript addresses a relevant and important topic by systematically examining the psychometric quality of instruments used to assess adverse childhood experiences. The use of PRISMA and COSMIN provides an appropriate methodological framework, and the revised version has improved the rationale for the review, the description of reviewer involvement, the discussion of cross-cultural validity, and the caution used in the conclusions. The manuscript also offers a potentially useful overview of the main ACE instruments currently used across different populations and cultural contexts. Nevertheless, several substantive issues should be addressed before the manuscript is suitable for publication.
First, the methodological application of COSMIN requires further clarification and possible revision. The manuscript appears to combine different stages of the COSMIN procedure, including methodological quality assessment, evaluation of measurement-property results, and grading of the certainty of evidence. These stages should be clearly distinguished. In particular, the authors should specify whether the COSMIN Risk of Bias Checklist was applied separately to each measurement property, how ratings were assigned, and how disagreements were resolved. The response to the previous concern regarding risk of bias does not directly address this issue, as it focuses instead on the absence of content-validity assessment. These are separate methodological concerns.
Second, the evidence-synthesis approach should be reconsidered. Table 5 reports percentages of “strong evidence” for individual studies, but COSMIN-based certainty grading should generally be conducted at the level of each instrument and each measurement property, using the total body of evidence. The current presentation may therefore be difficult to interpret and may overstate the comparative strength of instruments supported by only one study. The authors should reconstruct the synthesis so that each instrument is evaluated across the relevant psychometric domains, with the overall result and certainty of evidence reported for each domain.
Third, the search strategy remains insufficiently reproducible. The manuscript identifies the databases and provides general search terms, but it does not report the complete database-specific search strings, fields, filters, language restrictions, exact temporal limits, and final search date. These details should be included in supplementary material. The rationale for restricting the review to studies published within the previous 15 years should also be justified more convincingly, because this restriction may exclude foundational validation studies and introduce selection bias.
Fourth, the eligibility criteria and target population require clarification. The abstract refers to instruments used in adults, whereas the Methods and Results include adolescents and, in some studies, participants as young as 10 or 11 years. The population criteria should be stated consistently throughout the abstract, methods, results, and conclusions. The authors should also explain whether studies involving caregiver-assisted administration, modified instrument versions, and clinical versus general-population samples were eligible under predefined criteria.
Fifth, the manuscript contains several internal inconsistencies, citation errors, and labeling problems that should be corrected. Examples include incorrect reference numbers for some studies, the inclusion of reference [18] among the psychometric studies in Table 3, the incorrect heading “hypothesis testing for content validity” in Table 4, and inconsistent citations of the ACE-IQ-10 study. The reported totals of participants and the number of studies assessing measurement error should also be recalculated and aligned with the tables. Some table entries remain in Spanish and should be translated into English.
Sixth, the interpretation of the findings should be more cautious. Statements indicating that the reviewed instruments are generally reliable or suitable are too broad given that evidence is frequently moderate, unknown, or absent across several properties. The conclusion that ACE-IQ is the most suitable instrument should be reframed to indicate that it has the broadest evidence base among the instruments reviewed, while acknowledging that evidence quality varies by psychometric domain and that content validity, measurement error, responsiveness, and cultural adaptation remain insufficiently established.
Finally, the manuscript would benefit from a thorough language and editorial revision. Typographical errors, awkward constructions, inconsistent terminology, and formatting problems remain throughout the text and tables. Careful proofreading is necessary to improve readability, precision, and professional presentation.
Overall, the manuscript has clear potential and addresses a meaningful gap, but substantial methodological clarification, correction of the evidence synthesis, verification of tables and citations, and comprehensive language editing are required. I therefore recommend major revision before the manuscript can be considered for publication
Comments on the Quality of English LanguageThe overall quality of English is acceptable, and the manuscript is generally understandable. However, the text still requires careful professional language editing before publication. Several grammatical, typographical, stylistic, and terminological inconsistencies remain and occasionally reduce clarity.
Examples include typographical errors such as “assesment,” “blindy,” and “Aditionally,” as well as awkward expressions such as “global applicable tool.” Some sentences are excessively long and would benefit from restructuring to improve readability and precision. Punctuation is also inconsistent in several places, including missing commas in coordinated adjective sequences such as “inclusive, culturally sensitive, and accessible.”
Terminology should be standardized throughout the manuscript. For example, “cross-cultural,” “crosscultural,” “measurement invariance,” “content validity,” and “construct validity” are not always used consistently. Table 4 also contains the incorrect heading “hypothesis testing for content validity,” whereas the relevant COSMIN domain is hypothesis testing for construct validity.
Several table entries remain in Spanish, particularly in the descriptions of the Zanotti et al. and Hietamäki et al. samples, and these should be translated fully into English. In addition, author names, references, hyphenation, capitalization, and abbreviations should be checked systematically for consistency.
I therefore recommend comprehensive proofreading by a proficient academic English editor. The manuscript does not require complete rewriting, but it does need substantial linguistic and editorial revision to achieve the level of clarity, accuracy, and stylistic consistency expected for publication.
Author Response
Thank you for your suggestions
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors have addressed all concerns raised in the previous round of review.
Author Response
thank you
