Does Eyewitness Confidence Calibration Vary by Target Race?
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThis manuscript was interesting and examines an important topic. I think it is a good fit for this special issue. That being said, I did notice some things that could be improved:
- Perhaps most importantly, the introduction is very short (and has an overarching title which essentially states the conclusion of this paper rather than a research question). It does engage with some relevant literature but is brief. One important paper they have missed recently examined the confidence accuracy relationship in same and cross race line ups (but with White and Black ethnicity and simultaneous line ups) (Helm & Spearing, 2025). That paper, combined with other sources already cited in the introduction, already suggest that confidence calibration will be better in same race line ups than cross race line ups (although it is less clear whether this extends beyond white witnesses I think). This paper is useful in extending that work, but I think the introduction would work better in laying out existing work more fully, and then in identifying unanswered questions clearly, and how this work answers some of those questions (e.g., extending to Asian targets and witnesses, and extending to sequential lineups). In relation to these questions some discussion of prior expectations (e.g., why would you expect findings to replicate or not in these new contexts) seems important.
- The authors should clarify how many participants were recruited from each source (paid vs. student sample) and, if possible, should control for this in analyses.
- The Tredoux's E was much lower for the Asian line ups. It would be good to justify why findings should still be considered reliable in light of that (and why lineups weren't redesigned to be more comparable to the White lineups).
- The sample size to me seemed quite low for CAC curves.
- It would be good to know why the authors did not use ROC curves which, I think, would be a more typical way of measuring overall calibration quality.
Author Response
Dear Reviewer,
We sincerely thank you for your comments and suggestions. We found them extremely helpful in improving the quality and clarity of our manuscript. We have marked our responses to reviewer comments and the corresponding changes in the manuscript in purple text.
-The Authors
- Perhaps most importantly, the introduction is very short (and has an overarching title which essentially states the conclusion of this paper rather than a research question). It does engage with some relevant literature but is brief. One important paper they have missed recently examined the confidence accuracy relationship in same and cross race line ups (but with White and Black ethnicity and simultaneous line ups) (Helm & Spearing, 2025). That paper, combined with other sources already cited in the introduction, already suggest that confidence calibration will be better in same race line ups than cross race line ups (although it is less clear whether this extends beyond white witnesses I think)
This paper is useful in extending that work, but I think the introduction would work better in laying out existing work more fully, and then in identifying unanswered questions clearly, and how this work answers some of those questions (e.g., extending to Asian targets and witnesses, and extending to sequential lineups). In relation to these questions some discussion of prior expectations (e.g., why would you expect findings to replicate or not in these new contexts) seems important.
Thank you for your suggestion, we agree and we have now revised the title to “Does Eyewitness Confidence Calibration Vary by Target Race?” to frame the work as a research question rather than to imply a generalized conclusion.
We were not previously aware of the paper you shared and have now incorporated it into the Introduction, as it is highly relevant to the present work. We sincerely thank you for bringing it to our attention.
The Introduction was intentionally kept concise in order to engage directly with the most relevant literature while maintaining focus on our specific research question, rather than providing a broader narrative review. At present, the published literature examining the relationship between target race and the confidence–accuracy relationship remains limited to the studies discussed in the manuscript. Although each of these studies could be elaborated in greater detail, we believe that such an approach would be better suited to review articles that have already provided comprehensive coverage of the topic (e.g., Wixted & Wells, 2017). However, we acknowledge the need for clarifications, and we now included some more detail about the papers we reviewed.
Finally, we have strengthened both the Introduction and Discussion sections to more clearly articulate how the current study builds on and extends prior work, and how it addresses questions that had not been previously examined in the literature.
2. The authors should clarify how many participants were recruited from each source (paid vs. student sample) and, if possible, should control for this in analyses.
We used the same lineups in a previous (published) study that included White participants recruited both from paid online platforms (e.g., MTurk/CloudResearch) and from a university participant pool (drawn from the same institution as the present study). This allowed us to explore whether recruitment source influences lineup performances. If participants recruited via MTurk/CloudResearch performed better on the lineup task (e.g., by more accurately rejecting target-absent lineups or correctly identifying the target), they would be expected to yield lower Tredoux’s E values than university participants. However, the findings from that study showed that lineup fairness was comparable across recruitment sources (see the table below, which is also reported in a doctoral thesis that is available online for review, unmasked link available upon request). Consequently, there is neither a theoretical nor an empirical basis for conducting the proposed additional analysis, which would also reduce statistical power in the present study. We have now noted this in the manuscript and direct readers (or will do so following masked review) to the exact page of the doctoral thesis where this table is reported, which is freely available online.
Table 1
Lineup Fairness Analysis Based on the Recruitment Platform
|
|
MT/CR participants (n = 233) |
non-MT/CR participants (n = 170) |
|
White Lineup 1 |
3.52 [2.94, 4.37] |
3.02 [2.53, 3.74] |
|
White Lineup 2 |
3.31 [2.59, 4.58] |
3.82 [2.61, 7.10] |
|
Asian Lineup 1 |
2.27 [1.68, 3.48] |
1.67 [1.24, 2.57] |
|
Asian Lineup 2 |
2.10 [1.55, 3.25] |
2.55 [1.80, 4.37] |
Note. Tredoux’E results and 95% Confidence Intervals are reported.
3. The Tredoux's E was much lower for the Asian line ups. It would be good to justify why findings should still be considered reliable in light of that (and why lineups weren't redesigned to be more comparable to the White lineups).
The lineups were not redesigned because the present analyses were conducted after data collection had concluded (as we preregistered to do), and funding constraints did not permit the collection of an additional dataset. We have explicitly acknowledged this lineup fairness issue in the manuscript. Nevertheless, we encourage readers to consider an important nuance when interpreting the Tredoux’s E values reported here. Traditionally, Tredoux’s E is calculated using same-race participants as the suspect, because the presence of a cross-race effect (CRE) can itself influence estimates of lineup fairness. In the present study, however, Asian participants did not exhibit a CRE indicating an own-race advantage for Asian faces. As a result, these participants may have experienced difficulty comparable to that of cross-race observers when evaluating the lineups, which would yield Tredoux’s E values similar to those typically obtained from cross-race evaluators: low. We nonetheless acknowledge that ideally such effects should be examined using lineups with higher baseline fairness, and we recognize this as a limitation of the present work. Because it was not feasible to conduct an additional study, we felt that the current approach of noting this issue as a limitation represented the most appropriate response to the reviewer’s concern. Importantly, we believe the study remains informative, as we excluded the unfair Asian lineup and retained only those lineups that yielded acceptable fairness.
- The sample size to me seemed quite low for CAC curves.
To our knowledge, there are no established tools for conducting a priori power analyses for CAC curves. Sample sizes used for CAC analyses in the published literature vary widely, ranging from relatively small samples (e.g., Pezdek, Abed, & Reisberg, 2020, Total N = 114) to medium-samples (Kleider-Offutt, Stevens, Mickes, & Boogert, N = 237; Helm & Spearing, 2025, N = 205) substantially larger ones (e.g., Mansour, 2020, N = 968). Our sample size (N = 203 Asian and N= 202 White) aligns with those of published research.
- It would be good to know why the authors did not use ROC curves which, I think, would be a more typical way of measuring overall calibration quality.
A receiver operating characteristic (ROC) curve depicts the trade-off between the true positive rate (sensitivity) and the false positive rate (1 − specificity) across varying decision criteria. ROC curves are a standard in the field for examining eyewitness performance (accuracy) but not for examining confidence. The standard in the field for examining confidence are confidence–accuracy characteristic (CAC) curves firstly, and calibration curves secondly. Accordingly, we reported CAC and calibration indices.
Although it would, in principle, be possible to examine how partial area under the curve (pAUC) varies across the confidence bins used in ROC analyses for White and Asian targets, we have never seen such an analysis in the published literature. Indeed, all of the studies cited in our introduction which examined the confidence-accuracy relationship relied on CAC or calibration analyses rather than ROC curves. Moreover, ROC analyses—particularly those comparing pAUC across conditions—typically require substantially larger sample sizes to yield stable and interpretable estimates, which would not be feasible with the present dataset.
Reviewer 2 Report
Comments and Suggestions for AuthorsI truly enjoyed reading this manuscript. The study is incredibly interesting and the manuscript very well-written. I wanted to particularly commend you on the brief explanations on statistical approaches in the introduction. Too many articles assume readers are familiar with (or do not care that readers may not be familiar with) the common practices within their lab, region, or specific field. Seeing such a thorough sample size determination was also a breath of fresh air. Overall, you have a clear and pleasant writing style which was a joy to read. Know that, whenever this gets published, many an undergrad in my classroom will be forced to read at least your introduction as it provides a perfect example of a clear, well-structured introduction section.
I only have a handful of questions, comments, and suggestions—see below:
Questions:
Page 4. Demographics: I was wondering whether you could speak a bit to the fact that your Asian sample is mainly from the US and your White sample is mainly from the UK? It’s not acknowledged anywhere in the discussion and although I can’t think of any major red flags, it did catch my attention.
- “The position of the target and fillers was randomised across participants” Just to clarify, does this mean the position of the target was randomized for each lineup or just for each participant? I am interpreting the latter as the target being in the same position across all four lineups that a participant sees.
Page 8: My apologies—I am used to working with CIs, not SEs (in tables, at least). To clarify:
“Overlapping standard errors indicate a difference is not statistically significant, while non-overlapping confidence intervals indicate a statistically significant difference.” So for both, overlap means a lack of statistical significance?
Standard errors that do not overlap and confidence intervals that do overlap are uninformative about statistical significance.” But wouldn’t SEs that do not overlap indicate a significant difference, as you just mentioned overlapping SEs indicate a lack of significance?
Small edits:
Page 4. Demographics: Per APA 7, numbers under 10 should be spelled out.
- Target Race “:”
- Seems like this was maybe supposed to be formatted as a heading?
Page 10: this study was the first in the literature to found that White participants’ CA relationship differs for White and Asian targets using a sequential lineup.
Author Response
Dear Reviewer,
We sincerely thank you for your comments and suggestions. We found them extremely helpful in improving the quality and clarity of our manuscript. We have marked our responses to reviewer comments and the corresponding changes in the manuscript in purple text.
-The Authors
Page 4. Demographics: I was wondering whether you could speak a bit to the fact that your Asian sample is mainly from the US and your White sample is mainly from the UK? It’s not acknowledged anywhere in the discussion and although I can’t think of any major red flags, it did catch my attention.
Thank you for this comment. The sample characteristics that caught your attention may relate to the fact that we did not find a CRE for our Asian participants, a result that is consistent with prior findings in the literature, as discussed in the manuscript. However, the fact that the participants are both from a different country and a somewhat different culture is worth noting and we now do so.
“The position of the target and fillers was randomised across participants” Just to clarify, does this mean the position of the target was randomized for each lineup or just for each participant? I am interpreting the latter as the target being in the same position across all four lineups that a participant sees.
We have now clarified this point in the manuscript. Each lineup employed a randomized configuration of target/suspect position, and this configuration varied across participants.
Page 8: My apologies—I am used to working with CIs, not SEs (in tables, at least). To clarify: “Overlapping standard errors indicate a difference is not statistically significant, while non-overlapping confidence intervals indicate a statistically significant difference.” So for both, overlap means a lack of statistical significance? Standard errors that do not overlap and confidence intervals that do overlap are uninformative about statistical significance.” But wouldn’t SEs that do not overlap indicate a significant difference, as you just mentioned overlapping SEs indicate a lack of significance?
This is indeed a confusing set of relationships. To clarify, one cannot assume that because non-overlapping CIs imply a difference, overlapping CIs should imply no statistical difference. Likewise, one cannot assume that because overlapping SEs imply null differences, non-overlapping SEs should imply differences. To put it another way:
- Overlapping SEs: The groups are not significantly different.
- Overlapping CIs: Uninformative about whether the groups differ
- Non-overlapping SEs: Uninformative about whether the groups differ
- Non-overlapping CIs: The groups are significantly different.
Therefore, overlapping SEs indicate a non-significant difference and non-overlapping CIs indicate a significant difference. One requires both types of error bars to make a definitive judgment. However, the combination of overlapping SEs and non-overlapping CIs can provide the full picture. This is why we report both SEs and CIs. We have now provided a more detailed explanation on how the SEs and CIs should be interpreted.
Please also see this webpage (which we find useful) for more details: https://www.graphpad.com/support/faq/spanwhat-you-can-conclude-when-two-error-bars-overlap-or-dontspan/
Page 4. Demographics: Per APA 7, numbers under 10 should be spelled out.
We have now made these changes in the relevant section, where we believe they do not hinder readability.
Page 10: this study was the first in the literature to found that White participants’ CA relationship differs for White and Asian targets using a sequential lineup.
We have now corrected the grammar in this sentence. Thank you for noting this.
Reviewer 3 Report
Comments and Suggestions for AuthorsThank you for the opportunity to review the manuscript entitled “The Confidence of White Eyewitnesses is Better Calibrated 2 with White Targets than Asian Targets.” I like the research and believe it makes a meaningful contribution, but meaningful revisions need to be made as well.
The manuscript was generally well written, but perhaps overly technical and without proper emphasis on the meaning and applications of the research. I found the introduction to the nature and history of indices of reliability to be a bit brief and underdeveloped. To be fair it was concise while my writing style is comparatively rambling, but I think more detailed information on where these come from, how they are used in the criminal justice system, and deeper explanation (perhaps with narrative examples) of their importance would add a lot of legal and psychological context to a paper that sometimes reads a bit overly technical. I have similar thoughts on the discussion. The paper, as written, would be of interest to a forensic psychologist who specializes in the study of eyewitness identifications, but I have a hard time seeing anyone from the criminal justice system or another field of academia picking up this one and being able to follow along and understand why it is meaningful. How this relates to actual criminal evidence, legal procedure, and jury decision making is lost.
The definition of "Asian" should be put forward earlier, as it is only clear that you mean East Asian halfway through the manuscript. The term can be used by different people in different ways: South Asia (UK), East Asia or Southeast Asia (US), Central Asia, or West Asia/Middle East. People from these areas have different cultures, different levels of exposure to other races, and show phenotypic variation in facial features.
The wide range of places the participants came from creates a major confound, given that cross-race bias is highly affected by the amount of experience that the participants have had with members of the target race. This creates an issue within groups such as potential differences between Asian participants originally from the US vs those from Asian countries, it also creates problems for comparing the groups; most White participants were from the UK which has a White majority of 82% while the US has a White majority of 61.6%. Thus, the Asians from the US would have had much more experience with Whites and other races than Whites from the UK. This disparity of experience would have still been the case if all of the Asians were from the UK or all of the Whites were from the US, because then you could still draw conclusions about how the relationship plays out within a given society. Having both groups made up of participants spread out across the world with completely different levels and types of cross-race experiences makes it difficult to draw group level conclusions. This limitation was acknowledged, but only in regards to the Asian participants (not Whites from different countries) and not nearly strongly enough or in enough detail. If possible I would also like to see analyses with geographical variables included (at least for the geographical groups large enough to support those analyses).
The retention interval was extremely brief. This not only threatens generalizability due to mundane concerns of ecological validity or the strength of the results, but may have had an effect on the patterns of results themselves.
It would have been interesting to see if a group of White participants would have used the same criteria and chosen the same images during the match-to-description pre-testing as the East Asian participants did.
Any chance the lineups were recorded? Even though it would be a lot of extra work at this point, I would be interested to see if another (psychological) indicia of reliability was affected by the cross-race variable: response time.
Stats have never been my strongest area, but everything seems in order in terms of the choice of tests and their interpretation.
Author Response
Dear Reviewer,
We sincerely thank you for your comments and suggestions. We found them extremely helpful in improving the quality and clarity of our manuscript. We have marked our responses to reviewer comments and the corresponding changes in the manuscript in purple text.
-The Authors
The manuscript was generally well written, but perhaps overly technical and without proper emphasis on the meaning and applications of the research. I found the introduction to the nature and history of indices of reliability to be a bit brief and underdeveloped. To be fair it was concise while my writing style is comparatively rambling, but I think more detailed information on where these come from, how they are used in the criminal justice system, and deeper explanation (perhaps with narrative examples) of their importance would add a lot of legal and psychological context to a paper that sometimes reads a bit overly technical.
Thank you for this review. We have now added sections highlighting the importance of these research questions for both legal and psychological contexts, along with an illustrative example.
I have similar thoughts on the discussion. The paper, as written, would be of interest to a forensic psychologist who specializes in the study of eyewitness identifications, but I have a hard time seeing anyone from the criminal justice system or another field of academia picking up this one and being able to follow along and understand why it is meaningful. How this relates to actual criminal evidence, legal procedure, and jury decision making is lost.
Thank you for this review. We have now added sections to the Discussion explaining our findings in relation to legal decision-making.
The definition of "Asian" should be put forward earlier, as it is only clear that you mean East Asian halfway through the manuscript. The term can be used by different people in different ways: South Asia (UK), East Asia or Southeast Asia (US), Central Asia, or West Asia/Middle East. People from these areas have different cultures, different levels of exposure to other races, and show phenotypic variation in facial features.
We have now clarified at the end of the Introduction that our stimuli were East Asian.
The wide range of places the participants came from creates a major confound, given that cross-race bias is highly affected by the amount of experience that the participants have had with members of the target race. This creates an issue within groups such as potential differences between Asian participants originally from the US vs those from Asian countries, it also creates problems for comparing the groups; most White participants were from the UK which has a White majority of 82% while the US has a White majority of 61.6%. Thus, the Asians from the US would have had much more experience with Whites and other races than Whites from the UK. This disparity of experience would have still been the case if all of the Asians were from the UK or all of the Whites were from the US, because then you could still draw conclusions about how the relationship plays out within a given society. Having both groups made up of participants spread out across the world with completely different levels and types of cross-race experiences makes it difficult to draw group level conclusions. This limitation was acknowledged, but only in regards to the Asian participants (not Whites from different countries) and not nearly strongly enough or in enough detail. If possible I would also like to see analyses with geographical variables included (at least for the geographical groups large enough to support those analyses).
Thank you for this review. We have opted not to elaborate on a direct US–UK comparison in the manuscript to avoid detracting from the primary contributions of the study and we elaborate our thought process below. However, should the reviewer and editor deem this discussion essential, we would be willing to incorporate it.
Lee and Penrod (2022) reported a robust cross-race effect (CRE) across the literature, noting that 59.1% of the studies included White majority-race samples. Consequently, evidence for the CRE has been most consistently observed in studies with White participants. Töredi et al. (2024) reclassified all studies from Lee and Penrod’s (2022) meta-analysis that included White participants according to whether those participants constituted a racial majority in their country of recruitment, without restricting classification to a specific target-face race. Notably, in all such studies, White participants were members of the racial majority (as is the case in the present UK sample). Of these studies, 93% reported a CRE, whereas only 6% did not. Importantly, because these studies were conducted across multiple White-majority countries with varying levels of exposure to other-race faces, these findings suggest that recruiting White participants in the United States—rather than the United Kingdom—would be unlikely to alter the CRE observed among White participants in the present study.
With respect to the Asian participants, we agree with the reviewer that recruiting participants from East Asia, with more controlled exposure to White faces, would be ideal. However, such recruitment is difficult to achieve using online platforms (we have in fact tried to do this in other experiments). Data collection via MTurk necessarily limited the extent to which participant exposure histories could be controlled, a constraint common to much of the CRE literature, as cross-national collaborations that would enable large, tightly controlled samples remain relatively rare. We acknowledge that Asian participants’ exposure to White faces in the United States may have influenced our failure to observe a CRE in this group, a point we already address in the Discussion. Nevertheless, we do not expect that recruiting Asian participants from the United Kingdom, rather than the United States, would substantially change this pattern. Töredi et al.’s (2024) reanalysis of studies including Asian participants in White majority countries found that 58.62% reported a CRE, whereas 41.37% reported no CRE, indicating considerable variability—and near-chance outcomes—among non-White participants living in White-majority countries.
Lee, J., & Penrod, S. D. (2022). Three‐level meta‐analysis of the other‐race bias in facial identification. Applied Cognitive Psychology, 36(5), 1106–1130. https://doi.org/10.1002/acp.399
Töredi, D., Mansour, J. K., Jones, S. E., Skelton, F., & McIntyre, A. (2024). The impact of minority-race status on the cross-race effect: A critical review. Perspectives on psychological science: a journal of the Association for Psychological Science, https://doi.org/10.1177/17456916251345459
The retention interval was extremely brief. This not only threatens generalizability due to mundane concerns of ecological validity or the strength of the results, but may have had an effect on the patterns of results themselves.
We want to highlight that our retention interval is similar to that of published eyewitness memory research (e.g., Mansour, Beaudry, & Lindsay, 2017; Sauerland & Sporer, 2009; Smith, Wilford, Quigley-McBride, & Wells, 2019). Furthermore, the goal is to ensure that participants rely on long-term memory, and a 30-second cognitively engaging distractor task should be sufficient to achieve this by disrupting working memory reliance and the rehearsal of encoded faces. In real-world contexts, individuals rarely have an uninterrupted, close-range view for encoding, as is often provided in eyewitness paradigm mock-crime videos. If the retention interval is excessively long, however, this may create a situation in which witnesses’ memory strength is unrealistically low and therefore not ecologically valid.
It would have been interesting to see if a group of White participants would have used the same criteria and chosen the same images during the match-to-description pre-testing as the East Asian participants did.
We used the same match-to-description technique when constructing the White lineups (although this was done outside the present study). Specifically, White participants made the match-to-description judgments for White lineup members, just as East Asian participants did for East Asian lineup members in the current study. Thus, the lineup construction procedure was held constant across the target race manipulation. The lineup constructor was selected to be of the same race as the lineup members because the constructor’s own cross-race effect can influence selection decisions. We direct the reviewer to the following references, which have empirically examined this issue and support this approach, which are also now cited in the manuscript.
Brigham, J.C., & Ready, D.J. (1985). Own-race bias in lineup construction. Law Hum Behav 9, 415–424 https://doi.org/10.1007/BF01044480
Akan, M., Starns, J. J., Cohen, A. L., & Reid, A. E. (2025). The role of race in lineup construction. Journal of Experimental Psychology: General, 154(12), 3251–3283. https://doi.org/10.1037/xge0001836
Any chance the lineups were recorded? Even though it would be a lot of extra work at this point, I would be interested to see if another (psychological) indicia of reliability was affected by the cross-race variable: response time.
Response time has rarely been used as an index with sequential lineups because it is not clear how exactly response time can be measured in such cases. Less than a handful of published studies so far has measured response times with sequential lineups to the best of our knowledge (Sporer, 1993; Töredi et al., 2025) .
This research formed part of a broader doctoral thesis that also examined response time as an additional index; the corresponding results are available online for review (unmasked link available upon request). However, we chose to focus the present manuscript exclusively on confidence, as it is more heavily relied upon by jurors (Garrett et al., 2020) and exhibits greater and more consistent reflective validity than response time (Carlson et al., 2025). This focus is now explicitly noted in our Introduction.
Carlson, C. A., Lockamyeir, R. F., Carlson, M. A., Goodsell, C. A., Jones, A. R., Wooten, A. R., & Brand, Z. T. (2025). Comparing the strength of the confidence-accuracy versus response time-accuracy relationship for eyewitness identification. Scientific Reports, 15(1), 11064. https://doi.org/10.1038/s41598-025-96224-y
Garrett, B. L., Liu, A., Kafadar, K., Yaffe, J., & Dodson, C. S. (2020). Factoring the role of eyewitness evidence in the courtroom. Journal of Empirical Legal Studies, 17(3), 556-579. https://doi.org/10.1111/jels.12259
Round 2
Reviewer 3 Report
Comments and Suggestions for AuthorsThank you for the opportunity to review a revised version of this paper. Although I disagree with some of the authors' assertions in their response (described below), I found the revised version to much more clearly spell out how the research is relevant to criminal proceedings involving cross-race identifications. On that account I am satisfied.
In regards to the response to identification accuracy and confidence of Whites identifying East Asians varying by country due to different levels of exposure the authors' arguement did not completely satisfy me. Lee and Penrod (2022) conducted no analyses on variation in cross-race bias as a function of nation. Further, and as was pointed out by the authors, an examination of the studies in the meta-analysis Whites were from majority White countries, almost exclusivley in North America and Westen Europe. Thus, they were such a strong majority that it could have easily created a ceiling effect, and statistically significant results may have been found if they included countries where Whites were not the majority (e.g., South Africa, Namibia, Belize). Since you are comparing two White majority countries this may not have been an issue, save for two things: 1) some of the studies go back to the 1970s and since that time Whites in the US have gone from a nearly 90% majority to 61%, and thus the comparison may not be valid. 2) The authors' argument seems to equate the number of East Asian people in these nations as being directly inversely related to the number of White people. East Asians are a minority of 6.1% in the US (see US census) and in the UK somewhere between 0.8% and 2% (the UK census categories are a bit unclear). Thus, even if the proportion that Whites were a majority in these countries were equal, the size of the East Asian minorities would not be. So, we'll leave it to the editor's discretion as to request it be included as a limitation.
In regards to retention interval I can appreciate using methodology that has already been validated and accepted within the scientific community. However, this does not overcome the limitation that in an actual criminal investigation a person would never be asked by police officers to make an identification of a perpetrator a few seconds following initial exposure. It is not a plausible scenario; in my expert witness work I have been involved in cases where identifications happened anywhere between 3 months to 2 hours after the incident, but never sooner. Given what is known about memory decay and the forgetting curve the results may not be applicable to real-life eyewitness identifications. Again, we'll leave it to the editor's discretion as to request it be included as a limitation.
The comments about cross-racial match-to-description approaches and response time were merely random thoughts, not requested revisions.
Author Response
In regards to the response to identification accuracy and confidence of Whites identifying East Asians varying by country due to different levels of exposure the authors' arguement did not completely satisfy me. Lee and Penrod (2022) conducted no analyses on variation in cross-race bias as a function of nation. Further, and as was pointed out by the authors, an examination of the studies in the meta-analysis Whites were from majority White countries, almost exclusivley in North America and Westen Europe. Thus, they were such a strong majority that it could have easily created a ceiling effect, and statistically significant results may have been found if they included countries where Whites were not the majority (e.g., South Africa, Namibia, Belize). Since you are comparing two White majority countries this may not have been an issue, save for two things: 1) some of the studies go back to the 1970s and since that time Whites in the US have gone from a nearly 90% majority to 61%, and thus the comparison may not be valid. 2) The authors' argument seems to equate the number of East Asian people in these nations as being directly inversely related to the number of White people. East Asians are a minority of 6.1% in the US (see US census) and in the UK somewhere between 0.8% and 2% (the UK census categories are a bit unclear). Thus, even if the proportion that Whites were a majority in these countries were equal, the size of the East Asian minorities would not be. So, we'll leave it to the editor's discretion as to request it be included as a limitation.
As noted in our previous response, Lee and Penrod did not conduct this analysis. However, Töredi et al. (2024) examined the generalizability of the CRE across the studies included in Lee and Penrod’s meta-analysis and reported that the CRE literature predominantly comprised White-majority samples. Thus, there are no empirical data on White minority participants within that literature. Although we agree that majority vs. minority status could plausibly influence the CRE, there is currently insufficient empirical basis to make this claim in the manuscript. Moreover, Töredi et al. (2024) found that White participants reliably showed a CRE regardless of nationality and target race (e.g., despite that one would expect differences in contact with Black vs. Asian faces). This suggests that CRE differences between our UK and US White participants would be unlikely; therefore, we did not include this as a limitation. That said, we agree this is a more substantial limitation for our Asian participants, given evidence that the CRE is less robust in non-White samples and may be more sensitive to sociocultural context and contact (though contact accounts for <2% of variance in the CRE; see Singh et al., 2021). Accordingly, we now note this limitation in the manuscript (p. 21).
In regards to retention interval I can appreciate using methodology that has already been validated and accepted within the scientific community. However, this does not overcome the limitation that in an actual criminal investigation a person would never be asked by police officers to make an identification of a perpetrator a few seconds following initial exposure. It is not a plausible scenario; in my expert witness work I have been involved in cases where identifications happened anywhere between 3 months to 2 hours after the incident, but never sooner. Given what is known about memory decay and the forgetting curve the results may not be applicable to real-life eyewitness identifications. Again, we'll leave it to the editor's discretion as to request it be included as a limitation.
We understand this argument, and we have noted this limitation and the need for future research on page 21. This limitation is one that applies to perhaps most of the eyewitness identification literature—because of the resource demands of using more plausible delays—but it is a limitation, nonetheless.

