The Variation in IRT in Different Ethnic Groups in England—Implications for a Newborn Screening Programme for CF in Diverse Multiethnic Populations
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThis is a valuable and timely analysis highlighting important inequities in CF newborn screening. Clarifying the limitations, particularly around isoform mechanisms and small subgroup sizes, would strengthen the manuscript. The rationale for UK screening could be better described. CF is a genetic disease, but the authors seem to have little focus on this.
This study leverages a very large dataset to robustly assess ethnic variation in IRT values and provides a clear comparison between the two DELFIA analysers. The identification of clinically meaningful disparities, particularly the lower PPV and higher false‑positive rates in Black African and Indian infants, adds important value to the CF screening literature. The authors offer a plausible biological hypothesis regarding IRT isoforms, which may guide future research. However, the mechanistic explanation for analyser differences across ethnic groups remains speculative, as no direct biochemical assessment of IRT isoforms was performed.
Some ethnic subgroups have small numbers, limiting confidence in CF:CFSPID ratio comparisons (e.g., Pakistani group). The use of parent‑reported ethnicity introduces potential misclassification, especially for mixed‑heritage infants. As a retrospective study, conclusions are constrained by available data, and the generalisability beyond the UK population is uncertain.
The manuscript could be strengthened by briefly discussing whether next‑generation sequencing (NGS) might help mitigate the ethnic disparities described. NGS could reduce false positives in populations with physiologically higher IRT but low CF incidence, improve equity by detecting a broader and more diverse range of CFTR variants than current variant panels, and offer clearer stratification between CF and CFSPID outcomes.
NGS could be integrated as a tier‑2 test triggered by elevated IRT, complementing, rather than replacing, the existing protocol. While cost, turnaround time, and laboratory capacity would need consideration, such an approach could enhance diagnostic accuracy and reduce unnecessary follow‑up in minority ethnic groups. In this context one could mention briefly that AI‑supported NGS could help standardise pathogenicity classification across diverse ethnic groups. Including this forward looking perspective would add useful depth to the Discussion.
Author Response
Thank you for your valuable feedback. In particular around the genetics and the fact we have not mentioned NGS. We will address this.
Reviewer 2 Report
Comments and Suggestions for AuthorsThe manuscript describes cystic fibrosis newborn screening in England and more specifically the immunoreactive trypsinogen (IRT) cut-offs used and the variation seen in 95% centiles and therefore positive predictive values for screening different ethnic groups. The data is useful for physicians and genetic counselors when discussing risk associated with a positive newborn screen result. The laboratories use an interesting method to determine a national cut-off which presumably needs a lot of coordination.
Reviewer comments:
Are the cut-offs in Figure 1 examples of cut-offs (e.g., 58 and 65 µg/L)? Presumably these change every 3 months? Is the ‘safety net’ cut-off also recalculated each quarter? How was the value 120 µg/L decided on?
How do IRT kit lot changes affect the cut-offs that are determined for each quarter? Presumably not all labs use the same kit lots during the same quarter so how do the ‘national’ cut-offs work for all labs? Kit lot changes and seasonality have been shown to influence IRT values which is why in the US the recommendation is to have a floating IRT cut-off. What cut-off ranges have you set throughout the years of the study?
Are differences in IRT levels observed based on age of specimen collection? Are all newborn specimens that are submitted to your programs collected at 5 days of age? Are the specimens collected at the hospital of birth after discharge or at pediatricians’ offices?
It should be noted that traditional CFTR screening panels usually have lower detection rates in minority groups including black and Asian populations. This can lead to false negative screening results in these groups. Therefore, it could be possible that screening for only 50-100 variant panel may miss rare variants in these populations. Is it possible that CF cases in Table 2 are an underestimation? Related to this question, have any false negative results been reported?
Presumably if there is an increase in black African babies in the birth population, it would lead to an increase in the overall 99.5% centile value. Would this place other babies at risk of missing the cut-off?
The addition of CF and CFSPID incidence for the various ethnicities to Table 2 would be very helpful.
Author Response
Thank you for your valuable feedback.
With regards the 99.5th centile cut-off; the dashed line represents the 99.5th centile, irrespective of ethnicity, for this 4 year period of data collection.
Presumably these change (99.5th cut-offs) every 3 months?
No this isn't the case. We are reluctant to regularly change the cut-offs as we feel/know that it comes with a risk associated with Labs having to alter cut-offs in software packages. We are also aware that particular lot numbers may add a temporary bias. We change cut-offs only when a prolonged 2-3 quarters worth of data indicates it. We give particular weight to the number of samples referred for DNA testing which we believe should be close to 0.5%. As example, we set out in the AutoDeflia group with a cut-off of 62 in 2020. This was increased to 65 in 2023 but has not been altered again since. The GSP group has been altered x3 duirng the same period.
How was the value 120ug/L decided on?
This value was set in 2007 and at the time was thought to be close to the 99.9th centile. We have since shown that the 99.9th is actually much lower (closer to 100ug/l). However, rather than change the value to actual 99.5th the CF advisory board in the UK have thought that it has worked well at 120ug/l. If time/funding were available it would be good to carry out some research around this (this could be tied into the work required around false negatives- see below).
Presumably not all labs use the same kit lot
Correct. When we get the data submitted we insist on having the lot numbers used provided. For each group there tends to be 2-3 in use kit lots and by splitting the data into kit lots it helps us to identify if there is a particular lot number that looks out of alignment. Interestingly we have not noticed any significance difference with regards seasonality although this could be due to our reasonably boring climate!
Are differences in IRT levels observed based on age of specimen collection?
In 2020 when we started collecting this data, we did also look whether age of samples had an impact. After a year it had little impact and we focused on ethnicity where it was clear that there were large differences. We only included data points from babies who had samples taken days 4-21. Most samples are taken at home but there will be a small proportion who are still in hospital at day 5 and will have the sample taken then.
Is it possible that CF cases in Table 2 are an underestimation? Have any false negatives been reported?
Yes! In the UK reporting of false negatives could be improved. While paediatricians are quick to inform labs of CF cases that were screen normal, they are often then reluctant to fill in the required report so that it can be verified as a true CF case rather than a CFSPID. It has been estimated that the false negative rate is about 5%, however, it would be a very useful piece of work if someone could dig deeper into this and confirm it's accuracy.
An increase in African babies in the birth population, would lead to an increase in the overall 99.5th centile...would this place other babies at risk of missing the cut-off?
In theory yes. However, we have not so far seen the 99.5th cut-offs climbing. Although the AutoDelfia one has risen from 62 to 65, the GSP has hovered up and down between 56 and 58. It is a good point, I believe if in the future there was a trend emerging of the 99.5th rising we would look at ethnicity (on the back of this paper) alongside false negative reporting to judge it's impact.
Reviewer 3 Report
Comments and Suggestions for AuthorsThe variation in IRT in different ethnic groups in England – imlications for a newborn screening programme for CF in diverse multiethnic populations
Overall:
The manus addresses a very interesting and important issue as populations worldwide are getting increasingly mixed.
The NBS-CF algorithm follows ECFS standards and is well described and implemented.
Abstract:
Precise and clear
Introduction:
Explains very well the NBS programme and changes over time. It seems at if it started with a national 99.5th of 70 mikrogr/L I 2007, and then the following years adjusted by the various laboratories. Have the actual calculations been calculated for 2007 – for comparison? It is reasonable to include 4 years (2020-2024) from the introduction of a national cut-off.
Materials and methods:
The two different methods measuring IRT are described – as well as the ongoing follow-up.
Although not all laboratories reported ethnicity – a vast majority 10 of 13 were represented.
The CF outcome data over a 10-year period from 5 laboratories was chosen to assess ethnicity impact on PPV and CFSPID.
Results:
Data is well presented in Tables and Figures.
Figure 2 and 3: numbers, marking and dotted line most be more precise
GSPP has a narrower confidence interval – which is not commented on
The CF Outcome data shows that PPV>50 % is safe above ECFS guidelines. However, this does not include black African and Indian groups.
Discussion:
Differences between GSP and AD: any speculations to uniform across all laboratories?
Is likelihood of having CF in Europeans as compared to black Africans only 5 times?
Any speculations why Pakistanians with high PPV´s have that high CF:CFSPID ratio?
The low incidence of CF in black Africans and very low CF:CFSPID ratio is difficult to explain.
As all CF populations worldwide may become more diverse and mixed algorithms have to adjust along the way.
The suggested solution the let the counsellors take into account which ethnicity they are facing seems to be the best possible solution. However, it has to be remembered that ancestors may be of unknown origin (to the patient/parents). Recently new ethnicities were shown to have CF (inuit) – due to European ancestors many generations back.
Conclusions:
Black Africans are mentioned due to high IRT
– Indians (and black Africans) with very low PPV should be included in the conclusion
Author Response
Thank you for your valuable review.
Have the actual calculations been calculated for 2007 comparison?
No they haven't. The cut-off set in 2007 was based upon historical data and once we commenced screening it became clear that the actual 99.5th centile was lower than 70ug/l. This was exaggerated when many labs switched to the GSP analyser which as you can see in our data, runs more than 10% lower than the AutoDelfia.
Is likelihood of having CF in Europeans as compared to black Africans only 5 times?
This was taken from reference 10.
Any speculations why Pakistanians with high PPV have that high CF:CFSPID ratio?
It is interesting but with such small numbers it is difficult to assess it's importance. What we do know is that a number of the Pakistani cases of CF were from two particular cities in England and were homozygous for rare (in the UK population) mutations. We could speculate that these are the result of consanguineous marriages.
In your closing note Indians with very low PPV should be included in the conclusion.
We agree and will update this to include it.

