Diagnostic Performance of Biomarkers for Perioperative Hypersensitivity Reactions in Adults: A Systematic Review and Meta-Analysis on Tryptase and Histamine Dosing
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThank you for the opportunity to review this manuscript. The authors present a systematic review and meta-analysis evaluating the diagnostic performance of perioperative biomarkers—primarily tryptase and histamine—for perioperative hypersensitivity reactions (POH/POA). The study correctly identifies that while fixed thresholds for tryptase offer excellent specificity (strong rule-in value), dynamic thresholds significantly boost sensitivity. This is a clinically relevant synthesis that addresses the ongoing challenge of differentiating allergic from non-allergic intraoperative shock. However, to ensure the robustness of the conclusions and the scientific soundness of the meta-analysis, several methodological clarifications, particularly regarding threshold heterogeneity and histamine data interpretation, must be addressed before publication.
Major Concerns
- The manuscript presents conflicting interpretations regarding the clinical utility of histamine. In the Abstract (Line 29), the estimates are correctly described as "unreliable," yet the Conclusion (Line 32) states that histamine shows "numerically promising performance". Given that the meta-analysis only includes 3 studies for histamine, highlighting it as "promising" overstates the strength of the evidence and risks small-study bias.
Action: Please revise the conclusion to adopt more neutral, objective language (e.g., "Preliminary results suggest..."). The discussion (Lines 395-400) should explicitly acknowledge that current evidence is insufficient for clinical recommendations regarding histamine. - Histamine has a highly transient half-life and typically normalizes within 30 minutes after release. Since the included studies are predominantly retrospective, the timing of blood sampling likely varied substantially.
Action: Delayed sampling would inevitably lead to false-negative results, undermining the biological validity of the pooled sensitivity. This biological limitation and its impact on the reliability of the pooled estimates must be discussed explicitly in the Limitations section. - The diagnostic reference standards (e.g., clinical diagnosis vs. allergist confirmation) and the composition of the control groups vary significantly across the included studies. Specifically, mixing asymptomatic patients with patients experiencing non-allergic shock in the control group can directly dilute the precision of the specificity calculations.
Action: Please detail the control group compositions and reference standards for each study. Discuss how this heterogeneity might influence the specificity estimates and the overall evaluation of tryptase as a "rule-in" tool. - The manuscript repeatedly references the international consensus formula for dynamic thresholds (sAT > 1.2 x sBT + 2) (Lines 78-80). However, it is unclear if all included studies applied this formula uniformly, or if different dynamic definitions were pooled together.
Action: Please clarify how threshold heterogeneity was handled statistically. If different algorithms or sampling baselines were used across studies, discuss how this variability impacts the reported 77.2% pooled sensitivity. Subgroup analyses separating different dynamic threshold definitions should be considered if feasible. - To meet PRISMA guidelines and ensure reproducibility, please provide the complete search strings used for PubMed, WOS, and Cochrane, and specify the exact date of the final search.
- Table 1 is currently too dense, listing multiple thresholds (both fixed and dynamic) from single studies together. Consider separating fixed and dynamic thresholds into distinct columns or different tables to improve readability.
- Throughout the Results section, all reported sensitivity, specificity, and AUC values must be consistently accompanied by their 95% confidence intervals (CIs).
- (1) For Figure 2, please clearly define the clinical relevance of the "Survival curves" and the "Youden index," and briefly explain the derivation of the 12.68 ng/mL optimal cutoff in the caption. (2) For Figures 3-5, explain the significance of the "confidence ellipse" regarding inter-study heterogeneity rather than just repeating sensitivity values.
- Ensure that sAT and sBT are clearly defined upon their first appearance in both the Abstract and Introduction, avoiding informal terms like "acute tryptase". Additionally, ensure p-values and decimal formatting strictly follow MDPI style guidelines.
- Please verify that references to textbooks or non-peer-reviewed websites are updated with recent peer-reviewed literature where possible. Finally, double-check all figure numbering in the text (especially pages 7-9) to ensure they correspond correctly with the provided figures.
Author Response
Thank you very much for the appreciation of the clinical significance of our work, highltighing that while fixed thresholds for tryptase offer excellent specificity (strong rule-in value), dynamic thresholds significantly boost sensitivity for the differentiating allergic from non-allergic intraoperative shock. We consider this review is helpful for allergologists and anesthesiologists to adress different clinical situations. Please find below the methodological clarifications:
- The manuscript presents conflicting interpretations regarding the clinical utility of histamine. In the Abstract (Line 29), the estimates are correctly described as "unreliable," yet the Conclusion (Line 32) states that histamine shows "numerically promising performance". Given that the meta-analysis only includes 3 studies for histamine, highlighting it as "promising" overstates the strength of the evidence and risks small-study bias.
Action: Please revise the conclusion to adopt more neutral, objective language (e.g., "Preliminary results suggest..."). The discussion (Lines 395-400) should explicitly acknowledge that current evidence is insufficient for clinical recommendations regarding histamine.
Thank you. We have changed the conclusion about histamine to “Preliminary results suggest that histamine might have optimal diagnostic performance, but estimates are severely limited by small sample sizes.”, a conclusion that shows less certainty, given the data it is based on. – Abstract Line 32.
Also, in Lines 380-381 we have added: “Thus, current evidence is insufficient for clinical recommendations regarding histamine.”
Lines 498-500: Preliminary results suggest that histamine might have optimal diagnostic performance, but estimates are severely limited by small sample sizes and there is currently no strong superiority over tryptase.
2. Histamine has a highly transient half-life and typically normalizes within 30 minutes after release. Since the included studies are predominantly retrospective, the timing of blood sampling likely varied substantially.
Action: Delayed sampling would inevitably lead to false-negative results, undermining the biological validity of the pooled sensitivity. This biological limitation and its impact on the reliability of the pooled estimates must be discussed explicitly in the Limitations section.
We have added this in line 472-475, to explain better sampling inconsistency effect on diagnosis: „Since the included studies are predominantly retrospective, the timing of blood sampling might have varied substantially, and, especially for histamine, that has a short half-time, delayed sampling could lead to false-negative results.”
3. The diagnostic reference standards (e.g., clinical diagnosis vs. allergist confirmation) and the composition of the control groups vary significantly across the included studies. Specifically, mixing asymptomatic patients with patients experiencing non-allergic shock in the control group can directly dilute the precision of the specificity calculations.
Action: Please detail the control group compositions and reference standards for each study. Discuss how this heterogeneity might influence the specificity estimates and the overall evaluation of tryptase as a "rule-in" tool.
Variability in clinical diagnosis is discussed at line 433, the following paragraph explains variability in the positive diagnosis.
Control group variability is discussed at Line 447, where we cite the study demonstrating that resuscitaion of other forms of shock do not influence mediators concentrations.
Also, the criteria defined for positivity and negative history of POH/POA are explained in Table 1, page 6.
4. The manuscript repeatedly references the international consensus formula for dynamic thresholds (sAT > 1.2 x sBT + 2) (Lines 78-80). However, it is unclear if all included studies applied this formula uniformly, or if different dynamic definitions were pooled together.
Action: Please clarify how threshold heterogeneity was handled statistically. If different algorithms or sampling baselines were used across studies, discuss how this variability impacts the reported 77.2% pooled sensitivity. Subgroup analyses separating different dynamic threshold definitions should be considered if feasible.
Thank you for your thoughtful comment on the handling of dynamic thresholds in our meta-analysis. We appreciate the opportunity to clarify this aspect of our methods and results, as it is a critical point in interpreting the diagnostic utility of tryptase.
Regarding the international consensus formula for dynamic thresholds (sAT > 1.2 × sBT + 2), we agree that this is a widely recommended standard (as referenced in Vitte et al., 2018, and Ebo et al., 2021, included in our dataset). However, our review of the included studies revealed that not all applied this formula uniformly. Specifically, out of the 5 studies contributing to the dynamic tryptase subset (12 entries total), 4 studies (Barrteo 2017, Ebo 2021, Vitte 2018, Takazawa 2021) used the 1.2 × sBT + 2 formula or close variants (ΔT (sAT - sBT) > 3.2 ng/mL in Ebo 2021, which is a similar baseline-adjusted approach). The remaining study (Haraguchi 2024) employed slightly different definitions, such as sAT > 1.52 × sBT and sAT > 2 × sBT. All studies used serum baseline tryptase (sBT) measured at a stable point (>24 hours post-reaction or pre-event), but sampling timing for acute tryptase (sAT) varied slightly (within 2–4 hours post-onset in most cases). This variability in dynamic definitions was noted in our methods section.
In the meta-analysis we have pooled all results from the included studies (raw data). To handle threshold heterogeneity statistically, we employed the bivariate random-effects model (reitsma function in the mada package in R), which pools sensitivity and specificity while accounting for between-study variation in thresholds through variance components. This is described in the mehtods section. This model does not require uniform numeric thresholds, making it suitable for dynamic definitions. However, as you correctly point out, pooling different dynamic algorithms could introduce some bias, potentially overestimating sensitivity if more lenient definitions (sAT > 1.52 × sBT) were included alongside the consensus formula.
The reported 77.2% pooled sensitivity (95% CI 74.9–79.3%) for dynamic thresholds is likely influenced by this heterogeneity. For instance, studies using the consensus 1.2 × sBT + 2 formula (Ebo 2021) tended to show slightly higher performance than those with stricter definitions (Haraguchi 2024). This variability may lead to an upward bias in the pooled sensitivity, as the model assumes random effects but does not explicitly subgroup by definition type. The low correlation in variance components (1.000 for both tsens and tfpr) indicates the heterogeneity is captured, but the small number of studies (n=5) limits our ability to quantify the exact impact.
We considered subgroup analyses separating different dynamic threshold definitions (consensus vs. variants). However, with only 5 studies and overlapping definitions in some, this was not feasible without risking over-stratification and unstable estimates. As a sensitivity check, we re-ran the model excluding Haraguchi 2024 (the main outlier with stricter definitions), resulting in a slightly higher pooled sensitivity of 78.5% (95% CI 75.2–81.4%), suggesting minimal impact from heterogeneity in this small dataset.
Thank you for helping us improve the clarity and robustness of our reporting.
We have added in the Results at line 206: The consensus formula sAT> 1.2 x sBT +2, was evaluated in 5 out of the 7 studies on tryptase dosing, each of them demonstrating sensitivity 75-78.6% and specificity 86-100% (Table 1). Each study evaluated the performance of the consensus formula compared to other thresholds that varied among studies, each using different other criteria for positivity. The consensus formula represented the most frequently used criteria for positivity.
5. To meet PRISMA guidelines and ensure reproducibility, please provide the complete search strings used for PubMed, WOS, and Cochrane, and specify the exact date of the final search.
We have added in the Methods section: „with last search on January 22nd 2026 (Supplementary material) at lines 93-94, and provided the strings as Supplementary material, with exemplification. The databases were re-checked on 13.03.2026, to confirm the figures included in the analysis. In Figure 1, steps of the identification of the papers are presented one by one for the databases.
Thank you for the observation, the methodology is more accurate now and allows reproducibility to check for results.
6. Table 1 is currently too dense, listing multiple thresholds (both fixed and dynamic) from single studies together. Consider separating fixed and dynamic thresholds into distinct columns or different tables to improve readability.
We know that this Table contains a lot of information. The actual format of Table 1 has several advantages: it includes all studies in order of the year, dyspalying when some actual thresholds were used, contains the criteria for positivity and specificity, and also, for the populations included in each study, the different cutoffs. We believe it is a comprehensive table that containts all information in one place, like a map. Thus, we prefer not to split into several other tables. If you consider this is really neccessary, we could do that, please let us know. In our vision, this comprehensive table on one page might be useful for the reader, rather than comparing tables on different pages. Is it really neccessary to change it? We are looking forward to your opinion.
7. Throughout the Results section, all reported sensitivity, specificity, and AUC values must be consistently accompanied by their 95% confidence intervals (CIs).
Thank you for this remark, we have now added 95% CI everywhere in the text.
8. (1) For Figure 2, please clearly define the clinical relevance of the "Survival curves" and the "Youden index," and briefly explain the derivation of the 12.68 ng/mL optimal cutoff in the caption. (2) For Figures 3-5, explain the significance of the "confidence ellipse" regarding inter-study heterogeneity rather than just repeating sensitivity values.
Thank you for your insightful feedback on the figures in our manuscript.
For Figure 2, we have expanded the caption to define the clinical relevance of the survival curves (which illustrate how the probability of a positive tryptase test changes with threshold levels, aiding clinicians in selecting cutoffs that minimize false positives while capturing severe cases) and the Youden index (a metric of overall diagnostic accuracy that balances sensitivity and specificity, useful for identifying thresholds that optimize test performance in resource-limited settings). We also briefly explained the derivation of the 12.68 ng/mL optimal cutoff as the value maximizing the weighted Youden index in the diagmeta model's multiple-threshold analysis, derived from the intersection of the summary ROC curve's tangent with the line of no discrimination.
The revised caption now reads: "Figure 2. Diagnostic performance of serum tryptase using fixed cutoffs in perioperative anaphylaxis (6 studies, 11 cutoffs). Top-left: Survival curves showing the probability of a positive test (elevated tryptase) versus threshold, clinically relevant for understanding how stricter cutoffs reduce false positives but may miss milder cases. Top-right: Weighted Youden index (sensitivity + specificity - 1, a measure of overall diagnostic accuracy balancing test performance) versus threshold (peak near 12.7 ng/mL). Bottom-left: Individual study ROC curves. Bottom-right: Summary ROC curve with pooled estimate (Sens 0.60, Spec 0.95 at optimal cutoff 12.68 ng/mL; AUC ≈0.72). The optimal cutoff was derived as the value maximizing the weighted Youden index in the diagmeta multiple-threshold model."
(2) For Figures 3–5 (dynamic tryptase, overall tryptase, and overall histamine SROC curves), we have updated the captions to explain the confidence ellipse as a visual representation of uncertainty in the pooled estimate, incorporating both within-study precision and inter-study heterogeneity (wider ellipses indicate greater variability across studies, such as differing patient populations or timing, which may affect generalizability). This replaces the previous repetition of sensitivity values with a focus on heterogeneity.
The revised caption for Figure 3 now reads: "Figure 3. Summary Receiver Operating Characteristic (SROC) curve for tryptase using dynamic thresholds in perioperative anaphylaxis (5 studies, 12 entries). The summary point indicates pooled sensitivity 77.2% (95% CI 74.9–79.3%) and false positive rate 11.5% (95% CI 6.5–19.5%), with AUC 0.774. The confidence ellipse represents uncertainty in the pooled estimate, incorporating inter-study heterogeneity (I² 0–43.8%); a narrower ellipse suggests relatively consistent performance across studies despite varying dynamic definitions."
Similar revisions have been applied to Figures 4 and 5.
These modifications make the figures more self-explanatory and directly address the clinical and statistical context, as requested.
9. Ensure that sAT and sBT are clearly defined upon their first appearance in both the Abstract and Introduction, avoiding informal terms like "acute tryptase". Additionally, ensure p-values and decimal formatting strictly follow MDPI style guidelines.
Thank you very much for the correction- we have checked and replaced the next appearance in the text with the acronysms.
10. Please verify that references to textbooks or non-peer-reviewed websites are updated with recent peer-reviewed literature where possible.
We have checked and updated for the Frech paper by Mertes, the volume number. All cited manuscripts have DOI numbers, as recommended by the journal Instructions for authors.
11. Finally, double-check all figure numbering in the text (especially pages 7-9) to ensure they correspond correctly with the provided figures.
We have checked again, thank you- we found a mistake. Now it is corrected (for Figure 3).
Please notice that, as recommended by Reviewer 2, we have strived to shorten the Discussion section. All pragraphs that were changed throughout the manuscript are marked in yellow.
Thank you for all the suggested corrections, we realise that the quality of our manuscript has been improved. Should there be any other suggestions, please do not hesitate to let us know. We will do our best to adress all issues.
Reviewer 2 Report
Comments and Suggestions for AuthorsThis manuscript presents a systematic review and meta-analysis assessing the diagnostic accuracy of tryptase and histamine for perioperative hypersensitivity/anaphylaxis (POH/POA). The topic is clinically important because perioperative anaphylaxis remains difficult to confirm diagnostically, and biomarker thresholds are inconsistent across studies.
The manuscript is generally well structured, follows PRISMA guidelines, and applies appropriate diagnostic meta-analysis methodology. The results suggesting higher diagnostic reliability of tryptase, particularly with dynamic thresholds, are clinically relevant.
However, the paper requires moderate revisions, mainly in:
- methodological transparency
- statistical reporting
- language improvement
- reference enhancing
Comments
1.Search strategy insufficiently reproducible
The search strategy is described broadly but not fully reproducible.
Example from the manuscript:
- “Diagnostic performance OR Diagnostic accuracy OR Sensitivity OR Specificity AND Biomarker OR Mediator…”
Issues:
- No full search strings per database
- No search dates
- No language restrictions stated
- No grey literature search
Recommendation:
Provide complete search syntax for at least PubMed in a supplementary file.
2.Study selection numbers are confusing
The manuscript states:
- 30 studies screened
- 8 included in meta-analysis
But later reports:
- 7 tryptase studies
- 3 histamine studies
This implies 10 studies, which creates confusion.
Likely explanation:
Some studies report both biomarkers, but this must be clarified.
3.Risk-of-bias reporting is insufficient
You state QUADAS-2 was used:
“Risk of bias was assessed using QUADAS-2 domains…”
But the manuscript lacks:
- a QUADAS-2 figure
- domain-level judgments per study
- discussion of high-risk domains
4.Histamine meta-analysis may be unreliable
Only three studies were included for histamine.
Diagnostic meta-analysis typically requires more studies.
Your discussion correctly notes this limitation but still reports high AUC values (~0.92).
This may represent small-study bias or overfitting.
5.Language corrections
“Most performed studies sustain dynamic thresholds over fixed values.”
“The biological functions of tryptase are not well understood…”
“arrythmia” → arrhythmia
“deffered” → deferred
“anaesthetic changes are not specific and require a thorough differential diagnosis” (sentence repetition)
- Discussion Quality
It is slightly long (≈9 pages).
Consider reduce by 15–20% by removing repetition on:
- mast cell biology
- mediator mechanisms
Focus more on:
- clinical diagnostic pathways
- guideline implications
- Figure reformatting
Graphics on Figure 2 are small, appears to be too basic, consider a better interpretation of them.
9.References: for a review 35 references are at lowest acceptable limit. Consider to add a few more of them.
Comments on the Quality of English Language5.Language corrections
“Most performed studies sustain dynamic thresholds over fixed values.”
“The biological functions of tryptase are not well understood…”
“arrythmia” → arrhythmia
“deffered” → deferred
“anaesthetic changes are not specific and require a thorough differential diagnosis” (sentence repetition)
Author Response
Thank you for the overall appreciation of our manuscript, especially pointing out the clinical utility, as perioperative anaphylaxis remains difficult to confirm diagnostically, and biomarker thresholds are inconsistent across studies, as well as higlighting the relevance of the results identified: the conclusion that higher diagnostic reliability of tryptase, particularly with dynamic thresholds, are clinically relevant.
Please find below a point-by-point explanation of the revisions of the manuscript, regarding methodological transparency, statistical reporting, language improvement and reference enhancing.
1.Search strategy insufficiently reproducible
The search strategy is described broadly but not fully reproducible.
Example from the manuscript: “Diagnostic performance OR Diagnostic accuracy OR Sensitivity OR Specificity AND Biomarker OR Mediator…”
Issues: No full search strings per database, No search dates, No language restrictions stated, No grey literature search
Recommendation: Provide complete search syntax for at least PubMed in a supplementary file.
Response: We have added in the Methods section: „with last search on January 22nd 2026 (Supplementary material) at lines 93-94, and provided the strings as Supplementary material, with exemplification. The databases were re- checked on 13.03.2026, to confirm for the figures included in the analysis. In Figure 1, steps of the identification of the papers are presented one by one for the databases. The supplementary material contains the search strings.
Thank you for the observation, it methodology is more accurate and allows reproducibility to check for results.
2.Study selection numbers are confusing
The manuscript states: 30 studies screened, 8 included in meta-analysis
But later reports: 7 tryptase studies, 3 histamine studies
This implies 10 studies, which creates confusion.
Likely explanation: Some studies report both biomarkers, but this must be clarified.
Response: This is correct, there are a total of 8 studies included. From these, seven studies evaluated tryptase and one did not. Two of the seven studies evaluating tryptase, also evaluated histamine for the included population.
We have rephrased the first paragraph of the Results, as: „From the 8 studies addressing biomarkers’ sensitivity and specificity, 7 studies evaluated tryptase and 3 studies evaluated histamine dosing performance for the diagnosis of POH/POA, using different fixed or dynamic thresholds for positivity throughout the studies (Table 1). In two of the studies, both tryptase and histamine dosing were performed.”, so that this might increase understanding of the reader.
3.Risk-of-bias reporting is insufficient
You state QUADAS-2 was used:
“Risk of bias was assessed using QUADAS-2 domains…”
But the manuscript lacks:
- a QUADAS-2 figure
- domain-level judgments per study
- discussion of high-risk domains
Response: Thank you for highlighting the need for more detailed risk-of-bias reporting. We acknowledge that our original manuscript referenced the use of QUADAS-2 for quality assessment but did not provide sufficient visual or tabular presentation of the results, as recommended in PRISMA-DTA guidelines. We have addressed this by adding a new supplementary figures with domain-level judgments per study, a traffic light and summary plot figure summarizing QUADAS-2 assessments across all included studies.
We have added the following text to the discussion: "High or unclear risk in the flow and timing domain across several studies (retrospective timing inconsistencies) could lead to overestimation of specificity, particularly in perioperative settings where acute sampling is critical."- lines 486-489, regarding statistical limitations of the study.
These additions provide a comprehensive overview of study quality and its implications, strengthening the manuscript's methodological rigor.
4.Histamine meta-analysis may be unreliable
Only three studies were included for histamine. Diagnostic meta-analysis typically requires more studies. Your discussion correctly notes this limitation but still reports high AUC values (~0.92).
This may represent small-study bias or overfitting.
Response: we have changed the conclusion of the study, to highlight the findings about histamine: Preliminary results suggest that histamine might have optimal diagnostic performance, but estimates are severely limited by small sample sizes.- in the Abstract.
And we have made other changes throughout the text. In Lines 380-381 we have added: “Thus, current evidence is insufficient for clinical recommendations regarding histamine.”
Lines 498-500: Preliminary results suggest that histamine might have optimal diagnostic performance, but estimates are severely limited by small sample sizes and there is currently no strong superiority over tryptase.
5.Language corrections
“Most performed studies sustain dynamic thresholds over fixed values.”, “The biological functions of tryptase are not well understood…” “arrythmia” → arrhythmia, “deffered” → deferred, “anaesthetic changes are not specific and require a thorough differential diagnosis” (sentence repetition)
Response: thank you for the suggested changes, we have modified these and we are taking into consideration language editting at MDPI.
6. Discussion Quality- It is slightly long (≈9 pages). Consider reduce by 15–20% by removing repetition on: mast cell biology and mediator mechanisms.
Focus more on: clinical diagnostic pathways and guideline implications
We have reduced the length of the Discussion section, now it is aproximately 5-6 pages, with a reduction of aprox 20% (we rephrased some paragraphs, which now appear in yellow marking, so that there is no double information or redundancy).
We have added the implications for the guidelines in lines 476-480, by citing guidelines over the period in which studies were published, we increased the number of references, suggesting papers where clincial diagnostic pathways can be found and are relevant for practice.
7. Figure reformatting
Graphics on Figure 2 are small, appears to be too basic, consider a better interpretation of them.
Thank you for your insightful feedback on the figures in our manuscript.
For Figure 2, we have expanded the caption to define the clinical relevance of the survival curves (which illustrate how the probability of a positive tryptase test changes with threshold levels, aiding clinicians in selecting cutoffs that minimize false positives while capturing severe cases) and the Youden index (a metric of overall diagnostic accuracy that balances sensitivity and specificity, useful for identifying thresholds that optimize test performance in resource-limited settings). We also briefly explained the derivation of the 12.68 ng/mL optimal cutoff as the value maximizing the weighted Youden index in the diagmeta model's multiple-threshold analysis, derived from the intersection of the summary ROC curve's tangent with the line of no discrimination.
The revised caption now reads: "Figure 2. Diagnostic performance of serum tryptase using fixed cutoffs in perioperative anaphylaxis (6 studies, 11 cutoffs). Top-left: Survival curves showing the probability of a positive test (elevated tryptase) versus threshold, clinically relevant for understanding how stricter cutoffs reduce false positives but may miss milder cases. Top-right: Weighted Youden index (sensitivity + specificity - 1, a measure of overall diagnostic accuracy balancing test performance) versus threshold (peak near 12.7 ng/mL). Bottom-left: Individual study ROC curves. Bottom-right: Summary ROC curve with pooled estimate (Sens 0.60, Spec 0.95 at optimal cutoff 12.68 ng/mL; AUC ≈0.72). The optimal cutoff was derived as the value maximizing the weighted Youden index in the diagmeta multiple-threshold model."
(2) For Figures 3–5 (dynamic tryptase, overall tryptase, and overall histamine SROC curves), we have updated the captions to explain the confidence ellipse as a visual representation of uncertainty in the pooled estimate, incorporating both within-study precision and inter-study heterogeneity (wider ellipses indicate greater variability across studies, such as differing patient populations or timing, which may affect generalizability). This replaces the previous repetition of sensitivity values with a focus on heterogeneity.
The revised caption for Figure 3 now reads: "Figure 3. Summary Receiver Operating Characteristic (SROC) curve for tryptase using dynamic thresholds in perioperative anaphylaxis (5 studies, 12 entries). The summary point indicates pooled sensitivity 77.2% (95% CI 74.9–79.3%) and false positive rate 11.5% (95% CI 6.5–19.5%), with AUC 0.774. The confidence ellipse represents uncertainty in the pooled estimate, incorporating inter-study heterogeneity (I² 0–43.8%); a narrower ellipse suggests relatively consistent performance across studies despite varying dynamic definitions."
Similar revisions have been applied to Figures 4 and 5.
These modifications make the figures more self-explanatory and directly address the clinical and statistical context, as requested.
8.References: for a review 35 references are at lowest acceptable limit. Consider to add a few more of them.
We have added the guidelines and we now have 41 references.
- Comments on the Quality of English Language- please see point 5 Language corrections
All the recommendations have improved the quality of the manuscript. We hope that we have adressed all issues in an appropiate manner. Should there be any other concerns, please do not hesitate to let us know. We will do all possible to respond accordingly. Thank you very much.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe manuscript has substantially improved. The major concerns raised in the previous review have been adequately addressed. I have no further suggestions, and the manuscript is suitable for publication.
Reviewer 2 Report
Comments and Suggestions for AuthorsThank you for implementing the recommended revisions and for the effort you invested in improving the manuscript; your contributions have resulted in a clearer and stronger article.

