Evaluation of an Established Semi-Quantitative Chest CT Scoring System for Assessing the Severity of COVID-19 Pneumonia: What Is Its Diagnostic Value Regarding Patient Outcomes?
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThis article is a carefully conducted retrospective study demonstrating the diagnostic value of the Pan semi-quantitative CT scoring system for predicting the need for mechanical ventilation and ICU length of stay, but it has only moderate predictive power for in-hospital mortality.
However, several aspects of the design and interpretation of the results limit the generalizability of the findings and require more rigorous formulation and clarification in the text. The comments are presented point by point below.
- The introduction offers a clear and comprehensive overview of the epidemiological background of COVID‑19 and the characteristic chest CT patterns associated with SARS‑CoV‑2 pneumonia. However, the specific scientific contribution of the present study is not sufficiently highlighted. Given that the same group has already reported on the practicability and reproducibility of the Pan CT scoring system, including AI‑based assessment ([9] Neumann, E.; Movlilishvili, A.; Scherfeld, S.T. Visual and AI-based assessment of COVID-19 pneumonia: practicability and reproducibility of an established semi-quantitative chest CT scoring system. Diagnostics 2025, 15, 1987.). It would be important in this manuscript to delineate the incremental novelty more explicitly.
- The inclusion criteria (CT within ±3 days of the positive PCR result and age ≥18 years) are clearly defined. However, the exclusion of 38 patients due to image artefacts or prior extensive lung surgery and a further 218 patients because of incomplete clinical or laboratory data means that only 351 of the original 607 patients were ultimately analysed, implying substantial selection of the study population. It would therefore be important to provide a more detailed description of how the exclusion of patients with missing data may have influenced the reported mortality rate (21.4%) and the distribution of clinical outcomes.
- The authors provide a detailed description of the CT acquisition protocols, indicating that some examinations were performed as non contrast scans in the context of suspected COVID 19 pneumonia, whereas others were contrast enhanced for oncologic staging or suspected pulmonary embolism.However, although previous work by the same group has suggested that patients undergoing non‑contrast CT for symptomatic COVID‑19 infection tend to have higher Pan scores, the present study does not explicitly model the potential impact of protocol heterogeneity (including slice thickness between 0.8 and 5 mm and the use of contrast medium) on the accuracy of the Pan score or on its association with clinical outcomes. It would therefore be important either to adjust for these protocol differences in the multivariable analyses or, at a minimum, to provide a clear justification and discussion of why they are unlikely to have materially influenced the main findings.
- The authors provide a detailed description of the available laboratory parameters and transparently acknowledge the substantial amount of missing data, particularly for D dimer and pulmonary function testing. Given that the observed correlations with laboratory markers are weak and that pulmonary function was assessed in only 16 patients, the decision not to over interpret these results is appropriate. Nonetheless, the Discussion could more explicitly and succinctly state that the principal conclusions of the study relate to the association between the Pan score and the need for intensive respiratory support and prolonged ICU care, rather than to a comprehensive characterisation of the inflammatory or laboratory profile of COVID 19 pneumonia.
- The authors state that the Pan score is a simple and accessible tool that can be applied both by readers with different levels of radiological experience and by a fully automated AI‑based software solution. However, the practical added value of AI‑derived scoring in resource‑limited settings could be elaborated more concretely. In particular, the manuscript would benefit from clarifying how rapidly the software generates results and whether these outputs are suitable for real‑time triage in the emergency department, outlining how the tool is intended to be integrated into existing workflows, and providing a more detailed discussion of key limitations.
- To strengthen the clinical context of the Discussion, the authors could briefly address the long term airway sequelae observed after severe COVID 19, for example by referring to the series “Treatment of post resuscitation cicatricial tracheal stenosis after severe COVID 19 associated pneumonia”. Respiratory Medicine Case Reports, 2022, 40:101768. DOI 10.1016/j.rmcr.2022.101768, and emphasise that high Pan scores and prolonged mechanical ventilation may be linked not only to adverse acute outcomes but also to substantial late tracheal complications. In a similar vein, the broader potential of CT derived structural indices might be illustrated by citing studies that assess diaphragm thickness on CT with clinical and functional correlations, such as the recent report by Topolnitskiy E.B. et al. “Computed Tomography Assessment of Diaphragm Thickness in Myasthenia Gravis With Clinical and Functional Correlations” Cureus 2026, 18, e110567; doi:10.7759/cureus.110567. These examples would help to convey that semi quantitative parenchymal measures like the Pan score form part of a wider framework of CT based tools for evaluating airway injury and respiratory muscle involvement in complex thoracic pathology.
Author Response
Comment 1: The introduction offers a clear and comprehensive overview of the epidemiological background of COVID‑19 and the characteristic chest CT patterns associated with SARS‑CoV‑2 pneumonia. However, the specific scientific contribution of the present study is not sufficiently highlighted. Given that the same group has already reported on the practicability and reproducibility of the Pan CT scoring system, including AI‑based assessment ([9] Neumann, E.; Movlilishvili, A.; Scherfeld, S.T. Visual and AI-based assessment of COVID-19 pneumonia: practicability and reproducibility of an established semi-quantitative chest CT scoring system. Diagnostics 2025, 15, 1987.). It would be important in this manuscript to delineate the incremental novelty more explicitly.
Response 1: Thank you for pointing this out. We have added a corresponding paragraph to the introduction.
Comment 2: The inclusion criteria (CT within ±3 days of the positive PCR result and age ≥18 years) are clearly defined. However, the exclusion of 38 patients due to image artefacts or prior extensive lung surgery and a further 218 patients because of incomplete clinical or laboratory data means that only 351 of the original 607 patients were ultimately analysed, implying substantial selection of the study population. It would therefore be important to provide a more detailed description of how the exclusion of patients with missing data may have influenced the reported mortality rate (21.4%) and the distribution of clinical outcomes.
Response 2: We thank your for pressing on this point, as it led us to identify an error in how the exclusions had been described. The manuscript stated that 218 patients had been excluded because of incomplete clinical or laboratory data. This was not correct. On re-examination, these patients were excluded because the available chest CT examination did not fall within the predefined window of ±3 days relative to the positive PCR result, and they therefore did not meet inclusion criterion (b). This is an eligibility violation rather than a data availability problem, and the Methods section has been corrected accordingly. We apologise for the imprecision in the original submission. This distinction matters for the concern the reviewers raise. Reviewer 2 rightly notes that patients with incomplete records are often systematically different from those with complete records, for example because of shorter admissions or different care pathways. Eligibility with respect to the CT–PCR interval, by contrast, is determined solely by the timing of the examination relative to the diagnostic test. It was assessed against the full record for every case and without reference to clinical outcome. We therefore do not expect this exclusion to have introduced outcome-related selection bias, and we now state this explicitly in the Methods. The remaining 38 exclusions were due to motion or breathing artifacts, prior extensive lung surgery, incomplete CT data sets, or other technical difficulties preventing AI-based evaluation, as originally reported. We have considered adding the supplementary table comparing included and excluded patients that both reviewers requested. Since the exclusion was based on an eligibility criterion assessed without reference to clinical outcome, rather than on data availability, we believe such a comparison would not be informative in the way originally intended: the excluded patients did not differ in their clinical documentation but in the timing of their CT examination relative to the diagnostic test. We have instead addressed the underlying concern by describing the exclusion explicitly in the Methods, as Reviewer 2 offered as an alternative. We would of course be happy to provide the comparison table should the reviewers still consider it useful.
Comment 3: The authors provide a detailed description of the CT acquisition protocols, indicating that some examinations were performed as non contrast scans in the context of suspected COVID 19 pneumonia, whereas others were contrast enhanced for oncologic staging or suspected pulmonary embolism.However, although previous work by the same group has suggested that patients undergoing non‑contrast CT for symptomatic COVID‑19 infection tend to have higher Pan scores, the present study does not explicitly model the potential impact of protocol heterogeneity (including slice thickness between 0.8 and 5 mm and the use of contrast medium) on the accuracy of the Pan score or on its association with clinical outcomes. It would therefore be important either to adjust for these protocol differences in the multivariable analyses or, at a minimum, to provide a clear justification and discussion of why they are unlikely to have materially influenced the main findings.
Response 3: We thank youfor this point. You asked us either to adjust for protocol differences or to justify why they are unlikely to have influenced the findings; Reviewer 2 asked us to report the number of patients in each category and, if possible, to run a sensitivity analysis. We have done all of these. The distribution is now reported in the Results. Contrast-enhanced examinations were performed in 296 patients (84.3%) and non-contrast examinations in 55 (15.7%). Reconstructed slice thickness was 3 mm in 242 patients (68.9%), 5 mm in 59 (16.8%), 1 mm in 35 (10.0%), 2 mm in 10 (2.8%), and 0.8 mm in 5 (1.4%). We confirm that Pan scores differ between protocols. Patients scanned without contrast had higher scores than those scanned with contrast (median 10.8 vs. 8.1; p = 0.019), which is consistent with our earlier observation and expected, since these examinations were performed in symptomatic patients with suspected COVID-19 pneumonia. Scores also differed across slice thickness groups (p = 0.012). Crucially, however, this does not carry over to the association between the Pan score and clinical outcomes. We added both contrast medium administration and slice thickness as covariates to the primary multivariable models. For mechanical ventilation, the odds ratio per one-point increase in the Pan score was 1.29 (95% CI: 1.22–1.38) in the primary model and remained 1.29 (1.22–1.38) after adjustment for contrast medium and after additional adjustment for slice thickness. For prolonged ICU stay, the corresponding estimates were 1.27, 1.26, and 1.27. The estimates therefore move by at most 0.01. We also performed the restricted analysis Reviewer 2 suggested. Because the non-contrast group comprises only 55 patients, we report it as a robustness check rather than as a primary analysis, and only in unadjusted form. The estimates are consistent with the full cohort (mechanical ventilation: OR = 1.34, 95% CI: 1.17–1.59, 27 events; ICU stay ≥ 7 days: OR = 1.29, 95% CI: 1.14–1.52, 16 events). All of these analyses are presented in a new supplementary table, and the Discussion has been revised to state that protocol heterogeneity appears to influence the absolute level of the score rather than its relationship with clinical outcomes. We note that the direction of the reviewers' expectation was inverted in our cohort: non-contrast examinations performed for suspected COVID-19 constituted the minority rather than the majority of scans. This is why we prioritised adjustment over exclusion, as restricting the analysis to this subgroup alone would have substantially reduced statistical power.
Comment 4: The authors provide a detailed description of the available laboratory parameters and transparently acknowledge the substantial amount of missing data, particularly for D dimer and pulmonary function testing. Given that the observed correlations with laboratory markers are weak and that pulmonary function was assessed in only 16 patients, the decision not to over interpret these results is appropriate. Nonetheless, the Discussion could more explicitly and succinctly state that the principal conclusions of the study relate to the association between the Pan score and the need for intensive respiratory support and prolonged ICU care, rather than to a comprehensive characterisation of the inflammatory or laboratory profile of COVID 19 pneumonia.
Response 4: Important advice, thank you very much. We have adjusted the discussion and hope that the focus has become clearer now.
Comment 5: The authors state that the Pan score is a simple and accessible tool that can be applied both by readers with different levels of radiological experience and by a fully automated AI‑based software solution. However, the practical added value of AI‑derived scoring in resource‑limited settings could be elaborated more concretely. In particular, the manuscript would benefit from clarifying how rapidly the software generates results and whether these outputs are suitable for real‑time triage in the emergency department, outlining how the tool is intended to be integrated into existing workflows, and providing a more detailed discussion of key limitations.
Response 5: Thank you for pointing this out. We have added a corresponding paragraph to the discussion. As we did not measure the time it takes the AI to evaulate the images, this will have to be the subject of further studies.
Comment 6: To strengthen the clinical context of the Discussion, the authors could briefly address the long term airway sequelae observed after severe COVID 19, for example by referring to the series “Treatment of post resuscitation cicatricial tracheal stenosis after severe COVID 19 associated pneumonia”. Respiratory Medicine Case Reports, 2022, 40:101768. DOI 10.1016/j.rmcr.2022.101768, and emphasise that high Pan scores and prolonged mechanical ventilation may be linked not only to adverse acute outcomes but also to substantial late tracheal complications. In a similar vein, the broader potential of CT derived structural indices might be illustrated by citing studies that assess diaphragm thickness on CT with clinical and functional correlations, such as the recent report by Topolnitskiy E.B. et al. “Computed Tomography Assessment of Diaphragm Thickness in Myasthenia Gravis With Clinical and Functional Correlations” Cureus 2026, 18, e110567; doi:10.7759/cureus.110567. These examples would help to convey that semi quantitative parenchymal measures like the Pan score form part of a wider framework of CT based tools for evaluating airway injury and respiratory muscle involvement in complex thoracic pathology.
Response 6: Thank you for this point. We have added a corresponding paragraph to the discussion. Since the long-term consequences would provide sufficient grounds for a separate study, and the focus of our current study is clearly on the short-term clinical course, we decided not to discuss the sources you suggested in greater detail. This also helped us keep the manuscript from becoming overly long. Nevertheless, we greatly appreciate your suggestion.
Additional information:
During the revision we noticed that two coefficients had been transcribed incorrectly from Figure 9 into the corresponding Results paragraph. The correlation between the Pan score and LDH was reported as ρ = 0.41 but is ρ = 0.51, and the correlation with arterial pO2 was reported as ρ = −0.13 but is ρ = −0.19. The figure itself was correct in both instances. Both values have been corrected in the revised manuscript. In addition, the correlation with the duration of mechanical ventilation changed from ρ = 0.43 to ρ = 0.52 after we excluded patients who were documented as ventilated but for whom no ventilation duration had been recorded; these patients had previously been treated as having 0 hours of ventilation.
We re-examined the ventilation data and identified an inconsistency in the way ventilation status had been coded. Five patients who received non-invasive ventilation had been recorded as not ventilated. The dataset has been corrected accordingly, and the number of patients requiring mechanical ventilation is now 137 of 351 (39.0%) rather than 132 (37.6%).
We have taken this opportunity to define the endpoint more precisely. Mechanical ventilation is now defined throughout as any documented requirement for ventilatory support, comprising both invasive and non-invasive ventilation, and this is stated in the Methods. The corresponding descriptions in the Discussion have been amended, as one passage previously referred to invasive ventilation specifically.
All affected figures and estimates have been recalculated. The odds ratio for mechanical ventilation per one-point increase in the Pan score changed from 1.31 to 1.28 in the univariate model; the direction, magnitude, and significance of the association are unaffected. We further note that ventilation duration was documented for only 66 of the 137 ventilated patients; analyses of ventilation duration are based on this subset, and patients with documented ventilation but no recorded duration are now treated as missing rather than as zero hours. This is stated in the Methods and Limitations.
While addressing the reviewer's request to specify the ordinal regression model, we identified an error in how ventilation duration had been categorised. The 71 patients who were documented as ventilated but for whom no duration had been recorded had been assigned to the shortest duration category (< 24 hours) rather than treated as missing. The model has been refitted with these patients excluded, and the categories are now no ventilation (n = 214), < 24 hours (n = 13), 24 to < 48 hours (n = 6), 48 to < 72 hours (n = 5), 72 hours to < 7 days (n = 13), and ≥ 7 days (n = 29), leaving 280 patients. The coefficient changed from β = 0.276 to β = 0.295 (OR = 1.34 per one-point increase, 95% CI: 1.26–1.44; p < 0.001). We confirm that the model is a proportional-odds cumulative logit model, and the proportional-odds assumption was formally tested (likelihood ratio test against an unconstrained multinomial model: LR = 4.49, df = 4, p = 0.344).
The correction of the study period is described under Comment 7 above.
Reviewer 2 Report
Comments and Suggestions for AuthorsThis is a retrospective single-center study (n=351) evaluating whether the Pan semi-quantitative CT severity score — scored independently by a radiology specialist, a resident, a medical student, and an AI tool — correlates with laboratory markers and predicts mortality, mechanical ventilation, and ICU stay in COVID-19 pneumonia. The authors found modest performance for mortality (AUC 0.639) but good performance for ventilation need (AUC 0.804) and prolonged ICU stay (AUC 0.803).
Overall, this is a solid, clinically relevant piece of work with a reasonably large cohort and a commendably thorough statistical toolkit (ROC analysis, multivariable logistic/ordinal regression, spline-adjusted predicted probabilities, correlation matrices). It reads as a natural companion to the authors' prior reproducibility paper (ref. 9), extending that work toward clinical outcome prediction. That said, I have a number of concerns that I think need to be addressed before this is ready for publication — some are fairly substantive.
Major comments
1. Selection/attrition from the original cohort. You started with 607 patients and excluded 256 (38 for technical/AI-evaluability reasons, 218 for incomplete clinical/lab data) to arrive at 351 — a 42% attrition rate. That's a lot of patients to lose, and the 218 excluded for missing lab/clinical data are the group I'd worry about most: patients with incomplete records are often systematically different (e.g., shorter admissions, different care pathways, possibly less severe or, conversely, transferred/deceased quickly). Please add a supplementary table comparing baseline characteristics (age, sex, outcome if known, Pan score if available) of included vs. excluded patients, or at minimum discuss this attrition explicitly as a source of potential selection bias rather than only mentioning the retrospective design generally.
2. Mixed use of Pearson and Spearman correlation. Table 3 states Pearson's r was used throughout, while Figure 9 explicitly uses Spearman's ρ for what appears to be a very similar set of associations. Given that several of your key lab variables (CRP, D-dimer) are presented as median/IQR — a clear signal of skewed distributions — Pearson's r is not the ideal choice for Table 3. I'd recommend using Spearman (or another rank-based measure) consistently throughout, or providing a clear methodological justification for why two different correlation coefficients were used for what look like overlapping analyses. As it stands this reads as if the "right" test was chosen after the fact for whichever table.
3. Correlating a binary variable (biological sex) via Pearson's r. This is mathematically equivalent to point-biserial correlation, but it should be labeled as such, or better, sex differences should be assessed via a more standard categorical approach and effect size (e.g., odds ratio from logistic regression, which you're already doing elsewhere).
4. Multiple comparisons. You've run a large number of correlations and regression models, several of which sit right at the edge of significance (leukocytes p=0.045, sex p=0.047, several p-values in the 0.01–0.05 band in Figure 9/Table 3). With this many tests, I'd like to see either a correction (FDR/Bonferroni) applied to the exploratory correlation matrix, or at least an explicit statement that these are hypothesis-generating and unadjusted, so readers can calibrate how much weight to put on the borderline findings.
5. Missing data in multivariable models isn't quantified. D-dimer was available for only 155/351 patients (44%) and CRP for 297/351 (85%). When these are included as covariates in the "adjusted" logistic models, the effective analytic sample must shrink substantially via listwise deletion — and D-dimer especially is unlikely to be missing at random (it's often ordered selectively when PE is suspected, i.e., in sicker patients). Please report the actual n used in each multivariable model, and discuss how this non-random missingness might bias the adjusted estimates. If feasible, a sensitivity analysis with multiple imputation would strengthen this considerably.
6. Effect size language runs ahead of the numbers in a few places. The abstract and discussion describe r=0.54 (ventilation) as a "strong and consistent association." By conventional correlation-strength cutoffs (e.g., Cohen), 0.54 is a moderate association, not strong. Similarly, an AUC of 0.639 for mortality is closer to "poor-to-modest" than "modest" alone. I'd recommend recalibrating this language throughout (abstract, results, discussion) so the interpretive claims match the numbers a reader will see in the tables — a Q1/Q2 reviewer or reader is going to notice this mismatch.
7. Pandemic era / variant / treatment protocol not accounted for. Your inclusion window (March 2020–December 2021) spans multiple variants (wild-type, Alpha, likely early Delta) and a period of rapidly evolving standard-of-care (dexamethasone, remdesivir, evolving ICU protocols) — not to mention the beginning of vaccine rollout in your country during 2021. All of these plausibly affect both CT severity distributions and mortality independent of Pan score. At minimum this needs to be named explicitly as a limitation; ideally, calendar period (or pandemic wave) and vaccination status (if retrievable) should be considered as covariates or at least described descriptively (e.g., how many patients fell into early vs. later phases of the study period).
8. Inter-rater reliability isn't reported in this paper. You average four raters' scores into a single "mean Pan score" per patient, which is a reasonable approach, but the reliability of that average depends on how well the four raters agree — and that agreement is reported only in your companion paper (ref. 9), not here. Given that this paper's central exposure variable is that averaged score, I think at least a summary ICC (with reference to ref. 9) belongs in the Methods or Results here, rather than requiring the reader to pull the other paper.
9. Contrast vs. non-contrast CT protocols pooled together. You acknowledge this heterogeneity as a limitation (patients scanned for oncologic/PE indications got contrast, COVID-focused scans didn't), and cite your own prior finding that non-contrast COVID scans have higher Pan scores. Given you already have this information, please report how many patients in this cohort fall into each category, and ideally run a sensitivity analysis excluding the non-COVID-indication scans to show the main findings hold.
Minor / technical comments
- Abstract reports 128 females as 36.5%, Methods reports the same n as 36.4% — please reconcile the rounding.
- The AUC value of 0.639 for mortality is introduced somewhat ambiguously — it's unclear from the text and Figure 3 legend whether this refers to the mean (4-rater) score specifically or the best-performing individual rater ("the highest AUC reached 0.639"). Please clarify in-text and consider a supplementary table with each rater's individual AUC for mortality/ventilation/ICU stay, since you've clearly already computed these to produce Figure 3.
- Figure 8 caption has a duplicated phrase ("additive. additive.") — likely a copy-paste artifact.
- Methods, page 3: "organised in a dedicated list within in the radiological information system" — duplicated "within in."
- Please double check that the ordinal regression model (β = 0.276) is specified precisely — is this a proportional-odds cumulative logit model? If so, was the proportional-odds assumption tested? A one-line note on model specification would help readers evaluate it.
- Consider stating explicitly whether this study follows the STROBE reporting guideline for observational studies.
- The introduction leans fairly heavily on background/epidemiology before getting to the specific gap being addressed; a slightly tighter transition into your hypothesis would help.
Author Response
Comment 1: Selection/attrition from the original cohort. You started with 607 patients and excluded 256 (38 for technical/AI-evaluability reasons, 218 for incomplete clinical/lab data) to arrive at 351 — a 42% attrition rate. That's a lot of patients to lose, and the 218 excluded for missing lab/clinical data are the group I'd worry about most: patients with incomplete records are often systematically different (e.g., shorter admissions, different care pathways, possibly less severe or, conversely, transferred/deceased quickly). Please add a supplementary table comparing baseline characteristics (age, sex, outcome if known, Pan score if available) of included vs. excluded patients, or at minimum discuss this attrition explicitly as a source of potential selection bias rather than only mentioning the retrospective design generally.
Response 1: Thank you for pressing on this point, as it led us to identify an error in how the exclusions had been described.
The manuscript stated that 218 patients had been excluded because of incomplete clinical or laboratory data. This was not correct. On re-examination, these patients were excluded because the available chest CT examination did not fall within the predefined window of ±3 days relative to the positive PCR result, and they therefore did not meet inclusion criterion (b). This is an eligibility violation rather than a data availability problem, and the Methods section has been corrected accordingly. We apologise for the imprecision in the original submission.
This distinction matters for the concern the reviewers raise. Reviewer 2 rightly notes that patients with incomplete records are often systematically different from those with complete records, for example because of shorter admissions or different care pathways. Eligibility with respect to the CT–PCR interval, by contrast, is determined solely by the timing of the examination relative to the diagnostic test. It was assessed against the full record for every case and without reference to clinical outcome. We therefore do not expect this exclusion to have introduced outcome-related selection bias, and we now state this explicitly in the Methods.
The remaining 38 exclusions were due to motion or breathing artifacts, prior extensive lung surgery, incomplete CT data sets, or other technical difficulties preventing AI-based evaluation, as originally reported.
We have considered adding the supplementary table comparing included and excluded patients that both reviewers requested. Since the exclusion was based on an eligibility criterion assessed without reference to clinical outcome, rather than on data availability, we believe such a comparison would not be informative in the way originally intended: the excluded patients did not differ in their clinical documentation but in the timing of their CT examination relative to the diagnostic test. We have instead addressed the underlying concern by describing the exclusion explicitly in the Methods, as Reviewer 2 offered as an alternative. We would of course be happy to provide the comparison table should the reviewers still consider it useful.
Comment 2: Mixed use of Pearson and Spearman correlation. Table 3 states Pearson's r was used throughout, while Figure 9 explicitly uses Spearman's ρ for what appears to be a very similar set of associations. Given that several of your key lab variables (CRP, D-dimer) are presented as median/IQR — a clear signal of skewed distributions — Pearson's r is not the ideal choice for Table 3. I'd recommend using Spearman (or another rank-based measure) consistently throughout, or providing a clear methodological justification for why two different correlation coefficients were used for what look like overlapping analyses. As it stands this reads as if the "right" test was chosen after the fact for whichever table.
Response 2: Thank you for this observation, which is entirely justified. We agree that Pearson's correlation coefficient was not appropriate given the skewed distributions of several laboratory parameters. All correlation analyses have therefore been recalculated using Spearman's rank correlation coefficient, which is now applied consistently throughout the manuscript, including Table 3 and Figure 9. The Methods section has been amended accordingly and states the rationale explicitly.
This recalculation changed several values. Most importantly, the association between the Pan score and mechanical ventilation is now ρ = 0.49 (previously r = 0.54), and the association with ICU length of stay is ρ = 0.43 (previously r = 0.39). Two previously reported associations — leukocyte count with mortality and D-dimer with mechanical ventilation — are no longer statistically significant. All affected passages in the Abstract, Results, and Discussion have been revised, and the interpretive language has been adjusted accordingly.
Comment 3: Correlating a binary variable (biological sex) via Pearson's r. This is mathematically equivalent to point-biserial correlation, but it should be labeled as such, or better, sex differences should be assessed via a more standard categorical approach and effect size (e.g., odds ratio from logistic regression, which you're already doing elsewhere).
Response 3: We agree. Biological sex has been retained as a row in Table 3 for completeness, but the coefficient is now explicitly identified as a rank-biserial association measure, and the coding (male = 1, female = 0) is stated in the table legend. In addition, sex differences are now reported as odds ratios with 95% confidence intervals derived from logistic regression, consistent with the approach used elsewhere in the manuscript.
Comment 4: Multiple comparisons. You've run a large number of correlations and regression models, several of which sit right at the edge of significance (leukocytes p=0.045, sex p=0.047, several p-values in the 0.01–0.05 band in Figure 9/Table 3). With this many tests, I'd like to see either a correction (FDR/Bonferroni) applied to the exploratory correlation matrix, or at least an explicit statement that these are hypothesis-generating and unadjusted, so readers can calibrate how much weight to put on the borderline findings.
Response 4: Thank you for raising this point and have addressed it by applying a correction rather than by adding a caveat. All p-values from the correlation analyses have been adjusted using the Benjamini-Hochberg false discovery rate procedure, and adjusted q-values are now reported alongside the unadjusted p-values in Table 3 and in Figure 9. The correction was applied separately within each family of tests, as Table 3 and Figure 9 address distinct questions; this is stated in the Methods.
Of the ten associations in Table 3 that reached nominal significance, eight remained significant after correction (q < 0.05). Importantly, all three associations involving the Pan score — with mortality, mechanical ventilation, and ICU length of stay — withstood adjustment. The two associations that did not survive correction were biological sex with mortality (q = 0.099) and CRP with mechanical ventilation (q = 0.079); both are now described as non-significant in the revised text. We believe this provides a more informative basis for readers than an unadjusted analysis accompanied by a caveat.
The same correction was applied to the correlations underlying Figure 9. Of the twelve associations reaching nominal significance there, ten remained significant after adjustment; the two that did not (haemoglobin and prothrombin time) are no longer described as significant. The figure now displays significance levels based on adjusted q-values, as stated in the legend.
Comment 5: Missing data in multivariable models isn't quantified. D-dimer was available for only 155/351 patients (44%) and CRP for 297/351 (85%). When these are included as covariates in the "adjusted" logistic models, the effective analytic sample must shrink substantially via listwise deletion — and D-dimer especially is unlikely to be missing at random (it's often ordered selectively when PE is suspected, i.e., in sicker patients). Please report the actual n used in each multivariable model, and discuss how this non-random missingness might bias the adjusted estimates. If feasible, a sensitivity analysis with multiple imputation would strengthen this considerably.
Response 5: We agree that this information was missing and have added it. Table 4 now reports, for every model, the analytic sample after listwise deletion and the number of outcome events entering the model. The Methods section states explicitly that the analytic sample varies between models for this reason.
You are correct that the reduction is substantial. For mechanical ventilation, the sample decreases from 351 (unadjusted) to 297 when CRP is included and to 138 when D-dimer is added; for prolonged ICU stay, the corresponding figures are 346, 292, and 137. We also agree that D-dimer is unlikely to be missing at random, as it is typically ordered when pulmonary embolism is suspected, and we now state this explicitly in the Methods and Limitations.
To address the underlying concern, we have restructured the presentation of the multivariable models. The model adjusted for age, sex, and haematological parameters is now designated as the primary multivariable model, as it retains essentially the full analytic sample (n = 349 and n = 344, respectively) while providing an adequate number of events per variable. The models additionally including CRP and D-dimer are now explicitly labelled as sensitivity analyses.
We would draw your attention to the stability of the effect estimate across all levels of adjustment. For mechanical ventilation, the odds ratio per one-point increase in the Pan score ranges from 1.24 to 1.29 across the five models; for prolonged ICU stay, from 1.26 to 1.28. Despite the analytic sample falling by more than 60% in the most heavily adjusted models, the estimate does not move materially. We regard this as evidence that the incomplete availability of these laboratory parameters has not substantially biased the reported association.
We have also added an explicit caveat regarding the most heavily adjusted model for prolonged ICU stay, which is based on only 15 outcome events and is therefore reported as exploratory. We considered multiple imputation as suggested; given the stability of the estimates across adjustment levels and the fact that the primary model is based on virtually the complete cohort, we felt that the additional modelling assumptions would not materially strengthen the conclusions. We would of course be happy to add such an analysis should the reviewer consider it necessary.
Comment 6: Effect size language runs ahead of the numbers in a few places. The abstract and discussion describe r=0.54 (ventilation) as a "strong and consistent association." By conventional correlation-strength cutoffs (e.g., Cohen), 0.54 is a moderate association, not strong. Similarly, an AUC of 0.639 for mortality is closer to "poor-to-modest" than "modest" alone. I'd recommend recalibrating this language throughout (abstract, results, discussion) so the interpretive claims match the numbers a reader will see in the tables — a Q1/Q2 reviewer or reader is going to notice this mismatch.
Response 6: Thank you for pointing this out. We have adjusted the descriptions accordingly.
Comment 7: Pandemic era / variant / treatment protocol not accounted for. Your inclusion window (March 2020–December 2021) spans multiple variants (wild-type, Alpha, likely early Delta) and a period of rapidly evolving standard-of-care (dexamethasone, remdesivir, evolving ICU protocols) — not to mention the beginning of vaccine rollout in your country during 2021. All of these plausibly affect both CT severity distributions and mortality independent of Pan score. At minimum this needs to be named explicitly as a limitation; ideally, calendar period (or pandemic wave) and vaccination status (if retrievable) should be considered as covariates or at least described descriptively (e.g., how many patients fell into early vs. later phases of the study period).
Response 7: We agree that this required explicit treatment and have addressed it in three ways. First, the distribution of patients across the study period is now reported in the Results: 38 patients (10.8%) were included in 2020, 109 (31.1%) in 2021, and 204 (58.1%) in 2022.
Second, while revising this section we identified an error in the stated inclusion period. The manuscript reported inclusion between 21 March 2020 and 27 December 2021; the correct period, verified against the dataset, is 12 August 2020 to 30 December 2022. This has been corrected in the Abstract and the Methods. We apologise for this error. As a consequence, the study period is longer than originally stated and spans the wild-type, Alpha, Delta and Omicron eras rather than ending before Omicron. The Discussion has been amended accordingly and now names these phases explicitly.
Third, and most importantly, we examined whether this affects our findings. Calendar period was added as a covariate to the primary multivariable models. The odds ratio per one-point increase in the Pan score was essentially unchanged: 1.29 (95% CI: 1.22–1.38) versus 1.30 (1.23–1.39) for mechanical ventilation, and 1.27 (1.19–1.36) in both models for prolonged ICU stay. This is now stated in the Discussion. We therefore believe that, although the pandemic phase plausibly influenced disease severity and outcomes in absolute terms, as the reviewer rightly notes and as Tsakok et al. have shown for Omicron, the predictive value of the Pan score itself was stable across the periods covered.
As previously stated, vaccination status and prior infections were not routinely documented at the time of admission, so these specific variables could not be analysed retrospectively. This remains acknowledged as a limitation.
Comment 8: Inter-rater reliability isn't reported in this paper. You average four raters' scores into a single "mean Pan score" per patient, which is a reasonable approach, but the reliability of that average depends on how well the four raters agree — and that agreement is reported only in your companion paper (ref. 9), not here. Given that this paper's central exposure variable is that averaged score, I think at least a summary ICC (with reference to ref. 9) belongs in the Methods or Results here, rather than requiring the reader to pull the other paper.
Response 8: We agree that the reliability of the averaged score belongs in this manuscript rather than only in the companion paper. The intraclass correlation coefficient is now reported in the Results: ICC = 0.85 (95% CI: 0.81–0.88), indicating excellent agreement between the four readers, with reference to our previous work [9] for the full reliability analysis.
Comment 9: Contrast vs. non-contrast CT protocols pooled together. You acknowledge this heterogeneity as a limitation (patients scanned for oncologic/PE indications got contrast, COVID-focused scans didn't), and cite your own prior finding that non-contrast COVID scans have higher Pan scores. Given you already have this information, please report how many patients in this cohort fall into each category, and ideally run a sensitivity analysis excluding the non-COVID-indication scans to show the main findings hold.
Response 9: We thank both reviewers for this point. Reviewer 1 asked us either to adjust for protocol differences or to justify why they are unlikely to have influenced the findings; Reviewer 2 asked us to report the number of patients in each category and, if possible, to run a sensitivity analysis. We have done all of these.
The distribution is now reported in the Results. Contrast-enhanced examinations were performed in 296 patients (84.3%) and non-contrast examinations in 55 (15.7%). Reconstructed slice thickness was 3 mm in 242 patients (68.9%), 5 mm in 59 (16.8%), 1 mm in 35 (10.0%), 2 mm in 10 (2.8%), and 0.8 mm in 5 (1.4%).
We confirm that Pan scores differ between protocols. Patients scanned without contrast had higher scores than those scanned with contrast (median 10.8 vs. 8.1; p = 0.019), which is consistent with our earlier observation and expected, since these examinations were performed in symptomatic patients with suspected COVID-19 pneumonia. Scores also differed across slice thickness groups (p = 0.012).
Crucially, however, this does not carry over to the association between the Pan score and clinical outcomes. We added both contrast medium administration and slice thickness as covariates to the primary multivariable models. For mechanical ventilation, the odds ratio per one-point increase in the Pan score was 1.29 (95% CI: 1.22–1.38) in the primary model and remained 1.29 (1.22–1.38) after adjustment for contrast medium and after additional adjustment for slice thickness. For prolonged ICU stay, the corresponding estimates were 1.27, 1.26, and 1.27. The estimates therefore move by at most 0.01.
We also performed the restricted analysis Reviewer 2 suggested. Because the non-contrast group comprises only 55 patients, we report it as a robustness check rather than as a primary analysis, and only in unadjusted form. The estimates are consistent with the full cohort (mechanical ventilation: OR = 1.34, 95% CI: 1.17–1.59, 27 events; ICU stay ≥ 7 days: OR = 1.29, 95% CI: 1.14–1.52, 16 events).
All of these analyses are presented in a new supplementary table, and the Discussion has been revised to state that protocol heterogeneity appears to influence the absolute level of the score rather than its relationship with clinical outcomes.
We note that the direction of the reviewers' expectation was inverted in our cohort: non-contrast examinations performed for suspected COVID-19 constituted the minority rather than the majority of scans. This is why we prioritised adjustment over exclusion, as restricting the analysis to this subgroup alone would have substantially reduced statistical power.
Additional information:
During the revision we noticed that two coefficients had been transcribed incorrectly from Figure 9 into the corresponding Results paragraph. The correlation between the Pan score and LDH was reported as ρ = 0.41 but is ρ = 0.51, and the correlation with arterial pO2 was reported as ρ = −0.13 but is ρ = −0.19. The figure itself was correct in both instances. Both values have been corrected in the revised manuscript. In addition, the correlation with the duration of mechanical ventilation changed from ρ = 0.43 to ρ = 0.52 after we excluded patients who were documented as ventilated but for whom no ventilation duration had been recorded; these patients had previously been treated as having zero hours of ventilation.
We re-examined the ventilation data and identified an inconsistency in the way ventilation status had been coded. Five patients who received non-invasive ventilation had been recorded as not ventilated. The dataset has been corrected accordingly, and the number of patients requiring mechanical ventilation is now 137 of 351 (39.0%) rather than 132 (37.6%).
We have taken this opportunity to define the endpoint more precisely. Mechanical ventilation is now defined throughout as any documented requirement for ventilatory support, comprising both invasive and non-invasive ventilation, and this is stated in the Methods. The corresponding descriptions in the Discussion have been amended, as one passage previously referred to invasive ventilation specifically.
All affected figures and estimates have been recalculated. The odds ratio for mechanical ventilation per one-point increase in the Pan score changed from 1.31 to 1.28 in the univariate model; the direction, magnitude, and significance of the association are unaffected. We further note that ventilation duration was documented for only 66 of the 137 ventilated patients; analyses of ventilation duration are based on this subset, and patients with documented ventilation but no recorded duration are now treated as missing rather than as zero hours. This is stated in the Methods and Limitations.
While addressing the reviewer's request to specify the ordinal regression model, we identified an error in how ventilation duration had been categorised. The 71 patients who were documented as ventilated but for whom no duration had been recorded had been assigned to the shortest duration category (< 24 hours) rather than treated as missing. The model has been refitted with these patients excluded, and the categories are now no ventilation (n = 214), < 24 hours (n = 13), 24 to < 48 hours (n = 6), 48 to < 72 hours (n = 5), 72 hours to < 7 days (n = 13), and ≥ 7 days (n = 29), leaving 280 patients. The coefficient changed from β = 0.276 to β = 0.295 (OR = 1.34 per one-point increase, 95% CI: 1.26–1.44; p < 0.001). We confirm that the model is a proportional-odds cumulative logit model, and the proportional-odds assumption was formally tested (likelihood ratio test against an unconstrained multinomial model: LR = 4.49, df = 4, p = 0.344).
The correction of the study period is described under Comment 7 above.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe authors carefully considered the reviewers' comments and significantly improved the manuscript. The scientific novelty of the study and patient selection criteria were clarified; an initially inaccurate description of the reasons for excluding part of the cohort was corrected.
The ventilatory support data were re-verified, the coding of patients receiving NIV was corrected, and the endpoint definition was clarified to include both non-invasive and invasive ventilation. The ventilation duration analysis was also revised: patients without a documented duration are no longer counted as having a duration of zero. The authors additionally analyzed the impact of heterogeneity in CT protocols, including the use of contrast agent and slice thickness. Adjustment for these parameters and sensitivity analysis confirmed the robustness of the underlying association between the Pan score, the need for ventilatory support, and prolonged ICU stay. Overall, the changes improved the methodological transparency and reliability of the results. I recommend that authors conduct a thorough final review of the manuscript, including the text, tables, figures, figure and table captions, and statistical notations. Particular attention should be paid to ensuring complete consistency of numerical data, sample sizes, percentages, correlation coefficients, odds ratios, confidence intervals, and 𝑝-values across all sections of the manuscript.
Author Response
Thank you very much for your valuable and critical comments, which have contributed significantly to improving this paper.
Following a thorough revision of the manuscript in accordance with your comments and the points raised by the second reviewer, we have carefully proofread the text, tables, figures and captions once again. We very much hope that everything is now coherent and satisfactory.
Reviewer 2 Report
Comments and Suggestions for AuthorsI thank the authors for their careful and largely transparent point-by-point response. Substantively, the revision is a clear improvement over the original submission, and I appreciate that the authors used the review process to re-examine the dataset rather than simply rebutting the comments.
However, the manuscript is not yet ready for publication in its current form. The clean copy submitted for review still contains numerous track-changes/merging artefacts, including duplicated and contradictory sentences and, more importantly, what appear to be concatenated numerical values. As a result, several key results cannot be verified reliably from the submitted version. I have outlined both the substantive assessment and the remaining file-level problems below.
1. Assessment of the substantive responses
- Attrition/selection bias (Comment 1): I appreciate that revisiting this issue led to the identification and correction of a genuine coding error. The 218 excluded patients were excluded because they fell outside the CT–PCR timing window, rather than because of “incomplete data” as originally stated. This is a materially different and less concerning explanation, and the authors' explanation for why a baseline comparison table would not be particularly informative is reasonable. I would nevertheless suggest briefly acknowledging that CT timing may itself be related to the clinical trajectory. For example, patients who deteriorate very rapidly or are discharged very quickly may be systematically more or less likely to undergo CT within the ±3-day window. Thus, some residual selection bias cannot be completely excluded.
- Pearson → Spearman throughout (Comment 2): Adequately addressed.
- Point-biserial handling of sex (Comment 3): Adequately addressed. Reporting sex using both a rank-biserial coefficient and a logistic-regression odds ratio is a reasonable solution.
- Multiple-comparison correction (Comment 4): Adequately and transparently addressed, including the distinction between the two test families.
- Missing data / multivariable models (Comment 5): The addition of the per-model sample size and number of events, the distinction between the primary near-complete model and the CRP/D-dimer sensitivity analyses, and the discussion of the likely non-random missingness of D-dimer data all improve this section substantially.
- Effect-size language (Comment 6): Only partially addressed. As discussed below, remnants of the original “strong” wording remain alongside the revised “moderate” wording.
- Pandemic era/variant/protocol drift (Comment 7): This is the strongest part of the revision. The corrected inclusion window, calendar-year distribution, and calendar-period-adjusted sensitivity analysis address the requested concern well and support the stability of the Pan score's predictive value across different pandemic periods.
- Inter-rater reliability (Comment 8): Adequately addressed.
- Contrast/slice-thickness heterogeneity (Comment 9): Very thoroughly addressed, including the requested subgroup robustness analysis.
I also appreciate the authors' transparency in identifying several additional transcription and coding errors during the revision, including the LDH/pO₂ correlation values, ventilation-duration handling, and NIV coding. The willingness to identify and correct these issues is a positive aspect of the revision.
2. Major concern: the submitted manuscript is not a clean, internally consistent version
The main remaining problem is with the manuscript file itself. The PDF appears to have been generated before the tracked changes were fully accepted and the superseded text removed. In several places, the old and revised wording appear together, and some numerical values seem to have been concatenated at the character level. This makes it difficult, and in some cases impossible, to determine which values represent the authors' final results.
Examples include:
- Abstract: The mechanical-ventilation rates appear as both “56.0 vs. 32.1%” and “61.3 vs. 33.0%.” The paragraph also contains both the original “strong” wording and the revised “moderate” wording, together with different correlation coefficients. The inclusion dates likewise appear to have been merged: “December 27, 2021August 12, 2020, and December 30, 2022.” These need to be cleaned up so that only the final, verified values remain.
- Methods (exclusion criteria): The corrected explanation that 218 patients were excluded because they did not meet the CT–PCR timing criterion is immediately followed by the original statement that another 218 patients were excluded because of incomplete clinical and laboratory data. This directly conflicts with the revised explanation and should be corrected.
- Results — mechanical ventilation: The text currently states that mechanical ventilation was required in 1372 of 351 patients (397.06%), followed later by n = 1372 and n = 2149. These appear to be concatenation artefacts. Based on the authors' own response, the intended figures appear to be 137/351 (39.0%), with 214 patients not requiring ventilation, but the final values should be checked against the corrected dataset and R output rather than inferred from the manuscript.
- Table 2: The table still appears to contain the previous mechanical-ventilation rates (37.6% overall; 32.1% vs. 56.0%), whereas the revised text and abstract report 137 patients (39.0%) and 61.3% vs. 33.0%. Please regenerate the table from the corrected dataset rather than editing the existing table manually.
- Table 3: Several cells appear to contain concatenated old and new values, including the Pan-score/mortality correlation (“0.202” with “p < 0.01 (< 0.001)”), Pan-score/ventilation (“0.4954”), Pan-score/ICU stay (“0.4339”), age/ventilation (“0.91163”), and sex/ventilation (“0.64438”). These values need to be checked against the underlying analysis and presented only in their final form.
- Table 4: The older two-row table containing the original unadjusted ORs is still present immediately before the new expanded table with per-model sample sizes and events. The outdated table should be removed.
- Discussion: The phrase “strongly moderately correlated” is clearly a residual editing artefact. There also appears to be a discrepancy between the reported correlation for prolonged mechanical ventilation, with ρ = 0.52 appearing earlier and ρ = 0.43 later for what appears to be the same outcome. Please verify the correct value against the final analysis.
- Figure 9 caption/text: The statement “death after PCR: ρ = -0.37, q p = 0.004; n = 711” cannot be correct for a cohort of 351 patients with 75 deaths. The sample size and associated statistics should be rechecked and corrected.
- Unresolved author comment: An internal comment remains visible in the manuscript: “Commented [AH1]: Warum hier n = 138 und weiter oben n = 155?” along with another comment identifying a typo. The n = 138 versus n = 155 discrepancy for the D-dimer sensitivity model is important and should not simply be deleted with the comment. The authors need to determine which sample size is correct and explain the difference in the manuscript if appropriate.
I do not think these problems necessarily indicate flaws in the underlying analysis. They appear much more consistent with an incomplete tracked-changes cleanup or an incorrect export of the revised Word file. Nevertheless, in the current submitted version they materially affect the ability to verify the main quantitative findings. Several of the values that need to be confirmed before publication — including the mechanical-ventilation rate, several Table 3 correlations, and the Pan score/mortality correlation — are currently contradictory or unreadable.
3. Recommendation
I recommend that the authors undertake a final quality-control pass before resubmission. Specifically, they should:
- Provide a single clean manuscript with all tracked changes accepted and all superseded text removed.
- Regenerate Tables 2–4 directly from the final, corrected dataset rather than modifying the previous tables manually.
- Have a co-author independently cross-check the reported statistics in the abstract, main text, tables, and figure captions against the final R output.
- Resolve the n = 138 versus n = 155 discrepancy in the D-dimer-adjusted model and explain it clearly in the manuscript where necessary.
- Perform a final full-text check for duplicated words, merged numerical values, contradictory versions of the same result, and any remaining internal comments or tracked-change artefacts.
Once this quality-control step has been completed, I expect the manuscript to be in good shape for publication. The substantive responses to the original review comments are, overall, sound and thorough; the remaining issues are primarily about ensuring that the final manuscript accurately and consistently reflects those revisions.
Author Response
Thank you very much for your extremely valuable and critical comments, which have contributed significantly to improving this paper.
Following a thorough revision of the manuscript in accordance with your comments and the points raised by the second reviewer, we have carefully proofread the text, tables, figures and captions once again. We very much hope that everything is now coherent and satisfactory.
We regret that the changes made in the first round were not clear to you. They were displayed correctly on our end, so there must have been an error when the document was uploaded or when the PDF was generated automatically. We hope that the changes in the current version are now easier to follow.
Comment: I would nevertheless suggest briefly acknowledging that CT timing may itself be related to the clinical trajectory. For example, patients who deteriorate very rapidly or are discharged very quickly may be systematically more or less likely to undergo CT within the ±3-day window. Thus, some residual selection bias cannot be completely excluded.
Response: A corresponding paragraph has been added to the Discussion.
Comment: Resolve the n = 138 versus n = 155 discrepancy in the D-dimer-adjusted model and explain it clearly in the manuscript where necessary.
Response: We have added an explanation to the Results section clarifying the reasons for the different sample sizes for D-dimer, CRP, etc.
Round 3
Reviewer 2 Report
Comments and Suggestions for AuthorsThis is a much better manuscript than the last round. The authors clearly took the comments seriously — they didn't just patch things up, they went back and caught their own errors along the way (wrong exclusion reason, wrong study dates, Pearson used where Spearman was needed, a few miscopied correlation values, a ventilation-coding mistake). That's not nothing. Switching to Spearman throughout, adding FDR correction with q-values, reporting n and events for every model in Table 4, running the contrast-protocol and calendar-period sensitivity checks — all of this is exactly what was asked for, and it's done properly, not just cosmetically.
I do have one thing that needs fixing before this goes to print, though.
In the Results, the text says the ICU stay ≥7 days association "persisted after multivariate adjustment (adjusted OR = 1.26, 95% CI: 1.12–1.46)." Checked against Table 4, that number isn't the primary adjusted model — it's the D-dimer sensitivity model, which only has 15 events and which the authors themselves elsewhere describe as exploratory. The actual primary model (age/sex/haematology) gives OR = 1.27 (1.19–1.36). Right now the Results text quietly cites the underpowered number as if it were the headline result. Easy fix, but it needs fixing — a reader skimming the results would come away with the wrong impression of how solid that estimate is.
Smaller stuff:
- There's an inconsistency between sections on the same number: in the Results, the sentence correctly reads "...a predicted probability of >50% for an ICU stay of ≥7 days," but the equivalent sentence in the Discussion still has the old wording, "...a predicted probability of >50% or an ICU stay of ≥7 days." Looks like this was fixed in one place and missed in the other — worth reconciling.
- Table 4's unadjusted ICU-stay model uses n = 346, not the full 351 that the unadjusted ventilation model uses. Worth a one-line note on why (missing LOS data for 5 patients?), since the authors were otherwise very good this round about explaining every denominator.
Author Response
We thank the reviewer for the careful reading of the revised manuscript and for the encouraging assessment of our previous revision. The remaining points have been addressed as detailed below. Changes in the manuscript are marked using tracked changes.
Main point – Results text cited the sensitivity model rather than the primary model
In the Results, the text says the ICU stay ≥7 days association "persisted after multivariate adjustment (adjusted OR = 1.26, 95% CI: 1.12–1.46)." Checked against Table 4, that number isn't the primary adjusted model — it's the D-dimer sensitivity model, which only has 15 events […]. Right now the Results text quietly cites the underpowered number as if it were the headline result.
The reviewer is entirely correct, and we are grateful for the catch. This sentence predated the restructuring of Table 4 into primary and sensitivity models and was not updated accordingly. The passage has been rewritten so that the primary multivariable model is reported as the main result and the sensitivity analyses are clearly identified as such, including an explicit statement about the limited number of events in the D-dimer model. The revised text now reads:
Similarly, for an ICU stay ≥ 7 days, the univariate odds ratio per one-point increase in Pan score was 1.27 (95% CI: 1.19–1.36; n = 346, 56 events; p < 0.001), and the association persisted in the primary multivariable model (OR = 1.27, 95% CI: 1.19–1.36; n = 344, 55 events; p < 0.001). The estimate was likewise unchanged in sensitivity analyses including CRP (OR = 1.28, 95% CI: 1.19–1.38; n = 292) and D-dimer (OR = 1.26, 95% CI: 1.12–1.46; n = 137). The latter model was based on only 15 outcome events and should be regarded as exploratory (Table 4).
The corresponding passage for mechanical ventilation had already been structured in this way, so the two endpoints are now described consistently.
Minor point 1 – Inconsistent wording between Results and Discussion
In the Results, the sentence correctly reads "…a predicted probability of >50% for an ICU stay of ≥7 days," but the equivalent sentence in the Discussion still has the old wording, "…a predicted probability of >50% or an ICU stay of ≥7 days."
Corrected. The Discussion now reads "for an ICU stay of ≥ 7 days", matching the Results.
Minor point 2 – Denominator of the unadjusted ICU model
Table 4's unadjusted ICU-stay model uses n = 346, not the full 351 that the unadjusted ventilation model uses. Worth a one-line note on why (missing LOS data for 5 patients?)
The reviewer's inference is correct: the length of ICU stay was not documented for five patients, who are therefore excluded from all models for this endpoint. A note has been added to the legend of Table 4:
The analytic sample for the ICU stay endpoint is 346 rather than 351 because the length of ICU stay was not documented for five patients.
Additional correction identified by the authors
While implementing the above changes we noticed that a typographical error flagged in the previous round had not been carried over into the revised file: the Methods section read "organised in a dedicated list within in the radiological information system". The duplicated word has now been removed.
We thank the reviewer once again for the constructive and attentive review, which has substantially improved the manuscript.
