Changes in Symptom Networks During Inpatient Cancer Rehabilitation: A Retrospective Bayesian Gaussian Graphical Model Analysis of Real-World Patient-Reported Outcomes
Round 1
Reviewer 1 Report
Comments and Suggestions for Authors
Thank you for submitting this interesting paper on a highly relevant topic. I have a few suggestions that may help further strengthen the manuscript:
- The manuscript would benefit from a more detailed explanation of why a Bayesian network approach was preferred over conventional frequentist methods and what specific advantages this offered for the present analysis.
- The interpretation of network centrality measures should be discussed more cautiously, particularly regarding their use for identifying potential intervention targets, since centrality does not necessarily imply causal relevance.
- Please provide additional justification for the predefined symptom communities used in the Bridge Expected Influence analyses.
- It would be helpful to clarify whether assumptions relevant for network estimation (e.g., missing data mechanisms, distributional properties, or multicollinearity) were formally examined.
- Given the heterogeneity of the sample, a more detailed discussion of diagnosis-specific differences and potential influences of treatment status or disease stage would strengthen the interpretation of the findings.
- The discussion could address alternative explanations for the observed network stability in greater detail, including methodological aspects and the relatively short duration of the rehabilitation interval.
- The limitations section should be expanded, particularly with respect to the single-center design, observational methodology, and the absence of post-discharge follow-up assessments.
- In several places, the Results section already includes extensive interpretation. A clearer separation between Results and Discussion would improve readability and structure.
Comments on the Quality of English Language
The manuscript would benefit from substantial language editing to improve readability, reduce repetition, and shorten overly long or dense passages.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for Authors
Results Section
(Comment 1) Table 2. Mea symptom -> Mea“n”
(Comment 2) Please revise the table formatting below.
- line 269 Table 5 -> Table 2
- line 369 Table 2 -> Table 3
- line 407 Table 3 -> Table 4
- line 369 Table 2 -> Table 3
(Comment 1) There appears to be an inconsistency in the interpretation of the nausea/vomiting–anxiety edge. Section 3.5 states that this association became stronger at discharge, whereas a positive Δ (+0.076; T0 − T1), the "Weaker at T1" label in Table 3, and the reported PIP change (0.988 → 0.148) all seem to indicate a weaker association at T1. Please clarify whether the discrepancy arises from the text, the Δ definition, or the reported values, and provide a clearer explanation of the reported PIP values and their corresponding timepoints.
(Comment 4) In the Abstract, it is reported as a negative value, “decoupling of social functioning from financial difficulties (Δ = −0.112)”, whereas in the main text and Table 3 it is reported as a positive value (+0.112).
(Comment 5) The transdiagnostic (diagnostic subgroup) analysis is inconsistently reported across multiple sections of the manuscript and is not sufficiently described. The Abstract, Introduction, and Conclusions refer to different numbers of diagnostic subgroups (2 vs. 10), and the interpretations of the findings are also contradictory. In addition, the Statistical Analysis section (2.3) does not describe the methodology used for the subgroup network analysis at all, and the Results section does not include any tables, figures, or numerical outputs supporting this analysis. As a result, the reported findings regarding structural consistency and hematological malignancies cannot be verified within the main text of the manuscript. The following revisions are requested.
(a) reconcile the number of diagnostic subgroups across all sections,
(b) add a clear description of the subgroup analysis method to Section 2.3
(c) report the corresponding results (including the per-entity correlation coefficients and the hematological-malignancy network) in the Results with appropriate tables and/or figures. Alternatively, if this analysis was not in fact performed, the related statements should be removed from the Abstract and Conclusions.
Discussion Section
(Comment 6) Although the authors rightly acknowledge in the Limitations section that the lack of a control group and the cross-sectional nature of the network models "do not support causal inference", the text in the Results and Discussion sections frequently slips into definitive, causal language. For instance, the authors state that "fatigue reduces... which in turn increases depressive symptomatology" (lines 665–666)and that "breathlessness reduces physical capacity, reduced physical abilities fuel anxious arousal, and anxious hypervigilance amplifies..." (lines 680–682). Since Gaussian Graphical Models based on cross-sectional timepoints only capture conditional associations without specifying directionality, these active causal verbs ("reduces," "fuels," "amplifies," "produced") are methodologically overstated. Please systematically soften these causal claims throughout the manuscript to relational terms, such as "associated with" or "linked to."
Author Response
Please see the attachment
Author Response File:
Author Response.pdf
Reviewer 3 Report
Comments and Suggestions for Authors
This manuscript examines changes in symptoms and functioning networks among 5,066 cancer survivors undergoing a 21-day inpatient cancer rehabilitation program. The authors used routinely collected electronic patient-reported outcome data, including the EORTC QLQ-C30 and HADS, and estimated Bayesian Gaussian Graphical Models at admission and discharge. They report substantial mean-level improvements across all 17 domains, while the overall network architecture remained largely stable. Emotional functioning and anxiety were identified as central nodes and potential targets for rehabilitation.
The topic is clinically relevant, and the use of large-scale real-world patient-reported outcome data is a clear strength. The manuscript is also methodologically ambitious and attempts to move beyond conventional pre-post comparisons by examining symptom interconnections. However, the current version has several major limitations that need to be addressed before it can be considered for publication in a high-impact oncology journal.
Most importantly, the manuscript frequently makes causal and intervention-oriented claims that are not supported by the observational, uncontrolled, retrospective design. The study can describe changes observed during rehabilitation, but it cannot determine whether rehabilitation caused those changes. The manuscript also overinterprets cross-sectional network centrality as evidence for therapeutic targets. In addition, there are serious reporting inconsistencies, including a mismatch between the stated analytical sample size and the numbers presented in Table 1. These issues substantially weaken the credibility of the findings and must be corrected.
Major comments
- The causal language is too strong for the study design. The study is a single-center, retrospective, uncontrolled pre-post analysis. Therefore, the data can show that patients improved during the rehabilitation period, but they cannot establish that rehabilitation caused these improvements. However, the manuscript repeatedly states or implies that inpatient rehabilitation “produced” improvements, “reduced” symptom coupling, or “disrupted” maladaptive symptom networks. These statements exceed what the design can support. For example, the abstract concludes that inpatient cancer rehabilitation “produces large symptomatic improvements,” and the Simple Summary states that rehabilitation led to meaningful improvements across all domains. These statements should be revised to more cautious language, such as “patients showed improvements during rehabilitation” or “improvements were observed over the course of inpatient rehabilitation.” The limitation section acknowledges the absence of a control group, but this limitation is not adequately reflected in the abstract, results, or conclusion.
- The title and terminology are potentially misleading. The title uses the phrase “Bayesian Network Analysis.” This is potentially misleading because the authors used Bayesian Gaussian Graphical Models, which estimate undirected conditional association networks. The term “Bayesian network” often refers to directed acyclic graphical models and may suggest directional or causal relations. The title and the main text should clearly state that the analysis used Bayesian Gaussian Graphical Models. A more accurate title would be: Changes in Cancer-Related Symptom Networks During Inpatient Rehabilitation: A Retrospective Bayesian Gaussian Graphical Model Analysis of Real-World Patient-Reported Outcomes.
- Network centrality is overinterpreted as evidence for therapeutic targets. The manuscript identifies emotional functioning, anxiety, and physical functioning as important central or bridge nodes and describes them as priority therapeutic targets. This interpretation is too strong. In cross-sectional Gaussian graphical models, centrality indicates statistical connectivity within a conditional association structure. It does not show that modifying a central node will improve other symptoms, nor does it establish causal pathways. This is particularly important because the nodes are not individual symptoms but aggregated scale scores from EORTC QLQ-C30 and HADS. Their centrality may reflect measurement structure, conceptual overlap, scale direction, or unmeasured confounding rather than true mechanistic influence. The authors should reassign these findings as hypothesis-generating. Statements such as “emotional functioning is a primary treatment target” should be softened to “emotional functioning may be a clinically relevant domain for further investigation in intervention studies.”
- The analysis does not capture within-person dynamic symptom change. The authors estimate one network at admission and another at discharge, then compare edge weights. This approach compares between-person conditional association structures at two time points. It does not show how symptoms influence each other within individuals over time. Therefore, claims about symptom propagation, network disruption, or dynamic symptom mechanisms are not supported. The manuscript should clearly distinguish between between-person conditional associations at admission, between-person conditional associations at discharge, and within-person symptom dynamics, which were not assessed.
- There is a serious inconsistency in the sample size and Table 1. The manuscript states that 5,066 patients had complete data and were included in the analysis. However, the age categories in Table 1 sum to 5,571, and the sex categories also sum to 5,571. This contradicts the reported analytical sample size. Moreover, the note below Table 1 refers to missing data on performance status and ECOG variables, but these variables are not shown in the table. This is a major reporting error. The authors must clarify the number of patients initially assessed, the number excluded and reasons for exclusion, the final analytical sample, whether Table 1 describes the full eligible cohort or the complete-case analytical sample, and why variables mentioned in the table note are not presented. Until this inconsistency is resolved, the reliability of the reported results remains insufficient.
- Selection bias due to complete-case analysis needs more attention. The analysis included only patients with complete data at both admission and discharge. Patients with incomplete questionnaires, early termination, or long intervals between baseline assessment and admission may differ systematically from included patients. They may have worse symptoms, poorer functional status, more advanced disease, or lower treatment adherence. The authors should provide a comparison between included and excluded patients. If possible, they should also conduct sensitivity analyses using multiple imputation or another approach to examine whether missing data could have influenced the findings.
- Differences in assessment setting may have influenced the observed changes. The baseline assessment was completed at home through a web-based portal, whereas the discharge assessment was completed on tablets at the rehabilitation center. This difference in assessment setting may introduce systematic measurement bias. Responses at discharge may be influenced by social desirability, staff presence, patient expectations, or gratitude toward the rehabilitation program. The authors should describe the timing and context of both assessments in greater detail and discuss this as a potential source of bias.
- Clinical meaningfulness is not adequately evaluated. The authors treat improvements across all domains as clinically meaningful. However, a uniform Cohen’s d threshold of 0.5 is not sufficient for interpreting clinical relevance across EORTC QLQ-C30 and HADS domains. Some statistically supported changes are very small. For example, dyspnea showed a very small standardized effect size despite a Bayes Factor greater than 10. Given the large sample size, even trivial differences can produce strong statistical evidence. The authors should interpret changes using established minimal important differences for EORTC QLQ-C30 and HADS where available. They should distinguish clearly between statistically supported changes and clinically meaningful changes.
- The Bayesian paired t-test results are overemphasized. The very large Bayes Factors mainly reflect the large sample size. Reporting many values such as “>10³⁰⁰” is not clinically informative. The manuscript would be stronger if it focused on absolute changes, 95% HDIs, standardized effect sizes, and proportions of patients achieving clinically meaningful improvement.
- The handling of multiple edge comparisons needs further justification. The authors evaluated 136 edge-level changes and classified edges as changed when the 95% HDI excluded zero. They state that this Bayesian decision criterion does not require multiplicity correction. This statement is too simplistic. Even in Bayesian analyses, the interpretation of many exploratory edge-level comparisons requires caution. The authors should either pre-specify key edges of interest, use a more conservative posterior probability threshold, report ROPE-based analyses, or explicitly frame all edge-level changes as exploratory.
- “No credible change” should not be interpreted as evidence of stability. The manuscript states that because 112 of 136 edges did not show credible change, the network architecture was largely stable. However, failure to detect change is not equivalent to evidence of equivalence or stability. To support a claim of structural stability, the authors should define an equivalence region and perform an analysis designed to assess practical equivalence. Alternatively, the statement should be softened.
- The explanation of posterior inclusion probability appears inaccurate. The introduction describes PIP as if it reflects how consistently an edge appears across bootstrap samples. However, in Bayesian model averaging, PIP is the posterior probability that a given edge is included across models. These are not the same concept. The explanation should be corrected.
- The direction of scales complicates Expected Influence interpretation. The network includes functioning scales, where higher scores indicate better status, and symptom scales, where higher scores indicate worse status. The interpretation of Expected Influence, especially its sign, depends heavily on this coding direction. For example, negative Expected Influence for emotional functioning is interpreted as a buffering effect. However, this interpretation partly arises from the opposite direction of the functioning and symptom scales. The authors should consider re-coding all variables in a common direction and repeating the centrality analyses as a sensitivity analysis. At minimum, they should discuss this limitation clearly.
- The role of financial impact needs reconsideration. Financial impact is not a symptom in the same sense as fatigue, pain, or insomnia. Including it in the symptom cluster may artificially increase associations with social functioning and quality of life. The authors should justify its classification or conduct sensitivity analyses excluding financial impact or treating it as a separate social-domain node.
- Important clinical variables are unavailable. The authors note that cancer stage, time since diagnosis, disease status, and treatment history were unavailable. These are major limitations. Symptoms and functioning in cancer survivors are strongly influenced by treatment phase, cancer stage, recurrence status, systemic therapy, surgery, and radiotherapy. Without these variables, the interpretation of symptom networks is limited. The authors should avoid strong claims about transdiagnostic symptom architecture unless these limitations are more clearly acknowledged.
- Transdiagnostic generalizability is overstated. The manuscript states that diagnostic subgroup analyses confirmed high structural consistency. However, comparing diagnostic subgroup networks with the pooled network is not fully independent because each subgroup contributes to the pooled network. Moreover, diagnostic groups differ in age, sex, treatment patterns, prognosis, and symptom burden. These factors may confound apparent diagnostic differences or similarities. The subgroup analyses should be described as exploratory. The authors should avoid concluding that the findings are broadly transdiagnostic unless more rigorous subgroup comparisons are provided.
- The Symptom Prioritization Matrix should be presented as exploratory. The matrix combining Bridge Expected Influence and standardized improvement is visually useful, but its thresholds are arbitrary. A median split of Bridge EI and Cohen’s d = 0.5 do not establish clinically validated categories. The labels “major recovery drivers” and “high-priority therapeutic targets” are too strong. This figure should be described as an exploratory visualization of domains that combine higher network connectivity and larger observed change.
- The description of subgroup analyses is inconsistent. The introduction states that the authors compared network structures across the two largest diagnostic subgroups, whereas the abstract refers to ten diagnostic subgroups. This inconsistency should be corrected. The authors should clarify whether the subgroup analyses were pre-specified or post hoc, how diagnostic subgroups were selected, and what minimum sample size was required.
Minor comments
- The manuscript contains several typographical or encoding problems, such as “admiĴed,” “paĴern,” and “beĴer.” These should be corrected throughout.
- The author line appears to contain formatting problems, including “Riedl and D.” This should be checked carefully.
- The text refers to “Table 5” when discussing symptom-level changes, but the relevant table appears to be Table 2.
- The title of Table 2 contains a typographical error: “Mea symptom and functioning changes.”
- The notation for Bayes Factors should be standardized throughout the manuscript.
- “Posterior probability of direction = 100%” should probably be reported as “>99.9%,” unless exact 100% is justified.
- The term “clinically meaningful” should be used only when supported by accepted minimal important differences or clearly defined clinical thresholds.
- The conceptual overlap between HADS anxiety/depression and EORTC emotional functioning should be discussed.
- The authors should report the proportion of patients exceeding clinical HADS cutoffs at admission and discharge.
- The interval between T0 assessment and admission should be summarized using median and interquartile range.
- The timing of T1 assessment should be specified more precisely.
- Individual variation in rehabilitation dose and treatment modality should be described if available.
- Centrality estimates should be accompanied by uncertainty intervals or stability indices.
- The clinical importance of small edge weights should be discussed cautiously.
- The supplementary table of all 136 edges is useful, but the main manuscript should include a more focused summary of the most important edges.
- The statement on AI-assisted code review should specify the extent of AI use and confirm that the authors independently verified all analyses.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Round 2
Reviewer 3 Report
Comments and Suggestions for Authors
The authors satisfactorily addressed all concerns this reviewer raised. No further comments.

