Virological Failure and Mortality Among People Living with HIV Transitioned to Second- or Third-Line Antiretroviral Therapy in Rwanda: A Competing-Risks Survival Analysis of National Case-Based Surveillance Data, 2019–2025
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe topic is of interest with regard to the risk of virological failure in an African setting for PWH transitioning to other than first line therapy (mostly 2nd lines).
The main limitation, as the outcome is the virological success, is the absence of HIV drug resistance pattern, before and after transition.
Overall, the paper is well written, but I would not be so confident in the conclusions: INSTI are better than PI-based regimens in the context of an observational trial where adherence behaviour is not sufficiently quantified and drug resistance background is unknown. Despite this caveat, I think it reports relevant findings in a clinical approach So, I suggest to tone down the conclusions accordingly.
The Methods section is clear and described in details.
The statistical analysis is correct. The tables and figures are clear
The English language is appropriate
Certainly, several points deserve attention before publication:
WHich PI? Which INSTIs?. Did people change also the NRTI backbones? All this need to be detailed in a table
Line 84: “poor adherence” is REPEATED
Which criteria are used to define: “good, moderate and bad adhrence rate”? (to be included in the methods section as adherence assessment).
Which were the reasons for switching to PI-based or INSTI-based (69%) regimens? Up to availability, or payment or previous failures or historical genotypes, or at clinicians’ descrition?
How frequent was the virological monitoring after transition? Were the fequency similar by type of of regimen (INSTI vs PI)?
Any HIV resistance data before transition (historical) or at transition time or at time of failure?
Major limitations: resistance data at VF (which accounted for 10% of people), and adherence data available for 16% only. So it is difficult to interprete the drivers of the virological success.
Author Response
Reviewer 1
|
Reviewer Comment (verbatim) |
Response and Revised Manuscript Text |
|
1. WHich PI? Which INSTIs? Did people change also the NRTI backbones? All this need to be detailed in a table. |
Response: A new supplementary table now lists the exact regimen (drug class plus NRTI backbone) each patient switched to, with a drug-class abbreviation key. Revised manuscript text: “Table S4. Specific regimen (drug combination, including NRTI backbone) switched to at transition, among the full analytic cohort (N=778): AZT+3TC+ATV 333 (43%); AZT+3TC+DTG 407 (52%); AZT+3TC+LOP/r 17 (2.2%); RAL+DRV/r+TDF 8 (1.0%); RAL+ETV+DRV/r 13 (1.7%).” “NRTI: 3TC = lamivudine; AZT = zidovudine; TDF = tenofovir disoproxil fumarate; INSTI: DTG = dolutegravir; RAL = raltegravir; PI: ATV = atazanavir; DRV/r = darunavir/ritonavir; LOP/r = lopinavir/ritonavir ('/r' = ritonavir-boosted); NNRTI: ETV = etravirine.” |
|
2. Line 84: "poor adherence" is REPEATED |
Response: The duplicated wording was removed during the full-text revision of the Background section. Revised manuscript text: “Advanced WHO clinical stage, low CD4 cell count, tuberculosis co-infection, poor adherence, delayed regimen transition, and high viral load at transition have been associated with poor outcomes on second-line ART [9,14,17].” |
|
3. Which criteria are used to define: "good, moderate and bad adherence rate"? (to be included in the methods section as adherence assessment). |
Response: Clarified in Methods (Exposure variables) that adherence is a routine categorical clinician judgement recorded per national ART programme M&E tools, and that CBS does not capture the underlying counselling criteria — this data limitation is now stated explicitly rather than implied. Revised manuscript text: “Adherence was recorded as a categorical clinician assessment (good, moderate, or bad) per Rwanda's national ART programme M&E tools; CBS does not capture the specific counselling criteria (e.g., pill count or missed-dose recall) underlying this categorical judgement.” |
|
4. Which were the reasons for switching to PI-based or INSTI-based (69%) regimens? Up to availability, or payment or previous failures or historical genotypes, or at clinicians' discretion? |
Response: A new supplementary table reports the recorded reason for regimen change, stratified by the regimen class switched to, among the 13.1% of patients with a reason field populated in CRF2. Revised manuscript text: “Table S7. Recorded reason for ART regimen change, by regimen class switched to (N with reason = 102/778, 13.1%). Treatment failure: INSTI-based 9 (16%), PI-based 20 (49%), Other 3 (75%). Original regimen out of stock: INSTI-based 16 (28%), PI-based 2 (4.9%). New regimen introduced: INSTI-based 21 (37%), PI-based 9 (22%).” |
|
5. How frequent was the virological monitoring after transition? Were the fequency similar by type of of regimen (INSTI vs PI)? |
Response: A new supplementary table quantifies VL testing coverage and timing after transition. Monitoring intervals are protocol-defined nationally and do not vary by regimen class, so results are reported for the overall cohort rather than stratified by regimen. Revised manuscript text: “Table S5. Viral load monitoring frequency and timing after ART line transition. Patients with ≥1 eligible VL test after the 180-day window: 661/778 (85.0%). Eligible VL tests per patient, median (IQR): 4 (2–7). Days from transition to first eligible VL test, median (IQR): 1167 (619–1473). Approximately 15% of patients had no eligible VL test recorded, and among those tested, median time to first test exceeded the 6–12 month guideline interval, consistent with real-world monitoring gaps.” |
|
6. Any HIV resistance data before transition (historical) or at transition time or at time of failure? |
Response: No genotypic resistance data are captured in Rwanda's CBS system. This is now stated explicitly as a limitation and is the basis for the confounding-by-indication caveat placed on the regimen-outcome association. Revised manuscript text: “Adherence was missing for 84%, and resistance, prior regimen history, tuberculosis, comorbidities, and facility-level clinical factors were unavailable, leaving substantial potential for residual confounding and confounding by indication.” |
|
7. Major limitations: resistance data at VF (which accounted for 10% of people), and adherence data available for 16% only. So it is difficult to interprete the drivers of the virological success. |
Response: The Discussion and Conclusion now state explicitly that, absent resistance and adherence data, the regimen-outcome association should not be read causally and should instead motivate intensified monitoring and resistance-guided review. Revised manuscript text: “Because regimen allocation was not randomized and resistance data were unavailable, the regimen association should inform intensified monitoring and resistance-guided clinical review rather than be interpreted as a causal comparison.” |
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsBrief Summary
This paper utilizes data from the Rwanda CBS system to conduct a retrospective cohort study on the VF and all-cause mortality rates among PLHIV receiving second- or third-line ART between 2019 and 2025. The strengths of this study lie in its use of large-scale national-level surveillance data, a long follow-up period, and the introduction of the Fine-Gray competing risks model to distinguish between virological failure and the risk of death, which holds certain public health value. However, upon rigorous review, the manuscript exhibits severe flaws in its epidemiological design, causal inference logic, variable definitions, and data completeness. The authors must undertake major revisions.
Comments regarding general concepts
- The paper concludes that PI-based regimens are the strongest independent risk factor for VF. In real-world clinical situations, patients typically use PI regimens because they have already experienced prior treatment failure or harbor complex drug resistance mutations. In the absence of genotypic resistance data, the association between PI regimens and high failure rates is highly likely due to the fact that patients with more complex conditions and poorer adherence were prescribed PIs. The authors cannot directly infer that PI drugs caused the failure in an observational study lacking a causal inference design, nor should they propose policy interventions based on this.
- The most central determinant of ART outcomes is medication adherence. However, Table 1 indicates that this cohort has a staggering 84% (655/778) missing data rate for adherence. Failing to include this critical variable in the multivariable Cox and Fine-Gray models will lead to extremely severe bias.
- In the methodology, the authors defined the patients' adherence level as the "worst recorded category" during follow-up, and used the "most recent record" for marital and employment status. In survival analysis, patients with longer survival times and more follow-up visits naturally have a higher probability of being recorded with at least one instance of poor adherence or a change in status. These variables must be treated as time-varying covariates; otherwise, only baseline data (at the time of ART transition) can be used for analysis.
- The study treats loss to follow-up (LTFU) directly as non-informative censoring. In HIV cohorts, loss to follow-up is often highly correlated with treatment failure or unrecorded mortality. The authors need to supplement the study with sensitivity analyses using inverse probability of censoring weighting (IPCW) or extreme value assumptions to evaluate the potential distortion of survival estimates caused by loss to follow-up.
- The primary analysis defines VF as the "first" viral load greater than 1000 copies/mL, whereas the strict definition requiring "two consecutive" measurements in the sensitivity analysis caused the number of events to drop sharply from 81 to 44. This indicates that nearly half of the "failures" may merely represent a single viral blip. The authors need to discuss the impact of this definitional difference on actual clinical intervention metrics more deeply in the discussion section.
- The discussion section attributes the extremely low mortality rate of 1.7% primarily to the healthy survivor effect. As data based on a passive routine surveillance system, although ascertainment bias was mentioned, its systematic underestimation of the extremely low mortality rate was not fully considered in the conclusion. The current inferences remain overly absolute.
- Only 13 deaths were observed over 3,564 person-years of follow-up. When the incidence rate of a competing event is extremely low, the results derived from the conventional Cox model and the Fine-Gray model typically do not show significant mathematical differences. The authors overstated the actual clinical value of introducing the competing risks model for correcting this dataset in the discussion.
Specific comments
- The abstract appears incomplete. The text is truncated at the results section reading: "with WHO stage IV also independently associated 37 (sHR 3.56, 95% CI: 1.39–9.12.".
- When describing the Cox model, the authors explicitly stated the use of "facility-clustered robust standard errors", but they did not mention whether the same clustering adjustment was applied when describing the Fine-Gray model. If unadjusted, the confidence intervals might be underestimated, leading to false-positive conclusions.
- The article emphasizes that WHO stage IV is a strong predictor of the composite outcome, but baseline data show that only 25 patients (3.2%) were in WHO stage IV, and its aHR confidence interval is extremely wide (1.27 to 6.70). A limitation statement regarding the statistical power and estimation precision of this very small sample subgroup needs to be added to the discussion.
- Addressing the 84% missing adherence data, the authors must provide a comparison table of baseline demographic and clinical characteristics between the "missing adherence data group" and the "complete data group" in the supplementary materials to demonstrate whether the missingness is random or if a systematic bias exists.
- At the bottom of the Kaplan-Meier survival curves in Figure 1 and Figure 2 of the results section, "Number at risk" tables were not provided. This does not comply with reporting standards for survival analysis and must be added.
- In Table 4 displaying the competing risks, the SHR values and P-value areas corresponding to all reference categories (ref) are blank. Baseline values (such as 1.00) should be added to standardize the table presentation.
- Some CD4 units use "uL".
- "≥" and ">=" appear interchangeably. It is recommended to unify standard mathematical symbols throughout the text.
- "PLHIV" should be "PLWH".
Author Response
Reviewer 2
|
Reviewer Comment (verbatim) |
Response and Revised Manuscript Text |
|
Comments regarding general concepts |
|
|
1. The paper concludes that PI-based regimens are the strongest independent risk factor for VF. In real-world clinical situations, patients typically use PI regimens because they have already experienced prior treatment failure or harbor complex drug resistance mutations. In the absence of genotypic resistance data, the association between PI regimens and high failure rates is highly likely due to the fact that patients with more complex conditions and poorer adherence were prescribed PIs. The authors cannot directly infer that PI drugs caused the failure in an observational study lacking a causal inference design, nor should they propose policy interventions based on this. |
Response: Causal language has been removed throughout. The Discussion and Conclusion now explicitly attribute the PI-based association to possible confounding by indication, and clinical recommendations were reframed from a directive to switch regimen class to individualized, resistance-informed review. Revised manuscript text: “However, those trials do not establish that all patients receiving PI-based therapy in routine care should switch to an INSTI; the present association may be partly explained by confounding by indication, unmeasured resistance, prior regimen failure, or adherence.” “Because regimen allocation was not randomized and resistance data were unavailable, the regimen association should inform intensified monitoring and resistance-guided clinical review rather than be interpreted as a causal comparison.” |
|
2. The most central determinant of ART outcomes is medication adherence. However, Table 1 indicates that this cohort has a staggering 84% (655/778) missing data rate for adherence. Failing to include this critical variable in the multivariable Cox and Fine-Gray models will lead to extremely severe bias. |
Response: Adherence was deliberately excluded from the adjusted models (it is post-baseline and 84% missing) rather than being force-included and biasing the estimates. A new supplementary table compares patients with vs. without a recorded adherence value, and the resulting potential for residual confounding is now flagged explicitly in the Discussion. |
|
3. In the methodology, the authors defined the patients' adherence level as the "worst recorded category" during follow-up, and used the "most recent record" for marital and employment status. In survival analysis, patients with longer survival times and more follow-up visits naturally have a higher probability of being recorded with at least one instance of poor adherence or a change in status. These variables must be treated as time-varying covariates; otherwise, only baseline data (at the time of ART transition) can be used for analysis. |
Response: Marital and employment status are now taken from the CRF2 record nearest the transition date (not the most recent record), so they reflect a fixed baseline value rather than a follow-up-length-dependent one. Adherence, which cannot be reduced to a single baseline value, was excluded from the adjusted models entirely rather than forced into a biased fixed-covariate form. Revised manuscript text: “Marital and employment status were taken from the CRF2 record nearest the transition date, preferring a record on or before transition date and using the nearest subsequent record only if none was available beforehand, and were treated as fixed covariates.” |
|
4. The study treats loss to follow-up (LTFU) directly as non-informative censoring. In HIV cohorts, loss to follow-up is often highly correlated with treatment failure or unrecorded mortality. The authors need to supplement the study with sensitivity analyses using inverse probability of censoring weighting (IPCW) or extreme value assumptions to evaluate the potential distortion of survival estimates caused by loss to follow-up. |
Response: An extreme-value bounding sensitivity analysis (Table S2) was added, re-estimating the composite incidence rate under the assumption that 10% or 25% of censored patients experienced an unrecorded event. A full IPCW model was not fitted (unstable given the modest event count), but the bounding analysis directly quantifies the potential scale of informative-censoring bias and is now discussed as a limitation. Revised manuscript text: “Table S2: 0% unrecorded → 94 events, 2.64/100 PY; 10% unrecorded → 162 events, 4.54/100 PY; 25% unrecorded → 265 events, 7.43/100 PY.” “Loss to follow-up may have concealed deaths or VF; extreme-value bounding showed that even a 10% unrecorded event rate among censored patients would nearly double the composite incidence rate, indicating this remains an important source of uncertainty.” |
|
5. The primary analysis defines VF as the "first" viral load greater than 1000 copies/mL, whereas the strict definition requiring "two consecutive" measurements in the sensitivity analysis caused the number of events to drop sharply from 81 to 44. This indicates that nearly half of the "failures" may merely represent a single viral blip. The authors need to discuss the impact of this definitional difference on actual clinical intervention metrics more deeply in the discussion section. |
Response: The Discussion now explicitly links the sensitivity-definition results to the blip interpretation: the PI-based association is directionally consistent but attenuated under the confirmed-VF definition, indicating a portion of primary-endpoint events likely reflect transient viraemia rather than confirmed failure. Revised manuscript text: “Under the sensitivity VF definition requiring two consecutive elevated viral loads (Table S3), the PI-based regimen association was directionally consistent but attenuated (aHR 1.72, 95% CI 0.98–3.04) compared with the primary single-VL definition, suggesting that some primary-endpoint events reflect transient viraemia (“blips”) rather than confirmed failure.” |
|
6. The discussion section attributes the extremely low mortality rate of 1.7% primarily to the healthy survivor effect. As data based on a passive routine surveillance system, although ascertainment bias was mentioned, its systematic underestimation of the extremely low mortality rate was not fully considered in the conclusion. The current inferences remain overly absolute. |
Response: Language was softened throughout: the low mortality rate is now attributed jointly to survivor selection and incomplete death ascertainment, and the Conclusion adds an explicit caution against over-interpreting this estimate. Revised manuscript text: “The low recorded mortality rate may reflect Rwanda's strong treatment programme, but it may also be influenced by survivor selection and incomplete ascertainment of deaths occurring after loss to follow-up [21,22,29–32].” “The low recorded mortality rate should also be interpreted cautiously given the small number of events and possible under-ascertainment.” |
|
7. Only 13 deaths were observed over 3,564 person-years of follow-up. When the incidence rate of a competing event is extremely low, the results derived from the conventional Cox model and the Fine-Gray model typically do not show significant mathematical differences. The authors overstated the actual clinical value of introducing the competing risks model for correcting this dataset in the discussion. |
Response: The Results text now states plainly that the Fine-Gray and Cox models yielded a similar pattern, rather than claiming the competing-risks approach materially changed the conclusions — consistent with the low death count the reviewer notes. Revised manuscript text: “The Fine–Gray model for VF, with death treated as a competing event, yielded a similar pattern to the composite-outcome Cox model.” |
|
Specific comments |
|
|
1. The abstract appears incomplete. The text is truncated at the results section reading: "with WHO stage IV also independently associated 37 (sHR 3.56, 95% CI: 1.39--9.12.". |
Response: Fixed. The abstract now reports complete Results and Conclusions sections. Revised manuscript text: “In adjusted Cox regression, PI-based regimen (adjusted hazard ratio [aHR] 2.16, 95% CI 1.38–3.39), WHO stage III (aHR 1.92, 95% CI 1.10–3.33), and WHO stage IV (aHR 3.56, 95% CI 1.58–7.99) were associated with the composite outcome. In Fine–Gray analysis, PI-based regimen (subdistribution hazard ratio [sHR] 3.39, 95% CI 1.98–5.80) and WHO stage IV (sHR 4.41, 95% CI 1.88–10.3) were associated with VF.” |
|
2. When describing the Cox model, the authors explicitly stated the use of "facility-clustered robust standard errors", but they did not mention whether the same clustering adjustment was applied when describing the Fine-Gray model. If unadjusted, the confidence intervals might be underestimated, leading to false-positive conclusions. |
Response: Clarified. Facility-clustered robust standard errors were used for both models; the Methods now state this explicitly for the Fine-Gray model as well. Revised manuscript text: “For VF, cumulative incidence functions were estimated using the Aalen–Johansen estimator, and Fine–Gray regression, with delayed entry from day 180 and robust standard errors clustered by health facility, treated death as the competing event.” |
|
3. The article emphasizes that WHO stage IV is a strong predictor of the composite outcome, but baseline data show that only 25 patients (3.2%) were in WHO stage IV, and its aHR confidence interval is extremely wide (1.27 to 6.70). A limitation statement regarding the statistical power and estimation precision of this very small sample subgroup needs to be added to the discussion. |
Response: Events-per-parameter (EPV) is now reported for both regression models and flagged in the Discussion as the reason wide, unstable confidence intervals should be expected for lower-frequency categories such as WHO stage IV. Revised manuscript text: “The small numbers of deaths and third-line patients limited endpoint-specific analyses, and events-per-parameter was below the conventional threshold of 10 for both the composite (6.3) and VF-specific (5.4) models, indicating wider, less stable confidence intervals for lower-frequency categories.” |
|
4. Addressing the 84% missing adherence data, the authors must provide a comparison table of baseline demographic and clinical characteristics between the "missing adherence data group" and the "complete data group" in the supplementary materials to demonstrate whether the missingness is random or if a systematic bias exists. |
Response: Added as Supplementary Table S5, comparing baseline characteristics and outcomes between patients with a recorded vs. missing adherence value. Revised manuscript text: “Table S5. Baseline characteristics and outcomes by availability of a recorded adherence value (Recorded N=123 vs Missing N=655). Significant differences: age (p=0.007), marital status (p=0.04), CD4 category (p=0.01), virologic failure (17% vs 9.2%, p=0.008), composite outcome (20% vs 11%, p=0.006).” |
|
5. At the bottom of the Kaplan-Meier survival curves in Figure 1 and Figure 2 of the results section, "Number at risk" tables were not provided. This does not comply with reporting standards for survival analysis and must be added. |
Response: Added. Numbers-at-risk tables now appear below each Kaplan-Meier panel in Figures 1 and 2. Revised manuscript text: “Numbers at risk are displayed below each panel. Log-rank p-value shown for stratified comparison.” |
|
6. In Table 4 displaying the competing risks, the SHR values and P-value areas corresponding to all reference categories (ref) are blank. Baseline values (such as 1.00) should be added to standardize the table presentation. |
Response: Fixed. All reference categories in Tables 3 and 4 now display an explicit sHR/HR of 1.00. Revised manuscript text: “I (ref) — 1.00; II — 1.45 (0.72, 2.91), p=0.29; III — 1.79 (0.98, 3.26), p=0.06; IV — 4.41 (1.88, 10.3), p<0.001 (Table 4, WHO clinical stage rows).” |
|
7. Some CD4 units use "uL". |
Response: Standardized to µL throughout the manuscript and all tables. Revised manuscript text: “CD4 count at transition (<200, 200–499, ≥500 cells/µL, or unknown)” |
|
8. “≥” and “>=” appear interchangeably. It is recommended to unify standard mathematical symbols throughout the text. |
Response: Standardized to ≥ throughout. One residual ">=500" instance in the sensitivity Cox model table (Table S3) was identified and corrected in this revision. Revised manuscript text: “≥500 (ref) — Table S3, CD4 count row (corrected from ">=500 (ref)").” |
|
9. "PLHIV" should be "PLWH". |
Response: All instances changed to PLWH throughout the manuscript. Revised manuscript text: “people living with HIV (PLWH)” |
Author Response File:
Author Response.pdf
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsOverall the paper has much improved and clearer. Two additional comments:
In TABLE S4: Specific regimen (drug combination, including NRTI backbone) switched to ..
ATV is boosted (ATV-r) or unboosted (if so, at which daily dose?). If this is the case, given that ATV represents 43% of PI-based regimen people have been switched to (if I got correctly), this should be stressed in the “discussion section” as boosted PI is NOT unboosted PI in terms of genetic barrier. Actually DRV-r and LOP-r account for 4% only, a really marginal proportion.
So sentence in the line 228 needs adjustment when considering the underperformance of PI-based regimen (ATV-based actually)
BTW: Does it reflect the common PI-based regimen used in other African countries today?
Table S6. Viral load monitoring frequency and timing after ART line transition.
I suggest to add the important information in the table. : … Monitoring intervals are protocol-defined nationally and do not vary by regimen class.
However the point is not whether there is an indication for different VL monitoring by classes, but whether in real world the monitoring in the 2 groups were different. A higher VL monitoring in one group can capture a higher number of events.
Line 120 .. “a categorical clinician assessment (good, moderate, or bad) per Rwanda's national ART programme M&E tools; CBS 120 does not capture the specific counselling criteria (e.g., pill count or missed-dose recall) underlying this categorical judge-121 ment.”
That point is very important and should be discussed as the info is limited overall (available only for 16% of PWH) and the criteria are very discretional.
In TABLE S5: for age, median (IQR), Median (Q1–Q3) is reapeted
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors have basically addressed the issues raised in the previous review, but there are still a few minor formatting errors that need to be revised.
Specifically, ">=" is still used instead of "≥" in the missing data comparison section of Table 1. The abbreviation "PLHIV" still appears in the titles of Table 1 and Table 2 and should be updated to "PLWH".
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf

