1. Introduction
Obstructive sleep apnea (OSA) is a highly prevalent chronic respiratory disorder characterised by repetitive episodes of partial or complete upper airway obstruction during sleep, resulting in intermittent hypoxia, sleep fragmentation, and sustained activation of the sympathetic nervous system. Beyond its well-established cardiovascular consequences—hypertension, cardiac arrhythmias, coronary artery disease, stroke, and heart failure—OSA imposes a substantial and frequently underappreciated burden on the respiratory system through chronic intermittent hypoxaemia. A meta-analysis of six prospective observational studies encompassing nearly 12,000 patients reported that severe OSA was associated with a pooled hazard ratio of 1.90 (95% CI 1.29–2.81) for all-cause mortality and 2.65 (95% CI 1.82–3.85) for cardiovascular mortality [
1]. Whether effective treatment translates into a meaningful reduction in mortality remains one of the most clinically consequential and contested issues in respiratory and sleep medicine.
Continuous positive airway pressure (CPAP) remains the cornerstone of OSA therapy. Beyond reducing the apnea–hypopnea index (AHI) and alleviating daytime sleepiness, CPAP attenuates nocturnal blood pressure surges, reduces sympathetic overdrive, improves endothelial function, and mitigates systemic inflammation [
2,
3]. Observational evidence has supported a mortality benefit: untreated severe OSA was independently associated with cardiovascular death in elderly patients, a risk significantly reduced by adequate CPAP treatment [
4]. Landmark data in men confirmed higher rates of fatal and non-fatal cardiovascular events in untreated compared with CPAP-treated patients [
5], a pattern subsequently confirmed in female cohorts [
6].
The landscape shifted with large-scale randomised controlled trials. The SAVE trial, enrolling 2717 adults with moderate-to-severe OSA and established cardiovascular disease, found no significant reduction in the primary composite cardiovascular endpoint (HR 1.10, 95% CI 0.91–1.32) [
7]. A meta-analysis found CPAP associated with reduced events in observational studies (RR 0.61, 95% CI 0.39–0.94), but not in RCTs (RR 0.57, 95% CI 0.32–1.02) [
8]. This discrepancy has been attributed to the “healthy user” effect in observational cohorts, and insufficient CPAP adherence in pragmatic trials [
7,
8].
Treatment adherence appears critical: protective effects are consistently observed at ≥4 h per night, with a dose–response extending to ≥8 h [
9,
10,
11]. Benefits may also be delayed, becoming apparent only after years of sustained treatment [
12,
13], and studies enrolling patients at highest cardiovascular risk have demonstrated the largest absolute mortality differences [
4,
14,
15]. A recent meta-analysis of over one million patients has sought to reconcile these discordant findings [
16]. Whether the null RCT findings reflect a true absence of benefit or are a consequence of insufficient adherence, suboptimal patient selection, and inadequate follow-up remains unresolved.
A critical limitation common to nearly all prior work is the treatment of CPAP as a binary exposure (treated vs. untreated), which ignores the wide variation in the physiological response actually achieved during real-world home use. Two patients prescribed CPAP may differ substantially in residual disease burden and residual nocturnal hypoxaemia, yet conventional analyses assign them identical exposure status. We hypothesised that the quality of the therapeutic response—quantified objectively from device-recorded residual AHI across home therapy nights—would carry prognostic information beyond CPAP prescription alone, and that the survival benefit of CPAP would be concentrated in patients with more severe disease. In a single-centre cohort of 4368 adults followed for up to 20 years with complete national death-registry linkage, we therefore (i) quantified the association of CPAP therapy with all-cause, cardiovascular, and pulmonary mortality across the spectrum of OSA severity; (ii) introduced and evaluated a response-quality classification based on residual AHI; and (iii) used complementary survival models (cause-specific Cox, Fine–Gray competing risks) and machine learning to disentangle the prognostic roles of residual AHI control and nocturnal hypoxaemic burden, and, ultimately, to identify which patients are the true beneficiaries of CPAP therapy.
2. Materials and Methods
2.1. Study Population and Data Source
We analysed a single-centre cohort of 4368 consecutive patients referred for full-night attended polysomnography at the Department of Sleep Medicine and Metabolic Disorders, Medical University of Lodz, Poland. Cox proportional hazards regression and Random Survival Forest (RSF) mortality models were applied to both the full cohort (n = 4368) and a prespecified subgroup of patients with an apnea–hypopnea index ≥ 15 events/hour (n = 2304). Participants were followed for 8 to 20 years, during which 1017 individuals died and 3351 (76.72%) remained alive at the end of follow-up (
Table 1). In the AHI ≥ 15 subgroup, 737 deaths occurred (32.0%). Exclusion criteria included age < 18 years and BMI < 18 kg/m
2.
OSA severity was classified as normal (AHI < 5), mild (AHI 5–14.9), moderate (AHI 15–29.9), or severe (AHI ≥ 30). All-cause mortality was ascertained through linkage with national registry data; cause of death was coded using ICD-10 and grouped by organ system (
Table 2).
Among the 1017 deceased patients (23.28%), cardiovascular diseases were the leading cause of death (n = 335; 7.67% of the total cohort), followed by neoplasms (n = 243; 5.56%) and respiratory diseases (n = 101; 2.31%) (
Table 2).
2.2. CPAP Therapy
CPAP therapy was defined as ≥4 h per night on at least 75% of nights. Under this definition, 956 patients (41.49%) in the AHI ≥ 15 subgroup were classified as CPAP-treated. The median duration of CPAP therapy was 169 days (IQR 58–683; range 15–4708 days).
The median CPAP therapy duration of 169 days (IQR 58–683) refers to the period over which device-recorded residual-AHI data were available for each treated patient—the window used to characterise response quality—and does not represent the total lifetime duration of CPAP use. In the survival models, CPAP was modelled as a fixed baseline (ever-treated) exposure rather than as a time-varying covariate.
CPAP was offered to all patients meeting clinical criteria; the specific reason a given patient did not ultimately receive it (non-acceptance, intolerance, or clinical unsuitability) was not systematically recorded, so the No-CPAP group is heterogeneous and a degree of treatment-selection bias cannot be excluded.
2.3. Cox Proportional Hazards Regression
Cox proportional hazards (PH) regression was performed separately in the AHI ≥ 15 subgroup and the full cohort. Three multivariable models were fitted: (i) a binary CPAP therapy model (yes/no), (ii) a model incorporating a multiplicative AHI × CPAP interaction term to assess severity-dependent treatment effects, and (iii) a CPAP response classification model replacing the binary CPAP variable with three-category response (Good, Poor, No CPAP). All multivariable models were adjusted for age, sex, BMI, AHI, oxygen desaturation (Des ≤ 90% TST), smoking, hypertension, diabetes mellitus, and hypothyroidism. The proportional hazards assumption was verified using Schoenfeld residuals. Discrimination was assessed by the concordance index (C-index) and calibration by the Integrated Brier Score (IBS).
2.4. Cause-Specific Cox Regression
Cause-specific Cox proportional hazards models were fitted separately for cardiovascular and pulmonary mortality. In each model, deaths from all other causes were treated as censored observations at the time of their occurrence. This approach estimates the instantaneous hazard of the event of interest among individuals who have not yet experienced any event. Univariable models evaluated each clinical, demographic, and polysomnographic variable individually. Multivariable models were adjusted for age, sex, BMI, AHI, oxygen desaturation (Des ≤ 90% TST), smoking, hypertension, diabetes mellitus, and hypothyroidism. Separate multivariable models were fitted for binary CPAP therapy and CPAP response classification.
2.5. Fine–Gray Competing Risks Regression
To complement the cause-specific Cox models, Fine–Gray subdistribution hazard regression was performed for cardiovascular and pulmonary mortality. Unlike cause-specific hazards, the Fine–Gray model directly quantifies the effect of covariates on the cumulative incidence function (CIF) while retaining subjects who die of competing causes in the risk set. This is particularly relevant here, as non-cardiovascular and non-pulmonary deaths constitute a substantial proportion of events. Univariable models assessed each sleep-related and treatment variable individually; multivariable models were adjusted for age, sex, BMI, smoking, hypertension, diabetes, and dyslipidemia.
2.6. CPAP AHI Measurement and Response Classification
The residual AHI on CPAP (AHICPAP) was calculated as the mean AHI recorded by the CPAP device, averaged across all therapy nights on which the patient actively used the device at home. The AHI reduction was computed as [AHIdiagnostic − AHICPAP]/AHIdiagnostic × 100. Patients were classified into response categories as follows. Good response: AHI reduction ≥ 50% and residual AHI < 10 events/h. Poor response: AHI reduction < 50%, or AHI reduction ≥ 50%, but residual AHI ≥ 10. No CPAP: did not receive CPAP therapy (reference category).
The response thresholds were selected to reflect current clinical practice for defining adequate CPAP control. A residual AHI below 10 events/hour corresponds to the widely used target for effective therapy, whereas the additional requirement of at least a 50% reduction from the diagnostic AHI ensures that classification as a good responder reflects a substantial improvement relative to each patient’s own baseline rather than a low residual value arising merely from a modest starting AHI. Patients meeting only one of these two criteria were classified as poor responders. These cut-offs are pragmatic and, to our knowledge, not yet formally standardised; they are intended as a reproducible, clinically interpretable operationalisation of response quality and warrant external validation.
A Random Forest classifier (500 trees) and a regularised multinomial logistic regression were trained to predict CPAP response category (Good, Poor, No CPAP) using stratified cross-validation. Performance was summarised with accuracy, macro F1, and macro-averaged AUC.
2.7. Statistical Analysis
Spearman rank correlations were used for continuous variables; Mann–Whitney U and Kruskal–Wallis tests for categorical comparisons; post hoc pairwise tests used Dunn’s method with Bonferroni correction. Analyses were performed in Python (version 3.14) (lifelines, scikit-survival, scikit-learn). A two-sided p < 0.05 was considered statistically significant.
This study was conducted in compliance with the amended Declaration of Helsinki, and the Bioethics Committee of the Medical University of Łódź, Poland, approved the study protocol (RNN/23/15/KE; RNN/393/19/KE). This study was neither funded by an institutional grant nor by the pharmaceutical industry or a medical company.
4. Discussion
This study demonstrates that CPAP therapy confers a significant, severity-dependent survival benefit in patients with moderate-to-severe OSA, with the magnitude of benefit further determined by the quality of the therapeutic response. The central message is therefore one of patient selection: the true beneficiaries of CPAP are patients with severe disease who achieve effective control of residual events and nocturnal hypoxaemia during home use, rather than the OSA population at large. The AHI × CPAP interaction and stratified survival analyses converge on a consistent finding: the mortality benefit of CPAP is greatest in more severe disease. This severity-dependent pattern is consistent with landmark observational studies [
4,
5,
6] and with the biological plausibility of treating a condition whose pathophysiological sequelae scale with event frequency and hypoxaemic burden.
From a respiratory-medicine perspective, the most novel finding is the dissociation between the drivers of cardiovascular and pulmonary death. The Fine–Gray analyses reveal that the percentage of total sleep time spent below 90% oxygen saturation (Des < 90% TST), rather than CPAP-related variables, is the principal driver of cause-specific pulmonary mortality after competing risks are accounted for—emphasising the importance of hypoxaemic burden as a treatment target independent of the AHI. Nocturnal desaturation was the dominant predictor of pulmonary mortality in both populations, and retained independent prognostic importance in the full cohort after multivariable adjustment. The elevated subdistribution hazards associated with CPAP use in univariable Fine–Gray models reflect confounding by indication rather than a true adverse effect. The CPAP response classification introduced here extends the binary treatment paradigm by showing that the quality of AHI reduction on home CPAP therapy independently predicts survival; the dose–response gradient observed in the Kaplan–Meier curves (good responders > poor responders > no CPAP) corroborates observational findings of incremental mortality reduction with greater nightly usage [
9,
10,
11]. Caution is nonetheless warranted, as the response classification is based on the mean residual AHI recorded by the device across therapy nights, and longitudinal adherence data were not available.
The apparent divergence between the standard and competing-risks analyses reflects the different questions these models address rather than an inconsistency in the data. The cause-specific Cox model estimates the instantaneous rate of cardiovascular death among patients still at risk and indicates a protective association of CPAP, whereas the Fine–Gray subdistribution models estimate effects on the cumulative incidence of cardiovascular death while retaining patients who die of competing causes—chiefly neoplasms and respiratory disease—in the risk set. Under a high competing-risk burden, a reduction in the cause-specific hazard need not translate into a lower cumulative incidence, and stratifying CPAP by response category further reduces group sizes and widens confidence intervals. The cause-specific and subdistribution hazards should therefore be regarded as complementary, and the cardiovascular survival advantage interpreted with corresponding caution.
Our findings contribute to resolving the well-documented discordance between observational studies and RCTs on CPAP and mortality [
16]. The null findings of the SAVE trial [
7] and negative meta-analyses of RCTs [
17,
18] may be explained by several factors: CPAP adherence in our cohort was defined as ≥4 h/night on ≥75% of nights—substantially exceeding the 3.3-h mean in SAVE [
7]; our 8–20 year follow-up enabled detection of benefits that may require years of cumulative therapy to emerge [
12,
13]; and stratified analyses reveal that the benefit is confined to severe OSA, a subgroup effect inevitably diluted in intention-to-treat analyses of mixed-severity populations [
8,
14]. Regarding cardiovascular mortality specifically, a meta-analysis of nine studies reported a 66% reduction (HR 0.34, 95% CI 0.17–0.68) attributable predominantly to observational data [
19], consistent with our finding of a 29% cardiovascular mortality reduction in the AHI ≥ 15 subgroup.
These findings carry direct implications for clinical practice and population respiratory health. First, they support concentrating CPAP resources and adherence-support efforts on patients with severe OSA, in whom the absolute survival benefit is greatest, rather than distributing them uniformly across all OSA severities. Second, the prognostic value of response quality argues for routine monitoring of device-recorded residual AHI as a modifiable treatment target, shifting the clinical goal from merely initiating CPAP to verifying that effective disease control is achieved during home use. Third, the divergent drivers of cardiovascular and pulmonary mortality—residual AHI control versus nocturnal hypoxaemic burden—suggest that oximetric endpoints deserve explicit attention alongside the AHI when titrating therapy and designing trials. Because baseline polysomnographic and clinical features predicted response category with high accuracy, such stratification could in principle be applied at the point of diagnosis, enabling more efficient enrichment of future randomised trials and more personalised management pathways. Finally, because the survival benefit of CPAP is concentrated in patients who achieve effective disease control, those with poor adherence or an inadequate therapeutic response warrant complementary, interdisciplinary management: early screening for OSA in dental settings and the use of mandibular advancement devices as an alternative for patients with low CPAP compliance may broaden effective treatment and improve outcomes [
20].
Several limitations warrant consideration. The observational design precludes causal inference, and residual confounding—including the “healthy user” effect [
8]—cannot be fully excluded. The CPAP response classification is based on device-recorded residual AHI averaged across therapy nights and does not capture longitudinal changes in adherence or efficacy. The classification models achieved macro F1-scores of 0.57–0.63, reflecting class imbalance—particularly for good responders—and require external validation before clinical deployment. Finally, the single-centre design may limit generalisability, and external validation of the prediction models in independent respiratory cohorts is required.
A further limitation is the potential selection bias inherent to the observational design: patients who received CPAP may have differed systematically from those who did not in ways only partly captured by the measured covariates. Despite adjustment for the available demographic, polysomnographic, and comorbidity variables, residual confounding by unmeasured factors (overall health status, socioeconomic position, motivation, adherence support) cannot be excluded, so the survival benefit should be interpreted as an association—possibly reflecting in part a “healthy user” effect—rather than proof of causation.
Because continued adherence throughout the 8–20-year follow-up could not be verified, and CPAP was modelled as a baseline exposure rather than a time-varying covariate, immortal-time bias and a healthy-user effect cannot be fully excluded; the observed association is best interpreted as reflecting treatment initiation rather than proven sustained protection.
Information on several potentially relevant comorbidities—chronic lung disease, heart failure, chronic kidney disease, and malignancy—was not collected in the database and could not be included as covariates, representing a further potential source of residual confounding.