1. Introduction
Over the past quarter-century, the therapeutic landscape for axial spondyloarthritis (axSpA) has undergone substantial transformation. A pivotal driver of this progress has been the introduction of biological (b-) and targeted synthetic (ts-) disease-modifying anti-rheumatic drugs (DMARDs) [
1,
2]. These agents have revolutionized patient outcomes by targeting specific inflammatory pathways, offering new hope for individuals who previously had limited treatment options. The development of tumor necrosis factor inhibitors (TNFis) marked the first major breakthrough, followed by the emergence of agents with alternative mechanisms of action, including interleukin-17 inhibitors (IL-17is), interleukin-23 inhibitors (IL-23is), and Janus kinase inhibitors (JAKis) [
1]. This expansion has provided clinicians with an unprecedented array of therapeutic choices, but has also introduced new complexities in treatment sequencing and optimization.
Nevertheless, primary inefficacy or adverse events necessitate treatment discontinuation in approximately one-third of patients undergoing a first course of advanced therapy (b/tsDMARDs) [
3]. This high rate of treatment failure underscores the critical importance of understanding optimal strategies for subsequent lines of therapy. The reasons for discontinuation are multifaceted and may include inadequate efficacy, loss of response over time, intolerance, or the development of adverse events ranging from mild injection site reactions to serious infections or malignancies [
3]. Each treatment failure presents a clinical dilemma regarding the most appropriate next step.
Clinicians now have access to agents with diverse mechanisms of action (MOAs), including TNFis, IL-17is, IL-23is, and JAKis. TNFis, such as infliximab, adalimumab, etanercept, certolizumab pegol and golimumab, have been the mainstay of treatment for nearly two decades, with extensive real-world evidence supporting their efficacy and safety. IL-17is, including secukinumab, ixekizumab and bimekizumab, target the IL-17 pathway and have shown particular efficacy in axial manifestations. IL-23is, such as guselkumab, risankizumab and ustekinumab, while primarily indicated for psoriatic arthritis and psoriasis, are sometimes used in axSpA patients with psoriasis. JAKis, such as tofacitinib, baricitinib, filgotinib and upadacitinib, represent the newest class of oral advanced therapies, offering an alternative route of administration and distinct mechanism of action targeting intracellular signaling pathways [
2].
Following the failure of an advanced therapeutic line, two principal strategies are considered: cycling (switching to another agent with the same MOA) or swapping (switching to an agent with a different MOA). The choice between these strategies involves multiple considerations, including the reason for initial treatment failure, patient comorbidities, extra-articular manifestations, drug availability and cost, and patient preference. For instance, patients who fail a TNFi due to primary inefficacy may theoretically benefit from swapping to a drug with a different MOA, whereas those who fail due to secondary inefficacy or adverse events might be candidates for either cycling or swapping. The optimal approach remains a subject of ongoing debate in the absence of robust comparative data.
Given the expanding therapeutic arsenal, randomized controlled trials (RCTs) directly comparing the efficacy of cycling versus swapping strategies after advanced therapy failure in axSpA are lacking. While RCTs represent the gold standard for evidence generation, their execution in this context presents substantial challenges, including the need for large sample sizes, lengthy follow-up periods, and the ethical considerations of randomizing patients to potentially suboptimal strategies. Consequently, real-world observational studies have become essential sources of evidence to guide clinical decision-making [
4,
5,
6,
7,
8].
Existing observational data suggest that, following first-line TNFi failure, both strategies demonstrate comparable effectiveness [
4,
5,
6,
7,
8]. However, these studies have several limitations. Many are derived from registries that predate the availability of newer drugs with different MOAs such as IL-23is and JAKis. Others focus exclusively on TNFi to IL-17i swapping, without considering the full spectrum of available agents. Furthermore, few studies have comprehensively adjusted for potential confounders, including disease duration, line of therapy, and extra-articular manifestations. Importantly, no studies to date have systematically examined outcomes following failure of non-TNFi advanced therapies, a knowledge gap that our study aims to address.
The primary objective of this retrospective, observational study was to evaluate the comparative effectiveness, measured by treatment persistence, of cycling versus swapping strategies in axSpA patients after failure of any advanced therapy line. Treatment persistence, defined as the duration of time from treatment initiation to discontinuation, serves as a validated proxy for real-world effectiveness, incorporating both efficacy and tolerability into a single composite outcome. High persistence rates indicate treatments that are both effective and well-tolerated in routine clinical practice, while low persistence suggests limitations in either domain.
A secondary objective was to identify clinical characteristics associated with a higher probability of treatment discontinuation. Understanding patient-level predictors of treatment success or failure can inform personalized treatment selection, moving beyond the “one-size-fits-all” approach to a more nuanced, stratified medicine paradigm. Potential predictors of interest include demographic factors (age, sex), disease characteristics (radiographic versus non-radiographic axSpA, disease duration), extra-articular manifestations (psoriasis, inflammatory bowel disease), and treatment-related factors (drug class, line of therapy, year of initiation).
2. Materials and Methods
This review of medical records was conducted in accordance with the principles of the Declaration of Helsinki and received approval from the local Ethics Committee (Protocol No. 34713). Written informed consent was obtained from all participants. The retrospective observational study was designed and reported in accordance with the STROBE (Strengthening the Reporting of Observational Studies in Epidemiology) guidelines for observational research.
2.1. Study Population
We included all patients aged ≥18 years with a diagnosis of axSpA who commenced treatment with a bDMARD or tsDMARD, in accordance with local guidelines and the treating physician’s assessment, at our Rheumatology Unit in Parma (Italy) between January 2004 and July 2025 and who experienced failure of at least one line of advanced therapy. The diagnosis of axSpA was made by experienced rheumatologists based on the Assessment of SpondyloArthritis international Society (ASAS) classification criteria. Both radiographic (ankylosing spondylitis) and non-radiographic axSpA patients, as defined by ASAS classification criteria, were included to reflect the full spectrum of the disease encountered in clinical practice.
Failure was defined as discontinuation due to primary or secondary inefficacy or adverse events. Primary inefficacy was defined as inadequate response within the first 3–6 months of treatment despite adequate dosing and adherence. Secondary inefficacy was defined as loss of initially achieved response after at least 6 months of treatment. Adverse events leading to discontinuation were categorized as infections, malignancies, or other events based on clinical documentation. Treatment discontinuations due to patient choice, pregnancy planning, switch to biosimilar or loss to follow-up were excluded from the failure definition to maintain focus on clinically driven treatment changes. Similarly, patient lines switching from originator to biosimilar or vice versa were excluded from the study.
2.2. Effectiveness Assessment
Treatment persistence (drug survival) was employed as a proxy for real-world effectiveness. A detailed pharmacological history was collected for each patient. Data were extracted from electronic medical records and pharmacy databases by trained rheumatologists using a standardized data collection form. Treatment start and end dates were verified against prescription records and clinical notes to ensure accuracy. For patients who discontinued and later restarted the same agent, only the initial treatment course was considered in the primary analysis.
Line of therapy was defined chronologically, with each new b/tsDMARD—regardless of its mechanism of action—counted as a new line. Thus, switches between agents within the same drug class (e.g., between different TNF inhibitors) were considered subsequent lines of therapy. The first-ever b/tsDMARD prescribed was designated as line 1, irrespective of whether it had been initiated at our center or elsewhere. For patients referred after starting treatment at another institution, prior treatment records were retrieved whenever possible to accurately determine the correct line of therapy.
Agents were categorized by MOA: TNFis (including golimumab, certolizumab pegol, etanercept, adalimumab, infliximab, and their biosimilars), IL-23is (ustekinumab, guselkumab, risankizumab), IL-17is (secukinumab, ixekizumab), and JAKis (upadacitinib, tofacitinib). Biosimilars were considered equivalent to their originator biologics for MOA classification purposes. For combination therapies (e.g., TNFi plus methotrexate), only the b/tsDMARD component was considered for persistence analysis.
Each treatment course following an initial failure was classified: the patients that subsequently used an agent with the same MOA constituted the Cycling Group (CG); those who used an agent with a different MOAs constituted the Swapping Group (SG). The SG included patients that switched to a b/tsDMARD with a different MOA. Accordingly, TNFi courses in this group represent switches from a nonTNFi agent to a TNFi. In cases where patients had multiple treatment courses, each course was independently classified and included in the analysis, with appropriate statistical adjustments for within-patient correlation.
2.3. Statistical Analysis
Descriptive statistics are presented as medians with Inter Quartile Range (IQRs) or 95% confidence intervals (CIs), where appropriate. Between-group (CG vs. SG) comparisons for baseline characteristics were performed using the Chi-squared test or Kruskal–Wallis test, as appropriate. Normality of continuous variables was assessed using the Shapiro–Wilk test, which guided the choice of parametric versus non-parametric tests.
Treatment persistence was analyzed using Kaplan–Meier survival curves, with between-group differences assessed using the log-rank test. Patients were censored at the date of last follow-up, death, or administrative censoring (31 July 2025), whichever occurred first. Sensitivity analyses were performed excluding patients with less than 6 months of follow-up to assess the robustness of findings.
A Cox proportional hazards regression model was employed to identify factors independently associated with treatment discontinuation. The proportional hazards assumption was tested using Schoenfeld residuals. Variables included in the model were: age, sex, group (CG vs. SG), drug class, axSpA subtype (radiographic vs. non-radiographic), presence of inflammatory bowel disease (IBD), presence of psoriasis (PsO), disease duration, line of therapy, and year of treatment initiation. Variable selection was based on clinical relevance and univariate screening (p < 0.20). Moreover, a propensity score analysis was performed to adjust for potential baseline imbalances between CG and SG, and to ensure the robustness of the findings. The propensity score was estimated using a logistic regression model including all relevant baseline covariates, and incorporated into the analysis through covariate adjustment. Multiple imputation by chained equations was used to handle missing data for covariates with less than 10% missingness.
Because some patients contributed more than one treatment course, within-patient correlation was accounted for by using Cox proportional hazards models with robust standard errors clustered at the patient level.
A two-sided
p-value < 0.05 was considered statistically significant. Statistical analysis was performed using Jamovi (
https://www.jamovi.org, v 2.3). All analyses were independently verified by a second statistician blinded to group allocation.
3. Results
We analyzed 156 axSpA patients (59 radiographic axSpA, 97 non-radiographic axSpA) who failed at least one line of advanced therapy, corresponding to 343 treatment courses (CG: 213; SG: 130). The study cohort was predominantly middle-aged with a slight male predominance, reflecting the typical epidemiology of axSpA. The median follow-up duration was 4.7 years (IQR: 2.3–8.1 years), providing adequate observation time to assess medium-term treatment persistence.
Patients in the SG had a longer median disease duration, a higher prevalence of PsO, and had undergone more prior lines of advanced therapy compared to the CG. Patients in the CG had a higher prevalence of IBD. Baseline characteristics are detailed in
Table 1.
Regarding the distribution of drug classes, TNFis dominated in the CG (97.2%), consistent with their historical role as first-line agents. In contrast, the SG showed a more balanced distribution of drug classes, with IL-17is (48.5%) and TNFis (30.8%) being most common, followed by JAKis (13.0%) and IL-23is (7.7%).
The reasons for treatment failure were documented for 184 of 343 treatment courses (53.6%). Secondary failure was the most common cause in both groups (CG: 62/117, 529%; SG: 38/67, 56.7%), followed by primary failure (CG: 34/117, 29%; SG: 19/67, 28.3%). Adverse events accounted for a minority of discontinuations (CG:21/117, 17.9%; SG: 10/67, 14.9%).
The median drug survival was 970 days (95% CI: 685–1550) in the CG and 816 days (95% CI: 664–1547) in the SG (Hazard Ratio [HR]: 1.13, 95% CI: 0.83–1.53;
p = 0.442). Retention rates at 1, 2, and 3 years were 62.7%, 49.3%, and 39.2% for the CG, and 69.8%, 47.8%, and 31.8% for the SG, respectively (
Figure 1). The survival curves overlapped substantially throughout the follow-up period, with 95% confidence intervals showing considerable overlap at all time points. The log-rank test confirmed the absence of a statistically significant difference between groups (
p = 0.442).
Subgroup analyses stratified by line of therapy (second-line vs. third-line or later) yielded consistent findings, with no significant difference between cycling and swapping in either subgroup (p for interaction = 0.382). Similarly, analyses restricted to TNFi-exposed patients (the largest subgroup) showed comparable persistence between those cycling to another TNFi and those swapping to a non-TNFi agent (HR: 1.08, 95% CI: 0.79–1.48; p = 0.621). In a sensitivity analysis restricted to discontinuations due to inefficacy, the risk of treatment failure remained comparable between the CG and SG (HR: 1.18, 95% CI: 0.85–1.65; p = 0.314), consistent with the primary analysis. Similarly, in an analysis restricted to discontinuation due to adverse events the risk of treatment failure remained comparable between the CG and SG (HR: 0.86, 05% CI: 0.39–0.1.89; p = 0.7).
In the multivariable Cox regression model, the only variable significantly associated with a higher risk of treatment discontinuation was a more recent year of prescription (HR: 1.08 per year, 95% CI: 1.03–1.12;
p < 0.001) (
Figure 2). This finding corresponds to an approximately 8% increase in the hazard of discontinuation for each one-year increment in prescription date, a clinically meaningful effect. For example, compared to a treatment initiated in 2015, a treatment initiated in 2025 would have an approximately 80% higher hazard of discontinuation over the follow-up period (HR: 1.08^10 ≈ 2.16).
Similarly, the CG and SG showed comparable discontinuation risks in the propensity score analysis (HR: 1.09, 95% CI: 0.66–1.78;
p = 0.742) (
Figure 3).
The treatment strategy (cycling vs. swapping) was not a significant predictor of discontinuation (p = 0.442). Other non-significant predictors in the multivariable model included age (HR: 1.00, 95% CI: 0.99–1.01; p = 0.856), sex (HR: 0.92, 95% CI: 0.68–1.24; p = 0.582), disease duration (HR: 1.00, 95% CI: 0.99–1.01; p = 0.921), axSpA subtype (HR: 1.12, 95% CI: 0.83–1.51; p = 0.461), IBD (HR: 0.85, 95% CI: 0.57–1.27; p = 0.425), PsO (HR: 1.08, 95% CI: 0.79–1.48; p = 0.621), drug class (global p = 0.312), and line of therapy (global p = 0.284).
The strong association between prescription year and discontinuation risk persisted in multiple sensitivity analyses, including models excluding patients with less than one year of follow-up, models using calendar year as a categorical rather than continuous variable, and models adjusting for specific drug availability by year. This consistency suggests the finding is robust and unlikely to be explained by methodological artifacts.
4. Discussion
In this real-world, retrospective cohort study of axSpA patients who failed at least one line of advanced therapy, we found no significant difference in treatment persistence over three years between cycling and swapping strategies. This finding contributes to the growing body of evidence suggesting that both approaches represent valid options following treatment failure, and that factors other than the choice of switching strategy may be more important determinants of long-term outcomes.
Our results align with several large registry-based studies that have reported comparable effectiveness between cycling and swapping following TNFi failure [
4,
7,
9]. However, our study extends these observations in several important ways. First, we included failures of any advanced therapy, not solely TNFis, reflecting contemporary clinical practice where patients may be exposed to multiple MOAs over their treatment journey. Second, we included newer agents such as JAKis and IL-23is that have been underrepresented in prior comparative research. Third, we performed comprehensive multivariable adjustment for potential confounders, reducing the risk of bias inherent in unadjusted comparisons.
Historical evidence from registries such as DANBIO and NOR-DMARD established that cycling to a second TNFi after failure of a first remains a viable option, albeit with somewhat reduced response rates [
7,
9]. The Danish DANBIO registry, analyzing 432 axSpA patients cycling TNFis, reported 2-year drug survival rates of approximately 50% for second TNFi courses, comparable to the 49.3% we observed in our cycling group at 2 years [
9]. Similarly, the NOR-DMARD registry demonstrated that while response rates were numerically lower for a second TNFi compared to first, a substantial proportion of patients still achieved clinically meaningful improvement [
7].
With the advent of drugs with novel MOAs, recent observational studies have compared cycling and swapping post-TNFi failure. Data from the Korean College of Rheumatology Biologics registry [
5], the Swiss Clinical Quality Management cohort [
10], and a collaborative analysis of five Nordic registries [
4] reported comparable effectiveness between a second TNFi and swapping to an IL-17i. The Nordic study, which included over 10,000 treatment courses, found similar one-year treatment outcomes for secukinumab and TNFis in axSpA patients, regardless of prior TNFi exposure [
4]. The Swiss cohort study similarly reported comparable effectiveness of secukinumab versus an alternative TNFi in patients with prior TNFi exposure, with adjusted HRs close to 1.0 [
10].
Conversely, two mono-centric studies suggested superior retention for TNFis over IL-17is in patients with prior biologic exposure [
6,
11]. The study by van Es et al. in the Netherlands found that cycling to a second TNFi was associated with better persistence than swapping to an IL-17i in axSpA patients, although the authors acknowledged potential confounding by indication [
6]. Similarly, Yang et al. in Korea reported higher retention rates for TNFis compared to IL-17is in biologic-experienced patients [
11]. However, methodological considerations, such as patient stratification and confounding variables (e.g., number of prior biologics), complicate the interpretation of these latter findings [
6]. Our study, with its broader inclusion criteria and comprehensive adjustment, supports the equipoise between strategies.
The mechanisms underlying comparable effectiveness between cycling and swapping remain speculative. For cycling, the rationale is that patients may respond to a second agent within the same class despite failing the first, possibly due to differences in molecular structure, pharmacokinetics, or binding characteristics among agents sharing the same MOA [
12]. For swapping, the rationale is that targeting a different inflammatory pathway may overcome resistance mechanisms that limit the efficacy of the drug with the initial MOA [
13]. Our finding that both strategies yield similar outcomes suggests that neither mechanism consistently dominates in clinical practice.
The strong association between more recent prescription year and higher discontinuation risk warrants careful consideration. We hypothesize that this finding reflects evolving treatment paradigms rather than declining drug efficacy. Over the study period from 2004 to 2025, several important changes occurred in axSpA management: the introduction of treat-to-target strategies emphasizing rapid achievement of remission [
2]; the expansion of therapeutic options, reducing clinical inertia and lowering the threshold for switching; increased awareness of the importance of tight disease control; and the availability of more sensitive outcome measures to detect suboptimal responses [
12]. All these factors may have contributed to earlier discontinuation of therapies perceived as inadequately effective, resulting in shorter persistence for more recently initiated treatments.
Alternative explanations include potential changes in patient characteristics over time, such as the inclusion of patients with milder or earlier disease as diagnostic criteria evolved, or changes in prescribing patterns with preferential use of newer agents in more refractory patients [
13]. However, our multivariable model adjusted for line of therapy and disease duration, partially accounting for these factors. The persistence of the year effect after adjustment suggests that temporal trends in treatment philosophy may be the dominant explanation. Although the association persisted after adjustment for line of therapy, residual confounding related to unmeasured temporal factors cannot be fully excluded.
The comparable persistence between cycling and swapping, coupled with the absence of identified clinical predictors of differential response, supports an individualized approach to treatment selection following advanced therapy failure. In the absence of clear superiority for either strategy, clinical decision-making should be guided by patient-specific factors including extra-articular manifestations, comorbidity profile, patient preference, and drug availability [
14]. For example, patients with IBD may preferentially benefit from TNFis or IL-23is, which have proven efficacy for intestinal inflammation [
15]. Those with severe psoriasis might favor IL-17is or IL-23is given their robust efficacy for skin manifestations [
14]. Patients with recurrent infections might avoid JAKis given the class labeling regarding infection risk [
2].
A strength of this study is its inclusion of patients failing any advanced therapy line, not solely TNFis. Consequently, approximately one-third of swaps involved initiation of a TNFi after failure of a different MOA. This design reflects the reality of modern clinical practice where patients may cycle through drugs in multiple MOA classes over their disease course. By including all advanced therapies, we provide a more comprehensive picture of treatment sequencing than studies restricted to TNFi-experienced populations.
Furthermore, we included MOAs (JAKis, IL-23is) underrepresented in prior comparative studies. Although IL-23is lack formal indication for axial disease, their use in clinical practice for patients with prominent extra-articular manifestations is common, and emerging real-world data suggest potential axial benefit [
14,
15,
16,
17,
18]. Our inclusion of these agents, while limited by small numbers, provides preliminary evidence regarding their performance in the context of treatment switching and generates hypotheses for more focused research [
19,
20].
4.1. Limitations of the Study
Several limitations must be acknowledged. The observational, retrospective, single-center design introduces potential for selection bias and limits generalizability. Our findings may not be applicable to other healthcare settings with different patient populations, treatment availability, or prescribing practices. Replication in larger, multicenter cohorts and registries is needed to confirm our conclusions.
Treatment persistence is influenced by factors beyond efficacy, including tolerability. Although this limits its validity as a pure effectiveness proxy, our sensitivity analysis restricted to inefficacy-related discontinuations yielded results consistent with the primary analysis, supporting the robustness of our findings.
The absence of randomization means channeling bias cannot be excluded. Despite multivariable adjustment, residual confounding by unmeasured factors such as disease activity, functional status, quality of life, or socioeconomic factors may have influenced our results. Propensity score matching or instrumental variable analyses in larger datasets could provide additional reassurance.
To maintain analytical power, drugs were analyzed by class rather than individually. This approach assumes homogeneity within drug classes, which may not hold true. For instance, within the TNFi class, etanercept has different molecular structure and receptor-binding characteristics compared to monoclonal antibodies, potentially influencing its performance in cycling strategies [
12]. However, sample size constraints precluded agent-level analyses.
Disease activity scores (e.g., BASDAI) were not included as covariates due to high rates of missing data inherent to the retrospective design. This is an important limitation, as baseline disease activity may influence both treatment selection and outcomes. Prospective studies with standardized disease activity assessments at each treatment switch would address this gap.
Although non-drug-related discontinuations were censored, this approach is standard in drug survival analyses and minimizes bias because these reasons are unrelated to treatment performance. However, we acknowledge that censoring such cases may slightly affect generalizability, particularly in settings with high rates of nonmedical switching.
Finally, the mono-centric nature of the study may affect the external validity of the findings. Our center’s prescribing patterns, patient population, and treatment algorithms may differ from other settings. However, the consistency of our findings with registry-based studies from diverse geographic regions provides some reassurance regarding generalizability [
4,
7,
9,
10].
4.2. Further Research
The finding that more recent treatment initiation predicts discontinuation warrants further investigation into evolving therapeutic paradigms. Future research should explore whether this reflects improved treatment standards, changing patient expectations, or other factors. Longitudinal studies with detailed capture of disease activity, treatment response, and reasons for discontinuation could help elucidate the mechanisms underlying this temporal trend.
Comparative effectiveness research in this area would benefit from larger, prospective, multicenter cohorts with standardized data collection, including validated disease activity measures, patient-reported outcomes, and biomarker data. Such studies could enable more granular analyses, including agent-level comparisons, identification of predictors of differential response, and evaluation of longer-term outcomes beyond three years. Until such evidence becomes available, our findings support clinical equipoise between cycling and swapping strategies, empowering clinicians and patients to make individualized treatment decisions based on the totality of available evidence.