Next Article in Journal
ACross-Paradigm CNN–Swin Transformer Ensemble with Super-Resolution Enhancement for Multi-Class Alzheimer’s Disease Classification
Next Article in Special Issue
Decoding Visual Pathway Dysfunction with SERF-MEG: A Study in Patients with Optic Neuropathy
Previous Article in Journal
Domain-Adaptive Transfer Learning for HPV Lesion Classification in Whole Slide Images: A Patient-Level Pipeline Across the Cytology–Histology Continuum
Previous Article in Special Issue
Dual-SwinOrd: A Dual-Head Swin Transformer with Semantic Prior Injection for Ordinal Diabetic Retinopathy Grading
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Development and Internal Validation of an Explainable Machine Learning Model for Predicting Buttock Claudication After EVAR: A Dual-Center Cohort Study

1
Department of Vascular Surgery, Xuanwu Hospital, Capital Medical University, Beijing 100053, China
2
Department of Vascular Surgery, Fu Xing Hospital, Capital Medical University (FXH-CMU), Beijing 100038, China
*
Author to whom correspondence should be addressed.
Bioengineering 2026, 13(6), 665; https://doi.org/10.3390/bioengineering13060665
Submission received: 9 April 2026 / Revised: 26 May 2026 / Accepted: 3 June 2026 / Published: 8 June 2026
(This article belongs to the Special Issue AI-Driven Approaches to Diseases Detection and Diagnosis)

Abstract

Buttock claudication after endovascular aneurysm repair (EVAR) impairs recovery and quality of life, yet individualized preoperative risk tools are scarce. We conducted a retrospective dual-center cohort study of consecutive EVAR patients from Fuxing and Xuanwu Hospitals. The endpoint was new-onset postoperative buttock claudication. Missingness was quantified for each predictor and handled using complete-case analysis or model-based single imputation according to the extent of missingness. Data were split into training and held-out test sets at a 70:30 ratio with outcome stratification. Predictor screening, preprocessing, and hyperparameter tuning were performed within the training/resampling framework to minimize data leakage. Ten algorithms were tuned using stratified 10-fold cross-validation, and test set performance was assessed using discrimination, threshold-based metrics, calibration plots, calibration intercept/slope, Brier score, and decision-curve analysis. SHapley Additive exPlanations (SHAP) provided model-agnostic explanations. A web calculator was deployed. Among 272 patients, 71 (26.1%) developed claudication. Independent risk factors included aneurysm with iliac involvement (adjusted OR 4.04), male sex (3.26), unilateral (3.86) and bilateral internal iliac artery embolization (8.61), and hyperlipidemia (5.66); >2 distal internal iliac branches was protective (0.15). On the test set, the neural network achieved the highest AUROC (test ROC), with the highest sensitivity (0.810) and top F1 (0.557) at balanced specificity (0.617); CatBoost maximized accuracy (0.790) and specificity (0.900). Calibration was acceptable, and DCA showed positive net benefit across clinically plausible thresholds. SHAP confirmed physiologic directions and enabled case-level interpretation. An explainable machine learning framework accurately stratifies risk of buttock claudication after EVAR, highlighting the roles of internal iliac embolization, iliac involvement, and distal branch anatomy. The publicly available Shiny tool supports perfusion-aware planning and shared decision-making.

1. Introduction

Endovascular aneurysm repair (EVAR) has become the default approach for infrarenal abdominal and iliac aneurysms; however, pelvic ischemic symptoms—most notably buttock claudication—remain a meaningful complication [1,2]. These symptoms can reduce walking capacity, delay rehabilitation, and impair postoperative quality of life. Strategies used to achieve an adequate seal and appropriate limb alignment (intentional internal iliac artery [IIA] embolization being one prime example) may reduce pelvic inflow, whereas the adequacy of collateral circulation varies considerably across patients and is difficult to assess prospectively. Routine practice in making decisions regarding preservation of one or both IIAs, staging of such procedures, selective embolization, or even small adjustments to landing zones, relies heavily on anatomical feasibility and clinician judgment [3,4]. The incidence of buttock claudication itself is non-trivial; patient specific vulnerability combines both anatomic factors and various forms of systemic vascular burden. A robust preoperative risk stratification tool is therefore required; planning, IIA-preservation strategies, patient counseling, and all that follows from it, must be at least partially guided by such a tool.
Buttock claudication after EVAR is fundamentally a problem of pelvic hypoperfusion. The IIA supplies the superior/inferior gluteal and obturator systems; intentionally embolizing or covering one or both IIAs to secure an endograft seal forces the gluteal bed to rely on various collateral routes [5,6,7,8]. Under exertion these circuits can become insufficient, creating the well-described gradient of risk: no IIA interruption → unilateral → bilateral loss of inflow. Anatomic complexity further modulates this risk; iliac-involved aneurysms specifically often require longer landing zones, adjunctive embolization, or limb extensions that directly encroach upon pelvic inflow. The actual capacity for collateralization likely depends on distal IIA arborization—number and/or caliber of those branches representing some form of native reserve. Systemic vascular milieu also plays a part: peripheral arterial disease, dyslipidemia and all associated cumulative exposures impair both macro and microvascular function, decreasing all aspects of vasodilatory reserve; sex-related differences in muscle mass and subsequent oxygen demand during ambulation are another potential contributor [9,10]. General attributes like age, BMI, or simple aortic diameter show inconsistent associations when pelvic inflow/outflow is taken into account. Collectively, anatomic and clinical determinants of risk interact in near-constant, individualized ways. Risk stratification should and must be, to some degree, mechanism-based.
Despite decades of EVAR experience, most studies on buttock claudication remain single-center, small, and heterogeneous in design [11,12]. Variable outcome definitions and follow-up, limited adjustment for confounding, the near absence of calibration and formal clinical-utility assessment remain common limitations. Many models emphasize AUROC alone; post hoc threshold selection, risk train–test leakage, and class imbalance are commonly unaddressed problems. Anatomic details that likely matter—the number of distal internal iliac branches being one example, structured collateral scores another—are often missing or inconsistently coded; systemic vascular burden is coarsely captured. Where machine learning has been attempted, black-box models are the norm; transparent, patient-level explanations are uncommon, external validation and all subsequent ‘deployable’ tools are rare, and integration into clinical workflows has not been demonstrated. Clinicians still lack a rigorous, interpretable and usable preoperative risk tool to personalize IIA preservation, selective/staged embolization, and to some extent postoperative planning.
Risk for buttock claudication emerges from nonlinear interactions among anatomic (embolization status, iliac involvement, branch count) and clinical factors. Traditional regression captures average effects but misses higher-order patterns and all associated thresholds (bilateral vs. single, for instance, and their synergistic effects). A multi-algorithm ML strategy is therefore appropriate to model these complex boundaries; prespecified safeguards are required—screening (p < 0.05), at least a 70/30 train-test split, some form of cross-validation, locked thresholds, and comprehensive evaluation (AUROC, AUPRC, calibration, decision curves). Avoiding the ‘black-box’ adoption barriers, we embed model-agnostic explainability via SHapley Additive exPlanations (SHAP) to support all of the above-mentioned clinical decisions. Individualized and, to the end-user at least, partially transparent, risk stratification is a realizable goal in EVAR.
We aimed to build an individualized, clinically usable risk model for postoperative buttock claudication following EVAR using a dual-center cohort. Specifically, we prespecified feature screening within the training set at p < 0.05, then developed and benchmarked ten supervised algorithms under a 70/30 train–test split with stratified cross-validation and locked thresholds prior to formal test evaluation. Comprehensive performance was reported—AUROC, accuracy, various forms of sensitivity/specificity, precision, F1, calibration and decision-curve analysis; confusion matrices were used to visualize more subtle error structures. Global variable importance across models and both global and case-level SHAP explanations for the ‘top performing’ test set model were provided to ensure at least some degree of physiological coherence and patient-level transparency [13]. The present study was designed to develop and internally validate an explainable preoperative prediction model for postoperative buttock claudication after EVAR. Its main contribution is the integration of clinically relevant pelvic perfusion-related variables, including internal iliac artery embolization status, iliac artery involvement, and distal internal iliac branch anatomy, into a transparent risk-prediction workflow. The model was benchmarked against logistic regression and other machine learning algorithms, interpreted using SHAP, and implemented as a research-use Shiny prototype for individualized risk estimation.

2. Materials and Methods

2.1. Cohort and Baseline Characteristics

This retrospective, dual-center cohort study included consecutive adults who underwent EVAR at Fuxing Hospital, Capital Medical University, and Xuanwu Hospital, Capital Medical University. The time frame of inclusion was from 1 January 2017 to 1 July 2025. The protocol was approved by the Ethics Committee of Fuxing Hospital (approval No. 2025FXHEC-KSP054); waiver of consent was obtained. Prespecified eligibility criteria were used: inclusion required age ≥ 18 years, some form of index EVAR for infrarenal abdominal aortic aneurysm (AAA), iliac involvement of the AAA, or isolated iliac artery aneurysm; baseline perioperative data and subsequent follow-up sufficient to ascertain the desired endpoint were also required. Exclusions were open aortic repair, any thoracic/thoracoabdominal endovascular procedure without an EVAR component; prior EVAR (only the first was analyzed), repeat/revision EVAR during the study window; documented pre-existing buttock claudication or other clearly confounding alternative conditions (severe degenerative hip/knee disease being a prime example, neurogenic being another); perioperative death prior to outcome ascertainment; or missing primary outcome data or unavailable key anatomical or procedural predictor information required for model development. Missingness was assessed for each candidate predictor before model development. No missing values were observed in the variables used for final model development; therefore, imputation was not performed in the final analysis. Candidate predictors were extracted from electronic records and operative notes: Site_of_aneurysm, Gender, Peripheral_arterial_disease, Number_of_internal_iliac_arteries_embolized, Chronic_obstructive_pulmonary_disease, Chronic_kidney_disease, Antiplatelet, Hyperlipidemia, History_of_previous_abdominal_and_pelvic_surgery, Cardiovascular_disease, Marital_status, Smoking, Drinking, Cerebrovascular_disease, Hypertension, Diabetes, Number_of_distal_internal_iliac_artery_branches, Age, BMI, and Maximum_diameter_of_abdominal_aorta. The primary endpoint was new-onset buttock claudication after EVAR—exertional gluteal pain relieved by rest—adjudicated from standardized ward/clinic documentation within a prespecified postoperative window. Categorical variables were ordinal-encoded, and continuous variables were inspected for distributional characteristics and outliers. Missingness was first quantified for each candidate predictor before model development. Variables with low-level missingness were handled using complete-case analysis for the corresponding analyses, whereas variables with non-negligible missingness were handled using model-based single imputation based only on information available within the training data. To avoid information leakage, imputation, scaling, encoding, and any other preprocessing steps were incorporated into the resampling pipeline and were not estimated using the held-out test set.

2.2. Feature Screening in the Training Set

After eligibility, the data were randomly split into training (70%) and test (30%) sets; stratification by the endpoint and a fixed random seed was used to ensure adequate representation. Feature screening was performed exclusively within the training set. Univariable logistic regressions were fitted for each candidate predictor; variables meeting the prespecified screening threshold (p < 0.05) were entered into a subsequent multivariable model. A full multivariable logistic regression was then fitted, and multicollinearity was assessed using variance inflation factors(VIF < 5 being acceptable), and all adjusted odds ratios (ORs) with their corresponding 95% confidence intervals were reported. The specific variables that remained in the final model comprised the feature set for any downstream machine learning developments.

2.3. Model Development and Internal Validation (Training Set)

Ten supervised algorithms were trained on the multivariable-selected features: AdaBoost, CatBoost, K-nearest neighbors (KNNs), LightGBM, Logistic regression, neural network, Random Forest, Support Vector Machine (SVM), XGBoost, and Gradient Boosting Machine (GBM). Pipelines were tailored per model (e.g., standardization for KNN/SVM/neural network; no scaling for tree/boosting models), with class weighting where applicable. Hyperparameters were tuned using stratified 10-fold cross-validation within the training set, with the AUROC as the primary optimization metric and the AUPRC and F1 score used as secondary considerations. Calibration was evaluated using calibration plots, calibration intercept, calibration slope, and the Brier score. If substantial miscalibration was observed during internal validation, probability recalibration was performed within the same cross-validation framework. Classification thresholds were selected using training-set predictions, including Youden’s J statistic and F1-maximizing thresholds, and were locked before application to the held-out test set. Feature selection, preprocessing, hyperparameter tuning, calibration, and threshold selection were performed exclusively within the training/resampling framework, and the held-out test set was not used for any model fitting, preprocessing, tuning, or threshold selection.

2.4. Independent Test Set Evaluation

All tuned models were frozen and applied to the independent 30% test set without any further fitting or threshold adjustment. The same discrimination, calibration, and classification metrics were computed; ROC, calibration, and DCA plots were generated accordingly. Confusion matrices were created to visualize more specific false-positive/false-negative patterns relative to these locked thresholds. Quantifying uncertainty in the discrimination aspect was of particular interest; bootstrap 95% CIs for AUROC were obtained (1000 resamples, stratified by outcome). The algorithm achieving the highest test set AUROC—with both F1 and various forms of calibration as secondary considerations—was designated the ‘primary’ model; all subsequent interpretation and potential deployment followed from this choice.

2.5. Explainability and Clinical Deployment

We summarized global variable importance for each algorithm using model-appropriate approaches. The best-performing model on the test set—the neural network—was given more specific attention; SHAP values were computed to provide both global and local interpretability. Outputs included a global bar plot and associated beeswarm distribution to rank/illustrate feature contributions, various forms of dependence plots for key predictors, and more representative force and waterfall plots to decompose at least some individual predictions. The locked primary model was then implemented in a bedside Shiny web application: study predictors are entered as inputs, individualized predicted risk follows, along with an explained panel, all to some degree rooted in the previous SHAP work. All analyses were performed in R 4.5.0; version-controlled scripts were used throughout to maintain all aspects of reproducibility.

3. Results

3.1. Cohort and Baseline Characteristics

We included 272 EVAR patients; 71 (26.1%) developed postoperative buttock claudication and the remaining 201 did not. Variables are summarized in Table 1. Patients with buttock claudication more often had an abdominal aortic aneurysm with associated iliac artery aneurysm (47.9% vs. 26.4%; p = 0.001) and some form of peripheral arterial disease (54.9% vs. 34.8%; p = 0.005). Internal iliac artery embolization, whether unilateral or bilateral, was more frequent among patients with claudication (overall p < 0.001). Hyperlipidemia was also significantly more prevalent (87.3% vs. 66.2%), while a history of previous abdominal/pelvic surgery was less common (29.6% vs. 44.3%). A reduced number of distal internal iliac artery branches (≤2) was markedly enriched in the cases. Age (65.0 ± 8.3 vs. 65.7 ± 7.6 years), BMI (29.8 ± 4.5 kg/m2) and maximum abdominal aortic diameter (5.7 ± 0.5 cm) showed no significant difference between the two groups; all other comorbidities and near-identical lifestyle factors were balanced. The cohort was then randomly split 70%/30% into training and test sets to allow for both model development and subsequent independent evaluation.

3.2. Feature Screening in the Training Set

In the training set (n = 191; buttock claudication 50/191, 26.2%), both univariable and multivariable logistic regression identified several independent risk factors for postoperative buttock claudication (Table 2). On multivariable analysis, abdominal aortic aneurysm with associated iliac artery aneurysm (adjusted OR 4.04, 95% CI 1.73–9.40, p = 0.001), male sex (adjusted OR 3.26, 95% CI 1.20–8.86, p = 0.021), unilateral internal iliac artery embolization (adjusted OR 3.86, 95% CI 1.38–10.80, p = 0.010), bilateral internal iliac artery embolization (adjusted OR 8.61, 95% CI 1.78–41.68, p = 0.007), and hyperlipidemia (adjusted OR 5.66, 95% CI 1.84–17.37, p = 0.002) were independently associated with higher odds of postoperative buttock claudication, whereas having >2 distal internal iliac artery branches was the only independent protective factor (adjusted OR 0.15, 95% CI 0.05–0.41, p < 0.001). Other comorbidities and all continuous variables (age, BMI, and maximum abdominal aortic diameter) were not independently associated after adjustment. Variables retained in the final multivariable model were then used as inputs for subsequent machine learning development.

3.3. Model Development and Internal Validation (Training Set)

Using the multivariable-selected predictors, we trained ten algorithms on the 70% training split and reported performance at model-specific probability thresholds (Table 3). Operating points showed complementary trade-offs: LightGBM achieved the highest F1 (0.656), with a near-perfect balance of sensitivity and specificity; KNN and SVM followed closely thereafter. Random Forest yielded the highest accuracy (0.822), driven by extremely high specificity but low sensitivity, all other metrics being compromised; GBM prioritized some form of recall (sensitivity was ~0.88) at the expense of the previously mentioned specificity. Logistic regression and the neural network produced identical summaries at their chosen thresholds. Full thresholds and point estimates for every model are provided in Table 3.
Discrimination, calibration, and to a large extent clinical utility on the training data are visualized in Figure 1A–C (ROC, calibration, DCA); qualitatively these mirror the trade-offs previously described. Error profiles are shown as confusion matrices for each method (Figure 2, ordered by model): AdaBoost, CatBoost, KNN, LightGBM, Logistic, neural network, Random Forest, SVM, XGBoost, GBM. Focusing on the more ‘pure’ specificity or sensitivity favored models, and all intermediate cases, the various models show a range of clinically useful behaviors.

3.4. Independent Test Set Evaluation

On the held-out test set, overall performance varied across algorithms at their model-specific operating thresholds (Table 4). Accuracy spanned from 0.642 to 0.790; CatBoost achieved the highest accuracy (the latter being accompanied by near-perfect specificity, 0.900) though a more moderate F1 of 0.541. Random Forest maximized some form of specificity (0.917) and precision, sensitivity and consequently F1 being severely compromised (0.238/0.323). Logistic regression and the neural network delivered the highest sensitivity (both at 0.810) with top tier F1 scores; balanced specificity was also obtained for them (approximately 0.617). SVM, GBM, and to a large extent XGBoost showed very similar balanced profiles (accuracy, sensitivity and F1 all around 0.704–0.714), KNN being the outlier with lower F1. Discrimination, calibration and various clinical utility curves for the test set are detailed in Figure 1D–F (ROC, calibration, DCA); the neural network was selected as the primary model for downstream explainability based on these metrics. Per-model error patterns are visualized through the test set confusion matrices (Figure 3, order as above), where specificity-oriented models (Random Forest being a prime example) have, and should have, fewer false positives; sensitivity or ‘all round’ models are of secondary clinical importance.

3.5. Explainability and Clinical Deployment

Across various algorithms, the global variable importance profiles (Figure 4) consistently identified internal iliac artery embolization status (bilateral/unilateral), the number of distal internal iliac branches, and aneurysm site (specifically abdominal aortic with some form of iliac involvement) as the most influential predictors. Contributions were directionally concordant with multivariable logistic results: embolization and associated iliac involvement being risk-increasing, and having >2 distal branches risk-decreasing. Focusing on the top-performing neural network, SHAP analyses (Figure 5A–E) provided both global and near individual-level transparency; the bar plot and subsequent beeswarm confirmed the features mentioned above as dominant drivers; dependence plots illustrated stepwise (categorical) risk shifts—none → unilateral → bilateral being a prime example—and the previously noted monotonic decrease with increasing branch count; representative force/waterfall plots decomposed specific patient predictions, embolization, iliac factors, male sex and to a large extent hyperlipidemia pushing the prediction ‘up’ while all or part of the distal branch effect pulls it down. These effect directions were and remain clinically plausible. A bedside usable interactive web calculator was developed; Figure 6 shows the interface, currently available at https://cacsriskmodel.shinyapps.io/makee/, accessed on 7 October 2025. Individual risk estimates can be generated from any set of patient features, with all or part of the model explanation provided to the clinician.

4. Discussion

In this dual-center retrospective cohort study (2017–2025), postoperative buttock claudication occurred in 26.1% (71/272) following EVAR. In the training set, multivariable logistic regression identified more than two distal internal iliac artery branches as an independent protective factor for postoperative buttock claudication (aOR, 0.15; 95% CI, 0.05–0.41), whereas peripheral arterial disease showed only a borderline association. Across ten machine learning algorithms, training/internal validation revealed complementary trade-offs (e.g., LightGBM highest F1 = 0.656; Random Forest accuracy = 0.822 but poor sensitivity; GBM sensitivity most closely balanced). The held-out test set then permitted optimal performance: the neural network (combined with simple logistic regression) achieved the best sensitivity (0.810) and near-top F1, CatBoost followed in accuracy and all round ‘net benefit’. Calibration was acceptable and decision-curve analysis provided clinically meaningful net benefit at various plausible thresholds. Model-agnostic explainability using SHAP for the neural network corroborated all of the above: embolization status and/or iliac involvement push risk upward, richer distal branch networks pull it down, and to the clinician at least, risk can be transparently and practically managed.
The risk architecture we observed is coherent with pelvic perfusion biology. Internal iliac artery (IIA) embolization compromises direct inflow to the superior and inferior gluteal arteries; thus, risk increases in a well-defined stepwise manner—none → unilateral → bilateral occlusion—precisely replicated by the SHAP dependence plots. Sacrificing both IIAs forces the gluteal musculature to rely on more circuitous collateral routes (profunda femoris–descending branch–gluteal anastomoses, obturator and lumbar being prime examples); these are frequently inadequate during exertion, leading to exertional gluteal ischemia [7,11]. A greater number of distal IIA branches (>2) likely represents a richer native collateral network; this reserve reduces the hemodynamic drop across the pelvis, and the strong protective association seen along with negative SHAP contributions is a direct reflection of it.
Aneurysm with iliac involvement plausibly signals more extensive aorto-iliac disease and all associated manipulation (landing zones, device coverage, adjunctive embolization), increasing the odds of such pelvic hypoperfusion [4,7]. Male sex may capture higher absolute muscular oxygen demand and various sex-linked vascular risk clustering, pushing the ischemic threshold higher during ambulation [14,15]. Hyperlipidemia aligns with microvascular dysfunction and impaired endothelial nitric-oxide pathway vasodilation, worsening any pre-existing supply–demand mismatch in the gluteal bed [16,17]. The borderline PAD signal is directionally consistent with a systemic atherosclerotic milieu—limited runoff/collateralization being part of that process [18,19]. Variables not retained as independent predictors (age, BMI, maximum aortic diameter, etc.) will always have or at least partially have weaker effects on the final clinical outcome of buttock claudication. Preserving at least one IIA, or any form of anatomically based preoperative planning, is a real and present clinical need.
Our findings align with—and extend—the existing EVAR literature on pelvic ischemia following internal iliac artery (IIA) interruption. Prior observational cohorts consistently report higher rates of buttock claudication associated with bilateral compared to unilateral IIA embolization; the protective role of some form of preserved pelvic inflow is frequently emphasized; our data reproduce this stepwise risk gradient, with a more richly branched distal IIA network serving as the specific mitigating factor [20,21]. The association between iliac-involved aneurysms and this claudication risk is concordant with reports of more complex aorto-iliac anatomy, longer ‘landing’ zones, and various forms of adjunctive embolization all increasing the potential for pelvic hypoperfusion [22]. Signals for a systemic vascular burden—hyperlipidemia being a prime example, borderline PAD if measured—were directionally consistent with work previously linking endothelial dysfunction and to a large extent limited collateral capacity to at least partially exercise-induced ischemia [23,24]. Variables such as age, BMI, and maximum aortic diameter have shown more heterogeneous effects; in our analysis they were all attenuated when true pelvic inflow/outflow anatomy and all related comorbidities were taken into account. Primary risk, it seems, is best defined by the anatomical or functional state of the pelvis itself.
Methodologically, our study adds several elements that are less commonly found in earlier reports The main contribution of the present study is not the discovery of entirely new risk factors, but the integration of clinically expected pelvic perfusion-related variables into an explainable prediction workflow and a research-use Shiny prototype for individualized risk estimation. First, we benchmarked ten supervised learning algorithms against multivariable logistic regression; complementary operating characteristics were demonstrated, with a specific neural network identified as the ‘best’ test set discriminator under locked thresholds [25,26]. Second, beyond simple discrimination, calibration and decision-curve analysis were emphasized; a clinically interpretable view of net benefit across all possible thresholds is what we sought to report—this very specific type of assessment being near constant in its absence from the literature [27,28]. Third, model-agnostic explainability was supplied; mechanistic plausibility (embolization status, iliac involvement increasing risk; more distal branches being protective) could be at least partially reconciled with the more abstract, individual level prediction logic of the model. Prior black-box applications lacked this form of explainability [29]. Finally, operationalizing the model into a Shiny tool bridges the near-perpetual gap between methodological promise and some form of bedside utility; preoperative planning and all subsequent patient counseling, based on transparent risk estimates, are both improved by it [30].
Across algorithms, operating characteristics were complementary rather than uniformly superior. On internal validation, LightGBM maximized F1 (0.656) with a near-perfect balance of sensitivity and specificity, Random Forest achieved the highest accuracy (0.822) by pushing towards some form of specificity (low false positives), and GBM served as the recall prioritizer (sensitivity 0.880). The held-out test set confirmed these trade-offs: CatBoost had the best accuracy and near-maximal specificity, but a moderate F1; Random Forest itself pushed specificity even higher, sensitivity dropping considerably (0.238, F1~0.323). Neural network delivered the highest sensitivity, the top F1 being attained under this screening-oriented objective. Calibration analyses were generally acceptable, decision-curve analysis showing positive net benefit for at least some clinically plausible thresholds; real world utility beyond simple discrimination was supported. Model selection followed a locked-threshold protocol to prevent test leakage; AUROC, F1 and calibration all factored into the final decision, with the neural network as the primary model chosen. Practically two operating regimes are defensible: preoperative screening/counseling type sensitivity, and more ‘specific’ follow-up based on resource or intervention risk. Bootstrap 95% CIs for AUROC (1000 resamples) provided the last necessary uncertainty bound. Finally, while the neural network led in all aspects of performance, we believe a fully interpretable model is still the end goal of clinical modeling.
Model-agnostic SHAP analyses bridged the gap between statistical performance and bedside decision-making [31,32]. Globally, both bar and beeswarm plots consistently identified IIA embolization status, iliac involvement, and some measure of the distal IIA branch count as the dominant drivers of risk; directions of effect were concordant with known physiology and multivariable odds ratios. Dependence plots captured more clinically intuitive gradients (none → unilateral → bilateral embolization increasing risk, >2 distal branches decreasing it if at all), case-level force and waterfall plots then decomposed individual predictions into near-actionable contributors. These various forms of explanation support shared decision-making: patients flagged as high risk can have specific recommendations made to them. Clinicians may (i) prioritize IIA preservation (anatomically if possible), (ii) move towards staged or selective embolization, branched/parallel iliac strategies, or any form of alternative ‘landing’ plan, and (iii) tailor postoperative surveillance and to a point rehabilitation. Operationally, our Shiny tool packages this entire workflow—routine preoperative variables are accepted, a personalization of risk is returned with a concise explanation panel, and all or part of the decision-making process can be guided by either screening or confirmation style use cases.
Strengths include a dual-center cohort spanning eight and a half years, prespecified training-set screening (p < 0.05), evaluation of ten algorithms all under a locked-threshold protocol, comprehensive assessment beyond simple discrimination—calibration and decision-curve analysis—SHAP transparency, and a deployable web calculator. Several limitations warrant some form of caution. The retrospective design risks residual confounding and various types of outcome misclassification; standardized definitions help but do not fully eliminate this [33,34]. Center-specific devices, embolization techniques, and near-perfectly controlled perioperative pathways limit true generalizability. Because both centers shared similar institutional practice patterns, including patient selection, EVAR technique, embolization strategy, device use, and documentation habits, the model may not generalize to hospitals with different procedural approaches or patient populations. Some potentially informative imaging features (quantitative pelvic collateral scores, IIA diameter/flow type metrics) were not collected, symptom severity and/or duration were not modeled. The sample size and number of outcome events were limited, particularly in the held-out test set; therefore, comparisons among ten machine learning algorithms should be interpreted as exploratory, because a small number of reclassified patients could materially affect sensitivity, specificity, F1 score, and AUROC. Although repeated cross-validation and bootstrap AUROC confidence intervals were used to reduce uncertainty, larger cohorts and external validation are required to assess model stability and generalizability. Finally, DCA assumes constant threshold utilities across settings, cost-effectiveness being the ultimate clinical endpoint we would like to see analyzed.
Clinically, the model encourages preoperative perfusion-aware planning: preserve at least one IIA whenever possible, appraise distal branch anatomy proactively, and reserve bilateral embolization for carefully selected scenarios with some form of mitigation strategy in place. The tool can inform counseling regarding expected risk, guide the intensity of follow-up, and flag patients for early rehabilitation or more specific perfusion assessments. However, the model should be interpreted as a risk-prediction tool rather than evidence that modifying any individual predictor or changing a procedural strategy will necessarily reduce postoperative buttock claudication. Research priorities include (i) prospective multicenter external validation; (ii) enriching the predictors with imaging-derived pelvic collateral metrics and various device/technique details; (iii) extending beyond simple binary outcomes to time-to-event and severity-based outcomes; (iv) explicit clinical or health-economic utilities to define ‘optimal’ operating points; and (v) impact trials testing whether model-guided preservation of IIA’s reduces buttock claudication, with all other ischemic complications being the secondary endpoint. Integration into the electronic health record with automated data ingestion and real-time, interpretable SHAP style explanations is the practical next step in achieving sustainable clinical adoption.

5. Conclusions

In this dual-center retrospective cohort study, we developed a preliminary explainable prediction model for postoperative buttock claudication after EVAR. The model showed promising internal performance and identified clinically plausible pelvic perfusion-related predictors, including internal iliac artery embolization, iliac involvement, and distal internal iliac branch anatomy. However, because the study was based on retrospective data and internal validation, external validation, prospective testing, and clinical impact evaluation are required before the model or Shiny calculator can be used for routine clinical decision-making.

Author Contributions

All authors (Y.L., H.D. and Y.G.) made a significant contribution to the work reported, whether in the conception, study design, execution, acquisition of data, analysis and interpretation, or in all these areas; took part in drafting, revising or critically reviewing the article; gave final approval of the version to be published; agreed on the journal to which the article has been submitted; and agreed to be accountable for all aspects of the work. All authors have read and agreed to the published version of the manuscript.

Funding

This study was supported by the National Key Research and Development Program of China [2021YFC2500500].

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Review Board of Fuxing Hospital, Capital Medical University (protocol code 2025FXHEC-KSP054, 6 August 2025).

Informed Consent Statement

The requirement for informed consent was waived by the Institutional Review Board of Fuxing Hospital, Capital Medical University, because of the retrospective nature of the study.

Data Availability Statement

The data analyzed and the codes used during the current study are available from the corresponding author on reasonable request.

Acknowledgments

The authors would like to express their sincere gratitude to the editors and anonymous reviewers for their valuable comments and constructive suggestions, which greatly helped improve the quality and clarity of this manuscript. The authors also thank all individuals who provided support and assistance during the preparation and revision of this work.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AAAabdominal aortic aneurysm
BMIbody mass index
CIconfidence interval
DCAdecision-curve analysis
EVARendovascular aneurysm repair
MLmachine learning
ROCreceiver operating characteristic
ORodds ratio
SHAPSHapley Additive exPlanations
PADperipheral arterial disease

References

  1. Fatima, J.; Correa, M.P.; Mendes, B.C.; Oderich, G.S. Pelvic revascularization during endovascular aortic aneurysm repair. Perspect. Vasc. Surg. Endovasc. Ther. 2012, 24, 55–62. [Google Scholar] [CrossRef]
  2. Kudo, T. Surgical Complications after Open Abdominal Aortic Aneurysm Repair: Intestinal Ischemia, Buttock Claudication and Sexual Dysfunction. Ann. Vasc. Dis. 2019, 12, 157–162. [Google Scholar] [CrossRef]
  3. Ji, J.; Bi, J.; Chen, Y.; Zhang, X.; Zhao, B.; Liang, H.; Fan, J.; Dai, X. Mid-term outcomes of different treatments of internal iliac artery in endovascular aneurysm repair. Sci. Prog. 2024, 107, 368504241274998. [Google Scholar] [CrossRef]
  4. Xu, H.; Fan, H.; Li, Y.; Zhang, H.; Zhao, Z.; Li, L.; Liu, M.; Liu, J.; Guo, M. The safety and efficacy of complete preservation of internal iliac arteries in patients with aortoiliac aneurysms using iliac branch stent grafts: A multicenter retrospective comparative study. Sci. Prog. 2025, 108, 368504251353580. [Google Scholar] [CrossRef]
  5. Hayashi, N.; Yunoki, J.; Tanaka, A.; Shigetomi, K.; Baba, K.; Shichijo, M.; Jinnouchi, K.; Morokuma, H.; Itoh, M.; Kamohara, K. Efficacy of the Preloading Coil-In-Plug Method for Internal Iliac Artery Embolization. Ann. Vasc. Surg. 2025, 122, 627–634. [Google Scholar] [CrossRef] [PubMed]
  6. Kansal, V.; Jetty, P.; Kubelik, D.; Hajjar, G.; Hill, A.; Brandys, T.; Nagpal, S. Internal iliac coverage during endovascular repair of abdominal aortic aneurysms is a safe option: A preliminary study. Vascular 2017, 25, 28–35. [Google Scholar] [CrossRef] [PubMed]
  7. Perini, P.; Mariani, E.; Fanelli, M.; Ucci, A.; Rossi, G.; Massoni, C.B.; Freyrie, A. Surgical and Endovascular Management of Isolated Internal Iliac Artery Aneurysms: A Systematic Review and Meta-Analysis. Vasc. Endovasc. Surg. 2021, 55, 254–264. [Google Scholar] [CrossRef]
  8. Suzuki, S.; Akamatsu, D.; Goto, H.; Kakihana, T.; Sugawara, H.; Tsuchida, K.; Yoshida, Y.; Umetsu, M.; Kamei, T.; Unno, M. Prospective clinical study for claudication after endovascular aneurysm repair involving hypogastric artery embolization. Surg. Today 2022, 52, 1645–1652. [Google Scholar] [CrossRef] [PubMed]
  9. Davel, A.P.; Lu, Q.; Moss, M.E.; Rao, S.; Anwar, I.J.; DuPont, J.J.; Jaffe, I.Z. Sex-Specific Mechanisms of Resistance Vessel Endothelial Dysfunction Induced by Cardiometabolic Risk Factors. J. Am. Heart Assoc. 2018, 7, e007675. [Google Scholar] [CrossRef]
  10. Pabon, M.; Cheng, S.; Altin, S.E.; Sethi, S.S.; Nelson, M.D.; Moreau, K.L.; Hamburg, N.; Hess, C.N. Sex Differences in Peripheral Artery Disease. Circ. Res. 2022, 130, 496–511. [Google Scholar] [CrossRef]
  11. Kontopodis, N.; Tavlas, E.; Papadopoulos, G.; Galanakis, N.; Tsetis, D.; Ioannou, C.V. Embolization or Simple Coverage to Exclude the Internal Iliac Artery During Endovascular Repair of Aortoiliac Aneurysms? Systematic Review and Meta-analysis of Comparative Studies. J. Endovasc. Ther. 2017, 24, 47–56. [Google Scholar] [CrossRef]
  12. ‘t Mannetje, Y.W.; Broos, P.; Teijink, J.A.W.; Stokmans, R.A.; Cuypers, P.W.M.; van Sambeek, M. Midterm Results After Abandoning Routine Preemptive Coil Embolization of the Internal Iliac Artery During Endovascular Aneurysm Repair. J. Endovasc. Ther. 2019, 26, 238–244. [Google Scholar] [CrossRef]
  13. Liu, C.; Zhang, K.; Yang, X.; Meng, B.; Lou, J.; Liu, Y.; Cao, J.; Liu, K.; Mi, W.; Li, H. Development and Validation of an Explainable Machine Learning Model for Predicting Myocardial Injury After Noncardiac Surgery in Two Centers in China: Retrospective Study. JMIR Aging 2024, 7, e54872. [Google Scholar] [CrossRef]
  14. Ansdell, P.; Thomas, K.; Hicks, K.M.; Hunter, S.K.; Howatson, G.; Goodall, S. Physiological sex differences affect the integrative response to exercise: Acute and chronic implications. Exp. Physiol. 2020, 105, 2007–2021. [Google Scholar] [CrossRef] [PubMed]
  15. Martins, H.A.; Barbosa, J.G.; Seffrin, A.; Vivan, L.; Souza, V.; De Lira, C.A.B.; Weiss, K.; Knechtle, B.; Andrade, M.S. Sex Differences in Maximal Oxygen Uptake Adjusted for Skeletal Muscle Mass in Amateur Endurance Athletes: A Cross Sectional Study. Healthcare 2023, 11, 1502. [Google Scholar] [CrossRef] [PubMed]
  16. Gliozzi, M.; Scicchitano, M.; Bosco, F.; Musolino, V.; Carresi, C.; Scarano, F.; Maiuolo, J.; Nucera, S.; Maretta, A.; Paone, S.; et al. Modulation of Nitric Oxide Synthases by Oxidized LDLs: Role in Vascular Inflammation and Atherosclerosis Development. Int. J. Mol. Sci. 2019, 20, 3294. [Google Scholar] [CrossRef]
  17. Sorop, O.; van den Heuvel, M.; van Ditzhuijzen, N.S.; de Beer, V.J.; Heinonen, I.; van Duin, R.W.; Zhou, Z.; Koopmans, S.J.; Merkus, D.; van der Giessen, W.J.; et al. Coronary microvascular dysfunction after long-term diabetes and hypercholesterolemia. Am. J. Physiol. Heart Circ. Physiol. 2016, 311, H1339–H1351. [Google Scholar] [CrossRef]
  18. Annex, B.H.; Cooke, J.P. New Directions in Therapeutic Angiogenesis and Arteriogenesis in Peripheral Arterial Disease. Circ. Res. 2021, 128, 1944–1957. [Google Scholar] [CrossRef]
  19. Forsythe, R.O.; Brownrigg, J.; Hinchliffe, R.J. Peripheral arterial disease and revascularization of the diabetic foot. Diabetes Obes. Metab. 2015, 17, 435–444. [Google Scholar] [CrossRef] [PubMed]
  20. Kobayashi, Y.; Sakaki, M.; Yasuoka, T.; Iida, O.; Dohi, T.; Uematsu, M. Endovascular repair with contralateral external-to-internal iliac artery bypass grafting: A case series. BMC Res. Notes 2015, 8, 183. [Google Scholar] [CrossRef]
  21. Melas, N.; Saratzis, A.; Dixon, H.; Saratzis, N.; Lazaridis, J.; Perdikides, T.; Kiskinis, D. Isolated common iliac artery aneurysms: A revised classification to assist endovascular repair. J. Endovasc. Ther. 2011, 18, 697–715. [Google Scholar] [CrossRef] [PubMed]
  22. Qrareya, M.; Zuhaili, B. Management of Postoperative Complications Following Endovascular Aortic Aneurysm Repair. Surg. Clin. N. Am. 2021, 101, 785–798. [Google Scholar] [CrossRef]
  23. Olin, J.W.; White, C.J.; Armstrong, E.J.; Kadian-Dodov, D.; Hiatt, W.R. Peripheral Artery Disease: Evolving Role of Exercise, Medical Therapy, and Endovascular Options. J. Am. Coll. Cardiol. 2016, 67, 1338–1357. [Google Scholar] [CrossRef] [PubMed]
  24. Tucker, W.J.; Fegers-Wustrow, I.; Halle, M.; Haykowsky, M.J.; Chung, E.H.; Kovacic, J.C. Exercise for Primary and Secondary Prevention of Cardiovascular Disease: JACC Focus Seminar 1/4. J. Am. Coll. Cardiol. 2022, 80, 1091–1106. [Google Scholar] [CrossRef]
  25. Faisal, M.; Scally, A.; Howes, R.; Beatson, K.; Richardson, D.; Mohammed, M.A. A comparison of logistic regression models with alternative machine learning methods to predict the risk of in-hospital mortality in emergency medical admissions via external validation. Health Inform. J. 2020, 26, 34–44. [Google Scholar] [CrossRef] [PubMed]
  26. Lu, T.; Fang, Y.; Liu, H.; Chen, C.; Li, T.; Lu, M.; Song, D. Comparison of Machine Learning and Logic Regression Algorithms for Predicting Lymph Node Metastasis in Patients with Gastric Cancer: A two-Center Study. Technol. Cancer Res. Treat. 2024, 23, 15330338231222331. [Google Scholar] [CrossRef]
  27. Chalkou, K.; Vickers, A.J.; Pellegrini, F.; Manca, A.; Salanti, G. Decision Curve Analysis for Personalized Treatment Choice between Multiple Options. Med. Decis. Mak. 2023, 43, 337–349. [Google Scholar] [CrossRef]
  28. Gershman, B.; Thompson, R.H.; Boorjian, S.A.; Lohse, C.M.; Costello, B.A.; Cheville, J.C.; Leibovich, B.C. Radical Versus Partial Nephrectomy for cT1 Renal Cell Carcinoma. Eur. Urol. 2018, 74, 825–832. [Google Scholar] [CrossRef]
  29. Liu, J. Computational aspects of psychometric methods with R by Patricia Martinková and Adéla Hladká, Chapman & Hall/CRC, 2023, ISBN: 9781003054313, https://doi.org/10.1201/9781003054313. Biometrics 2025, 82, ujaf132. [Google Scholar] [CrossRef]
  30. Yang, T.; Wei, J.; Wang, Z.; Li, S.; Zhang, T.; Jia, S.; Meng, C. Analysis of risk factors and prediction model construction of deep vein thrombosis in patients with lumbar degenerative diseases before surgery. Sci. Rep. 2025, 15, 26069. [Google Scholar] [CrossRef]
  31. Abbas, Q.; Jeong, W.; Lee, S.W. Explainable AI in Clinical Decision Support Systems: A Meta-Analysis of Methods, Applications, and Usability Challenges. Healthcare 2025, 13, 2154. [Google Scholar] [CrossRef]
  32. Li, C.; Xu, C.; Xu, J.; Song, W.; Yu, Z.; Zhang, Z.; Wei, D.; Li, W.; Qian, Y.; Lei, D. Development and validation of an interpretable shap-based machine learning model for predicting postoperative complications in laryngeal cancer. BMC Surg. 2025, 25, 469. [Google Scholar] [CrossRef] [PubMed]
  33. England, B.R.; Baker, J.F.; George, M.D.; Johnson, T.M.; Yang, Y.; Roul, P.; Frideres, H.; Sayles, H.; Yu, F.; Matson, S.M.; et al. Advanced therapies in US veterans with rheumatoid arthritis-associated interstitial lung disease: A retrospective, active-comparator, new-user, cohort study. Lancet Rheumatol. 2025, 7, e166–e177. [Google Scholar] [CrossRef] [PubMed]
  34. Sun, J.; Mehta, H.B.; Segal, J.B.; Alexander, G.C. Comparative safety and effectiveness of apixaban and rivaroxaban for treatment of cancer-associated venous thromboembolism: A retrospective cohort study. PLoS Med. 2025, 22, e1004754. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Discrimination, calibration, and clinical utility of the models. (AC) Training/internal-validation set: (A) ROC curves with AUROC for each algorithm; (B) calibration plots comparing predicted versus observed risk (smoothed curve and ideal 45° line); (C) decision-curve analysis (DCA) showing net benefit across threshold probabilities. (DF) Held-out test set: (D) ROC curves; (E) calibration plots; (F) DCA. Thresholds used in DCA and confusion matrices were locked from the training phase. Higher curves indicate better performance. Abbreviations: AUROC, area under the ROC curve; DCA, decision-curve analysis.
Figure 1. Discrimination, calibration, and clinical utility of the models. (AC) Training/internal-validation set: (A) ROC curves with AUROC for each algorithm; (B) calibration plots comparing predicted versus observed risk (smoothed curve and ideal 45° line); (C) decision-curve analysis (DCA) showing net benefit across threshold probabilities. (DF) Held-out test set: (D) ROC curves; (E) calibration plots; (F) DCA. Thresholds used in DCA and confusion matrices were locked from the training phase. Higher curves indicate better performance. Abbreviations: AUROC, area under the ROC curve; DCA, decision-curve analysis.
Bioengineering 13 00665 g001
Figure 2. Confusion matrices for the training/internal validation set (positive class = postoperative buttock claudication). Matrices depict true/false positives and true/false negatives at the locked operating threshold for each algorithm, ordered as: (A) AdaBoost, (B) CatBoost, (C) KNN, (D) LightGBM, (E) logistic regression, (F) neural network, (G) Random Forest, (H) SVM, (I) XGBoost, (J) GBM. Darker cells indicate higher counts. These visualizations highlight models that favor sensitivity versus specificity and reveal dominant error modes, and the blue color gradient reflects cell frequency, with darker blue indicating a higher count/proportion.
Figure 2. Confusion matrices for the training/internal validation set (positive class = postoperative buttock claudication). Matrices depict true/false positives and true/false negatives at the locked operating threshold for each algorithm, ordered as: (A) AdaBoost, (B) CatBoost, (C) KNN, (D) LightGBM, (E) logistic regression, (F) neural network, (G) Random Forest, (H) SVM, (I) XGBoost, (J) GBM. Darker cells indicate higher counts. These visualizations highlight models that favor sensitivity versus specificity and reveal dominant error modes, and the blue color gradient reflects cell frequency, with darker blue indicating a higher count/proportion.
Bioengineering 13 00665 g002
Figure 3. Confusion matrices for the held-out test set (positive class = postoperative buttock claudication). (A) AdaBoost, (B) CatBoost, (C) KNN, (D) LightGBM, (E) logistic regression, (F) neural network, (G) Random Forest, (H) SVM, (I) XGBoost, (J) GBM. Thresholds are identical to those fixed on the training data. Patterns illustrate generalizability of each model’s error profile, and the blue color gradient reflects cell frequency, with darker blue indicating a higher count/proportion.
Figure 3. Confusion matrices for the held-out test set (positive class = postoperative buttock claudication). (A) AdaBoost, (B) CatBoost, (C) KNN, (D) LightGBM, (E) logistic regression, (F) neural network, (G) Random Forest, (H) SVM, (I) XGBoost, (J) GBM. Thresholds are identical to those fixed on the training data. Patterns illustrate generalizability of each model’s error profile, and the blue color gradient reflects cell frequency, with darker blue indicating a higher count/proportion.
Bioengineering 13 00665 g003
Figure 4. Global variable importance across algorithms. Importance summarized on a comparable scale. IIA embolization status, distal IIA branch count, and iliac-involved aneurysm consistently rank highest. (A) AdaBoost; (B) CatBoost; (C) GBM; (D) KNN; (E) LightGBM; (F) logistic regression; (G) neural network; (H) Random Forest; (I) SVM; (J) XGBoost.
Figure 4. Global variable importance across algorithms. Importance summarized on a comparable scale. IIA embolization status, distal IIA branch count, and iliac-involved aneurysm consistently rank highest. (A) AdaBoost; (B) CatBoost; (C) GBM; (D) KNN; (E) LightGBM; (F) logistic regression; (G) neural network; (H) Random Forest; (I) SVM; (J) XGBoost.
Bioengineering 13 00665 g004
Figure 5. SHAP explainability for the best-performing model (neural network). (A) Global SHAP bar plot ranking features by mean absolute SHAP value (overall contribution). (B) Beeswarm plot: each point represents a patient; horizontal position indicates SHAP value (impact on risk), and color encodes feature value (low to high). (C) SHAP dependence plots for key predictors (e.g., IIA embolization status, distal branch count, iliac involvement), optionally with interaction hints. (D) Representative SHAP force plot decomposing an individual prediction into prediction-increasing positive contributions (yellow) and prediction-decreasing negative contributions (purple). (E) Waterfall plot illustrating cumulative SHAP contributions from baseline risk to the final predicted probability. Positive SHAP values raise, and negative values lower, the predicted risk.
Figure 5. SHAP explainability for the best-performing model (neural network). (A) Global SHAP bar plot ranking features by mean absolute SHAP value (overall contribution). (B) Beeswarm plot: each point represents a patient; horizontal position indicates SHAP value (impact on risk), and color encodes feature value (low to high). (C) SHAP dependence plots for key predictors (e.g., IIA embolization status, distal branch count, iliac involvement), optionally with interaction hints. (D) Representative SHAP force plot decomposing an individual prediction into prediction-increasing positive contributions (yellow) and prediction-decreasing negative contributions (purple). (E) Waterfall plot illustrating cumulative SHAP contributions from baseline risk to the final predicted probability. Positive SHAP values raise, and negative values lower, the predicted risk.
Bioengineering 13 00665 g005
Figure 6. Web-based clinical calculator for individualized risk estimation. Screenshot of the deployed Shiny application showing input fields for study predictors and the resulting predicted probability of postoperative buttock claudication with an accompanying explanation panel aligned to SHAP outputs. The tool is accessible at https://cacsriskmodel.shinyapps.io/makee/, accessed on 7 October 2025.
Figure 6. Web-based clinical calculator for individualized risk estimation. Screenshot of the deployed Shiny application showing input fields for study predictors and the resulting predicted probability of postoperative buttock claudication with an accompanying explanation panel aligned to SHAP outputs. The tool is accessible at https://cacsriskmodel.shinyapps.io/makee/, accessed on 7 October 2025.
Bioengineering 13 00665 g006
Table 1. Characteristics of included patients.
Table 1. Characteristics of included patients.
Demographic CharacteristicsDescNo Buttock Claudication (N = 201)Buttock Claudication (N = 71)p
Site_of_aneurysmAbdominal_aortic_aneurysm_with_iliac_artery_aneurysm53 (26.4%)34 (47.9%)0.001
Abdominal_aortic_aneurysm_without_iliac_artery_aneurysm148 (73.6%)37 (52.1%)
GenderFemale67 (33.3%)15 (21.1%)0.076
Male134 (66.7%)56 (78.9%)
Peripheral_arterial_diseaseNo131 (65.2%)32 (45.1%)0.005
Yes70 (34.8%)39 (54.9%)
Number_of_internal_iliac_arteries_embolizedBilateral_embolization14 (7%)10 (14.1%)<0.001
None80 (39.8%)10 (14.1%)
Unilateral_embolism107 (53.2%)51 (71.8%)
Chronic_obstructive_pulmonary_diseaseNo187 (93%)69 (97.2%)0.325
Yes14 (7%)2 (2.8%)
Chronic_kidney_diseaseNo159 (79.1%)55 (77.5%)0.903
Yes42 (20.9%)16 (22.5%)
AntiplateletNo63 (31.3%)26 (36.6%)0.504
Yes138 (68.7%)45 (63.4%)
HyperlipidemiaNo68 (33.8%)9 (12.7%)0.001
Yes133 (66.2%)62 (87.3%)
History_of_previous_abdominal_and_pelvic_surgeryNo112 (55.7%)50 (70.4%)0.042
Yes89 (44.3%)21 (29.6%)
Cardiovascular_diseaseNo126 (62.7%)43 (60.6%)0.861
Yes75 (37.3%)28 (39.4%)
Marital_statusMarried92 (45.8%)32 (45.1%)1
Other109 (54.2%)39 (54.9%)
SmokingNo160 (79.6%)58 (81.7%)0.837
Yes41 (20.4%)13 (18.3%)
DrinkingNo128 (63.7%)45 (63.4%)1
Yes73 (36.3%)26 (36.6%)
Cerebrovascular_diseaseNo106 (52.7%)33 (46.5%)0.442
Yes95 (47.3%)38 (53.5%)
HypertensionNo103 (51.2%)42 (59.2%)0.312
Yes98 (48.8%)29 (40.8%)
DiabetesNo136 (67.7%)43 (60.6%)0.348
Yes65 (32.3%)28 (39.4%)
Number_of_distal_internal_iliac_artery_branches≤2104 (51.7%)63 (88.7%)<0.001
>297 (48.3%)8 (11.3%)
AgeMean ± SD65.7 ± 7.665.0 ± 8.30.484
BMIMean ± SD30.0 ± 3.829.8 ± 4.50.725
Maximum_diameter_of_abdominal_aortaMean ± SD5.8 ± 0.65.7 ± 0.50.094
Table 2. Training set univariate and multivariate logistic regression results.
Table 2. Training set univariate and multivariate logistic regression results.
Demographic CharacteristicsDescNo Buttock Claudication (N = 141)Buttock Claudication (N = 50)OR (Univariable)OR (Multivariable)
Site_of_aneurysmAbdominal_aortic_aneurysm_without_iliac_artery_aneurysm107 (75.9%)26 (52%)
Abdominal_aortic_aneurysm_with_iliac_artery_aneurysm34 (24.1%)24 (48%)2.90 (1.48–5.71, p = 0.002)4.04 (1.73–9.40, p = 0.001)
GenderFemale45 (31.9%)7 (14%)
Male96 (68.1%)43 (86%)2.88 (1.20–6.90, p = 0.018)3.26 (1.20–8.86, p = 0.021)
Peripheral_arterial_diseaseNo94 (66.7%)21 (42%)
Yes47 (33.3%)29 (58%)2.76 (1.42–5.35, p = 0.003)2.19 (0.98–4.89, p = 0.055)
Number_of_internal_iliac_arteries_embolizedNone58 (41.1%)7 (14%)
Unilateral_embolism76 (53.9%)36 (72%)3.92 (1.63–9.45, p = 0.002)3.86 (1.38–10.80, p = 0.010)
Bilateral_embolization7 (5%)7 (14%)8.29 (2.24–30.67, p = 0.002)8.61 (1.78–41.68, p = 0.007)
Chronic_obstructive_pulmonary_diseaseNo131 (92.9%)48 (96%)
Yes10 (7.1%)2 (4%)0.55 (0.12–2.58, p = 0.445)
Chronic_kidney_diseaseNo111 (78.7%)41 (82%)
Yes30 (21.3%)9 (18%)0.81 (0.36–1.86, p = 0.622)
AntiplateletNo47 (33.3%)22 (44%)
Yes94 (66.7%)28 (56%)0.64 (0.33–1.23, p = 0.179)
HyperlipidemiaNo50 (35.5%)5 (10%)
Yes91 (64.5%)45 (90%)4.95 (1.84–13.26, p = 0.002)5.66 (1.84–17.37, p = 0.002)
History_of_previous_abdominal_and_pelvic_surgeryNo77 (54.6%)33 (66%)
Yes64 (45.4%)17 (34%)0.62 (0.32–1.21, p = 0.163)
Cardiovascular_diseaseNo87 (61.7%)27 (54%)
Yes54 (38.3%)23 (46%)1.37 (0.72–2.63, p = 0.341)
Marital_statusother81 (57.4%)27 (54%)
married60 (42.6%)23 (46%)1.15 (0.60–2.20, p = 0.673)
SmokingNo112 (79.4%)42 (84%)
Yes29 (20.6%)8 (16%)0.74 (0.31–1.74, p = 0.484)
DrinkingNo95 (67.4%)33 (66%)
Yes46 (32.6%)17 (34%)1.06 (0.54–2.11, p = 0.859)
Cerebrovascular_diseaseNo82 (58.2%)25 (50%)
Yes59 (41.8%)25 (50%)1.39 (0.73–2.66, p = 0.319)
HypertensionNo78 (55.3%)30 (60%)
Yes63 (44.7%)20 (40%)0.83 (0.43–1.59, p = 0.566)
DiabetesNo94 (66.7%)31 (62%)
Yes47 (33.3%)19 (38%)1.23 (0.63–2.40, p = 0.551)
Number_of_distal_internal_iliac_artery_branches≤272 (51.1%)43 (86%)
>269 (48.9%)7 (14%)0.17 (0.07–0.40, p < 0.001)0.15 (0.05–0.41, p < 0.001)
AgeMean ± SD66.1 ± 7.164.4 ± 8.90.97 (0.93–1.01, p = 0.162)
BMIMean ± SD30.0 ± 3.830.3 ± 4.01.02 (0.94–1.12, p = 0.596)
Maximum_diameter_of_abdominal_aortaMean ± SD5.8 ± 0.65.7 ± 0.50.82 (0.46–1.46, p = 0.501)
Table 3. Training set evaluation metrics.
Table 3. Training set evaluation metrics.
ModelThresholdAccuracySensitivitySpecificityPrecisionF1
Logistic0.2551790633360430.7540.80.7380.5190.63
SVM0.3837688546711910.7590.80.7450.5260.635
GBM0.1767103216097960.7230.880.6670.4840.624
NeuralNetwork0.2839618582478190.7540.80.7380.5190.63
RandomForest0.50.8220.420.9650.8080.553
Xgboost0.2284727096557620.7330.860.6880.4940.628
KNN0.3680376822770790.780.760.7870.5590.644
AdaBoost0.3781568895704570.7280.860.6810.4890.623
LightGBM0.2682908104665960.7640.860.730.5310.656
CatBoost0.5648150318210620.6910.840.6380.4520.587
Table 4. Test set evaluation metrics.
Table 4. Test set evaluation metrics.
ModelThresholdAccuracySensitivitySpecificityPrecisionF1
Logistic0.2177651163033090.6670.810.6170.4250.557
SVM0.2850176907533190.7040.7140.70.4550.556
GBM0.2503739669621870.7040.7140.70.4550.556
NeuralNetwork0.2557766182800990.6670.810.6170.4250.557
RandomForest0.50.7410.2380.9170.50.323
Xgboost0.2368519902229310.7040.7140.70.4550.556
KNN0.2841180411398830.6910.6190.7170.4330.51
AdaBoost0.4123475430183690.6910.7140.6830.4410.545
LightGBM0.2274970648229740.6420.6670.6330.3890.491
CatBoost0.5683412303978050.790.4760.90.6250.541
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Li, Y.; Deng, H.; Gu, Y. Development and Internal Validation of an Explainable Machine Learning Model for Predicting Buttock Claudication After EVAR: A Dual-Center Cohort Study. Bioengineering 2026, 13, 665. https://doi.org/10.3390/bioengineering13060665

AMA Style

Li Y, Deng H, Gu Y. Development and Internal Validation of an Explainable Machine Learning Model for Predicting Buttock Claudication After EVAR: A Dual-Center Cohort Study. Bioengineering. 2026; 13(6):665. https://doi.org/10.3390/bioengineering13060665

Chicago/Turabian Style

Li, Yajing, Hongru Deng, and Yongquan Gu. 2026. "Development and Internal Validation of an Explainable Machine Learning Model for Predicting Buttock Claudication After EVAR: A Dual-Center Cohort Study" Bioengineering 13, no. 6: 665. https://doi.org/10.3390/bioengineering13060665

APA Style

Li, Y., Deng, H., & Gu, Y. (2026). Development and Internal Validation of an Explainable Machine Learning Model for Predicting Buttock Claudication After EVAR: A Dual-Center Cohort Study. Bioengineering, 13(6), 665. https://doi.org/10.3390/bioengineering13060665

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop