Background and Objectives: Osteoporotic vertebral compression fractures affect approximately one in four postmenopausal women and carry substantial morbidity, yet established clinical tools such as dual-energy X-ray absorptiometry (DXA) provide only patient-level risk and do not identify which specific vertebra is most likely to fail. Computed tomography (CT) acquired for unrelated indications is the most widely available three-dimensional substrate for opportunistic screening, but published machine learning models for vertebral fracture risk almost universally operate at the patient level. The present study aimed to develop and rigorously validate a per-vertebra prediction pipeline applicable to both routine clinical lumbar-spine CT and opportunistic abdominal CT, both acquired for indications unrelated to osteoporosis screening.
Materials and Methods: Two independent retrospective cohorts were assembled from a single academic centre: a routine clinical lumbar-spine CT cohort of 106 patients yielding 478 evaluable vertebrae, and a routine abdominal CT cohort of 126 patients yielding 589 evaluable vertebrae. Vertebral bodies were segmented automatically with TotalSegmentator v2 and the trabecular core isolated by morphological erosion. A panel of 505 quantitative imaging biomarkers compliant with Image Biomarker Standardisation Initiative recommendations was extracted, covering trabecular density, vertebral morphometry, classical texture, trabecular network architecture, sub-endplate vulnerability, low-density topology, radial heterogeneity and adjacent muscle quality. Within-patient feature engineering expanded the input pool to 1293 contextual descriptors. Three model families were evaluated under fully nested leave-one-patient-out cross-validation: ElasticNet logistic regression, a softmax-ranking approximation of conditional logistic regression, and a Two-Stage model combining a patient-level fragility score with a within-patient vertebral outlier score. Patient-level bootstrap resampling (2000 iterations) was used to obtain 95% confidence intervals.
Results: On routine clinical lumbar-spine CT the Two-Stage model achieved a per-vertebra AUC of 0.750 (95% CI 0.704 to 0.795), an F1 of 0.549, a within-patient concordance index of 0.693, an expected calibration error of 0.044, and Hit@3 of 0.934. It was the only model evaluated that returned calibrated probabilities; the softmax-ranking and ElasticNet baselines gave expected calibration errors of 0.232 and 0.218 respectively. On opportunistic abdominal CT, the softmax-ranking model gave AUC 0.672 (95% CI 0.615 to 0.727). Selected biomarkers were dominated by regional trabecular density and trabecular network architecture; a stable core of lumbar features entered the model in 100% of cross-validation folds, indicating high reproducibility. The closest prior per-vertebra CT-based predictor in primary, non-surgical patients (Muehlematter and colleagues, 58-patient cohort) reported a per-vertebra AUC of 0.64, which is one of several reference points for the present results. Ten methodological variants and sensitivity analyses, including rank fusion, internal tissue normalisation and additional biomechanical features, did not provide statistically significant gains, indicating that the binding constraint at this sample size is data volume rather than methodology.
Conclusions: A two-stage decomposition that separates systemic skeletal fragility from within-patient vertebral outlier status produces well-calibrated per-vertebra fracture-risk estimates from routine clinical lumbar spine CT and was the only model evaluated to do so, which is what permits a per-vertebra output to be reported as an absolute risk rather than as an ordering alone; a within-patient ranking model is preferable for opportunistic abdominal CT. The discrimination advantage of the decomposition over that baseline is numerical and consistent but not statistically established at this sample size, and the work is presented as a transparent and reproducible single-centre benchmark for the still under-developed per-vertebra prediction task. Its clearest near-term value is opportunistic, namely flagging elevated per-vertebra fracture risk on CTs already acquired for unrelated indications without additional radiation, cost or a dedicated densitometric study. External multi-centre validation is the necessary next step.
Full article