Automated Single-Slice Lumbar QCT HU Value Measurement with Clinical Workflow
Abstract
1. Introduction
2. Materials and Methods
2.1. Study Population
2.2. CT Acquisition Protocols
2.3. Automated Pipeline
2.3.1. Pre-Processing
2.3.2. Eligibility Gate: ViT-B/16 Slice Prescreening
2.3.3. QC-Envelope Segmentation: Whole-Vertebra U-Net–ResNet34
2.3.4. Intra-Patient Quality Ranking and Best-Slice Selection: PairRank-Swin
2.3.5. HU Calculation and Optional Exploratory vBMD Proxy Conversion
2.4. Model Training and Validation
2.5. Evaluation Metrics and Statistical Analysis
2.5.1. Eligibility Gate (ViT-B/16)
2.5.2. Intra-Patient Quality Ranking and Best-Slice Selection (PairRank-Swin)
2.5.3. Segmentation Metrics
2.5.4. HU Agreement (Primary Outcome)
2.5.5. Assessment of Intra-Rater Repeatability of Manual HU Measurement
2.5.6. General Settings
2.5.7. Open Module-Level Baseline Comparison
3. Results
3.1. Performance of the Eligibility Gate
3.2. QC-Envelope and Trabecular Segmentation
3.3. Intra-Patient Quality Ranking (PairRank-Swin)
3.4. Intra-Rater Repeatability of Manual HU Measurement
3.5. Agreement Between Automated and Expert HU Values
3.6. Ablation Analysis of the End-to-End Workflow
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Pisani, P.; Renna, M.D.; Conversano, F.; Casciaro, E.; Di Paola, M.; Quarta, E.; Muratore, M.; Casciaro, S. Major osteoporotic fragility fractures: Risk factor updates and societal impact. World J. Orthop. 2016, 7, 171–181. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- El Maghraoui, A.; Roux, C. DXA scanning in clinical practice. QJM 2008, 101, 605–617. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gruenewald, L.D.; Koch, V.; Martin, S.S.; Yel, I.; Eichler, K.; Gruber-Rouh, T.; Lenga, L.; Wichmann, J.L.; Alizadeh, L.S.; Albrecht, M.H.; et al. Diagnostic accuracy of quantitative dual-energy CT-based volumetric bone mineral density assessment for the prediction of osteoporosis-associated fractures. Eur. Radiol. 2022, 32, 3076–3084. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yoon, H.; Kim, J.-H.; Ryu, D.-S.; Yoon, S.-H. What Causes the Discrepancy between Quantitative Computed Tomography and Dual Energy X-Ray Absorptiometry? Nerve 2021, 7, 64–70. [Google Scholar] [CrossRef] [Scilit]
- Jones, G.; Nguyen, T.; Sambrook, P.N.; Kelly, P.J.; Eisman, J.A. A longitudinal study of the effect of spinal degenerative disease on bone density in the elderly. J. Rheumatol. 1995, 22, 932–936. [Google Scholar] [PubMed]
- Park, H.; Kang, W.Y.; Woo, O.H.; Lee, J.; Yang, Z.; Oh, S. Automated deep learning-based bone mineral density assessment for opportunistic osteoporosis screening using various CT protocols with multi-vendor scanners. Sci. Rep. 2024, 14, 25014. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fusco, S.; Spadafora, P.; Gallazzi, E.; Ghiara, C.; Albano, D.; Sconfienza, L.M.; Messina, C. Comparison Between Quantitative Computed Tomography-Based Bone Mineral Density Values and Dual-Energy X-Ray Absorptiometry-Based Parameters of Bone Density and Microarchitecture: A Lumbar Spine Study. Appl. Sci. 2025, 15, 3248. [Google Scholar] [CrossRef] [Scilit]
- Oliveira, M.A.; Moraes, R.; Castanha, E.B.; Prevedello, A.S.; Vieira Filho, J.; Bussolaro, F.A.; García Cava, D. Osteoporosis Screening: Applied Methods and Technological Trends. Med. Eng. Phys. 2022, 108, 103887. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, J.; Zeng, Y.; Yu, W. Criteria for osteoporosis diagnosis: A systematic review and meta-analysis of osteoporosis diagnostic studies with DXA and QCT. EClinicalMedicine 2025, 83, 103244. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pickhardt, P.J.; Pooler, B.D.; Lauder, T.; del Rio, A.M.; Bruce, R.J.; Binkley, N. Opportunistic screening for osteoporosis using abdominal computed tomography scans obtained for other indications. Ann. Intern. Med. 2013, 158, 588–595. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jang, S.; Graffy, P.M.; Ziemlewicz, T.J.; Lee, S.J.; Summers, R.M.; Pickhardt, P.J. Opportunistic Osteoporosis Screening at Routine Abdominal and Thoracic CT: Normative L1 Trabecular Attenuation Values in More than 20 000 Adults. Radiology 2019, 291, 360–367. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vadera, S.; Osborne, T.; Shah, V.; Stephenson, J.A. Opportunistic screening for osteoporosis by abdominal CT in a British population. Insights Imaging 2023, 14, 57. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Westerhoff, M.; Gyftopoulos, S.; Dane, B.; Vega, E.; Murdock, D.; Lindow, N.; Herter, F.; Bousabarah, K.; Recht, M.P.; Bredella, M.A. Deep Learning-based Opportunistic CT Osteoporosis Screening and the Establishment of Normative Values. Radiology 2025, 317, e250917. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hamouda, A.M.; Pennington, Z.; Astudillo Potes, M.; Shafi, M.; Mikula, A.L.; Lakomkin, N.; Martini, M.L.; Bydon, M.; Kennel, K.A.; Drake, M.T.; et al. Impact of contrast administration and CT reconstruction plane on Hounsfield units for assessing underlying bone quality in the lumbar spine. J. Neurosurg. Spine 2025, 42, 331–339. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liebl, H.; Schinz, D.; Sekuboyina, A.; Malagutti, L.; Löffler, M.T.; Bayat, A.; El Husseini, M.; Tetteh, G.; Grau, K.; Niederreiter, E.; et al. A computed tomography vertebral segmentation dataset with anatomical variations and multi-vendor scanner data. Sci. Data 2021, 8, 284. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Guo, M.; Zhang, Y.; Gu, X.; Liu, X.; Peng, F.; Zhang, Z.; Jing, M.; Fu, Y. A comparative study of bone density in elderly people measured with AI and QCT. Front. Artif. Intell. 2025, 8, 1582960. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, S.; Tong, X.; Cheng, Q.; Xiao, Q.; Cui, J.; Li, J.; Liu, Y.; Fang, X. Fully automated deep learning system for osteoporosis screening using chest computed tomography images. Quant. Imaging Med. Surg. 2024, 14, 2816–2827. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pan, J.; Lin, P.C.; Gong, S.C.; Wang, Z.; Cao, R.; Lv, Y.; Zhang, K.; Wang, L. Effectiveness of opportunistic osteoporosis screening on chest CT using the DCNN model. BMC Musculoskelet. Disord. 2024, 25, 176. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Niu, X.; Huang, Y.; Li, X.; Yan, W.; Lu, X.; Jia, X.; Li, J.; Hu, J.; Sun, T.; Jing, W.; et al. Development and validation of a fully automated system using deep learning for opportunistic osteoporosis screening using low-dose computed tomography scans. Quant. Imaging Med. Surg. 2023, 13, 5294–5305. [Google Scholar] [CrossRef] [Scilit] [PubMed]







| Variable | Development Set | External Validation Cohort |
|---|---|---|
| Age (years), mean ± SD | 64.6 ± 14.4 | 62.1 ± 12.4 |
| Sex (female/male), n (%) | 56 (61.5%)/35 (38.5%) | 31 (62.0%)/19 (38.0%) |
| Total number of slices, n (L1–L3, DICOM) | 2625 | 1482 |
| “Problematic slices” (severe deformity/metal artifact), n (%) | 737 (28.1%) | 297 (20.0%) |
| Cohort/Subset | Patients (n) | Slices (n) | Purpose | Notes |
|---|---|---|---|---|
| Development (overall) | 91 | 2625 | Source pool | Patient-level 5-fold CV |
| Segmentation subset (QC-Envelope + trabecular) | 38 | 922 | Dual-target U-Net training (QC-Envelope + trabecular) | Learning curve plateaued |
| Eligibility Gate | 27 | 736 | ViT-B/16 classifier | ratio ≈ 9:1 |
| Intra-patient Quality Ranking subset | 91 | 2360 | PairRank-Swin training | Intra-patient positive-negative pairing |
| Parameter | Value |
|---|---|
| Tube voltage (kVp) | 120 (n = 60), 135 (n = 31) |
| Tube current (mA) | 310 ± 163 (range, 80–638) |
| Slice thickness (mm) | 3.0 (n = 64), 5.0 (n = 23), other/irregular settings (n = 4) |
| Matrix | 810 × 810 (n = 33), 512 × 512 (n = 55), 512 × 518 (n = 3) |
| Field of view (mm) | 161.1 ± 29.2 |
| Cohort (Evaluation) | Slices | Accuracy (%) | Precision (%) | Recall (Sensitivity) (%) | Specificity (%) | F1 (%) |
|---|---|---|---|---|---|---|
| Development (5-fold, aggregated) | 736 | 99.05 (95% CI, 98.05–99.54) | 99.40 (95% CI, 98.46–99.76) | 99.55 (95% CI, 98.67–99.85) | 94.67 (95% CI, 87.07–97.91) | 99.47 |
| External (independent) | 1482 | 99.26 (95% CI, 98.68–99.59) | 99.62 (95% CI, 99.10–99.84) | 99.54 (95% CI, 99.00–99.79) | 97.22 (95% CI, 93.66–98.81) | 99.58 |
| Cohort | Segmentation Target | Dice (Mean ± SD) | 95% CI (Dice) | Slices Used |
|---|---|---|---|---|
| Development (CV) | QC-Envelope (whole-vertebra) | 0.9596 ± 0.0042 | 0.9593–0.9599 | 922 labeled slices |
| Development (CV) | Trabecular (cancellous compartment) | 0.9668 ± 0.0082 | 0.9663–0.9673 | 922 labeled slices |
| External (independent) | QC-Envelope (whole-vertebra) | 0.9572 ± 0.0060 | 0.9460–0.9670 | post-Eligibility Gate n = 1302 |
| External (independent) | Trabecular (cancellous compartment) | 0.9710 ± 0.0075 | 0.9600–0.9830 | post-Eligibility Gate n = 1302 |
| Model | N Slices | Dice | HD95 | ASSD |
|---|---|---|---|---|
| QC-Envelope (ours) | 342 | 0.977 ± 0.016 | 6.42 ± 7.45 | 2.16 ± 1.61 |
| TotalSegmentator (lumbar union) | 342 | 0.430 ± 0.256 | 86.76 ± 47.89 | 34.45 ± 26.65 |
| Cohort | Post-Eligibility Gate Slices | Threshold | Precision% | Recall% | F1% | PR-AUC |
|---|---|---|---|---|---|---|
| Development (5-fold CV) | 2360 | 0.55 (fixed) | 84 ± 6 | 83 ± 5 | 84 ± 2 | 0.87 ± 0.04 |
| External (independent) | 1302 | 0.55 (fixed) | 91.99 | 80.02 | 85.59 | 0.88 ± 0.05 |
| Metric | Estimate | 95% CI | p-Value | Notes |
|---|---|---|---|---|
| Pearson’s r | 0.987 | 0.976 to 0.993 | — | Correlation between automated and manual HU |
| Spearman’s ρ | 0.980 | 0.946 to 0.989 | — | Rank correlation between automated and manual HU |
| Lin’s CCC | 0.985 | 0.978 to 0.989 | — | Bootstrap CI |
| Mean bias (HU) | −0.44 | −2.62 to +1.71 | 0.691 | Bland–Altman bias; H0: bias = 0 |
| Limits of agreement (HU) | −14.88 to +13.99 | — | — | Lower LoA CI: −18.20 to −11.10; Upper LoA CI: +10.36 to +17.00 |
| Within ±10 HU | 37/44 (84.1%) | 70.6% to 92.1% | — | Wilson CI |
| Within ±15 HU | 43/44 (97.7%) | 88.2% to 99.6% | — | Wilson CI |
| OLS slope, b | 1.057 | 1.004 to 1.111 | 0.036 | Regression: Machine = a + b × Manual; H0: b = 1 |
| OLS intercept, a | −5.31 | −10.34 to −0.28 | 0.039 | Regression: Machine = a + b × Manual; H0: a = 0 |
| Setting | N Paired Cases | Pearson r | Lin’s CCC | MAE (HU) | Bias (HU) | Within ±10 HU | Within ±15 HU |
|---|---|---|---|---|---|---|---|
| FULL | 44 | 0.987 | 0.985 | 6.10 | −0.44 | 84.1% | 97.7% |
| NO_GATE | 44 | 0.987 | 0.985 | 5.92 | −0.27 | 86.4% | 97.7% |
| NO_RANK | 44 | 0.491 | 0.250 | 70.04 | +65.36 | 15.9% | 20.5% |
| ROI_NO_ELLIPSE | 44 | 0.968 | 0.967 | 7.98 | +1.77 | 72.7% | 86.4% |
| ROI_NO_EROSION | 44 | 0.902 | 0.713 | 32.05 | +31.25 | 6.8% | 15.9% |
| Study | Input Granularity | QC Mechanism (Slice-Level) | Intermediate-Output Traceability | External Evaluation | Notes |
|---|---|---|---|---|---|
| Niu et al. [19] | Low-dose chest CT (multi-slice) | No explicit slice-level QC reported | Intermediate artifacts not emphasized as stored reviewable outputs | Large, multi-slice, single-center | Strong automated screening performance, but not specialized for lumbar QCT or operator-style single-slice measurement |
| Wang et al. [17] | Chest CT (multi-slice) | No explicit slice-level QC reported | Multi-task outputs reported, but not framed as preserved stepwise review artifacts | Large, multi-slice, single-center | Localization/segmentation/classification framework using multi-slice inputs |
| Westerhoff et al. [13] | Multi-slice/volumetric CT | Not emphasized as an explicit slice-level gate (focus on automated volumetric ROI placement and protocol-aware harmonization) | ROI logic described, but not designed as a stepwise review-preserving workflow | multi-scanner/multi-protocol cohort | Automated 3D trabecular ROI, harmonization across scanner models and tube voltages, and establishment of normative values/screening thresholds |
| This study | Lumbar QCT (single-slice) | Eligibility Gate + intra-patient quality ranking (PairRank-Swin) + best-slice selection | Reviewable intermediate outputs preserved, including QC-Envelope overlays, trabecular masks/ROI, QC decisions, and attention maps | Small, single-center, cross-device test | Stepwise QC-first workflow with locked thresholds and explicit rejection of non-evaluable or borderline-quality inputs |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ye, Z.-Y.; Peng, J.-M.; Lu, B.-Q.; Kamishima, T. Automated Single-Slice Lumbar QCT HU Value Measurement with Clinical Workflow. Mach. Learn. Knowl. Extr. 2026, 8, 77. https://doi.org/10.3390/make8030077
Ye Z-Y, Peng J-M, Lu B-Q, Kamishima T. Automated Single-Slice Lumbar QCT HU Value Measurement with Clinical Workflow. Machine Learning and Knowledge Extraction. 2026; 8(3):77. https://doi.org/10.3390/make8030077
Chicago/Turabian StyleYe, Zhe-Yu, Jun-Mu Peng, Bing-Qian Lu, and Tamotsu Kamishima. 2026. "Automated Single-Slice Lumbar QCT HU Value Measurement with Clinical Workflow" Machine Learning and Knowledge Extraction 8, no. 3: 77. https://doi.org/10.3390/make8030077
APA StyleYe, Z.-Y., Peng, J.-M., Lu, B.-Q., & Kamishima, T. (2026). Automated Single-Slice Lumbar QCT HU Value Measurement with Clinical Workflow. Machine Learning and Knowledge Extraction, 8(3), 77. https://doi.org/10.3390/make8030077

