Next Article in Journal
Research and Advances in Miniaturization of Optical Spectrometers
Previous Article in Journal
Research on Mode-Locked Erbium-Doped Fiber Laser Based on the Lyot Filter
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Machine-Learning Augmented Optical Tissue Sensing for Intraoperative Guidance in Spine Surgery: A Systematic Review and Meta-Analysis

1
School of Medicine, Baylor College of Medicine, 1 Baylor Plaza, Houston, TX 77030, USA
2
T.H. Chan School of Medicine, UMass Chan Medical School, Worcester, MA 01655, USA
3
College of Natural Sciences, The University of Texas at Austin, Austin, TX 78712, USA
4
Department of Biomedical Sciences, Texas A&M University, College Station, TX 77843, USA
5
Human-Machine Perception Laboratory, Department of Computer Science, University of Nevada, Reno, NV 89557, USA
6
Department of Medical Genetics, University of Cambridge, Cambridge CB2 0QQ, UK
7
Department of Ophthalmology, Harvard Medical School, Massachusetts Eye and Ear, Boston, MA 02114, USA
8
Department of Computer Science, Cornell University, Ithaca, NY 14853, USA
9
Department of Orthopaedic Surgery, Baylor University Medical Center, Dallas, TX 75246, USA
*
Authors to whom correspondence should be addressed.
Optics 2026, 7(5), 62; https://doi.org/10.3390/opt7050062
Submission received: 12 July 2026 / Revised: 13 August 2026 / Accepted: 24 August 2026 / Published: 27 August 2026
(This article belongs to the Section Biomedical Optics)

Abstract

Background: Pedicle-screw malposition and canal breach can cause major neurologic and vascular complications. Fluoroscopy, navigation, and robotics guide trajectory, but they do not directly identify the tissue in front of the instrument. We evaluated machine learning (ML)-augmented optical sensing for tissue discrimination in spine-relevant settings. Methods: We conducted a PRISMA-DTA systematic review under a prospectively registered PROSPERO protocol and searched five databases. Classification accuracy was pooled with random-effects models and reported with 95% confidence intervals (CI) and prediction intervals (PI). Study quality was assessed with QUADAS-2 and separate ML-specific signaling questions. Results: Eight studies met the inclusion criteria. Five single-group ex vivo or cadaveric classification studies entered the meta-analysis, four using porcine tissue. Across individual spectra, image patches, or images, pooled classification accuracy was 95.9% (95% CI, 93.5–97.7%), with substantial between-study variation (I2 = 99.4%; 95% PI, 89.8–99.4%). A sensitivity analysis using subject counts as the denominator yielded 97.36% (95% CI, 87.46–100.00%). Because outcomes were not reported at the subject-level, this estimate should not be read as patient-level accuracy. No study tested intraoperative tissue classification in living humans or performed external validation, and one had a high risk of data leakage. Two non-ML perfusion studies were retained only to describe a prespecified evidence gap. Conclusions: ML-augmented optical sensing can distinguish spine-relevant tissues in controlled preclinical settings. Current evidence does not establish intraoperative or patient-level performance. The literature remains small, heterogeneous, predominantly preclinical, and without external validation. Prospective in vivo human studies with independent validation are needed before these systems can support intraoperative decision-making.

1. Introduction

Safe placement of pedicle-screw fixation in spine surgery remains difficult because the pedicle is a narrow bony corridor bordered by neural and vascular structures. A cortical breach or malpositioned screw can injure the spinal cord, nerve roots, or adjacent vessels, with medial and inferior breaches carrying the greatest risk of neurological injury. Fluoroscopy, navigation, and robotics improve placement accuracy by showing where an instrument is relative to preoperative or intraoperative images. However, there is no technology that can tell the surgeon what tissue is directly in front of an advancing probe or screw. Registration error, anatomical shift, or a small trajectory deviation can therefore carry an otherwise well-aligned instrument through the cortex and into the epidural space or spinal canal without real-time warning [1,2,3,4,5,6].
Optical sensing offers a different kind of feedback because it samples tissue directly at the tool’s tip. For example, diffuse reflectance spectroscopy (DRS) measures wavelength-dependent absorption and scattering and can help distinguish cancellous from cortical bone before a breach occurs [2]; Raman spectroscopy detects molecular vibrational signatures that may separate bone from soft tissue and neural structures; and optical coherence tomography (OCT), including polarization-sensitive OCT, provides micron-scale structural detail, while birefringence can highlight collagen-rich tissues such as ligamentum flavum and dura. Other optical techniques like near-infrared spectroscopy with photoplethysmography and laser speckle contrast imaging focus on spinal cord oxygenation and perfusion rather than tissue identification. Together, these methods map to four main intraoperative uses: pedicle trajectory and breach avoidance, bone quality and screw purchase, detection of the bone–dura or epidural interface, and monitoring of spinal cord perfusion or ischemia (Figure 1).
Machine learning (ML) and deep learning (DL) can convert optical spectra and images into clinically useful readouts, such as cancellous versus cortical bone or bone versus soft tissue. That creates the possibility of real-time feedback during instrumentation and parallels the broader use of ML in spine surgery and musculoskeletal diagnostics [8,9,10,11,12,13]. Reported performance, however, can appear better than it actually is when models are developed in a single cohort, measurements from the same subject are used for both training and test sets, data leakage occurs, or external validation is absent [14,15,16,17,18].
Despite its growing interest, ML-assisted optical sensing in spine surgery has not been systematically evaluated. The relevant studies are spread across a multitude of specialties, including engineering, biophotonics, and the surgical literature, and each differs in the optical modality, tissue target, reference standard, and validation approach discussed. We therefore conducted a PRISMA-DTA systematic review and diagnostic accuracy meta-analysis to identify eligible studies, pool classification accuracy, examine heterogeneity and prediction intervals, and assess risk of bias with ML-specific criteria. Our aim was to determine how well these systems currently perform and how much confidence those performance estimates deserve.

2. Materials and Methods

2.1. Eligibility Criteria

We followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses-Diagnostic Test Accuracy (PRISMA-DTA) guidelines and prospectively registered the protocol in PROSPERO (CRD420261447883) [19]. The complete protocol, including the full search strategy, is available through PROSPERO. Eligibility criteria, the search strategy, and the analysis plan were set before data extraction and followed the PICOT framework. Eligible populations included human, high-fidelity ex vivo, or animal spine or spine-relevant bone/tissue models. Index tests were optical modalities, including DRS, Raman, OCT, PS-OCT, NIRS, and LSCI, coupled with ML or DL. Accepted reference standards included histology, micro-computed tomography, cone-beam CT anatomical labeling, mechanical testing, expert anatomical identification, or a controlled induced condition. Outcomes were diagnostic accuracy, regression performance, or processing latency in tissue-guidance settings. We excluded reviews, commentaries, phantom-only or purely in vitro studies, and studies without tissue-differentiation or property-prediction endpoints. The protocol also prespecified a narrow narrative extension for spine-specific non-ML optical studies when they addressed an otherwise unrepresented clinical use case, spinal cord perfusion/ischemia monitoring. These studies were kept separate from the PRISMA-DTA diagnostic accuracy evidence: they were excluded from quantitative synthesis and diagnostic accuracy eligibility assessment and were used only to document that evidence gap.

2.2. Information Sources and Search

We searched PubMed, Embase, IEEE Xplore, Web of Science, and Scopus on 1 June 2026 using three Boolean strategies. Across the five databases, the searches returned 4835 records. Boolean limits required trimming the optics facet in focused searches, although every included modality was mapped to retained terms. We also performed backward and forward citation chaining. Bayhaqi 2023 [20] and Unal 2025 [21] were identified only through citation chaining because both studied femoral rather than vertebral tissue and lacked spine indexing. Full search strings, counts, de-duplication, and screening decisions are provided in the search log.

2.3. Study Selection and Extraction

Records were screened for both an optical modality and an ML/DL component, except for the pre-specified perfusion extension. Three reviewers independently screened titles and abstracts, reviewed full texts against the PICOT criteria, and resolved disagreements by consensus. The same reviewers extracted study characteristics, modality, target tissue, reference standard, ML architecture, validation scheme, performance metrics, and risk-of-bias items. Quantitative values were checked against the primary reports. When an abstract and table disagreed, we used the tabulated value; two discrepancies were resolved this way (see Section 3).

2.4. Risk of Bias and Synthesis

We assessed the five studies in the accuracy pool with QUADAS-2 and added ML-specific questions on subject-level partitioning, external validation, living-human data, thresholds/hyperparameters, unit independence, and blinded independent reference standards. Overall ratings followed QUADAS-2 roll-up rules. The six ML-specific signaling items, along with the adapted or limited appraisals of the regression and physiological studies, are reported.
We grouped studies by clinical use case. Classification accuracy was pooled using the Freeman–Tukey double-arcsine transformation and REML random-effects models, and reported as back-transformed accuracy, 95% CI, 95% PI, I2, τ2, and Cochran’s Q. Subgroup analyses were descriptive. Robustness analyses included leave-one-out re-estimation, denominator-standardized analyses, and a Burström-alternative input. We did not perform publication-bias tests because fewer than 10 studies make funnel-plot asymmetry testing unreliable, consistent with the Cochrane Handbook and the exemplar [14]. Regression and perfusion outcomes were synthesized narratively. Analyses were performed in R 4.6.0 with metafor, and certainty was judged using a GRADE-informed diagnostic accuracy approach.

2.5. Unit of Analysis and Commensurability

Raw classification accuracy was the only performance metric reported by all five pooled studies, so it was the only measure we could synthesize. No pooled study reported an AUC, and balanced accuracy and Cohen’s κ were reported too inconsistently to allow pooling. Because the classification tasks ranged from binary discrimination to six-class tissue labeling, raw accuracy is not fully comparable across studies; the chance baseline changes with the number of classes. We explored this limitation with chance-corrected accuracy, calculated as (accuracy − 1/C)/(1 − 1/C), where C is the number of classes, and pooled the corrected values with the same Freeman–Tukey and REML model. The pooled value was 93.86 on the corrected scale (Supplementary Materials, Section S1.5). This is a 0 to 100 scale on which 0 represents chance performance, and 100 represents perfect performance, so it should not be read as raw accuracy. The correction also assumes a chance baseline of 1/C, which is valid only when classes are balanced. No included study established balanced class prevalence; where classes are imbalanced, the true baseline is higher, and this correction is optimistic.
The included studies also report performance at two very different levels. The lower level is the observation, meaning an individual spectrum, image patch, or image. The upper level is the subject, meaning the cadaver, animal, or donor from which those observations were drawn. Denominators of 64 to 48,000 observations came from far fewer subjects, so an observation-level denominator reflects sampling effort more than independent evidence. That distinction matters clinically: a surgeon acts on a patient, not on a spectrum. We therefore treated the subject as the clinically relevant unit and reported the subject-level estimate as co-primary alongside the observation-level pooled accuracy. The subject-level analysis preserves each study’s reported accuracy and reallocates it to a denominator of independent subjects; it re-expresses the same performance at the subject scale rather than measuring per-subject classification directly. Integer rounding places three of the five studies at exactly 100% under this construction (6 of 6, 6 of 6, 14 of 15, 8 of 8, and 5 of 6 subjects), so the estimate has a wide interval with an upper limit at the 100% ceiling. For Bayhaqi 2023, we counted the 15 independently split tissue samples rather than the 5 source animals. Counting the 5 animals instead would raise the subject-level estimate to 99.01%, making the reported approach the more conservative choice.
We also labeled each analysis by whether it was prespecified or added later. PROSPERO field 28 prespecified the leave-one-out analyses, including the leakage-free pool, the subject-level analysis, denominator-standardized re-analyses, a subject-level design-effect analysis, and a logit-normal binomial GLMM. Processing latency was prespecified as an additional outcome under PROSPERO field 25 and was summarized narratively rather than pooled. The spine-only pool, restricted core pool, within-use-case pooling, chance-corrected analysis, and design-effect scenarios under assumed intraclass correlations were post hoc and were requested during peer review. Full inputs, model outputs, and a plain-language interpretation of each analysis are provided in Supplementary Materials.

3. Results

3.1. Search and Selection

The database searches returned 4835 records. After 1845 duplicates were removed, 2990 unique records underwent title and abstract screening. Citation chaining identified two additional eligible studies, Bayhaqi 2023 [20] and Unal 2025 [21], which the database searches missed because they were indexed under long-bone rather than spine terms. Twenty-two reports underwent full-text review, and eight met the inclusion criteria (Figure 2).
The most common reason for exclusion was optical sensing without an ML/DL component. This group included three Swamy-group DRS reports [22,23,24], an angled-fiber DRS study [25], a Doppler PS-OCT needle-visualization study [26], and an OCT study of rat spinal cord [27]. Because no ML-based study evaluated spinal cord perfusion, we retained two non-ML perfusion studies for the pre-specified narrative extension. Full-text exclusions were categorized according to PRISMA 2020.

3.2. Study Characteristics

Table 1 and Table 2 summarize the eight included studies across four clinical use cases: pedicle trajectory and breach avoidance, bone quality and screw purchase, bone–dura or epidural-interface detection, and spinal cord perfusion or ischemia. Five studies reported classification accuracy and were included in the meta-analysis: Burström 2019, Kosik 2023, Bayhaqi 2023, Wang 2022, and Wang 2024 [20,28,29,30,31]. All were single-group proof-of-concept studies without a comparison arm. Four used porcine tissue and one used a human cadaveric spine. Two studies used non-spine bone as a surrogate for spinal applications: Bayhaqi 2023 [20] examined porcine femur, and Unal 2025 [21] examined human femoral cortical bone. Sample sizes ranged from 64 to 48,000 analyzable spectra, image patches, or images, with reported accuracies ranging from 91.53% to 97.6%. None of the pooled studies reported AUC, and only Burström 2019 [28] provided clearly defined subject-level sensitivity and specificity.

3.3. Pedicle Trajectory and Breach Avoidance

Burström 2019 studied DRS in six human cadaveric spines with 45 CBCT-verified pedicle-screw breaches and 1615 spectra [28]. Using leave-one-cadaver-out validation, a support vector machine distinguished cancellous from cortical bone with 97.6% accuracy, 98.3% sensitivity, and 97.7% specificity. This was the only leakage-free human-cadaver study to report sensitivity and specificity. The tissue, however, was non-perfused, and prior non-ML DRS studies from the same group suggest that this use case is more mature than most others reviewed. Kosik 2023 evaluated near-infrared Raman spectroscopy in six swine for six-class tissue classification and binary bone-versus-soft-tissue discrimination [29]. The binary classifier reached 100% accuracy, and the six-class model reached 96.9%, but the ex vivo holdout appeared to be spectrum-level rather than subject-level, creating a substantial leakage concern. Its in situ and in vivo drilling phases were qualitative only, and Raman detection times exceeded the approximate 100 ms target for real-time guidance.

3.4. Bone Quality and Screw Purchase

Unal 2025 used Raman spectroscopy in 118 ex vivo human femoral cortical-bone specimens to predict fracture-toughness properties against mechanical R-curve testing [21]. The best model explained a substantial share of variance in crack-propagation energy (J-integral R2 = 0.737, XGBoost) and a more modest share in crack-initiation toughness (K_init R2 = 0.623, Extra Trees). The abstract attributed the best K_init model to an ensemble, but the full-text tables did not support this attribution. Because these outcomes were continuous regression metrics rather than classification proportions, Unal 2025 was summarized separately and excluded from the accuracy pool.

3.5. Bone–Dura and Epidural-Interface Detection

Three OCT studies addressed canal-boundary or epidural-interface detection. Bayhaqi 2023 trained DenseNet121 on ex vivo porcine femurs to classify bone versus marrow during closed-loop laser osteotomy [20]. It reported 96.28% accuracy across 48,000 patches using a subject-level split, with 45.96 ms detection latency. Its main limitation was that single-tissue patch training did not reproduce the multilayer interface required clinically. Wang 2022 used forward-view endoscopic OCT and ResNet50 in eight porcine backbones, breaking epidural-layer recognition into four sequential binary classifiers with 96.65% cross-testing accuracy [30]. A separate Inception model estimated dura-to-probe distance with a mean absolute percentage error of 3.05% ± 0.55% and mean absolute error of 34.1 µm. Wang 2024 extended the approach with PS-OCT in six porcine backbones [30]. DOPU mode achieved the best cross-testing accuracy at 91.53%, with an epidural space F1 of 99.56%; multimode fusion was proposed only as future work.

3.6. Spinal Cord Perfusion and Ischemia

Two studies examined spinal cord perfusion, but neither used ML. Mainard (2022) demonstrated the feasibility of multiwavelength NIRS/PPG on the exposed spinal cord of one lesion-free pig, using analytical signal processing rather than classification [32]. Ren 2025 used laser speckle contrast imaging in 31 rabbits undergoing pedicle subtraction osteotomy and found significant perfusion reductions across surgical phases, including a fall from 519.22 ± 137.87 to 315.00 ± 50.24 perfusion units after osteotomy and to 237.73 ± 40.46 after dural removal [33]. Together with the excluded qualitative OCT spinal cord study [27], these studies define a clear gap: optical perfusion hardware is feasible, but no study has paired it with a trained diagnostic model.

3.7. Meta-Analysis of Classification Accuracy

The five pooled studies are shown in Table 2 and Figure 3. Two results are most useful for interpretation. First, the 95% prediction interval, which estimates the range a future comparable study might report, was 89.8–99.4%. Second, the co-primary subject-level estimate was 97.36% (95% CI 87.46–100.00%), with no residual heterogeneity (I2 = 0%, τ2 = 0.0000, Q(4) = 1.83, p = 0.767). This analysis preserves each study’s reported accuracy and reallocates it to a denominator of independent subjects; it does not measure per-subject classification directly. Integer rounding places three of the five studies at exactly 100%; their upper limit reaches the 100% ceiling, and the interval is wide because relatively few independent subjects support it. The observation-level random-effects pooled accuracy, calculated over spectra and image patches, was 95.9% (95% CI 93.5–97.7%). We interpret that value as a proof-of-concept summary of tissue separability rather than patient-level performance. Observation-level heterogeneity was near-total (I2 = 99.4%, τ2 = 0.0030, Q(4) = 262.6, p < 0.001), driven largely by very large imaging denominators that produced near-zero sampling variance. Raw accuracies themselves spanned only about six percentage points, and resetting denominators to independent subjects reduced I2 to 0%, showing how strongly the heterogeneity depends on clustered observations. Descriptive subgroup analyses were hypothesis-generating only: point-spectroscopy studies pooled at 97.8% versus 95.1% for cross-sectional imaging, and human-cadaver evidence was limited to Burström 2019. The between-group moderator test was not significant (Q = 0.83, df = 1, p = 0.361). Clinical use case, optical modality, and model class also produced the same partition of the five studies because spectroscopy studies used classical ML and imaging studies used CNNs. A use-case subgroup analysis therefore relabels the same split rather than isolating a true use-case effect (Sections S1.4 and S1.9).
The prespecified robustness analyses gave similar pooled estimates while making the uncertainty more visible. Leave-one-out estimates ranged from 95.4% to 96.7%, and equal-weight and denominator-capped analyses produced nearly identical results. The common-effect model generated an implausibly narrow CI because Bayhaqi 2023 [20] and Wang 2022 [30] dominated the weighting; random-effects weights were much more balanced. Other robustness analyses were concordant: a t-based prediction interval widened to 84.9–100%, Knapp–Hartung correction gave 95.9% accuracy (95% CI 92.7–98.2%), and a logit-normal binomial GLMM returned 96.0% (95% CI 93.8–97.4%). Restricting the pool to spine studies, which removed only Bayhaqi 2023 [20] because Unal 2025 reported a regression endpoint and was never pooled, yielded 95.78% (95% CI 92.58–98.14%) with a prediction interval of 88.23–99.70% (k = 4; Section S1.1). Removing Kosik 2023 [29], the study with likely spectrum-level leakage, yielded 95.70% (95% CI 92.98–97.79%) with a prediction interval of 88.93–99.40% (k = 4; Section S1.2), indicating that this study did not drive the pooled value or prediction interval. Removing both gave 95.52% (95% CI 91.46–98.33%) with a prediction interval of 86.42–99.79% (k = 3; Section S1.3). Chance-corrected pooling returned 93.86% (95% CI 91.09–96.16%) on a 0 to 100 scale in which zero represents chance performance (Section S1.5). That value is not an accuracy, and the correction is optimistic because it assumes balanced classes. Across assumed intraclass correlations of 0.01 to 0.50, design-effect scenarios moved the pooled point estimate by no more than 1.59 percentage points. I2, however, fell from 99.4% to 82.9% at an assumed correlation of 0.01 and to 0% from 0.10 upward, again showing how much the apparent heterogeneity depends on treating every spectrum and image patch as independent (Section S1.6). The prespecified subject-level co-primary analysis returned 97.36% (95% CI 87.46–100.00%) with I2 = 0% (Q(4) = 1.83, p = 0.767; Section S1.7). Because the tasks ranged from binary to multiclass classification, balanced accuracy or Cohen’s κ would be preferable to raw accuracy if future studies report them consistently.

3.8. Risk of Bias and Certainty

QUADAS-2 judgments are summarized in Table 3, the ML-specific signaling items in Table 4, and both are shown in Figure 4. Kosik 2023 was rated as high risk of bias because its ex vivo split likely allowed spectrum-level leakage [28]. Burström 2019, Bayhaqi 2023, Wang 2022, and Wang 2024 were rated unclear overall because key safeguards, such as prespecified thresholds, blinded reference standards, or fully independent testing, were incompletely reported [20,28,30,31]. All five pooled studies lacked external validation and living-human intraoperative testing. Applicability concerns were greater for the four animal studies and lower for the human-cadaver study. The ML-specific assessment also highlighted a recurring problem: studies often reported thousands of spectra or image patches even though those observations came from only a small number of biological subjects. That structure increases the risk of unit-of-analysis inflation when observations are treated as independent.
Using a GRADE-informed diagnostic accuracy framework, certainty was very low for the observation-level pooled classification estimate, the co-primary subject-level estimate, and the prediction interval. The main reasons were indirectness, small numbers of independent subjects, extreme observation-level heterogeneity, risk of data leakage, and the absence of external or live-human validation.

4. Discussion

The main signal is encouraging, but the evidence is much less mature than the headline accuracy suggests. Across the five classification studies, reported accuracy clustered near 96%. The 95% prediction interval of 89.8–99.4% and the subject-level estimate of 97.36% (95% CI 87.46–100.00%) are more clinically informative because they better reflect uncertainty and account for the number of independent subjects. The observation-level pooled value of 95.9% summarizes how well spectra and image patches were separated under study conditions; it is not evidence of patient-level or clinical-grade performance. That distinction matters because the entire pooled evidence base comes from five single-group proof-of-concept studies, four in porcine tissue and none in living humans, with small subject counts, extreme observation-level heterogeneity, and meaningful risk of optimistic bias. Taken together, the literature shows that optical signals paired with ML can separate spine-relevant tissues under controlled preclinical conditions. It does not yet show that these systems are ready for intraoperative use.
Generalizability is the main limitation. Every pooled study evaluated models on the same type of tissue used for development, usually ex vivo porcine bone or spine tissue, or non-perfused human cadaveric tissue. None provided external validation in an independent site or cohort, and none tested classification in living, perfused human tissue during surgery (Table 4). Accuracy in these experiments therefore reflects internal tissue separability under controlled conditions. The operative field is different: bleeding, motion, cautery artifact, fluid, tissue deformation, and surgeon–instrument interaction can all change the optical signal.
The live operative field is a much harder optical environment than ex vivo tissue. Burström 2019 fitted hemoglobin and methemoglobin into its chromophore model but excluded them from the final analysis because the cadaveric field had no perfusion [28]. The classifier therefore never had to contend with the absorber that would dominate much of the visible and near-infrared signal in a perfused pedicle, leaving its effect on breach warning unknown. Raman spectroscopy faces a related problem because its signal is weak relative to background from blood; Kosik 2023 reported greater inter-measurement variance during the in vivo phase and attributed it to bleeding interference [29]. Coherence-based methods have a different vulnerability: motion decorrelates speckle, so respiration, cardiac pulsation, retraction, and drilling can degrade the contrast on which the classifier depends. Cautery plume, irrigation, and bone dust can add further scattering or absorption at the probe tip. These concerns already appear in the live perfusion literature. Mainard 2022 described heavily noisy raw signals and motion artifact on the exposed cord [32], while Ren 2025 cited blood-artifact interference and limited laser penetration [33]. Ex vivo and cadaveric accuracy therefore establishes tissue separability under favorable optical conditions. Clinical safety will depend on performance under the very confounders those experimental designs remove.
The heterogeneity requires similar caution. The observation-level confidence interval was narrow (93.5% to 97.7%), yet I2 was approximately 99%, making the pooled average incomplete on its own. Much of that heterogeneity comes from very large spectrum-level denominators nested within a small number of subjects, which makes within-study variance look artificially small. The design-effect scenarios make this dependence clear. With an assumed intraclass correlation of 0.01, I2 fell from 99.4% to 82.9%; from an assumed correlation of 0.10 upward, it fell to 0%, while the pooled point estimate moved by at most 1.59 percentage points across the full grid. In other words, the near-total heterogeneity depends heavily on treating each spectrum or image patch as independent. That is why the prediction interval and subject-level estimate carry more weight in our interpretation. The reported accuracies themselves covered a relatively narrow range, while the 95% prediction interval remained wide (89.8% to 99.4%) and the subject-level model showed no residual heterogeneity. These results are more useful for understanding what a future comparable study might report than the observation-level pooled value alone.
Several study-design choices can also make performance look better than it is. Kosik’s likely spectrum-level split is the clearest example: correlated spectra from the same animal may have appeared in both training and test sets [29]. The same problem can arise whenever many measurements are sampled from only a few subjects. Nominal denominators ranged from 64 to 48,000 observations, but these observations were clustered within only 6 to 15 independent subjects or samples per study; thus, the effective sample size was much smaller. Near-ceiling accuracy on relatively simple tasks, such as bone versus marrow or bone versus soft tissue, also leaves little room to distinguish between methods. Finally, the lack of prospective protocols, external validation, and standardized reporting makes selective analysis and optimistic reporting difficult to rule out, consistent with concerns seen more broadly in deep-learning diagnostic imaging [14]. Leave-one-out and denominator-standardized analyses showed that the pooled average was statistically stable. They cannot correct the underlying preclinical, single-cohort, and potentially leakage-prone study designs.
The subgroup results should be read as descriptive. Point spectroscopy pooled higher than cross-sectional imaging (97.8% versus 95.1%), but modality and model class were completely confounded: the spectroscopy studies used classical SVMs, whereas the imaging studies used CNNs. The comparison was underpowered and non-significant, and it cannot separate an optical-modality effect from an algorithm effect. These data do not establish that spectroscopy outperforms imaging or that classical ML outperforms deep learning.
Spinal cord perfusion and ischemia is a separate, clinically urgent application. Existing NIRS/PPG and LSCI studies already generate quantitative optical signals related to cord perfusion [32,33]. Neither study used a trained ML model, and both treated ML as future work. This leaves a practical research gap: perfusion-plus-ML systems need to distinguish true ischemia from bleeding, motion artifact, and normal physiologic fluctuation in real time.
Clinical translation will require a different evidentiary standard from the current proof-of-concept literature. Future studies should separate training and testing at the subject level, report the true number of biological subjects and independent test units, and validate models externally before making claims about clinical performance. In vivo human studies are essential because the operative environment introduces confounders that ex vivo tissue cannot reproduce. Speed also matters. Systems intended to warn a surgeon or control a robot must respond quickly enough to change an action before injury occurs. Closed-loop applications add another requirement: outputs must be interpretable and auditable if a model is expected to stop a tool, redirect a trajectory, or warn of impending breach or ischemia. These needs fit the broader enthusiasm for optical sensing and AI in spine surgery [1,2,8,9], while also explaining why the current evidence is not ready for clinical adoption. Related work in spine and musculoskeletal AI points to the same problem. Recent scoping reviews emphasize that model performance has limited clinical value without external validation, calibration, interpretability, workflow integration, and prospective evidence of utility [34,35,36,37,38,39,40]. Large-cohort prediction and comparative-effectiveness studies likewise show that decision-support tools eventually have to be judged against meaningful outcomes such as discharge disposition, reoperation, revision, longitudinal outcomes, and treatment durability [41,42,43,44,45]. Optical-ML systems will also need to fit with bone-quality assessment, implant and construct design, surgeon education, and embedded feedback systems before they can move from tissue classification to safe intraoperative decision support [45,46,47,48]. Musculoskeletal deep-learning studies outside the spine show that CNN-based orthopedic image interpretation is feasible, but they reinforce the same need for transparent validation before deployment [49,50,51,52,53].
Speed is another practical constraint that most studies did not fully report. A warning that arrives after the drill has advanced is not a useful warning, and the working target for drilling feedback is on the order of 100 ms [29]. Only two of the five pooled studies reported a time that included both acquisition and inference (Section S1.8). Bayhaqi 2023 [20] reported 45.96 ms for OCT detection, including acquisition and inference, and was the only study with an end-to-end value below the 100 ms target [20]. Kosik 2023 reported Raman integration times of 0.4 to 20 s, with detection typically taking about 1 s, which is roughly an order of magnitude too slow for continuous drilling feedback [29]. The other three studies reported only part of the timing pathway. Burström 2019 reported 50 to 150 ms for spectral acquisition, but no inference time, so total latency is unknown [28]. Wang 2022 reported 13 ms of inference per image, and Wang 2024 reported 0.0314 s per image, but their front-end values of 200 kHz and 5.5 to 76 kHz are A-scan line rates rather than per-image acquisition times [30,31]. Total per-frame latency therefore cannot be calculated from the published data. Future studies should report acquisition, transfer, and inference time separately, then report the end-to-end total. They should also test the solutions already proposed in this literature rather than treat them as solved: shorter Raman integration times and band-limited feature sets [29], lighter models through weight pruning or knowledge distillation [30], and hardware acceleration for inference [20].
This review has several limitations. Only five studies contributed to the pooled accuracy analysis. Sensitivity analyses with capped denominators and an alternative Burström denominator did not materially change the pooled estimate, although author-supplied confusion matrices would allow more precise re-analysis. Several analyses added during peer review were not prespecified, including the spine-only pool, restricted core pool, within-use-case pooling, chance-corrected analysis, and design-effect scenarios; therefore, we report them as post hoc rather than protocol-driven confirmations. The chance-corrected analysis assumes a chance baseline of one divided by the number of classes, which is valid only with balanced class prevalence. No included study established balanced prevalence, so the correction is optimistic when classes are imbalanced. The design-effect analysis is also a scenario exercise rather than an empirical estimate because none of the included studies reported an intraclass correlation, and one cannot be calculated from the published summaries. The assumed values of 0.01 to 0.50 are intended to bracket a plausible range, not to estimate the clustering actually present.
Four priorities stand out. First, future studies should use subject-wise splitting and report subject-level performance so that leakage and unit-of-analysis inflation are minimized. Second, models need external validation, followed by testing in living human operative tissue. Third, AI reporting should follow contemporary standards such as TRIPOD-AI and CLAIM, with prospective protocols and transparent analysis plans. Fourth, spinal cord perfusion monitoring deserves development as a dedicated optical-ML application by pairing NIRS/PPG or LSCI with trained ischemia classifiers. These steps would move the field from promising preclinical accuracy toward evidence that can support clinical translation.

5. Conclusions

Machine learning-augmented optical tissue sensing is promising, with consistently high reported accuracy for distinguishing bone and adjacent tissues in spine-relevant settings. The most clinically informative summaries are the 95% prediction interval (89.8% to 99.4%) and the subject-level estimate (97.36%, 95% CI 87.46% to 100.00%), which re-expresses reported performance using independent subjects as the denominator. The observation-level pooled accuracy of 95.9% is best viewed as a proof-of-concept measure of separability across spectra and image patches. The evidence base is still small, heterogeneous, and almost entirely preclinical, and it is vulnerable to optimistic bias from data leakage, unit-of-analysis inflation, and single-cohort validation. Before this technology can be translated responsibly into spine surgery, it needs subject-level evaluation, external validation, and in vivo testing in living human tissue under contemporary AI reporting standards. Spinal cord perfusion and ischemia monitoring is the clearest actionable gap: optical hardware already exists, but trained ML models have not yet been applied.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/opt7050062/s1, Supplementary S1: sensitivity and robustness analyses, Sections S1.1 to S1.11.

Author Contributions

S.G.S. and R.A.P. aided in conceptualization. S.G.S., R.A.P., R.K., N.P., Z.G.S., R.Z., and S.S. aided in methodology and the preparation of the original manuscript draft. S.G.S., R.A.P., and R.K. aided in formal analysis, data curation, and manuscript review and editing. A.T. (Alireza Tavakkoli), E.W., J.O., S.G., A.T. (Ajay Tripuraneni), and J.R. aided in study supervision and critical manuscript revision. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data supporting this review are contained within the article and in Supplementary Materials.

Acknowledgments

During the preparation of this manuscript/study, the authors used ChatGPT 5.4 to improve grammar accuracy and refine style. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Guha, D.; Yang, V.X.D. Perspective review on applications of optics in spinal surgery. J. Biomed. Opt. 2018, 23, 060601. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Losch, M.S.; Heintz, J.D.; Edström, E.; Elmi-Terander, A.; Dankelman, J.; Hendriks, B.H.W. Fiber-Optic Pedicle Probes to Advance Spine Surgery through Diffuse Reflectance Spectroscopy. Bioengineering 2024, 11, 61. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Fatima, N.; Massaad, E.; Hadzipasic, M.; Shankar, G.M.; Shin, J.H. Safety and accuracy of robot-assisted placement of pedicle screws compared to conventional free-hand technique: A systematic review and meta-analysis. Spine J. 2021, 21, 181–192. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Naik, A.; Smith, A.D.; Shaffer, A.; Krist, D.T.; Moawad, C.M.; MacInnis, B.R.; Teal, K.; Hassaneen, W.; Arnold, P.M. Evaluating robotic pedicle screw placement against conventional modalities: A systematic review and network meta-analysis. Neurosurg. Focus 2022, 52, E10. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Li, H.M.; Zhang, R.J.; Shen, C.L. Accuracy of pedicle screw placement and clinical outcomes of robot-assisted technique versus conventional freehand technique in spine surgery from nine randomized controlled trials: A meta-analysis. Spine 2020, 45, E111–E119. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Wei, F.-L.; Gao, Q.-Y.; Heng, W.; Zhu, K.-L.; Yang, F.; Du, M.-R.; Zhou, C.-P.; Qian, J.-X.; Yan, X.-D. Association of robot-assisted techniques with the accuracy rates of pedicle screw placement: A network pooling analysis. eClinicalMedicine 2022, 48, 101421. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Kumar, R.; Bouras, A.; Phadke, R.; Salman, S.; Kaur, H.; Sporn, K.; Damian, A.; Paladugu, P.; Vaja, S.; Tavakkoli, A.; et al. Dynamic spine stabilization through mechanically tuned constructs and embedded biomechanical feedback systems: A narrative review. Discov. Sens. 2026, 2, 26. [Google Scholar] [CrossRef] [Scilit]
  8. Kalanjiyam, G.P.; Chandramohan, T.; Raman, M.; Kalyanasundaram, H. Artificial intelligence: A new cutting-edge tool in spine surgery. Asian Spine J. 2024, 18, 458–471. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Kumar, R.; Gowda, C.; Sekhar, T.C.; Vaja, S.; Hage, T.; Sporn, K.; Waisberg, E.; Ong, J.; Zaman, N.; Tavakkoli, A. Advancements in Machine Learning for Precision Diagnostics and Surgical Interventions in Interconnected Musculoskeletal and Visual Systems. J. Clin. Med. 2025, 14, 3669. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Lopez, C.D.; Boddapati, V.; Lombardi, J.M.; Lee, N.J.; Mathew, J.; Danford, N.C.; Iyer, R.R.; Dyrszka, M.D.; Sardar, Z.M.; Lenke, L.G.; et al. Artificial learning and machine learning applications in spine surgery: A systematic review. Glob. Spine J. 2022, 12, 1561–1572. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Lee, N.J.; Lombardi, J.M.; Lehman, R.A. Artificial intelligence and machine learning applications in spine surgery. Int. J. Spine Surg. 2023, 17, S18–S25. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Hornung, A.L.; Hornung, C.M.; Mallow, G.M.; Barajas, J.N.; Rush, A.; Sayari, A.J.; Galbusera, F.; Wilke, H.-J.; Colman, M.; Phillips, F.M.; et al. Artificial intelligence in spine care: Current applications and future utility. Eur. Spine J. 2022, 31, 2057–2081. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Tragaris, T.; Benetos, I.S.; Vlamis, J.; Pneumaticos, S.G. Machine learning applications in spine surgery. Cureus 2023, 15, e48078. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Aggarwal, R.; Sounderajah, V.; Martin, G.; Ting, D.S.W.; Karthikesalingam, A.; King, D.; Ashrafian, H.; Darzi, A. Diagnostic accuracy of deep learning in medical imaging: A systematic review and meta-analysis. npj Digit. Med. 2021, 4, 65. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Ly, C.O.; Unnikrishnan, B.; Tadic, T.; Patel, T.; Duhamel, J.; Kandel, S.; Moayedi, Y.; Brudno, M.; Hope, A.; Ross, H.; et al. Shortcut learning in medical AI hinders generalization: Method for estimating AI model generalization without external data. npj Digit. Med. 2024, 7, 124. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Maleki, F.; Ovens, K.; Gupta, R.; Reinhold, C.; Spatz, A.; Forghani, R. Generalizability of machine learning models: Quantitative evaluation of three methodological pitfalls. Radiol. Artif. Intell. 2023, 5, e220028. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Varoquaux, G.; Cheplygina, V. Machine learning for medical imaging: Methodological failures and recommendations for the future. npj Digit. Med. 2022, 5, 48. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Gidwani, M.; Chang, K.; Patel, J.B.; Hoebel, K.V.; Ahmed, S.R.; Singh, P.; Fuller, C.D.; Kalpathy-Cramer, J. Inconsistent partitioning and unproductive feature associations yield idealized radiomic models. Radiology 2023, 307, e220715. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Salman, S.; Phadke, R.; Kumar, R.; Panwalker, N. Machine Learning-Augmented Optical Tissue Sensing for Intraoperative Guidance in Spine Surgery: A Systematic Review and Meta-Analysis. PROSPERO 2026 CRD420261447883. Available online: https://www.crd.york.ac.uk/PROSPERO/view/CRD420261447883 (accessed on 11 July 2026).
  20. Bayhaqi, Y.A.; Hamidi, A.; Navarini, A.A.; Cattin, P.C.; Canbaz, F.; Zam, A. Real-time closed-loop tissue-specific laser osteotomy using deep-learning-assisted optical coherence tomography. Biomed. Opt. Express 2023, 14, 2986–3002. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Unal, M.; Unlu, R.; Uppuganti, S.; Nyman, J.S. Prediction of biomechanical properties of ex vivo human femoral cortical bone using Raman spectroscopy and machine learning algorithms. Bone Rep. 2025, 26, 101870. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Swamy, A.; Burström, G.; Spliethoff, J.W.; Babic, D.; Reich, C.; Groen, J.; Edström, E.; Terander, A.E.; Racadio, J.M.; Dankelman, J.; et al. Diffuse reflectance spectroscopy, a potential optical sensing technology for the detection of cortical breaches during spinal screw placement. J. Biomed. Opt. 2019, 24, 1–11. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Swamy, A.; Spliethoff, J.W.; Burström, G.; Babic, D.; Reich, C.; Groen, J.; Edström, E.; Elmi-Terander, A.; Racadio, J.M.; Dankelman, J.; et al. Diffuse reflectance spectroscopy for breach detection during pedicle screw placement: A first in vivo investigation in a porcine model. Biomed. Eng. Online 2020, 19, 47. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Swamy, A.; Burström, G.; Spliethoff, J.W.; Babic, D.; Ruschke, S.; Racadio, J.M.; Edström, E.; Elmi-Terander, A.; Dankelman, J.; Hendriks, B.H.W. Validation of diffuse reflectance spectroscopy with magnetic resonance imaging for accurate vertebral bone fat fraction quantification. Biomed. Opt. Express 2019, 10, 4316–4328. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Losch, M.S.; Kardux, F.; Dankelman, J.; Hendriks, B.H.W. Diffuse reflectance spectroscopy of the spine: Improved breach detection with angulated fibers. Biomed. Opt. Express 2023, 14, 739–750. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Harper, D.J.; Kim, Y.; Gómez-Ramírez, A.; Vakoc, B.J. Needle guidance with Doppler-tracked polarization-sensitive optical coherence tomography. J. Biomed. Opt. 2023, 28, 102910. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Giardini, M.E.; Zippo, A.G.; Valente, M.; Krstajic, N.; Biella, G.E. Electrophysiological and Anatomical Correlates of Spinal Cord Optical Coherence Tomography. PLoS ONE 2016, 11, e0152539. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Burström, G.; Swamy, A.; Spliethoff, J.W.; Reich, C.; Babic, D.; Hendriks, B.H.W.; Skulason, H.; Persson, O.; Terander, A.E.; Edström, E. Diffuse reflectance spectroscopy accurately identifies the pre-cortical zone to avoid impending pedicle screw breach in spinal fixation surgery. Biomed. Opt. Express 2019, 10, 5905–5920. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Kosik, I.; Dallaire, F.; Pires, L.; Tran, T.; Leblond, F.; Wilson, B. Preclinical evaluation of Raman spectroscopy for pedicular screw insertion surgical guidance in a porcine spine model. J. Biomed. Opt. 2023, 28, 057003. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Wang, C.; Calle, P.; Reynolds, J.C.; Ton, S.; Yan, F.; Donaldson, A.M.; Ladymon, A.D.; Roberts, P.R.; de Armendi, A.J.; Fung, K.-M.; et al. Epidural anesthesia needle guidance by forward-view endoscopic optical coherence tomography and deep learning. Sci. Rep. 2022, 12, 9057. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Wang, C.; Liu, Y.; Calle, P.; Li, X.; Liu, R.; Zhang, Q.; Yan, F.; Fung, K.; Conner, A.K.; Chen, S.; et al. Enhancing epidural needle guidance using a polarization-sensitive optical coherence tomography probe with convolutional neural networks. J. Biophotonics 2024, 17, e202300330. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Mainard, N.; Tsiakaka, O.; Li, S.; Denoulet, J.; Messaoudene, K.; Vialle, R.; Feruglio, S. Intraoperative Optical Monitoring of Spinal Cord Hemodynamics Using Multiwavelength Imaging System. Sensors 2022, 22, 3840. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Ren, Z.; Wang, J.; Ye, X.; Ma, Y. Real-time monitoring of spinal cord hemodynamics with laser speckle contrast imaging during pedicle subtraction osteotomy in rabbits. Front. Surg. 2025, 12, 1578420. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Salman, S.; Phadke, R.; Kumar, R.; Momin, A.; Tavakkoli, A. Risk prediction in spine surgery: A scoping review of traditional models, artificial intelligence, and the challenge of clinical translation. Spine Deform. 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Salman, S.G.; Phadke, R.; Kumar, R.; Zaman, N.; Tavakkoli, A. Digital twins and multimodal artificial intelligence in spine care: A scoping review of concepts, evidence, and translational barriers. Spine Deform. 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Collins, G.S.; Dhiman, P.; Ma, J.; Schlussel, M.M.; Archer, L.; Van Calster, B.; E Harrell, F.; Martin, G.P.; Moons, K.G.M.; van Smeden, M.; et al. Evaluation of clinical prediction models (part 1): From development to external validation. BMJ 2024, 384, e074819. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Salman, S.G.; Phadke, R.; Kumar, R.; Momin, A.; Tavakkoli, A. Response to: Comment on ‘Risk prediction in spine surgery: A scoping review of traditional models, artificial intelligence, and the challenge of clinical translation’. Spine Deform. 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Kelly, C.J.; Karthikesalingam, A.; Suleyman, M.; Corrado, G.; King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 2019, 17, 195. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Salman, S.G.; Phadke, R.A.; Kumar, R.; Zaman, N.; Tavakkoli, A. Response to: Comment on ‘digital twins and multimodal artificial intelligence in spine care: A scoping review of concepts, evidence, and translational barriers’. Spine Deform. 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  40. Vasey, B.; Nagendran, M.; Campbell, B.; Clifton, D.A.; Collins, G.S.; Denaxas, S.; Denniston, A.K.; Faes, L.; Geerts, B.; Ibrahim, M.; et al. Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ 2022, 377, e070904. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Salman, S.G.; Phadke, R.; Carlin, T.; Rana, A.; Dawson, J.R.; Fitzgerald, C.A.; Seger, C.P.; Zielinski, M.D.; Dumas, R.P. The fracture orthopedic risk of non-home discharge (FORD) score: A novel bedside predictive tool for non-home discharge in orthopedic trauma patients. Injury 2026, 57, 113301. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Salman, S.G.; Phadke, R.; Kumar, R.; Gill, K.; Vaja, S.; Lee, N.J.; Bono, C. Long-Term Outcomes of Lumbar Total Disc Arthroplasty and Hybrid Constructs: A Systematic Review. Spine J. 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  43. Phadke, R.; Salman, S.; Kumar, R.; Paidisetty, V.; Matthews, B.; Srinivas, R.; Vaja, S.; Lee, N.J. Endoscopic and percutaneous minimally invasive repair of pars interarticularis defects: A systematic review of clinical outcomes. Spine Deform. 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Srinivas, R.; Phadke, R.; Salman, S.; Hazem, D.; Kaur, H.; Kumar, R.; Vaja, S.; Lee, N.J. Endoscopic versus open lumbar decompression: A retrospective cohort study of 31,000 patients with 90-day follow-up. Neurosurg. Rev. 2026, 49, 240. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  45. Phadke, R.A.; Salman, S.G.; Salman, Z.G.; Yedupati, S.M.; Ong, J.; Tavakkoli, A.; Galhotra, S.; Tripuraneni, A.; Rizkalla, J. Deep Learning-Based Multi-Class Pediatric Wrist Fracture Subtype Classification: A Pilot Study Comparing Convolutional Neural Network Architectures. J. Imaging 2026, 12, 307. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Salman, S.G.; Phadke, R.; Burnett, J.; Walsh, J. Sequential Versus Step-Therapy Approaches for Osteoporosis Management in Orthopedic Subspecialties. Curr. Osteoporos. Rep. 2026, 24, 19. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Kumar, R.; Phadke, R.; Salman, S. Advancing AI literacy in Canadian orthopedic education: A framework for equitable and inclusive training. Can. Med. Educ. J. 2025, 16, 36–38. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Phadke, R.; Salman, S.; Bansal, A.; Kumar, R.; Paladugu, P.; Ong, J.; Waisberg, E.; Lee, A.G. Musculoskeletal degeneration in space: Translational parallels with spaceflight-associated ocular pathology and surgical implications. Life Sci. Space Res. 2026. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  49. Tampu, I.E.; Eklund, A.; Haj-Hosseini, N. Inflation of test accuracy due to data leakage in deep learning-based classification of OCT images. Sci. Data 2022, 9, 580. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  50. Surya, A.; Salman, S.; Phadke, R.; Sporn, K.; Yaldo, L.; Kumar, R.; Paladugu, P.; Ong, J.; Waisberg, E.; Masalkhi, M.; et al. Perioperative ischemic optic neuropathy: A comprehensive review of anesthetic implications, hemodynamic pathophysiology, and risk stratification. Int. Ophthalmol. 2026, 46, 155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  51. Olczak, J.; Fahlberg, N.; Maki, A.; Razavian, A.S.; Jilert, A.; Stark, A.; Sköldenberg, O.; Gordon, M. Artificial intelligence for analyzing orthopedic trauma radiographs. Acta Orthop. 2017, 88, 581–586. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. Jones, R.M.; Sharma, A.; Hotchkiss, R.; Sperling, J.W.; Hamburger, J.; Ledig, C.; O’toole, R.; Gardner, M.; Venkatesh, S.; Roberts, M.M.; et al. Assessment of a deep-learning system for fracture detection in musculoskeletal radiographs. npj Digit. Med. 2020, 3, 144. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Langerhuizen, D.W.G.; Bulstra, A.E.J.; Janssen, S.J.; Ring, D.; Kerkhoffs, G.M.M.J.; Jaarsma, R.L.; Doornberg, J.N. Is deep learning on par with human observers for detection of radiographically visible and occult fractures of the scaphoid? Clin. Orthop. Relat. Res. 2020, 478, 2653–2659. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Conceptual framework for machine learning-augmented optical tissue sensing in spine surgery. The four upper panels each map one clinically relevant use case to its optical modality, intended intraoperative task, and ML readout, with arrows connecting each panel to the corresponding site on the lumbar spine. (A) Pedicle trajectory and breach avoidance, using diffuse reflectance spectroscopy and Raman spectroscopy to distinguish cancellous from cortical bone at the probe tip and keep the pilot hole inside the pedicle. (B) Bone quality and screw purchase, using Raman spectroscopy to estimate mineralization and collagen–matrix properties and to predict biomechanical and fracture-toughness metrics by regression. (C) Bone–dura and epidural-interface detection, using OCT and polarization-sensitive OCT with depth-resolved layered-tissue classification on B-scans as a breach warning. (D) Spinal cord perfusion and ischemia monitoring, using multiwavelength NIRS and laser speckle contrast imaging to map cord perfusion and hemodynamics; this panel is drawn with a dashed border and labeled as an evidence gap because the included perfusion studies applied optical monitoring without ML or DL. The lower panel illustrates the proposed intraoperative pipeline, in which an optical probe acquires spectra or images, ML/DL models convert these signals into tissue classifications, regression outputs, or perfusion maps, and the resulting readout informs intraoperative decision-making. Created in BioRender. Kumar, R. (2026) [7] https://BioRender.com/b6gpr3w.
Figure 1. Conceptual framework for machine learning-augmented optical tissue sensing in spine surgery. The four upper panels each map one clinically relevant use case to its optical modality, intended intraoperative task, and ML readout, with arrows connecting each panel to the corresponding site on the lumbar spine. (A) Pedicle trajectory and breach avoidance, using diffuse reflectance spectroscopy and Raman spectroscopy to distinguish cancellous from cortical bone at the probe tip and keep the pilot hole inside the pedicle. (B) Bone quality and screw purchase, using Raman spectroscopy to estimate mineralization and collagen–matrix properties and to predict biomechanical and fracture-toughness metrics by regression. (C) Bone–dura and epidural-interface detection, using OCT and polarization-sensitive OCT with depth-resolved layered-tissue classification on B-scans as a breach warning. (D) Spinal cord perfusion and ischemia monitoring, using multiwavelength NIRS and laser speckle contrast imaging to map cord perfusion and hemodynamics; this panel is drawn with a dashed border and labeled as an evidence gap because the included perfusion studies applied optical monitoring without ML or DL. The lower panel illustrates the proposed intraoperative pipeline, in which an optical probe acquires spectra or images, ML/DL models convert these signals into tissue classifications, regression outputs, or perfusion maps, and the resulting readout informs intraoperative decision-making. Created in BioRender. Kumar, R. (2026) [7] https://BioRender.com/b6gpr3w.
Optics 07 00062 g001
Figure 2. PRISMA flow diagram of study identification and selection. Database searching identified 4835 records across Embase, PubMed, Web of Science, Scopus, and IEEE Xplore. After removal of 1845 duplicates, 2990 records were screened by title and abstract, and 2968 were excluded. Twenty-two full-text reports were assessed for eligibility, of which 14 were excluded: optical sensing without an ML/DL component (n = 6), not spine or wrong target (n = 5), and other reasons (n = 3). Eight studies were included in the review: five studies in the classification-accuracy meta-analysis pool, one regression study, and two no-ML perfusion studies retained as a prespecified narrative extension (not PRISMA-DTA eligible). Citation chaining identified two reports that were included within the full-text assessment count. Created in BioRender. Kumar, R. (2026) [7] https://BioRender.com/eig566y.
Figure 2. PRISMA flow diagram of study identification and selection. Database searching identified 4835 records across Embase, PubMed, Web of Science, Scopus, and IEEE Xplore. After removal of 1845 duplicates, 2990 records were screened by title and abstract, and 2968 were excluded. Twenty-two full-text reports were assessed for eligibility, of which 14 were excluded: optical sensing without an ML/DL component (n = 6), not spine or wrong target (n = 5), and other reasons (n = 3). Eight studies were included in the review: five studies in the classification-accuracy meta-analysis pool, one regression study, and two no-ML perfusion studies retained as a prespecified narrative extension (not PRISMA-DTA eligible). Citation chaining identified two reports that were included within the full-text assessment count. Created in BioRender. Kumar, R. (2026) [7] https://BioRender.com/eig566y.
Optics 07 00062 g002
Figure 3. Forest plot of pooled classification accuracy. The plot shows back-transformed per-study accuracies with 95% confidence intervals for the five pooled studies [19,27,28,29,30], grouped into descriptive subgroup blocks (point spectroscopy versus cross-sectional imaging; equivalently classical ML versus deep learning). A weight column reports each study’s random-effects weight, and the plotted square scales accordingly. Those weights are near-balanced (10.2–23.6%), whereas a common-effect model would spread the same five studies across 0.1–50.8% and let two large-denominator studies dominate. Each block carries its own heterogeneity statistics beneath its subtotal diamond (point spectroscopy I2 = 0.00%, τ2 = 0.00000, Q(1) = 0.33, p = 0.5647; cross-sectional imaging I2 = 99.74%, τ2 = 0.00375, Q(2) = 260.29, p < 0.001); the subtotal diamonds show no prediction interval, because τ2 is zero in the point-spectroscopy block and a coinciding interval would read there as a finding rather than as an artifact of k = 2. Two summary rows sit below the blocks. The upper row gives the Freeman–Tukey double-arcsine random-effects (REML) pooled accuracy at the observation level, 95.9% (95% CI 93.5–97.7%), and the dashed whiskers with their printed end labels mark the 95% prediction interval of 89.81–99.37%; between-study heterogeneity was near-total (I2 = 99.4%, τ2 = 0.0030, Q(4) = 262.6, p < 0.001). The lower shaded row gives the co-primary subject-level estimate, which preserves each study’s reported accuracy and reallocates it to a denominator of independent subjects: 97.36% (95% CI 87.46–100.00%), I2 = 0.00%, τ2 = 0.00000, Q(4) = 1.83, p = 0.7674. Because τ2 equals zero in that model, its prediction interval coincides with its confidence interval, and its upper limit sits on the 100% ceiling. The prediction interval remains the primary interpretive quantity and the subject-level estimate stands beside it as co-primary, while the observation-level pooled value serves as a proof-of-concept summary over spectra and image patches rather than over patients. Per-study values are Freeman–Tukey back-transformed at the harmonic-mean sample size (271.1) and may therefore differ slightly from the raw accuracies in Table 2 (e.g., Kosik 96.9% raw vs. 96.4% as displayed [29]). The accuracy axis is truncated at 85%; no interval extends below it.
Figure 3. Forest plot of pooled classification accuracy. The plot shows back-transformed per-study accuracies with 95% confidence intervals for the five pooled studies [19,27,28,29,30], grouped into descriptive subgroup blocks (point spectroscopy versus cross-sectional imaging; equivalently classical ML versus deep learning). A weight column reports each study’s random-effects weight, and the plotted square scales accordingly. Those weights are near-balanced (10.2–23.6%), whereas a common-effect model would spread the same five studies across 0.1–50.8% and let two large-denominator studies dominate. Each block carries its own heterogeneity statistics beneath its subtotal diamond (point spectroscopy I2 = 0.00%, τ2 = 0.00000, Q(1) = 0.33, p = 0.5647; cross-sectional imaging I2 = 99.74%, τ2 = 0.00375, Q(2) = 260.29, p < 0.001); the subtotal diamonds show no prediction interval, because τ2 is zero in the point-spectroscopy block and a coinciding interval would read there as a finding rather than as an artifact of k = 2. Two summary rows sit below the blocks. The upper row gives the Freeman–Tukey double-arcsine random-effects (REML) pooled accuracy at the observation level, 95.9% (95% CI 93.5–97.7%), and the dashed whiskers with their printed end labels mark the 95% prediction interval of 89.81–99.37%; between-study heterogeneity was near-total (I2 = 99.4%, τ2 = 0.0030, Q(4) = 262.6, p < 0.001). The lower shaded row gives the co-primary subject-level estimate, which preserves each study’s reported accuracy and reallocates it to a denominator of independent subjects: 97.36% (95% CI 87.46–100.00%), I2 = 0.00%, τ2 = 0.00000, Q(4) = 1.83, p = 0.7674. Because τ2 equals zero in that model, its prediction interval coincides with its confidence interval, and its upper limit sits on the 100% ceiling. The prediction interval remains the primary interpretive quantity and the subject-level estimate stands beside it as co-primary, while the observation-level pooled value serves as a proof-of-concept summary over spectra and image patches rather than over patients. Per-study values are Freeman–Tukey back-transformed at the harmonic-mean sample size (271.1) and may therefore differ slightly from the raw accuracies in Table 2 (e.g., Kosik 96.9% raw vs. 96.4% as displayed [29]). The accuracy axis is truncated at 85%; no interval extends below it.
Optics 07 00062 g003
Figure 4. QUADAS-2 risk-of-bias and applicability assessment. The figure has four panels. (A) Risk of bias, a traffic-light plot of per-study judgments across the four QUADAS-2 domains (patient/sample selection, index test, reference standard, flow and timing) with an overall rating. (B) Applicability: the same plot across the three standard applicability domains plus the adapted flow-and-timing column, with an overall rating. (C) AI signaling questions: the six ML-specific items (subject-level split, external validation, human in vivo data, a priori thresholds, unit independence, blinded reference standard) rated Yes, No, or Unclear. (D) Summary: a stacked bar chart giving the proportion of judgments per domain across the five accuracy-pool studies [20,28,29,30,31]. Green denotes low risk, low concern, or Yes; yellow denotes unclear; and red denotes high risk, high concern, or No. Panels A and B correspond to Table 3, and panel C corresponds to Table 4. Kosik 2023 [29] is rated High overall for risk of bias because of data leakage; all five studies [20,28,29,30,31] share an absence of external validation and of living-human in vivo data, and applicability concerns are High for the four animal studies [20,29,30,31]. Created in BioRender. Kumar, R. (2026) [7] https://BioRender.com/m9ea5xw.
Figure 4. QUADAS-2 risk-of-bias and applicability assessment. The figure has four panels. (A) Risk of bias, a traffic-light plot of per-study judgments across the four QUADAS-2 domains (patient/sample selection, index test, reference standard, flow and timing) with an overall rating. (B) Applicability: the same plot across the three standard applicability domains plus the adapted flow-and-timing column, with an overall rating. (C) AI signaling questions: the six ML-specific items (subject-level split, external validation, human in vivo data, a priori thresholds, unit independence, blinded reference standard) rated Yes, No, or Unclear. (D) Summary: a stacked bar chart giving the proportion of judgments per domain across the five accuracy-pool studies [20,28,29,30,31]. Green denotes low risk, low concern, or Yes; yellow denotes unclear; and red denotes high risk, high concern, or No. Panels A and B correspond to Table 3, and panel C corresponds to Table 4. Kosik 2023 [29] is rated High overall for risk of bias because of data leakage; all five studies [20,28,29,30,31] share an absence of external validation and of living-human in vivo data, and applicability concerns are High for the four animal studies [20,29,30,31]. Created in BioRender. Kumar, R. (2026) [7] https://BioRender.com/m9ea5xw.
Optics 07 00062 g004
Table 1. Study and technical characteristics of the eight included studies are shown.
Table 1. Study and technical characteristics of the eight included studies are shown.
StudyPMIDCountryDesignSubjectsAnalyzable UnitsModalityAcquisition ParametersAnatomical TargetReference StandardML ArchitectureValidation Scheme
A. Pedicle trajectory and breach avoidance
Burström 2019 [28]31799054SwedenEx vivo human cadaver proof-of-concept study6 (4 usable for leave-one-out)1615 spectraDiffuse reflectance spectroscopy (DRS)400–1600 nm; halogen source; 2 fibers 1.22 mm apart; 50–150 ms acquisitionThoracic/lumbar pedicle (cortical-cancellous interface)CBCT anatomical labeling by a blinded physicianSVM (radial kernel; LIBSVM/e1071)Subject-level leave-one-cadaver-out (4 folds); + 66/33 spectrum-level holdout
Kosik 2023 [29]37265877CanadaPreclinical porcine study (ex vivo + in situ + in vivo)6 swine (3 ex vivo/2 in situ/1 in vivo)162 ex vivo spectra (+132 in situ/in vivo)Raman spectroscopy785 nm diode 100 mW; 400–2000 cm−1; ~1.8 cm−1 res; 0.5 mm spot; 0.4–20 s integrationPorcine vertebra/pedicle; bone vs. soft tissue & spinal cordExpert anatomical identification; fluoroscopy/3D-CT (in situ/in vivo)Linear SVM (L1-Lasso features + L2-Ridge)Ex vivo 60/40 holdout (pooled; likely spectrum-level); external in situ/in vivo
B. Bone quality/screw purchase
Unal 2025 [21]40917467USAEx vivo human cadaveric cortical-bone regression study118 donors (58 M/60 F; age 21–101)118 specimensRaman spectroscopy785 nm (58 donors, Xplora)/830 nm (60 donors, inVia) ~35 mW; 20× obj; 785–1800 cm−1Femoral mid-diaphysis cortical boneMechanical R-curve testing (SENB 3-point bending)SVR, XGBoost, Extra Trees, EnsembleSpecimen-level 80/20 (n = 94/24); inner 5-fold CV; no external validation
C. Bone–dura/epidural-interface detection
Bayhaqi 2023 [20]37342720SwitzerlandEx vivo porcine real-time closed-loop laser-osteotomy study5 pigs (15 samples; +6 for ablation eval)48,000 test patchesOptical coherence tomography (OCT)1310 nm, 61.5 nm BW; 104.17 kHz A-scan; 26.2 mm range; 26 um lat/18 um ax resPorcine femur (bone vs. bone marrow)Manual tissue labeling; micro-CT (ablation eval only)DenseNet121 (CNN)Subject-level split (2 train/1 val/2 test pigs)
Wang 2022 [30]35641505USAEx vivo porcine endoscopic-OCT + deep-learning study8 pigs40,000 OCT images (+24,000 for regression)Optical coherence tomography (OCT)1300 nm, 100 nm BW; 200 kHz A-scan; axial res 10.6 um; images cropped 181 × 241 pxEpidural space; 5 spinal tissue layers (needle path)Anatomical ID (2 raters) + H&E histologyResNet50 (classification); Inception (regression); Xception comparedSubject-level 8-fold cross-testing + nested CV
Wang 2024 [31]37833242USAEx vivo porcine PS-OCT + deep-learning study6 porcine backbone samples6000 DOPU images (24,000 total)Polarization-sensitive OCT (PS-OCT)1300 nm, 170 nm BW; axial res 5.5 um; 5.5–76 kHz; image 430 × 950 px, 3 um pixelEpidural space; 5 spinal tissue layersH&E histology + anatomical identificationResNet50 (CNN) per imaging modeSubject-level 6-fold cross-testing (leave-one-sample-out) + nested CV
D. Spinal cord perfusion/ischemia monitoring: prespecified narrative extension (not PRISMA-DTA eligible)
Mainard 2022 [32]35632249FranceIn vivo porcine device-feasibility study (no ML)1 pig (lesion-free)Continuous NIRS/PPG signals (1 pig)Near-infrared spectroscopy (NIRS) + PPGMulti-wavelength (4 light sources); transmission + reflection modesSpinal cord (L3 laminectomy; dural surface)None (proof-of-concept; no ground truth)None (analytical/statistical only)N/A (single-animal feasibility; no train/test)
Ren 2025 [33]40703431ChinaIn vivo rabbit PSO perfusion study (no ML)31 male New Zealand white rabbitsLSCI perfusion maps (31 rabbits, triplicate/phase)Laser speckle contrast imaging (LSCI)780 nm, 5 ms exposure; RFLSI III; >30 fpsSpinal cord & arteries (PSO)None (within-animal phase comparison)None (analytical/statistical only)N/A (no ML)
Notes: Studies are grouped by intended clinical use case, matching the Results section: (A) pedicle trajectory and breach avoidance, (B) bone quality and screw purchase, (C) bone–dura or epidural-interface detection, and (D) spinal cord perfusion or ischemia monitoring. The table summarizes each study’s design, sample and analyzable-unit structure, optical modality, acquisition parameters, anatomical target, reference standard, machine learning architecture, and validation approach. Mainard 2022 [32] and Ren 2025 [33] did not use machine learning and are reported as prespecified narrative extensions (not PRISMA-DTA eligible) because they address spinal cord perfusion monitoring, a clinically relevant optical-sensing application not yet linked to ML/DL validation in the included literature. Abbreviations: ax, axial; BW, bandwidth; CBCT, cone-beam computed tomography; CNN, convolutional neural network; CV, cross-validation; DOPU, degree of polarization uniformity; DRS, diffuse reflectance spectroscopy; ETR, extra trees regression; H&E, hematoxylin and eosin; LOO, leave-one-out cross-validation; LSCI, laser speckle contrast imaging; ML, machine learning; NIRS, near-infrared spectroscopy; OCT, optical coherence tomography; PPG, photoplethysmography; PS-OCT, polarization-sensitive optical coherence tomography; PSO, pedicle subtraction osteotomy; RS, Raman spectroscopy; SENB, single-edge-notched bend; SVM, support vector machine; SVR, support vector regression; XGB, extreme gradient boosting.
Table 2. Diagnostic and performance outcomes of the eight included studies are shown below.
Table 2. Diagnostic and performance outcomes of the eight included studies are shown below.
StudyEndpointPrimary TaskAccuracySensitivitySpecificityAUCOther MetricsIn PoolNotes
A. Pedicle trajectory and breach avoidance
Burström 2019 [28]ClassificationCancellous vs. cortical bone (binary breach warning)97.6%98.3%97.7% YesClean subject-level LOO; only human cohort
Kosik 2023 [29]Classification6-class tissue + binary bone-vs-soft/spinal cord100% binary; 96.9% 6-class (62/64) YesLikely spectrum-level split (leakage)
B. Bone quality/screw purchase
Unal 2025 [21]RegressionFracture toughness (Kinit & J-integral) J-integral R2 = 0.737 (XGBoost); Kinit R2 = 0.623 (Extra Trees)NoSpecimen-level split; no external validation
C. Bone–dura/epidural-interface detection
Bayhaqi 2023 [20]ClassificationBone vs. bone marrow (binary)96.28% YesSubject-level split; 48,000 test patches
Wang 2022 [30]Classification (+ regression)5-layer epidural via 4 sequential binary + dura-distance regression96.65% (ResNet50) Dura-distance MAPE 3.05% ± 0.55% (Inception)YesSubject-level cross-testing; Inception used for regression only
Wang 2024 [31]Classification5-class epidural tissue (PS-OCT)91.53% (DOPU) YesSubject-level cross-testing; porcine only, no human
D. Spinal cord perfusion/ischemia monitoring: prespecified narrative extension (not PRISMA-DTA eligible)
Mainard 2022 [32]Physiologic/no-MLSpinal cord oxygenation/perfusion monitoring NIRS/PPG feasibility, 1 pig, lesion-free (no accuracy)NoNo ML, evidence gap; 1 pig, lesion-free
Ren 2025 [33]Physiologic/no-MLSpinal cord perfusion across PSO stages Cord perfusion 519 → 315 PU across PSO stages (paired t-tests)NoNo ML, evidence gap; perfusion only
Note: Studies are grouped by clinical use case, matching Table 1 and the Results. Endpoint type distinguishes classification, regression, and physiologic studies without machine learning. Five classification studies contributed to the diagnostic accuracy meta-analysis pool: Burström 2019 [28], Kosik 2023 [29], Bayhaqi 2023 [20], Wang 2022 [30], and Wang 2024 [31]. Unal 2025 [21], the dura-distance regression endpoint in Wang 2022 [30], and the physiologic perfusion studies by Mainard 2022 [32] and Ren 2025 [33] were summarized descriptively and were not included in the pooled classification-accuracy analysis. Blank cells indicate that the metric was not reported and should not be interpreted as zero values. Only Burström 2019 [28] reported subject-level sensitivity and specificity. No included study reported an AUC. Kosik 2023 [29] reported near-ceiling performance, including 100% binary bone-versus-soft-tissue accuracy, but its ex vivo holdout split was not stated to be subject-level and may therefore be vulnerable to spectrum-level data leakage. AUC, area under the curve; DOPU, degree of polarization uniformity; LOO, leave-one-out cross-validation; MAPE, mean absolute percentage error; ML, machine learning; PSO, pedicle subtraction osteotomy; PU, perfusion units; R2, coefficient of determination; XGBoost, extreme gradient boosting. ResNet50, DenseNet121, and Inception refer to convolutional neural network architectures.
Table 3. QUADAS-2 risk of bias and applicability concerns in the five diagnostic accuracy studies.
Table 3. QUADAS-2 risk of bias and applicability concerns in the five diagnostic accuracy studies.
Study (Modality; Model)Patient Selection (RoB)Index Test (RoB)Reference Standard (RoB)Flow & Timing (RoB)Overall RoBPatient Selection (Applic.)Index Test (Applic.)Reference Standard (Applic.)Flow & Timing (Applic.), Adapted †Overall Applicability
Burström 2019 [28] (DRS; human cadaver; SVM)LowLowLowLowLowLowUnclearLowHighUnclear
Kosik 2023 [29] (Raman; porcine; SVM)UnclearHighUnclearUnclearHighHighHighUnclearHighHigh
Bayhaqi 2023 [20] (OCT; porcine; DenseNet121)LowLowUnclearLowUnclearHighHighLowHighHigh
Wang 2022 [30] (OCT; porcine; ResNet50)LowLowLowLowLowHighUnclearLowHighHigh
Wang 2024 [31] (PS-OCT; porcine; ResNet50)LowLowLowLowLowHighUnclearLowHighHigh
Note: The table reports the QUADAS-2 assessment of the five studies included in the classification-accuracy meta-analysis. The machine learning-specific signaling items that previously appeared in this table now appear in Table 4. Risk of bias was rated across the four standard QUADAS-2 domains: patient/sample selection, index test, reference standard, and flow/timing. Applicability was rated across the three standard QUADAS-2 domains: patient/sample selection, index test, and reference standard, with the review target defined as ML-augmented optical tissue sensing for real-time guidance in living-human spine surgery. Judgments are Low, High, or Unclear. Overall risk of bias was rated High if any domain was High, Unclear if any domain was Unclear and none was High, and Low only if all domains were Low. Overall applicability was assigned using the same rule across the three standard applicability domains. † Adapted column. Standard QUADAS-2 carries four risk-of-bias domains but only three applicability domains, and it assigns no applicability judgment to flow and timing. The flow and timing applicability column is therefore an adaptation introduced for this review and is not a standard QUADAS-2 judgment. It flags the ex vivo, cadaveric, or animal-model gap uniformly, and it is High for all five studies because every study measured tissue outside the live operative field. It is excluded from the overall applicability roll-up. Domain definitions: Patient/sample selection assessed whether subjects, specimens, animals, and sampled tissue sites could support inference to live human spine surgery. Index test assessed conduct and interpretation of the optical modality and ML classifier, including train/test separation, threshold or hyperparameter specification, leakage risk, and clinical deployability. Reference standard assessed the rigor, independence, and blinding of tissue ground truth, with histology considered strongest, followed by micro-CT/CBCT anatomy and expert visual identification. Flow/timing assessed whether all samples received the same reference standard, whether all eligible samples were analyzed, and whether index-test and reference-standard registration was adequate. Study-specific reasons for the judgments: Burström 2019 [28], a physician labeled the CBCT reference blind to the spectral readings; index-test applicability is Unclear because the 1.22 mm probe fiber spacing is wide relative to the 0 to 3 mm pre-cortical zone the device must warn on, and flow/timing applicability is High because the cadaveric field carried no perfusion. Kosik 2023 [29], index-test risk is High because the ex vivo 60/40 holdout was pooled across swine with no stated subject-level split (Table 4) and the 13-band feature set was selected to reach the >96% target; the reference standard was expert visual identification with blinding not reported; flow/timing risk is Unclear because the reference standard differed across the ex vivo, in situ, and in vivo phases and only the ex vivo phase was quantified; and applicability is High because the model is porcine and Raman detection takes about 1 s against a real-time target near 100 ms. Bayhaqi 2023 [20]: reference-standard risk is Unclear because tissue labels were defined by patch location on the same OCT images rather than by an independent blinded standard, with micro-CT reserved for ablation-damage evaluation; index-test applicability is High because the model was trained on single-tissue patches and is insensitive to the multilayered bone and marrow interface it must resolve clinically. Wang 2022 [30], H&E histology with anatomical identification supports a Low reference-standard risk although blinding is not reported; index-test applicability is Unclear because ex vivo endoscopic OCT showed high inter-subject variability, with one fold reaching only 67.3%; the pooled 96.65% accuracy is the ResNet50 result, the average of four sequential binary classifiers under subject-level 8-fold cross-testing, while the Inception architecture supplied the dura-to-probe distance regression, a continuous endpoint that contributed nothing to the pooled accuracy. Wang 2024 [31], corresponding H&E histology again supports a Low reference-standard risk with blinding not reported; the reported estimate comes from the DOPU channel, the best of four imaging modes, and index-test applicability is Unclear because a single input polarization state yields cumulative polarization parameters without local birefringence. Unal 2025 [21], Mainard 2022 [32], and Ren 2025 [33] received no QUADAS-2 traffic-light judgment because they did not report reference-standard-based classification accuracy. Table 4 records their adapted and out-of-scope appraisals, together with the machine learning signaling items for all eight included studies. Table 3 and Table 4 are both set in landscape orientation. Abbreviations: QUADAS-2, Quality Assessment of Diagnostic Accuracy Studies-2; RoB, risk of bias; Applic., applicability; ML, machine learning; DRS, diffuse reflectance spectroscopy; OCT, optical coherence tomography; PS-OCT, polarization-sensitive optical coherence tomography; DOPU, degree of polarization uniformity; SVM, support vector machine; CBCT, cone-beam computed tomography; H&E, hematoxylin and eosin; micro-CT, micro-computed tomography. DenseNet121 and ResNet50 are convolutional neural network (CNN) architectures.
Table 4. Machine learning-specific signaling items and adapted appraisals for the eight included studies.
Table 4. Machine learning-specific signaling items and adapted appraisals for the eight included studies.
Study (Modality; Model)Subject-Level SplitExternal ValidationHuman In Vivo DataA Priori Threshold/HyperparametersUnit IndependenceBlinded Independent Reference Standard
Burström 2019 [28] (DRS; human cadaver; SVM)YesNoNo (cadaver)YesNoYes
Kosik 2023 [29] (Raman; porcine; SVM)NoNoNo (porcine)UnclearNoUnclear
Bayhaqi 2023 [20] (OCT; porcine; DenseNet121)YesNoNo (porcine)YesNoUnclear
Wang 2022 [30] (OCT; porcine; ResNet50)YesNoNo (porcine)YesNoUnclear
Wang 2024 [31] (PS-OCT; porcine; ResNet50)YesNoNo (porcine)YesNoUnclear
Unal 2025 [21] (Raman; human femur ex vivo; XGBoost/Extra Trees)Yes (specimen-level)NoNo (ex vivo)AdaptedYesAdapted
Mainard 2022 [32] (NIRS/PPG; in vivo porcine; no ML)Out of scopeOut of scopeOut of scopeOut of scopeOut of scopeOut of scope
Ren 2025 [33] (LSCI; in vivo rabbit; no ML)Out of scopeOut of scopeOut of scopeOut of scopeOut of scopeOut of scope
Note: The table reports the six machine learning-specific signaling items recorded for every included study, together with the adapted appraisal of Unal 2025 [21] and the out-of-scope appraisal of the two perfusion studies. QUADAS-2 risk-of-bias and applicability judgments for the five diagnostic accuracy studies appear in Table 3. Item definitions: Subject-level split asks whether the train/test partition was made at the subject-level, meaning the cadaver, animal, or specimen, rather than at the spectrum or image level. External validation asks whether performance was confirmed in an independent cohort or site. Human in vivo data asks whether living-human data were used, so cadaveric and animal data count as No. A priori threshold/hyperparameters asks whether classification thresholds and hyperparameters were fixed in advance rather than tuned on the test set. Unit independence asks whether the analyzed units were independent rather than many spectra or image patches drawn from the same subject. Blinded independent reference standard asks whether the reference standard was unambiguous, independent of the index test, and applied blind to it. Yes means the study met the item, No means it did not, and Unclear means the study did not report enough detail to judge. A No or Unclear judgment on subject-level splitting signals a data-leakage risk. Findings across the five diagnostic accuracy studies: No study performed external validation, none tested ML-augmented optical sensing in living-human perfused intraoperative spine tissue, four used porcine models, and the only human study was cadaveric. Kosik 2023 [29] was judged high risk for data leakage because its ex vivo holdout appeared to be spectrum-level rather than subject-level; its in situ and in vivo porcine phases were external but qualitative and reported no accuracy, so the external-validation item remains No. All five studies were vulnerable to unit-of-analysis inflation because performance denominators were clustered spectra or image patches derived from a small number of subjects. Adapted appraisal of Unal 2025: Unal 2025 evaluated regression of fracture-toughness measures, crack-initiation toughness (K_init) and the J-integral, rather than tissue classification, so it did not enter the accuracy pool and received no QUADAS-2 traffic-light judgment [21]. Adapted means that the item was applied in modified form because the endpoint is regression. The 80/20 train/test split (n = 94 train, 24 test) was made at the specimen-level with inner 5-fold cross-validation for tuning. Because the study contributed 118 specimens from 118 donors and averaged the 16 to 20 Raman spectra acquired per specimen into a single composite spectrum, the specimen-level is also the donor level, and the analyzed units are independent. The a priori item is marked Adapted because a regression endpoint carries no classification threshold, although hyperparameters were tuned inside the training set only. The reference-standard item is marked Adapted because the reference standard was mechanical R-curve testing by single-edge-notched bend rather than a tissue diagnosis, and blinding was not reported. Leakage risk is low to moderate: the split is specimen-level, but the test set is small (n = 24), and the study performed no external validation. Out-of-scope appraisal of Mainard 2022 [32] and Ren 2025 [33]: out of scope means that neither QUADAS-2 nor the machine learning signaling items apply. Neither study used a machine learning classifier, and neither reported a reference-standard-based diagnostic accuracy endpoint. Both measured spinal cord perfusion: Mainard 2022 [32] with NIRS and PPG in one lesion-free pig and Ren 2025 [33] with LSCI in 31 rabbits, compared by paired t-tests. Both are reported as a prespecified narrative extension (not PRISMA-DTA eligible) and are excluded from all quantitative synthesis. Table 3 and Table 4 are both set in landscape orientation. Abbreviations: ML, machine learning; QUADAS-2, Quality Assessment of Diagnostic Accuracy Studies-2; PRISMA-DTA, Preferred Reporting Items for Systematic Reviews and Meta-Analyses-Diagnostic Test Accuracy; DRS, diffuse reflectance spectroscopy; OCT, optical coherence tomography; PS-OCT, polarization-sensitive optical coherence tomography; SVM, support vector machine; XGBoost, extreme gradient boosting; NIRS, near-infrared spectroscopy; PPG, photoplethysmography; LSCI, laser speckle contrast imaging; CV, cross-validation; K_init, crack-initiation toughness. DenseNet121, ResNet50, and Inception refer to convolutional neural network (CNN) architectures.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Salman, S.G.; Phadke, R.A.; Kumar, R.; Panwalker, N.; Salman, Z.G.; Zeitouny, R.; Sarnala, S.; Tavakkoli, A.; Waisberg, E.; Ong, J.; et al. Machine-Learning Augmented Optical Tissue Sensing for Intraoperative Guidance in Spine Surgery: A Systematic Review and Meta-Analysis. Optics 2026, 7, 62. https://doi.org/10.3390/opt7050062

AMA Style

Salman SG, Phadke RA, Kumar R, Panwalker N, Salman ZG, Zeitouny R, Sarnala S, Tavakkoli A, Waisberg E, Ong J, et al. Machine-Learning Augmented Optical Tissue Sensing for Intraoperative Guidance in Spine Surgery: A Systematic Review and Meta-Analysis. Optics. 2026; 7(5):62. https://doi.org/10.3390/opt7050062

Chicago/Turabian Style

Salman, Samer G., Rohan A. Phadke, Rahul Kumar, Neil Panwalker, Zane G. Salman, Ryan Zeitouny, Sai Sarnala, Alireza Tavakkoli, Ethan Waisberg, Joshua Ong, and et al. 2026. "Machine-Learning Augmented Optical Tissue Sensing for Intraoperative Guidance in Spine Surgery: A Systematic Review and Meta-Analysis" Optics 7, no. 5: 62. https://doi.org/10.3390/opt7050062

APA Style

Salman, S. G., Phadke, R. A., Kumar, R., Panwalker, N., Salman, Z. G., Zeitouny, R., Sarnala, S., Tavakkoli, A., Waisberg, E., Ong, J., Galhotra, S., Tripuraneni, A., & Rizkalla, J. (2026). Machine-Learning Augmented Optical Tissue Sensing for Intraoperative Guidance in Spine Surgery: A Systematic Review and Meta-Analysis. Optics, 7(5), 62. https://doi.org/10.3390/opt7050062

Article Metrics

Back to TopTop