Review Reports
- Carlos Bedolla 1,
- Jose M. Gonzalez 1 and
- Eric J. Snider 1,5,*
- et al.
Reviewer 1: Anonymous Reviewer 2: Anonymous Reviewer 3: Anonymous
Round 1
Reviewer 1 Report
Comments and Suggestions for Authors- Data were compared from 20 participants through the duration of the experiment. However, it is very little data to apply ML and corelate the performance. I suggest add more participants.
- Literature review is very short. Add literature related to CRM and CRI. Make it 25 Nos.
- Provide experimental setup figure.
- Provide a schematic view of sensor-device configurations. Use real images.
- Add Mean absolute error also.
- In Figure 4, The paired test result was highest between 2 and 3 min, why?
- Results and discussion section should be merged.
- Comparative analysis should be placed in seprate section.
- Conclusion should be point wise. Add future scope also for readers.
Author Response
Data were compared from 20 participants through the duration of the experiment. However, it is very little data to apply ML and corelate the performance. I suggest add more participants.
We thank the reviewer for this comment. We would first clarify that this study does not train or fit a machine learning model on the present cohort; both the CRM and CRI systems are fixed, pre-developed implementations, and the 20 participants are used to compare their outputs rather than to develop an algorithm. The sample size therefore pertains to the statistical power of the device comparison rather than to model training. For that comparison, the data were sufficient to reach statistically significant conclusions for each of our pre-specified analyses. Pooled correlation across all subjects and LBNP steps yielded strong linear agreement (R2 = 0.859, p < 0.0001) over 8,599 paired observations. Subject-level mean compensatory reserve was statistically equivalent between CRM (64.57%) and CRI (64.20%) within a ±5% margin (p = 0.00021), and CRM early-detection time was superior to CRI (20.52 ± 5.82 vs. 17.92 ± 7.95 minutes; p = 0.01811). The convergence of significant results indicates that the present cohort was adequately powered for the device comparison performed. We agree that larger cohorts remain valuable, and future LBNP protocols can continue to use these systems to gather additional participants datasets to further strengthen and generalize these findings.
Literature review is very short. Add literature related to CRM and CRI. Make it 25 Nos.
Thank you for this suggestion, we have expanded the literature review relating to CRI and CRM. We have added more references to previous studies validating the CRM and CRI algorithms.
Provide experimental setup figure. Provide a schematic view of sensor-device configurations. Use real images.
Thank you for this suggestion, we have added an additional figure, Figure 1, which presents the physical LBNP setup with color-coded schematics of the sensor-device configuration. We have used actual images of the sensor-device configurations and our patient monitoring devices. We note that the individual shown is a study author who volunteered for the photograph and was not an enrolled participant; consent for publication was obtained. This statement was also added to “Informed Consent Statement” for awareness that appropriate consent was obtained from the author for the publication of this figure.
Add mean absolute error also.
We thank the reviewer for this suggestion. The magnitude-based error requested is captured in our performance metrics (Section 2.4). Specifically, we reported median absolute error as our measure of overall accuracy independent of bias, together with median error for systematic bias and Bland-Altman mean bias and 95% limits of agreement for spread of disagreement. We selected the median form of absolute error rather than mean because the relevant distributions departed from normality, making median the more robust and representative summary of error magnitude. We believe the requested error characterization is already captured by the present metrics.
In Figure 4, The paired test result was highest between 2 and 3 min, why?
The comparison in figure 4 is with regard to early detection time compared to experimental stop or decompensation. The authors unfortunately do not understand the reviewer's comment relating to 2 and 3 minutes being highest. We are happy to clarify but do not understand the point. The figure is looking at distribution of when slope criteria hit for early detection time and that typically was 20 to 30 minutes in advance of decompensation. The slope detection is highly sensitive, so these times being high are expected and in line with prior publications.
Results and discussion section should be merged.
The results and discussion are split based on journal formatting requirements. For this, we defer to the journal editorial staff on best practices for formatting the journal with regard to these sections.
Comparative analysis should be placed in seprate section.
Thank you for bringing this point up, our comparative analysis section was named “Compensatory Reserve Performance”, we have updated Section 3.2 to be named Comparative Analysis as that more appropriately reflects the content in Section 3.2.
Conclusion should be point wise. Add future scope also for readers.
Thank you for the suggestion. However, if by point wise, the reviewer means bullet points, we would also defer to the editorial staff on best formatting practices. We are happy to format the conclusion in bullet format if that is preferred. Future research directions have been added to the conclusion per the reviewer recommendation.
Reviewer 2 Report
Comments and Suggestions for AuthorsThe Bland‑Altman analysis shows 95% limits of agreement from –19.85% to +20.13% on a 0–100% scale. While the mean bias is near zero, individual differences between CRM and CRI can exceed 20 percentage points. In a clinical setting where compensatory reserve values guide triage decisions (e.g., a drop from 80% to 60% may indicate impending decompensation), such wide limits imply that the two devices are not interchangeable for individual patient monitoring. The authors should discuss the clinical implications of this variability and whether averaging or trend analysis mitigates the risk.
Equivalence testing uses a margin of ±5% relative to CRI, but no clinical or regulatory justification is provided for this value. More critically, the non‑inferiority test for early detection time uses an inferiority limit of 0% (i.e., CRM must not be slower than CRI). Standard non‑inferiority trials require a pre‑specified negative margin (e.g., –2 minutes) to account for acceptable loss of performance. Testing against zero is essentially a superiority test for non‑inferiority. The authors should revise this analysis or provide a strong rationale for a zero margin. Missing data and exclusion of 12 participants without both devices recording
Of 35 recruited participants, 12 (34%) were excluded because not both devices recorded data throughout the study. The manuscript does not describe the nature of these recording failures (e.g., sensor disconnection, software crash, PPG signal quality loss). Such a high exclusion rate can introduce selection bias if the failures correlate with physiological parameters (e.g., low blood pressure, motion artifact). The authors should report the reasons for missing data and perform a sensitivity analysis comparing baseline characteristics of included vs. excluded participants to assess potential bias.
Author Response
The Bland‑Altman analysis shows 95% limits of agreement from –19.85% to +20.13% on a 0–100% scale. While the mean bias is near zero, individual differences between CRM and CRI can exceed 20 percentage points. In a clinical setting where compensatory reserve values guide triage decisions (e.g., a drop from 80% to 60% may indicate impending decompensation), such wide limits imply that the two devices are not interchangeable for individual patient monitoring. The authors should discuss the clinical implications of this variability and whether averaging or trend analysis mitigates the risk.
Thank you for making this point. We have further expanded on this point in the discussion. We have addressed the clinical implications of the limits of agreement. We have clarified that CRM and CRI are not interchangeable for single timepoint readings but show equivalent averaged values and non-inferior trend-based detection, supporting their preferred use as a trend monitoring tool rather than an individual, absolute value tools with no information on trend over time.
Equivalence testing uses a margin of ±5% relative to CRI, but no clinical or regulatory justification is provided for this value. More critically, the non‑inferiority test for early detection time uses an inferiority limit of 0% (i.e., CRM must not be slower than CRI). Standard non‑inferiority trials require a pre‑specified negative margin (e.g., –2 minutes) to account for acceptable loss of performance. Testing against zero is essentially a superiority test for non‑inferiority. The authors should revise this analysis or provide a strong rationale for a zero margin.
We appreciate the critical review of the manuscript. The equivalence test margin was set at a similar threshold to ANSI criterion for medical devices for heart rate measurement, 5 bpm. CRM measures in percentile so 5% was used. References for justification of this point have been added in the methods. The justification for non-inferiority was that CRM needs to be equivalent or better than CRI with regards to prediction time which is why the margin was set at 0 as it should not be worse. We agree with the reviewer's point and have modified the statistical test to assess for superiority for this test as it is more appropriate when no inferiority margin is provided. Methods have been updated to reflect this change.
Of 35 recruited participants, 12 (34%) were excluded because not both devices recorded data throughout the study. The manuscript does not describe the nature of these recording failures (e.g., sensor disconnection, software crash, PPG signal quality loss). Such a high exclusion rate can introduce selection bias if the failures correlate with physiological parameters (e.g., low blood pressure, motion artifact). The authors should report the reasons for missing data and perform a sensitivity analysis comparing baseline characteristics of included vs. excluded participants to assess potential bias.
We appreciate you pointing this out and allowing us to clarify. The 12 participants referred to here were not excluded from this study, but part of a larger data collection effort for a different objective. Only 20 participants out of the 35 recruited for the larger study were prospectively collected for the purpose of comparing CRM and CRI. The manuscript text has been modified to reflect this.
Reviewer 3 Report
Comments and Suggestions for AuthorsComments:
This study evaluates a new deep learning-based device by comparing it with an FDA-cleared benchmark. Testing both systems on 20 healthy subjects during a progressive LBNP protocol is a practical way to see how they align. The results show a solid linear agreement, a very low mean bias and even an earlier detection time. The logic is good, and the result is comprehensive, but it would be better accepted if these concerns could be addressed.
The authors claimed that the CRM algorithm was originally trained using a historical dataset of 194 LBNP subjects. To ensure the validity of the evaluation, could the authors explicitly clarify whether the 20 healthy subjects included in this real-time protocol were completely independent of the historical training, validation, and hyperparameter-tuning cohorts? Any overlap between the training pool and the prospective testing pool must be ruled out to ensure no data leakage occurred.
The final analysis is based on a sample of 20 subjects. Is the sample size of 20 sufficient to provide adequate statistical power for the non-inferiority test?
The CRM system processes raw PPG waveforms sampled at 31.25 Hz. In many environments, motion artifacts are highly common and typically introduce severe baseline shifts or high-frequency distortion to raw optical signals. How does the 1-D CNN model maintain inference accuracy when the raw input waveform is compromised by movement, especially the output is only 1Hz. Also, does the current CRM system include a Signal Quality Index (SQI) or a filtering layer to reject noisy data?
Did the authors evaluate alternative machine learning or deep learning architectures for the development of CRM?
Author Response
This study evaluates a new deep learning-based device by comparing it with an FDA-cleared benchmark. Testing both systems on 20 healthy subjects during a progressive LBNP protocol is a practical way to see how they align. The results show a solid linear agreement, a very low mean bias and even an earlier detection time. The logic is good, and the result is comprehensive, but it would be better accepted if these concerns could be addressed.
We appreciate you taking the time to review our manuscript. Please see responses for each point below.
The authors claimed that the CRM algorithm was originally trained using a historical dataset of 194 LBNP subjects. To ensure the validity of the evaluation, could the authors explicitly clarify whether the 20 healthy subjects included in this real-time protocol were completely independent of the historical training, validation, and hyperparameter-tuning cohorts? Any overlap between the training pool and the prospective testing pool must be ruled out to ensure no data leakage occurred.
Thank you for the thorough review of our manuscript. The data analyzed in this study from the 20 subjects was prospectively captured. However, during the recruitment of these 20 participants the study team did not control for any potential previous participation in LBNP experiments which could have contributed to the original 194 subjects. The manuscript text has been clarified to reflect this.
The final analysis is based on a sample of 20 subjects. Is the sample size of 20 sufficient to provide adequate statistical power for the non-inferiority test?
We appreciate you pointing this out. We believe that 20 is sufficient to provide adequate statistical power for the non-inferiority test when taken into account with all the other statistical analysis performed. Pooled correlation across all subjects and LBNP steps yielded strong linear agreement (R2 = 0.859, p < 0.0001) over 8,599 paired observations. Subject-level mean compensatory reserve was statistically equivalent between CRM (64.57%) and CRI (64.20%) within a ±5% margin (p = 0.00021), and CRM early-detection time was non-inferior to CRI (20.52 ± 5.82 vs. 17.92 ± 7.95 minutes; p = 0.01811). The convergence of all three of these statistical analyses performed demonstrates that the non-inferiority test is sufficient and is supported by the two remaining statistical analyses.
The CRM system processes raw PPG waveforms sampled at 31.25 Hz. In many environments, motion artifacts are highly common and typically introduce severe baseline shifts or high-frequency distortion to raw optical signals. How does the 1-D CNN model maintain inference accuracy when the raw input waveform is compromised by movement, especially the output is only 1Hz. Also, does the current CRM system include a Signal Quality Index (SQI) or a filtering layer to reject noisy data?
The controlled laboratory setting of this study does not provide the type of signal distortion that a real-world scenario would introduce. The CRM algorithm was also developed using similarly “clean” data, so it is unable to identify such signal noise before inference. Future work would need to characterize data that appears in noisy real-world scenarios so that methodologies can be developed that work alongside the CRM algorithm. Comments have been added to the Discussion section to address this concern.
Did the authors evaluate alternative machine learning or deep learning architectures for the development of CRM?
We thank the reviewer for this question. We would like to clarify that the present study is a device-comparison and validation study rather than an algorithm-development study. Both the CRM and CRI systems were used as finished, fixed implementations: CRI as the FDA-cleared reference and CRM as the system under evaluation. Accordingly, the development and selection of the CRM estimation algorithm, including any consideration of alternative model architectures, was outside the goals of this current work. Our goal focused specifically on the head-to-head agreement and predictive performance of the two compensatory reserve systems as deployed.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsLiterature review is still very poor. Please add separate a literature review section and summarise critical research gap and add point wise objective of this research.
Comments on the Quality of English LanguageFine
Author Response
Thank you for your feedback and the opportunity to respond again. We have expanded the introduction to include a more in-depth literature review of both the CRM and CRI algorithms in a dedicated section. In addition, we have added a point wise objective at the closing of the introduction section.
Reviewer 2 Report
Comments and Suggestions for AuthorsI have no more comments.
Author Response
Thank you for taking the time to review our manuscript.