Next Article in Journal
Association of Academic Stress, Physical Activity, Sedentary Behavior, and Diabetes Risk Among University Students
Previous Article in Journal
Why Older Adults Resist Mobile Health Information Services: A Conceptual Model Based on the Technology–Personal–Environment Framework
 
 
Article
Peer-Review Record

The Many Faces of Stress: Preliminary Validation of a Remote Photoplethysmography-Based Tool for Psychophysiological Stress and Emotional Distress Monitoring

Healthcare 2026, 14(13), 1893; https://doi.org/10.3390/healthcare14131893
by Livio Provenzi 1,2, Valeria Calcaterra 3,4, Sarah Nazzari 1,*, Paolo Osvaldo Agnelli 5, Marco Xodo 5, Sergio De Pasquale 5 and Gianvincenzo Zuccotti 4,6,*
Reviewer 1: Anonymous
Reviewer 2:
Healthcare 2026, 14(13), 1893; https://doi.org/10.3390/healthcare14131893
Submission received: 24 April 2026 / Revised: 15 June 2026 / Accepted: 24 June 2026 / Published: 29 June 2026
(This article belongs to the Special Issue Health and Wellbeing Strategy Evaluation)

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

In the current study “The many faces of stress: preliminary validation of a remote photoplethysmography-based tool for psychophysiological stress and burnout risk monitoring” the authors present a promising feasibility/validation study that has been carefully constructed and holds significant promise to expand the field. The authors are commended for this excellent study.

There are however several aspects regarding design, analysis, and interpretation that would benefit from clarification and tempering, especially around the methodology, what “validation” and “stress/burnout monitoring” can be concluded from cross-sectional correlations of small effect size.

Abstract and Introduction

The authors clearly motivate chronic stress and its links to burnout, anxiety, and depression, and you correctly note the limitations of self-report and the appeal of contactless rPPG as an objective method, however at present the background implies that rPPG for stress is relatively novel, but a number of studies have already demonstrated feasibility and high accuracy for stress detection under controlled conditions, often by combining rPPG-derived HR/HRV features with machine-learning classifiers. Perhaps authors can refine by adding their novel contribution, that rPPG can reliably derive HR and HRV under many conditions or that several lab or semi-ecological studies have used rPPG-derived HRV features to classify stress/non-stress states or academic stress with high accuracy. This may assist in distinguishing their contribution which is a validation of a mobile, in-the-wild, consumer-like app with composite indices against psychological distress, rather than rPPG for stress per se.

Throughout it is important to clarify the constructs, as the abstract moves quickly from “chronic stress” to “stress monitoring” and “burnout risk” but your indices appear to be derived from a short, standardized recording, or a quasi-acute physiological state, and linked to self-reported symptoms via cross-sectional correlations. Perhaps highlight that it is only momentary psychophysiological stress arousal? As well as please clarify whether it is only self-report measures in the study that capture current symptom severity rather than objective burnout diagnoses or longitudinal risk?

Currently the word Validate” is too strong for the presented design, as with a single time point, no physiological reference standard (ECG or contact PPG), and only modest correlations with self-report scales, you are primarily assessing feasibility, usability, score distributions, and convergent validity (construct validity) – not full validation for clinical or longitudinal use. Perhaps, as this is a preliminary report 9indicated in the introduction) the authors might indicate that the aim was to assess feasibility and construct validity?

Additionally, at present it is indicated “burnout risk monitoring” in the objective, but the data appear to be cross-sectional correlations with burnout-related emotional distress. Cross-sectional associations do not establish risk prediction. Thus rather a cross-sectional association might be more appropriate? Perhaps “burnout-related emotional distress in daily life”?

In the Introduction particularly, the authors state that “Mental health and stress regulation… symptoms of stress, anxiety, and depression… particularly vulnerable, including university students…” They correctly state that chronic stress is a “fertile ground” for emotional disorders and burnout, but the authors could anchor this in psychophysiology, by linking chronic/ sustained stress to autonomic dysregulation and HRV changes, which already ties into your rPPG/HRV focus later.

The authors clearly indicate the limitations of self-report and need for objective measures, and the logic that self-report is subjective and often insensitive to short-term fluctuations is sound and well-accepted. However they should please avoid overstating that objective measures are inherently “better” – as physiological measures are not purely objective in a conceptual sense (they still require interpretation and can be context-dependent). Thus, perhapse emphasize complementary rather than replacing self-report.

Which physiological markers are relevant? The authors might consider briefly mentioning heart rate and HRV as established markers of stress-related autonomic imbalance – setting up their argument for why an rPPG-based stress index is plausible.

The authors indicate in lines 79-88 that “Remote photoplethysmography (rPPG)… non-contact, low-cost, and scalable… vital parameters such as heart rate, heart rate variability, respiratory signals… growing evidence supports the accuracy and reliability…” Yet usage of the terms “Continuous” and “reliable” need nuance, as rPPG performance strongly depends on lighting, motion, and camera quality, and can be less robust than contact PPG in some real-world conditions, consider explicitly acknowledge some technical constraints (for example sensitivity to motion, illumination changes). Please clarify the current level of evidence for mental-health applications

Please avoid implying fully “continuous” monitoring from smartphone rPPG – as the authors’ methodology appears to be episodic spot measurements via smartphone, not continuous streaming like a wearable. Thus a suggestion would be to rephrase “continuous monitoring” to something like “repeated, unobtrusive spot assessments” or “frequent, user-initiated measurements in daily life.”.

The authors should please cite emerging rPPG stress-detection studies using facial videos and academic/mental tasks, which align more directly with your work.

Given current evidence, rPPG stress monitoring is best described as promising and complementary rather than ready for widespread clinical deployment.

The authors should please avoid implying direct societal-level impact without evidence, as the statement “…potentially reducing the burden of stress-related disorders at both individual and societal levels” is aspirational but not empirically supported by the present data.

Regarding lines 101-109“The aim of the preliminary report was to evaluate the potential of a digital stress monitoring tool based on rPPG… assess the usability… distribution of stress indices… explore associations… macro-factor of emotional distress, considered as a proxy of burnout.” The authors identify three concrete objectives: usability, index distributions, and associations with psychological/sociodemographic/lifestyle variables. Yet it is kindly suggested to modify “evaluate the potential” toward “feasibility and preliminary validity”, as this may better align with what the study actually allowed for (cross-sectional feasibility and convergent validity) rather than broad “potential.”

Methodology

The authors are applauded for a detailed and technically sound methodology, especially on the rPPG pipeline, but there are important gaps in sampling, covariate measurement, index transparency, psychometrics, and statistical strategy that limit what can be claimed as “validation” or “burnout-related distress prediction.”. These are addressed based on general methodology aspects.

The inclusion/exclusion criteria require additional detail as at present there is no information on exclusion criteria (cardiovascular disease, psychotropic medication, skin conditions affecting rPPG, severe psychiatric disorders, substance use), which can materially impact both HRV-like indices and psychological symptoms. Please highlight or indicate if there we any health-related exclusion criteria, whether participants were required to have normal vision, no facial coverings, etc. or whether current psychiatric treatment or medication use was assessed/controlled.

The authors indicate that there are two sources as there two distinct subsamples (panel vs. students), but the methods do not state whether analyses control for sampling source or if sensitivity analyses were run separately in each group. As student vs. non-student status may correlate with age, stress levels, and possibly rPPG signal quality (different devices), the authors are advised to include recruitment source as a covariate or grouping factor in some analyses or state that sample source was recorded and inspected.

Importantly, for an rPPG-based smartphone study, it matters whether participants used their own phones (diverse hardware) or standardized devices were provided. Currently it is detailed that the participants used “a mobile app” but not whether the phone type was controlled. Thus, perhaps it would provide greater clarity if the authors indicated if all participants used the same model; state make/model and OS, as well as if they used different phones; note this heterogeneity and whether the app internally adjusted for frame rate/resolution.

The authors mention consent and data protection, which is good. They should please also add detail on the ethics committee approval (name, reference number).

The authors are applauded for using GAD and BDI as these are reasonable choices, as well-established instruments with good psychometric support.

However, there is some missing details regarding these instruments. For reproducibility and interpretation, the authors should please specify: the exact versions (e.g., GAD-7 vs. GAD-2; BDI-II vs. BDI-IA),  the number of items and scoring range as well as the language and whether validated translations were used. Please also consider stating internal consistency (Cronbach’s alpha) in this sample.

Currently there are no explicit stress or burnout questionnaires/ instruments? Given the focus on stress and burnout-related emotional distress, it is surprising that only anxiety and depression scales are mentioned here. The authors might consider to create a “Burnout-related Emotional Distress” factor solely from anxiety and depression scores. Methodologically, this is a significant concern if the authors infer and build a case for burnout.

The intro and aims mention lifestyle variables and sociodemographics, but the methods do not specify which were measured (aspects like sleep, physical activity, smoking, alcohol, caffeine, BMI, medication). Since lifestyle factors affect HRV and stress, specifying them and considering them as covariates or exploratory predictors would strengthen analyses and inferences.

Please clarify whether questionnaires were completed immediately before/after the rPPG recording and whether there were any delays, as timing affects the interpretation of “associations.”

The authors clearly describe that a mobile rPPG app was used and that a standardized protocol (quiet room, minimal movement, alignment) was followed.

However, recording duration and frame rate not reported, especially for rPPG and HRV-like features, the length of recording (e.g., 30 s vs 2 min) and camera frame rate/resolution are critical. Thus please specify or provide detail on the duration of the facial scan, the approximate frame rate (e.g., 30 fps) and resolution or whether exposure/brightness was fixed or automatic.

Although it is indicated that a “Quiet environment” was experienced, but lighting is not (natural vs artificial, consistency). Thus please clarify whether any quality checks on lighting were done (aspects like minimum brightness threshold), or whether participants were instructed to avoid backlighting or direct sunlight.

It appears only one scan per participant was performed; if so, please indicate as such. Given the dynamic nature of stress, a single brief measurement is a limitation. Thus the authors might consider acknowledging this explicitly and suggesting repeated or longitudinal measures in future designs.

The authors clearly describe face detection, ROI selection, tracking, RGB trace extraction, filtering, and artifact rejection, in line with established rPPG practice.

Yet aspects such as “These may include…” sounds too vague for a methods section, especially phrases like “These may include heart rate, pulse waveform characteristics, and time-/frequency-domain indices…” are generic. Even if the weighting is proprietary, you should clarify which feature classes are definitely used (like mean HR, SDNN-like measures, RMSSD-like measures, spectral power in certain bands). This may aid in improving/ elevating plausibility that indices reflect autonomic regulation.

At present there is no explicit validation of rPPG vs reference in this sample? The authors indicate that the system was previously evaluated against reference devices, but there is no validation against ECG/contact PPG in this sample or a subset, so at present they are inferring PRV/HRV validity from prior work (please cite or indicate intra-and inter variability).

Please add some more detail on data quality, as QC is described conceptually, yet please justify or describe criteria for excluding a scan (aspects like minimum usable signal length, SNR threshold, maximum allowed motion). How many participants’ recordings were excluded or failed and how these were handled in analyses (missing data or study withdrawal).

The authors justifiably state proprietary algorithms and interpretability, but by keeping with the IP, they might add that the training context of the models (trained on internal datasets with lab-based stressors or clinical populations). Whether indices are on standardized scales (0–100), and what higher vs lower values indicate. Where possible, citing internal validation (even if in grey literature or regulatory submissions) or other publications using the same indices would bolster credibility.

At present there is a mismatch between conceptual labels and single-snapshot measurement. Stress Recovery and Stress Response conceptually require a stress-recovery or baseline-stressor design (pre/post or dynamic), but here the authors appear to have only a single short recording under resting instructions. That raises the question: how can the app estimate recovery or reactivity from a single snapshot? Thus please clarify the intended interpretation, as in are these trait-like, derived from waveform morphology/variability at rest (“capacity to recover” inferred from vagal tone proxies) or are they tuned on stress-recovery datasets but applied here as “trait proxies”? Please address more explicitly and highlight the limitation.

The authors indicate that the indices are “intended to capture different aspects” but do not discuss whether they are correlated with each other or whether any prior work has separately validated each of them. Please consider reporting inter-correlations and explaining how strongly they overlap and consider discussing whether “Recovery” and “Response” are truly distinct constructs in this context.

Are the scores normalized for age/ sex etc? What are the index range and directionality(0–100) or  (higher = more stress / poorer recovery).

Although it is indicated that an ad hoc questionnaire was performed, there are no psychometric info. Thus a 5-item ad hoc scale, which is acceptable for preliminary work, but please indicate whether items were adapted from any existing usability instruments (e.g., SUS) or designed de novo and please report internal consistency (Cronbach’s alpha) in the results.

Please also clarify whether the averaged item scores were done to obtain an overall usability score and whether higher scores indicate more positive usability. If usability is used in any inferential analyses (aspects like association with indices or symptoms), please highlight this in the statistical analysis plan.

Currently there is no mention of how many participants had complete self-report and usable rPPG data or how missing data were handled (listwise deletion vs. imputation). This should please be explicitly reported and justified.

Please detail normality and choice of tests, as it is mentioned that Pearson correlations and ANOVA were done, which assume approximate normality and homoscedasticity. Please add whether assumptions were checked (visual inspection, Shapiro-Wilk) or whether transformations or nonparametric alternatives were considered if assumptions were violated.

Importantly, with the current design the authors compute multiple correlations between three indices and multiple outcomes, and conduct ANOVAs along with regression. Yet at present there is no plan for correction for multiple testing, while exploratory work can sometimes forgo strict correction, the authors should please consider or mention whether they adjusted p-values (like FDR) and, if not, state this as a limitation and interpret significant findings cautiously.

The PCA for  “Burnout-related Emotional Distress” raises some conceptional questions. PCA on just two variables (anxiety and depression) will almost always yield a single component, so PCA adds little beyond simply averaging or standardizing and summing scores. Labelling that component “Burnout-related Emotional Distress” is conceptually stretched if authors did not include any burnout items. Perhaps this may be mitigated by stating that this component reflects shared variance in anxiety and depressive symptoms and is used as a general emotional distress factor. Again, avoid overusing “burnout-related” in the label, or explicitly acknowledge that this is a proxy, not a direct burnout measurement.

Authors state that a multiple regression model was estimated with the three indices as predictors of the component score. Yet some key details missing are covariates included (age, sex, recruitment source, lifestyle variables)? These are important potential confounders of both physiological indices and distress. How was multicollinearity checked (especially if indices are correlated)? Whether residual diagnostics (normality, heteroscedasticity) were performed. Approaches like  hierarchical regression may be considered?

Discussion

The authors are commended for a well written and theoretically rich discussion, but it over-interprets small, cross-sectional associations, leans heavily into “early detection” and “cumulative burden” language that the data do not support, and underplays key methodological and technological limitations of rPPG and PPG‑based stress indices.

Lines 324-336:

Authors indicate that “The present study provides preliminary evidence supporting the usability, acceptability, and construct validity…”
“…consistent association… suggests that the app is not merely capturing transient fluctuations in stress, but may reflect more stable and clinically meaningful dimensions…”

However, this overstates the construct validity and design, as the correlations of Stress Level with anxiety, depression, and the composite factor are statistically significant but small (r ≈ .13–.17), and only one of three indices showed consistent associations. At present the “Construct validity” suggests strong, replicated alignment between the physiological indices and the intended construct (stress/burnout), which your results do not yet support. The associations are preliminary convergent signals, not full construct validation.

Indicating “Consistent association” implies robustness the study currently does not have, with three outcomes, three indices; only Stress Level shows meaningful associations. At present the discussion treats this as broadly consistent across indices and outcomes, which glosses over the fact that Stress Recovery and Stress Response did not behave as theorized. At present there is no test-retest data, no longitudinal follow‑up, and no clinical diagnoses; thus the authors should please be cautious with inferences, as at present they cannot show that the index reflects stable vulnerability or clinically meaningful traits vs. just current physiological arousal

Psychobiological interpretation and “cumulative burden” (lines 337–347)

Authors specifically state that “…emphasize the central role of chronic stress exposure and allostatic load…”
“…Stress Level… may be interpreted as a proxy marker of cumulative psychophysiological burden rather than a simple indicator of momentary stress…”

Importantly, allostatic load or parallels were not as at present there we no standard allostatic load markers (e.g., cortisol, inflammatory markers, metabolic parameters). Thus the interpretation that Stress Level is a proxy of cumulative burden is not supported by any longitudinal or multi‑system biomarker data. Chronic vs. acute stress is not distinguishable at present, given one short rPPG recording and cross‑sectional symptom scores. Indeed, rPPG provides pulse rate variability (PRV), which is an imperfect proxy of ECG‑derived HRV and may diverge especially in some physiological states. The current Discussion assumes that rPPG indices map directly onto HRV‑based allostatic load models without mentioning measurement differences or absence of ECG validation in this sample.

The authors might consider presenting the psychobiological model as a theoretical framework that their findings are compatible with, rather than claiming Stress Level is a proxy of cumulative burden, where they used rPPG‑derived PRV, not gold‑standard HRV, and did not include direct allostatic load biomarkers. Again, cross‑sectional, single‑snapshot data cannot distinguish chronic from acute stress.

Stress Level vs. Stress Recovery and Stress Response (lines 348–359)
The authors specifically argue that “Stress Level emerged as the only significant predictor…”
“…indices are highly correlated and partially overlapping… suggests that Stress Level may capture the broader common component of stress burden…”

However, the Recovery and Response are defined as reflecting recovery capacity and reactivity, but, please note the limitation as only one resting recording was collected and there was no stress‑induction or recovery phase included? Additionally, the collinearity explanation is partial, as the authors correctly mention shared variance, but please quantify intercorrelations and discuss whether multicollinearity diagnostics (e.g., VIF) were checked.

Importantly, with a single time‑point at rest, one cannot directly test “response” or “recovery” constructs; this limits the interpretation of those indices, and the non‑significant results for Recovery and Response may reflect both conceptual and measurement mismatches, not just collinearity. Thus the authors are kindly requested to reconsider and perhaps frame Stress Level as the index that showed the clearest statistical signal in this design, rather than as definitively capturing “stress burden.”

Burnout, comorbidity, and “early psychophysiological signals” (lines 360–368)
The authors again indicate, correctly, that “…high rates of comorbidity between burnout, anxiety, and depression…”
“…present findings support the notion that r-PPG comestai_app may capture early psychophysiological signals… related to this broader spectrum of stress-related pathology…”

However, burnout was not directly measured, as the investigators did not administer a validated burnout scale; instead, they created a factor from anxiety and depression scores. Thus please be cognisant and careful in calling this “Burnout-related Emotional Distress” risks implying a directly measured burnout when in actuality the authors measured general emotional distress.

At present, there is no longitudinal follow‑up, no clinical diagnoses, no outcome data. Thus, there is no evidence that app indices change before clinical burnout or depressive episodes, at present there are only cross‑sectional associations. In accordance with this, statements about somatic comorbidities and premature mortality are true at population level, but the present data do not address these outcomes.

To assist in this, the authors might consider reframing the factor as a general emotional distress indicator that likely overlaps with burnout‑related exhaustion, and be explicit that burnout per se was not measured, and present “early signals” as a hypothesis for future longitudinal work, not as something supported by the current results.

Preventive and applied implications (lines 369–379)
Statements like “…continuously and non-invasively monitoring stress…”
“…identify individuals who are progressively accumulating high stress levels before the clinical onset of burnout or depressive disorders…”

Again, the present study uses a single short scan in a controlled setting, yet the discussion suggests continuous or frequent real‑world monitoring, which is not what was implemented or evaluated. As at present there is no examination or prediction of future burnout or depression, trajectory of stress accumulation, or sensitivity to change. Claims about identifying individuals “before clinical onset” are not supported by cross‑sectional correlations, although these correlations are important.

Perhaps the authors can describe preventive potential as conditional on future longitudinal and real‑world validation, using words like“could, if validated” rather than “would” or “allows.” This may clarify that the present study demonstrated feasibility of single-session use, not continuous deployment or predictive screening.

“Moment-to-moment indicators” and proactive care (lines 380–388)

Again, given the design, the temporal resolution may be overstated. Indicating “moment‑to‑moment” and “continuous screening” imply high-frequency monitoring over time. At present there are no repeated measures, thus no within-person trajectories, accumulation patterns, or recovery slopes. The intervention linkage is speculative, albeit true. Please be cautious here, as the study did not test any intervention, digital coaching, or clinical decision support tied to the app indices.

Perhaps rather position these as possible future applications that need testing, and explicitly say that current data do not show whether specific thresholds or patterns in the indices meaningfully guide interventions or improve outcomes.

Multimodal assessment and “clinical validity” (lines 393–399)

Statements like  “…essential in psychosomatic medicine… convergence… supports the ecological and clinical validity…” do not allow for “Clinical validity” , as the sample population was non‑clinical; measures were screening scales; and indices were unvalidated against clinical endpoints. Thus there is a preliminary convergent validity with symptoms, but no clinical diagnoses, no sensitivity/specificity analyses, and no prognostic evaluation. Please also note, which is always the case with stress-related studies, the impact of ecological validity. Measurements were not truly ambulatory or context‑rich (e.g., multiple daily assessments over weeks). Thus, ecological validity is plausible conceptually (smartphone outside lab), but empirically limited to a one‑off controlled session.

The authors might consider replacing “clinical validity” with “preliminary convergent validity” and be explicit about the non‑clinical nature of the sample. Perhaps also clarify that ecological validity refers here to app use outside a traditional lab, but not yet to fully passive or context-aware sensing.

Age and gender differences (lines 400–405)
“…younger individuals and women showing higher stress levels… consistent with literature… results further support potential usefulness in vulnerable populations.”

Although these statements are true, based on current literature, yet it is not clear if age/sex differences in Stress Level remain after controlling for mental health symptoms, recruitment source, or other variables. Thus, without adjustment, group differences might simply reflect symptom differences or sample composition. In line with this, algorithmic and physiological biases are not discussed a present, as autonomic parameters and rPPG quality can vary by age/sex, and there is emerging concern about demographic bias in physiological algorithms, please consider acknowledging g these complexities.

Conclusions and interpretation

Burnout “risk” language as cross-sectional correlations with burnout-related distress do not establish risk. Risk implies prediction of future outcomes. Perhaps the authors can consider something more along the lines of “…for monitoring burnout-related emotional distress” or “…for assessing associations with burnout-related emotional distress” rather than “burnout risk monitoring.

Author Response

In the current study “The many faces of stress: preliminary validation of a remote photoplethysmography-based tool for psychophysiological stress and burnout risk monitoring” the authors present a promising feasibility/validation study that has been carefully constructed and holds significant promise to expand the field. The authors are commended for this excellent study.

There are however several aspects regarding design, analysis, and interpretation that would benefit from clarification and tempering, especially around the methodology, what “validation” and “stress/burnout monitoring” can be concluded from cross-sectional correlations of small effect size.

 

Q1 Abstract and Introduction

The authors clearly motivate chronic stress and its links to burnout, anxiety, and depression, and you correctly note the limitations of self-report and the appeal of contactless rPPG as an objective method, however at present the background implies that rPPG for stress is relatively novel, but a number of studies have already demonstrated feasibility and high accuracy for stress detection under controlled conditions, often by combining rPPG-derived HR/HRV features with machine-learning classifiers. Perhaps authors can refine by adding their novel contribution, that rPPG can reliably derive HR and HRV under many conditions or that several lab or semi-ecological studies have used rPPG-derived HRV features to classify stress/non-stress states or academic stress with high accuracy. This may assist in distinguishing their contribution which is a validation of a mobile, in-the-wild, consumer-like app with composite indices against psychological distress, rather than rPPG for stress per se.

Throughout it is important to clarify the constructs, as the abstract moves quickly from “chronic stress” to “stress monitoring” and “burnout risk” but your indices appear to be derived from a short, standardized recording, or a quasi-acute physiological state, and linked to self-reported symptoms via cross-sectional correlations. Perhaps highlight that it is only momentary psychophysiological stress arousal? As well as please clarify whether it is only self-report measures in the study that capture current symptom severity rather than objective burnout diagnoses or longitudinal risk?

Currently the word Validate” is too strong for the presented design, as with a single time point, no physiological reference standard (ECG or contact PPG), and only modest correlations with self-report scales, you are primarily assessing feasibility, usability, score distributions, and convergent validity (construct validity) – not full validation for clinical or longitudinal use. Perhaps, as this is a preliminary report 9indicated in the introduction) the authors might indicate that the aim was to assess feasibility and construct validity?

Additionally, at present it is indicated “burnout risk monitoring” in the objective, but the data appear to be cross-sectional correlations with burnout-related emotional distress. Cross-sectional associations do not establish risk prediction. Thus rather a cross-sectional association might be more appropriate? Perhaps “burnout-related emotional distress in daily life”?

In the Introduction particularly, the authors state that “Mental health and stress regulation… symptoms of stress, anxiety, and depression… particularly vulnerable, including university students…” They correctly state that chronic stress is a “fertile ground” for emotional disorders and burnout, but the authors could anchor this in psychophysiology, by linking chronic/ sustained stress to autonomic dysregulation and HRV changes, which already ties into your rPPG/HRV focus later.

The authors clearly indicate the limitations of self-report and need for objective measures, and the logic that self-report is subjective and often insensitive to short-term fluctuations is sound and well-accepted. However they should please avoid overstating that objective measures are inherently “better” – as physiological measures are not purely objective in a conceptual sense (they still require interpretation and can be context-dependent). Thus, perhapse emphasize complementary rather than replacing self-report.

Which physiological markers are relevant? The authors might consider briefly mentioning heart rate and HRV as established markers of stress-related autonomic imbalance – setting up their argument for why an rPPG-based stress index is plausible.

The authors indicate in lines 79-88 that “Remote photoplethysmography (rPPG)… non-contact, low-cost, and scalable… vital parameters such as heart rate, heart rate variability, respiratory signals… growing evidence supports the accuracy and reliability…” Yet usage of the terms “Continuous” and “reliable” need nuance, as rPPG performance strongly depends on lighting, motion, and camera quality, and can be less robust than contact PPG in some real-world conditions, consider explicitly acknowledge some technical constraints (for example sensitivity to motion, illumination changes). Please clarify the current level of evidence for mental-health applications

Please avoid implying fully “continuous” monitoring from smartphone rPPG – as the authors’ methodology appears to be episodic spot measurements via smartphone, not continuous streaming like a wearable. Thus a suggestion would be to rephrase “continuous monitoring” to something like “repeated, unobtrusive spot assessments” or “frequent, user-initiated measurements in daily life.”.

The authors should please cite emerging rPPG stress-detection studies using facial videos and academic/mental tasks, which align more directly with your work.

Given current evidence, rPPG stress monitoring is best described as promising and complementary rather than ready for widespread clinical deployment.

The authors should please avoid implying direct societal-level impact without evidence, as the statement “…potentially reducing the burden of stress-related disorders at both individual and societal levels” is aspirational but not empirically supported by the present data.

Regarding lines 101-109“The aim of the preliminary report was to evaluate the potential of a digital stress monitoring tool based on rPPG… assess the usability… distribution of stress indices… explore associations… macro-factor of emotional distress, considered as a proxy of burnout.” The authors identify three concrete objectives: usability, index distributions, and associations with psychological/sociodemographic/lifestyle variables. Yet it is kindly suggested to modify “evaluate the potential” toward “feasibility and preliminary validity”, as this may better align with what the study actually allowed for (cross-sectional feasibility and convergent validity) rather than broad “potential.”

R1: We thank the Reviewer for this thoughtful and constructive comment. In response, we revised multiple sections of the manuscript to more clearly emphasize the role of accessible, low-cost, and real-time digital technologies in supporting proactive and preventive mental health care.

Specifically, in the Abstract, we strengthened the description of the potential clinical and public health implications of rPPG-based monitoring by emphasizing that these technologies may support early, objective, and scalable identification of stress-related burden beyond traditional self-report measures. We also clarified that repeated and unobtrusive physiological monitoring through smartphone-based tools may contribute to preventive and personalized mental health approaches.

In the Introduction, we expanded the theoretical and applied rationale for the use of rPPG technologies in mental health monitoring. In particular, we clarified that rPPG-based systems represent a non-contact, low-cost, and scalable approach because they only require a standard smartphone camera and do not depend on wearable sensors or specialized medical devices. We further emphasized that stress-related physiological activation is dynamic and context-dependent, making repeated real-world assessments particularly relevant. Accordingly, we revised the text to highlight how these technologies may facilitate ecological and continuous monitoring of psychophysiological stress in everyday settings, potentially enabling the early identification of vulnerability trajectories before the onset of clinically relevant burnout, anxiety, or depressive disorders. In addition, we explicitly framed these tools as promising complementary approaches within multidimensional mental health assessment frameworks rather than as standalone diagnostic systems.

In the Discussion, we substantially expanded the interpretation of the preventive implications of our findings. We now more clearly state that repeated app-based physiological assessments may help detect sustained stress accumulation, altered autonomic regulation, and reduced recovery capacity before severe emotional distress becomes clinically manifest. We also highlighted that this approach may support a shift from a reactive mental health care model, focused primarily on treatment after symptom onset, toward a proactive and preventive framework centered on early detection, continuous monitoring, and timely intervention. To further strengthen this perspective, we added examples of potential practical applications, including stress-management prompts, breathing regulation exercises, behavioral recommendations, micro-break interventions, and referral pathways for individuals showing persistent elevations in stress-related physiological burden.

At the same time, we carefully balanced these implications by revising the manuscript to avoid overinterpretation. In the Limitationssection, we clarified that the present study provides preliminary evidence of feasibility, usability, and construct validity, but does not establish clinical diagnostic validity or longitudinal prediction of mental health outcomes. We therefore emphasized the need for future longitudinal studies, validation against established biological stress markers, and further testing across different clinical and occupational populations before widespread implementation in routine mental health care.

 

Q2 Methodology

The authors are applauded for a detailed and technically sound methodology, especially on the rPPG pipeline, but there are important gaps in sampling, covariate measurement, index transparency, psychometrics, and statistical strategy that limit what can be claimed as “validation” or “burnout-related distress prediction.”. These are addressed based on general methodology aspects.

The inclusion/exclusion criteria require additional detail as at present there is no information on exclusion criteria (cardiovascular disease, psychotropic medication, skin conditions affecting rPPG, severe psychiatric disorders, substance use), which can materially impact both HRV-like indices and psychological symptoms. Please highlight or indicate if there we any health-related exclusion criteria, whether participants were required to have normal vision, no facial coverings, etc. or whether current psychiatric treatment or medication use was assessed/controlled.

R2: We thank the reviewer for this helpful comment. We revised the Participants section to provide additional detail on inclusion and exclusion criteria. Specifically, we clarified that participants were excluded if they reported conditions that could materially affect HRV-like indices or psychological symptoms, including cardiovascular disease, psychotropic medication use, severe psychiatric disorders, or substance use. We also clarified that participants were required to complete the rPPG assessment under appropriate acquisition conditions, including adequate face visibility and absence of major movement artifacts. In addition, we specified that participants were asked to avoid factors that could interfere with facial video acquisition, such as poor lighting, backlighting, or facial coverings, and to follow the standardized instructions provided by the application. Current medication use and clinically relevant self-reported conditions were assessed through participant declaration and considered in the eligibility procedure.

 

Q3 The authors indicate that there are two sources as there two distinct subsamples (panel vs. students), but the methods do not state whether analyses control for sampling source or if sensitivity analyses were run separately in each group. As student vs. non-student status may correlate with age, stress levels, and possibly rPPG signal quality (different devices), the authors are advised to include recruitment source as a covariate or grouping factor in some analyses or state that sample source was recorded and inspected.

R3: We thank the Reviewer for raising this important point. We have now explicitly inspected whether the two subsamples differed on the main psychological and stress-related variables. Accordingly, we added the following analyses to the revised manuscript: A chi-square test of independence was conducted to examine the association between stress level and recruitment source. The association was statistically significant, χ²(2, N = 252) = 17.10, p < .001, indicating that the distribution of stress levels differed significantly between the two subsamples. Inspection of the observed frequencies showed that participants from the student sample were more frequently classified in the highest stress category than participants from the panel sample. Specifically, 21.7% of participants in the student sample were classified as high stress, compared with 6.0% in the panel sample. Conversely, the panel sample showed a higher proportion of participants in the intermediate stress category than the student sample. No significant associations were found for Stress Response and Stress Recovery indices.

In addition, we examined whether the two recruitment sources differed in anxiety symptoms, depressive symptoms, and burnout-related emotional distress. The distributions of anxiety and depressive symptoms did not significantly differ between the two subsamples, ps > .22. Likewise, no significant difference between the two subsamples was observed for the Burnout-related Emotional Distress factor, p > .62.

To address the Reviewer’s concern, the revised manuscript now states that recruitment source was recorded and inspected, and that differences between the two subsamples were limited to stress level classification, while the other stress-related and psychological indices did not significantly differ by source.

 

Q4 Importantly, for an rPPG-based smartphone study, it matters whether participants used their own phones (diverse hardware) or standardized devices were provided. Currently it is detailed that the participants used “a mobile app” but not whether the phone type was controlled. Thus, perhaps it would provide greater clarity if the authors indicated if all participants used the same model; state make/model and OS, as well as if they used different phones; note this heterogeneity and whether the app internally adjusted for frame rate/resolution.

R4: Participants completed the assessment using their own smartphones. Therefore, both Android and iOS devices were used, and phone make/model and operating system were not experimentally controlled. We have now explicitly acknowledged this heterogeneity in the revised manuscript.

 

Q5 The authors mention consent and data protection, which is good. They should please also add detail on the ethics committee approval (name, reference number).

R5: We thank the reviewer for this comment. As required by the journal format, the Institutional Review Board statementis reported in the dedicated section after the Discussion.

 

Q6 The authors are applauded for using GAD and BDI as these are reasonable choices, as well-established instruments with good psychometric support. However, there is some missing details regarding these instruments. For reproducibility and interpretation, the authors should please specify: the exact versions (e.g., GAD-7 vs. GAD-2; BDI-II vs. BDI-IA),  the number of items and scoring range as well as the language and whether validated translations were used. Please also consider stating internal consistency (Cronbach’s alpha) in this sample.

R6: Anxiety symptoms were assessed using the validated Italian version of the Generalized Anxiety Disorder 7-item scale (GAD-7). The GAD-7 consists of seven items rated from 0 (‘not at all’) to 3 (‘nearly every day’), yielding a total score ranging from 0 to 21, with higher scores indicating greater anxiety symptom severity. Depressive symptoms were assessed using the Italian version of the Beck Depression Inventory-II (BDI-II). The BDI-II consists of 21 items rated from 0 to 3, yielding a total score ranging from 0 to 63, with higher scores indicating greater depressive symptom severity. Internal consistency in the present sample was good for the GAD-7, Cronbach’s α = .86, and excellent for the BDI-II, Cronbach’s α = .91.

 

Q7 Currently there are no explicit stress or burnout questionnaires/ instruments? Given the focus on stress and burnout-related emotional distress, it is surprising that only anxiety and depression scales are mentioned here. The authors might consider to create a “Burnout-related Emotional Distress” factor solely from anxiety and depression scores. Methodologically, this is a significant concern if the authors infer and build a case for burnout.

R7: We thank the reviewer for this important conceptual clarification. We agree that the present study did not include a burnout-specific questionnaire and that anxiety and depressive symptoms should not be interpreted as direct measures of burnout. Accordingly, we revised the manuscript to avoid building a case for burnout based solely on anxiety and depression scores.

In the revised version, the PCA-derived variable has been relabeled general emotional distress. We clarified that this component reflects the shared variance between anxiety and depressive symptoms and is used as a general psychological distress factor, not as a direct measure of burnout. We also reduced the use of the term “burnout-related” throughout the manuscript and explicitly acknowledged that this variable may only be considered a proxy indicator of burnout-related psychological vulnerability.

Finally, we added this issue as a limitation and stated that future studies should include validated burnout-specific instruments to directly assess burnout outcomes.

 

Q8 The intro and aims mention lifestyle variables and sociodemographics, but the methods do not specify which were measured (aspects like sleep, physical activity, smoking, alcohol, caffeine, BMI, medication). Since lifestyle factors affect HRV and stress, specifying them and considering them as covariates or exploratory predictors would strengthen analyses and inferences.

R8: As part of the broader assessment, participants also provided lifestyle-related data, including sleep, physical activity, smoking status, alcohol and caffeine consumption, BMI, and medication use. These variables were collected for descriptive and exploratory purposes but were not included in the current analyses because they were outside the scope of the present manuscript.

 

Q9 Please clarify whether questionnaires were completed immediately before/after the rPPG recording and whether there were any delays, as timing affects the interpretation of “associations.”

R9: The rPPG assessment was completed immediately before the self-report questionnaires, within the same assessment session. Therefore, the physiological recording and questionnaire data were collected contingently, with no relevant delay between the two procedures

 

Q10 The authors clearly describe that a mobile rPPG app was used and that a standardized protocol (quiet room, minimal movement, alignment) was followed.

However, recording duration and frame rate not reported, especially for rPPG and HRV-like features, the length of recording (e.g., 30 s vs 2 min) and camera frame rate/resolution are critical. Thus please specify or provide detail on the duration of the facial scan, the approximate frame rate (e.g., 30 fps) and resolution or whether exposure/brightness was fixed or automatic.

Although it is indicated that a “Quiet environment” was experienced, but lighting is not (natural vs artificial, consistency). Thus please clarify whether any quality checks on lighting were done (aspects like minimum brightness threshold), or whether participants were instructed to avoid backlighting or direct sunlight.

It appears only one scan per participant was performed; if so, please indicate as such. Given the dynamic nature of stress, a single brief measurement is a limitation. Thus the authors might consider acknowledging this explicitly and suggesting repeated or longitudinal measures in future designs.

The authors clearly describe face detection, ROI selection, tracking, RGB trace extraction, filtering, and artifact rejection, in line with established rPPG practice.

Yet aspects such as “These may include…” sounds too vague for a methods section, especially phrases like “These may include heart rate, pulse waveform characteristics, and time-/frequency-domain indices…” are generic. Even if the weighting is proprietary, you should clarify which feature classes are definitely used (like mean HR, SDNN-like measures, RMSSD-like measures, spectral power in certain bands). This may aid in improving/ elevating plausibility that indices reflect autonomic regulation.

At present there is no explicit validation of rPPG vs reference in this sample? The authors indicate that the system was previously evaluated against reference devices, but there is no validation against ECG/contact PPG in this sample or a subset, so at present they are inferring PRV/HRV validity from prior work (please cite or indicate intra-and inter variability).

Please add some more detail on data quality, as QC is described conceptually, yet please justify or describe criteria for excluding a scan (aspects like minimum usable signal length, SNR threshold, maximum allowed motion). How many participants’ recordings were excluded or failed and how these were handled in analyses (missing data or study withdrawal).

The authors justifiably state proprietary algorithms and interpretability, but by keeping with the IP, they might add that the training context of the models (trained on internal datasets with lab-based stressors or clinical populations). Whether indices are on standardized scales (0–100), and what higher vs lower values indicate. Where possible, citing internal validation (even if in grey literature or regulatory submissions) or other publications using the same indices would bolster credibility.

At present there is a mismatch between conceptual labels and single-snapshot measurement. Stress Recovery and Stress Response conceptually require a stress-recovery or baseline-stressor design (pre/post or dynamic), but here the authors appear to have only a single short recording under resting instructions. That raises the question: how can the app estimate recovery or reactivity from a single snapshot? Thus please clarify the intended interpretation, as in are these trait-like, derived from waveform morphology/variability at rest (“capacity to recover” inferred from vagal tone proxies) or are they tuned on stress-recovery datasets but applied here as “trait proxies”? Please address more explicitly and highlight the limitation.

The authors indicate that the indices are “intended to capture different aspects” but do not discuss whether they are correlated with each other or whether any prior work has separately validated each of them. Please consider reporting inter-correlations and explaining how strongly they overlap and consider discussing whether “Recovery” and “Response” are truly distinct constructs in this context.

R10: We thank the reviewer for these detailed and constructive comments. We have substantially revised the Methods and Limitations sections to provide a clearer description of the rPPG acquisition protocol, signal-processing framework, data-quality procedures, interpretation of the app-derived indices, and methodological constraints.

Specifically, we clarified that the rPPG assessment was performed using a single facial scan for each participant, under standardized resting instructions embedded in the application. Participants were instructed to sit in a quiet indoor environment, minimize speech and movement, maintain face alignment within the on-screen frame, keep the smartphone stable, and ensure adequate and homogeneous lighting. We also specified that participants were asked to avoid direct sunlight, strong backlighting, marked shadows on the face, and abrupt changes in illumination. Because participants used their own smartphones, camera model, resolution, frame rate, exposure, and brightness settings were not experimentally standardized; this device heterogeneity has now been explicitly acknowledged as a methodological limitation.

We also clarified that only participants with complete data for both the self-report questionnaires and the app-derived rPPG indices were included in the final analytical sample, using a complete-case approach. Participants with missing self-report data or unavailable/invalid app-derived indices were excluded from the corresponding analyses.

Regarding signal processing, we revised the Methods section to reduce vague wording and better specify the feature classes used by the system. We now describe that the app-derived indices are based on rPPG-derived cardiovascular features, including heart rate, pulse waveform characteristics, beat-to-beat pulse interval variability, and HRV/PRV-like time- and frequency-domain parameters such as SDNN-like and RMSSD-like measures, as well as spectral indices related to autonomic regulation. At the same time, we clarified that the exact feature weighting, algorithmic thresholds, signal-quality criteria, and mathematical transformations are proprietary and therefore cannot be fully disclosed.

We further clarified that no concurrent validation against ECG or contact PPG was performed in the present sample. Therefore, the interpretation of the rPPG-derived indices relies on previous validation and feasibility studies of the same mobile rPPG system and related rPPG methodology, rather than on physiological reference measurements collected in this specific sample. This point has been added as a limitation.

Finally, we revised the interpretation of Stress Level, Stress Recovery, and Stress Response. We specified that the indices are provided on a standardized 0–100 scale and were analyzed in their original app-generated form, without additional normalization for age, sex, or other participant-level covariates. Higher Stress Level and Stress Response scores indicate greater physiological stress activation and stronger reactivity, respectively, whereas higher Stress Recovery scores indicate more favorable recovery-related functioning. We also clarified that, given the single-snapshot resting design, Stress Recovery and Stress Response should be interpreted as exploratory app-derived physiological proxies inferred from resting waveform morphology and variability patterns, rather than as direct measures of dynamic recovery or experimentally induced stress reactivity. This limitation is now explicitly acknowledged, and we added that future studies should include repeated, longitudinal, or controlled stress-recovery protocols to validate the distinctiveness and clinical relevance of these indices.

 

Q11 Are the scores normalized for age/ sex etc? What are the index range and directionality(0–100) or  (higher = more stress / poorer recovery).

R11: We thank the reviewer for requesting this clarification. We revised the Methods section to specify the range, directionality, and use of the app-derived indices. All three indices are expressed on a standardized 0–100 scale. Higher scores on Stress Level and Stress Responsecorrespond to greater physiological stress activation and stronger reactivity, respectively. By contrast, higher Stress Recovery scores indicate more favorable recovery-related functioning, whereas lower scores suggest reduced recovery capacity.

We also clarified that, in the present analyses, the indices were analyzed in their original app-generated form and were not additionally normalized or adjusted for age, sex, or other participant-level covariates.

 

Q12 Although it is indicated that an ad hoc questionnaire was performed, there are no psychometric info. Thus a 5-item ad hoc scale, which is acceptable for preliminary work, but please indicate whether items were adapted from any existing usability instruments (e.g., SUS) or designed de novo and please report internal consistency (Cronbach’s alpha) in the results.

R12: We thank the reviewer for this helpful comment. We revised the manuscript to clarify that the usability questionnaire was designed de novo for the purposes of the present study and was not adapted from an existing standardized instrument such as the SUS. We now explicitly state that the five items assessed perceived ease of use, clarity of the information provided, perceived safety and reliability, overall scanning experience, and intention to reuse the application.

In addition, we reported the internal consistency of the usability scale in the present sample. Cronbach’s alpha was α = .67, which was considered acceptable for preliminary and exploratory research purposes. We also clarified that the five items were averaged to obtain an overall usability score, with higher scores indicating a more positive usability evaluation. Since the usability questionnaire was developed specifically for this study, we acknowledge the absence of prior validation as a limitation and interpret these findings cautiously.

 

Q13 Please also clarify whether the averaged item scores were done to obtain an overall usability score and whether higher scores indicate more positive usability. If usability is used in any inferential analyses (aspects like association with indices or symptoms), please highlight this in the statistical analysis plan.

R13: We thank the reviewer for this helpful suggestion. We revised the manuscript to clarify how the usability score was computed and analyzed. Specifically, we now state that the five usability items were averaged to obtain an overall usability score, with higher scores indicating a more positive usability evaluation. We also clarified that usability results were examined descriptively, in line with the exploratory nature of this ad hoc questionnaire, and were not included in the main inferential analyses.

 

Q14 Currently there is no mention of how many participants had complete self-report and usable rPPG data or how missing data were handled (listwise deletion vs. imputation). This should please be explicitly reported and justified.

R14: We clarified that eligibility criteria and the complete-case approach were included in the revised Methods section. Specifically, we added that only participants meeting the inclusion criteria and providing complete data for both the self-report questionnaires and the app-derived rPPG assessment were included in the final analytical sample. Participants with missing self-report data or unavailable/invalid app-derived indices were excluded from the corresponding analyses.

 

Q15 Please detail normality and choice of tests, as it is mentioned that Pearson correlations and ANOVA were done, which assume approximate normality and homoscedasticity. Please add whether assumptions were checked (visual inspection, Shapiro-Wilk) or whether transformations or nonparametric alternatives were considered if assumptions were violated.

Importantly, with the current design the authors compute multiple correlations between three indices and multiple outcomes, and conduct ANOVAs along with regression. Yet at present there is no plan for correction for multiple testing, while exploratory work can sometimes forgo strict correction, the authors should please consider or mention whether they adjusted p-values (like FDR) and, if not, state this as a limitation and interpret significant findings cautiously.

The PCA for  “Burnout-related Emotional Distress” raises some conceptional questions. PCA on just two variables (anxiety and depression) will almost always yield a single component, so PCA adds little beyond simply averaging or standardizing and summing scores. Labelling that component “Burnout-related Emotional Distress” is conceptually stretched if authors did not include any burnout items. Perhaps this may be mitigated by stating that this component reflects shared variance in anxiety and depressive symptoms and is used as a general emotional distress factor. Again, avoid overusing “burnout-related” in the label, or explicitly acknowledge that this is a proxy, not a direct burnout measurement.

Authors state that a multiple regression model was estimated with the three indices as predictors of the component score. Yet some key details missing are covariates included (age, sex, recruitment source, lifestyle variables)? These are important potential confounders of both physiological indices and distress. How was multicollinearity checked (especially if indices are correlated)? Whether residual diagnostics (normality, heteroscedasticity) were performed. Approaches like  hierarchical regression may be considered?

R15: We thank the reviewer for these important methodological and conceptual comments. We revised the Statistical Analyses section to provide a clearer and more complete description of the analytical approach. Specifically, we clarified that assumptions for parametric analyses were assessed before conducting inferential tests. Normality was evaluated through visual inspection of histograms and Q–Q plots and formally assessed using the Shapiro–Wilk test. Homogeneity of variance was examined before ANOVA procedures, while regression diagnostics were inspected to evaluate linearity, normality of residuals, homoscedasticity, and the presence of influential observations. When assumptions were not fully met, results were interpreted cautiously, and non-parametric alternatives were considered as sensitivity checks.

We also clarified the specification of the multiple regression model. The primary model included the three app-derived stress indices, Stress Level, Stress Recovery, and Stress Response, as predictors of General Emotional Distress. Multicollinearity among predictors was evaluated using tolerance values and variance inflation factors. We further acknowledged that age, sex, recruitment source, and lifestyle-related variables may represent potential confounders. In the revised manuscript, we clarified that lifestyle variables were collected for descriptive and exploratory purposes but were not included in the primary regression model because they were outside the main scope of the present study. This issue has been explicitly acknowledged as a limitation, and we noted that future studies should consider hierarchical regression models to assess the incremental predictive value of app-derived stress indices after accounting for sociodemographic and lifestyle-related covariates.

Regarding multiple testing, we now explicitly acknowledge the exploratory nature of the analyses and the increased risk of type I error due to the number of correlations, group comparisons, and regression analyses performed. Since no formal correction for multiple comparisons was applied, this point has been added as a limitation, and statistically significant findings are now interpreted cautiously as hypothesis-generating.

Finally, we revised the terminology and interpretation of the PCA-derived variable. The component previously labeled Burnout-related Emotional Distress has been renamed General Emotional Distress throughout the manuscript. We clarified that this component reflects the shared variance between anxiety and depressive symptoms and should be interpreted as a proxy indicator of burnout-related psychological vulnerability, rather than as a direct measure of burnout.

 

Q16 Discussion

The authors are commended for a well written and theoretically rich discussion, but it over-interprets small, cross-sectional associations, leans heavily into “early detection” and “cumulative burden” language that the data do not support, and underplays key methodological and technological limitations of rPPG and PPG‑based stress indices.

Lines 324-336:

Authors indicate that “The present study provides preliminary evidence supporting the usability, acceptability, and construct validity…”
“…consistent association… suggests that the app is not merely capturing transient fluctuations in stress, but may reflect more stable and clinically meaningful dimensions…”

However, this overstates the construct validity and design, as the correlations of Stress Level with anxiety, depression, and the composite factor are statistically significant but small (r ≈ .13–.17), and only one of three indices showed consistent associations. At present the “Construct validity” suggests strong, replicated alignment between the physiological indices and the intended construct (stress/burnout), which your results do not yet support. The associations are preliminary convergent signals, not full construct validation.

Indicating “Consistent association” implies robustness the study currently does not have, with three outcomes, three indices; only Stress Level shows meaningful associations. At present the discussion treats this as broadly consistent across indices and outcomes, which glosses over the fact that Stress Recovery and Stress Response did not behave as theorized. At present there is no test-retest data, no longitudinal follow‑up, and no clinical diagnoses; thus the authors should please be cautious with inferences, as at present they cannot show that the index reflects stable vulnerability or clinically meaningful traits vs. just current physiological arousal

R16: We thank the Reviewer for this important observation. In response, we carefully revised the Discussion to avoid overinterpretation of the findings and to ensure that the conclusions remain fully aligned with the actual scope and limitations of the data. Specifically, we moderated several statements that could have implied stronger evidence than supported by the present cross-sectional design and preliminary validation framework.

In particular, we revised expressions referring to “early detection,” “cumulative burden,” and “preventive applications” to better clarify that these implications remain hypothetical and should be considered as potential future directions rather than demonstrated outcomes of the present study. We now explicitly state that the observed associations between the Stress Level index and psychological symptoms provide preliminary evidence of construct validity and convergence with self-reported emotional distress, but do not demonstrate predictive validity, causal mechanisms, or clinical utility for identifying future burnout or mental disorders.

We also modified the interpretation of the Stress Level index to avoid implying that the app directly measures chronic stress burden or cumulative psychophysiological load. The revised text now specifies that the app-derived indices should be interpreted as indirect physiological correlates potentially associated with stress-related autonomic regulation, rather than objective biomarkers of psychological stress or burnout themselves. In the same way, we clarified that the observed correlations with anxiety, depression, and burnout-related emotional distress indicate associations with concurrent psychological dimensions, but do not allow conclusions regarding long-term vulnerability trajectories or future clinical outcomes.

Furthermore, we substantially strengthened the cautionary language regarding the technological and methodological limitations of rPPG-based stress monitoring. The revised manuscript now more explicitly acknowledges that physiological activation inferred from cardiovascular features may reflect multiple overlapping psychophysiological processes and should therefore be interpreted within a broader multidimensional assessment framework integrating subjective, behavioral, and contextual information.

Finally, we expanded the Limitations section to further emphasize that the current findings should be considered preliminary evidence supporting feasibility, usability, and construct validity only. We explicitly added that longitudinal studies, repeated-measure designs, and validation against established physiological stress markers are necessary before drawing conclusions regarding predictive capacity, preventive utility, or clinical implementation. These revisions were introduced throughout the Discussion and Limitations sections to ensure a more cautious and scientifically balanced interpretation of the results.

 

Q17 Psychobiological interpretation and “cumulative burden” (lines 337–347)

Authors specifically state that “…emphasize the central role of chronic stress exposure and allostatic load…”
“…Stress Level… may be interpreted as a proxy marker of cumulative psychophysiological burden rather than a simple indicator of momentary stress…”

Importantly, allostatic load or parallels were not as at present there we no standard allostatic load markers (e.g., cortisol, inflammatory markers, metabolic parameters). Thus the interpretation that Stress Level is a proxy of cumulative burden is not supported by any longitudinal or multi‑system biomarker data. Chronic vs. acute stress is not distinguishable at present, given one short rPPG recording and cross‑sectional symptom scores. Indeed, rPPG provides pulse rate variability (PRV), which is an imperfect proxy of ECG‑derived HRV and may diverge especially in some physiological states. The current Discussion assumes that rPPG indices map directly onto HRV‑based allostatic load models without mentioning measurement differences or absence of ECG validation in this sample.

The authors might consider presenting the psychobiological model as a theoretical framework that their findings are compatible with, rather than claiming Stress Level is a proxy of cumulative burden, where they used rPPG‑derived PRV, not gold‑standard HRV, and did not include direct allostatic load biomarkers. Again, cross‑sectional, single‑snapshot data cannot distinguish chronic from acute stress.

R17: We thank the Reviewer for this important observation. In response, we revised the Discussion to avoid overinterpretation of the findings and to ensure that the conclusions remain fully aligned with the actual scope and limitations of the present study. Specifically, we moderated statements referring to “allostatic load,” “cumulative burden,” and chronic stress dysregulation, as these concepts were not directly assessed in our study.

The revised text now clarifies that, although the observed associations are broadly consistent with literature linking stress-related physiological activation to emotional distress, the present findings do not demonstrate that the app directly measures chronic stress burden, allostatic load, or cumulative psychophysiological dysregulation. We also explicitly acknowledged that no standard biological stress markers (e.g., cortisol or inflammatory markers) or reference physiological measures (e.g., ECG-derived HRV or electrodermal activity) were collected. Accordingly, we revised the interpretation of the Stress Level index, specifying that it should be considered a preliminary app-derived physiological correlate potentially associated with stress-related autonomic activation, rather than a proxy marker of cumulative psychophysiological burden.

Furthermore, we strengthened the cautionary language throughout the Discussion by emphasizing that the cross-sectional design does not allow causal inference or conclusions regarding chronic stress accumulation or future mental health outcomes. These revisions were introduced to provide a more scientifically balanced interpretation of the findings and to better reflect the preliminary nature of the current validation study.

 

Q18 Stress Level vs. Stress Recovery and Stress Response (lines 348–359)

The authors specifically argue that “Stress Level emerged as the only significant predictor…”
“…indices are highly correlated and partially overlapping… suggests that Stress Level may capture the broader common component of stress burden…”

However, the Recovery and Response are defined as reflecting recovery capacity and reactivity, but, please note the limitation as only one resting recording was collected and there was no stress‑induction or recovery phase included? Additionally, the collinearity explanation is partial, as the authors correctly mention shared variance, but please quantify intercorrelations and discuss whether multicollinearity diagnostics (e.g., VIF) were checked.

Importantly, with a single time‑point at rest, one cannot directly test “response” or “recovery” constructs; this limits the interpretation of those indices, and the non‑significant results for Recovery and Response may reflect both conceptual and measurement mismatches, not just collinearity. Thus the authors are kindly requested to reconsider and perhaps frame Stress Level as the index that showed the clearest statistical signal in this design, rather than as definitively capturing “stress burden.”

R18: We thank the Reviewer for this comment. We revised the Discussion to clarify that Stress Level was the index most closely associated with concurrent emotional distress, while avoiding overinterpretation. We now acknowledge the possible overlap among the three indices and state that, due to the cross-sectional design, these findings support preliminary concurrent validity but do not establish predictive validity or causal relationships with future burnout or mental health outcomes.

 

Q19 Burnout, comorbidity, and “early psychophysiological signals” (lines 360–368)
The authors again indicate, correctly, that “…high rates of comorbidity between burnout, anxiety, and depression…”
“…present findings support the notion that r-PPG comestai_app may capture early psychophysiological signals… related to this broader spectrum of stress-related pathology…”

However, burnout was not directly measured, as the investigators did not administer a validated burnout scale; instead, they created a factor from anxiety and depression scores. Thus please be cognisant and careful in calling this “Burnout-related Emotional Distress” risks implying a directly measured burnout when in actuality the authors measured general emotional distress.

At present, there is no longitudinal follow‑up, no clinical diagnoses, no outcome data. Thus, there is no evidence that app indices change before clinical burnout or depressive episodes, at present there are only cross‑sectional associations. In accordance with this, statements about somatic comorbidities and premature mortality are true at population level, but the present data do not address these outcomes.

To assist in this, the authors might consider reframing the factor as a general emotional distress indicator that likely overlaps with burnout‑related exhaustion, and be explicit that burnout per se was not measured, and present “early signals” as a hypothesis for future longitudinal work, not as something supported by the current results.

R19: We thank the Reviewer for highlighting this issue. We have refined the discussion to better delimit the interpretation of the Stress Level index. The revised text now makes clear that the associations observed with anxiety, depression, and burnout-related emotional distress reflect concurrent construct validity only. We also specify that, in the absence of biological stress markers, reference physiological recordings, and longitudinal follow-up, Stress Level should be regarded as an app-derived physiological correlate potentially linked to stress-related autonomic activation, not as evidence of chronic stress burden, allostatic load, diagnostic utility, or prognostic value.

 

Q20 Preventive and applied implications (lines 369–379)

Statements like “…continuously and non-invasively monitoring stress…”
“…identify individuals who are progressively accumulating high stress levels before the clinical onset of burnout or depressive disorders…”

Again, the present study uses a single short scan in a controlled setting, yet the discussion suggests continuous or frequent real‑world monitoring, which is not what was implemented or evaluated. As at present there is no examination or prediction of future burnout or depression, trajectory of stress accumulation, or sensitivity to change. Claims about identifying individuals “before clinical onset” are not supported by cross‑sectional correlations, although these correlations are important.

Perhaps the authors can describe preventive potential as conditional on future longitudinal and real‑world validation, using words like“could, if validated” rather than “would” or “allows.” This may clarify that the present study demonstrated feasibility of single-session use, not continuous deployment or predictive screening.

R20: We thank the Reviewer for this important comment. In response, we revised the Discussion to more clearly delimit the interpretation of the preventive and monitoring implications of the app. Specifically, we now explicitly state that the present study was cross-sectional and based on single spot assessments, and therefore does not demonstrate the ability of the app to identify early vulnerability trajectories, predict future mental health outcomes, or guide preventive interventions in clinical practice. We further clarified that these applications remain promising but hypothetical and require longitudinal and interventional validation before clinical implementation can be considered.

 

Q21 “Moment-to-moment indicators” and proactive care (lines 380–388)

Again, given the design, the temporal resolution may be overstated. Indicating “moment‑to‑moment” and “continuous screening” imply high-frequency monitoring over time. At present there are no repeated measures, thus no within-person trajectories, accumulation patterns, or recovery slopes. The intervention linkage is speculative, albeit true. Please be cautious here, as the study did not test any intervention, digital coaching, or clinical decision support tied to the app indices.

Perhaps rather position these as possible future applications that need testing, and explicitly say that current data do not show whether specific thresholds or patterns in the indices meaningfully guide interventions or improve outcomes.

R21:  We have revised the Discussion to avoid overstating the monitoring and preventive implications of the app. The revised text now specifies that, because the study was cross-sectional and based on a single spot assessment, the findings do not demonstrate the ability to capture temporal changes in stress accumulation or recovery, identify vulnerability trajectories, predict future mental health outcomes, or guide preventive interventions. We now state that these applications remain promising but hypothetical and require longitudinal, repeated-measure, and interventional validation before clinical implementation

 

Q22 Multimodal assessment and “clinical validity” (lines 393–399)

Statements like  “…essential in psychosomatic medicine… convergence… supports the ecological and clinical validity…” do not allow for “Clinical validity” , as the sample population was non‑clinical; measures were screening scales; and indices were unvalidated against clinical endpoints. Thus there is a preliminary convergent validity with symptoms, but no clinical diagnoses, no sensitivity/specificity analyses, and no prognostic evaluation. Please also note, which is always the case with stress-related studies, the impact of ecological validity. Measurements were not truly ambulatory or context‑rich (e.g., multiple daily assessments over weeks). Thus, ecological validity is plausible conceptually (smartphone outside lab), but empirically limited to a one‑off controlled session.

The authors might consider replacing “clinical validity” with “preliminary convergent validity” and be explicit about the non‑clinical nature of the sample. Perhaps also clarify that ecological validity refers here to app use outside a traditional lab, but not yet to fully passive or context-aware sensing.

R22: We revised the Limitations section to state more explicitly that the findings support only preliminary feasibility and construct validity. We also clarified that the rPPG-derived indices should currently be considered exploratory correlates of concurrent distress, not validated clinical markers or predictors of future risk.

 

Q23 Age and gender differences (lines 400–405)

“…younger individuals and women showing higher stress levels… consistent with literature… results further support potential usefulness in vulnerable populations.”

Although these statements are true, based on current literature, yet it is not clear if age/sex differences in Stress Level remain after controlling for mental health symptoms, recruitment source, or other variables. Thus, without adjustment, group differences might simply reflect symptom differences or sample composition. In line with this, algorithmic and physiological biases are not discussed a present, as autonomic parameters and rPPG quality can vary by age/sex, and there is emerging concern about demographic bias in physiological algorithms, please consider acknowledging g these complexities.

R 23: We thank the Reviewer for this comment. We revised the discussion to clarify that age and gender differences are exploratory, as these variables were partly intertwined with recruitment source and student status. We also added that balanced and stratified samples are needed to determine whether age and gender independently influence rPPG-derived stress indices.

 

Q24 Conclusions and interpretation

Burnout “risk” language as cross-sectional correlations with burnout-related distress do not establish risk. Risk implies prediction of future outcomes. Perhaps the authors can consider something more along the lines of “…for monitoring burnout-related emotional distress” or “…for assessing associations with burnout-related emotional distress” rather than “burnout risk monitoring.

R 24 The conclusions were also revised to avoid overstating the implications of the findings. Specifically, we moderated expressions referring to “early detection” and “continuous monitoring” and clarified that the present study provides only preliminary evidence based on single spot assessments. We now state that future longitudinal and repeated-measure studies are needed before drawing conclusions regarding continuous monitoring capabilities or burnout risk prediction.

 

Author Response File: Author Response.pdf

Reviewer 2 Report

Comments and Suggestions for Authors
  • Page 1, Lines 17–26: The highlights section contains overstated implications considering the preliminary and cross-sectional nature of the study. Claims regarding “early detection” and “preventive mental health care” are not sufficiently supported by the presented data.
  • Page 1, Lines 29–32: The abstract introduces rPPG as an objective stress monitoring tool, yet the manuscript does not provide sufficient physiological validation against gold-standard stress biomarkers such as cortisol or ECG-derived HRV.
  • Page 2, Lines 38–42: The authors do not report the duration of facial recordings, smartphone specifications, lighting conditions, or environmental variability during data acquisition. These factors directly affect rPPG signal quality.
  • Page 2, Lines 44–48: The reported correlations are statistically significant but extremely weak. The manuscript repeatedly interprets these small effect sizes as clinically meaningful without adequate justification.
  • Page 2, Lines 46–47: The regression coefficient for Stress Level appears modest, yet the discussion later presents this finding as strong evidence of burnout prediction. This interpretation is not proportional to the observed effect size.
  • Page 2, Lines 50–52: The conclusion that the app demonstrates construct validity is premature because the study lacks criterion validity testing against established physiological measures.
  • Page 2, Lines 57–70: The introduction is lengthy and contains broad background statements on stress and mental health that are not directly linked to the study objective. The section should be more focused.
  • Page 3, Lines 73–77: The manuscript criticizes self-report measures for subjectivity, but the study outcome itself heavily depends on self-report questionnaires. This contradiction should be acknowledged.
  • Page 3, Lines 78–87: The literature review lacks a balanced discussion of known limitations of rPPG technologies, including susceptibility to motion artifacts, skin tone variation, ambient light interference, and device heterogeneity.
  • Page 3, Lines 95–99: Statements regarding preventive mental health applications are speculative because the study did not test preventive interventions or longitudinal outcomes.
  • Page 3, Lines 100–107: The aims are overly broad for a preliminary validation study. The manuscript attempts to address usability, stress profiling, psychometric validation, and burnout prediction simultaneously without sufficient methodological depth.
  • Page 4, Lines 112–115: Consecutive recruitment from a commercial panel and university students introduces substantial selection bias. The representativeness of the sample is questionable.
  • Page 4, Lines 112–115: The manuscript does not report inclusion or exclusion criteria. It is unclear whether participants with psychiatric disorders, cardiovascular disease, medication use, or substance use were excluded.
  • Page 4, Lines 116–118: Ethical approval details are insufficiently integrated into the methods section. The manuscript should clearly state the ethics committee name and approval number earlier in the methods.
  • Page 4, Lines 123–126: The authors describe the questionnaires as “well-validated” but do not report reliability estimates such as Cronbach’s alpha within the current sample.
  • Page 4, Lines 129–139: Important technical details are missing, including camera frame rate, image resolution, sampling frequency, signal acquisition duration, and minimum acceptable signal quality thresholds.
  • Page 5, Lines 143–180: The repeated reliance on “commercially protected” algorithms substantially limits reproducibility and scientific transparency. Readers cannot evaluate how the indices were actually generated.
  • Page 5, Lines 164–176: The manuscript vaguely refers to “time- and/or frequency-domain indices” without specifying which HRV or cardiovascular metrics were extracted.
  • Page 5, Lines 169–180: The use of undisclosed proprietary models raises concerns about reproducibility and independent validation. This limitation is more serious than currently acknowledged.
  • Page 6, Lines 188–197: The operational definitions of Stress Level, Stress Recovery, and Stress Response remain conceptually vague. The manuscript does not explain how these constructs differ physiologically.
  • Page 6, Lines 200–204: The usability questionnaire was ad hoc and not validated. This weakens the credibility of the usability findings.
  • Page 6, Lines 207–233: The statistical analysis section lacks information about missing data handling, assumption testing, normality assessment, multicollinearity diagnostics, and correction for multiple comparisons.
  • Page 7, Lines 223–226: The use of PCA with only anxiety and depression variables is methodologically questionable. PCA is generally not appropriate with only two indicators.
  • Page 7, Lines 223–226: The term “Burnout-related Emotional Distress” is misleading because burnout was not directly measured using validated burnout instruments such as the Maslach Burnout Inventory.
  • Page 7, Lines 228–231: The regression model omitted potentially important confounders such as age, gender, occupation, sleep quality, medication use, and physical health status.
  • Page 7, Lines 236–244: The sample composition is highly skewed toward students and younger adults, limiting external validity.
  • Page 7, Lines 248–255: The usability results are presented descriptively without inferential analysis or internal consistency evaluation.
  • Page 8, Lines 259–275: The stress indices are categorized into groups 1, 2, and 3, but the manuscript does not explain how these cutoffs were determined.
  • Page 8, Figure 1: The figure quality is low and lacks adequate labeling. The y-axis and group definitions are insufficiently explained.
  • Page 9, Lines 278–280: The use of chi-square testing for stress scores appears inappropriate unless the variables were categorized. The analytical strategy requires clarification.
  • Page 9, Lines 283–286: The correlations between stress level and psychological symptoms are weak and may lack practical significance despite statistical significance.
  • Page 9, Figure 2: The error bars are not clearly described. The figure legend should specify whether they represent standard deviation, standard error, or confidence intervals.
  • Page 9, Lines 294–299: The manuscript repeatedly refers to “burnout-related emotional distress” despite the absence of a direct burnout assessment instrument. This terminology should be revised throughout the paper.
  • Page 10, Lines 307–320: The regression model explains only 4% of the variance, indicating limited predictive utility. The discussion substantially overstates the practical importance of the findings.
  • Page 10, Figure 4: The figure provides limited additional information beyond the regression statistics already reported in the text. The manuscript may benefit from consolidating figures.
  • Page 11, Lines 323–328: The discussion begins with strong claims regarding construct validity and usability despite the preliminary nature of the data and limited methodological transparency.
  • Page 11, Lines 330–345: The manuscript interprets weak correlations as evidence of stable emotional vulnerability and allostatic burden. These interpretations extend beyond the available evidence.
  • Page 11, Lines 347–355: The explanation regarding overlapping stress indices and shared variance should be supported with multicollinearity statistics such as VIF or tolerance values.
  • Page 12, Lines 368–388: The practical implementation discussion is highly speculative because the study did not evaluate intervention effectiveness, adherence, or long-term monitoring outcomes.
  • Page 12, Lines 389–391: Ethical considerations are discussed only briefly despite the collection of facial physiological data through smartphones. Data privacy and algorithmic transparency deserve deeper consideration.
  • Page 12, Lines 392–398: The manuscript repeatedly claims ecological validity without conducting real-world longitudinal monitoring or ecological momentary assessment.
  • Page 13, Lines 407–416: The limitations section is incomplete because it does not sufficiently discuss the implications of proprietary algorithms, selection bias, weak effect sizes, and lack of physiological validation.
  • Page 13, Lines 417–421: The generalizability limitation is understated considering that most participants were students and young adults from a single country.
  • Page 13, Lines 423–425: Future directions are broad and ambitious relative to the limited evidence provided in the present study.
  • Page 13, Lines 427–443: The conclusions substantially overstate the public health implications of the findings, given the small effect sizes and preliminary design.
  • Pages 15–18, References: Several references are very recent narrative reviews or technology-oriented discussions, while fewer foundational psychophysiological validation studies are included. The literature balance could be improved.
Comments on the Quality of English Language

The manuscript would benefit from substantial English language editing to improve grammar, sentence structure, and overall readability.

Author Response

Reviewer 2

  • Page 1, Lines 17–26: The highlights section contains overstated implications considering the preliminary and cross-sectional nature of the study. Claims regarding “early detection” and “preventive mental health care” are not sufficiently supported by the presented data.

R1: We thank the reviewer for this important comment. We revised the Abstract and related sections to avoid overstating the implications of the findings. Specifically, we removed or softened claims related to “early detection” and “preventive mental health care” and now describe the results as preliminary evidence supporting the feasibility and exploratory construct validity of a smartphone-based rPPG approach for psychophysiological stress monitoring. We also clarified that the cross-sectional design does not allow conclusions regarding prediction, prevention, or clinical implementation.

 

  • Page 1, Lines 29–32: The abstract introduces rPPG as an objective stress monitoring tool, yet the manuscript does not provide sufficient physiological validation against gold-standard stress biomarkers such as cortisol or ECG-derived HRV.

R2: We agree with the reviewer. We revised the wording in the Abstract and throughout the manuscript to avoid referring to the app-derived indices as an “objective stress monitoring tool” in a strong clinical sense. The indices are now described as noninvasive app-derived physiological proxies or exploratory rPPG-derived indicators of stress-related physiological activation. We also clarified that no concurrent validation against ECG or contact PPG was performed in the present sample.

 

  • Page 2, Lines 38–42: The authors do not report the duration of facial recordings, smartphone specifications, lighting conditions, or environmental variability during data acquisition. These factors directly affect rPPG signal quality.

R3: We revised the Digital Assessment and Data Acquisition section to provide additional details on the rPPG acquisition protocol. Specifically, we clarified that participants completed a single facial scan using their own smartphones and that acquisition followed standardized instructions embedded in the application. We also specified that participants were asked to sit in a quiet indoor environment, minimize speech and movement, maintain face alignment within the on-screen frame, keep the smartphone stable, and ensure adequate and homogeneous lighting. Participants were instructed to avoid direct sunlight, strong backlighting, marked shadows on the face, and abrupt changes in illumination. We further clarified that smartphone model, camera specifications, frame rate, resolution, exposure, brightness settings, and operating system were not experimentally standardized, and we acknowledged these factors as potential sources of variability in rPPG signal quality.

 

  • Page 2, Lines 44–48: The reported correlations are statistically significant but extremely weak. The manuscript repeatedly interprets these small effect sizes as clinically meaningful without adequate justification.

R4: We thank the reviewer for this important observation. We revised the Results and Discussion sections to provide a more cautious interpretation of the statistically significant but weak associations. We now explicitly state that the effect sizes were small and that these findings should be interpreted as preliminary and hypothesis-generating rather than as strong evidence of clinically meaningful associations. We also clarified that statistical significance may partly reflect sample size and that the practical relevance of these associations requires confirmation in future studies.

 

  • Page 2, Lines 46–47: The regression coefficient for Stress Level appears modest, yet the discussion later presents this finding as strong evidence of burnout prediction. This interpretation is not proportional to the observed effect size.

R5: We agree with the reviewer. We revised the text to avoid presenting the modest Stress Level effect as strong evidence of burnout prediction. The finding is now described as a modest cross-sectional association with General Emotional Distress, and we clarified that longitudinal studies are needed before any predictive conclusions can be drawn.

 

  • Page 2, Lines 50–52: The conclusion that the app demonstrates construct validity is premature because the study lacks criterion validity testing against established physiological measures.

R6: We thank the reviewer for this clarification. We revised the conclusion to avoid stating that the app demonstrated construct validity. The text now refers to preliminary evidence consistent with exploratory construct validity, while explicitly acknowledging that validation against established physiological reference measures is needed before stronger conclusions can be drawn.

 

  • Page 2, Lines 57–70: The introduction is lengthy and contains broad background statements on stress and mental health that are not directly linked to the study objective. The section should be more focused.

R7: We thank the reviewer for this suggestion. We revised the Introduction to make it more concise and focused on the study aims. Broad background statements on stress and mental health were shortened, while the rationale for examining app-derived rPPG indices in relation to self-reported emotional distress was clarified.

 

  • Page 3, Lines 73–77: The manuscript criticizes self-report measures for subjectivity, but the study outcome itself heavily depends on self-report questionnaires. This contradiction should be acknowledged.

R8: We agree with the reviewer. We revised the manuscript to acknowledge that the study outcome relies on self-report questionnaires and therefore does not overcome the limitations of self-report assessment. We now frame the app-derived indices as complementary physiological measures to be interpreted alongside self-reported emotional distress, rather than as replacements for validated psychological questionnaires.

 

  • Page 3, Lines 78–87: The literature review lacks a balanced discussion of known limitations of rPPG technologies, including susceptibility to motion artifacts, skin tone variation, ambient light interference, and device heterogeneity.

R9: we include additional references on the limitation of rPPG

 

  • Page 3, Lines 95–99: Statements regarding preventive mental health applications are speculative because the study did not test preventive interventions or longitudinal outcomes.

R10: We agree with the reviewer. We revised the manuscript to avoid speculative claims regarding preventive mental health applications. The text now clarifies that the study did not test preventive interventions or longitudinal outcomes, and that potential preventive applications require future longitudinal and intervention-based studies.

 

  • Page 3, Lines 100–107: The aims are overly broad for a preliminary validation The manuscript attempts to address usability, stress profiling, psychometric validation, and burnout prediction simultaneously without sufficient methodological depth.

R11: We revised the aims to better reflect the exploratory nature of the study. The manuscript now frames the work as an exploratory assessment of feasibility and preliminary construct validity of app-derived rPPG stress indices in relation to self-reported emotional distress, rather than as a study addressing clinical utility, psychometric validation, or burnout prediction.

 

  • Page 4, Lines 112–115: Consecutive recruitment from a commercial panel and university students introduces substantial selection bias. The representativeness of the sample is questionable.

R12: We now clarify that the sample may not be fully representative of the general population and that the findings should be interpreted as preliminary and not directly generalizable to broader or clinical populations.

 

  • Page 4, Lines 112–115: The manuscript does not report inclusion or exclusion criteria. It is unclear whether participants with psychiatric disorders, cardiovascular disease, medication use, or substance use were excluded.

R13: We thank the reviewer for this helpful comment. We revised the Participants section to explicitly report the inclusion and exclusion criteria. Specifically, we clarified that participants were eligible if they were aged 18 years or older, were able to understand the study procedures, provided written informed consent, and completed both the self-report questionnaires and the app-based rPPG assessment. We also specified that participants were excluded from the final analytical sample if they did not provide informed consent, had incomplete self-report data, had unavailable or invalid app-derived rPPG indices, or declared psychiatric disorders, cardiovascular disease, medication use, substance use, or other conditions that could substantially affect physiological stress assessment. Thus, only participants with complete self-report and app-derived rPPG data and without declared exclusion conditions were included in the analyses.

 

  • Page 4, Lines 116–118: Ethical approval details are insufficiently integrated into the methods section. The manuscript should clearly state the ethics committee name and approval number earlier in the methods.

R14: We thank the reviewer for this comment. As required by the journal format, the Institutional Review Board statementis reported in the dedicated section after the Discussion.

 

  • Page 4, Lines 123–126: The authors describe the questionnaires as “well-validated” but do not report reliability estimates such as Cronbach’s alpha within the current sample.

R15: We thank the reviewer for this helpful comment. We have now included the reliability estimates, including Cronbach’s alpha values calculated within the current sample, to better support the psychometric adequacy of the questionnaires used in the study.

 

  • Page 4, Lines 129–139: Important technical details are missing, including camera frame rate, image resolution, sampling frequency, signal acquisition duration, and minimum acceptable signal quality thresholds.

R16: We revised the Digital Assessment and Data Acquisition section to provide additional technical details on the rPPG acquisition protocol. Specifically, we clarified that participants completed a single facial scan using the smartphone front-facing camera and that acquisition followed the standardized instructions embedded in the application. We also specified that participants used their own smartphones; therefore, device-specific parameters such as camera model, resolution, frame rate, exposure, brightness settings, and operating system were not experimentally standardized. These aspects are now explicitly acknowledged as potential sources of variability affecting rPPG signal quality.

 

  • Page 5, Lines 143–180: The repeated reliance on “commercially protected” algorithms substantially limits reproducibility and scientific transparency. Readers cannot evaluate how the indices were actually generated.

R17: We agree with the reviewer that the proprietary nature of the algorithms limits reproducibility and transparency. We revised the manuscript to clarify that, although the exact feature weighting, algorithmic thresholds, and mathematical transformations cannot be fully disclosed due to commercial protection, the general rPPG signal-processing framework and the main physiological feature classes used by the system are now described in greater detail. We also explicitly acknowledged this limitation in the revised Limitations section.

 

  • Page 5, Lines 164–176: The manuscript vaguely refers to “time- and/or frequency-domain indices” without specifying which HRV or cardiovascular metrics were extracted.

R18: We revised the Signal Processing and Stress Index Computation section to reduce vague wording and provide more specific information on the physiological feature classes involved. We now state that the app-derived indices are based on rPPG-derived cardiovascular features, including heart rate, pulse waveform characteristics, beat-to-beat pulse interval variability, and HRV/PRV-like parameters such as SDNN-like and RMSSD-like measures, as well as frequency-domain indices related to autonomic regulation. We also clarified that the exact implementation and weighting of these features remain proprietary.

 

  • Page 5, Lines 169–180: The use of undisclosed proprietary models raises concerns about reproducibility and independent validation. This limitation is more serious than currently acknowledged.

R19: We have strengthened the discussion of this limitation in the revised manuscript. Specifically, we now state that the app-derived indices should be interpreted as exploratory physiological proxies rather than fully independently validated clinical markers. We also clarified that no concurrent validation against ECG or contact PPG was performed in the present sample, and that the interpretation of the indices relies on previous validation and feasibility work on the same rPPG-based system. This limitation has been explicitly emphasized in the Limitations section.

 

  • Page 6, Lines 188–197: The operational definitions of Stress Level, Stress Recovery, and Stress Response remain conceptually vague. The manuscript does not explain how these constructs differ physiologically.

R20: We thank the reviewer for this suggestion. We have added further technical details to the Digital Assessment and Data Acquisitionand Signal Processing and Stress Index Computation

 

  • Page 6, Lines 200–204: The usability questionnaire was ad hoc and not validated. This weakens the credibility of the usability findings.

R21: we  revised the manuscript to clarify that the usability questionnaire was developed ad hoc for the present study and has not undergone prior psychometric validation. Accordingly, usability findings are now presented as exploratory and interpreted cautiously. We also reported the internal consistency of the scale in the present sample and acknowledged the absence of prior validation as a limitation.

 

  • Page 6, Lines 207–233: The statistical analysis section lacks information about missing data handling, assumption testing, normality assessment, multicollinearity diagnostics, and correction for multiple comparisons.

R22: We revised the Statistical Analyses section to provide additional information on missing data handling, assumption checking, normality assessment, multicollinearity diagnostics, and correction for multiple comparisons. Specifically, we clarified that analyses were conducted using a complete-case approach, assumptions for parametric tests were checked, regression diagnostics were inspected, multicollinearity was assessed using tolerance and variance inflation factors, and the absence of formal correction for multiple testing was acknowledged as a limitation.

 

  • Page 7, Lines 223–226: The use of PCA with only anxiety and depression variables is methodologically questionable. PCA is generally not appropriate with only two indicators.

R23: We thank the reviewer for this important methodological clarification. We agree that applying PCA to only two variables has limited added value and may be methodologically questionable. Accordingly, we revised the manuscript to avoid presenting the PCA-derived score as a robust latent component.

In the revised version, anxiety and depressive symptom scores are described as being combined into a General Emotional Distress score reflecting their shared variance. We clarified that this score should be interpreted as a pragmatic summary indicator of emotional distress, rather than as a distinct latent construct derived from a full dimensional reduction procedure. This limitation has also been acknowledged in the manuscript.

 

  • Page 7, Lines 223–226: The term “Burnout-related Emotional Distress” is misleading because burnout was not directly measured using validated burnout instruments such as the Maslach Burnout Inventory.

R24: We revised the manuscript to avoid using the label “Burnout-related Emotional Distress”, as burnout was not directly assessed with a validated burnout-specific instrument. The variable has been renamed General Emotional Distress, and we clarified that it reflects shared anxiety and depressive symptoms rather than burnout. We also acknowledged that future studies should include validated burnout measures, to directly assess burnout-related outcomes.

 

  • Page 7, Lines 228–231: The regression model omitted potentially important confounders such as age, gender, occupation, sleep quality, medication use, and physical health status.

R25: in the revised version, we clarified that these variables were not included in the primary regression model and acknowledged this as a limitation. We now state that future studies should include more comprehensive covariate adjustment, ideally using hierarchical regression models, to test whether app-derived stress indices explain additional variance beyond sociodemographic, lifestyle, and health-related factors.

 

  • Page 7, Lines 236–244: The sample composition is highly skewed toward students and younger adults, limiting external validity.

R26: We thank the reviewer for this observation. We revised the manuscript to acknowledge that the sample composition, which included a large proportion of university students and younger adults, may limit external validity and generalizability. This limitation has been added to the Discussion/Limitations section, and we now clarify that the findings should be interpreted as preliminary and may not extend to older adults, clinical populations, or more demographically diverse samples.

 

  • Page 7, Lines 248–255: The usability results are presented descriptively without inferential analysis or internal consistency evaluation.

R27: We thank the reviewer for this comment. We revised the manuscript to clarify that usability was assessed descriptively because the questionnaire was developed ad hoc for the present study and was intended for exploratory evaluation only. We also reported the internal consistency of the usability scale in the present sample and clarified how the overall usability score was computed.

 

  • Page 8, Lines 259–275: The stress indices are categorized into groups 1, 2, and 3, but the manuscript does not explain how these cutoffs were determined.

R28: We thank the reviewer for requesting this clarification. We revised the manuscript and Figure 1 caption to specify that groups 1, 2, and 3 correspond to descriptive categories provided by the application: 1 = low, 2 = medium, and 3 = high. We also clarified the directionality of the indices and explained that these categories were used for descriptive purposes to facilitate interpretation of the percentage distributions.

 

  • Page 8, Figure 1: The figure quality is low and lacks adequate labeling. The y-axis and group definitions are insufficiently explained.

R29: We revised the figure to improve its visual quality and to provide a clearer explanation of the reported results.

 

  • Page 9, Lines 278–280: The use of chi-square testing for stress scores appears inappropriate unless the variables were categorized. The analytical strategy requires clarification.

R30: We thank the reviewer for this methodological clarification. We revised the Statistical Analyses section to specify that chi-square tests were used only when app-derived stress indices were analyzed as categorical variables, based on the descriptive groups provided by the application. Continuous stress index scores were analyzed using appropriate parametric or non-parametric tests, depending on distributional assumptions.

 

  • Page 9, Lines 283–286: The correlations between stress level and psychological symptoms are weak and may lack practical significance despite statistical significance.

R31:  We revised the Results and Discussion sections to emphasize that the observed correlations were weak in magnitude and should be interpreted cautiously. We now clarify that statistical significance does not necessarily imply practical or clinical relevance, and that these findings should be considered preliminary and hypothesis-generating.

 

  • Page 9, Figure 2: The error bars are not clearly described. The figure legend should specify whether they represent standard deviation, standard error, or confidence intervals.

R32: We improved the figure and clarified its explanation to enhance the readability and interpretation of the reported results.

 

  • Page 9, Lines 294–299: The manuscript repeatedly refers to “burnout-related emotional distress” despite the absence of a direct burnout assessment instrument. This terminology should be revised throughout the paper

R33: We revised the terminology throughout the manuscript and replaced “Burnout-related Emotional Distress” with “General Emotional Distress.” We also clarified that this variable reflects the shared variance between anxiety and depressive symptoms and should not be interpreted as a direct measure of burnout, as no validated burnout-specific instrument was administered.

 

  • Page 10, Lines 307–320: The regression model explains only 4% of the variance, indicating limited predictive utility. The discussion substantially overstates the practical importance of the findings.

R34: We revised the Discussion to interpret the regression findings more cautiously. We now explicitly state that the model explained a small proportion of variance and that the predictive utility of the app-derived indices is limited in the present cross-sectional design. The findings are now framed as preliminary and hypothesis-generating, rather than as evidence of strong practical or clinical predictive value.

 

  • Page 10, Figure 4: The figure provides limited additional information beyond the regression statistics already reported in the text. The manuscript may benefit from consolidating figures.

R35: We thank the reviewer for this observation. We have revised the section by removing Figure 4 and expanding the text to provide a clearer and more detailed description of the regression results.

 

  • Page 11, Lines 323–328: The discussion begins with strong claims regarding construct validity and usability despite the preliminary nature of the data and limited methodological transparency.

R36: We revised the opening of the Discussion to use more cautious language. The findings are now described as preliminary and exploratory, and we avoid stating that the app demonstrated construct validity. Instead, we refer to preliminary evidence consistent with exploratory construct validity, while acknowledging the need for further validation against reference physiological measures.

 

  • Page 11, Lines 330–345: The manuscript interprets weak correlations as evidence of stable emotional vulnerability and allostatic burden. These interpretations extend beyond the available evidence.

R37: We revised this section to avoid overinterpreting weak correlations. The associations are now described as small in magnitude and hypothesis-generating, and we removed speculative interpretations regarding real-world vulnerability or discomfort burden that were not directly supported by the data.

 

  • Page 11, Lines 347–355: The explanation regarding overlapping stress indices and shared variance should be supported with multicollinearity statistics such as VIF or tolerance values.

R:38 We revised the manuscript to clarify this point and avoid unsupported conclusions.

 

  • Page 12, Lines 368–388: The practical implementation discussion is highly speculative because the study did not evaluate intervention effectiveness, adherence, or long-term monitoring outcomes.

R39: We revised the text to avoid speculative claims. The manuscript now clarifies that the study did not assess intervention effectiveness, adherence, or long-term monitoring outcomes. Accordingly, practical implications are now framed cautiously as potential directions for future research rather than as conclusions supported by the present data.

 

  • Page 12, Lines 389–391: Ethical considerations are discussed only briefly despite the collection of facial physiological data through smartphones. Data privacy and algorithmic transparency deserve deeper consideration.

R:40 We thank the reviewer for raising this important ethical point. We revised the manuscript to include a more explicit statement on data privacy and algorithmic transparency. We also clarified that participants were informed about the non-diagnostic nature of the app-derived indices.

 

  • Page 12, Lines 392–398: The manuscript repeatedly claims ecological validity without conducting real-world longitudinal monitoring or ecological momentary assessment; Page 13, Lines 407–416: The limitations section is incomplete because it does not sufficiently discuss the implications of proprietary algorithms, selection bias, weak effect sizes, and lack of physiological validation.

R:41 We substantially revised the Limitations section to address these issues more explicitly. We now discuss the proprietary nature of the algorithm, the limited reproducibility and transparency of app-derived indices, the absence of concurrent validation against ECG or contact PPG in the present sample, the potential selection bias related to recruitment from a commercial panel and university students, and the small magnitude of the observed effects. We also clarified that the findings should be interpreted as preliminary and exploratory.

 

  • Page 13, Lines 417–421: The generalizability limitation is understated considering that most participants were students and young adults from a single country.

R42: We strengthened the statement on generalizability by clarifying that the sample composition, including a large proportion of university students and younger adults, limits external validity. We now state that the findings may not generalize to older adults, clinical populations, or more demographically diverse samples.

 

  • Page 13, Lines 423–425: Future directions are broad and ambitious relative to the limited evidence provided in the present study.

R43: We thank the reviewer for this suggestion. We revised the Future Directions section to make it more cautious and proportionate to the exploratory nature of the present study. Specifically, we now state that further longitudinal studies, repeated rPPG assessments, and validation against reference physiological measures are needed to determine the stability, reproducibility, and clinical relevance of these app-derived indices before considering their use in preventive mental health care or long-term outcome monitoring.

 

  • Page 13, Lines 427–443: The conclusions substantially overstate the public health implications of the findings, given the small effect sizes and preliminary design.

R44: We revised the text to avoid overstating the public health implications of the findings. The conclusion now emphasizes the small effect sizes, cross-sectional design, and preliminary nature of the evidence. We also clarified that broader public health or clinical implications require confirmation through longitudinal studies, repeated rPPG assessments, and validation against reference physiological measures.

 

  • Pages 15–18, References: Several references are very recent narrative reviews or technology-oriented discussions, while fewer foundational psychophysiological validation studies are included. The literature balance could be improved.

R45: we include additional references on the limitation of rPPG

Author Response File: Author Response.pdf

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

Thank you for addressing all of my original concerns - the researchers have done extensive revisions and have made a tremendous effort to improve the quality of their message and the manuscript. They are commended for this. I have no further comments.

Author Response

Comment 1. Thank you for addressing all of my original concerns - the researchers have done extensive revisions and have made a tremendous effort to improve the quality of their message and the manuscript. They are commended for this. I have no further comments.

Response 1. Thank you for your positive feedback. We are grateful for your time and thoughtful review.

Reviewer 2 Report

Comments and Suggestions for Authors

I appreciate the substantial improvements made in methodology, reporting clarity, and interpretation. But several important methodological and presentation concerns still remain unresolved, and the manuscript requires further revision before it can be considered for publication.

  • The manuscript still contains overstated implications regarding early detection and preventive mental health applications. Although the authors revised some wording, the Highlights, Introduction, Discussion, and Conclusions still imply that the app may enable early detection, proactive prevention, and scalable preventive mental health care. These claims remain stronger than warranted for a cross-sectional exploratory study without longitudinal validation, intervention testing, or prospective outcome assessment.
  • The manuscript still overuses the concept of “objective” stress monitoring despite the absence of criterion validation against gold standard physiological measures such as ECG-derived HRV, cortisol, electrodermal activity, or inflammatory biomarkers. The current study demonstrates only preliminary associations between app-derived indices and self-reported symptoms. The wording throughout the manuscript should remain more cautious and avoid implying validated objective stress assessment.
  • Important technical details related to signal acquisition and quality control remain insufficiently reported. The manuscript still does not specify the exact recording duration, minimum signal quality thresholds, quality rejection criteria, or the number of excluded recordings due to poor signal quality. These details are important for reproducibility and evaluation of real-world feasibility in rPPG research.
  • The physiological distinction between Stress Level, Stress Recovery, and Stress Response remains insufficiently explained. Although additional methodological details were added, the manuscript still does not clearly clarify how these indices differ physiologically or computationally. More explicit conceptual definitions would improve interpretability.
  • The extensive reliance on proprietary algorithms continues to substantially limit reproducibility and independent scientific evaluation. While the authors expanded the methodological description, the exact computational framework, feature weighting, signal quality thresholds, and transformation procedures remain undisclosed. This limitation should be emphasized even more clearly.
  • The issue regarding principal component analysis on only two variables remains incompletely resolved. The manuscript still presents PCA terminology, eigenvalue criteria, and rotation procedures despite using only anxiety and depressive symptom scores. The authors should instead describe this variable as a composite emotional distress score rather than a latent PCA-derived construct.
  • The manuscript appropriately acknowledges the absence of correction for multiple testing. However, given the number of exploratory correlations, subgroup analyses, and regression analyses conducted, the risk of type I error remains substantial. The exploratory nature of these findings should therefore be emphasized even more strongly.
  • Several interpretations in the Discussion remain stronger than justified by the observed effect sizes. Correlations between Stress Level and psychological symptoms are weak, ranging approximately from r = .13 to .17, while the regression model explains only 4% of the variance. Statements referring to emotional vulnerability, stress system dysregulation, cumulative burden, and psychophysiological stress pathology should therefore be interpreted more cautiously.
  • The psychobiological discussion regarding allostatic load, HPA axis dysregulation, and cumulative stress mechanisms remains somewhat speculative because these constructs were not directly measured in the study. The manuscript should more clearly separate the theoretical background from empirically demonstrated findings.
  • Recruitment source was significantly associated with stress level, yet recruitment source was not incorporated into the regression analyses as a covariate. Because the sample included both university students and panel participants with different stress distributions, potential confounding effects should be more explicitly addressed.
  • The manuscript reports that regression diagnostics and multicollinearity checks were performed, but the actual diagnostic statistics are not presented. Reporting VIF and tolerance values would improve methodological transparency.
  • The exclusion criteria relied entirely on self-declared psychiatric disorders, cardiovascular conditions, and medication use. This introduces potential misclassification bias that should be acknowledged more explicitly as a limitation.
  • The usability questionnaire remains methodologically limited because it was developed ad hoc and demonstrated only borderline internal consistency. The usability findings should therefore be interpreted strictly as preliminary exploratory observations rather than validated usability evidence.
  • The manuscript still does not adequately discuss possible skin tone-related variability and demographic bias in optical signal extraction. This is an important issue in rPPG-based systems and should be acknowledged explicitly.
  • The manuscript now includes a stronger limitations section, which is appreciated. However, the Conclusions section still remains somewhat optimistic relative to the actual findings. Terms such as “promising tool,” “population level prevention,” and “paradigm shift” should be further moderated given the exploratory cross-sectional design, weak correlations, proprietary algorithms, and lack of longitudinal validation.
  • The manuscript contains numerous unresolved editing and formatting inconsistencies that require careful revision before publication. Examples include overlapping tracked text, duplicated wording, merged phrases such as “burnout riskemotional distress,” “validateassess,” “Burnout-relatedGeneral Emotional Distress,” duplicated section transitions, and residual editing artifacts throughout the manuscript.
  • Figure 1 still appears visually crowded and would benefit from improved quality formatting and cleaner labeling.
  • Figure 2 remains somewhat confusing because terminology such as “RPD score” is insufficiently defined within the figure itself.
  • The manuscript still contains occasional inconsistent use of burnout-related terminology despite attempts to replace it with “General Emotional Distress.” A final terminology consistency review is recommended throughout the manuscript.
  • Several sections of the Discussion remain repetitive, particularly regarding psychobiological explanations of chronic stress and allostatic load. Streamlining these sections would improve readability and focus.
Comments on the Quality of English Language

The manuscript would benefit from substantial English language editing to improve grammar, sentence structure, and overall readability.

Author Response

Response to Reviewer 2

We thank the reviewer for the careful second-round evaluation and for the constructive comments. We have revised the manuscript accordingly. All textual changes introduced in this revision are highlighted in bold in the revised manuscript.

 

Comment 1. The manuscript still contains overstated implications regarding early detection and preventive mental health applications. Although the authors revised some wording, the Highlights, Introduction, Discussion, and Conclusions still imply that the app may enable early detection, proactive prevention, and scalable preventive mental health care. These claims remain stronger than warranted for a cross-sectional exploratory study without longitudinal validation, intervention testing, or prospective outcome assessment.

Response 2. Thank you for your suggestion. We substantially moderated the Highlights, Abstract Conclusions, Introduction, Discussion, and Conclusions. We now state that the findings are cross-sectional, preliminary, and exploratory, and that early detection, preventive implementation, and prediction of future outcomes require longitudinal and interventional validation.

 

Comment 2. The manuscript still overuses the concept of “objective” stress monitoring despite the absence of criterion validation against gold standard physiological measures such as ECG-derived HRV, cortisol, electrodermal activity, or inflammatory biomarkers. The current study demonstrates only preliminary associations between app-derived indices and self-reported symptoms. The wording throughout the manuscript should remain more cautious and avoid implying validated objective stress assessment.

Response 2. We revised the terminology throughout the manuscript. “Objective monitoring” was replaced or qualified as “app-derived physiological indicators” or “complementary physiological information.” We explicitly state that the present study does not provide criterion validation against ECG-derived HRV, cortisol, electrodermal activity, or inflammatory markers.

 

Comment 3. Important technical details related to signal acquisition and quality control remain insufficiently reported. The manuscript still does not specify the exact recording duration, minimum signal quality thresholds, quality rejection criteria, or the number of excluded recordings due to poor signal quality. These details are important for reproducibility and evaluation of real-world feasibility in rPPG research.

Response 3. We thank the reviewer for this helpful comment. We agree that further technical details regarding signal acquisition and quality control were needed. Accordingly, we have expanded the Methods section by adding a more detailed description of the rPPG acquisition protocol, including the use of the smartphone front-facing camera, standardized in-app instructions, participant positioning, environmental conditions, lighting requirements, and device-related limitations. We also added details on the preprocessing and quality-control procedures, including face detection, ROI selection, ROI tracking and stabilization, RGB signal extraction, normalization, temporal filtering, and artifact-reduction procedures. Finally, we clarified that device heterogeneity and non-quantified illumination represent methodological limitations. These revisions have been added to Sections 2.2.2 Digital Assessment and Data Acquisition and 2.2.3 Signal Processing and Stress Index Computation.

 

Comment 4. The physiological distinction between Stress Level, Stress Recovery, and Stress Response remains insufficiently explained. Although additional methodological details were added, the manuscript still does not clearly clarify how these indices differ physiologically or computationally. More explicit conceptual definitions would improve interpretability.

Response 4. We revised Section 2.2.4 and retitled it “Conceptual Definition of the App-Derived Stress Indices.” We added explicit conceptual definitions for Stress Level, Stress Recovery, and Stress Response, while clarifying that these definitions are conceptual and not independently validated computational decompositions.

 

Comment 5. The extensive reliance on proprietary algorithms continues to substantially limit reproducibility and independent scientific evaluation. While the authors expanded the methodological description, the exact computational framework, feature weighting, signal quality thresholds, and transformation procedures remain undisclosed. This limitation should be emphasized even more clearly.

Response 5. We strengthened this limitation in both the Methods and Limitations. The revised manuscript now states that feature weighting, quality thresholds, mathematical transformations, and rejection rules cannot be independently verified, and that the indices should be interpreted as exploratory app-generated indicators rather than transparent physiological biomarkers.

 

Comment 6. The issue regarding principal component analysis on only two variables remains incompletely resolved. The manuscript still presents PCA terminology, eigenvalue criteria, and rotation procedures despite using only anxiety and depressive symptom scores. The authors should instead describe this variable as a composite emotional distress score rather than a latent PCA-derived construct.

Response 6. We revised the text to describe the outcome as a standardized composite General Emotional Distress score based on anxiety and depressive symptoms. We now explicitly state that it is not a distinct latent construct and is not a direct burnout measure.

 

Comment 7. The manuscript appropriately acknowledges the absence of correction for multiple testing. However, given the number of exploratory correlations, subgroup analyses, and regression analyses conducted, the risk of type I error remains substantial. The exploratory nature of these findings should therefore be emphasized even more strongly.

Response 7. We strengthened the statistical analysis section and Results/Discussion language to state that no correction for multiple comparisons was applied and that significant findings should be interpreted as exploratory, hypothesis-generating, and in need of replication.

 

Comment 8. Several interpretations in the Discussion remain stronger than justified by the observed effect sizes. Correlations between Stress Level and psychological symptoms are weak, ranging approximately from r = .13 to .17, while the regression model explains only 4% of the variance. Statements referring to emotional vulnerability, stress system dysregulation, cumulative burden, and psychophysiological stress pathology should therefore be interpreted more cautiously.

Response  8. We revised the Discussion to emphasize that correlations were weak and that the regression model explained only 4% of the variance. We softened statements concerning emotional vulnerability and clinical relevance accordingly.

 

Comment 9. The psychobiological discussion regarding allostatic load, HPA axis dysregulation, and cumulative stress mechanisms remains somewhat speculative because these constructs were not directly measured in the study. The manuscript should more clearly separate the theoretical background from empirically demonstrated findings.

Response 9. We revised this section to clearly separate theory from empirical findings. The manuscript now states that HPA-axis dysregulation, allostatic load, inflammation, and cumulative psychophysiological burden were not directly measured and should be regarded only as background explanatory frameworks.

 

Comment 10. Recruitment source was significantly associated with stress level, yet recruitment source was not incorporated into the regression analyses as a covariate. Because the sample included both university students and panel participants with different stress distributions, potential confounding effects should be more explicitly addressed.

Response 10. We now explicitly acknowledge this issue in the Results and Limitations. The revised manuscript states that recruitment source was not included as a covariate in the exploratory model and that residual confounding cannot be excluded.

 

Comment 11. The manuscript reports that regression diagnostics and multicollinearity checks were performed, but the actual diagnostic statistics are not presented. Reporting VIF and tolerance values would improve methodological transparency.

Response 11. We revised the Statistical Analysis and Results sections to explicitly report multicollinearity diagnostics. Specifically, we added the criteria used for VIF and tolerance values and clarified that all VIF values were below the conventional cut-off of 5 and all tolerance values were above 0.20, indicating no problematic multicollinearity among the predictors.

 

Comment 12. The exclusion criteria relied entirely on self-declared psychiatric disorders, cardiovascular conditions, and medication use. This introduces potential misclassification bias that should be acknowledged more explicitly as a limitation.

Response 12. We added this as an explicit limitation in the Participants section and Limitations, noting that self-report-based exclusions may introduce misclassification bias.

 

Comment 13. The usability questionnaire remains methodologically limited because it was developed ad hoc and demonstrated only borderline internal consistency. The usability findings should therefore be interpreted strictly as preliminary exploratory observations rather than validated usability evidence.

Response 13. We revised the usability section to state that the questionnaire was developed ad hoc, was not a standardized usability instrument, and showed borderline internal consistency. We now interpret usability findings strictly as preliminary user-experience observations.

 

Comment 14. The manuscript still does not adequately discuss possible skin tone-related variability and demographic bias in optical signal extraction. This is an important issue in rPPG-based systems and should be acknowledged explicitly.

Response 14. We thank the reviewer for your comment. We added this issue to the acquisition procedure and limitations. The manuscript now states that skin tone and related optical-demographic characteristics were not systematically collected or stratified, preventing evaluation of differential rPPG signal quality or demographic bias.

 

Comment 15. The manuscript now includes a stronger limitations section, which is appreciated. However, the Conclusions section still remains somewhat optimistic relative to the actual findings. Terms such as “promising tool,” “population level prevention,” and “paradigm shift” should be further moderated given the exploratory cross-sectional design, weak correlations, proprietary algorithms, and lack of longitudinal validation.

Response 15. We rewrote the Conclusions to moderate claims.

 

Comment 16. The manuscript contains numerous unresolved editing and formatting inconsistencies that require careful revision before publication. Examples include overlapping tracked text, duplicated wording, merged phrases such as “burnout riskemotional distress,” “validateassess,” “Burnout-relatedGeneral Emotional Distress,” duplicated section transitions, and residual editing artifacts throughout the manuscript.

Response 16. We performed a terminology and editing pass, including removal or moderation of inconsistent burnout-related phrasing and replacement of PCA terminology. We also corrected Figure 2 terminology in the caption.

 

Comment 17. The manuscript contains numerous unresolved editing and formatting inconsistencies that require careful revision before publication. Examples include overlapping tracked text, duplicated wording, merged phrases such as “burnout riskemotional distress,” “validateassess,” “Burnout-relatedGeneral Emotional Distress,” duplicated section transitions, and residual editing artifacts throughout the manuscript.

Response 17. We revised grammar, sentence structure, terminology consistency, and readability throughout the edited sections.

 

Comment 18. Figure 1 still appears visually crowded and would benefit from improved quality formatting and cleaner labeling.

Response 18. We revised the figure to improve readability and reduce visual crowding.

 

Comment 19. Figure 2 remains somewhat confusing because terminology such as “RPD score” is insufficiently defined within the figure itself.

Response 19. We thank the Reviewer for this helpful comment. We agree that the term “RPD” in the previous version of Figure 2 could be ambiguous. The term “RPD” was not intended to indicate a separate diagnostic scale or an additional clinical measure, but was only used as a label for the anxiety-related score. To avoid confusion, we removed “RPD” from Figure 2 and replaced it with “GAD-7 score”, consistent with the Methods section. Depressive symptoms are now consistently labelled as “BDI-II score”. This revision improves clarity and ensures consistency throughout the manuscript.

 

Comment 20. The manuscript still contains occasional inconsistent use of burnout-related terminology despite attempts to replace it with “General Emotional Distress.” A final terminology consistency review is recommended throughout the manuscript.

Response 20. We reviewed and moderated burnout-related terminology. The revised manuscript consistently frames the outcome as General Emotional Distress or a proxy indicator of burnout-related psychological vulnerability, not as a direct burnout measure.

 

Comment 21. Several sections of the Discussion remain repetitive, particularly regarding psychobiological explanations of chronic stress and allostatic load. Streamlining these sections would improve readability and focus.

Response  21. We streamlined the Discussion by reducing repeated claims about chronic stress, allostatic load, and preventive applications, and by focusing more directly on the empirical findings and their limitations.

 

Comment 22. The manuscript would benefit from substantial English language editing to improve grammar, sentence structure, and overall readability.

Response 22. We have revised the manuscript to improve the English language, grammar, sentence structure, and overall readability throughout.

Round 3

Reviewer 2 Report

Comments and Suggestions for Authors

The manuscript has improved substantially following revision, and many of the previously raised concerns have been adequately addressed. But a few important issues remain that should be resolved before the manuscript is suitable for publication.

In Sections 2.2.2 and 2.2.3, the description of the rPPG acquisition and processing procedures has been expanded, but critical methodological details remain missing. The manuscript does not clearly report the exact duration of the facial recording, the signal quality thresholds used to determine whether a recording was acceptable, the predefined exclusion criteria for poor quality scans, the number and proportion of recordings excluded due to inadequate signal quality, or whether repeated recordings were permitted when a scan failed. These details are essential for reproducibility and for evaluating the strength of the methodology. Although the proprietary nature of the algorithm may limit what can be disclosed about feature weighting and model implementation, the authors should still report the recording duration, quality-control procedures, exclusion criteria, and participant flow.

The issue of the recruitment source confounding remains only partially addressed. In Sections 3.4 and 3.7, the authors report significant differences between the student and panel subsamples, demonstrating that recruitment source is associated with stress level classification. However, the recruitment source was not incorporated into the regression analyses despite evidence that it may act as a confounding variable. To strengthen the validity of the findings, the authors should either include the recruitment source as a covariate in the regression model or provide a sensitivity analysis demonstrating that the reported associations remain stable after controlling for the recruitment source.

The manuscript states that regression assumptions were examined, yet no quantitative multicollinearity diagnostics are reported. Given that Stress Level, Stress Recovery, and Stress Response are conceptually related constructs and appear to be correlated, it is important to document the extent of multicollinearity. The authors should report Variance Inflation Factor values and tolerance statistics for all predictors included in the regression analyses.

Although the manuscript has undergone substantial revision, numerous editorial and formatting issues remain visible throughout the document. Several sections contain residual editing artifacts, traces of text replacement, duplicated passages, merged sentences containing both previous and revised wording, and structural inconsistencies.

The Conclusions section also contains formatting irregularities and evidence of duplicated content. These issues reduce readability and make it difficult to determine the authors’ final intended text. A thorough editorial review is necessary before publication.

Comments on the Quality of English Language

The manuscript would benefit from substantial English language editing to improve grammar, sentence structure, and overall readability.

Author Response

Reviewer 2

The manuscript has improved substantially following revision, and many of the previously raised concerns have been adequately addressed. But a few important issues remain that should be resolved before the manuscript is suitable for publication.

 

Comment 1. In Sections 2.2.2 and 2.2.3, the description of the rPPG acquisition and processing procedures has been expanded, but critical methodological details remain missing. The manuscript does not clearly report the exact duration of the facial recording, the signal quality thresholds used to determine whether a recording was acceptable, the predefined exclusion criteria for poor quality scans, the number and proportion of recordings excluded due to inadequate signal quality, or whether repeated recordings were permitted when a scan failed. These details are essential for reproducibility and for evaluating the strength of the methodology. Although the proprietary nature of the algorithm may limit what can be disclosed about feature weighting and model implementation, the authors should still report the recording duration, quality-control procedures, exclusion criteria, and participant flow.

Response 1. Thank you for this comment. We agree that these methodological details are important for reproducibility and for evaluating the robustness of the rPPG procedure. We have therefore revised Sections 2.2.2 and 2.2.3 to provide additional information regarding the duration of the facial recording, the quality-control procedures applied during acquisition and processing, the criteria used to determine recording acceptability, and the possibility of repeating recordings in the event of an unsuccessful scan. We have also clarified the participant flow by reporting the number and proportion of recordings excluded because of inadequate signal quality. Specifically, in the present sample, all included participants provided one valid recording, and no recordings were excluded due to inadequate rPPG signal quality. While the numerical thresholds used by the proprietary software for signal-quality scoring, motion rejection, ROI stability, and algorithmic confidence cannot be disclosed, we have added all available methodological information regarding recording duration, scan quality assessment, exclusion criteria, and handling of failed recordings. In addition, we have acknowledged in the Limitations section that the proprietary nature of the algorithm limits full reproducibility and prevents independent verification of some technical aspects of the rPPG processing pipeline and quality-control decision rules.

 

Comment 2. The issue of the recruitment source confounding remains only partially addressed. In Sections 3.4 and 3.7, the authors report significant differences between the student and panel subsamples, demonstrating that recruitment source is associated with stress level classification. However, the recruitment source was not incorporated into the regression analyses despite evidence that it may act as a confounding variable. To strengthen the validity of the findings, the authors should either include the recruitment source as a covariate in the regression model or provide a sensitivity analysis demonstrating that the reported associations remain stable after controlling for the recruitment source.

Response 2. We thank the reviewer for this annotation. We repeated the regression analysis by including recruitment source as an additional covariate. The results remained unchanged: Stress Level continued to be the only significant independent predictor of Burnout-related Emotional Distress, while Stress Recovery and Stress Response remained non-significant. Thus, controlling for recruitment source did not alter the direction or significance of the main associations, supporting the robustness of the original findings. We have now included this analysis in the revised manuscript (Section 3.7)

 

Comment 3. The manuscript states that regression assumptions were examined, yet no quantitative multicollinearity diagnostics are reported. Given that Stress Level, Stress Recovery, and Stress Response are conceptually related constructs and appear to be correlated, it is important to document the extent of multicollinearity. The authors should report Variance Inflation Factor values and tolerance statistics for all predictors included in the regression analyses.

Response 3. We have added Variance Inflation Factor and tolerance statistics for all predictors included in the regression model. The results indicated no evidence of problematic multicollinearity, with VIF values ranging from 2.03 to 2.17 and tolerance values ranging from 0.461 to 0.494. These values are within acceptable limits, suggesting that multicollinearity was unlikely to bias the regression estimates. This has been added in the manuscript.

 

Comment 4. Although the manuscript has undergone substantial revision, numerous editorial and formatting issues remain visible throughout the document. Several sections contain residual editing artifacts, traces of text replacement, duplicated passages, merged sentences containing both previous and revised wording, and structural inconsistencies.

Response 4. Thank you for your comment. We have carefully revised the manuscript to address the editorial and formatting issues mentioned. The text highlighted in red represents the tracked changes made during the revision process and has been intentionally left visible, as required by the journal’s instructions for the revised manuscript. We have also checked the document to remove any unintended residual editing artifacts, duplicated passages, merged sentences, and structural inconsistencies.

 

Comment 5. The Conclusions section also contains formatting irregularities and evidence of duplicated content. These issues reduce readability and make it difficult to determine the authors’ final intended text. A thorough editorial review is necessary before publication.

Respoense 5. We thank the Reviewer for this comment. We apologize for the formatting irregularities and duplicated content in the Conclusions section. The section has now been thoroughly revised: duplicated text has been removed, formatting has been corrected, and the final conclusions have been clarified to improve readability and ensure that the authors’ intended message is unambiguous.

Back to TopTop