1. Introduction
Self-report measures are one of the most popular if not the most popular method of measurement in the behavioral, social, and organizational sciences [
1,
2]. Their ubiquity is justified by their several advantages. First, self-report measures are relatively easy to develop compared to other methods and can be designed to assess a wide range of constructs. Their administration also requires minimal expertise and resources, and they can be administered simultaneously to large groups of participants [
3]. Finally, self-report measures are unique in their ability to capture subjective and/or covert information, which can be difficult if not impossible to access via other methods (e.g., self-perceptions; [
3,
4]).
However, self-report measures have multiple shortcomings that stem from this very inherent reliance on the participant’s subjective judgement or recall. Indeed, they may be susceptible to multiple response biases such as careless responding, which takes place when participants respond with little to no attention to the item content [
5], or socially desirable responding, which refers to participants’ tendency to misrepresent themselves generally in a favorable manner [
6]. A final limitation of self-reports is that they can be intrusive to real-time tasks if completed simultaneously [
7]. That is, asking participants to provide self-reports while performing a task is likely to cause disruption to the performance of the task.
The limitations of self-report measures are vividly highlighted in the domain of stress research, where unobtrusive measures have been proposed as alternatives that circumvent the noted limitations. Indeed, obtaining unobtrusive measures of stress is a research direction that has seen a burgeoning increase in popularity in the fields of psychology, medicine, and engineering. Due to the increasing physical, cognitive, and emotional demands of the modern world, stress is becoming a global epidemic with a 29% prevalence worldwide [
8]. Stress is generally defined as an individual’s physiological and psychological response to a challenge or threat from the environment that exceeds their coping capacity [
9]. Based on its duration, it can be categorized as chronic or acute. Chronic stress arises from prolonged exposure and has been linked to negative health outcomes. In contrast, acute stress results from short-term events (e.g., cognitive demands, interpersonal conflicts) but can potentially lead to chronic stress if the physiological responses remain unresolved over time [
10]. Acute stress is typically assessed via self-reports, which are limited by subjectivity, careless responding, socially desirable responding, and low temporal resolution. Being both objective and passive, unobtrusive measures of stress have the potential to address some of these limitations, while also enabling continuous stress assessment and subsequently, detection. This, in turn, potentially enables the design of timely interventions to reduce the negative effects of stress by administering them at opportune moments [
11]. Common types of unobtrusive measures that have been used for detecting stress include computer interaction measures which capture the interaction between the user and the computer via a mouse and keyboard clicks, facial measures which capture facial expressions via a camera, and physiological measures which capture the body’s internal responses to stressor stimuli.
Stress can affect the
autonomic nervous system (ANS), also known as the involuntary nervous system since its functions tend to be unconscious without voluntary control [
12]. For example, the sympathetic nervous system (SNS) is a component of the ANS that is responsible for the body’s “fight-or-flight” response to dangerous or stressful situations. Therefore, as it pertains to physiological measures, the activation of the SNS due to a stressor event can cause increased activity of the sweat glands that is manifested in changes in electrodermal activity (EDA) [
13], pupil dilation [
14], the narrowing of blood vessels that leads to a decrease in the blood volume pulse (BVP) [
15], and an increase in heart rate [
16]. On the other hand,
the parasympathetic nervous system (PNS) is responsible for the “rest and digest” function of the body and indicates a state of recovery from the stressful stimuli manifested as a decrease in sweat activity, slower heart rate, and constricted pupils [
17]. One’s ability to effectively modulate their heart response can result in increased heart rate variability [
18,
19], which can be quantified using time- and frequency-domain features of the electrocardiogram (ECG) signal [
20].
A stressful event can also impact facial and head movement. For instance, increased stress can potentially cause more frequent and more rapid head movement, asymmetric lip deformations, and reduced frequency of mouth opening [
21,
22]. Stress may further lead to increased tension in the speech musculature and compromise one’s ability to control one’s vocal expressions, thus impacting speech output and resulting in changes in vocal measures [
23,
24].
Studies have explored features from multiple modalities for stress detection, primarily employing supervised machine learning algorithms. A comprehensive review of the features and models used for automatic stress detection can be found in [
22]. Physiological measures generally yield moderate classification performance in automatic stress detection. For example, in a study where stress was induced using the Montreal Imaging Stress Task, EDA features achieved a binary classification accuracy between 60.5% and 72.3% for classifying between the absence and presence of self-reported stress [
25]. Similarly, analyzing data from 18 participants who completed the Trier Social Stress Test resulted in an 88% binary classification accuracy using EDA features and heart rate features [
26]. Despite these positive results from simulated stress in controlled settings, detecting stress in realistic task settings and environments remains challenging due to the weaker links between daily life stressors and physiological responses compared to behavioral responses [
27]. For instance, physiological measures alone achieved a 64% accuracy in binary classification accuracy when estimating stress from naturalistic office-related work tasks, while computer interaction features produced a similar accuracy of 65% for the same task [
28]. Some researchers have also developed stress detection systems based on facial action units. For instance, a support vector machine (SVM) classifier utilizing facial action units to identify stress achieved an accuracy of 77% [
29]. More recent studies, such as [
30], have identified common facial patterns associated with stress, offering a more interpretable method for stress detection. Other studies have combined facial action unit measures with features that capture head orientation, overall facial movements, and emotion, resulting in 75% binary classification accuracies [
28].
The present study examined stress in the context of performance on a task-type known as synthetic task environments (STEs). STEs, which are extensively used in the individual and team performance literatures, are complex performance tasks that simulate and model the cognitive, information-processing, psychological, and in team contexts, the social processes and demands that are present in operational environments [
31,
32]. STEs permit the examination of complex performance processes in controlled experiments in a manner that is typically not feasible in actual operational settings. Thus, STEs, which include but are not limited to computer-based simulators and desktop trainers, can be configured to simulate a wide range of conditions and events, and also allow for the collection of detailed performance data. As reflected in [
32], computer-based simulations of this sort have clear research informative value and utility because “STEs provide a valuable compromise between the complexity of the real world, which is important and critical for establishing externally valid results, and experimental control, which is necessary to establish internally valid results.” [
33]. These advantages explain their widespread use in training and complex skill acquisition research [
32].
Given the limitations of self-report measures and the increasing popularity of unobtrusive measures, the goal of the present study is to examine the extent to which unobtrusive measures of stress can serve as valid substitutes for self-report measures of stress. Specifically, using a concurrent design, we investigate the criterion-related validity of three types of unobtrusive measures of stress (i.e., physiological, computer interaction, and facial data), that is, the extent to which they can predict two outcomes: (1) the stress condition to which participants are assigned (i.e.,
no-stress,
low-stress, and
high-stress) and (2) performance on a complex synthetic task, better than self-report state measures (i.e., state anxiety, task load, and momentary stress). The use of these two different outcomes allows for a more comprehensive comparative evaluation of the efficacy of unobtrusive measures as valid alternatives for self-report measures of stress. To address this question, we analyzed data collected in a synthetic task environment called
Crisis in the Kodiak: Oilrig Search and Rescue [
31,
34], which is a prototypical example of the types of complex synthetic task environments used in lab-based training and performance research.
This study contributes to the literature in several ways. First, the majority of studies on automated stress detection have explored different stress operationalizations individually [
22], leaving the potential of comparing and contrasting among the types of stress outcomes largely underexplored. To address this gap, the present study uses two complementary stress outcomes (i.e., the experimentally assigned stress condition and task performance) to assess the comparative criterion-related validity of self-reports and different types of unobtrusive measures (i.e., physiological, computer interaction, and facial data). The use of these two different outcomes allows for a more comprehensive comparative evaluation of the utility of the different modalities as measures of stress. Second, unlike prior studies that typically evaluate stress detection at a single time point [
25,
26], the data are collected at multiple time points (i.e., 4 trials) which allows for an examination of the temporal stability of their criterion-related validity. Third, data analyzed in this study were collected via a synthetic task environment (STE), which is a noteworthy contribution to the extant literature because most studies in this field have focused on simulated stress during tasks that do not realistically simulate real-life stress conditions (e.g., Stroop Test, Mental Arithmetic Task) [
21,
25,
26], and there are mixed results regarding the feasibility of automated stress detection from real-world data.
4. Discussion
The study sought to examine the extent to which unobtrusive measures of stress can serve as valid substitutes for self-report measures. To achieve this objective, the study used features from four different modalities (i.e., self-report, physiological, computer interaction, and facial) to predict two outcomes (the stress condition and task performance) using data from multiple time points (i.e., 4 trials).
Based on the results, a number of summary statements and discussion points can be noted. First, the analyses by modality indicated that the stress condition was predicted by the self-report measures (), whereas task performance was predicted by physiological () and mouse interaction () measures. The observed relationship between task performance and physiological and mouse interaction measures suggests that the criterion-related validity of these two modalities is not limited to brief controlled tasks (e.g., Stroop Color-Word Test), which are commonly examined in the literature but also extend to performance on more complex tasks. Pertaining to the stress condition, the finding that this outcome was predicted exclusively by self-report measures suggests that in spite of their recognized limitations as reviewed in the introduction, self-report measures can in some instances have higher criterion-related validity than other modalities. Taken together, the different results of the stress condition and task performance also indicate that the validity and utility of modalities depend on the outcome of interest. One should note that the two outcomes examined in the present study are qualitatively different, which may explain why they were predicted by different modalities. The stress condition is a variable that was manipulated to introduce meaningful variance in the level of stress experienced by participants. On the other hand, performance on complex tasks such as that used in the present study is the result of cognitive, emotional, and behavioral processes that participants engaged in, and this performance could potentially be influenced by the levels of stress they experienced.
One may also speculate about possible reasons for the absence of relationships between specific modalities and outcomes. For instance, the finding that physiological measures did not predict the stress condition is at odds with prior studies that reported an association between physiological measures and stress [
22]. This null result may, however, be due to the type of task used in the present study. Indeed, much of the prior work on stress manipulation has relied on controlled tasks such as mental arithmetic tests and the Stroop Color-Word Test, which are significantly shorter in duration and more specialized in specific skills or abilities compared to
Crisis, which is relatively longer (6 to 10 min), open-looped, and involves various psychological processes. Another surprising result worth noting is that self-report measures did not predict task performance. Prior work suggests that weak correlations between self-report and behavioral (or performance-based) measures may be due to the poor reliability of behavioral measures [
45]. However, this explanation is unlikely to be the case in the present study as
Crisis performance displayed high test-retest reliability estimates (i.e., >0.65; see
Section 2). Thus, a more plausible explanation for the absence of a relationship between self-report and task performance in this instance is likely the limited variability in
Crisis performance scores resulting from the difficulty of the task. As reflected in
Table 2, the maximum possible score is 600, whereas the
mean task performance scores across trials ranged from 4.32 to 136.97.
Second, the analyses of all modalities by trial indicated that model predictions of the outcomes varied from one trial to another. Specifically, for task performance, the best predictions were obtained in Trial 1 () and Trial 4 (). Similarly for the stress condition, among the trials where stress was manipulated (i.e., Trial 2 to 4), classification was best in Trial 4 (). These findings could be related to the unique experiences of participants in these two trials. That is, Trial 1 is different from other trials in that the design of the study and the experience of the task are still novel to the participants, and therefore, this trial requires more cognitive processing. Trial 4 is also unique in that it is the last trial of the task. Specifically, by Trial 4, participants might have had sufficient practice to get better at the task. In addition, being aware that Trial 4 is the last one, participants may also have exerted all their efforts in this trial to maximize their chances of receiving the additional performance-based compensation. More generally, the variability in the results by trial indicates that the criterion-related validity of stress features is trial-dependent with potentially different psychological underpinnings at different time points.
Third, the analyses of all modalities and trials predicted both the stress condition (), and task performance (). This suggests that features from different modalities and trials as a whole contain information relevant to stress, which is consonant with the previously discussed results of the analyses by modality and trial. However, interestingly, the predictions using all modalities and trials were not the best across all the analyses conducted. Specifically, the model with all modalities and trials () predicted the stress condition worse than both the model with only self-report measures () and the model with only Trial 1 (). Similarly, pertaining to task performance, the model with all modalities and trials () predicted this outcome worse than the model with only Trial 4 (). These results suggest that more modalities are not necessarily better and may instead introduce noise to the predictions. Practically, this also means that practitioners may not need to measure stress using all available modalities and time windows to maximize predictions of the outcomes of interest. Another possible explanation is that all modalities contribute to a high-dimensional feature space, which may be challenging for the models to learn effectively given the limited number of data samples relative to the number of features. A more targeted and parsimonious approach may be more advantageous, especially given the logistical challenges and costs of data collection.
Limitations and Suggestions for Future Research
Some limitations of the present study serve as the basis for future research. The use of a sample of college students who performed a single synthetic complex task in a protocol that lasted approximately 2.5 h raises potential generalizability concerns. Future research should seek to replicate the findings reported in other settings. To better assess the predictive power of stress modalities, future work may also consider using a stronger manipulation of stress. The same types of stress demands as those used in the present study (i.e., time constraints and distracting audio) could be manipulated with stronger contrasts. Alternatively, one can also introduce other types of demands to the task environment such as lighting [
46] or temperature [
47].
Furthermore, prior work has indicated that stress is a multivariate construct [
28]. Therefore, although the present study examined multiple modalities of stress (i.e., self-report, physiological, computer interaction, and facial), additional features from other modalities such as oculometry and speech could also be considered. And because such features can be unobtrusively collected from wearable sensors, microphones, and cameras, their inclusion alongside self-report measures would allow for a more comprehensive examination of the extent to which unobtrusive measures of stress can serve as valid substitutes for self-report measures.
Third, the finding that model predictions of the outcomes in the present study varied by trial calls for further research to identify the optimal time and the frequency of measurement that maximize predictions. Relatedly, because predictions of the outcomes also varied by modality, another pertinent research question is whether predictions are influenced by an interaction of both the time and the modality of measurement.
Finally, the present study employed relatively simple multimodal fusion approaches to establish a benchmark for evaluating the ability of multiple modalities to predict stress-related outcomes using the collected data. This design enabled differences in predictive performance to be more directly attributed to the measurement modalities rather than to variations in model complexity. Advanced multimodal fusion approaches, hybrid/fuzzy classification techniques, and probabilistic approaches to model uncertainty in the various multimodal measurements should be explored as future work to further improve predictive performance while preserving the relative strengths of the different stress measurement modalities observed in this study.