1. Introduction
The objective assessment of task complexity and human functional state during task performance remains a major challenge in neuroergonomics and biomedical engineering, as modern work environments increasingly require continuous monitoring of physiological state to prevent overload, optimize human performance, and reduce the risk of human-factor-related errors [
1]. In recent years, growing attention has been devoted to the identification of physiological measures capable of assessing human functional state directly during task performance [
2,
3,
4]. Such approaches are particularly relevant for operator performance monitoring, personnel training, adaptive human–machine systems, and digital learning environments, where timely detection of increasing task demands may improve both safety and efficiency. Recent advances in physiological sensing technologies enable continuous physiological monitoring of human state and create new opportunities for the development of intelligent physiological computing systems capable of adapting to the current condition of the user [
5]. Consequently, there has been growing interest in the automatic classification of task complexity based on physiological measurements.
Recent advances in neuroergonomics have accelerated the integration of machine learning with multimodal physiological monitoring for continuous assessment of human functional state. Current research is increasingly focused not only on estimating workload but also on developing adaptive systems capable of responding to ongoing physiological changes in real time [
6]. Nevertheless, recent reviews emphasize substantial methodological heterogeneity in feature extraction, model evaluation, and reporting practices, which complicates direct comparison between studies and limits the reproducibility and generalizability of published models. Traditionally, workload and task complexity have been assessed using subjective rating scales, such as the NASA Task Load Index (NASA-TLX), performance-based metrics, and expert evaluations [
7]. Although these approaches provide valuable information, they have several important limitations. Subjective rating scales are typically retrospective and may interfere with ongoing task performance. Furthermore, they are inherently influenced by individual perception and interpretation. Performance-based measures often become available only after task completion and provide limited information about the physiological resources expended during task execution. For applications requiring continuous assessment, such as operator training and performance monitoring, physiological markers are particularly valuable because they can be recorded continuously and non-invasively during task execution [
5,
6,
7,
8].
Among peripheral physiological signals, heart rate variability (HRV), electrodermal activity, photoplethysmography (PPG), and respiratory parameters are the most used measures for assessing workload and functional state. HRV reflects autonomic regulation of the cardiovascular system and provides time-domain, frequency-domain, and nonlinear metrics associated with physiological flexibility, sympathovagal balance, and cardiac vagal control [
9,
10,
11]. Electrodermal activity is highly sensitive to sympathetic nervous system activation and is widely used as an indicator of mental effort, arousal, and task-related activation [
12,
13]. Photoplethysmographic and respiratory measures provide complementary information regarding autonomic and vascular responses to changing workload conditions [
14]. More recent studies have demonstrated that PPG-derived cardiovascular features are sensitive to changes in cognitive workload and provide informative biomarkers for physiological workload assessment during mental tasks [
15]. Systematic reviews consistently demonstrate that no single physiological parameter is sufficient for accurately characterizing workload-related changes, whereas multimodal combinations of physiological signals generally provide better discrimination of human functional state [
8,
16,
17]. This observation is particularly relevant for task complexity classification because increasing task demands rarely manifest as isolated changes in a single physiological variable. Instead, they are typically accompanied by coordinated reconfiguration across cardiovascular, electrodermal, respiratory, and peripheral vascular systems. To identify such complex physiological patterns, machine learning approaches have become increasingly important for the assessment of workload, stress, fatigue, and cognitive state [
4]. Early neuroergonomics studies demonstrated that psychophysiological measures could successfully classify operator functional state and estimate workload in real time [
18]. Subsequent research expanded this concept by integrating multimodal neurophysiological and physiological signals into supervised machine learning frameworks, including Support Vector Machines, Random Forests, artificial neural networks, and other classification algorithms [
6,
19,
20]. Machine learning methods are particularly attractive because they capture nonlinear relationships between physiological variables and task demands that cannot easily be identified using conventional statistical approaches. At the same time, biomedical applications require not only high predictive performance but also physiological interpretability of model outputs. Consequently, classical machine learning algorithms such as Random Forests and Support Vector Machines remain highly relevant, particularly for datasets of moderate size where feature-importance analysis and physiological interpretation are essential [
21].
Recent studies further demonstrate that multimodal physiological monitoring combined with machine learning enables accurate estimation of cognitive workload across a variety of experimental conditions [
22]. In particular, Pontiggia et al. (2026) [
23] emphasized the importance of developing robust multimodal classifiers capable of maintaining reliable performance across different physiological conditions and recording scenarios. Nevertheless, the primary emphasis of these investigations has been on maximizing classification performance, whereas comparatively less attention has been paid to identifying generalizable physiological mechanisms underlying adaptation to increasing task complexity.
Despite substantial progress in physiological workload classification, several important limitations remain. First, most recent studies have evaluated physiological responses within a single experimental paradigm, including driving simulation, aviation, vigilance, office work, or laboratory cognitive tasks, without examining whether the identified physiological signatures generalize across fundamentally different forms of human activity [
22,
24]. Furthermore, systematic reviews emphasize considerable heterogeneity in experimental protocols, physiological features, and machine learning methodologies, making direct comparison between studies difficult and limiting the identification of common physiological markers of task complexity [
6,
8].
Second, a large proportion of existing ML models rely on static physiological features calculated over predefined task intervals, whereas transient physiological responses occurring during transitions between consecutive task stages remain largely unexplored. Such transition periods are of particular interest because physiological regulation is inherently dynamic, reflecting continuous autonomic adaptation rather than stable physiological states. Third, the literature often considers workload as a unitary construct, although increasing task complexity may involve not only cognitive demands but also motor coordination, action switching, executive control, and sensorimotor integration. Finally, although recent ML models frequently achieve high classification accuracy, many rely on complex feature representations or black-box approaches that provide limited insight into the physiological mechanisms underlying model decisions. Consequently, there remains a need for physiologically interpretable models capable of identifying whether similar adaptive physiological features contribute across different forms of human activity [
6,
25,
26].
One promising approach to addressing these limitations is the use of dynamic physiological features, which characterize changes in physiological regulation during transitions between consecutive task conditions rather than describing isolated physiological states. Absolute physiological measurements are strongly influenced by interindividual variability, including baseline autonomic activity, age, sex, physical fitness, emotional state, and other subject-specific factors. In contrast, relative physiological changes may directly reflect autonomic adaptation to changing task demands by reducing the influence of baseline differences between individuals [
25]. Consequently, dynamic physiological representations may characterize regulatory responses associated with predefined changes in task demand. However, they should not automatically be interpreted as context-independent markers of complexity when the preceding condition differs between classes.
From a physiological perspective, transitions between simple and complex task conditions are expected to evoke rapid autonomic reconfiguration involving coordinated cardiovascular, electrodermal, respiratory, and peripheral vascular responses. These adaptive processes may not be adequately represented by average physiological values calculated over individual task periods. From a machine-learning perspective, transition-based physiological features may provide more robust descriptors of task-related adaptation by emphasizing relative physiological responses while reducing the influence of interindividual baseline variability [
10]. Such features may therefore improve not only classification performance but also the physiological interpretability of ML models.
To our knowledge, few studies have directly compared conventional absolute physiological measurements with transition-based features calculated for repeatedly reproduced task-demand transitions across cognitive and motor-cognitive activities using a unified protocol and identical machine learning framework. Although recent studies have improved the robustness of multimodal physiological workload classification [
23], the comparative value of absolute versus transition-based physiological representations across different activity modalities remains largely unexplored. Furthermore, previous investigations have primarily addressed physiological workload within individual task domains, whereas the existence of common physiological markers of increasing task complexity across different activity modalities remains largely unexplored. Addressing this question is essential for developing generalized physiological classifiers that extend beyond a single experimental paradigm and provide physiologically interpretable assessment of human functional state.
The present study aimed to evaluate the feasibility of machine-learning-based classification of task complexity using physiological signals recorded during cognitive and motor-cognitive activities. Specifically, we compared models trained on absolute physiological values recorded during simple and complex task stages with models trained on dynamic physiological features calculated for baseline/rest-to-simple and simple-to-complex transitions. The fixed sequence was selected deliberately to obtain three comparable observations of each transition type for every participant. The study was therefore designed to assess classification within this standardized sequence rather than to establish a context-independent classifier of task complexity. Unlike previous studies, the proposed framework enables direct comparison of physiological adaptation across two distinct activity modalities while preserving model interpretability through physiologically meaningful features and Random Forest feature-importance analysis.
We hypothesized that (1) dynamic physiological features would provide better discrimination between the two standardized transition types than absolute physiological values would provide between the corresponding task stages; (2) nonlinear machine learning algorithms would outperform linear classification models owing to the nonlinear nature of physiological adaptation; and (3) similar physiological feature patterns would be observed across both cognitive and motor-cognitive tasks, suggesting partial overlap in the physiological variables contributing to classification. By addressing these questions, the present study contributes to the development of physiologically interpretable and generalizable machine learning approaches for objective assessment of task complexity across different forms of human activity.
2. Materials and Methods
2.1. Participants
The study included 69 healthy volunteers aged 18–35 years (mean age: 23.69 ± 4.9 years): 34 females and 35 males. All participants had normal or corrected-to-normal vision and reported no history of neurological, psychiatric, or cardiovascular disorders.
All participants provided written informed consent prior to participation. The study was conducted in accordance with the Declaration of Helsinki and approved by the local ethics committee.
No formal a priori power analysis was performed because the sample size was determined by the number of participants available within the predefined experimental protocol. The final sample size is comparable to those reported in previous physiological workload studies.
2.2. Experimental Design
This study included two experimental protocols designed to create varying levels of cognitive and motor-cognitive complexity during continuous task performance.
All experimental sessions were conducted in standardized laboratory conditions during daylight hours, under controlled ambient temperature and lighting. Physiological signals were continuously recorded throughout each session.
Segmentation of the task into simple and complex stages was synchronized using software markers inserted during data collection (
Figure 1).
The order of experimental stages was fixed by design. The principal purpose of this sequence was to obtain three repeated observations of two predefined increases in task demand: baseline/rest-to-simple and simple-to-complex. Randomization of individual stages would have generated additional transition types and substantially reduced the number of comparable observations per transition. Obtaining an equivalent number of observations for all possible transition combinations would also have required a longer protocol, potentially increasing fatigue and time-on-task effects.
2.2.1. Cognitive (“Reading”) Protocol
The cognitive protocol consisted of alternating blocks of simple and complex reading tasks. Each task block lasted 1 min and was separated by 1 min rest intervals.
The simple condition involved reading standard horizontally oriented text, whereas the complex condition involved reading text rotated by 180°. Text stimuli were presented on a 15-inch monitor using automated slide presentation software.
The protocol comprised repeated cycles of (1) simple reading, (2) inverted-text reading, and (3) rest. The sequence was repeated three times. Before the beginning of the task series, a baseline resting recording was obtained.
The inverted-text condition was used to increase cognitive demand by introducing additional visuospatial transformation and perceptual processing requirements while preserving linguistic content.
2.2.2. Motor-Cognitive (“Movement”) Protocol
The motor-cognitive protocol included tasks with different levels of motor and cognitive complexity.
The simple condition consisted of the sequential motor task “fist-edge-palm”, while the complex condition involved a coordinated motor-cognitive sequence (“bunny through ring”), requiring continuous switching between hand configurations and increased executive control.
Each task block lasted 1 min and was separated by 1 min rest intervals. The experimental sequence was repeated three times following an initial resting baseline recording.
Prior to each task block, participants were briefly shown a schematic instruction on the monitor. Task execution was then performed from memory without additional visual guidance. The complex task was designed to increase cognitive-motor integration demands, motor coordination complexity, and executive control requirements.
A detailed demonstration of the cognitive and motor-cognitive tasks is provided in
Video S1 (cognitive and motor-cognitive tasks).
All participants were right-handed according to the demographic questionnaire. The motor task was performed with the non-dominant (left) hand, whereas the PPG sensor and skin-conductance electrodes were attached to the dominant (right) hand, which remained relaxed and did not participate in task execution.
2.3. Physiological Signal Acquisition
Physiological recordings were performed using a Thought Technology Ltd. (Montreal, QC, Canada) system, comprising a FlexComp Infiniti recording module and BioGraph Infiniti software (version 6.9; Thought Technology Ltd., Montreal West, QC, Canada). Data were transmitted wirelessly via Bluetooth Low Energy (BLE, 2.4 GHz) to a personal computer in real time, which minimized movement-related artifacts caused by cable tension (
Figure 2).
The following physiological parameters were continuously recorded:
All physiological channels (PPG, SC and respiration) were acquired synchronously using the FlexComp Infiniti encoder at 2048 Hz.
Skin conductance was recorded using the Skin Conductance Flex/Pro sensor (SA9309M; Thought Technology Ltd., Montreal, QC, Canada) with Ag/AgCl finger electrodes attached to the distal phalanges of the second and fourth fingers. Skin conductance was measured using the exosomatic constant-voltage method (0.5–1.0 V), with a measurement range of 0–30 µS, in accordance with the manufacturer’s specifications. For each one-minute stage, the mean skin conductance level (SCL, µS) was calculated automatically and used for subsequent statistical and machine-learning analyses.
Photoplethysmography was recorded using the HR/BVP Flex/Pro photoplethysmography sensor (SA9308M; Thought Technology Ltd., Montreal, QC, Canada) attached to the thumb. The sensor provides the blood volume pulse (BVP) waveform, BVP amplitude, heart rate, and pulse-to-pulse (PP) intervals, from which heart rate variability (HRV) parameters were subsequently derived.
During the motor-cognitive protocol, both the SC and PPG sensors were attached to the hand that did not perform the movement task to minimize motion-related artifacts associated with repetitive hand movements.
Respiration was recorded using the Respiration Flex/Pro sensor (SA9311M; Thought Technology Ltd., Montreal, QC, Canada), a sensitive stretch-based girth sensor positioned around the abdomen. The sensor detects abdominal expansion and contraction and provides the respiration waveform, respiration amplitude, and respiration rate. The respiration signal represents a relative measure of abdominal expansion in accordance with the manufacturer’s specifications. For each one-minute stage, the mean respiration amplitude and mean respiration rate were calculated automatically and used for subsequent statistical and machine-learning analyses. Abdominal breathing was selected because diaphragmatic respiratory movements are closely associated with autonomic regulation and constitute a major component of slow deep breathing patterns that influence parasympathetic activity [
27].
During recording, signal quality was visually monitored in real time. To minimize artifacts, participants were instructed to remain relatively still, avoid unnecessary hand and body movements, and maintain a calm, natural breathing rhythm outside the task-specific conditions. Recording segments containing pronounced movement-related or technical artifacts (e.g., sensor contact loss, abrupt non-physiological signal jumps, or temporary signal dropout) were excluded from further analysis.
After PPG acquisition, the signal was processed to analyze heart rate variability (HRV). It has been shown that spectral characteristics of HRV derived from pulse wave (PPG) are comparable to those calculated from electrocardiography [
28,
29]; therefore, PPG can be used as an alternative data source for HRV assessment.
From the raw PPG signal, a sequence of pulse-to-pulse (PP) intervals was extracted using the built-in functions of BioGraph Infiniti software. The resulting PP interval series was exported in a format compatible with Kubios HRV software, where it underwent preprocessing, including visual inspection, artifact detection, and exclusion of physiologically implausible intervals. Conventional heart rate variability (HRV) indices were subsequently calculated in Kubios HRV from the PPG-derived interval series using its standard processing pipeline. These indices included the mean RR interval (derived from the PPG-derived PP interval series), heart rate (HR), SDNN, RMSSD, Total Power, SD1, SD2, and related parameters.
Although standard HRV analysis commonly uses longer recordings, previous studies and current HRV recommendations indicate that ultra-short (1 min) recordings can provide reliable estimates of time-domain indices such as SDNN and RMSSD, while frequency-domain measures should be interpreted with greater caution. Total Power and SD2 were therefore included as relative descriptors of autonomic dynamics rather than as clinical HRV estimates [
30].
Pulse peaks were detected using the built-in BioGraph Infiniti peak-detection algorithm. PP intervals were exported to Kubios HRV Standard (version 3.5.0) [
11], where automatic medium-level artifact correction and cubic spline interpolation were applied.
The following physiological parameters were used for model building (
Table 1).
2.4. Data Preprocessing and Feature Extraction
Physiological data preprocessing was performed using Python (version 3.12.13) in the Google Colab environment. PP interval preprocessing, including artifact correction and detrending, was performed using Kubios HRV Standard as described in
Section 2.3.
Missing values, physiologically implausible observations, and artifacts were identified before analysis. Empty cells, textual representations of missing data, and zero values that were physiologically implausible for the corresponding variables were treated as missing values (NaN). Outliers were detected using a robust criterion based on the median absolute deviation (MAD; median ± 3.5 × 1.4826 × MAD) [
31]. When MAD approached zero, the interquartile range (IQR) criterion was applied. Outliers were corrected using winsorization by replacing values outside the acceptable range with the nearest boundary value. Missing values were imputed using the median of the corresponding variable. Dynamic (Δ%) features were then calculated from the preprocessed dataset. Finally, all predictor variables were standardized using z-score normalization before machine learning model training.
Two groups of physiological features were extracted:
Dynamic features were calculated as relative changes between consecutive experimental stages associated with increased task difficulty using the following formula:
where X
previous corresponds to the value from the immediately preceding stage with lower load, and X
current corresponds to the value from the subsequent stage with higher load. Specifically, dynamic features were computed for the following transitions:
Transitions involving a decrease in difficulty (e.g., complex task → rest) were not used for dynamic feature extraction in the classification of high-load stages.
Accordingly, the class labels for the dynamic-feature models represented these two transition types. The simple-transition class corresponded to task engagement following baseline/rest, whereas the complex-transition class corresponded to an increase in demand during ongoing task performance. Thus, transition type, preceding condition, and task-demand level were not experimentally independent in the dynamic-feature dataset. Dynamic-model performance was therefore interpreted as discrimination between these two standardized transition-related physiological responses rather than as context-independent classification of task complexity.
All recordings were visually inspected prior to feature extraction, and segments with obvious technical artifacts or temporary sensor-contact loss were excluded from further analysis.
2.5. Machine Learning Classification
Machine learning analysis was performed in Python (Google Colab) using the scikit-learn and statsmodels libraries. Three supervised classification algorithms representing complementary modeling approaches were evaluated: Logistic Regression (LR), Random Forest (RF), and Support Vector Machine (SVM). Logistic Regression served as an interpretable linear baseline classifier, whereas Random Forest and Support Vector Machine were selected because of their ability to capture nonlinear relationships frequently observed in physiological data while maintaining robust performance on datasets of moderate size [
24,
32]. Deep learning approaches were intentionally not considered because the primary objective of the present study was not to maximize predictive performance but to compare physiologically interpretable feature representations using transparent and reproducible machine learning methods. Unless otherwise specified, all classifiers were implemented using the default hyperparameters provided by the scikit-learn library to ensure fair comparison between algorithms and minimize model-specific optimization bias.
Hyperparameter optimization was intentionally not performed because the objective of the study was to compare feature representations rather than maximize predictive performance. Gradient boosting methods (e.g., XGBoost and LightGBM) were not included because the objective of the study was not algorithm benchmarking but comparison of physiologically interpretable feature representations using widely established machine-learning approaches.
Classification models were developed independently for (1) the cognitive protocol, (2) the motor-cognitive protocol, and (3) the combined dataset comprising both activity modalities. For each dataset, two alternative feature representations were evaluated. The first consisted of absolute physiological values (Abs) recorded at each experimental stage.
The second representation comprised dynamic physiological features (Δ%), calculated as the relative percentage change in each physiological parameter with respect to the immediately preceding experimental stage. In this dataset, class 0 represented the baseline/rest-to-simple transition and class 1 represented the simple-to-complex transition. Dynamic features were therefore intended to distinguish physiological responses associated with these two repeatedly reproduced transition types while reducing the influence of interindividual baseline variability.
Prior to model training, all predictor variables were standardized using z-score normalization. The dataset was randomly divided into training (80%) and testing (20%) subsets at the participant level, ensuring that physiological recordings from the same participant were never simultaneously included in both subsets and thereby preventing subject-related data leakage. A single participant-level split was used for the reported analysis; cross-validation was not performed. The same training procedure was applied to all evaluated classifiers to ensure an unbiased comparison of model performance.
Model performance was primarily evaluated using ROC-AUC because it is less sensitive to class imbalance and provides a threshold-independent estimate of discriminative performance. Classification accuracy is reported as a complementary performance measure.
To improve the physiological interpretability of the classification models, feature importance was estimated for the Random Forest classifier using the built-in Gini impurity-based importance measure. This analysis quantified the relative contribution of individual physiological variables to task-complexity classification and enabled identification of the physiological systems contributing most strongly to model performance. Particular attention was paid to physiologically interpretable cardiovascular and autonomic regulation features, including HRV, SC, PPG, and respiratory parameters, as these represented the primary physiological markers investigated in the present study.
Finally, classification performance obtained using models trained on absolute physiological values was compared with that of models trained on dynamic physiological features to compare stage-based discrimination using absolute physiological values with transition-based discrimination using dynamic physiological features within the fixed experimental sequence.
2.6. Assessment of Subjective Workload Using NASA-TLX
After completing each protocol, participants rated the subjective difficulty of the simple and complex task segments using the Raw NASA Task Load Index (NASA TLX). Ratings were obtained separately for the cognitive and motor-cognitive protocols. The scale comprised six dimensions (Mental, Physical, Temporal demand, Performance, Effort, Frustration), each rated on a 100-point scale. Performance was reverse-scored. No pairwise weighting was applied. Raw scores were summed to produce a global workload score per task segment per protocol per participant. These ratings were used to verify the effectiveness of the difficulty manipulation and to support interpretation of the physiological classification results. Subjective workload was assessed using a Russian-language adaptation of the NASA Task Load Index (NASA-TLX) developed and validated for Russian-speaking participants [
33].
2.7. Objective Task-Performance Assessment
Objective task performance was assessed retrospectively from video recordings by two independent experts with expertise in neurophysiology. The experts were not involved in conducting the experiment and were not present during data collection. Each recording was evaluated independently using a standardized scoring form and predefined assessment criteria. In cases of disagreement, the experts reviewed the recording jointly and reached a consensus.
Performance was evaluated separately for each one-minute task stage. For the Reading protocol, reading speed was calculated as the number of words read per minute. For the Movement protocol, execution speed was calculated as the number of completed movement elements per minute. A movement element was defined as one hand configuration within the predefined sequence. Thus, one complete “fist–edge–palm” sequence comprised three elements, whereas one complete “bunny-through-ring” sequence comprised thirteen elements. Element-based scoring was used because the two exercises differed substantially in sequence length and completion time, making the number of complete sequences per minute unsuitable for direct comparison.
Incorrect and incomplete responses were coded separately by the experts using predefined criteria. Because the present study focused on task-execution speed as an objective measure supporting the experimental manipulation, error-related outcomes were not included in the current analysis and will be examined separately.
4. Discussion
The present study evaluated the feasibility of ML-based classification of task complexity using physiological signals recorded during cognitive and motor-cognitive activities.
Both subjective (NASA-TLX) and objective (task-execution speed) measures confirmed the effectiveness of the experimental manipulation, supporting the interpretation that the observed physiological differences primarily reflected changes in task demands rather than random fluctuations in performance.
The main finding was that models trained on dynamic physiological features consistently outperformed those based on absolute physiological values in the reading and movement protocols and in the combined dataset.
However, the dynamic classes represented baseline/rest-to-simple and simple-to-complex transitions within a fixed experimental sequence. Their superior performance should therefore be interpreted as improved discrimination of these standardized transition-related responses rather than as evidence of context-independent task-complexity classification.
The study extends previous workload research by demonstrating that transition-based features provide consistent advantages across two distinct activity modalities. The partially similar feature patterns observed in the cognitive and motor-cognitive protocols suggest that comparable autonomic systems contribute to responses during standardized increases in task demand. Whether these patterns generalize to other transition contexts or randomized task sequences remains to be established.
One possible explanation is that absolute physiological values are highly susceptible to interindividual variability. Baseline autonomic activity varies with age, sex, fitness, lifestyle, and other factors, resulting in substantial between-subject variability in HRV and related indices [
10]. Consequently, identical task demands may elicit different absolute responses across participants. By contrast, dynamic characteristics describe relative changes during transitions between task conditions and may therefore directly reflect task-related adaptation. This interpretation is consistent with previous research emphasizing workload transitions and adaptive physiological responses [
34]. The findings support the view that autonomic adaptation should be interpreted as a dynamic process rather than as a sequence of independent states. Similar observations have been reported in real-world workload-monitoring studies, where individual calibration and change patterns were more informative than absolute physiological levels alone [
25].
Feature-importance analysis provided additional insight into the physiological variables contributing to classification. In the combined model, the highest-ranked predictors included the PPG-derived mean RR interval, Total Power, abdominal respiration amplitude, HR range, skin conductance, and SD2. Their importance was relatively evenly distributed across physiological domains, with no single variable showing overwhelming dominance.
This pattern suggests that increasing task demands are associated with a distributed autonomic response involving several physiological systems. HRV indicators, including the mean RR intervals, Total Power, HR range, SD2, RMSSD, and SDNN, were consistently represented among the most informative predictors, indicating an important contribution of cardiac autonomic regulation. Respiratory amplitude and skin conductance also contributed substantially, particularly in the combined and motor-cognitive datasets, suggesting that respiratory and sympathetic electrodermal responses provide complementary information.
The balanced distribution of feature importance across cardiovascular, electrodermal, respiratory, and peripheral vascular variables argues against a single dominant marker of task complexity. Instead, adaptation appears to emerge from coordinated interactions among multiple autonomic subsystems. This supports the development of multimodal monitoring systems that integrate complementary biosignals rather than relying on a single measure. Previous studies have similarly shown that workload-related responses arise from interactions among several physiological systems [
8,
16], and that multimodal approaches improve cognitive-load assessment by capturing different aspects of adaptation to changing task requirements [
26].
Taken together, the results indicate that task complexity is represented by a multimodal adaptation pattern involving coordinated cardiovascular, respiratory, electrodermal, and peripheral vascular adjustments.
The recurrence of similar feature patterns across the two protocols suggests partial overlap in the autonomic systems contributing to classification. However, because cross-protocol transfer was not evaluated, these findings should not be interpreted as evidence that a model trained on one activity modality would generalize to another.
The superior performance of Random Forest and Support Vector Machine compared with Logistic Regression suggests that the relationship between autonomic responses and the standardized task-stage or transition labels is inherently nonlinear. Similar findings have been reported in previous workload studies, in which nonlinear ML methods outperformed linear approaches when classifying cognitive states and workload levels [
19,
20]. Logistic Regression provided an interpretable linear baseline, whereas Random Forest and Support Vector Machine were better suited to model multidimensional interactions among autonomic variables.
The best-performing classifier differed between activity modalities. Support Vector Machine achieved the highest performance in the cognitive protocol, whereas Random Forest performed best in the motor-cognitive protocol. Although these differences should be interpreted cautiously, they suggest that optimal model performance may depend on the structure of the underlying responses. Cognitive tasks may have generated relatively compact response patterns that favored Support Vector Machine classification, whereas the greater interindividual variability and heterogeneity associated with motor-cognitive activity may have been better captured by the ensemble-based Random Forest classifier. This interpretation is consistent with evidence that different ML algorithms capture complementary aspects of physiological regulation depending on task characteristics and data structure [
18,
20].
An important practical advantage is that the reported performance was achieved using physiologically interpretable features and relatively simple ML models. The results therefore support the feasibility of monitoring systems capable of detecting responses resembling the standardized task-demand transitions examined here. Nevertheless, validation in randomized, counterbalanced, and real-world task sequences is required before the models can be used for context-independent detection of task complexity.
Many of the highest-ranked predictors can be obtained using wearable cardiovascular monitoring systems without complex neurophysiological instrumentation. This enhances the translational potential of the approach for occupational, educational, and rehabilitation settings. The prominent contribution of the PPG-derived mean RR interval and other HRV-related features is particularly encouraging because cardiovascular signals can be acquired continuously using widely available wearable sensors. Such systems may support adaptive training, operator monitoring, rehabilitation technologies, and other applications requiring objective assessment of functional state [
3].
The reported performance estimates were obtained from a single held-out participant-level test set. Future studies should employ repeated validation procedures to estimate confidence intervals and quantify model variability.
Several limitations should be acknowledged. The sample consisted of healthy young adults, which may limit generalizability to other populations. The experimental tasks were laboratory models of cognitive and motor-cognitive activity and may not reproduce the variability of real operational environments. Real-world settings involve greater variation in task context, environmental conditions, concurrent demands, and individual behavioral strategies, all of which may affect model performance. Furthermore, although classification performance was consistently above chance level, it remained moderate. Larger datasets, additional modalities, and advanced temporal models may improve performance.
A major design limitation is that task order was fixed in both protocols. Dynamic-feature class labels were therefore intrinsically associated with both transition type and preceding condition: baseline/rest-to-simple for one class and simple-to-complex for the other. Consequently, the models may have captured effects related to task engagement, preceding autonomic state, carryover, learning, fatigue, or protocol position in addition to changes in task demand.
The fixed sequence was selected deliberately to provide three repeated and directly comparable observations of each predefined transition type. Randomizing individual stages would have generated additional transition combinations and reduced the number of observations available for each transition. Obtaining an equivalent number of observations in a fully counterbalanced design would have required a substantially longer protocol, potentially increasing fatigue and time-on-task effects. Nevertheless, the present findings cannot establish context-independent classification of task complexity. Future studies should evaluate simple and complex tasks after the same preceding condition and use randomized or counterbalanced sequences designed to separate task difficulty from transition context.
Cross-protocol transferability was not evaluated. Separate models were trained and tested within each protocol, whereas the combined model included data from both protocols during training and testing. Therefore, similarities in feature importance and model performance indicate overlapping physiological contributions but do not demonstrate that a classifier trained on one activity modality can generalize to another.
Only classical ML algorithms were evaluated. These models were selected because of their interpretability and suitability for moderate-sized datasets; however, future studies should determine whether temporal deep-learning architectures provide additional benefits without compromising physiological interpretability.
Quantitative condition-specific signal-quality indices were not retained during preprocessing; therefore, artifact rates could not be retrospectively compared between simple and complex motor conditions. Future studies should incorporate automated signal-quality metrics and condition-specific artifact reporting.
Feature importance was estimated using the impurity-based measure implemented in Random Forest. Although widely used, this method may be influenced by predictor characteristics and does not provide model-independent estimates. Future studies should evaluate permutation-based or SHAP-based explanations within repeated validation frameworks.
Further research should directly assess cross-protocol transfer by training models on one activity type and testing them on another. Continuous sliding-window approaches should also be investigated for real-time detection of transitions between activity states. Finally, the robustness of transition-based features should be evaluated in more ecologically valid operational scenarios and across task domains.