1. Introduction
Human locomotion, particularly walking, represents a cornerstone of daily human function, requiring seamless integration among musculoskeletal, neuromotor, and sensory systems; alterations in gait biomechanics are highly sensitive biomarkers for underlying physiological and neurological pathologies [
1,
2]. The accelerating global demographic shift toward aging has precipitated a surge in gait-associated morbidities, encompassing osteoarthritis, post-stroke hemiparesis, Parkinsonian gait freezing, and sarcopenic decline [
3,
4]. Such disorders profoundly impair mobility, erode autonomy, and elevate susceptibility to falls, institutionalization, and premature death [
5]. Consequently, proactive gait surveillance and early anomaly detection emerge as pivotal strategies in precision preventive medicine and neurorehabilitation [
2,
6].
Gait assessment traditionally relies on laboratory-based motion capture systems, force platforms, and instrumented treadmills [
7,
8]. While these tools provide precise biomechanical measurements such as joint angles, ground reaction forces, and step timing, their deployment is limited by high cost, technical complexity, and the inability to capture naturalistic, daily-life walking patterns [
9,
10]. Consequently, many clinically relevant gait abnormalities remain undetected until substantial functional decline occurs [
11]. Wearable inertial measurement units (IMUs) have emerged as practical alternatives, capable of capturing continuous tri-axial acceleration and angular velocity data from key body segments such as the feet, shanks, or trunk [
12,
13]. These devices allow high-frequency motion data collection in real-world environments, enabling the detection of subtle gait deviations that may not manifest in laboratory assessments [
2]. Recent studies validate the accuracy of IMUs against optical motion capture systems, demonstrating reliable measurement of stride length, cadence, and variability [
14].
Despite their advantages, IMU datasets are high-dimensional, temporally complex, and often noisy, necessitating advanced computational approaches for meaningful interpretation [
14,
15]. Classical machine learning approaches such as support vector machines, random forests, and k-nearest neighbors rely heavily on handcrafted features, limiting their capacity to model complex, nonlinear temporal dependencies across multiple sensor axes [
14]. Deep learning approaches, including convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory (LSTM) networks, and Transformer architectures, have demonstrated superior capability in learning hierarchical features from raw time-series data [
7,
16,
17]. CNNs are particularly effective for learning local temporal and spatial features, while LSTMs excel in modeling long-term temporal dependencies [
16,
17]. Transformers, utilizing self-attention mechanisms, can capture global dependencies across sequences but demand larger datasets and substantial computational resources [
18,
19].
Temporal Convolutional Networks (TCNs) offer an efficient alternative, employing dilated causal convolutions to capture long-range temporal dependencies while maintaining manageable computational cost [
20,
21]. However, standard TCNs often collapse multi-axis sensor data, losing critical spatial relationships such as inter-limb coordination and cross-joint interactions [
14,
22]. Preserving both spatial and temporal correlations is essential for accurate classification of gait conditions, particularly when subtle deviations reflect early-stage neurological or musculoskeletal impairments [
11]. Insights from human-in-the-loop exoskeleton studies further underscore the importance of detailed spatiotemporal modeling. Exoskeleton-assisted walking has demonstrated that optimizing torque profiles based on real-time feedback can improve walking economy, reduce metabolic cost, and enhance gait symmetry [
23,
24]. Notably, adaptive multi-joint assistance based on continuous monitoring of joint angles and muscle activation has enabled personalized intervention strategies that improve mobility outcomes [
25]. These findings emphasize the necessity of capturing multi-scale spatial and temporal features when designing AI-based gait assessment models.
The Temporal Convolutional Network (TCN) represents a recent advancement in deep learning for multivariate time-series classification, particularly in EEG for emotion recognition [
26,
27], and for action segmentation and detection [
20]. The TCN extends conventional TCNs by incorporating multiple temporal anchors with distinct receptive fields, preserving both local and global temporal information while maintaining spatial correlations across sensor channels [
28,
29]. The model further integrates attention mechanisms to dynamically weigh contributions from each temporal and spatial anchor, enhancing feature interpretability and robustness [
30,
31]. This architecture has demonstrated superior performance in classifying complex spatial and temporal datasets compared to conventional CNN, LSTM, and standard TCN models [
14,
20].
In addition to spatial and temporal datasets, TCN and related architectures are increasingly applied in brain–computer interface (BCI) systems for emotion recognition, cognitive workload estimation, and neural decoding [
28,
29,
32]. These applications highlight the versatility of multi-scale, space-aware temporal models in extracting meaningful features from high-dimensional, temporally structured signals [
28,
33]. For instance, deep convolutional and graph convolutional neural networks have been successfully applied to EEG datasets, capturing local–global interactions essential for reliable decoding of neural signals [
32,
34]. Transformer-based architectures have further improved performance by leveraging hierarchical attention mechanisms, enabling simultaneous modeling of spatial and temporal dependencies [
27,
29,
35,
36].
Despite its strong performance in modeling complex neural dynamics, Multi-Anchor Space-Aware Temporal Convolutional Network (MASA-TCN) has been used exclusively in EEG-based emotion recognition and has never been applied to human gait analysis [
26]. This represents a notable gap, particularly given the critical role gait plays as an indicator of health, aging, fall risk, frailty, and neurological decline [
1,
2,
3,
4,
5,
6]. Wearable IMU-based gait analysis has become central to mobility assessment, enabling the extraction of stride parameters, joint coordination patterns, and dynamic balance metrics that are clinically meaningful and responsive to disease progression [
7,
8,
9,
10,
11,
12,
13,
14,
15,
37].
Introducing MASA-TCN into this domain offers several key advantages. First, its multi-anchor temporal convolutions are well suited to capture the fine-grained temporal structure of gait cycles, including stride timing, swing–stance transitions, and stride-to-stride variability features that are strongly linked to gait stability and fall risk [
1,
2,
9]. Second, the model’s spatial-attention fusion provides a principled way to identify the most informative IMU channels and motion axes, supporting interpretable analysis of how specific body segments contribute to normal or pathological gait patterns, consistent with findings across sensor-based gait studies [
10,
11,
12,
13,
14,
15]. Third, MASA-TCN’s lightweight, TCN-driven architecture enables efficient real-time inference, aligning with advances in wearable-sensor gait tracking and activity recognition systems designed for continuous monitoring in real-world settings [
21,
38,
39,
40].
Taken together, these factors position MASA-TCN as a novel and technically well-justified architecture for gait analysis. Its proven ability to model multi-scale temporal–spatial patterns in EEG, combined with the clinical importance of gait and the growing maturity of wearable sensing, provides a strong foundation for applying MASA-TCN to gait datasets for the first time. This also paves the way for integrating MASA-TCN with XAI methods to deliver transparent, clinically interpretable gait assessments, an essential requirement for future healthcare and mobility-monitoring applications.
Wearable exoskeleton studies have further informed the design and evaluation of AI-based gait models. For instance, adaptive hip exosuits have been shown to reduce metabolic cost during walking and running, while ankle and knee exoskeletons improve stability and joint-specific assistance [
24,
25,
41]. Human-in-the-loop optimization enables dynamic adjustment of torque profiles based on user-specific biomechanics, emphasizing the importance of accurate multi-channel sensor modeling [
42]. These findings reinforce the notion that effective AI-based gait analysis must account for both local joint-level dynamics and global locomotor coordination, which MASA-TCN accomplishes through its multi-anchor, space-aware design [
26,
43]. This study’s contributions, as shown in
Figure 1, are the following:
Development of G-MASA-TCN, a unified deep learning model for multi-condition gait classification and individual identification.
Space-aware temporal layers for spatial–spectral learning across multiple IMU channels.
Multi-anchor attentive fusion for capturing dynamic gait events at multiple temporal scales.
Evaluation on 260 real patients with eight clinically confirmed conditions, demonstrating high classification and identification accuracy.
Comparative analysis with standard TCN, Gated Recurrent Unit (GRU), and Transformer neural network-based architectures.
G-MASA-TCN offers a unified approach for analyzing human gait, capturing complex spatiotemporal patterns, and enabling accurate, clinically relevant classification and identification. It advances automated gait analysis by overcoming limitations of prior deep learning methods, providing both rich spatial–temporal modeling and high recognition accuracy.
The article is organized as follows.
Section 2 describes the experimental design, participants, gait signal data, and preprocessing methods. It also presents the G-MASA-TCN with the deep neural network architecture.
Section 3 and
Section 4 present and analyze the results obtained. Finally,
Section 5 offers conclusions and suggests potential directions for future research.
3. Results
To ensure that the evaluation of all models was statistically reliable, reproducible, and free from artifacts caused by dataset shuffling, each experiment was repeated five times using different random seeds and a subject-level 5-fold cross-validation, with 70% of the data used for training, 10% for validation, and 20% for testing in each fold. This combined procedure allows assessment of model variance, robustness, and sensitivity to initialization, while maintaining rigorous, non-overlapping test partitions for unbiased evaluation. Model training was conducted on Google Colab cloud infrastructure using an NVIDIA GPU, with an average training time of approximately 5 min, while testing and inference evaluations were performed on a local workstation equipped with an Intel Core i7 (10th generation) CPU to assess deployment feasibility on standard hardware. For each experiment, two complementary sets of results are reported: (1) the best-performing run in terms of validation accuracy, including the confusion matrix, precision, recall, F1-scores, overall accuracy, and full training/validation accuracy curves; (2) aggregate statistics across all five runs, represented by the mean confusion matrix and corresponding standard deviations. To further visualize the model’s ability to cluster gait dynamics across the eight classes (HSs, HOA, KOA, ACL, PD, CVA, CIPN, and RIL), a t-SNE projection of the learned embeddings is provided for each experiment.
Because diagnostic applications require not only strong predictive performance but also transparent decision-making, these quantitative results are complemented by a comprehensive explainable AI (XAI) analysis. This interpretability framework reveals the internal spatiotemporal representations learned by each model and identifies the most influential sensor channels and gait-cycle segments contributing to classification. Together, the statistical evaluation and XAI visualizations provide a complete and trustworthy assessment of both model accuracy and interpretability in biomechanical and clinical contexts. In the following sections, the results of the training and evaluation strategies introduced, followed by an evaluation using the Integrated Gradients (IG) XAI method.
3.1. Experiment 1: Zero-Padding Strategy
This experiment investigated the effect of zero padding on modeling variability in gait duration among participants. Since real-world gait recordings vary in length due to differences in stride cycles, cadence, mobility level, and pathological conditions, all sequences were standardized to the global maximum length
by padding shorter sequences with zeros. This ensured that architectures such as the TCN, Transformer, and the GRU could process all trials in a unified tensor format while preserving complete temporal dynamics. However, the impact of padding on learning stability varied substantially across models. Across the five independent seeds (see
Table 3), the baseline TCN reached
, confirming that conventional dilated convolutions can capture long-range temporal dependencies but remain sensitive to variation in gait rhythm. The Transformer, however, performed considerably worse at
, likely due to its positional-encoding dependence and limited ability to regulate noisy or padded temporal segments. The GRU exhibited the weakest performance at
, which suggests that recurrent architectures face difficulty maintaining stable hidden-state transitions when padding disrupts stride periodicity.
In sharp contrast, the proposed G-MASA-TCN achieved 95.2% ± 0.8%, not only outperforming all baseline models but also exhibiting extremely low variance across all five seeds. This demonstrates the architecture’s stability under variable-length gait conditions and its ability to extract discriminative temporal cues from complete, natural gait cycles. The multi-anchor temporal filters of MASA-TCN allow the model to analyze gait at multiple temporal scales simultaneously, enabling it to capture subtle abnormalities such as reduced stride length in HOA and KOA, elevated gait asymmetry in ACL and CVA, and tremor-induced high-frequency fluctuations in PD. Moreover, the spatial-attention fusion mechanism highlights the most informative IMU channels and axes when distinguishing neuropathic gait alterations in CIPN and central motor disruption patterns in RIL. The results of this experiment include the confusion matrix for the best seed run, the mean confusion matrix across all runs, accuracy/precision/recall/F1 metrics, and the complete training and validation accuracy curves are summarized in
Figure 4. The data show that G-MASA-TCN manages temporal irregularities caused by pathology-specific gait disruptions, supporting its advantage over architectures with limited spatial adaptivity or weaker temporal stability.
3.2. Experiment 2: Fixed-Length Segmentation (500-Frame Windows)
This experiment evaluated the impact of fixed-length segmentation, which eliminated zero padding altogether. Gait signals were windowed into non-overlapping 500-frame segments, corresponding to approximately 5 s of motion. This segmentation reduced temporal variability between samples, thereby stabilizing the learning process by presenting the models with consistent temporal structures. Under this standardized alignment, the baseline models improved significantly across all five seeds, as shown in
Table 3. The traditional TCN reached
, the Transformer achieved
, and the GRU improved markedly to
. These gains reflect the benefit of ensuring that each input segment contained a comparable number of gait cycles, a crucial factor given the distinct stride dynamics across HSs, HOA, KOA, and PD and the irregular timing patterns in ACL, CVA, CIPN, and RIL.
The G-MASA-TCN once again delivered the strongest performance, achieving 96.8% ± 1.2% across the five seeds. This marks an improvement over Experiment 1 and indicates that MASA-TCN benefited from clean, uniformly segmented gait windows in which stride-level micro-patterns were preserved without padding artifacts. The multi-scale anchoring mechanism was especially advantageous here: it captured the interaction between low-frequency gait-cycle components (e.g., stance–swing transitions) and high-frequency pathological markers (e.g., PD tremor- or CIPN-related proprioceptive instability). The spatial-attention module also became more effective in this setting, consistently highlighting hip-mounted and shank-mounted IMU axes that strongly differentiate HOA, KOA, and ACL conditions. The results of this experiment, including detailed confusion matrices for the best run, the aggregated mean matrices, and the learning curves, are displayed in
Figure 4. The improvement across all architectures confirms that fixed-length segmentation fostered more stable temporal learning, but G-MASA-TCN maintained a decisive performance advantage, approaching near-perfect discrimination across all eight gait classes.
The models were initially trained and validated using 70% of the available dataset, with 60% allocated for training and 10% reserved for validation, while the remaining 30% was held out for testing. This initial split enabled optimization of model parameters and hyperparameters while providing an unbiased evaluation of performance on unseen data. Given that G-MASA-TCN achieved the highest predictive performance among all architectures, further experiments were conducted to evaluate its robustness across different training-to-testing ratios. Specifically, the model was retrained on 60% of the data and tested on the remaining 40%, then trained on 40% and tested on 60%, and finally trained on 30% with testing on 70% of the dataset. Across all these splits, the observed changes in model accuracy were minimal and statistically insignificant, which can be attributed to the large size of the dataset comprising 7455 samples and the use of fixed-length segmentation of 500 samples per segment, which ensured consistent input representations for training and testing. The results for each split are presented not only in terms of standard performance metrics but also visually through t-SNE clustering of the predicted eight gait classes (HSs, HOA, KOA, ACL, CVA, PD, CIPN, RIL), demonstrating clear separation and alignment with clinical groupings as shown in
Figure 5. These analyses collectively confirm the stability, robustness, and generalization capability of G-MASA-TCN across varying training data sizes.
3.3. Experiment 3: Subject-Level Identification via Gait Signatures
This experiment explored a fundamentally different problem setup: subject identification from gait signals. Unlike pathology classification, the models were tasked with identifying 260 individual subjects, requiring each architecture to learn unique, person-specific gait signatures. This scenario emphasized subtle spatiotemporal features that remained consistent across trials, despite variations in pathology or sensor noise. Using 500-frame segments, as shown in
Table 3, the GRU achieved 76% ± 3.3%, outperforming the Transformer (62% ± 4.1%) but slightly below the TCN (79% ± 3.5%). The GRU’s performance reflects its suitability for capturing individualized rhythmic patterns. However, the G-MASA-TCN again achieved the highest performance, reaching 96.4% ± 2.1% across the five seeds.
To further evaluate generalization and mimic biometric identification scenarios, an additional test was conducted in which 10 subjects were completely removed from the training set and treated as unknown “imposters” during testing. G-MASA-TCN returned zero predictions for these unseen subjects, indicating that it did not falsely classify unknown individuals as any of the enrolled subjects. This result mirrors real-world biometric systems, where unknown individuals are correctly recognized as imposters, demonstrating that G-MASA-TCN reliably captures subject-specific gait signatures without overfitting to the training population.
This result highlights architecture’s exceptional capability to learn high-resolution gait signatures, capturing discriminative temporal cues that differ across individuals even within the same pathology class. Such signatures include personalized cadence, foot–ground contact profiles, habitual asymmetries, and joint-coordination rhythms. The multi-anchor mechanism, by analyzing the gait signals at varying temporal scales, enables the model to discover nuanced personal stride features that conventional TCN, Transformer, and GRU architectures fail to isolate consistently. The spatial attention module complements this by selecting the most identity-relevant IMU axes, often focusing on lateral acceleration and angular velocity channels associated with personalized stride lateralization. The results for this experiment, including the best-run confusion matrix, mean confusion matrix, and complete evaluation metrics, are presented in
Figure 6. The t-SNE plots reveal well-separated clusters, demonstrating that G-MASA-TCN learns highly discriminative representations of individual gait signatures, confirming its promise for gait biometrics.
3.4. Experiment 4: 5-Fold Cross Validation
In this experiment, all models were trained and evaluated using 5-fold cross-validation. In each fold, two subjects from each class were held out for testing, while the remaining subjects were used for training. The test subjects were randomly selected under the constraint that no subject was included in the testing set more than once across all five folds.
The choice of using two test subjects per class was motivated by the limited number of available subjects in the HO, KOA, and ACL classes, each containing fewer than 20 subjects. Selecting two subjects per fold provided a more reliable and balanced evaluation under these constraints. The testing results, reported in
Table 3, indicate that the proposed G-MASA-TCN with fixed-length inputs achieved the best overall performance. Specifically, it attained a peak accuracy of 96.1%, with a mean accuracy of 94.3% ± 1.8% across the 5-fold cross-validation. This outcome further confirms the consistency of the results obtained in Experiments 1 and 2.
3.5. Experiment 5: Subject-Level Identification via Real-World Gait Signature Data
To evaluate the generalization capability of the proposed G-MASA-TCN model, experiments were conducted on two real-world IMU-based gait datasets (see
Section 2.2.5): irregular and uneven surface walking 30 subjects and natural everyday walking in an urban environment. For this dataset, the models was trained twice under two evaluation protocols. In the first experiment, the model was trained to classify all subjects, where each subject was treated as an independent class. This setup evaluated the model’s ability to capture subject-specific gait signatures under non-clinical, real-world walking conditions. The obtained classification performance demonstrates that the proposed architecture can effectively learn discriminative gait representations from IMU data recorded on uneven and irregular surfaces as well as during natural urban walking. The results, as reported in the confusion matrix in
Figure 7, show that the best-performing model achieved 93% accuracy in predicting the classes of each subject. Across the five random-seed repetitions, the model obtained a mean accuracy of 91% ± 2%, demonstrating consistent performance and reliability across different data splits.
In the second experiment, an imposter detection protocol was employed to further assess G-MASA-TCN model robustness. Five subjects were randomly selected and completely excluded from the training set. During testing, the model was required to identify these unseen subjects as imposters, i.e., subjects not belonging to any of the known classes learned during training and it return zero F1 score. This experiment evaluated the model’s ability to generalize beyond trained identities and to detect unfamiliar gait patterns encountered in real-world scenarios.
The primary objective of this validation was twofold: to assess the ability of G-MASA-TCN to capture human gait characteristics outside the clinical environment and to evaluate its robustness in modeling gait patterns during walking on uneven, irregular, and uncontrolled surfaces. The quantitative results for the dataset experimental protocol are summarized in
Table 3. Overall, the results confirm that the proposed G-MASA-TCN model maintains strong performance in real-world settings, highlighting its potential applicability for robust gait modeling beyond controlled clinical conditions.
3.6. XAI Evaluation
The explainability analysis based on Integrated Gradients (IG) offers a comprehensive view of how the G-MASA-TCN model interprets sensor data to classify gait patterns across eight distinct conditions. By incorporating multiple visualization techniques including sample-level IG heatmaps, global feature importance bar plots, and temporal relevance curves, the analysis reveals the spatial and temporal characteristics that guide the model’s decision-making. These interpretations provide insight into how specific sensor modalities, representing different anatomical locations and biomechanical movements, contribute to distinguishing healthy gait from pathological gait patterns. The features used in this study include a diverse set of acceleration and gyroscope measurements collected from the head (HE_Acc, HE_FreeAcc, HE_Gyr), lower back (LB_Acc, LB_FreeAcc, LB_Gyr), and bilateral foot sensors (LF_Acc, LF_FreeAcc, LF_Gyr for the left foot; RF_Acc, RF_FreeAcc, RF_Gyr for the right foot). Together, these measurements capture a rich representation of whole-body dynamics during gait.
The first layer of interpretability arises from the IG heatmaps, which illustrate how the model allocates attention to different sensor channels over time for each selected representative sample. Each heatmap contains twelve rows, each corresponding to one of the twelve sensor features, with columns spanning the full temporal sequence of the gait cycle. The color intensities encode the magnitude of the IG attributions, allowing the viewer to identify which sensor–time combinations most influence the model’s prediction. The point of maximum attribution is marked in each heatmap, serving as an indicator of the single most influential sensor reading in the sample. Through these visualizations, it becomes evident that the model’s feature utilization is not uniform across classes; rather, each pathology exhibits its own unique pattern of spatial and temporal emphasis.
For HSs and participants with CIPN, the heatmaps in
Figure 8 consistently highlight the lower-back acceleration sensor (LB_Acc). This feature captures the linear acceleration of the trunk, which is a central component of the gait process and highly sensitive to deviations in postural control or stability. In both HS and CIPN samples, LB_Acc displays a pronounced attribution concentration around the mid-stance region of the gait cycle. This similarity may reflect the fundamental importance of trunk dynamics in maintaining balance and gait rhythm, even though CIPN subjects often exhibit subtle alterations due to sensory deficits. Nevertheless, the model successfully captures class-specific signatures, as reflected in the distinct intensity distributions across time. The fact that both groups rely heavily on the LB_Acc sensor suggests that trunk kinematics contain robust diagnostic information for discriminating against normal gait patterns from those affected by neuropathic conditions.
In contrast to the HS and CIPN groups, the SVA and PD (see
Figure 9) cohorts show different attribution patterns aligned with their respective biomechanical abnormalities. For individuals with SVA, the model assigns strongest importance to LF_Gyr, the angular velocity of the left foot. This sensor primarily reflects rotational foot movements, which may be particularly relevant for identifying stability-related deviations characteristic of SVA. The IG heatmaps show that LF_Gyr accumulates most of its attribution early in the gait cycle, around frame 10. This finding suggests that the onset of stance or the initial contact phase carries discriminative cues for this population. For subjects with PD, however, the model’s focus shifts toward RF_Acc, the acceleration of the right foot. Parkinsonian gait is known for asymmetries, reduced foot clearance, and irregular acceleration patterns. The concentration of attribution around frame 200 for PD subjects reflects characteristic movement disruptions occurring during the mid-stance or transition phases, when foot stability is challenged. These class-specific patterns demonstrate that the G-MASA-TCN model is attuned to the biomechanical signatures associated with distinct pathologies.
The RIL and ACL groups (See
Figure 10) further exhibit unique attribution patterns centered around lower-limb accelerations. For RIL, the model’s attention is most strongly associated with LF_Acc, indicating that the linear acceleration of the left foot serves as a key differentiator for this condition. The temporal peak around frame 100 suggests that mid-stance alterations play a critical role in distinguishing this group. In ACL subjects, however, the dominant feature reverts back to LB_Acc. The attribution for this cohort is delayed relative to others, reaching its highest intensity around frame 380, near the end of the gait cycle. This timing suggests that late-stance and pre-swing phases when the knee undergoes significant load-bearing and extension may contain the strongest indicators of compensation or deficit resulting from ACL.
Similarly distinct patterns arise in the HOA and KOA groups (see
Figure 11). In HOA samples, the head kinematics specifically HE_FreeAcc (head acceleration minus gravity) are highlighted as the most influential feature. This sensor captures subtle head-movement adjustments associated with discomfort or instability resulting from HOA. The temporal importance peak around frame 220 suggests that the model identifies instability or compensatory adjustments during late mid-stance or early push-off as the key discriminative moment. For KOA subjects, the model again prioritizes RF_Acc, the right-foot acceleration, with its temporal peak around frame 190. KOA commonly influences loading, propulsion, and weight transfer, which explains the reliance on foot acceleration measurements during the mid-gait phases.
While the IG heatmaps provide sample-level insights, the global feature importance plots aggregate these attributions across time to identify which sensors contribute most consistently across the gait cycle. For each class, the three highest-ranking features reflect the sensors most frequently relied upon by the model. Importantly, the dominant features in the heatmaps always correspond to those with highest mean importance globally. Thus, the importance assigned to LB_Acc in CIPN and ACL, LF_Gyr in SVA, RF_Acc in PD and KOA, LF_Acc in RIL, and HE_FreeAcc in HOA is not limited to single time steps but represents a consistent pattern across the entire gait sequence. These global visualizations help reveal the biomechanical drivers characteristic of each pathology: trunk acceleration for neuropathy-related or ligament injuries, foot angular velocity for balance-related abnormalities, head acceleration for hip-related impairments, and foot acceleration for Parkinsonian and osteoarthritic gait deviations.
The third visualization modality, the temporal importance plot, provides complementary insights by summarizing how the model’s attention evolves across the gait cycle. By averaging attributions across all sensor channels in each time step, the plot identifies critical phases that the model deems most informative. The temporal patterns vary substantially across conditions. HSs and CIPN show their peaks near frame 200, SVA near frame 10, PD again near frame 200, RIL around frame 100, ACL close to frame 380, HOA near frame 220, and KOA around frame 190. These differences emphasize that each pathology manifests discriminative movement signatures at distinct gait phases. By identifying when the model focuses on particular gait events initial contact, mid-stance, push-off, or late-swing, researchers can correlate these moments with clinical gait abnormalities.
Taken together, these three forms of visual explanation offer a rich, multilayered interpretation of the G-MASA-TCN model’s predictions. The IG heatmaps reveal which sensors the model treats as most informative, the global feature importance plots summarize the overall influence of each sensor across time, and the temporal relevance plots pinpoint when in the gait cycle the critical diagnostic information emerges. This integrated analysis enhances understanding of the model’s internal logic, enabling clinicians and researchers to identify not only the anatomical locations most relevant for classification but also the specific phases of movement that contain the strongest biomechanical indicators of pathology. Ultimately, these insights contribute to a more transparent and interpretable gait-classification system, supporting the reliability, clinical credibility, and potential translational application of the model across diverse gait disorders.
4. Discussion
4.1. Result Discussion
The results of this study demonstrate the effectiveness of integrating multi-sensor wearable gait data with advanced temporal deep learning architectures to capture clinically relevant movement signatures across a broad spectrum of neurological and musculoskeletal conditions. By leveraging information from four strategically placed IMUs located on the head, lower back, and both feet, the model was able to exploit both global and segment-specific movement patterns, resulting in highly accurate classification across diverse cohorts. The consistently strong performance across all experimental conditions underscores the richness of wearable inertial data and the suitability of temporal convolutional representations for gait analysis.
A key outcome of this work is the demonstration that gait dynamics, when represented as multi-channel time series, contain sufficiently distinctive patterns to differentiate among multiple pathologies as well as between individuals. The large, clinically annotated dataset used in this study provided a wide range of gait behaviors, including patterns associated with neuropathy, osteoarthritis, ligament injuries, and neurological impairment, alongside healthy control data. The high classification accuracy obtained across repeated cross-validation suggests that these conditions manifest in reproducible sensor-level and temporal signatures. This supports the broader view within the gait analysis community that wearable kinematic data can serve as a reliable proxy for more complex biomechanical measurements traditionally obtained using laboratory motion-capture systems.
Beyond accuracy, the interpretability component of this work provided valuable insights into how different conditions influence movement patterns. The Integrated Gradients analysis revealed that certain sensor channels particularly those associated with lower-back acceleration, foot rotations, and free-acceleration measurements played a dominant role in many classifications. These findings align with established biomechanical understanding: conditions such as neuropathy and osteoarthritis often alter foot–ground interaction forces, trunk stability, and gait symmetry, all of which are reflected in the accelerometer and gyroscope signals captured by wearable IMUs. Temporal attribution further highlighted that specific phases of the gait cycle, such as loading response, mid-stance, or push-off, were more informative for distinguishing pathological gait from normal movement. The convergence of these data-driven insights with clinical expectations strengthens confidence in the relevance and reliability of the extracted gait features. Further, the strong accuracy obtained on real-world gait data highlights the model’s ability to generalize effectively to uncontrolled walking conditions, capturing discriminative subject-specific gait signatures while maintaining robustness across different evaluation protocols.
Another notable result is the model’s ability to identify individuals with high accuracy in the subject-identification experiment. This indicates that gait, when captured at sufficient temporal resolution, functions as a unique biometric signature. Such findings have important implications for personalized healthcare, long-term patient monitoring, and the development of individualized rehabilitation protocols. They also highlight the robustness of the temporal representation learned by the model, which was able to generalize beyond pathology classification to a fundamentally different task requiring fine-grained discrimination.
Overall, the findings advance the field by showing that high-performing, interpretable gait classification is achievable using lightweight wearable sensors and modern deep learning techniques. The integration of explainable AI ensures that predictions are not treated as black-box outputs but are instead grounded in meaningful sensor-level and temporal evidence. This combination of performance, robustness, and transparency positions the proposed framework as a promising tool for future clinical applications, including automated screening, fall-risk assessment, disease progression monitoring, and personalized rehabilitation. Future work may explore real-time deployment, continuous gait tracking in free-living environments, and the use of domain adaptation to enhance applicability across populations and sensor configurations.
4.2. Comparison with the Original MASA-TCN (EEG Emotion Recognition)
Although our gait classification model retains the core structure of the original MASA-TCN proposed in [
26] for EEG-based emotion recognition, several adaptations were introduced to better suit the characteristics of inertial gait data. In the original work, the input consisted of relative power spectral density (rPSD) features computed across multiple frequency bands (typically 5–6), resulting in a tensor of the shape
, where
denotes the number of spectral bands. The SAT layer was designed to operate on a per-channel, per-frequency basis, using a context kernel of the size
and stride of
to extract band-specific temporal patterns, followed by spatial fusion across electrode channels via a kernel of the size
.
In contrast, our implementation processes raw time-domain IMU signals with 36 channels and no explicit frequency decomposition. The SAT context convolution therefore uses a kernel of the size , treating the full set of sensor streams as a single spatial dimension. This modification preserves the spirit of space-aware modeling but reinterprets “space” as anatomical sensor placement rather than spectral content. While the original model supported both continuous emotion regression (using Concordance Correlation Coefficient loss) and discrete classification, our application focuses exclusively on multi-class discrete classification via cross-entropy, necessitating only the mean-pooling classification head.
Activation functions also differ: while the original model used PReLU throughout, ReLU activation followed by batch normalization is applied in the MAAF block, which has shown superior stability when training on raw acceleration and angular velocity signals prone to high variance. The TCN dilation schedule is fixed at
in our version, yielding a consistent receptive field across experiments, whereas the original dynamically adjusted dilations starting from the SAT layer’s
. Despite these differences, the fundamental contribution—adaptive fusion of multi-scale, spatially informed temporal features—is retained and transferred from brain signals to full-body movement dynamics [
26].
4.3. Comparison with State-of-the-Art Gait Classifiers
To contextualize the performance and design of our G-MASA-TCN, it is instructive to compare it with leading deep learning architectures previously applied to IMU-based gait classification. One prominent approach is the DeepConvLSTM framework introduced by Ordóñez and Roggen (2016) [
38], which combines four convolutional layers for local feature extraction with two bidirectional LSTM layers to model long-range temporal dependencies. While effective, this hybrid architecture relies on sequential processing in the recurrent component, resulting in higher computational cost and memory usage during inference particularly challenging for real-time clinical applications. Moreover, the LSTM component introduces future context through bidirectionality, violating causality and preventing online deployment.
Another competitive baseline is the attention-augmented TCN-BiGRU model proposed in [
40], which stacks dilated temporal convolutional layers with bidirectional GRUs and attention mechanisms. This design achieves a large receptive field of approximately 340
via exponential dilation in the TCN:
This yields strong performance on activity recognition tasks; however, the inclusion of bidirectional recurrence again precludes causal inference, and the parameter count remains high (~1.0 M) due to the recurrent module and attention softmax computations. In contrast, our G-MASA-TCN eliminates recurrence entirely, relying instead on dilated convolutions and multi-scale parallelism to achieve an effective receptive field exceeding 1.2 s, sufficient to encompass multiple gait cycles while maintaining full causality and parallel computation.
More recently, Xiong Wei and Zifan Wan (2024) [
39] introduced a TCN-Attention model that augments a temporal convolutional backbone with channel-wise and temporal attention mechanisms to dynamically weigh sensor streams and key timestamps. This approach improves interpretability and adaptability to varying signal quality but introduces additional computational overhead from the attention scoring and softmax operations. Our G-MASA-TCN achieves comparable sensor prioritization implicitly through the learned 1 × 1 fusion weights in the MAAF block and the spatially joint 2D convolutions in SAT layers, without requiring explicit attention modules. This results in a simpler, faster model with fewer hyperparameters.
In terms of parameter efficiency, our model (0.92 M parameters) offers a favorable trade-off between capacity and deployability compared to heavier recurrent hybrids (>1.2 M) and attention-augmented models (~0.8–1.0 M). Crucially, the explicit multi-scale anchor design and built-in spatial fusion provide structured inductive biases tailored to gait kinematics, potentially leading to better generalization across heterogeneous patient populations compared to generic attention or recurrent mechanisms [
38,
39,
40].
Although direct comparison with prior methods on the same dataset is not possible due to the dataset’s recent release (October 2025), the existing literature highlights several limitations of commonly used baseline models for IMU-based gait analysis. Recurrent models such as GRUs often struggle with long-range temporal dependencies and are sensitive to sequence length and noise [
38,
51]. Transformer-based models, while powerful, typically require large-scale datasets and incur high computational cost, limiting their effectiveness in moderate-sized clinical datasets [
52,
53]. Similarly, conventional TCNs rely on fixed dilation patterns, which may restrict their ability to capture heterogeneous temporal dynamics across different gait pathologies [
45].
In contrast, the proposed G-MASA-TCN integrates multi-scale temporal modeling and adaptive attention mechanisms, enabling robust learning of both short- and long-term gait characteristics across diverse clinical conditions. This architectural design aligns with recent trends in state-of-the-art IMU-based gait analysis [
54] while extending them to a more challenging multi-cohort setting.
4.4. Computational Complexity and Power Consumption Analysis
The proposed G-MASA-TCN integrates multi-scale dilated temporal convolutions, channel-wise attention, and residual temporal modeling. While these design choices enhance representational capacity, they may increase computational demand. Therefore, a detailed complexity analysis in terms of Gigaoperations (GOPs) and an estimation of inference-time power consumption are provided to assess practical deployability.
4.4.1. Operation Count (GOP Analysis)
The dominant computational cost of G-MASA-TCN arises from one-dimensional convolutional layers. For a 1D convolution with input channels,
, output channels,
, the kernel size
, and the temporal length
, the number of floating-point operations (FLOPs) is approximated as
where the factor of 2 accounts for multiplication and addition. Using the fixed-length input setting adopted in this work (
,
), the total complexity per forward pass is summarized as follows:
SAT layers (multi-scale temporal extraction)
SAT-1: ; ; : ≈0.007 GOPs.
SAT-2: ; ; : ≈0.041 GOPs.
SAT-3: ; ; : ≈0.229 GOPs.
Attention mechanism
Temporal convolution block with residual connection
TCN-1: ; ; : ≈0.098 GOPs.
TCN-2: ; ; : ≈0.025 GOPs.
Residual projection (1 × 1 convolution): ≈0.008 GOPs.
Fully connected classification head
4.4.2. Power Consumption Estimation
To estimate inference-time energy consumption, we adopt commonly reported hardware efficiency values. Modern GPUs and embedded AI accelerators typically consume approximately 0.5–1.0 nJ per FLOP, while mobile CPUs require slightly higher energy per operation. Assuming a conservative value of 1 nJ/FLOP, the estimated energy per inference is
This low per-sample energy requirement indicates that G-MASA-TCN is suitable for real-time gait analysis and continuous monitoring scenarios. Importantly, inference is causal and convolution-based, avoiding the quadratic complexity associated with self-attention in Transformer models.
4.4.3. Discussion of Efficiency
Compared with baseline architectures evaluated in this study, G-MASA-TCN exhibits a favorable trade-off between accuracy and efficiency. Transformer-based models typically scale as with sequence length, leading to substantially higher computational and memory costs for long gait sequences. Recurrent architectures such as the GRU require sequential processing, limiting parallelization and increasing latency. In contrast, G-MASA-TCN scales linearly with temporal length () and supports full parallel inference. Despite integrating multi-scale temporal fusion and attention mechanisms, the model remains computationally lightweight, enabling deployment on wearable, edge, or clinical monitoring systems without excessive power consumption.
4.5. Real-Time Inference and Latency Analysis
Although G-MASA-TCN requires several hours for offline training due to cross-validation and subject-level learning, real-time deployment depends solely on inference-time complexity. Therefore, latency and throughput are analyzed using GOP-based benchmarking.
As derived in
Section 4.4, the proposed G-MASA-TCN requires approximately 0.41 GOPs per forward pass for a 5 s gait segment (500 samples at 100 Hz). Assuming a conservative inference with the input of 100 GOPs on a standard CPU, the expected inference latency is
On embedded AI platforms (300–1000 GOPs), latency reduces to 1.4–0.4 ms, while on modern GPUs (>5000 GOPs), sub-millisecond inference is achieved. These values are significantly lower than the data acquisition interval (5 s), confirming that G-MASA-TCN comfortably satisfies real-time gait analysis requirements. In contrast, Transformer-based architectures scale quadratically with sequence length (), leading to higher latency and memory overhead. Recurrent models such as the GRU require sequential processing, further increasing inference delay. G-MASA-TCN, by relying exclusively on parallel dilated convolutions, achieves low latency and deterministic inference time. These results demonstrate that the proposed framework is suitable for real-time clinical gait monitoring and wearable deployment.
4.6. Implications for Human Activity Recognition
The analytical framework developed in this study carries significant implications for the broader human activity recognition (HAR) domain. By integrating interpretable deep learning mechanisms with multi-sensor temporal modeling, the approach demonstrates that robust activity classification can be achieved without sacrificing transparency. This is particularly valuable in healthcare and safety-critical applications, where model accountability and clinical validation are essential. The ability of the framework to quantify sensor-level importance allows practitioners to understand not only what classification outcome was reached but also why, enabling deeper insight into the underlying physiological or behavioral patterns present in the data.
A central implication for the HAR field lies in the identification of key sensors and movement phases that consistently contribute to accurate predictions. Such insights can guide system designers in optimizing sensor placement, reducing instrumentation costs, and improving the feasibility of long-term monitoring systems. For example, the observation that trunk acceleration or foot rotational data frequently dominate model decisions suggests that monitoring these anatomical regions may be sufficient for capturing essential movement characteristics in certain applications. Moreover, temporal attribution highlights the specific phases of a movement cycle that carry the greatest informational value, supporting more refined feature engineering and potentially enabling segment-aware detection pipelines.
More broadly, the interpretability results demonstrate that complex human activities, whether related to gait, sports performance, workplace ergonomics, or daily-life monitoring, exhibit structured temporal patterns that can be captured reliably through lightweight wearable sensors. The combination of high accuracy and transparent decision pathways provides a foundation for translating HAR models into practical domains such as physical rehabilitation, athletic performance analysis, behavior tracking, occupational safety, and smart-home environments. In these contexts, understanding the reasoning behind system outputs is essential for ensuring acceptance by practitioners, promoting user trust, and facilitating regulatory approval when required.
4.7. Limitations and Generalization Considerations
While the framework demonstrates strong performance and interpretability, it is important to acknowledge several factors that influence the reliability and generalization of the findings. First, the Integrated Gradients (IG) method depends on smooth transitions between the baseline and actual inputs. Gait and other natural human activities often contain abrupt changes such as impacts, rapid directional shifts, or irregular compensatory movements which may challenge the smoothness assumption. In such situations, attributions may emphasize or de-emphasize certain regions of the signal, potentially affecting the interpretive clarity. Recognizing these conditions is important when applying IG to activities that exhibit high kinetic variability or irregular timing patterns.
Another consideration relates to the use of temporally averaged attribution measures. While these provide intuitive summaries of the most influential phases of an activity, they may mask more nuanced cross-sensor interactions or interdependencies that are crucial for understanding multidimensional human activities. Many real-world behaviors, whether in clinical populations or non-clinical settings, involve coordinated multi-joint movement patterns in which the interplay between different body segments conveys essential information. Averaging across features, although useful for visualization, may overlook these relationships and should therefore be interpreted with caution.
Generalizability across different populations, activity types, and environmental contexts also warrants consideration. Although the dataset used in this study is large and clinically diverse, real-world HAR applications often involve variations in sensor placement, individual movement styles, footwear, terrain, and daily-life unpredictability. These factors introduce variability that may affect the transferability of the learned representations. As such, extending the evaluation to larger, more heterogeneous datasets including free-living environments outside controlled laboratory settings would strengthen confidence in the model’s applicability to broader HAR tasks.
The interpretability insights generated here suggest opportunities for applying similar analytical frameworks in other fields that rely on multivariate physiological or behavioral time series. Domains such as cardiology, sleep analysis, workplace fatigue monitoring, and sports biomechanics all involve structured temporal signals where identifying key sensors and critical time windows can enhance both model performance and domain understanding. The observations from this study underscore the importance of integrating transparent learning mechanisms with rich sensor data to support informed decision-making in technologically assisted human-monitoring systems.
Importantly, this study serves as a pillar for future research, encouraging the community to extend data collection to more diverse and real-world scenarios. Potential directions include incorporating children, rare gait disorders, and everyday walking activities, such as stair climbing, navigating uneven terrain, or carrying objects. By providing this baseline, the current work lays the groundwork for advancing IMU-based gait analysis in complex and clinically relevant environments and inspires further innovations in low-cost, wearable sensor applications.