Sensor-Based Classification of Post-Stroke Motor Impairment Using Fugl-Meyer Lower Extremity Scores
Round 1
Reviewer 1 Report
Comments and Suggestions for Authors-
Regarding the wording, we believe it is necessary to define some of the terms expressed as acronyms beforehand, with the terminology used throughout the rest of the document appearing in parentheses. This is the general practice, but there are some exceptions, even though the authors summarize the nomenclature used in the work at the end. In this sense, since it is surface electromyography, we suggest using the terminology “sEMG” in the document.
-
The writing, however, is somewhat vague and at times confusing, which, in this reviewer's opinion, is related to the supplementary material presented. The main questions this reviewer had while reading the paper are perfectly explained there. I suggest that aspects related to the sEMG measurements and the use of the data in training, testing, and validation be included in the main body of the paper for greater clarity. Similarly, there are elements that would support the results and should also be included in the document, not as supplementary material.
-
This reviewer is unclear as to which stage or stages the correlations presented correspond to the Fugl-Myer instrument or scale. The procedure is primarily used to evaluate progress in the rehabilitation of post-stroke patients. This reviewer is uncertain about the timing of the measurements or whether they correspond to multiple sessions. Furthermore, the clinical diagnosis of the stroke cases, beyond ischemic or hemorrhagic, is not specified, nor is the time frame for patient evolution or the level of spasticity indicated. What does the authors' use of the term "controllable spasticity" mean in terms of the spasticity scale?
-
As a point of interest for this reviewer, why decision trees and not CNNs? In this same vein, given the advancements in Computer Vision models with Transformer blocks and self-care nodes, explored by this reviewer, and "contactless" systems, why not these? What is the authors' opinion on the hypothesis that, beyond measuring the "now," algorithms using the patient's initial Fugl-Meyer score, demographic data, and biomarkers (such as MRI scans of the lesion or responses to transcranial magnetic stimulation), machine learning could predict the patient's Fugl-Meyer score after 6 or 12 weeks of therapy?
-
In this regard, does the authors' proposal clinically explain the reason for the score obtained—is it due to synergistic isolation, reflex speed, or trunk compensation? Answers to these questions will foster confidence in these applications among medical personnel.
-
On the other hand, do the authors consider the sample used to be suitable for generalization and free of bias? Are there any future predictions of how much the indicators presented may be depressed in "noisier" clinical environments?
-
This reviewer, with some clinical rehabilitation experience, wonders: is the proposed method capable of detecting deceptive motor compensations that are visually apparent to the physiotherapist? In this regard, computer vision-based algorithms appear to be less invasive and more suitable for detecting motor compensation.
-
Finally, this reviewer is concerned about sensor migration. If the sensor moves or rotates even a couple of centimeters, the data changes drastically, introducing noise and confusing the machine learning algorithm. In this reviewer's experience, transformer models are an option to consider.
Author Response
Regarding the wording, we believe it is necessary to define some of the terms expressed as acronyms beforehand, with the terminology used throughout the rest of the document appearing in parentheses. This is the general practice, but there are some exceptions, even though the authors summarize the nomenclature used in the work at the end. In this sense, since it is surface electromyography, we suggest using the terminology “sEMG” in the document.
Response: We thank the reviewer for this suggestion. The manuscript has been revised accordingly to ensure that all acronyms are properly defined at first occurrence, followed by the abbreviated form in parentheses, and used consistently throughout the text. In particular, surface electromyography has now been introduced as sEMG at its first occurrence, and the acronym sEMG is used consistently thereafter across the manuscript.
The writing, however, is somewhat vague and at times confusing, which, in this reviewer's opinion, is related to the supplementary material presented. The main questions this reviewer had while reading the paper are perfectly explained there. I suggest that aspects related to the sEMG measurements and the use of the data in training, testing, and validation be included in the main body of the paper for greater clarity. Similarly, there are elements that would support the results and should also be included in the document, not as supplementary material.
Response: We thank the reviewer for this observation. We agree that several methodological details were previously described only in the Supplementary Material, which may have reduced clarity in the main manuscript.
In the revised version, we have expanded the main text to integrate key methodological details regarding the acquisition protocol, signal processing, feature extraction procedures (lines 91-127), and the model training, validation, and testing procedures (lines 167-187 and 202-220).
We believe that these revisions improve the clarity and reproducibility of the manuscript by ensuring that all essential methodological steps are now described within the main text.
This reviewer is unclear as to which stage or stages the correlations presented correspond to the Fugl-Myer instrument or scale. The procedure is primarily used to evaluate progress in the rehabilitation of post-stroke patients. This reviewer is uncertain about the timing of the measurements or whether they correspond to multiple sessions. Furthermore, the clinical diagnosis of the stroke cases, beyond ischemic or hemorrhagic, is not specified, nor is the time frame for patient evolution or the level of spasticity indicated. What does the authors' use of the term "controllable spasticity" mean in terms of the spasticity scale?
Response: We thank the reviewer for this comment. Participants were categorized into low, mid, and high impairment levels derived from the FMA-LE motor score. These are based on established thresholds from Smith et al. [18] and Kwong et al. [19], with high impairment defined as FMA-LE < 9 (non-ambulatory patients, which were not included in the study cohort due to the dataset inclusion criteria), and mid/low impairment defined within the ranges 9–21 and ≥21, respectively (lines 133-140). The FMA-LE score used for labelling corresponds to a single cross-sectional clinical assessment per participant, not a longitudinal evaluation across rehabilitation stages or multiple sessions (lines 131-132).
Regarding the dataset composition, for the cohort collected at the Hospital of Braga, stroke etiology was available (3 ischemic), with a post-stroke time since event of 12 ± 7 months (line 81-82). For the remaining participants from the publicly available ARRA dataset, detailed clinical metadata such as stroke subtype and time since stroke were not available[16]. The lack of detailed clinical metadata for the ARRA dataset, including stroke subtype and time since stroke, has also been explicitly acknowledged as a limitation in the revised Discussion (lines 379-381). Furthermore, we acknowledge that spasticity level was not quantitatively assessed using a standardized clinical scale (e.g., Modified Ashworth Scale), and therefore it is not included in the present analysis. This has been clarified in the revised manuscript as a limitation (lines 381-385). Finally, the criteria “lower limb muscle spasticity medically controlled” refers to cases where lower-limb spasticity was clinically managed by the treating physician (e.g., pharmacological intervention), without implying a specific numerical classification on a standardized spasticity scale. This clarification has been added to the manuscript to avoid ambiguity (lines 83-84).
As a point of interest for this reviewer, why decision trees and not CNNs? In this same vein, given the advancements in Computer Vision models with Transformer blocks and self-care nodes, explored by this reviewer, and "contactless" systems, why not these? What is the authors' opinion on the hypothesis that, beyond measuring the "now," algorithms using the patient's initial Fugl-Meyer score, demographic data, and biomarkers (such as MRI scans of the lesion or responses to transcranial magnetic stimulation), machine learning could predict the patient's Fugl-Meyer score after 6 or 12 weeks of therapy?
Response: We thank the reviewer for this insightful set of questions. Regarding the choice of classifier, the proposed study is based on features aggregated at the trial level, where each walking trial is represented by a single averaged gait cycle curve. The available data used in this study were harmonized at the trial-feature level since the ARRA dataset was not available as raw multi-cycle time-series suitable for robust training of deep temporal architectures such as CNNs or transformers. Given the limited sample size (84 trials from 32 participants), an interpretable model with controlled complexity (decision tree with constrained depth) was preferred, while higher-capacity deep learning models such as CNNs would be expected to present an increased risk of overfitting under these data conditions [1-2].
Regarding more advanced architecture and contactless systems (e.g., vision transformers or deep CNN-based pose estimation), we agree that these represent promising directions for future work, particularly when large-scale datasets with raw kinematic or video data are available. In the present study, kinematic information was derived from a contactless camera-based system but was subsequently reduced to spatiotemporal summary features (only summary features were publicly available), which showed lower classification performance compared to sEMG features.
Regarding longitudinal prediction of FMA-LE evolution, we agree this may be a highly relevant and clinically valuable direction (lines 403-405). Prior work such as Tozlu et al.[11] demonstrated that multimodal models combining clinical, transcranial magnetic stimulation (TMS) responses over 18 sessions, and imaging biomarkers achieved strong performance for upper-limb assessment. However, evaluating the applicability of such approaches for lower-limb assessment requires longitudinal datasets together with multimodal clinical, imaging, and TMS data, which were not available in the present study. Additionally, for practical longitudinal monitoring, we argue that sEMG-derived features may represent a more scalable and repeatable biomarker than MRI or TMS, which are generally more resource-intensive and less suitable for frequent assessments.
In the present study, the decision tree was not intended to compete with high-capacity deep learning models, but to provide an interpretable baseline for evaluating which sensor-derived biomarkers are informative for FMA-LE-derived impairment classification under limited data conditions.
[1] Goodfellow, I.; Bengio, Y.; Courville, A. Deep Learning; MIT Press, Ed.; 2016; ISBN 9780262035613.
[2] Breiman, L. Classification And Regression Trees; TAYLOR & FRANCIS LTD, Ed.; 1984; ISBN 9780412048418.
In this regard, does the authors' proposal clinically explain the reason for the score obtained—is it due to synergistic isolation, reflex speed, or trunk compensation? Answers to these questions will foster confidence in these applications among medical personnel.
Response: We thank the reviewer for this important question. The decision tree model provides transparent, rule-based decision boundaries derived from sensor features. This allows inspection of which features and thresholds contribute to the motor impairment classes derived from FMA-LE total motor score. However, because the model was trained using FMA-LE-derived total motor score rather than individual FMA-LE item scores, it cannot determine whether a given classification is specifically driven by specific motor impairment components (e.g., synergy control, reflex activity, or coordination). Moreover, because the FMA-LE does not quantify compensatory movement strategies, such a model would still be unable to identify whether a patient's performance reflects motor recovery or compensation.
Therefore, while the model supports interpretability at the feature-decision level, mapping these features to specific neurophysiological mechanisms requires additional clinical validation studies. For example, quantitative measures of trunk compensation could be used to assess whether the rule-based decision boundaries produced by the model are associated with compensatory movement patterns. This would help support the neurophysiological interpretation of the model outputs. We have clarified this limitation in the revised manuscript and highlighted clinician-in-the-loop interpretation as an important direction for future work (lines 399-401).
On the other hand, do the authors consider the sample used to be suitable for generalization and free of bias? Are there any future predictions of how much the indicators presented may be depressed in "noisier" clinical environments?
Response: We thank the reviewer for this important question. Our results showed a reduction in performance when the model was evaluated on the hybrid test set, which included data from the external Hospital of Braga cohort in addition to ARRA test data (MCC dropped from 0.7 to 0.6). This indicates that the current model may not fully generalize across different acquisition settings and populations (lines 369-370). However, the primary objective of this study was not to develop a final, fully generalizable classification model, but rather to evaluate multiple feature sets composed of sensor-based biomarkers acquired during walking for automated estimation of post-stroke motor impairment levels derived from FMA-LE. In this context, the classification framework serves as a tool to assess the relevance of the extracted features and guide the future development of a robust, generalizable clinical decision-support model.
Regarding robustness in noisier clinical environments, we investigated the effect of signal perturbations through a noise augmentation strategy applied during training. Specifically, we introduced white and pink noise to emulate common sources of variability in sEMG signals, such as sensor amplification noise and electrode–skin interface variability, respectively. We observed that models trained with noise-augmented data showed improved performance on the test set, suggesting improved robustness to the specific signal perturbations simulated in this study (MCC increased from 0.4 to 0.7) (lines 351-357). It is also important to note that the data were collected in a hospital environment rather than under strictly controlled laboratory conditions. Therefore, the recordings already included sources of variability commonly encountered in clinical practice. Moreover, the sEMG Delsys acquisition system complies with requirements of the Medical Device Regulation EU 2017/745, supporting the reliability and quality of the signal acquisition. However, we acknowledge that this does not guarantee performance stability under broader clinical conditions (e.g., electrode displacement). Further validation using larger, multi-site datasets collected under heterogeneous clinical settings will be required to fully quantify robustness and generalizability (lines 388-390).
This reviewer, with some clinical rehabilitation experience, wonders: is the proposed method capable of detecting deceptive motor compensations that are visually apparent to the physiotherapist? In this regard, computer vision-based algorithms appear to be less invasive and more suitable for detecting motor compensation.
Response: We thank the reviewer for this comment. We agree that computer vision-based approaches are less invasive and closely resemble the visual assessment routinely performed by physiotherapists, making them well suited for detecting observable compensatory movement patterns. However, sEMG provides physiological information on muscle activation, enabling the assessment of neuromuscular coordination that may not be directly observable from movement alone and is often impaired in post-stroke individuals [3]. Therefore, these physiological characteristics may provide complementary information to kinematic observations.
In its current form, the proposed method was not intended to explicitly detect or quantify compensatory movement strategies, including those that are visually apparent to physiotherapists. Instead, the sEMG and kinematic signals were used as input features for classifying motor impairment levels derived from FMA-LE motor score. Nevertheless, investigating the ability of these signals to identify specific compensatory strategies, and comparing their performance with computer vision-based methods, is an important direction for future research (lines 399-400).
[3] Den Otter, A. R., Geurts, A. C., Mulder, T., & Duysens, J. (2007). Abnormalities in the temporal patterning of lower extremity muscle activity in hemiparetic gait. Gait & posture, 25(3), 342–352. https://doi.org/10.1016/j.gaitpost.2006.04.007
Finally, this reviewer is concerned about sensor migration. If the sensor moves or rotates even a couple of centimeters, the data changes drastically, introducing noise and confusing the machine learning algorithm. In this reviewer's experience, transformer models are an option to consider.
Response: We thank the reviewer for highlighting the important issue of sensor migration. We acknowledge that electrode displacement or rotation can introduce signal variability, which may affect model generalization.
In the Hospital of Braga dataset, electrode placement followed SENIAM recommendations, with skin preparation and supervised acquisition to minimize placement-related variability. For the ARRA dataset, sEMG recordings were used as provided by the original dataset authors after their preprocessing pipeline. In addition, feature extraction was performed at the gait-cycle level using normalized averaged curves, which helps reduce cycle-to-cycle variability and residual noise while preserving the overall muscle activation patterns.
We acknowledge that the present noise-augmentation strategy does not explicitly simulate electrode displacement or rotation. Therefore, robustness to sensor migration remains to be specifically evaluated in future studies, either through controlled electrode-shift experiments or augmentation strategies designed to emulate sensor displacement.
Given the relatively limited sample size (32 participants), we prioritized models with controlled complexity and higher interpretability, such as a decision tree with constrained depth. In this context, deep learning approaches, including transformer architectures, were not explored due to the increased risk of overfitting under these data conditions [1–2]. We acknowledge that deep learning approaches may offer advantages in scenarios with larger and more diverse datasets, and we consider their investigation an interesting direction for future work (lines 405-408).
Reviewer 2 Report
Comments and Suggestions for AuthorsThe manuscript explores several sets of sensor-based biomarkers in a process to estimate motor impairment in post-stroke patients. The research draws on a Decision-Tree classifier to analyze open-source and collected data sets to optimize/improve the Fugl-Meyer Assessment for Lower Extremity, with the aim of assisting clinicians in proposing rehabilitation decisions for patients. The inclusion of EMG parameters appears to enhance objective assessment, particularly when trying to distinguish between subtle differences in impairment.
The manuscript provides excellent insight into the use of EMG features to improve model performance, and acknowledges that stride and step length spatiotemporal parameters also correlate with impairment -- inclusion of demographics also enhanced the validation scoring and thus the work sucessfully models data of quite different frquency characteristics. Overall the work provides a compelling argument for the proposed model to enhance FMA-LE.
Some items that may be helpful to consider:
- The results for EMG are highly encouraging, but it sems that conventional EMG is prone to movement artifacts, external electrical noise and skin-contact impedance changes. It would be helpful to understand if any of these issues arose or were of concern during data collection.
- EMG does not necessarily focus on a single muscle given the possibility of cross-talk from adjacent muscles. It may be helpful elaborate or comment on any issues such as cross-talk that may confound individual muscle data collection. This may have been addressed in manuscript but overlooked during this review.
- Relative to the number of EMG signals and kinematic parameters, it seems there is the possibility of overfitting the data on what might be considered a somewhat small dataset. This seems not to have been a concern as suggested in Lines 299-303; though it may be useful for some comment be included on the evident lack of overfitting.
- At the point where the minimization of sensor input is limited to two muscles, computational efficiency does, indeed, improve and encourages efficient translation for clinical use. However, there might seem to be overlooked any compensatory muscles a stroke survivor might be used that is not included in the existing data set.
- It would help to more fully understand any concerns or data intergration challenges that arose when blending high-frequency EMG signals with static demographic data and lower-frequency gait metrics.
- The manuscript is clear that the preliminary results using retrospective open-source data sets is a limitation, and the authors identify that prospective clinical validation is needed for furthers tests of the model. This does not take away from what appears to be a successful effort to strategically improve the FMA-LE derived classes.
Author Response
The manuscript explores several sets of sensor-based biomarkers in a process to estimate motor impairment in post-stroke patients. The research draws on a Decision-Tree classifier to analyze open-source and collected data sets to optimize/improve the Fugl-Meyer Assessment for Lower Extremity, with the aim of assisting clinicians in proposing rehabilitation decisions for patients. The inclusion of EMG parameters appears to enhance objective assessment, particularly when trying to distinguish between subtle differences in impairment.
The manuscript provides excellent insight into the use of EMG features to improve model performance, and acknowledges that stride and step length spatiotemporal parameters also correlate with impairment -- inclusion of demographics also enhanced the validation scoring and thus the work sucessfully models data of quite different frquency characteristics. Overall the work provides a compelling argument for the proposed model to enhance FMA-LE.
Response: We sincerely thank the reviewer for the positive evaluation of our work.
Some items that may be helpful to consider:
The results for EMG are highly encouraging, but it sems that conventional EMG is prone to movement artifacts, external electrical noise and skin-contact impedance changes. It would be helpful to understand if any of these issues arose or were of concern during data collection.
Response: We thank the reviewer for this comment. For the Hospital of Braga dataset, sEMG recordings were acquired under a controlled experimental protocol designed to minimize these effects. Specifically, electrode placement followed the SENIAM recommendations, appropriate skin preparation was performed prior to sensor placement, and all walking trials were supervised to ensure proper sensor attachment and consistent data acquisition (lines 108-111).
Following acquisition, the sEMG signals were pre-processed using the filtering pipeline to reduce baseline drift and motion-related artifacts. In addition, the analysis was performed using normalized averaged sEMG curves at the gait-cycle level. This representation helps reduce the influence of cycle-to-cycle variability and residual noise while preserving the overall muscle activation patterns (lines 111-112).
Regarding the publicly available ARRA dataset, the sEMG signals were provided as pre-processed recordings. As described by the original dataset authors, the signals had already undergone filtering, and the dataset provided normalized averaged sEMG waveforms at the gait-cycle level. Therefore, the analyses in both datasets were performed using comparable processed sEMG curves (lines 96-99).
Nevertheless, we recognize that residual movement artifacts, external electrical noise, and electrode-skin impedance variations cannot be fully excluded in dynamic walking recordings.
EMG does not necessarily focus on a single muscle given the possibility of cross-talk from adjacent muscles. It may be helpful elaborate or comment on any issues such as cross-talk that may confound individual muscle data collection. This may have been addressed in manuscript but overlooked during this review.
Response: In the present study, electrode placement followed standard anatomical SENIAM guidelines to target primary muscle groups involved in lower-limb gait, and signal preprocessing was applied to improve signal quality. However, we recognize that these measures do not fully eliminate crosstalk, particularly in dynamic walking conditions. Thus, the identified relevance of individual muscles should not be interpreted as a strictly muscle-isolated physiological effect. This limitation has now been explicitly added to the manuscript, and we clarify that muscle-specific interpretations should be made with caution (lines 324-327).
Relative to the number of EMG signals and kinematic parameters, it seems there is the possibility of overfitting the data on what might be considered a somewhat small dataset. This seems not to have been a concern as suggested in Lines 299-303; though it may be useful for some comment be included on the evident lack of overfitting.
Response: We thank the reviewer for this important observation. To assess model overfitting, we evaluated the relationship between training and validation performance (MCC) across different values of the decision tree maximum depth parameter which represents the model complexity (Figure S3). The selected configuration (max depth = 4) resulted in a training MCC of 0.77 and a validation MCC of 0.70, indicating a relatively small performance gap and no strong indication of overfitting within the cross-validation framework used, although this does not exclude overfitting or instability due to the limited sample size. We have clarified in the revised manuscript that this hyperparameter setting was chosen as it represents the best trade-off between model complexity and performance (lines 359-362).
However, we acknowledge that cross-validation results alone do not guarantee generalization to unseen populations. Indeed, when evaluating the model on an independent dataset (from Hospital of Braga), a reduction in performance was observed. This suggests that, while overfitting within the training/validation procedure appears limited, the overall generalizability of the model may be constrained by dataset size and population heterogeneity. This limitation has been explicitly added to the revised manuscript (lines 386-390).
At the point where the minimization of sensor input is limited to two muscles, computational efficiency does, indeed, improve and encourages efficient translation for clinical use. However, there might seem to be overlooked any compensatory muscles a stroke survivor might be used that is not included in the existing data set.
Response: We thank the reviewer for this important observation regarding the potential omission of relevant compensatory muscle activity when reducing the number of recorded sEMG channels. We agree that limiting the analysis to a reduced subset of muscles improves practicality and computational efficiency but may also reduce the ability to capture the compensatory activation patterns that can occur in post-stroke gait. In the present study, the selection of a smaller muscle set was treated as a feature reduction strategy, aimed at evaluating the trade-off between model performance and sensor burden rather than providing a complete characterization of all possible compensatory mechanisms. Therefore, the reduced-muscle configuration should be interpreted as a preliminary feasibility finding rather than as a definitive recommendation for clinical deployment. We acknowledge that future studies should investigate the trade-off between sensor reduction and physiological completeness to ensure that clinically relevant compensatory strategies are not overlooked (lines 335-337).
It would help to more fully understand any concerns or data intergration challenges that arose when blending high-frequency EMG signals with static demographic data and lower-frequency gait metrics.
Response: We thank the reviewer for this comment. In the present study, high-frequency sEMG signals were first processed into frequency-domain features (mean frequency, median frequency, and peak frequency), while gait features (step length, stride length, stride time, and cadence) were extracted from low-frequency data (kinematic signals and ground reaction force) (lines 114-124). Demographic variables (age, gender, body mass, and paretic side) were included as static covariates (lines 124-127). As a result, each walking trial was represented by a feature vector, where each element corresponded to one extracted feature (sEMG, gait, or demographic), providing a numerical representation of the trial for machine learning classification.
Because all modalities were converted into trial-level features before classification, the model did not directly fuse signals at their native sampling frequencies. Therefore, the main integration challenge was the loss of temporal information rather than temporal synchronization during model training. We acknowledge that this may lead to the loss of potentially informative dynamic patterns across time. If synchronized time-series datasets become available, future work could investigate sequence-based models that preserve the temporal structure of sEMG and kinematic signals to assess whether additional meaningful information can be extracted beyond the current feature-based representation (lines 405-408).
