1. Introduction
Objective quantification of upper-limb force production is increasingly relevant to wearable movement monitoring, rehabilitation engineering, and individualized motor training, because force-level changes cannot always be captured by visual observation or coarse ordinal assessment alone [
1,
2]. Manual muscle testing and related ordinal scales are simple and clinically accessible, but their limited score resolution and potential inter-rater variability make it difficult to identify small changes in force generation during repeated or high-score movements [
3,
4]. Wearable biosignals provide quantitative measurements of muscle electrical activation, mechanical vibration, and movement-related responses during standardized tasks [
5,
6]. Surface electromyography (sEMG) has been widely used to characterize muscle activation and motor-unit behavior, although its amplitude and spectral features can be influenced by electrode–skin impedance, crosstalk, motion artifacts, and preprocessing choices [
7,
8,
9,
10]. Mechanomyography (MMG) provides a complementary mechanical view of muscle contraction by capturing surface vibration, contractile oscillation, and load-related changes that are not directly represented in sEMG [
5,
11,
12].
Recent work in multimodal wearable sensing suggests that combining neuromuscular, mechanical, and movement-derived information can improve sensor-based movement and fatigue assessment [
13,
14,
15]. In shoulder and elbow motion estimation, biosignal fusion improved the coefficient of determination by 0.041–0.053 compared with MMG-only inputs [
13]. A manual-lifting study also reported high fatigue-classification accuracy when sEMG and MMG information was jointly modeled [
16]. Experimental studies using MVC and submaximal loading further support MMG as a mechanical counterpart to sEMG in force-related tasks [
11,
12]. However, previous work has mainly addressed joint acceleration, fatigue, gesture recognition, or load-lifting behavior. Evidence remains limited on whether sEMG and MMG provide complementary information for discrete MVC-normalized shoulder-abduction force levels, particularly when single-modality and fused feature sets are compared directly.
We hypothesized that the feature-level fusion of sEMG and MMG would improve four-level MVC-normalized shoulder-abduction classification and reduce ordinal prediction error compared with either modality alone. To test this hypothesis, we benchmarked eight supervised classifiers and compared sEMG-only, MMG-only, directly concatenated, and correlation-augmented fusion feature sets. Group-aware cross-validation at the action-trial level was used to reduce leakage between overlapping windows, and class-specific errors were examined to characterize the limits of force-level separability [
17,
18].
2. Materials and Methods
This healthy-volunteer study evaluated whether the feature-level fusion of surface electromyography (sEMG) and mechanomyography (MMG) signals could improve the classification of maximum voluntary contraction (MVC)-normalized shoulder-abduction force levels. Shoulder abduction was selected, because it is relevant to upper-limb function and rehabilitation training and permits controlled loading across relative force levels. The middle deltoid was selected as a major superficial agonist during shoulder abduction and as an accessible site for colocated noninvasive sEMG and MMG acquisition [
19,
20]. The overall workflow included multimodal signal acquisition, preprocessing, sliding-window feature extraction, feature-level fusion, supervised classification, group-aware validation at the action-trial level, and feature-set ablation (
Figure 1).
2.1. Participants
Ten healthy adults were recruited, including six males and four females. All participants were right-handed. Their age, height, and body weight were years, cm, and kg, respectively. None reported upper-limb neuromuscular or musculoskeletal disorders. Participants were instructed to avoid fitness training, resistance exercise, and other activities that imposed substantial upper-limb loading during the week before testing. All participants were informed of the experimental procedure and signed informed consent forms before data collection. Because the experiment involved healthy adults, the labels were interpreted as MVC-normalized force-level categories rather than clinically assigned manual muscle testing (MMT) grades.
2.2. Experimental Protocol and Force-Level Labeling
Participants sat upright with the trunk maintained in a neutral position, and the right upper limb was used as the test side. Before formal acquisition, each participant practiced the shoulder-abduction task under the supervision of a rehabilitation physician. Each repetition began with the upper limb in the neutral position. The participant then abducted the arm while holding the assigned adjustable dumbbell and maintained the abducted posture for 2 s. Each movement was repeated 10 times, and each participant completed three sets. A 2 min rest interval was provided between sets to reduce fatigue-related changes in the signals.
Before formal data collection, each participant completed five shoulder-abduction trials at the maximum load achievable with the adjustable dumbbell. The maximal abducted posture was maintained for 2 s in each trial, and the mean load was used as the individual MVC reference. Target dumbbell loads were then set at approximately 10%, 30%, 60%, and 90% of this reference. These conditions were assigned to Labels 2, 3, 4, and 5, respectively (
Table 1). The labels therefore represent four MVC-normalized external-load conditions under a controlled healthy-volunteer protocol and should not be interpreted as direct clinical MMT scores.
2.3. Signal Acquisition and Preprocessing
sEMG and MMG signals were recorded using a custom-developed wearable acquisition device. The sEMG channel used dry bipolar surface electrodes with a center-to-center spacing of 20 mm, an input range of mV, a gain of 1000, and input-referred noise of no more than 5 V RMS. The MMG channel incorporated an ADXL355 three-axis digital microelectromechanical systems (MEMS) accelerometer (Analog Devices, Inc. Wilmington, MA, USA). According to the manufacturer’s specifications, the ADXL355 supports user-selectable g, g, and g ranges, uses a 20-bit analog-to-digital converter, and has a typical sensitivity of 256,000 LSB/g in the g range. In the present protocol, the sEMG and MMG channels were synchronously sampled from the middle deltoid at 1000 Hz using a shared timestamp. The signals were transmitted to a laptop via Bluetooth and subsequently exported for feature extraction and modeling.
Before sensor placement, the skin surface was cleaned with alcohol wipes to reduce skin impedance. Sensor placement followed the general SENIAM recommendations for the middle deltoid [
20]. The sensor was placed near the midpoint between the acromion and the lateral epicondyle of the humerus, and its longitudinal axis was aligned with the muscle-fiber direction (
Figure 2).
The raw signals were preprocessed before feature extraction. For MMG, a fourth-order Butterworth band-pass filter with cutoff frequencies of 10 and 100 Hz was applied. Wavelet denoising was then performed using the sym5 wavelet basis with level-3 decomposition and soft-thresholding to reduce motion artifacts while retaining mechanical vibration information. For sEMG, a fourth-order Butterworth band-pass filter with cutoff frequencies of 20 and 450 Hz was used. No separate 50-Hz power-line notch filter was applied.
Figure 3 shows representative preprocessing results for the sEMG and MMG signals.
2.4. Feature Extraction and Feature-Level Fusion
Preprocessed continuous signals were segmented using a sliding-window strategy. The window length was set to 500 ms, and the stride was set to 150 ms. For each window, the predefined 26-feature set was retained for the main analysis, including 10 sEMG-derived features, 15 MMG-derived features, and one correlation-based fusion feature (
Table 2). Three inertial measurement unit (IMU)-derived kinematic features in the feature table were not included, because the present study focused on sEMG-MMG feature-level fusion.
The 10 sEMG-derived features included time-domain energy and complexity features, a frequency-domain feature, and model-based features: root mean square (EMG_RMS), waveform length (EMG_WL), zero crossing (EMG_ZC), slope sign change (EMG_SSC), Willison amplitude (EMG_WAMP), mean frequency (EMG_MNF), and fourth-order autoregressive coefficients (EMG_AR1–EMG_AR4). The 15 MMG-derived features included MMG root mean square (MMG_RMS_Mech), kurtosis (MMG_Kurtosis), mean power frequency (MMG_MPF), and 12 Mel-frequency cepstral coefficients (MMG_MFCC_1–MMG_MFCC_12). The additional Fusion_Corr_MA feature was retained from the predefined feature table as a window-level correlation-based feature designed to represent local sEMG-MMG amplitude coupling.
Feature-level fusion was implemented by concatenating the sEMG-derived, MMG-derived, and correlation-based features into a unified feature vector before model training. Let
denote the sEMG feature vector and
denote the MMG feature vector. The predefined 26-feature fusion vector was defined as
where
represents the Fusion_Corr_MA feature. This window-level electromechanical coupling feature was calculated as the Pearson correlation between the time-aligned moving-average amplitude sequences of rectified sEMG and MMG:
where
and
denote the paired moving-average amplitude sequences within the same 500 ms analysis window. The feature was intended to summarize local covariation between electrical activation and mechanical vibration. For scale-sensitive models, including LR, KNN, SVM-RBF, and MLP, z-score standardization was fitted only on the training folds and then applied to the corresponding validation folds. Tree-based models were trained using the extracted feature values directly.
2.5. Feature-Set Ablation Design
To examine whether the performance gain was attributable to multimodal fusion rather than a single signal source, four feature sets were compared in the ablation analysis (
Table 3). The main feature set was the predefined 26-feature fusion set used in the main analysis. The 25-feature sEMG+MMG set was additionally analyzed to examine whether the correlation-based fusion feature changed model performance beyond direct concatenation of sEMG and MMG features.
2.6. Classification Models
Eight representative supervised classifiers were evaluated to compare the performance of the feature-level fusion strategy: logistic regression (LR), k-nearest neighbors (KNN), decision tree (DT), support vector machine with radial basis function kernel (SVM-RBF), random forest (RF), extremely randomized trees (ET), histogram-based gradient boosting decision tree (HGBDT), and multi-layer perceptron (MLP). These models were selected to cover linear, distance-based, single-tree, margin-based, bagging ensemble, randomized-tree ensemble, boosting ensemble, and shallow neural-network classifiers.
Hyperparameters were selected empirically through repeated preliminary comparisons before the formal evaluation. The selected parameter set was then held fixed across all five cross-validation folds.
LR was implemented with L2 regularization, , class-balanced weighting, and a maximum of 5000 iterations. KNN used seven neighbors with distance weighting. DT used the Gini criterion with class-balanced weighting. SVM-RBF used , gamma set to “scale”, and class-balanced weighting. RF and ET each used 100 trees, class-balanced weighting, and unrestricted maximum depth. HGBDT was configured with 120 boosting iterations, a learning rate of 0.05, 31 maximum leaf nodes, L2 regularization of 0.01, and early stopping. MLP used two hidden layers with 64 and 32 neurons, rectified linear unit activation, Adam optimization, L2 regularization coefficient of 0.001, adaptive learning rate, and early stopping. The random seed was set to 42 for reproducibility when applicable.
2.7. Validation Strategy and Performance Metrics
Because the feature samples were generated from overlapping sliding windows, random row-level splitting could overestimate the classification performance by placing highly correlated neighboring windows in both training and validation sets. Therefore, group-aware cross-validation at the action-trial level was used as the primary validation strategy. The action-trial identifier was retained from the analysis feature table, and all windows from the same action trial were assigned to the same fold. Five folds were used. This design evaluated recognition of unseen action trials and reduced leakage between neighboring windows. It did not ensure that all trials from one participant were restricted to a single fold.
For each fold, models were trained on the training action-trial groups and evaluated on unseen action-trial groups. The performance was summarized using accuracy, balanced accuracy, macro-precision, macro-recall, macro-F1, weighted-F1, linear weighted Cohen’s kappa, quadratic weighted Cohen’s kappa (QWK), mean absolute grade error (MAE), and within-one-grade accuracy. Weighted kappa and MAE were included, because the target labels represented ordered force-level categories. Within-one-grade accuracy was defined as the proportion of samples for which the absolute difference between the predicted and true labels was no greater than 1.