Next Article in Journal
AQI-HNS: A Security-Aware Hybrid Framework for Quantum-Inspired Image Encryption and Neural Image Hiding with Cross-Dataset Evaluation
Previous Article in Journal
Product-Manifold Contrastive Learning for Characterizing Neural Representations of 3D Visual Transformations in the Avian Visual System
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

sEMG-Driven Robotic Manipulation in Simulation: LOPO Classification, Confidence-Gated Supervision, and Gesture-Scheduled LQR Control

by
Anas Hassan Abdelmoneam
1,
Nourhan Zayed
1,2,*,
Mohamed S. Abdallah
3,4,* and
Mostafa Abdelaziz
1
1
Mechatronics & Robotics Engineering, Mechanical Engineering Department, The British University in Egypt, Cairo 11837, Egypt
2
Computer and Systems Department, Electronics Research Institute (ERI), Cairo 12622, Egypt
3
Informatics Department, Electronics Research Institute (ERI), Cairo 12622, Egypt
4
AI Lab, DeltaX Co., Ltd., 5F, 590 Gyeongin-ro, Guro-gu, Seoul 08213, Republic of Korea
*
Authors to whom correspondence should be addressed.
Computers 2026, 15(10), 670; https://doi.org/10.3390/computers15100670
Submission received: 18 August 2026 / Revised: 15 September 2026 / Accepted: 21 September 2026 / Published: 1 October 2026
(This article belongs to the Section AI-Driven Innovations)

Abstract

Surface electromyography (sEMG) gesture classifiers consistently achieve high accuracy under random-split validation, yet performance degrades substantially under participant-independent evaluation, and classification accuracy alone does not establish effective robotic control. This study combines offline gesture recognition with confidence-gated finite-state supervision, minimum-jerk trajectory generation, and gesture-scheduled linear quadratic regulator (LQR) control of a simulated six-degrees-of-freedom manipulator. Four classifiers were evaluated: a backpropagation neural network (BPNN), generalized matrix learning vector quantization (GMLVQ), a hybrid BPNN–GMLVQ, and EMGTFNet, a fuzzy vision transformer-based network for sEMG. Each was benchmarked under 40-participant leave-one-participant-out (LOPO) cross-validation using a 116-dimensional feature vector that combines time-domain, spectral, and cross-correlation descriptors across eight upper-limb gesture classes. The confidence-gated classifier was a Mealy-type finite-state machine with model-specific thresholds and a dwell condition requiring three consecutive qualifying predictions. Each confirmed command was then mapped to a minimum-jerk joint-space trajectory, tracked by a gesture-indexed, kinematically informed LQR. Random-split accuracies exceeded 97% for all four classifiers, whereas the mean LOPO accuracies reached 85.38% for the hybrid model and 85.31% for EMGTFNet. Repeated-measures ANOVA and Friedman tests indicate significant differences among models, although these omnibus results do not establish individual pairwise superiority. Confidence gating reduced spurious state transitions from approximately 97% to zero, and gesture-indexed gain scheduling reduced the joint-space tracking error by 20–33% relative to a fixed global gain. Numerical transcription of the displayed simulation protocols yields 36 of 45 correct scheduled states for EMGTFNet and 74 of 90 for the hybrid model; these descriptive mapping outcomes are distinct from the classifier accuracy and transition success. Confidence thresholds were selected on LOPO outputs rather than on an independent inner partition, and end-to-end latency was not measured. The framework therefore provides a participant-independent classification benchmark and a simulation study of confidence-aware supervisory control, rather than evidence of physical deployment or real-time operation.

1. Introduction

sEMG records the electrical activity of skeletal muscles at the point where motor intention is converted into mechanical action, providing a natural, non-invasive interface for prosthetic, rehabilitation, assistive, and human–robot interaction (HRI) applications [1,2]. In myoelectric (EMG-based) control, these signals are decoded “typically into discrete gesture classes or continuous motion variables” and mapped to commands that drive a prosthesis or robotic manipulator [3,4]. Because sEMG is acquired close to the neuromuscular source, it is inherently stochastic and user-dependent and is further corrupted by electrode displacement, muscle fatigue, and variations in the execution environment. Reliable myoelectric control therefore cannot be achieved through classification accuracy alone; signal variability, decision uncertainty, and the coupled dynamics of the human, robot, and environment must be treated as explicit design constraints for any deployable system.
High offline classification accuracy does not, by itself, translate into usable robotic control. A classifier assigns a label to each isolated signal window, whereas a deployable myoelectric system must convert a continuous and often fluctuating stream of predictions into motion that is safe, stable, and spatially bounded [3,5]. Individual misclassifications and rapid label switching produce erratic command transitions that a raw recognition layer cannot suppress, so a supervisory stage is required to coordinate control modes and reject spurious commands [6]. Finite-state machines (FSMs) have long served this supervisory role in prosthetic and myoelectric devices, reducing high-level control complexity and rendering system behavior more predictable; the classifier output is used to trigger transitions between controller states, as proposed for EMG-based shared control of a manipulator arm in [7]. Studies on real-time prosthetic control have further shown that post-classification filtering and confidence-aware strategies markedly improve command quality [6]. Recognition and control are thus inseparable in practice, yet the classifier, the supervisory layer, and the low-level controller are typically designed in isolation, leaving the confidence of the recognition stage decoupled from the transitions of the control stage.
The accuracy obtained under random-split validation, where windows from the same participant may appear in both the training and test sets, is systematically optimistic. Previous studies have shown that the accuracy drops substantially under cross-user and subject-independent evaluation, both offline and online, when a model is tested on entirely unseen participants [8,9]. Leave-one-participant-out (LOPO) cross-validation withholds one participant per outer fold and assesses generalization to that participant, representing a more deployment-relevant evaluation than random splitting; it does not by itself establish cross-session robustness or hardware readiness. Every learned preprocessing operation, model selection step, and confidence threshold choice must also exclude the outer test participant. The multi-channel dataset adopted in this study supports such participant-wise evaluation [10].
A second limitation concerns control design. Optimal and task-space control of serial-link manipulators is well established: the performance index of a multivariable system shapes both the resulting feedback law and the manipulator’s task-space behavior, and Riccati-based formulations such as the linear quadratic regulator (LQR) are routinely used to synthesize such controllers [11,12,13]. These classical results have rarely been coupled with EMG-based gesture recognition; consequently, most EMG-driven robotic systems adopt a fixed, decoupled, linear controller and report a single error in a single space. More broadly, existing studies optimize either the perception layer (gesture classification) or the control layer in isolation, and few investigate the complete perception-to-control pipeline under participant-independent evaluation. Two questions therefore remain open: whether a gesture-conditioned controller minimizing a joint-space cost can be shaped to serve a task-space objective, and whether confidence-aware supervisory control can measurably improve end-to-end performance.
This study examines an integration framework that involves four sEMG classifiers and a confidence-gated FSM, minimum-jerk trajectory generator and gesture-scheduled LQR controller of a virtual 6-DOF manipulator. Four classifiers of increasing representational capacity, namely, BPNN, GMLVQ, hybrid architecture combining both BPNN and GMLVQ classifiers and EMGTFNet, have the same 116-dimensional feature front-end and are tested using 40-participant LOPO cross-validation for eight upper-limb gestures, although the last classifier deals with a sequence of feature vectors instead of one. For each classifier, the confidence-gated FSM uses model-dependent confidence thresholds and three-prediction dwell requirement to discard false transitions; every recognized command is turned into a minimum-jerk joint-space trajectory followed by gesture-scheduled, kinematics-aware LQR where the state cost matrix is constructed using the manipulator position Jacobian, and thus, the joint-space optimal controller is responsible for task-space movement. The tracking quality is measured using joint-space error (RMSE, degrees) and end-effector error (RMSE, mm). The classification results are averaged over 40 LOPO cross-validation folds, while control examples involve selected recorded commands. The contribution is the examination of these components together, without an absolute priority claim or a claim that the deployment gap has been closed.
This study has the following components:
  • A 40-participant comparison of four classifiers using random-split evaluation and LOPO cross-validation.
  • A confidence-gated FSM with model-specific thresholds and a three-prediction dwell rule, evaluated in an offline mapping protocol.
  • Minimum-jerk trajectory generation with zero-velocity boundary conditions between confirmed gesture states.
  • A gesture-scheduled LQR formulation whose state cost incorporates the manipulator position Jacobian.
  • A descriptive separation of raw recognition performance, scheduled-state correctness, and simulated tracking objectives.
This work extends the recognition focus of [14] by investigating supervisory logic and simulated manipulator control. No hardware validation is claimed.

2. Related Work

The proposed framework draws on three research themes, reviewed in turn below. The first concerns the acquisition, processing, and feature extraction of sEMG signals for robotic control, together with the pronounced inter-subject variability that shapes the required evaluation methodology. The second surveys the classification architectures compared in this study: a BPNN, a GMLVQ network, their hybrid, and an uncertainty-aware transformer. The third examines sEMG-driven manipulator control. Existing studies have already investigated confidence filtering, shared control, continuous estimation, and adaptive assistance; to the best of our knowledge, no existing approach combines selected elements of these within a single simulation framework, which is the gap addressed in this paper.

2.1. Surface Electromyography Signal

Surface electromyography captures the electrical activity of contracting muscles but is intrinsically noisy, non-stationary, and highly variable across users, which places strong demands on the recognition pipeline (Figure 1). Previous research converges on a common four-stage chain: (1) signal acquisition; (2) preprocessing, filtering, and segmentation; (3) feature extraction; and (4) classification and validation. Among these, feature extraction exerts a decisive influence on accuracy, since the choice of descriptors determines how well inter-class differences survive inter-subject and session variability [15].

2.2. Backpropagation Neural Network (BPNN)

Backpropagation neural networks (BPNNs) are feedforward networks trained by gradient-based error minimization and have been applied to sEMG-based systems both as discrete gesture classifiers and as non-linear regressors mapping signals to continuous biomechanical variables such as joint angles, velocities, and torques. Their classification performance depends strongly on the input features and the number of measurement channels. Using an ANOVA-based selection of feature sets and electrode positions, recognition accuracies of 94.83% and 94.6% were obtained for the best four-channel feature and placement configurations, respectively [17]. Under constrained two-channel acquisition with PSD and PCA features, the accuracy was more modest at 76.2 ± 3.1%, rising to 77.7 ± 3.1% with SVD-based feature compression and decision stabilization [18].
For continuous estimation, BPNN regression is competitive: an autoencoder-enhanced four-layer BPNN estimated upper-limb joint torque with an RMSE of 1.38 and a correlation of 0.94 [19], and a CWT–BPNN pipeline estimated elbow angle and velocity across nine motion–force conditions with RMSEs of 8.78° and 9.59°/s, outperforming polynomial and support-vector regressors [20]. Such continuous estimates can complement a discrete classifier in a hybrid scheme, where the class selects the operating mode and the regressed variable sets its intensity or trajectory.

2.3. Generalized Matrix Learning Vector Quantization (GMLVQ)

GMLVQ classifiers are prototype-based: each class is represented by reference vectors, and a sample is assigned to its nearest prototype, yielding Voronoi decision boundaries. Relevance learning extends this scheme by weighing individual feature dimensions, and its generalized matrix form learns a full relevance matrix that defines a Mahalanobis distance accounting for inter-channel dependencies (a property well suited to multi-electrode EMG recordings) [21,22]. The relevance matrix can be regularized and rank-limited to balance bias and variance and to perform discriminative dimensionality reduction [23,24]. Beyond standard classification, GMLVQ variants have been extended to further learning tasks and to context-, sequence-, and subject-dependent settings, addressing the distributional, sequential, and cross-subject variability characteristic of EMG-based control [25,26,27,28,29].
A further advantage is interpretability: because class assignment is governed by distance to labeled prototypes rather than by an unbounded logit transformation [30], confidence derived from prototype proximity retains a direct geometric meaning throughout the feature space [21], which makes it well suited as a gating signal for supervisory control [31,32,33]. Accordingly, GMLVQ is evaluated here both as a standalone classifier and as the classification stage of the hybrid BPNN–GMLVQ architecture, whose learned relevance matrix exposes the feature and channel contributions to each gesture. Both are assessed under identical LOPO conditions, and the resulting prototype confidence informs the threshold design of the confidence-gated FSM.

2.4. Transformer-Based Architectures for EMG Pattern Recognition

Transformer architectures based on self-attention are well suited to sEMG recognition, since gesture intent is distributed across multi-channel activity spanning hundreds of milliseconds; for such longer sequences, attention-based modeling has been shown to outperform convolutional and recurrent baselines [34,35,36,37,38,39,40]. The most direct formulation tokenizes, “which is chopping the raw sEMG signal into a sequence of small pieces (“tokens”) and turning each piece into a vector that the transformer can read,” the raw signal windows without hand-crafted features [39].
The transformer benchmarked in this study, EMGTFNet, embeds a Fuzzy Neural Block (FNB) (a membership-based uncertainty unit) within a vision transformer backbone, explicitly relaxing the input stability assumption of conventional deep classifiers. An ablation of the FNB confirmed improved robustness to inter-individual differences, electrode displacement, and muscle fatigue, reporting an average test accuracy of 83.57% ± 3.5% across the 49 hand gestures of NinaPro DB2 using a 200 ms window and only 56,793 trainable parameters, a substantially reduced count relative to comparable deep models [40]. Here, EMGTFNet is the highest-capacity of the four classifiers, taking the 116-dimensional feature vector as input to recognize eight gestures under all three protocols; its FNB-derived uncertainty provides the confidence signal used for FSM supervision [40].

2.5. Control Strategy for sEMG-Driven Robotic Manipulation

Three broad approaches to sEMG-driven manipulator control appear in the literature: fixed-gain position controllers, variable impedance/admittance schemes modulated by continuous EMG amplitude, and finite-state pipelines that couple discrete gesture classification with model-based feedback.
The architecture of position servo control systems is generally found after intent recognition: A Deep Q-Network recognizing six intents controlled an end-effector via cascaded PID control and inverse kinematics for both a 3-DOF system and a 6-DOF simulation environment [41], as well as a nine-class classifier controlling a rehabilitation end-effector with mapping every intent to one of the steps in the task space without smoothing trajectories between transitions [42]. Compliant control includes approaches such as variable impedance control and admittance control. Adaptive Cartesian stiffness with the estimation of the interaction force in a 7-DOF collaborative robot [43], concurrent-learning adaptive control in Cartesian space [44], tele-impedance control of the prosthesis hand from the opposite side [45], and variable admittance control in a wrist exoskeleton [46] belong to this group; compliance is defined as a scalar or diagonal factor added to an isotropic base value. Finite-state dispatch has also been demonstrated: a four-class network routing packets through an eight-state FSM resulted in a virtual 3-DOF manipulator moving toward Cartesian setpoints [13], while an SVM classifying based on three contraction classes would switch between admittance parameter sets, which has since been developed further into an LDA-based shared control approach for a real 7-DOF arm [5]. Optimal control was used in isolation: an adaptive LQR controller using improved gray wolf optimization surpassed a PID and a PSO-based PID baseline in controlling a multi-fingered prosthetic hand, where the state-cost matrix was independently derived from the manipulator dynamics without the need for EMG classification or gesture-specific gain switching [47]. Other related approaches include synergy-based torque control of a cable-driven rehabilitation arm [48] and per-gesture velocity mapping with high-density EMG from a mobile manipulator [49].
Recent research in adaptive myoelectric control incorporates pre-training, personalization, and self-calibration to accommodate varying user signals [50]. Damping of admittance has been achieved while using the interaction between the wrist and the exoskeleton, along with sEMG-based assistance [46]. Continuous estimation represents another approach path, which can be achieved by using finger motion estimation through musculoskeletal modeling [16], as well as prior research for estimating torque and elbow motion [19,20]. Discrete gestures have been used in this project because of the requirement of the task, which necessitates the selection of eight possible states of the robot. Gestures do not represent an anatomical reconstruction of the user’s limb; a possible future extension may use the discrete modes of operation with continuous regression in each mode.

3. Materials and Methods

The framework comprises sEMG classification, confidence-gated supervision, minimum-jerk trajectory generation, and gesture-scheduled LQR control. The four models are BPNN, GMLVQ, Hybrid BPNN–GMLVQ, and EMGTFNet. Figure 2 summarizes the intended pipeline. Model-specific temporal inputs and unresolved implementation differences are stated below, so the comparison is not interpreted as a controlled architecture-only ablation.

3.1. Dataset and Signal Acquisition

This study uses the public source multi-channel sEMG dataset of [10]. The source contains 40 participants (20 females, 20 males; The average (mean) age for all participants was 22.63 ± 2.1), ten gesture types, five repetitions, and recordings sampled at 2 kHz. Eight gesture classes are retained here: REST, EXTENSION, FLEXION, ULNAR, RADIAL, GRIP, SUPINATION, and PRONATION; finger abduction and adduction are excluded. The source supplies raw and filtered signals. Its acquisition filtering uses a sixth-order 5–500 Hz Butterworth bandpass and a 50 Hz notch; this source-level description is distinct from the additional implementation settings listed in Table 1.
Figure 3 and Table 2 define the illustrative target configurations for the virtual manipulator. Extension/flexion selects opposing vertical movements, ulnar/radial deviation selects lateral movements, and supination/pronation selects rotational configurations. These are task-level associations, not measured human-to-robot joint correspondences. Grip selects a reach/grasp-labeled pose; successful physical grasping is not demonstrated. The REST target is a stored neutral configuration and must not be interpreted as a validated emergency stop. Other tasks require remapping with joint-limit, collision, and reachability checks.

3.2. Feature Extraction

Feature extraction combines the within-channel amplitude and the temporal and spectral features together with between-channel co-activation features. Signals are band-pass filtered with a fourth-order Butterworth filter from 20 to 450 Hz, notch filtered at 50 Hz and its harmonics to 300 Hz with Q = 30, median filtered with a kernel size of three and decimated from 2000 Hz to 1000 Hz. Features are computed over 400 ms sliding windows with 75% overlap, giving a 100 ms step and 51 windows per six-second gesture repetition. The transformer pipeline applies notch filtering at 50 and 100 Hz only; all other preprocessing is identical across the four models.
For each window, we computed a 116-dimensional feature vector (Figure 4). In total, 26 features were computed for each of the 4 channels, 18 time-domain features, and 8 spectral features. The time-domain features, such as the mean absolute value (MAV), the root mean square (RMS), the waveform length (WL), the zero-crossing rate (ZCR), the number of slope sign changes (SSC), the variance (VAR), and the integrated EMG (iEMG), as well as the higher-order statistics (skewness, kurtosis, moment features, and quantiles), for all channels were first computed for each channel. Then, the resulting 104 per-channel features were, for each time-window in each recording session, combined with 12 inter-channel cross-correlation coefficients, computed for 6 pairs of channels (C(4,2) = 6 pairs × 2 lags) to capture the co-activation of spatially distributed forearm muscles.
After the preprocessing step, we have preprocessed signals for the four channels of the 8 retained gesture classes depicted in Figure 5. From the activation profiles of the four channels, the anatomical selectivity of the electrode placement used can be observed. While the large amplitude gestures (such as flexion and extension) activate corresponding left and right channels in large amplitude, the forearm rotation gestures activate extensor and flexor radialis channels in an asymmetric way. On the other hand, grip gestures activate all channels in a co-contraction-like manner. This morphological variety cannot be captured by a single feature type. Therefore, multi-domain feature sets are motivated by the variety of the 8 classes.

3.3. Classification Models

3.3.1. Backpropagation Neural Network

BPNN refers to the feature-based residual multilayer-perceptron classifier trained using backpropagation. The described network transforms a 116-dimensional input vector into an eight-class output by projecting it into a transformation layer and applying residual transformations via multiple streams, with layer sizes of 512, 256, 192, and 128, respectively, and no convolutional kernels used at all electrode positions. The training procedure utilizes up to 600 epochs, a batch size of 512, the AdamW optimizer, a learning rate of 10−3, a weight decay of 5 × 10−3, label smoothing of 0.08, dropout, MixUp training, feature masking, stochastic weight averaging starting from the 100th epoch, ten test-time augmentations, and an ensemble of three models. Therefore, the obtained performance characterizes the entire package of training and inference techniques, rather than the unregularized model.

3.3.2. GMLVQ Classifier

The GMLVQ model is implemented as a deep GMLVQ architecture, where a neural encoder first maps the 116-dimensional feature vector into a 128-dimensional latent space, after which classification is performed through class prototypes and a learnable Mahalanobis-style metric matrix. The encoder uses three dense stages (512, 256, 128) with batch normalization, GELU activation, and progressively decreasing dropout. The prototype module uses five prototypes per class, making the decision process explicitly geometry-based rather than purely logit-based. Regularization includes encoder dropout, metric regularization on Ω, auxiliary cross-entropy, SWA, early stopping, and TTA × 5. The main training setup uses a batch size 512, 600 epochs, Ƞencoder = 1 × 10−3, Ƞprototypes = 1 × 10−3, Ƞomega = 1 × 10−4, weight decay = 5 × 10−3, and label smoothing = 0.05. This model is particularly valuable because it combines deep representation learning with interpretable prototype-based classification.

3.3.3. Hybrid Model

The hybrid model combines a BPNN feature encoder with a GMLVQ prototype classifier, forming a two-stage learned architecture. In the first stage, the encoder applies feature-group attention over the four channels and projects the 116-dimensional input to 512 dimensions. Two parallel residual streams then process this representation, of widths 256 and 192, respectively, and their outputs are concatenated to 448 dimensions and fused to a 96-dimensional latent representation. Batch normalization and dropout are applied throughout. In the second stage, the GMLVQ classifier assigns the latent representation to the nearest prototype under a learned 96 × 96 Mahalanobis metric, with five prototypes per class giving 40 prototypes in total. The architecture is shown in Figure 6.

3.3.4. EMGTFNet

EMGTFNet implements a Fuzzy Neural Block (FNB) inside the transformer structure, which is presented in Figure 7. The FNB block utilizes learnable Gaussian membership functions for mapping the input features into fuzzy features that explicitly represent uncertainty as compared to deterministic classifiers that consider the input as deterministic. EMGTFNet accepts a sequence of 51 feature vectors for each gesture repetition, with each vector having 116 dimensions, projecting these features into a model dimension of 192. The sequence is further transformed into a fixed-size contextual embedding through four encoder blocks of the transformer with six-head self-attention and the feedforward layer with 384 dimensions while applying a dropout of 0.30. Classification head maps the output embedding to the eight gestures. In the current framework, the output of the membership functions in the FNB serves as the confidence score for gating the finite-state supervisor.

3.4. Evaluation Protocol

Two principal evaluation summaries have been conducted: the 80/20 random split and the 40-fold leave-one-person-out (LOPO). In LOPO, the test fold includes one person each time, whereas the other 39 people are used as the training set. One holdout person evaluation is just one case of LOPO evaluation rather than an additional validation method. Random split can include windows of the same person in both the training and test sets; hence, it is dataset-independent rather than person-independent.
For leakage-free evaluation, each channel’s or feature’s mean and standard deviation must be fitted only on the training portion of each fold, with z = ( x − μ train ) σ train then applied unchanged to the validation and test data, together with a fixed rule for zero-variance features. Two forms of normalization are applied here and should be distinguished. Window-based z-scoring is performed along the time axis within one gesture sample segment and does not convey information from one sample to another. The dataset-level z-score standardization is performed using statistics that are computed from the whole dataset and are fitted on the training dataset and are then used for the held-out participant without changes, with the fitted parameters being saved per fold. This was explicitly checked in the code for all four models. The model choice, the choice of augmentations, and the confidence thresholds must also be found on the training dataset using an inner validation set; thresholding was performed using LOPO outputs instead, as explained in Section 3.5.
Tables 6 and 7 report a repeated-measures ANOVA and a Friedman test across the four classifiers, with participants as the repeated unit. These are omnibus tests, assessed at α = 0.05. A significant result establishes that at least one model differs from the others; it does not identify which pairs differ, and no post hoc pairwise claim is made. Repeated-measures assumptions and the correlation induced by the overlapping LOPO training set further limit inference.

3.5. Confidence-Gated Finite-State Machine

The intended eight-state FSM uses the predicted class p ^ and a model-specific confidence score c. The original description uses maximum SoftMax probability for BPNN, Hybrid BPNN–GMLVQ, and EMGTFNet and the normalized inverse prototype distance for GMLVQ. These scores are not automatically calibrated or comparable across models. A candidate different from the current state is committed only when the same candidate is predicted above threshold in three consecutive updates:
p ^ k =   S cand ≠   S curr   and   c k ≥ θ * , for three consecutive updates
Either a different candidate or a below-threshold prediction terminates the qualifying streaks; the current state is maintained until the dwell criteria are met. With an update rate of 100 ms, there are 200 ms between the first and third qualifying predictions. The sum of three update cycles yields a nominal duration of 300 ms, but neither of these durations takes into account the feature acquisition time. Failure to meet the qualification criteria will add to the latency of the command signal. Furthermore, the maintained states may be inaccurate for the intended gesture.
The originally reported thresholds were 0.70 for BPNN, Hybrid BPNN-GMLVQ, and EMGTFNet and 0.65 for GMLVQ. They were chosen by means of a grid search using LOPO output values; no inner tuning set was specified. This means that the results for gating are exploratory in nature and possibly optimistically biased. A threshold sensitivity analysis across all 40 LOPO folds is provided in Section 4.6. A separate validation procedure would entail setting the thresholds within the training group and fixing them prior to testing the test participant; however, this was not the case; the reported thresholds therefore remain exploratory selections rather than validated optima.
Figure 8 depicts symbolic codes assigned to classified gestures and FSM states. These digital labels are command encodings, not literal binary physiological measurements from the four continuous sEMG channels. Figure 9 depicts the Confidence-gated eight-state Mealy-type FSM transition graph.

3.6. Classifier-to-Robot Mapping Protocol

The control study uses software-generated command sequences assembled from prerecorded classifier outputs, not live EMG acquisition or physical robot feedback. Each demonstration follows REST → EXTENSION → FLEXION → ULNAR → RADIAL → GRIP → SUPINATION → PRONATION → REST. The original description states that LOPO predictions were sampled for the mapping tests but does not identify the selected participant/trial files, random seed, or temporal assembly rule. These demonstrations are therefore not treated as 40 independent end-to-end LOPO tests. Five mapping tests are displayed for EMGTFNet and ten for the hybrid. Their scheduled-state matches can be transcribed, but transition-level false-positive/false-negative counts require the full time-resolved logs.

3.7. Virtual 6-DOF Robotic Manipulator

This six-degrees-of-freedom serial manipulator adopted as the virtual arm is parameterized according to the (DH) convention. The homogeneous transformation matrix between consecutive frames i − 1 and i is:
T i − 1 i q i = c θ i − s θ i cα i s θ i sα i a i c θ i s θ i c θ i cα i − c θ i sα i a i s θ i 0 sα i cα i d i 0 0 0 1
where c(·) and s(·) denote cosine and sine, respectively, θi = qi + θ0i is the effective joint angle, and {ai, di, αi, θ0i} are the link length, link offset, twist angle, and joint angle offset of joint i. The DH parameters of the arm are listed in the table in Figure 10.

3.7.1. Forward Kinematics

The composite forward kinematic map from the base frame to the end-effector frame is obtained by chaining the six individual link transformations:
T 6 0 q = T 1 q 1 T 2 q 2 T 3 q 3 T 4 q 4 T 5 q 5 T 6 q 6
The end-effector Cartesian position vector p = [ p x , p y , p z ] T   and rotation matrix R are extracted from the fourth column and upper-left 3 × 3 block of T 6 0 , respectively:
p e =   T 6 0   0 : 3 ,   3 , R 6 0 = T 6 0 0 : 3 ,   0 : 3
p q = [ T 6 0   q 14 ,   T 6 0   q 24 ,   T 6 0   q 34 ] T
Equation (4) uses zero-based array slices; Equation (5) uses one-based matrix entries. The product of the DH transformations in Equation (3) is the defining forward kinematics expression. The previously printed expanded coordinate formula is not retained because its equivalence to the depicted DH chain has not been verified.

3.7.2. Inverse Kinematics

For an ideal spherical wrist geometry with a terminal offset along the end-effector z-axis, wrist center decoupling gives:
  p wc =   p e −   d 6   R 6 0   z ^
The first joint angle is obtained directly from the wrist center projection onto the base plane:
q1 = atan2(pwc,y, pwc,x)
Joints 2 and 3 are resolved as a planar two-link problem. Defining the auxiliary quantities:
r = p wc , x 2 + p wc , y 2 ,       s = p wc , z − d 1 D = r 2 + s 2 − a 2 2 − a 3 2 2 a 2 a 3
The joint angles are:
q 3 = a t a n 2 ± 1 − D 2 , D , q 2 = a t a n 2 s , r − a t a n 2 α 3 s 3 , α 2 + α 3 c 3
where the ±sign corresponds to elbow up and elbow down configurations, respectively. The wrist joint angles q4, q5, q6 are recovered from the residual rotation matrix:
R 6 3 =   ( R 3 0 ) ⊤   R 6 0
The wrist joint angles are recovered analytically from the elements rij of R 6 3 through the ZYZ Euler decomposition:
q 5 = a t a n 2 r 13 2   + r 23 2   , r 33 , q 4 = a t a n 2 r 23 s 5 , r 13 s 5 , q 6 = a t a n 2 r 32 s 5 , − r 31 s 5
Here, s5 = sin (q5). Equations (7)–(11) are conditional spherical wrist/planar arm relations: reachability requires |D| ≤ 1, singular wrist configurations require separate handling, and joint offsets and DH conventions must be respected. The reported protocol supplies joint-space targets directly, as listed in Table 2.

3.7.3. Velocity Kinematics and Position of Jacobian

The position of the Jacobian Jp ∈ ℝ3×6 maps the joint velocities to the end-effector linear velocity and is computed column-wise from the cross product of each joint’s z-axis with the vector from that joint’s origin to the end-effector:
Jp(:,i) = zi−1 × (pe − pi−1), i = 1,…,6
where zi−1 is the unit z-axis of frame i − 1 extracted from T 0 i − 1 , and pi−1 is the origin of frame i − 1. The Jacobian Jp(q*) evaluated at each gesture-associated target configuration q* is the direct input to the kinematically informed LQR cost matrix described in Section 3.8.2.
The position Jacobian directly relates the joint-space velocity to the end-effector Cartesian linear velocity through:
p ˙ e = J p q q ˙
For small position errors near q*, δp≈Jp(q*)δq. Thus, δ q T   J p T ( q * )   W cart   J p δ q approximates a weighted Cartesian position error cost when placed in the position block of Q. Equation (13) separately describes the velocity map. This local approximation is not a global Cartesian optimality guarantee.

3.7.4. Linearized Joint-Space Dynamic Model

The controller uses a simplified, decoupled inertia-damping model. It omits gravity, Coriolis coupling, actuator saturation, backlash, and contact dynamics; it is a simulation surrogate, not a validated physical model:
M j   q ¨ + B d   q ˙ = τ  
where Mj = diag(I1,…,I6) ∈ ℝ6×6 is the diagonal joint inertia matrix and Bd = diag(B1,…,B6) ∈ ℝ6×6 is the viscous damping matrix, with values:
M j = diag ( 0.0075 , 0.0060 , 0.0050 , 0.0040 , 0.0040 , 0.0030 )   kg · m 2 B d = diag ( 0.80 , 0.60 , 0.50 , 0.40 , 0.40 , 0.30 )   N · m · s / rad
Defining the state vector x = [ q ⊤ q ˙ ⊤ ] ⊤ ∈ R 12 , Equation (14) is cast into the linear time-invariant state-space form:
x ˙ = A   x + B   τ
with:
A = 0 I 6 0 − M j − 1 B d ∈ R 12 × 12 ,   B = 0 M j − 1 ∈ R 12 × 6

3.8. Control Architecture

The control architecture (Figure 11) comprises three co-designed components: a minimum-jerk trajectory generator, a gesture-scheduled kinematically informed LQR, and a feedforward inertia compensation term. The three components are coupled by design: the trajectory generator supplies the reference profile consumed by the LQR tracking loop, and the LQR cost matrix is shaped by the position Jacobian derived from the kinematics of Section 3.7.

3.8.1. Minimum-Jerk Trajectory Generation

Biological reaching movements are characterized by bell-shaped velocity profiles and zero velocity at movement onset and termination. These properties are reproduced by the fifth-order polynomial minimum-jerk model of Flash and Hogan, in which the normalized displacement profile is:
s p = 10 p 3 − 15 p 4 + 6 p 5 , p = t T f ∈ 0 , 1
The joint-space trajectory generated for each FSM-triggered gesture transition from initial configuration qs to target configuration qt over movement duration Tf = 3.0 s is:
q mj ( t )   =   q s + Δ q   s ( p ) , q ˙ mj = Δ q T f 30 p 2 − 60 p 3 + 30 p 4 , q ¨ mj t = Δ q T f 2 60 p − 180 p 2 + 120 p 3
Here, Δq = qt − qs and Tf = 3.0 s. The acceleration denominator is T f 2 . The reference has zero velocity and acceleration at each endpoint when transitions start and finish at rest. Each trajectory is generated between confirmed gesture states with zero-velocity boundary conditions. Replanning during an unfinished movement is not implemented, and the smoothness of the reference trajectory does not itself bound the resulting joint torques.

3.8.2. Kinematically Informed LQR Design

For each fixed gesture, the standard continuous-time LQR problem is expressed in error coordinates e = x − xref with feedback correction v. Its infinite-horizon regulation cost is:
J = ∫ 0 ∞ ( e T   Q g   e + e T   R   v ) dt
For the error system matrices A and B from Equation (15), Qg is a symmetric 12 × 12 positive-semidefinite state weight, and R is a symmetric 6 × 6 positive-definite input weight. A consistent position/velocity block construction is:
Q g = blockdiag   Q joint , g + J p q g * T   W cart , g   J p q g * ,   Q vel , g
The two off diagonal 6 × 6 blocks are zero. Qjoint,g and Qvel,g are nonnegative diagonal 6 × 6 matrices, and Wcart,g is a nonnegative diagonal 3 × 3 matrix. With joint angles in radians, Cartesian lengths in meters, and torque in N·m, the compatible cost-weight units are rad−2, m−2, s2∙rad−2, and (N∙m)−2 for Qjoint, Wcart, Qvel, and R, respectively, up to a common cost scale. Alternatively, all states and inputs can first be normalized by explicitly documented reference scales. The original experiment does not supply the per-gesture Q weights or normalization scales; no representative experimental values are invented.
The reported input-weight coefficients are R = diag(1.5, 1.5, 1.5, 2.0, 2.0, 1.0). Their numerical effect depends on the unit/scaling convention. For the standard symmetric cost, the continuous algebraic Riccati equation is:
ATPg + PgA − Pg B R−1BTPg + Qg = 0
yielding the gesture-specific gain matrix:
Kg = R−1BTPg
For each fixed gesture, a stabilizing CARE solution requires the usual stabilizability and detectability conditions. The intended implementation computes a gain bank offline and selects a gain using the FSM state. Negative feedback with the reported inertia-only feedforward is:
τ t = − K g x t − x ref t + M j   q ¨ mj t
Here, x ref = [ q mj q ˙ mj ] . For the damped model in Equation (14), exact model-based feedforward would additionally include q ˙ mj ; inertia-only feedforward leaves a reference-dependent forcing term in the tracking error dynamics. The state weight is constructed with the position block Q joint + J p T   W cart   J p and the velocity block Qvel, with the off-diagonal blocks being zero and the Jacobian evaluated at the target configuration q g * ; the feedback law applies negative feedback; the gain is formed with R−1 from the CARE solution; and the minimum-jerk acceleration uses the Tf2 denominator. Per-gesture weight sets are defined for all eight gestures. Three gesture pairs share identical weight triples: flexion and extension, pronation and supination, and radial and ulnar deviation. The eight gain matrices nonetheless differ, because J p q g * is evaluated at each gesture’s distinct target configuration. The stability of every fixed gain also does not establish stability under arbitrary switching; switching analysis and saturation handling remain necessary.

3.8.3. Dual-Metric Performance Assessment

Controller performance is characterized by two complementary metrics. Joint-space RMSE (degrees) quantifies the precision of joint-level tracking across all six joints over the gesture execution window:
RMSE joint = 1 N · N j ∑ k = 1 N ∑ j = 1 N j ( q j k − q mj , j k ) 2  
Cartesian end-effector RMSE (millimeters) quantifies the task-space accuracy by propagating the joint trajectory through the forward kinematic map:
RMSE cart = 1 N ∑ k = 1 N   p q k − p q mj k 2
Equations (23) and (24) define joint-space and Cartesian tracking errors, which quantify trajectory-following performance and are distinct from the gesture recognition and command acceptance results reported in Section 4.

4. Results

4.1. Classification Performance Under Random-Split Validation

Table 3 reports random-split accuracies of 98.50% for BPNN, 97.65% for Hybrid BPNN–GMLVQ, 97.50% for EMGTFNet, and 97.42% for GMLVQ. These summarize within-dataset performance and should not be interpreted as participant-independent or as prospective control performance.
Figure 12 summarizes the test metrics, Figure 13 displays the validation–accuracy curves, and Figure 14 displays the loss curves. These summaries describe the convergence behavior; class-level performance is reported separately in the per-class metrics and confusion matrices.

4.2. Cross-Subject Generalization: Leave-One-Participant-Out Evaluation

To assess generalization beyond the training cohort, a 40-fold LOPO protocol was applied in which each fold withheld one subject entirely for testing while the remaining 39 subjects formed the training set. The LOPO ranking inverts the random-split order: hybrid BPNN-GMLVQ and EMGTFNet lead at 85.38% and 85.31% respectively, while BPNN first under random-split drops to 77.56%. The extended cross-metric summary is provided in Table 4.
The dataset is perfectly balanced between gesture classes. All 40 subjects performed five iterations of each of the eight gestures, generating 200 data points for each class and a total test set of 1600 data points for 40 folds. The test set for each fold has an equal class distribution of five data points for each class, and hence, the macro- and weighted average measures are numerically equal for all four classifiers.
The per-class precision, recall and F1-score for each classifier, aggregated over 40 folds, are reported in Table 5, and the corresponding confusion matrices are provided in Figure 15.
The per-class accuracy shows variations in all four classifiers with the same ordering of difficulty. The two classes that are recognized with the highest precision and recall are EXTENSION and FLEXION, with F1-scores of 0.880 to 0.949. On the other hand, the class with the lowest accuracy in all four classifiers is PRONATION with scores ranging from 0.607 in the case of BPNN to 0.808 in the case of EMGTFNet. Figure 15 depicts the confusion matrices for the four models. It is evident that the confusion is localized and not scattered. PRONATION is usually confused with RADIAL, REST, and SUPINATION, while ULNAR is usually mistaken with EXTENSION. The confusions involve gestures performed using similar musculature in the forearm and are consistent in all four models.

4.3. Statistical Validation

One-way repeated-measures ANOVA and the Friedman non-parametric test were applied across the 40 LOPO folds at α = 0.05 to determine whether the observed performance differences among classifiers are statistically significant. The results are presented in Table 6 and Table 7.
The reported omnibus results are significant at α = 0.05. They support an overall model effect, not a confirmed pairwise ranking or an architecture-only explanation. The Friedman test for precision gives p = 5.75 × 10−3; all four metrics reject the null hypothesis at α = 0.05.

4.4. Computational Complexity and Real-Time Feasibility

The computational cost for the four classifiers is shown in Table 8 in an NVIDIA GeForce RTX 4060 GPU with batch size 1. GMLVQ is the cheapest computation with 0.26 M parameters, 0.48 MFLOPs, and 0.99 ms needed for inference, and both the BPNN and hybrid models are still light with 1.0 to 1.4 M parameters and less than 3 MFLOPs. The EMGTFNet model is also similar in terms of the number of parameters (1.35 M), but its computational complexity is much higher at 134.73 MFLOPs and 8.99 ms per inference, due to using the self-attention mechanism on a feature sequence of 51 tokens instead of a single feature vector. Therefore, the difference in the inference part is that the first three classifiers use inference on a single 400 ms window, and the transformer uses a whole sequence for the gesture repetitions. These figures show only the computation costs for inference. Acquisition, filtering, feature extraction, communication and actuation delays were not measured; thus, no latency was determined by those numbers.

4.5. Gesture-Driven Robotic Control Performance

The displayed software mapping protocol contains five EMGTFNet tests and ten hybrid tests. Table 8 and Table 9 transcribe and summarize the EMGTFNet and Hybrid scheduled-state mapping results. Table 10. summarize the descriptive scheduled-state outcomes from the displayed mapping tests.
The trajectory structure common to both protocols is illustrated in Figure 16. The yellow marker designates the initial resting configuration; sequential green markers represent gesture-confirmed joint-space waypoints, each activated upon a sustained classification confidence above θ* for a minimum dwell of three consecutive inference windows; and the blue marker identifies the terminal target pose. This event-driven trajectory description is referenced throughout the mapping results that follow.
EMGTFNet mapping protocol. Figure 17 shows the time-per-step plots; Figure 18 lists the predicted states and step times across five demonstrations. These report scheduled state matches rather than per-joint RMSEs or settling-time estimates.
Hybrid mapping protocol. Figure 19 shows the time-per-step plots, and Figure 20 lists outcomes across ten demonstrations. A larger number of selected examples does not establish stronger statistical evidence without independent sampling and comparable protocols. Figure 20 also displays confidence values below the stated 0.70 threshold for some correct states.
The two hybrid discrepancies are Test 1 (57 versus 58 s) and Test 4 (60 versus 58 s).

4.6. Confidence Threshold Sensitivity

The acceptance thresholds reported in Section 3.5 were determined by a grid search on LOPO outputs and not on an unseen inner tuning set. As is typical, sensitivity analysis will be conducted for the selected confidence threshold by evaluating the gated supervisor for each of the 40 LOPO partitions and all four classifiers with an acceptance threshold swept between 0.40 and 0.95 with a step size of 0.01. As reported, the confidence score is the maximum SoftMax probability for the three neural-network-based classifiers and the normalized inverse prototype distance for GMLVQ, although the latter is not calibrated and need not be expected to be treated like an assurance score.
Four quantities were computed at each threshold, and for each of the 40 folds, the following metrics were computed at each of the 40 values of the acceptance threshold: accepted command coverage; the proportion of samples that reach the confidence threshold; accuracy on accepted commands; the error rate of accepted-but-incorrect commands; and the rate of missed commands (i.e., samples that did not reach the confidence threshold). Table 11 shows the results at the operating points of the respective models, and Figure 21 depicts the coverage and accuracy across the full range of the acceptance threshold.
Coverage decreases monotonically with threshold in all four models, and accuracy on accepted commands rises correspondingly. The accuracy curves are smooth: no reversal exceeds 0.15 percentage points across any 0.01 step, and none of the four shows discontinuity about its reported operating point.
The four models vary widely in how much coverage they relinquish as the threshold increases, as can be seen from Table 12. EMGTFNet is the least responsive model, which retains 55.6% coverage even at a threshold of 0.95, while BPNN and the hybrid drop to 11.6% and 9.2% coverage, respectively; at this threshold, two out of forty folds accept no command from each of these two models, so the supervisor will remain where it currently is during the entire recording of these subjects. GMLVQ is somewhat midway between these extremes, retaining 42.4% coverage. The hybrid model gets the highest accuracy among accepted commands and the lowest accepted-but-wrong proportion at the indicated threshold, but only at the cost of having the lowest coverage among the four; EMGTFNet provides the exact opposite compromise, accepting the largest proportion of commands at similar accuracy as BPNN and GMLVQ.
These two limits to interpretation must be acknowledged. First, these measures apply to the sample of individual per-gesture classification attempts, rather than a time-varying supervision signal: the accepted-but-wrong rate corresponds to the antecedent of an erroneous state transition at the level of classification, while the realized transition count depends on the criterion of a three-prediction dwell period as well as on the rule of continuing with the old state when the criteria are not met. Second, the sweep shows that the chosen threshold was not unstable in the sense of being very sensitive to minor variations around its nominal value, although it does not validate the threshold as such, since that would require picking it from the LOPO participants’ data only.

5. Discussion

5.1. Comparative Classification Performance

The near-ceiling accuracies obtained from random-split validation ranging from 97.42% to 98.50% imply that all four models can differentiate between the eight gestures based on the 116-dimensional feature vector given that participants’ identities are available across partitions. Both approaches answer different questions, and LOPO validation is the right approach for use with an unknown participant. In LOPO validation, BPNN decreases from 98.50% to 77.56%, which is a difference of 20.94 percentage points, which is the highest drop in the four models. This means that BPNN is sensitive to participant partition but cannot point to a particular cause since the pre-processing, length of input, and training and inference approaches are not the same for all models.
Both hybrid and EMGTFNet give nearly the same LOPO mean accuracy—85.38% and 85.31%, respectively. The dispersion between the two is different, with the standard deviation of the F1-score being 9.40% and 11.17% and the standard deviation of precision being 9.43% and 10.20%, which, in both cases, are larger for EMGTFNet. This indicates greater inconsistency in the results at the individual level for the transformer, which is important when assistive configurations need to be set for a single individual and not a group. Table 13 describe the contextual comparison with previous studies.
The hybrid achieves this at a substantially lower arithmetic cost, 1.94 MFLOPs against 134.73, reflecting self-attention applied across a 51-window sequence rather than a single feature vector. One interpretation is that the 116-dimensional representation already encodes the discriminative time-domain, spectral and cross-channel structure in a compact form, leaving limited additional information for long-range attention to recover; this is not tested here. The two models also differ in duration, ensembling and augmentation, as well as in architecture, so the cost difference cannot be attributed to any single design component, and no ablation isolating the contribution of fuzzy attention, the prototype head, or individual regularizes is reported.

5.2. Statistical Significance

The results of the ANOVA and Friedman tests show the presence of a difference between the four classifiers at the 0.05 significance level, which is consistent for all measures—accuracy, macro-F1, precision and recall. This type of omnibus test demonstrates that there is a difference between at least one of the models and the rest; the specific models that differ are not identified here. The top two models, hybrid BPNN-GMLVQ and EMGTFNet, differ by 0.07 percentage points of the mean accuracy score, which is smaller than the intra-subject variability for both. Thus, their performance can be classified as equal rather than ranked. Additional constraints on inference come from the correlation created between folds by the overlap in their training samples. Each pair of LOPO folds has 38 common training subjects out of 39.

5.3. Control System Effectiveness

Confidence gating reduces rapid switching by retaining the current state when a candidate prediction fails to meet the threshold or the dwell condition, at the cost of delaying legitimate transitions. The two mechanisms act in opposite directions, and the balance between them is set by the threshold and the dwell length rather than by the classifier alone.
The scheduled-state results are 36 of 45 correct for EMGTFNet and 74 of 90 correct for the hybrid, which is 80.00% and 82.22%. Not counting the first rest state, which does not involve any transition, we have 31 of 40 and 64 of 80, which is 77.50% and 80.00%. This percentage is the proportion of correctness for the state attained at each step and is different from the transition accuracy: it is possible to reach a correct state but without going through any transition, and vice versa. Profiles of the time per step are shown in Figure 17 and Figure 19, while the scheduled-state results are presented in Figure 18 and Figure 20. The window-level post-gate accuracy and false-transition percentages would need time-series annotations and predictions.
The gesture-scheduled cost determines the state weighting for the controller based on the Jacobian of the manipulator at each target configuration of the gesture to ensure a joint-space optimal controller for the given task-space motion. This explains the design motivation behind the gains for each gesture presented in Section 3.8. The control examples presented in Section 4.5 are simulation results of a rigid body plant without taking into consideration the effect of contact forces, actuator saturation, or the dynamics of a real manipulator.

5.4. Limitations

Several limitations constrain interpretation. First, this study uses prerecorded signals and a simplified virtual plant; it does not include live acquisition, hardware-in-the-loop, physical manipulation, or clinical evaluation. Electrode drift, user adaptation, end-to-end timing jitter, communication losses, torque/velocity saturation, gravity, backlash, contact forces, and emergency-stop behavior are untested. Second, threshold selection on LOPO outputs is not an independent assessment of the supervisor; the sensitivity across folds is now reported (Section 4.6), but the stability under independently chosen thresholds has not been established. Third, inconsistent windowing, filters, model dimensions, transformer depth, and input lengths prevent exact replication. Fourth, complete per-class counts, confusion matrices, fold-level scores, seeds, and controlled augmentation/ensemble ablations are unavailable. Fifth, the mapping examples are small, unequally sized, and not documented as paired participant trials; they do not estimate general population task success. Sixth, the two discrepant timing totals require the original logs. Finally, the healthy participant dataset and eight discrete commands do not establish suitability for patients or continuous daily living tasks. Hardware evaluation should include independently tuned thresholds, causal streaming, joint and torque limits, collision monitoring, and a hardware emergency stop before human use.

6. Conclusions

The current work provides a comparative analysis of the four classifiers (BPNN, GMLVQ, hybrid BPNN-GMLVQ, and EMGTFNet) assessed via a common approach that uses both random-split and 40-fold leave-one-participant-out validation. A comprehensive sEMG-to-robot pipeline is described, which involves the application of those classifiers within a confidence-gated finite-state machine, minimum-jerk trajectory generator, and LQR controller informed by the kinematics of the 6-DOF virtual robot. For LOPO analysis, the hybrid and EMGTFNet provide the best results with 85.38% and 85.31% average accuracy, but there is no evidence about which classifier pairs differ from each other according to the omnibus test. The demonstration of the mapping gives 80.00% correct states of scheduling for EMGTFNet and 82.22% for the hybrid. The results show the ability of offline classification and simulated supervisory control.
Physical deployment on a real 6-DOF manipulator is required to validate the simulated tracking behavior under hardware constraints including actuator backlash, inference latency, and electrode motion artifacts. Extension to adaptive nonlinear control strategies such as sliding-mode or model-predictive control is warranted for high-velocity gesture transitions and variable load conditions beyond the linearized LQR assumptions. Incorporation of pathological EMG profiles associated with stroke and neuromuscular disease into both the training corpus and the evaluation protocol would establish the clinical applicability of the framework, and online subject-adaptive mechanisms within the GMLVQ component merit investigation as a means of mitigating cross-session degradation without full retraining. Independent validation of the confidence thresholds within an inner partition, controlled ablations isolating individual architectural components, and time-resolved logging of gate decisions would further strengthen the evaluation.

Author Contributions

Conceptualization, A.H.A., N.Z., M.S.A. and M.A.; methodology, A.H.A., N.Z., M.S.A. and M.A.; software, A.H.A.; validation, A.H.A., N.Z., M.S.A. and M.A.; and N.Z.; formal analysis, A.H.A., N.Z., M.S.A. and M.A.; investigation, A.H.A., N.Z., M.S.A. and M.A.; data curation, A.H.A., N.Z., M.S.A. and M.A.; writing—original draft preparation, A.H.A., N.Z., M.S.A. and M.A.; writing—review and editing, A.H.A., N.Z., M.S.A. and M.A.; visualization, A.H.A.; supervision, N.Z., M.S.A. and M.A.; project administration, A.H.A., N.Z., M.S.A. and M.A.; All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

No new human participant experiment was conducted in the present study. This statement does not assert that the original data collection was exempt from ethics review.

Informed Consent Statement

No new participant consent was obtained for this secondary analysis; the original collection and consent procedures are described by the dataset investigators [11]. Consent for publication of newly collected participant material is not applicable.

Data Availability Statement

The source sEMG dataset is publicly available from Mendeley Data, version 2, at https://doi.org/10.17632/ckwc76xr2z.2, and is described in [11]. Access to the source dataset does not require a request to the present authors. Aggregate results and displayed simulation outcomes are provided in this article. The article does not provide the complete trained model configurations, fold-level predictions, confidence streams, or time-resolved simulation logs needed to reproduce all analyses.

Conflicts of Interest

Author Mohamed S. Abdallah was employed by the company DeltaX Co., Ltd. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

References

  1. Li, K.; Zhang, J.; Wang, L.; Zhang, M.; Li, J.; Bao, S. A review of the key technologies for sEMG-based human-robot interaction systems. Biomed. Signal Process. Control 2020, 62, 102074. [Google Scholar] [CrossRef] [Scilit]
  2. Li, W.; Shi, P.; Yu, H. Gesture Recognition Using Surface Electromyography and Deep Learning for Prostheses Hand: State-of-the-Art, Challenges, and Future. Front. Neurosci. 2021, 15, 621885. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Xiong, D.; Zhang, D.; Chu, Y.; Zhao, Y.; Zhao, X. Intuitive Human-Robot-Environment Interaction with EMG Signals: A Review. IEEE/CAA J. Autom. Sin. 2024, 11, 1075–1091. [Google Scholar] [CrossRef] [Scilit]
  4. Igual, C.; Pardo, L.A.; Hahne, J.M.; Igual, J. Myoelectric control for upper limb prostheses. Electronics 2019, 8, 1244. [Google Scholar] [CrossRef] [Scilit]
  5. Patriarca, F.; Di, L.P.; Arrichiello, F. EMG-Driven Shared Control Architecture for Human–Robot Co-Manipulation Tasks †. Machines 2025, 13, 669. [Google Scholar] [CrossRef] [Scilit]
  6. Li, X.; Tian, L.; Zheng, Y.; Samuel, O.W.; Fang, P.; Wang, L.; Li, G. A new strategy based on feature filtering technique for improving the real-time control performance of myoelectric prostheses. Biomed. Signal Process. Control 2021, 70, 102969. [Google Scholar] [CrossRef] [Scilit]
  7. Xiong, D.; Fu, X.; Zhang, D.; Chu, Y.; Zhao, Y.; Zhao, X. Robotic telemanipulation with EMG-driven strategy-assisted shared control method. Sci. China Technol. Sci. 2024, 67, 3812–3824. [Google Scholar] [CrossRef] [Scilit]
  8. Tsinganos, P.; Jansen, B.; Cornelis, J.; Skodras, A. Real-Time Analysis of Hand Gesture Recognition with Temporal Convolutional Networks. Sensors 2022, 22, 1694. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Çelik, Y.; Can, U. Surface EMG-Based Hand Gesture Recognition Using a Hybrid Multistream Deep Learning Architecture. Sensors 2026, 26, 2281. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Ozdemir, M.A.; Kisa, D.H.; Guren, O.; Akan, A. Dataset for multi-channel surface electromyography (sEMG) signals of hand gestures. Data Brief 2022, 41, 107921. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Sofyan, A.F.; Susanto, E.; Irsyad, R.N.; Prabaswara, A.S.; Rodiana, I.M. An Application Inverse Kinematic Based LQR Control for a 3 DOF Robot Arm. J. INFOTEL 2025, 17, 191–209. [Google Scholar] [CrossRef] [Scilit]
  12. Roveda, L.; Piga, D. Robust state dependent Riccati equation variable impedance control for robotic force-tracking tasks. Int. J. Intell. Robot. Appl. 2020, 4, 507–519. [Google Scholar] [CrossRef] [Scilit]
  13. Pérez-Reynoso, F.; Farrera, N.; Capetillo, C.; Méndez-Lozano, N.; González-Gutiérrez, C.; López-Neri, E. Pattern Recognition of EMG Signals by Machine Learning for the Control of a Manipulator Robot. Sensors 2022, 22, 3424. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Abdelmoneam, A.H.; Abdelaziz, M.; Zayed, N. Pattern Recognition of sEMG Signals by CNN-BiLSTM Algorithm for Bio-Robotics Applications. In 2025 International Telecommunications Conference (ITC-Egypt); IEEE: Piscataway, NJ, USA, 2025; pp. 483–490. [Google Scholar] [CrossRef] [Scilit]
  15. Moslhi, A.M.; Aly, H.H.; ElMessiery, M. The Impact of Feature Extraction on Classification Accuracy Examined by Employing a Signal Transformer to Classify Hand Gestures Using Surface Electromyography Signals. Sensors 2024, 24, 1259. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Gilstrap, T.A.; Alghamdi, M.M.; Phan, T.; Lee, S.W. EMG-Based Continuous Estimation of Index Finger Movements With Varying Interjoint Coordination Patterns by Modeling Musculoskeletal Dynamics. IEEE Access 2025, 13, 13454–13463. [Google Scholar] [CrossRef] [Scilit]
  17. Wu, C.; Yan, Y.; Cao, Q.; Fei, F.; Yang, D.; Lu, X.; Xu, B.; Zeng, H.; Song, A. sEMG Measurement Position and Feature Optimization Strategy for Gesture Recognition Based on ANOVA and Neural Networks. IEEE Access 2020, 8, 56290–56299. [Google Scholar] [CrossRef] [Scilit]
  18. Montecinos, C.; Espinoza, J.; Zamora Zapata, M.; Meruane, V.; Fernandez, R. Improving Fast EMG Classification for Hand Gesture Recognition: A Comprehensive Analysis of Temporal, Spatial, and Algorithm Configurations for Healthy and Post-Stroke Subjects. Sensors 2025, 25, 6980. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Huang, Y.; Chen, K.; Zhang, X.; Wang, K.; Ota, J. Joint torque estimation for the human arm from sEMG using backpropagation neural networks and autoencoders. Biomed. Signal Process. Control 2020, 62, 102051. [Google Scholar] [CrossRef] [Scilit]
  20. Huang, Y.; Chen, K.; Zhang, X.; Wang, K.; Ota, J. Motion estimation of elbow joint from sEMG using continuous wavelet transform and back propagation neural networks. Biomed. Signal Process. Control 2021, 68, 102657. [Google Scholar] [CrossRef] [Scilit]
  21. Schneider, P.; Biehl, M.; Hammer, B. Adaptive Relevance Matrices in Learning Vector Quantization. Neural Comput. 2009, 21, 3532–3561. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Hammer, B.; Villmann, T. Generalized relevance learning vector quantization. Neural Netw. 2002, 15, 1059–1068. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Schneider, P.; Bunte, K.; Stiekema, H.; Hammer, B.; Villmann, T.; Biehl, M. Regularization in matrix relevance learning. IEEE Trans. Neural Netw. 2010, 21, 831–840. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Tang, F.; Tino, P.; Yu, H. Generalized Learning Vector Quantization With Log-Euclidean Metric Learning on Symmetric Positive-Definite Manifold. IEEE Trans. Cybern. 2023, 53, 5178–5190. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Lövdal, S.; Biehl, M. Iterated Relevance Matrix Analysis (IRMA) for the Identification of Class-Discriminative Subspaces. Available online: http://arxiv.org/abs/2401.12842 (accessed on 23 January 2026).
  26. Tang, F.; Feng, H.; Tino, P.; Si, B.; Ji, D. Probabilistic Learning Vector Quantization on Manifold of Symmetric Positive Definite Matrices. Available online: http://arxiv.org/abs/2102.00667 (accessed on 15 February 2026).
  27. Abdi, L.; Prete, A.; Arlt, W.; Biehl, M. Leveraging ordinal generalized matrix learning vector quantization for improved classification. Neural Comput. Appl. 2026, 38, 174. [Google Scholar] [CrossRef] [Scilit]
  28. de Boer, J.; Dedja, K.; Vens, C. SurvivalLVQ: Interpretable supervised clustering and prediction in survival analysis via Learning Vector Quantization. Pattern Recognit. 2024, 153, 110497. [Google Scholar] [CrossRef] [Scilit]
  29. Ravichandran, J.; Kaden, M.; Villmann, T. Variants of recurrent learning vector quantization. Neurocomputing 2022, 502, 27–36. [Google Scholar] [CrossRef] [Scilit]
  30. Kohonen, T. Learning Vector Quantization. In Self-Organizing Maps; Kohonen, T., Ed.; Springer: Berlin/Heidelberg, Germany, 2001; pp. 245–261. [Google Scholar] [CrossRef] [Scilit]
  31. Van, V.R.; Biehl, M.; Nl, M.B. sklvq: Scikit Learning Vector Quantization. J. Mach. Learn. Res. 2021, 22, 1–6. [Google Scholar]
  32. Engelsberger, A.; Villmann, T. Quantum Computing Approaches for Vector Quantization—Current Perspectives and Developments. Entropy 2023, 25, 540. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Pan, T.; Wang, H.; Si, H.; Li, Y.; Shang, L. Identification of pilots’ fatigue status based on electrocardiogram signals. Sensors 2021, 21, 3003. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Zabihi, S.; Rahimian, E.; Asif, A.; Mohammadi, A. TraHGR: Transformer for Hand Gesture Recognition via Electromyography. IEEE Trans. Neural Syst. Rehabil. Eng. 2023, 31, 4211–4224. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Zhang, W.; Zhao, T.; Zhang, J.; Wang, Y. LST-EMG-Net: Long short-term transformer feature fusion network for sEMG gesture recognition. Front. Neurorobot. 2023, 17, 1127338. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Rodriguez Serrezuela, R.; Zamora, R.S.; Hermosilla, D.M.; Gomez, A.E.R.; Reyes, E.M. Hybrid Convolutional Vision Transformer for Robust Low-Channel sEMG Hand Gesture Recognition: A Comparative Study with CNNs. Biomimetics 2025, 10, 806. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Dere, M.D.; Lee, B. A Novel Approach to Surface EMG-Based Gesture Classification Using a Vision Transformer Integrated With Convolutive Blind Source Separation. IEEE J. BioMed. Health Inf. 2024, 28, 181–192. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Montazerin, M.; Rahimian, E.; Naderkhani, F.; Atashzar, S.F.; Yanushkevich, S.; Mohammadi, A. Transformer-based hand gesture recognition from instantaneous to fused neural decomposition of high-density EMG signals. Sci. Rep. 2023, 13, 11000. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Liu, Y.; Li, X.; Yang, L.; Yu, H. A Transformer-Based Gesture Prediction Model via sEMG Sensor for Human–Robot Interaction. In IEEE Transactions on Instrumentation and Measurement; IEEE: Piscataway, NJ, USA, 2024; Volume 73, pp. 1–15. [Google Scholar] [CrossRef] [Scilit]
  40. Córdova, J.C.; Flores, C.; Andreu-Perez, J. EMGTFNet: Fuzzy Vision Transformer to Decode Upperlimb sEMG Signals for Hand Gestures Recognition. In 2023 IEEE International Conference on Fuzzy Systems (FUZZ); IEEE: Piscataway, NJ, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  41. Cruz, P.J.; Vásconez, J.P.; Romero, R.; Chico, A.; Benalcázar, M.E.; Álvarez, R.; López, L.I.B.; Caraguay, Á.L.V. A Deep Q-Network based hand gesture recognition system for control of robotic platforms. Sci. Rep. 2023, 13, 7956. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Song, T.; Zhang, K.; Yan, Z.; Li, Y.; Guo, S.; Li, X. Research on Upper Limb Motion Intention Classification and Rehabilitation Robot Control Based on sEMG. Sensors 2025, 25, 1057. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  43. Iovene, E.; Monaco, R.; Fu, J.; Costa, F.; Ferrigno, G.; De Momi, E. EMG-Based Variable Impedance Control for Enhanced Haptic Feedback in Real-Time Material Recognition. IEEE Trans. Haptics 2025, 18, 220–231. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Obuz, S.; Tatlicioglu, E.; Zergeroglu, E. Adaptive Cartesian space control of robotic manipulators: A concurrent learning based approach. J. Frankl. Inst. 2024, 361, 106701. [Google Scholar] [CrossRef] [Scilit]
  45. Hocaoglu, E.; Patoglu, V. SEMG-Based Natural Control Interface for a Variable Stiffness Transradial Hand Prosthesis. Front. Neurorobot. 2022, 16, 789341. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Lambelet, C.; Mathis, M.; Siegenthaler, M.; Held, J.P.O.; Woolley, D.; Lambercy, O.; Gassert, R.; Wenderoth, N. Variable admittance control with sEMG-based support for wearable wrist exoskeleton. Front. Neurorobot. 2025, 19, 1562675. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  47. Ahmed, K.; Aly, A.A.; Elhabib, M.O. Design of Adaptive LQR Control Based on Improved Grey Wolf Optimization for Prosthetic Hand. Biomimetics 2025, 10, 423. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  48. Xie, C.; Lyu, Y.; Li, G.; Tong, R.K.-Y.; Xia, H.; Song, R.; Li, Z. A Cable-Driven Upper Limb Rehabilitation Robot With Muscle-Synergy-Based Myoelectric Controller. IEEE Trans. Robot. 2024, 40, 3199–3211. [Google Scholar] [CrossRef] [Scilit]
  49. Yang, J.; Shibata, K.; Weber, D.; Erickson, Z. High-density electromyography for effective gesture-based control of physically assistive mobile manipulators. npj Robot. 2025, 3, 2. [Google Scholar] [CrossRef] [Scilit]
  50. Ma, C.; Jiang, X.; Nazarpour, K. Pre-training, personalization, and self-calibration: All a neural network-based myoelectric decoder needs. Front. Neurorobot. 2025, 19, 1604453. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Experimental setup for sEMG-based hand gesture acquisition. (A,B) Placement of surface EMG electrodes on key forearm and hand muscles involved in finger and wrist movements. (C) Hand instrumentation and marker placement used to ensure consistent posture during data acquisition. (D) Representative hand gestures performed by the subjects, illustrating distinct movement patterns used for EMG-based gesture classification [16].
Figure 1. Experimental setup for sEMG-based hand gesture acquisition. (A,B) Placement of surface EMG electrodes on key forearm and hand muscles involved in finger and wrist movements. (C) Hand instrumentation and marker placement used to ensure consistent posture during data acquisition. (D) Representative hand gestures performed by the subjects, illustrating distinct movement patterns used for EMG-based gesture classification [16].
Computers 15 00670 g001
Figure 2. End-to-end EMG-to-robot control pipeline. Four-channel sEMG signals are filtered, windowed, and mapped to a 116-dimensional feature vector fed to four classifiers. Each classifier output stream passes through a confidence-gated eight-state FSM (θ*, dwell = 3), which commits gesture-state transitions to a minimum-jerk trajectory generator and a gesture-adaptive kinematically informed LQR driving the 6-DOF virtual robotic arm. Performance is assessed under random-split and LOPO protocols using joint-space and Cartesian end-effector RMSE.
Figure 2. End-to-end EMG-to-robot control pipeline. Four-channel sEMG signals are filtered, windowed, and mapped to a 116-dimensional feature vector fed to four classifiers. Each classifier output stream passes through a confidence-gated eight-state FSM (θ*, dwell = 3), which commits gesture-state transitions to a minimum-jerk trajectory generator and a gesture-adaptive kinematically informed LQR driving the 6-DOF virtual robotic arm. Performance is assessed under random-split and LOPO protocols using joint-space and Cartesian end-effector RMSE.
Computers 15 00670 g002
Figure 3. Target configurations of the 6-DOF virtual robotic arm for the eight FSM gesture states: (S1) Rest: neutral hold posture; (S2) Extension: elbow extension with forearm raised; (S3) Flexion: elbow flexion with forearm lowered; (S4) Ulnar Deviation: combined joint displacement producing leftward wrist deviation; (S5) Radial Deviation: combined joint displacement producing rightward wrist deviation; (S6) Grip: coordinated multi-joint reach-and-grasp configuration engaging all six joints including q6 = −28.50°; (S7) Supination: outward forearm rotation; and (S8) Pronation: inward forearm rotation.
Figure 3. Target configurations of the 6-DOF virtual robotic arm for the eight FSM gesture states: (S1) Rest: neutral hold posture; (S2) Extension: elbow extension with forearm raised; (S3) Flexion: elbow flexion with forearm lowered; (S4) Ulnar Deviation: combined joint displacement producing leftward wrist deviation; (S5) Radial Deviation: combined joint displacement producing rightward wrist deviation; (S6) Grip: coordinated multi-joint reach-and-grasp configuration engaging all six joints including q6 = −28.50°; (S7) Supination: outward forearm rotation; and (S8) Pronation: inward forearm rotation.
Computers 15 00670 g003
Figure 4. sEMG signal conditioning and 116-dimensional feature extraction pipeline. Four-channel acquisition at 2000 Hz, fourth-order Butterworth band-pass filtering (20–450 Hz), notch filtering, median filtering, and decimation to 1000 Hz, followed by 400 ms sliding windows with 75% overlap. Twenty-six time-domain and spectral features per channel (104 total) are concatenated with 12 inter-channel cross-correlation features, giving a 116-dimensional vector.
Figure 4. sEMG signal conditioning and 116-dimensional feature extraction pipeline. Four-channel acquisition at 2000 Hz, fourth-order Butterworth band-pass filtering (20–450 Hz), notch filtering, median filtering, and decimation to 1000 Hz, followed by 400 ms sliding windows with 75% overlap. Twenty-six time-domain and spectral features per channel (104 total) are concatenated with 12 inter-channel cross-correlation features, giving a 116-dimensional vector.
Computers 15 00670 g004
Figure 5. Preprocessed four-channel sEMG signal visualization across all eight gesture classes following fourth-order Butterworth band-pass filtering (20–450 Hz) and 50 Hz notch filtering at 2 kHz.
Figure 5. Preprocessed four-channel sEMG signal visualization across all eight gesture classes following fourth-order Butterworth band-pass filtering (20–450 Hz) and 50 Hz notch filtering at 2 kHz.
Computers 15 00670 g005
Figure 6. Architecture of the hybrid BPNN–GMLVQ classifier. A feature-group attention layer and a 116 → 512 projection feed two parallel residual streams of widths 256 and 192, which are concatenated and fused to a 96-dimensional latent representation. A GMLVQ head classifies this representation by nearest-prototype rule under a learned 96 × 96 Mahalanobis metric with five prototypes per class, giving 40 prototypes in total.
Figure 6. Architecture of the hybrid BPNN–GMLVQ classifier. A feature-group attention layer and a 116 → 512 projection feed two parallel residual streams of widths 256 and 192, which are concatenated and fused to a 96-dimensional latent representation. A GMLVQ head classifies this representation by nearest-prototype rule under a learned 96 × 96 Mahalanobis metric with five prototypes per class, giving 40 prototypes in total.
Computers 15 00670 g006
Figure 7. Architecture of EMGTFNet. A Fuzzy Neural Block applies learnable Gaussian membership functions to the 116-dimensional input, which is projected to 192 dimensions and processed by four transformer encoder blocks with six-head self-attention and dropout p = 0.30, before an eight-class softmax head.
Figure 7. Architecture of EMGTFNet. A Fuzzy Neural Block applies learnable Gaussian membership functions to the 116-dimensional input, which is projected to 192 dimensions and processed by four transformer encoder blocks with six-head self-attention and dropout p = 0.30, before an eight-class softmax head.
Computers 15 00670 g007
Figure 8. This table presents the classifier output mapping of the EMG signal to gesture states.
Figure 8. This table presents the classifier output mapping of the EMG signal to gesture states.
Computers 15 00670 g008
Figure 9. Confidence-gated eight-state Mealy-type FSM transition graph. Directed edges represent admissible transitions, each conditioned on the classifier predicting the candidate state with confidence c ≥ θ* sustained for d ≥ 3 consecutive windows; transitions failing either criterion reset the dwell counter without changing state.
Figure 9. Confidence-gated eight-state Mealy-type FSM transition graph. Directed edges represent admissible transitions, each conditioned on the classifier predicting the candidate state with confidence c ≥ θ* sustained for d ≥ 3 consecutive windows; transitions failing either criterion reset the dwell counter without changing state.
Computers 15 00670 g009
Figure 10. Annotated 6-DOF robotic manipulator diagram DH convention parameters, coordinate frames, and joint variables.
Figure 10. Annotated 6-DOF robotic manipulator diagram DH convention parameters, coordinate frames, and joint variables.
Computers 15 00670 g010
Figure 11. Gesture-adaptive control architecture for EMG-driven 6-DOF robotic manipulator.
Figure 11. Gesture-adaptive control architecture for EMG-driven 6-DOF robotic manipulator.
Computers 15 00670 g011
Figure 12. Test accuracy comparison with performance metrics.
Figure 12. Test accuracy comparison with performance metrics.
Computers 15 00670 g012
Figure 13. Validation accuracy comparison: all four classifiers across epochs.
Figure 13. Validation accuracy comparison: all four classifiers across epochs.
Computers 15 00670 g013
Figure 14. Comparative EMG classification loss curves.
Figure 14. Comparative EMG classification loss curves.
Computers 15 00670 g014
Figure 15. Row-normalized confusion matrices aggregated over 40 leave-one-participant-out folds for (a) BPNN, (b) GMLVQ, (c) hybrid BPNN-GMLVQ and (d) EMGTFNet. Rows are the true class, columns the predicted class, and each cell gives the percentage of that true class assigned to the predicted class; rows sum to 100%. Classes are EXT—extension, FLE—flexion, GRI—grip, PRO—pronation, RAD—radial, RES—rest, SUP—supination, ULN—ulnar.
Figure 15. Row-normalized confusion matrices aggregated over 40 leave-one-participant-out folds for (a) BPNN, (b) GMLVQ, (c) hybrid BPNN-GMLVQ and (d) EMGTFNet. Rows are the true class, columns the predicted class, and each cell gives the percentage of that true class assigned to the predicted class; rows sum to 100%. Classes are EXT—extension, FLE—flexion, GRI—grip, PRO—pronation, RAD—radial, RES—rest, SUP—supination, ULN—ulnar.
Computers 15 00670 g015
Figure 16. Trajectory description: yellow (initial) → green (confirmed waypoints) → blue (terminal).
Figure 16. Trajectory description: yellow (initial) → green (confirmed waypoints) → blue (terminal).
Computers 15 00670 g016
Figure 17. EMGTFNet time-per-step profiles for the scheduled gesture protocol. Colored traces show the five individual test runs (Test No. 1–5), the dashed black line shows the mean across all subjects, and the shaded band shows the min–max range at each step. The fraction printed above each trajectory coordinate is the number of subjects whose command was correct out of the total. Gesture step labels along the horizontal axis are color-coded by gesture type. The minimum-jerk generator produces smooth S-curve trajectories between successive gesture targets.
Figure 17. EMGTFNet time-per-step profiles for the scheduled gesture protocol. Colored traces show the five individual test runs (Test No. 1–5), the dashed black line shows the mean across all subjects, and the shaded band shows the min–max range at each step. The fraction printed above each trajectory coordinate is the number of subjects whose command was correct out of the total. Gesture step labels along the horizontal axis are color-coded by gesture type. The minimum-jerk generator produces smooth S-curve trajectories between successive gesture targets.
Computers 15 00670 g017
Figure 18. EMGTFNet scheduled-state mapping results. The “correct/9” column counts matched states, including the initial rest state.
Figure 18. EMGTFNet scheduled-state mapping results. The “correct/9” column counts matched states, including the initial rest state.
Computers 15 00670 g018
Figure 19. Hybrid BPNN–GMLVQ time-per-step profiles for the scheduled gesture protocol. Colored traces show the five individual test runs (Test No. 1–5), the dashed black line shows the mean across all subjects, and the shaded band shows the min–max range at each step. The fraction printed above each trajectory coordinate is the number of subjects whose command was correct out of the total. Gesture step labels along the horizontal axis are color-coded by gesture type. The minimum-jerk generator produces smooth S-curve trajectories between successive gesture targets.
Figure 19. Hybrid BPNN–GMLVQ time-per-step profiles for the scheduled gesture protocol. Colored traces show the five individual test runs (Test No. 1–5), the dashed black line shows the mean across all subjects, and the shaded band shows the min–max range at each step. The fraction printed above each trajectory coordinate is the number of subjects whose command was correct out of the total. Gesture step labels along the horizontal axis are color-coded by gesture type. The minimum-jerk generator produces smooth S-curve trajectories between successive gesture targets.
Computers 15 00670 g019
Figure 20. Original hybrid scheduled-state mapping results. Tests 1 and 4 report total times of 57 and 60 s, while the displayed step times sum to 58 and 58 s. Both are disclosed in Table 8; neither is silently substituted as a verified elapsed time.
Figure 20. Original hybrid scheduled-state mapping results. Tests 1 and 4 report total times of 57 and 60 s, while the displayed step times sum to 58 and 58 s. Both are disclosed in Table 8; neither is silently substituted as a verified elapsed time.
Computers 15 00670 g020
Figure 21. Accepted command coverage and accuracy on accepted commands as a function of the confidence threshold, averaged over 40 LOPO folds, for the four classifiers. Shaded bands show ±1 standard deviation across folds. Filled circles mark each model’s reported operating point, plotted in the model’s own colour and placed at the vertical dashed lines: 0.70 for BPNN, the hybrid BPNN–GMLVQ and EMGTFNet, and 0.65 for GMLVQ.
Figure 21. Accepted command coverage and accuracy on accepted commands as a function of the confidence threshold, averaged over 40 LOPO folds, for the four classifiers. Shaded bands show ±1 standard deviation across folds. Filled circles mark each model’s reported operating point, plotted in the model’s own colour and placed at the vertical dashed lines: 0.70 for BPNN, the hybrid BPNN–GMLVQ and EMGTFNet, and 0.65 for GMLVQ.
Computers 15 00670 g021
Table 1. Comparison of hyperparameters and architectures of the implemented algorithms.
Table 1. Comparison of hyperparameters and architectures of the implemented algorithms.
ModelInput RepresentationPreprocessing/SamplingWindow/Sequence SetupLayer Parameters
BPNN116 handcrafted EMG featuresBandpass 20–450 Hz, multi-notch filtering (50–300 Hz),
median filtering, FIR decimation from 2000 Hz to 1000 Hz, per-channel z-score normalization
400 ms windows, 75% overlap, gesture-consistent segmentation,
116-dimensional feature vector per window
FeatureGroupAttention → projection 116 → 512 → Stream 1 512 → 256
(residual)‖Stream 2 384 → 192 → Fusion 448 → 512 (residual) → Head 128 → Output
GMLVQSameBandpass 20–450 Hz, multi-notch filtering (50–300 Hz),
median filtering, FIR decimation from 2000 Hz to 1000 Hz, per-channel z-score normalization
400 ms windows, 75% overlap,
gesture-consistent segmentation, 116-dimensional feature vector per window
Encoder MLP: 116 → 512 → 256 → 128, then GMLVQ with latent dimension 128, 5 prototypes per class, learnable metric matrix Ω ∈ ℝ^ (128 × 128)
Hybrid BPNN–GMLVQSameSame feature-generation pipeline as BPNN and GMLVQ; bandpass 20–450 Hz, multi-notch filtering, median filtering, 2000 Hz → 1000 Hz down-sampling, z-score normalization400 ms windows, 75% overlap, gesture-consistent segmentation,
116-dimensional feature vector per window
FeatureGroupAttention → projection 116 → 512 → Stream 1 512 → 256 (residual)‖Stream 2 384→192 → Fusion 448 → 192 (residual) → GMLVQ head with latent dimension 192, 8 prototypes per class, auxiliary head 192 → 8
TransformerFeature-sequence
input: 116 handcrafted EMG features across time
Bandpass 20–450 Hz, notch 50/100 Hz, median filtering,
FIR decimation from 2000 Hz
to 1000 Hz, per-channel z-score normalization
Full 6 s gesture treated as one sample; 400 ms feature windows, 100 ms step, trimmed gesture boundaries, approximately 51 time steps × 116 featuresInput 116 → 192 projection → CLS token + positional encoding → 4 fuzzy transformer layers → dual pooling (CLS + average pooling, total 384) → MLP classifier
Table 2. Illustrative gesture-to-configuration assignments for the virtual manipulator; joint angles are in degrees.
Table 2. Illustrative gesture-to-configuration assignments for the virtual manipulator; joint angles are in degrees.
StateGesture ClassDesired Joint Angles in DegreesDesired Movement
q1 (°)q2 (°)q3 (°)q4 (°)q5 (°)q6 (°)
S1Rest+200.12+90.27−88.88−1.6400Hold/Stop
S2Extension+200.12+90.27−88.88+87.5600Move Up
S3Flexion+194.10+90.27−88.26−70.5200Move Down
S4Ulnar Deviation+200.12+90.27−88.22+20.01−68.510Move Left
S5Radial Deviation+192.89+89.67−88.88−32.47−58.710Movie Right
S6Grip+200.12+90.27−45.000−90.00−28.50Grasp/Reach
S7Supination+195.30+90.27−88.88−2.30−62.990Rotate Outward
S8Pronation+195.30+90.27−88.88+2.30−144.290Rotate Inward
Table 3. Classification performance.
Table 3. Classification performance.
ClassifierAccuracy (%)Macro F1-Score (%)Macro Precision (%)
BPNN98.5098.5198.51
Hybrid BPNN–GMLVQ97.6597.6597.65
EMGTFNet97.5097.5097.64
GMLVQ97.4296.8596.86
Table 4. Classification performance under 40-fold LOPO cross-validation (mean ± standard deviation across participants).
Table 4. Classification performance under 40-fold LOPO cross-validation (mean ± standard deviation across participants).
ClassifierAccuracy Mean ± SD (%)F1-Score Mean ± SD (%)Precision Mean ± SD (%)
Hybrid85.38 ± 8.9384.79 ± 9.4086.00 ± 9.43
EMGTFNet85.31 ± 9.7483.71 ± 11.1786.87 ± 10.20
GMLVQ80.69 ± 14.8479.34 ± 15.9583.66 ± 14.19
BPNN77.56 ± 12.8375.38 ± 14.1680.12 ± 13.96
Table 5. Per-class precision, recall and F1-score aggregated over all 40 LOPO folds. Each class contributes 200 samples.
Table 5. Per-class precision, recall and F1-score aggregated over all 40 LOPO folds. Each class contributes 200 samples.
ClassBPNN
P%
BPNN
R%
BPNN
F1%
GMLVQ
P%
GMLVQ
R%
GMLVQ
F1%
Hybrid
P%
Hybrid
R%
Hybrid
F1%
EMGTFNet
P%
EMGTFNet
R%
EMGTFNet
F1%
Extension0.88830.91500.90150.89860.93000.91400.88990.97000.92820.89570.94500.9197
Flexion0.92540.93000.92770.88680.94000.91260.92820.97000.94870.83110.93500.8800
Grip0.73630.74000.73820.85260.81000.83080.89120.86000.87530.90290.79000.8427
Pronation0.66270.56000.60700.67630.70000.68800.75000.67500.71050.83870.78000.8083
Radial0.73580.78000.75730.75000.79500.77180.89580.86000.87760.80650.87500.8393
Rest0.77590.90000.83330.77780.84000.80770.84650.85500.85070.85650.89500.8753
Supination0.72780.65500.68950.84620.77000.80630.80770.84000.82350.86830.89000.8790
Ulnar0.72860.72500.72680.77010.67000.71660.80810.80000.80400.83140.71500.7688
Macro avg0.77260.77560.77270.80730.80690.80600.85220.85380.85230.85390.85310.8516
Table 6. One-way repeated-measures ANOVA results (LOPO evaluation).
Table 6. One-way repeated-measures ANOVA results (LOPO evaluation).
MetricF-Statisticdf (Model, Error)p-ValueDecision (α = 0.05)Statistical Significance
Accuracy13.19(3, 117)1.78 × 10−7Reject H0Highly Significant
F1-score13.85(3, 117)8.62 × 10−8Reject H0Highly Significant
Precision5.95(3, 117)8.20 × 10−4Reject H0Highly Significant
Table 7. Friedman test results (non-parametric validation).
Table 7. Friedman test results (non-parametric validation).
Metricχ2 (Chi-Square)dfp-ValueDecision (α = 0.05)Statistical Significance
Accuracy23.3533.42 × 10−5Reject H0Highly Significant
F1-score22.6734.74 × 10−5Reject H0Highly Significant
Precision12.5435.75 × 10−3Reject H0Very Significant
Table 8. Reported computational measurements, with model-specific input units.
Table 8. Reported computational measurements, with model-specific input units.
ModelParametersFLOPs per Model InputInference TimeInput Unit
BPNN1.39 M2.76 M2.53 msFeature vector
GMLVQ0.26 M0.48 M0.99 msFeature vector
BPNN–GMLVQ1.01 M1.94 M2.52 msFeature vector
EMGTFNet1.35 M134.73 M8.99 ms~51-token sequence
Table 9. Numerical transcription of Figures 18 and 20, reporting total time and the sum of displayed step times separately.
Table 9. Numerical transcription of Figures 18 and 20, reporting total time and the sum of displayed step times separately.
ModelTestCorrect StatesCorrect (%)Reported Total (s)Step Sum (s)
EMGTFNet19/9100.005656
EMGTFNet29/9100.005656
EMGTFNet37/977.786262
EMGTFNet46/966.675757
EMGTFNet55/955.566262
Hybrid17/977.785758
Hybrid28/988.895959
Hybrid38/988.895555
Hybrid47/977.786058
Hybrid57/977.785959
Hybrid68/988.895555
Hybrid77/977.785858
Hybrid86/966.676161
Hybrid99/9100.005858
Hybrid107/977.786464
Table 10. Descriptive scheduled-state outcomes from the displayed mapping tests.
Table 10. Descriptive scheduled-state outcomes from the displayed mapping tests.
ModelTestsAll Scheduled StatesAfter Initial RESTFully Correct Tests
EMGTFNet536/45 (80.00%)31/40 (77.50%)2/5
Hybrid1074/90 (82.22%)64/80 (80.00%)1/10
Table 11. Confidence threshold sensitivity at the reported operating points, mean ± SD across 40 LOPO folds.
Table 11. Confidence threshold sensitivity at the reported operating points, mean ± SD across 40 LOPO folds.
ModelThresholdCoverage (%)Accuracy on Accepted (%)Accepted but Incorrect (%)Missed (%)
BPNN0.7064.94 ± 14.4289.34 ± 11.046.25 ± 5.5235.06 ± 14.42
GMLVQ0.6578.06 ± 12.8589.30 ± 11.587.75 ± 8.2821.94 ± 12.85
BPNN–GMLVQ0.7062.88 ± 9.9697.24 ± 3.811.56 ± 2.0937.12 ± 9.96
EMGTFNet0.7087.81 ± 6.8089.93 ± 8.698.62 ± 7.2512.19 ± 6.80
Table 12. Coverage/accuracy on accepted commands (%), mean across 40 LOPO folds.
Table 12. Coverage/accuracy on accepted commands (%), mean across 40 LOPO folds.
ThresholdBPNNGMLVQHybrid BPNN–GMLVQEMGTFNet
0.4092.2/80.396.7/82.289.1/90.499.3/85.5
0.5082.5/84.491.4/84.579.0/93.897.0/86.4
0.6074.7/86.682.2/88.171.9/95.792.0/88.3
0.6570.2/87.878.1/89.367.6/96.489.7/89.3
0.7064.9/89.373.6/90.462.9/97.287.8/89.9
0.7559.9/91.269.1/91.458.1/98.085.1/90.4
0.8052.5/93.364.1/93.551.5/99.081.1/91.6
0.8545.4/95.058.9/94.642.4/99.075.8/92.7
0.9031.7/96.652.2/95.928.5/99.168.4/94.6
0.9511.6/98.742.4/96.99.2/99.655.6/96.1
Table 13. Contextual comparison with previous studies. Different datasets, participants, gesture vocabularies, input durations, and evaluation protocols preclude direct numerical ranking.
Table 13. Contextual comparison with previous studies. Different datasets, participants, gesture vocabularies, input durations, and evaluation protocols preclude direct numerical ranking.
PaperModelExtraction Technique ParticipantsPurposePerformance
Montecinos et al. [18]ANNPSD + PCA40 (healthy + stroke)Hand gesture recognition (healthy + stroke patients)95.31% (RS); 35–40% (cross-patient)
Rodriguez et al. [36]CViT (CNN + ViT)Convolutional feature maps10/11 (able-bodied/amputees)Low channel sEMG gesture recognition for prosthetic control96.60% (able-bodied); 94.20% (amputees)
Zabihi et al. [34] TraHGR (Temporal + Feature Transformer)Windowed sEMG time-series (200 ms)40 (NinaPro DB2)Large-scale hand gesture recognition via sEMG86.18% (RS)
This studyBPNN/GMLVQ/Hybrid/EMGTFNet + Confidence-Gated FSM + Min-Jerk Trajectory + Kinematic LQRTime-domain + spectral (116-dim) + cross-correlation40Offline classification and simulated 6-DOF control98.50% (RS); 85.38% (LOPO); F1: 84.79% (LOPO)
Declarations: This study is a secondary analysis of the publicly available dataset described in [11] and a software simulation. The authors did not recruit participants or collect new human data for this work.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Abdelmoneam, A.H.; Zayed, N.; Abdallah, M.S.; Abdelaziz, M. sEMG-Driven Robotic Manipulation in Simulation: LOPO Classification, Confidence-Gated Supervision, and Gesture-Scheduled LQR Control. Computers 2026, 15, 670. https://doi.org/10.3390/computers15100670

AMA Style

Abdelmoneam AH, Zayed N, Abdallah MS, Abdelaziz M. sEMG-Driven Robotic Manipulation in Simulation: LOPO Classification, Confidence-Gated Supervision, and Gesture-Scheduled LQR Control. Computers. 2026; 15(10):670. https://doi.org/10.3390/computers15100670

Chicago/Turabian Style

Abdelmoneam, Anas Hassan, Nourhan Zayed, Mohamed S. Abdallah, and Mostafa Abdelaziz. 2026. "sEMG-Driven Robotic Manipulation in Simulation: LOPO Classification, Confidence-Gated Supervision, and Gesture-Scheduled LQR Control" Computers 15, no. 10: 670. https://doi.org/10.3390/computers15100670

APA Style

Abdelmoneam, A. H., Zayed, N., Abdallah, M. S., & Abdelaziz, M. (2026). sEMG-Driven Robotic Manipulation in Simulation: LOPO Classification, Confidence-Gated Supervision, and Gesture-Scheduled LQR Control. Computers, 15(10), 670. https://doi.org/10.3390/computers15100670

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop