Next Article in Journal
Comparative Study on the Surface Properties of Synthetic Carbonated Hydroxyapatite and Natural Hydroxyapatite Before and After Contact with Solutions with de- and Remineralization Activity
Previous Article in Journal
Analysis of Erosive Wear in Pipe Elbows and Biomimetic Protection Strategies
 
 
Article
Peer-Review Record

TransTCNet: Transformer-Based Temporal-Contextual Network for Low-Latency Typing Interfaces on Edge Devices

Biomimetics 2026, 11(5), 337; https://doi.org/10.3390/biomimetics11050337
by Asif Ullah 1,2, Zhendong Song 2,*, Waqar Riaz 3, Yizhi Shao 2 and Xiaozhi Qi 2,*
Reviewer 1:
Reviewer 2: Anonymous
Biomimetics 2026, 11(5), 337; https://doi.org/10.3390/biomimetics11050337
Submission received: 15 April 2026 / Revised: 9 May 2026 / Accepted: 9 May 2026 / Published: 12 May 2026
(This article belongs to the Section Bioinspired Sensorics, Information Processing and Control)

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

Comments to Authors:

  1. Introduction: While the introduction adequately positions the study, a more explicit discussion of challenges faced by character-level sEMG recognition (e.g., temporal variability, muscle signal fusion, sensor shifts) would strengthen motivation. Also, clarifying why transformer architectures are particularly well-suited here (against CNNs or RNNs) could add impact.
  2. Methods: Please consolidate preprocessing details (filter specifications, normalization, window size) and model hyperparameters clearly. Details such as the number of transformer layers, attention heads, dropout rates, and learning rate schedules would greatly enhance reproducibility.
  3. Experimental Setup: It would be instructive to describe data augmentation or regularization methods used (if any). Also, please clarify if the same model weights were used across all participants or individualized.
  4. Results: Statistical tests are used effectively. Consider including per-class confidence intervals or standard deviations for performance metrics, and expand on the implications of misclassified pairs for practical UI design.
  5. Figures/Tables: Increase font size on figures for readability. Captions should include interpretations beyond mere descriptions. Table formatting could be improved for clarity.
  6. Limitations and Future Work: The discussion is thorough and balanced. Strongly recommend expanding on strategies to reduce inter-subject variability and integrating continuous vs. isolated keypress modeling for real-world deployment.
Comments on the Quality of English Language

The English could be improved to more clearly express the research. The manuscript is understandable but contains several awkward phrasings, grammatical inconsistencies, and minor typographical errors. Careful proofreading by a native English speaker or professional language editor could improve clarity and flow. Certain technical descriptions could be rewritten for precision and conciseness.

Below are specific recommendations to improve the English language usage for clarity, conciseness, and readability:

  • Improve sentence conciseness and flow:
    • Example: "Such developments may have serious effects on human-computer interaction (HCI)" → "These developments could significantly impact human-computer interaction (HCI)."
    • Avoid redundancy, e.g., "may have serious effects" and "serious" can often be replaced with "significant."
    • Use active voice where possible to improve readability.
  • Correct minor grammatical issues:
    • Consistently use singular or plural nouns, e.g., "several machine learning and deep learning techniques have been used" (correct) rather than mixing tenses.
    • Recheck subject-verb agreement, especially in complex sentences.
  • Review technical terminology for consistency:
    • "sEMG" is correctly used throughout; ensure it is defined at first use if not already.
    • Terms like "cross-subject," "cross-participant," and "cross-session" should be consistent and hyphenated if used as compound adjectives.
  • Clearer expression of statistical and experimental results:
    • For example, instead of "the standard deviation of participants' individual accuracies (0.13%) is very small" → "participants’ accuracies showed a low standard deviation (0.13%), indicating high consistency."
    • Avoid phrases like "very small" or "considerably better"—specify quantitative comparisons or use "significantly" if supported statistically.
  • Simplify complex sentences:
    • Some sentences are long and packed with information (e.g., in the introduction and discussion sections). Breaking them into two or more sentences would aid clarity.
  • Punctuation and typographical consistency:
    • Use consistent spacing around dashes and commas.
    • Ensure consistent use of percentage formats (e.g., 87.57% not 87.57 %).
  • Reduction of passive voice where clarity suffers:
    • For example: "Such recordings contain 16 channels…" → "The recordings include 16 EMG channels…"
  • Avoid unnecessary jargon or repetition:
    • Where possible, explain complex terms in simpler language for broader accessibility.
  • Check figure and table references:
    • Ensure uniformity in naming figures (e.g., "Figure 6" vs. "Fig. 6") and consistent tense: "Figure 6 compares" instead of "Figure 6 comparison shows."
  • Consistent tense usage:
    • Prefer past tense for describing completed experiments and present tense for general truths or ongoing relevance.

Author Response

Reviewer 1

  1. Introduction: While the introduction adequately positions the study, a more explicit discussion of challenges faced by character-level sEMG recognition (e.g., temporal variability, muscle signal fusion, sensor shifts) would strengthen motivation. Also, clarifying why transformer architectures are particularly well-suited here (against CNNs or RNNs) could add impact.

Response:

We thank the reviewer for this insightful suggestion. We agree that explicitly discussing the unique challenges of character-level sEMG recognition relative to coarse gesture recognition strengthens the study's motivation. We also agree that a clearer justification for using transformer-based architectures helps clarify the rationale behind the proposed TransTCNet design.

Action Taken:

We have revised Section 1, Introduction, by adding a dedicated discussion of the main challenges in character-level sEMG typing recognition, including temporal variability in keystroke execution, muscle signal fusion among biomechanically similar keys, and sensitivity to electrode shifts across sessions. We also added a comparative explanation of why transformer-based attention is well-suited for this task, particularly compared with conventional CNN- and RNN-based approaches.

Changes in the Manuscript:

Unlike coarse gesture recognition, character-level typing faces three distinct challenges: temporal variability in keystroke duration and force, muscle signal fusion among biomechanically similar keys (e.g., E–D, J–N), and sensitivity to electrode shifts across sessions.

Transformers are particularly suited for this task. Unlike CNNs with limited receptive fields or RNNs with sequential processing and vanishing gradients, multi-head self-attention captures long-range temporal dependencies in parallel, enabling content-aware weighting across the entire 400-timepoint window while remaining compatible with low-latency implementation on resource-constrained devices.

 

  1. Methods: Please consolidate preprocessing details (filter specifications, normalization, window size) and model hyperparameters clearly. Details such as the number of transformer layers, attention heads, dropout rates, and learning rate schedules would greatly enhance reproducibility.

Response:

We thank the reviewer for this valuable comment. We agree that consolidating the preprocessing procedures and model hyperparameters improves the clarity and reproducibility of the study. In the revised manuscript, we have clarified the signal preprocessing pipeline, including filter specifications, window size, augmentation, normalization, and tensor formation. The model and training hyperparameters are already comprehensively reported in Table 1, including the number of transformer layers, attention heads, dropout rate, embedding dimension, optimizer settings, learning rate, gradient clipping, and model parameters.

Action Taken:

We revised Sections 2.1.1 and 2.2.2 to clarify the description of preprocessing. The normalization and tensor-formatting details are provided in Section 2.2.3, and the model/training hyperparameters are reported in Table 1.

Changes in the Manuscript: (Note: The changes are highlighted in red)

2.1.1 Keyboard Typing sEMG Dataset

The bilateral sEMGs (two forearms) were captured at 2000 Hz and bandpass filtered from 10 to 500 Hz using a 4th-order Butterworth filter during the original dataset acquisition stage to reduce motion artifact and high-frequency noise while preserving the frequency content of interest for motor control.

2.2.2 Data Augmentation

To improve model generalization and preserve the physiological significance of the surface EMG samples, all original sample windows were augmented with an augmentation factor of 3: each original window was retained, and two synthetic copies were generated. Following the original 10–500 Hz acquisition-stage filtering described in Section 2.1.1, an additional conservative 50–450 Hz bandpass filtering step was applied to the segmented windows during augmentation/preprocessing. Small-amplitude transient noise was then simulated by adding zero-mean controlled Gaussian noise with σ = 0.01 to each sample.

 

  1. Experimental Setup: It would be instructive to describe data augmentation or regularization methods used (if any). Also, please clarify if the same model weights were used across all participants or individualized.

Response:

We thank the reviewer for this helpful comment. We agree that clarifying the augmentation and regularization strategies, as well as whether the model was shared or individualized across participants, improves the reproducibility and interpretation of the experimental setup. The data augmentation procedure was already clarified in Section 2.2.2 in response to the previous comment. In the present revision, we further clarified the regularization strategies. We explicitly stated that a single global TransTCNet model was trained using pooled participant data and evaluated across all participants without participant-specific fine-tuning or individualized model weights.

Action Taken:

We referred to the revised augmentation description in Section 2.2.2 and further revised Sections 3.1 and 3.3 to clarify the model-weight protocol and regularization strategy.

Changes in the Manuscript:

Training and validation split

A single global TransTCNet model was trained on pooled training data from all participants, and the same model weights were used for evaluation across participants. No participant-specific fine-tuning, calibration, or individualized classifier heads were used. 

Hardware/software

Regularization was implemented through dropout in the model architecture, gradient clipping during optimization, and conservative data augmentation, as described in Section 2.2.2.

  1. Results: Statistical tests are used effectively. Consider including per-class confidence intervals or standard deviations for performance metrics, and expand on the implications of misclassified pairs for practical UI design.

Response:

We thank the reviewer for this constructive suggestion. We agree that the implications of misclassified key pairs are important for the design of practical sEMG-based typing interfaces. Since the current manuscript already reports class-wise precision, recall, F1-score, and confusion-matrix-based error patterns, we retained the existing performance reporting format. However, we expanded the Limitations and Future Work section to discuss how biomechanically similar misclassifications of key pairs may affect real-world interface design and how future systems could mitigate these errors.

Action Taken:

We revised the Limitation and Future Work section to expand the practical implications of misclassified key pairs for future sEMG-based typing interface design.

Changes in the Manuscript:

These frequently confused key pairs also have practical implications for the design of future sEMG-based typing interfaces. In real-world use, errors may not occur uniformly across all characters but may concentrate among keys requiring similar finger movements or overlapping forearm muscle activation patterns. Therefore, future interfaces could incorporate confidence-aware decision rules, adaptive recalibration, context-aware language correction, or confirmation prompts for high-risk key pairs. For example, when the classifier produces low-confidence predictions for commonly confused pairs, the interface could delay commitment or use linguistic context to reduce incorrect character entry.

  1. Figures/Tables: Increase font size on figures for readability. Captions should include interpretations beyond mere descriptions. Table formatting could be improved for clarity.

Response:

We thank the reviewer for this helpful suggestion. We agree that improving figure readability, caption informativeness, and table formatting enhances the manuscript's clarity and accessibility. In the revised manuscript, we increased the font size in the figures to improve readability, revised the figure captions to include brief interpretive statements rather than merely descriptive labels, and improved table formatting to present model settings and results more clearly.

 

 

Action Taken:

Figure captions were revised where appropriate to include brief interpretive statements, and figure quality and font sizes were improved for readability.

Changes in the Manuscript:

For the revised figure captions, improved figure quality, and enlarged figure font sizes, please refer to the updated manuscript, where figure captions were revised where appropriate to improve readability and include brief interpretations of the corresponding results.

  1. Limitations and Future Work: The discussion is thorough and balanced. Strongly recommend expanding on strategies to reduce inter-subject variability and integrating continuous vs. isolated keypress modeling for real-world deployment.

Response:

We thank the reviewer for this positive and constructive comment. We agree that inter-subject variability and the transition from isolated keypress recognition to continuous typing are critical challenges for real-world deployment. In the revised manuscript, we expanded the Limitation and Future Work section to more clearly describe possible strategies for reducing inter-subject variability and for extending the current isolated-keypress framework toward continuous sEMG-based typing systems.

Action Taken:

We revised the Limitation and Future Work section.

Changes in the Manuscript:

Potential strategies include subject-adaptive normalization, transfer learning, domain-adversarial training, few-shot calibration, and lightweight personalization layers that can adapt the global model to new users with minimal additional data.

Future continuous-typing models should also integrate temporal segmentation or sequence-decoding methods to distinguish idle periods, keypress onset and offset, and transitions between consecutive characters, thereby bridging the gap between isolated keypress classification and real-time text-entry deployment.

 

The English could be improved to more clearly express the research. The manuscript is understandable but contains several awkward phrasings, grammatical inconsistencies, and minor typographical errors. Careful proofreading by a native English speaker or professional language editor could improve clarity and flow. Certain technical descriptions could be rewritten for precision and conciseness.

Below are specific recommendations to improve the English language usage for clarity, conciseness, and readability:

  • Improve sentence conciseness and flow:
    • Example: "Such developments may have serious effects on human-computer interaction (HCI)" → "These developments could significantly impact human-computer interaction (HCI)."
    • Avoid redundancy, e.g., "may have serious effects" and "serious" can often be replaced with "significant."
    • Use active voice where possible to improve readability.
  • Correct minor grammatical issues:
    • Consistently use singular or plural nouns, e.g., "several machine learning and deep learning techniques have been used" (correct) rather than mixing tenses.
    • Recheck subject-verb agreement, especially in complex sentences.
  • Review technical terminology for consistency:
    • "sEMG" is correctly used throughout; ensure it is defined at first use if not already.
    • Terms like "cross-subject," "cross-participant," and "cross-session" should be consistent and hyphenated if used as compound adjectives.
  • Clearer expression of statistical and experimental results:
    • For example, instead of "the standard deviation of participants' individual accuracies (0.13%) is very small" → "participants' accuracies showed a low standard deviation (0.13%), indicating high consistency."
    • Avoid phrases like "very small" or "considerably better"—specify quantitative comparisons or use "significantly" if supported statistically.
  • Simplify complex sentences:
    • Some sentences are long and packed with information (e.g., in the introduction and discussion sections). Breaking them into two or more sentences would aid clarity.
  • Punctuation and typographical consistency:
    • Use consistent spacing around dashes and commas.
    • Ensure consistent use of percentage formats (e.g., 87.57% not 87.57 %).
  • Reduction of passive voice where clarity suffers:
    • For example: "Such recordings contain 16 channels…" → "The recordings include 16 EMG channels…"
  • Avoid unnecessary jargon or repetition:
    • Where possible, explain complex terms in simpler language for broader accessibility.
  • Check figure and table references:
    • Ensure uniformity in naming figures (e.g., "Figure 6" vs. "Fig. 6") and consistent tense: "Figure 6 compares" instead of "Figure 6 comparison shows."
  • Consistent tense usage:
    • Prefer past tense for describing completed experiments and present tense for general truths or ongoing relevance.

Response:

We thank the reviewer for these detailed and helpful suggestions regarding English language quality, clarity, conciseness, and technical consistency. We agree that improving sentence flow, grammar, consistency of terminology, figure/table references, tense usage, punctuation, and statistical phrasing enhances the manuscript's readability and professionalism. In the revised manuscript, we carefully proofread the text, revised awkward or redundant expressions, improved sentence conciseness, corrected grammatical and typographical issues, standardized technical terminology, and enhanced the clarity of statistical and experimental descriptions.

Action Taken:

We revised the manuscript throughout, with particular attention to the Introduction, Methods, Results, Limitations, and Future Work sections, figure captions, table formatting, and figure/table references.

Changes in the Manuscript:

The manuscript was carefully proofread and revised for English-language clarity, grammar, conciseness, consistent terminology, punctuation, tense usage, and formatting consistency. Specific revisions include improving awkward sentence structure, reducing redundant wording, standardizing terms (e.g., cross-participant/cross-subject usage where appropriate), ensuring consistent percentage formatting, improving figure and table references, and rewriting technical descriptions for greater precision and readability. For the detailed language and formatting revisions, please refer to the updated manuscript.

Author Response File: Author Response.pdf

Reviewer 2 Report

Comments and Suggestions for Authors

A. Summary of the manuscript and its key contributions
The manuscript presents an interesting study on character-level typing recognition using surface electromyography (sEMG) signals through a novel deep learning architecture called TransTCNet. The paper is generally well-structured and addresses an important topic in Human-Computer Interaction and neural interfaces, specifically aiming to provide a silent, hands-free typing solution.
The main contribution lies in the proposed two-stage neural network design, which integrates dilated causal convolutions for extracting local temporal features with a transformer-based encoder for capturing long-range global dependencies. The results indicate a high validation accuracy of 96.53%, significantly outperforming previous baseline models like SVMs and MLPs on the same 26-class dataset. Overall, the study offers a viable solution for wearable assistive technologies, prosthetic control, and immersive AR/VR environments where traditional input peripherals are impractical. 

B. Detailed evaluation of the methodology, analyses, and conclusions
The manuscript is rigorously structured, facilitating a clear progression from theoretical foundation to experimental validation.
The authors provide a comprehensive overview of the state of the art in sEMG-based gesture and activity recognition, noting the evolution from classical processing methods to deep learning architectures, identifying a key limitation in current models: the difficulty of simultaneously capturing local temporal features and global signal dependencies.
The research objective is clearly defined, aiming to develop a highly reliable classification model of keyboard presses using bilateral forearm sEMG data. 
The cited references are recent and relevant to the investigated topic.
The Materials and Methods section is logically structured and well-documented. It begins with the dataset description, detailing a publicly available dataset (reference [12]) of 19 subjects where 16 bipolar electrodes captured muscle activity at 2000 Hz during a metronome-paced typing task. The section continues with data preprocessing steps, including data curation, signal segmentation, signal augmentation, signal normalization, and signal reshaping. The augmentation strategy, involving Gaussian noise and sensor dropout simulation, is a valuable addition that enhances the model's robustness against real-world signal interference. The neural network architecture is described in detail, explaining the specific roles of the Temporal Convolutional Module (extracting local features via dilated convolutions) and the Global Dependency Module (utilizing Multi-Head Attention for long-range relationships).
The Experimental Setup section provides the necessary transparency for reproducibility and presents the training hyperparameters, evaluation metrics, and the computational environment using an NVIDIA RTX 5080 GPU.
The analysis of the results is exhaustive and utilizes multiple performance metrics, including accuracy/loss curves, confusion matrices, and participant-wise performance. An ablation study validates the hybrid design; it demonstrates that removing either the Temporal or Global module leads to a drastic drop in performance. The inclusion of t-SNE and PCA techniques to visualize class separability in the feature space adds an important qualitative dimension to the quantitative analysis.
Conclusions are coherent and provide a concise synthesis of the main findings, confirming that integrating attention mechanisms with causal convolutions is optimal for complex sEMG signals.
The authors acknowledge current limitations and propose relevant future directions, such as real-time implementation and optimization for new, unseen subjects.

C. Constructive feedback for the authors, highlighting areas for improvement
The manuscript is very well-prepared and demonstrates a high scientific standard. The authors have presented a deep learning approach that is both technically sound and clearly articulated. The methodology is robust, and the paper contributes significantly to the field of neural interfaces.
To further elevate the impact of this high-quality work, the following points could be addressed:

1. Given that the study utilizes the dataset presented in Reference [12], it is important for the authors to more clearly delineate which subsections containing dataset information originate from that source and which specific sections describe their original contributions to the preparation and refinement of the data for training the TransTCNet model.

2. While the benefit of the modules is clear, a deeper discussion on why the 1D-CNN baseline performed significantly lower (48.66%) compared to the Temporal module (72.67%) would provide better insight into the feature extraction process.  

3. It would be highly valuable if the authors discussed their vision for deploying this solution on platforms capable of running the application in real-time. Including a detailed analysis of the computational cost, specifically the number of trainable parameters and floating-point operations (FLOPs), would significantly increase the paper's value by helping to evaluate whether the architecture can be sustained by mobile or edge processing units.

4. To move from a laboratory prototype to a real-world application, the authors should consider that a 26-character alphabet is insufficient for text composition. The study would be much more impactful if it included a discussion on how the model could be expanded to recognize control commands such as Delete, Backspace, or Enter. It would be important for the authors to reflect on whether this expansion would require a more complex architecture or if the current TransTCNet model has the capacity to scale its vocabulary without a significant loss in accuracy. Including these considerations would clarify the path toward a fully functional "silent typing" interface.

5. While the metronome-paced task (75 BPM) provides a controlled environment, the study would benefit from a discussion on how the model might perform with natural, irregular typing rhythms or during prolonged sessions where muscle fatigue becomes a factor. 

Author Response

Reviewer 2:

  1. Summary of the manuscript and its key contributions

The manuscript presents an interesting study on character-level typing recognition using surface electromyography (sEMG) signals through a novel deep learning architecture called TransTCNet. The paper is generally well-structured and addresses an important topic in Human-Computer Interaction and neural interfaces, specifically aiming to provide a silent, hands-free typing solution.
The main contribution lies in the proposed two-stage neural network design, which integrates dilated causal convolutions for extracting local temporal features with a transformer-based encoder for capturing long-range global dependencies. The results indicate a high validation accuracy of 96.53%, significantly outperforming previous baseline models like SVMs and MLPs on the same 26-class dataset. Overall, the study offers a viable solution for wearable assistive technologies, prosthetic control, and immersive AR/VR environments where traditional input peripherals are impractical. 

  1. Detailed evaluation of the methodology, analyses, and conclusions
    The manuscript is rigorously structured, facilitating a clear progression from theoretical foundation to experimental validation.

The authors provide a comprehensive overview of the state of the art in sEMG-based gesture and activity recognition, noting the evolution from classical processing methods to deep learning architectures, identifying a key limitation in current models: the difficulty of simultaneously capturing local temporal features and global signal dependencies.
The research objective is clearly defined, aiming to develop a highly reliable classification model of keyboard presses using bilateral forearm sEMG data.

The cited references are recent and relevant to the investigated topic.
The Materials and Methods section is logically structured and well-documented. It begins with the dataset description, detailing a publicly available dataset (reference [12]) of 19 subjects where 16 bipolar electrodes captured muscle activity at 2000 Hz during a metronome-paced typing task. The section continues with data preprocessing steps, including data curation, signal segmentation, signal augmentation, signal normalization, and signal reshaping. The augmentation strategy, involving Gaussian noise and sensor dropout simulation, is a valuable addition that enhances the model's robustness against real-world signal interference. The neural network architecture is described in detail, explaining the specific roles of the Temporal Convolutional Module (extracting local features via dilated convolutions) and the Global Dependency Module (utilizing Multi-Head Attention for long-range relationships).
The Experimental Setup section provides the necessary transparency for reproducibility and presents the training hyperparameters, evaluation metrics, and the computational environment using an NVIDIA RTX 5080 GPU.

The analysis of the results is exhaustive and utilizes multiple performance metrics, including accuracy/loss curves, confusion matrices, and participant-wise performance. An ablation study validates the hybrid design; it demonstrates that removing either the Temporal or Global module leads to a drastic drop in performance. The inclusion of t-SNE and PCA techniques to visualize class separability in the feature space adds an important qualitative dimension to the quantitative analysis.
Conclusions are coherent and provide a concise synthesis of the main findings, confirming that integrating attention mechanisms with causal convolutions is optimal for complex sEMG signals.
The authors acknowledge current limitations and propose relevant future directions, such as real-time implementation and optimization for new, unseen subjects.

  1. Constructive feedback for the authors, highlighting areas for improvement
    The manuscript is very well-prepared and demonstrates a high scientific standard. The authors have presented a deep learning approach that is both technically sound and clearly articulated. The methodology is robust, and the paper contributes significantly to the field of neural interfaces.
    To further elevate the impact of this high-quality work, the following points could be addressed:

Response:

We sincerely thank the reviewer for the positive and encouraging evaluation of our manuscript. We are grateful that the reviewer recognized the significance of character-level sEMG typing recognition, the contribution of the proposed TransTCNet architecture, the clarity of the methodology, the comprehensiveness of the experimental analysis, and the relevance of the future research directions. We also appreciate the reviewer's constructive suggestions for further improving the manuscript. In response, we have revised the manuscript to improve clarity, reproducibility, figure/table readability, discussion of practical implications, and the presentation of limitations and future work.

Action Taken:

We revised the manuscript throughout to address the reviewer's constructive feedback and to improve overall clarity, organization, presentation quality, and the discussion of practical deployment considerations.

Changes in the Manuscript:

Please refer to the revised manuscript for the detailed changes made in response to the reviewer's comments, including improvements to the Introduction, Methods, Experimental Setup, Results presentation, figure captions, table formatting, and Limitations and Future Work section.

  1. Given that the study utilizes the dataset presented in reference [12], it is important for the authors to more clearly delineate which subsections containing dataset information originate from that source and which specific sections describe their original contributions to the preparation and refinement of the data for training the TransTCNet model.

Response:

We thank the reviewer for this important clarification. We agree that, because the study uses the publicly available dataset reported in Reference [12], the manuscript should clearly distinguish between dataset information originating from the source and the data preparation steps performed in the present study. In the revised manuscript, we clarified that the participant protocol, electrode configuration, sampling rate, and acquisition-stage filtering information are based on Reference [12]. We also clarified that the subsequent data curation, segmentation, augmentation, normalization, tensor formatting, and training/validation preparation were performed in this study to train and evaluate the proposed TransTCNet model.

Action Taken:

We revised Sections 2.1.1 and 2.2 to clearly distinguish the dataset information obtained from Reference [12] from the data preparation and refinement procedures performed in the present study.

Changes in the Manuscript:

Section 2.1.1: The dataset description in this subsection, including the participant protocol, electrode configuration, sampling frequency, typing task, and acquisition-stage filtering, is based on the original dataset reported in the reference [14].

Section 2.2: The procedures described in this section represent the data preparation and refinement steps performed in the present study after obtaining the publicly available dataset from the reference [14].

  1. While the benefit of the modules is clear, a deeper discussion on why the 1D-CNN baseline performed significantly lower (48.66%) compared to the Temporal module (72.67%) would provide better insight into the feature extraction process.  

Response:

We thank the reviewer for this insightful comment. We agree that the performance gap between the standard 1D-CNN baseline and the Temporal module should be explained more clearly, as it provides important insight into the role of temporal feature extraction in character-level sEMG recognition. In the revised manuscript, we clarified that the lower performance of the 1D-CNN baseline is likely due to its limited temporal receptive field and reduced ability to capture fine-grained timing variations in short sEMG windows. We also explained that the Temporal module improves performance by using dilated causal convolutions to capture multiscale muscle-activation patterns while preserving the sequential dynamics of keypresses.

Action Taken:

We revised Sections 4.9.1 and 4.9.2 to expand the explanation of the performance difference between the 1D-CNN baseline and the Temporal module.

Changes in the Manuscript:

Section 4.9.1: The relatively low performance suggests that a simple 1D convolutional model was insufficient for modeling the fine-grained timing variations in short character-level sEMG windows. Although the baseline could extract local signal patterns, its limited temporal receptive field made it less effective at capturing multiscale muscle activation dynamics associated with different keypresses.

Section 4.9.2: Compared with the standard 1D-CNN baseline, the Temporal module better preserved sequential activation patterns and captured both short- and long-range temporal dependencies within the 0.2-second sEMG window, accounting for the substantial performance improvement from 48.66% to 72.67%.

  1. It would be highly valuable if the authors discussed their vision for deploying this solution on platforms capable of running the application in real-time. Including a detailed analysis of the computational cost, specifically the number of trainable parameters and floating-point operations (FLOPs), would significantly increase the paper's value by helping to evaluate whether the architecture can be sustained by mobile or edge processing units.

Response:

We thank the reviewer for this valuable suggestion. We agree that discussing real-time deployment and computational cost is important for evaluating the practical feasibility of TransTCNet on mobile and edge platforms. In the revised manuscript, we expanded the Limitation and Future Work section to describe a possible streaming, window-based inference framework for real-time sEMG typing. We also clarified that the trainable parameter count reported in Table 1 provides an initial indication of the model's memory requirements. At the same time, detailed hardware-specific evaluation of FLOPs, inference latency, power consumption, and continuous idle-state handling should be performed on target mobile or embedded processors in future work.

Action Taken:

We revised the Limitations and Future Work section to expand the discussion of real-time deployment and to evaluate future hardware-specific computational costs.

Changes in the Manuscript:

For real-time deployment, TransTCNet could be implemented as a streaming, window-based inference system, in which each 0.2-second sEMG segment is processed sequentially to generate character predictions. The trainable parameter count reported in Table 1 provides an initial indication of the model's memory footprint; however, practical deployment on mobile or embedded processors also requires hardware-specific evaluation of FLOPs, inference latency, power consumption, and continuous idle-state handling during real-world typing.

  1. To move from a laboratory prototype to a real-world application, the authors should consider that a 26-character alphabet is insufficient for text composition. The study would be much more impactful if it included a discussion on how the model could be expanded to recognize control commands such as Delete, Backspace, or Enter. It would be important for the authors to reflect on whether this expansion would require a more complex architecture or if the current TransTCNet model has the capacity to scale its vocabulary without a significant loss in accuracy. Including these considerations would clarify the path toward a fully functional "silent typing" interface.

Response:

We thank the reviewer for this insightful suggestion. We agree that recognizing only the 26 alphabetic characters is not sufficient for a fully functional silent typing interface. Practical text composition would require additional non-alphabetic commands such as Delete, Backspace, Enter, Space, punctuation, and possibly mode-switching commands. In the revised manuscript, we expanded the Limitations and Future Work section to discuss how TransTCNet could be extended beyond the current 26-class alphabet recognition task. We clarified that the current architecture could, in principle, be scaled by expanding the output classification layer and training with additional labeled command classes. Still, that real-world performance would depend on the separability of the added gestures, the availability of sufficient training data, and the integration of sequence-level correction or command-handling mechanisms.

Action Taken:

We revised the Limitations and Future Work section to discuss expanding beyond 26 alphabetic characters toward a more complete silent typing interface that includes control commands.

Changes in the Manuscript:

Third, the current study focused only on 26 alphabetic keypress classes. In contrast, a fully functional silent-typing interface would also require non-alphabetic control commands, such as Space, Delete, Backspace, Enter, and punctuation. The current TransTCNet architecture could be extended to these additional classes by expanding the final classification layer and training the model with labeled sEMG examples for each command. However, scaling the vocabulary may introduce additional class overlap and could reduce accuracy if the added commands produce biomechanically similar muscle activation patterns. Future work should therefore evaluate whether the existing temporal-convolution and transformer-attention backbone can maintain performance as the vocabulary expands, or whether additional mechanisms such as hierarchical classification, command-specific calibration, language-model correction, or sequence-level decoding are needed for a complete real-time silent typing system.

  1. While the metronome-paced task (75 BPM) provides a controlled environment, the study would benefit from a discussion on how the model might perform with natural, irregular typing rhythms or during prolonged sessions where muscle fatigue becomes a factor. 

Response:

We thank the reviewer for this valuable suggestion. We agree that the metronome-paced task provides a controlled experimental setting but does not fully capture the variability of natural typing behavior or the effects of prolonged use. In the revised manuscript, we expanded the Limitations and Future Work section to discuss how irregular typing rhythms, variable keypress timing, and muscle fatigue may affect model performance in real-world deployment. We also clarified that future studies should evaluate TransTCNet under natural typing conditions and prolonged sessions, with possible fatigue-aware adaptation strategies.

Action Taken:

We revised the Limitations and Future Work section to discuss natural typing rhythms, prolonged use, and muscle fatigue as important considerations for real-world deployment.

Changes in the Manuscript:

Because the original typing task was metronome-paced at 75 BPM, future work should also evaluate TransTCNet under natural, irregular typing rhythms, in which keypress timing, force, and inter-keystroke intervals may vary more substantially. In addition, prolonged typing sessions may introduce muscle fatigue, which can alter sEMG amplitude and frequency characteristics over time [36]. Therefore, future longitudinal studies should examine model stability under extended use and consider fatigue-aware recalibration or adaptive normalization strategies.

Author Response File: Author Response.pdf

Back to TopTop