Next Article in Journal
Effect of Substrate Compliance on the Jumping Mechanism of the Tree Frog (Polypedates dennys)
Next Article in Special Issue
Predicting and Synchronising Co-Speech Gestures for Enhancing Human–Robot Interactions Using Deep Learning Models
Previous Article in Journal
Bioinspired Approaches and Their Philosophical–Ethical Dimensions: A Narrative Review
Previous Article in Special Issue
Robotic Removal and Collection of Screws in Collaborative Disassembly of End-of-Life Electric Vehicle Batteries
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Self-Supervised Voice Denoising Network for Multi-Scenario Human–Robot Interaction

by
Mu Li
1,
Wenjin Xu
1,
Chao Zeng
2 and
Ning Wang
3,*
1
Key Laboratory of Autonomous Systems and Networked Control, Ministry of Education, College of Automation Science and Engineering, South China University of Technology, Guangzhou 510640, China
2
Department of Computer Science, University of Liverpool, Liverpool L69 3BX, UK
3
School of Computing and Digital Technologies, Sheffield Hallam University, Sheffield S1 2NU, UK
*
Author to whom correspondence should be addressed.
Biomimetics 2025, 10(9), 603; https://doi.org/10.3390/biomimetics10090603
Submission received: 15 June 2025 / Revised: 24 August 2025 / Accepted: 6 September 2025 / Published: 9 September 2025
(This article belongs to the Special Issue Intelligent Human–Robot Interaction: 4th Edition)

Abstract

Human–robot interaction (HRI) via voice command has significantly advanced in recent years, with large Vision–Language–Action (VLA) models demonstrating particular promise in human–robot voice interaction. However, these systems still struggle with environmental noise contamination during voice interaction and lack a specialized denoising network for multi-speaker command isolation in an overlapping speech scenario. To overcome these challenges, we introduce a method to enhance voice command-based HRI in noisy environments, leveraging synthetic data and a self-supervised denoising network to enhance its real-world applicability. Our approach focuses on improving self-supervised network performance in denoising mixed-noise audio through training data scaling. Extensive experiments show our method outperforms existing approaches in simulation and achieves 7.5% higher accuracy than the state-of-the-art method in noisy real-world environments, enhancing voice-guided robot control.
Keywords: human–robot interaction; voice denoising; self-supervised learning; data synthesis human–robot interaction; voice denoising; self-supervised learning; data synthesis

Share and Cite

MDPI and ACS Style

Li, M.; Xu, W.; Zeng, C.; Wang, N. Self-Supervised Voice Denoising Network for Multi-Scenario Human–Robot Interaction. Biomimetics 2025, 10, 603. https://doi.org/10.3390/biomimetics10090603

AMA Style

Li M, Xu W, Zeng C, Wang N. Self-Supervised Voice Denoising Network for Multi-Scenario Human–Robot Interaction. Biomimetics. 2025; 10(9):603. https://doi.org/10.3390/biomimetics10090603

Chicago/Turabian Style

Li, Mu, Wenjin Xu, Chao Zeng, and Ning Wang. 2025. "Self-Supervised Voice Denoising Network for Multi-Scenario Human–Robot Interaction" Biomimetics 10, no. 9: 603. https://doi.org/10.3390/biomimetics10090603

APA Style

Li, M., Xu, W., Zeng, C., & Wang, N. (2025). Self-Supervised Voice Denoising Network for Multi-Scenario Human–Robot Interaction. Biomimetics, 10(9), 603. https://doi.org/10.3390/biomimetics10090603

Article Metrics

Back to TopTop