Next Article in Journal
Study on the Energy Absorption Performance of Triply Periodic Minimal Surface (TPMS) Structures at Different Load-Bearing Angles
Next Article in Special Issue
Perceptive Recommendation Robot: Enhancing Receptivity of Product Suggestions Based on Customers’ Nonverbal Cues
Previous Article in Journal
From Nature to Technology: Exploring the Potential of Plant-Based Materials and Modified Plants in Biomimetics, Bionics, and Green Innovations
Previous Article in Special Issue
Whole-Body Dynamics for Humanoid Robot Fall Protection Trajectory Generation with Wall Support
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Assessment of Pepper Robot’s Speech Recognition System through the Lens of Machine Learning

Educational Technology Laboratory, Intelligent System and Analytics Group, Department of Computer Science (IDI), Norwegian University of Science and Technology, 2815 Gjøvik, Norway
*
Author to whom correspondence should be addressed.
Biomimetics 2024, 9(7), 391; https://doi.org/10.3390/biomimetics9070391
Submission received: 24 April 2024 / Revised: 5 June 2024 / Accepted: 24 June 2024 / Published: 27 June 2024
(This article belongs to the Special Issue Intelligent Human-Robot Interaction: 2nd Edition)

Abstract

Speech comprehension can be challenging due to multiple factors, causing inconvenience for both the speaker and the listener. In such situations, using a humanoid robot, Pepper, can be beneficial as it can display the corresponding text on its screen. However, prior to that, it is essential to carefully assess the accuracy of the audio recordings captured by Pepper. Therefore, in this study, an experiment is conducted with eight participants with the primary objective of examining Pepper’s speech recognition system with the help of audio features such as Mel-Frequency Cepstral Coefficients, spectral centroid, spectral flatness, the Zero-Crossing Rate, pitch, and energy. Furthermore, the K-means algorithm was employed to create clusters based on these features with the aim of selecting the most suitable cluster with the help of the speech-to-text conversion tool Whisper. The selection of the best cluster is accomplished by finding the maximum accuracy data points lying in a cluster. A criterion of discarding data points with values of WER above 0.3 is imposed to achieve this. The findings of this study suggest that a distance of up to one meter from the humanoid robot Pepper is suitable for capturing the best speech recordings. In contrast, age and gender do not influence the accuracy of recorded speech. The proposed system will provide a significant strength in settings where subtitles are required to improve the comprehension of spoken statements.
Keywords: Pepper robot; speech recognition; audio features; evaluation metrics; K-means clustering Pepper robot; speech recognition; audio features; evaluation metrics; K-means clustering
Graphical Abstract

Share and Cite

MDPI and ACS Style

Pande, A.; Mishra, D. Assessment of Pepper Robot’s Speech Recognition System through the Lens of Machine Learning. Biomimetics 2024, 9, 391. https://doi.org/10.3390/biomimetics9070391

AMA Style

Pande A, Mishra D. Assessment of Pepper Robot’s Speech Recognition System through the Lens of Machine Learning. Biomimetics. 2024; 9(7):391. https://doi.org/10.3390/biomimetics9070391

Chicago/Turabian Style

Pande, Akshara, and Deepti Mishra. 2024. "Assessment of Pepper Robot’s Speech Recognition System through the Lens of Machine Learning" Biomimetics 9, no. 7: 391. https://doi.org/10.3390/biomimetics9070391

APA Style

Pande, A., & Mishra, D. (2024). Assessment of Pepper Robot’s Speech Recognition System through the Lens of Machine Learning. Biomimetics, 9(7), 391. https://doi.org/10.3390/biomimetics9070391

Article Metrics

Back to TopTop