Sensors 2013, 13(12), 16533-16550; doi:10.3390/s131216533

A Novel Voice Sensor for the Detection of Speech Signals

Received: 3 November 2013; in revised form: 26 November 2013 / Accepted: 27 November 2013 / Published: 2 December 2013
(This article belongs to the Section Physical Sensors)
This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Abstract: In order to develop a novel voice sensor to detect human voices, the use of features which are more robust to noise is an important issue. Voice sensor is also called voice activity detection (VAD). Due to that the inherent nature of the formant structure only occurred on the speech spectrogram (well-known as voiceprint), Wu et al. were the first to use band-spectral entropy (BSE) to describe the characteristics of voiceprints. However, the performance of VAD based on BSE feature was degraded in colored noise (or voiceprint-like noise) environments. In order to solve this problem, we propose the two-dimensional part-band energy entropy (TD-PBEE) parameter based on two variables: part-band partition number upon frequency index and long-term window size upon time index to further improve the BSE-based VAD algorithm. The two variables can efficiently represent the characteristics of voiceprints on each critical frequency band and use long-term information for noisy speech spectrograms, respectively. The TD-PBEE parameter can be regarded as a PBEE parameter over time. First, the strength of voiceprints can be partly enhanced by using four entropies applied to four part-bands. We can use the four part-band energy entropies for describing the voiceprints in detail. Due to the characteristics of non-stationary for speech and various noises, we will then use long-term information processing to refine the PBEE, so the voice-like noise can be distinguished from noisy speech through the concept of PBEE with long-term information. Our experiments show that the proposed feature extraction with the TD-PBEE parameter is quite insensitive to background noise. The proposed TD-PBEE-based VAD algorithm is evaluated for four types of noises and five signal-to-noise ratio (SNR) levels. We find that the accuracy of the proposed TD-PBEE-based VAD algorithm averaged over all noises and all SNR levels is better than that of other considered VAD algorithms.
Keywords: voice sensor; two-dimensional; part-band energy entropy; long-term information analysis; Mel-scaled filter bank
PDF Full-text Download PDF Full-Text [817 KB, uploaded 21 June 2014 10:37 CEST]

Export to BibTeX |

MDPI and ACS Style

Wang, K.-C. A Novel Voice Sensor for the Detection of Speech Signals. Sensors 2013, 13, 16533-16550.

AMA Style

Wang K-C. A Novel Voice Sensor for the Detection of Speech Signals. Sensors. 2013; 13(12):16533-16550.

Chicago/Turabian Style

Wang, Kun-Ching. 2013. "A Novel Voice Sensor for the Detection of Speech Signals." Sensors 13, no. 12: 16533-16550.

Sensors EISSN 1424-8220 Published by MDPI AG, Basel, Switzerland RSS E-Mail Table of Contents Alert