Next Article in Journal
A Fast Loss Model for Cascode GaN-FETs and Real-Time Degradation-Sensitive Control of Solid-State Transformers
Next Article in Special Issue
Nonlinear Regularization Decoding Method for Speech Recognition
Previous Article in Journal
Dual-Energy Processing of X-ray Images of Beryl in Muscovite Obtained Using Pulsed X-ray Sources
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Characterization of Deep Learning-Based Speech-Enhancement Techniques in Online Audio Processing Applications

Computer Science Department, Instituto de Investigaciones en Matematicas Aplicadas y en Sistemas, Universidad Nacional Autonoma de Mexico, Mexico City 3000, Mexico
Sensors 2023, 23(9), 4394; https://doi.org/10.3390/s23094394
Submission received: 7 March 2023 / Revised: 24 April 2023 / Accepted: 28 April 2023 / Published: 29 April 2023

Abstract

Deep learning-based speech-enhancement techniques have recently been an area of growing interest, since their impressive performance can potentially benefit a wide variety of digital voice communication systems. However, such performance has been evaluated mostly in offline audio-processing scenarios (i.e., feeding the model, in one go, a complete audio recording, which may extend several seconds). It is of significant interest to evaluate and characterize the current state-of-the-art in applications that process audio online (i.e., feeding the model a sequence of segments of audio data, concatenating the results at the output end). Although evaluations and comparisons between speech-enhancement techniques have been carried out before, as far as the author knows, the work presented here is the first that evaluates the performance of such techniques in relation to their online applicability. This means that this work measures how the output signal-to-interference ratio (as a separation metric), the response time, and memory usage (as online metrics) are impacted by the input length (the size of audio segments), in addition to the amount of noise, amount and number of interferences, and amount of reverberation. Three popular models were evaluated, given their availability on public repositories and online viability, MetricGAN+, Spectral Feature Mapping with Mimic Loss, and Demucs-Denoiser. The characterization was carried out using a systematic evaluation protocol based on the Speechbrain framework. Several intuitions are presented and discussed, and some recommendations for future work are proposed.
Keywords: speech enhancement; online applicability; real-time factor speech enhancement; online applicability; real-time factor

Share and Cite

MDPI and ACS Style

Rascon, C. Characterization of Deep Learning-Based Speech-Enhancement Techniques in Online Audio Processing Applications. Sensors 2023, 23, 4394. https://doi.org/10.3390/s23094394

AMA Style

Rascon C. Characterization of Deep Learning-Based Speech-Enhancement Techniques in Online Audio Processing Applications. Sensors. 2023; 23(9):4394. https://doi.org/10.3390/s23094394

Chicago/Turabian Style

Rascon, Caleb. 2023. "Characterization of Deep Learning-Based Speech-Enhancement Techniques in Online Audio Processing Applications" Sensors 23, no. 9: 4394. https://doi.org/10.3390/s23094394

APA Style

Rascon, C. (2023). Characterization of Deep Learning-Based Speech-Enhancement Techniques in Online Audio Processing Applications. Sensors, 23(9), 4394. https://doi.org/10.3390/s23094394

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop