Next Article in Journal
Variability of Air Pollutants in the Indoor Air of a General Store
Next Article in Special Issue
Unmasking Nasality to Assess Hypernasality
Previous Article in Journal
A Stability Analysis of an Abandoned Gypsum Mine Based on Numerical Simulation Using the Itasca Model for Advanced Strain Softening Constitutive Model
Previous Article in Special Issue
Orthogonalization of the Sensing Matrix Through Dominant Columns in Compressive Sensing for Speech Enhancement
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Environment-Aware Knowledge Distillation for Improved Resource-Constrained Edge Speech Recognition

1
Institut National de la Recherche Scientifique (INRS-EMT), Université du Québec, Montreal, QC H5A 1K6, Canada
2
INRS-UQO Mixed Research Unit on Cybersecurity, Gatineau, QC J8X 3X7, Canada
*
Author to whom correspondence should be addressed.
Appl. Sci. 2023, 13(23), 12571; https://doi.org/10.3390/app132312571
Submission received: 18 October 2023 / Revised: 18 November 2023 / Accepted: 20 November 2023 / Published: 22 November 2023
(This article belongs to the Special Issue Advances in Speech and Language Processing)

Abstract

Recent advances in self-supervised learning have allowed automatic speech recognition (ASR) systems to achieve state-of-the-art (SOTA) word error rates (WER) while requiring only a fraction of the labeled data needed by its predecessors. Notwithstanding, while such models achieve SOTA results in matched train/test scenarios, their performance degrades substantially when tested in unseen conditions. To overcome this problem, strategies such as data augmentation and/or domain adaptation have been explored. Available models, however, are still too large to be considered for edge speech applications on resource-constrained devices; thus, model compression tools, such as knowledge distillation, are needed. In this paper, we propose three innovations on top of the existing DistilHuBERT distillation recipe: optimize the prediction heads, employ a targeted data augmentation method for different environmental scenarios, and employ a real-time environment estimator to choose between compressed models for inference. Experiments with the LibriSpeech dataset, corrupted with varying noise types and reverberation levels, show the proposed method outperforming several benchmark methods, both original and compressed, by as much as 48.4% and 89.2% in the word error reduction rate in extremely noisy and reverberant conditions, respectively, while reducing by 50% the number of parameters. Thus, the proposed method is well suited for resource-constrained edge speech recognition applications.
Keywords: automatic speech recognition; knowledge distillation; self-supervised learning; modulation spectrum; context awareness automatic speech recognition; knowledge distillation; self-supervised learning; modulation spectrum; context awareness

Share and Cite

MDPI and ACS Style

Pimentel, A.; Guimarães, H.R.; Avila, A.; Falk, T.H. Environment-Aware Knowledge Distillation for Improved Resource-Constrained Edge Speech Recognition. Appl. Sci. 2023, 13, 12571. https://doi.org/10.3390/app132312571

AMA Style

Pimentel A, Guimarães HR, Avila A, Falk TH. Environment-Aware Knowledge Distillation for Improved Resource-Constrained Edge Speech Recognition. Applied Sciences. 2023; 13(23):12571. https://doi.org/10.3390/app132312571

Chicago/Turabian Style

Pimentel, Arthur, Heitor R. Guimarães, Anderson Avila, and Tiago H. Falk. 2023. "Environment-Aware Knowledge Distillation for Improved Resource-Constrained Edge Speech Recognition" Applied Sciences 13, no. 23: 12571. https://doi.org/10.3390/app132312571

APA Style

Pimentel, A., Guimarães, H. R., Avila, A., & Falk, T. H. (2023). Environment-Aware Knowledge Distillation for Improved Resource-Constrained Edge Speech Recognition. Applied Sciences, 13(23), 12571. https://doi.org/10.3390/app132312571

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop