Next Article in Journal
An Intelligent Monitoring System for the Driving Environment of Explosives Transport Vehicles Based on Consumer-Grade Cameras
Previous Article in Journal
Enhanced Blockchain-Based Data Poisoning Defense Mechanism
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Enhancing Far-Field Speech Recognition with Mixer: A Novel Data Augmentation Approach

1
School of Information Systems Engineering, Information Engineering University, Zhengzhou 450001, China
2
Laboratory for Advanced Computing and Intelligence Engineering, Wuxi 214000, China
3
Research and Development Department, Zhengzhou Xinda Institute of Advanced Technology, Zhengzhou 450001, China
*
Authors to whom correspondence should be addressed.
Appl. Sci. 2025, 15(7), 4073; https://doi.org/10.3390/app15074073
Submission received: 4 March 2025 / Revised: 31 March 2025 / Accepted: 3 April 2025 / Published: 7 April 2025
(This article belongs to the Section Computing and Artificial Intelligence)

Abstract

Recent advancements in end-to-end (E2E) modeling have notably improved automatic speech recognition (ASR) systems; however, far-field speech recognition (FSR) remains challenging due to signal degradation from factors such as low signal-to-noise ratio, reverberation, and interfering sounds. This requires richer training data and multi-channel speech enhancement. To address this gap, we introduce Mixer, a novel data augmentation technique designed to further enhance the performance of large-scale pre-trained models for FSR. Mixer interpolates and mixes feature representations of speech samples and their corresponding losses, extending the MixSpeech framework to intermediate layers of Whisper. Additionally, we propose Mixer-C, which further leverages multi-channel information by combining speech from different microphone channels using a channel selector. Experimental results demonstrate that Mixer significantly outperforms existing methods, including SpecAugment, achieving a relative word error rate (WER) reduction of 3.6% compared to the baseline. Furthermore, Mixer-C offers an additional WER improvement of 2.2%, showcasing its efficacy in improving FSR accuracy.
Keywords: far-field ASR; data augmentation; mixer; whisper far-field ASR; data augmentation; mixer; whisper

Share and Cite

MDPI and ACS Style

Niu, T.; Chen, Y.; Qu, D.; Hu, H. Enhancing Far-Field Speech Recognition with Mixer: A Novel Data Augmentation Approach. Appl. Sci. 2025, 15, 4073. https://doi.org/10.3390/app15074073

AMA Style

Niu T, Chen Y, Qu D, Hu H. Enhancing Far-Field Speech Recognition with Mixer: A Novel Data Augmentation Approach. Applied Sciences. 2025; 15(7):4073. https://doi.org/10.3390/app15074073

Chicago/Turabian Style

Niu, Tong, Yaqi Chen, Dan Qu, and Hengbo Hu. 2025. "Enhancing Far-Field Speech Recognition with Mixer: A Novel Data Augmentation Approach" Applied Sciences 15, no. 7: 4073. https://doi.org/10.3390/app15074073

APA Style

Niu, T., Chen, Y., Qu, D., & Hu, H. (2025). Enhancing Far-Field Speech Recognition with Mixer: A Novel Data Augmentation Approach. Applied Sciences, 15(7), 4073. https://doi.org/10.3390/app15074073

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop