Next Article in Journal
Advancements in Medical Imaging and Image-Guided Procedures: A Potential—Or Rather Likely—Paradigm Shift in Diagnosis and Therapy: Understand Disruption and Take Advantage of It!
Next Article in Special Issue
Lip2Speech: Lightweight Multi-Speaker Speech Reconstruction with Gabor Features
Previous Article in Journal
Influence Analysis of Liquefiable Interlayer on Seismic Response of Underground Station Structure
Previous Article in Special Issue
Amplitude and Phase Information Interaction for Speech Enhancement Method
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Exploring Multi-Stage GAN with Self-Attention for Speech Enhancement

by
Bismark Kweku Asiedu Asante
1,*,
Clifford Broni-Bediako
2 and
Hiroki Imamura
1,*
1
Graduate School of Science and Engineering, Soka University, Hachioji City 192-8577, Japan
2
RIKEN Center for Advanced Intelligence Project, Nihonbashi, Chuo City 103-0027, Japan
*
Authors to whom correspondence should be addressed.
Appl. Sci. 2023, 13(16), 9217; https://doi.org/10.3390/app13169217
Submission received: 23 June 2023 / Revised: 23 July 2023 / Accepted: 25 July 2023 / Published: 14 August 2023
(This article belongs to the Special Issue Advanced Technology in Speech and Acoustic Signal Processing)

Abstract

Multi-stage or multi-generator generative adversarial networks (GANs) have recently been demonstrated to be effective for speech enhancement. The existing multi-generator GANs for speech enhancement only use convolutional layers for synthesising clean speech signals. This reliance on convolution operation may result in masking the temporal dependencies within the signal sequence. This study explores self-attention to address the temporal dependency issue in multi-generator speech enhancement GANs to improve their enhancement performance. We empirically study the effect of integrating a self-attention mechanism into the convolutional layers of the multiple generators in multi-stage or multi-generator speech enhancement GANs, specifically, the ISEGAN and the DSEGAN networks. The experimental results show that introducing a self-attention mechanism into ISEGAN and DSEGAN leads to improvements in their speech enhancement quality and intelligibility across the objective evaluation metrics. Furthermore, we observe that adding self-attention to the ISEGAN’s generators does not only improves its enhancement performance but also bridges the performance gap between the ISEGAN and the DSEGAN with a smaller model footprint. Overall, our findings highlight the potential of self-attention in improving the enhancement performance of multi-generator speech enhancement GANs.
Keywords: speech enhancement; generative adversarial network (GAN); multi-stage GAN; multi-generator GAN; self-attention mechanism speech enhancement; generative adversarial network (GAN); multi-stage GAN; multi-generator GAN; self-attention mechanism

Share and Cite

MDPI and ACS Style

Asiedu Asante, B.K.; Broni-Bediako, C.; Imamura, H. Exploring Multi-Stage GAN with Self-Attention for Speech Enhancement. Appl. Sci. 2023, 13, 9217. https://doi.org/10.3390/app13169217

AMA Style

Asiedu Asante BK, Broni-Bediako C, Imamura H. Exploring Multi-Stage GAN with Self-Attention for Speech Enhancement. Applied Sciences. 2023; 13(16):9217. https://doi.org/10.3390/app13169217

Chicago/Turabian Style

Asiedu Asante, Bismark Kweku, Clifford Broni-Bediako, and Hiroki Imamura. 2023. "Exploring Multi-Stage GAN with Self-Attention for Speech Enhancement" Applied Sciences 13, no. 16: 9217. https://doi.org/10.3390/app13169217

APA Style

Asiedu Asante, B. K., Broni-Bediako, C., & Imamura, H. (2023). Exploring Multi-Stage GAN with Self-Attention for Speech Enhancement. Applied Sciences, 13(16), 9217. https://doi.org/10.3390/app13169217

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop