Next Article in Journal
3D Phased Array Enabling Extended Field of View in Mobile Satcom Applications
Previous Article in Journal
Channel Switching Algorithms for a Robust Networked Control System with a Delay and Packet Errors
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Target Speaker Extraction Using Attention-Enhanced Temporal Convolutional Network

1
School of Computer Science and Technology, Shandong University of Technology, Zibo 255000, China
2
Department of Communication Engineering, National Central University, Taoyuan 32001, Taiwan
3
Department of Computer Science and Information Engineering, Providence University, Taichung 43301, Taiwan
4
Department of Computer Science and Information Engineering, Fu Jen Catholic University, New Taipei City 24205, Taiwan
5
Faculty of Digital Technology, The University of Danang—University of Technology and Education, Danang 550000, Vietnam
6
AI Research Center, Hon Hai Research Institute, New Taipei City 236, Taiwan
7
Department of Computer Science and Information Engineering, National Central University, Taoyuan 32001, Taiwan
*
Authors to whom correspondence should be addressed.
Electronics 2024, 13(2), 307; https://doi.org/10.3390/electronics13020307
Submission received: 25 November 2023 / Revised: 18 December 2023 / Accepted: 28 December 2023 / Published: 10 January 2024
(This article belongs to the Section Artificial Intelligence)

Abstract

When recording conversations, there may be multiple people talking at once. While our human ears can filter out unwanted sounds, this can be challenging for automatic speech recognition (ASR) systems, leading to reduced accuracy. To address this issue, preprocessing mechanisms such as speech separation and targeted speaker extraction are necessary to separate each person’s speech. With the development of deep learning, the quality of separated speech has improved significantly. Our objective is to focus on speaker extraction, which entails implementing a primary system for speech extraction and a secondary subsystem for delivering target information. To accomplish this, we have chosen a temporal convolutional network (TCN) architecture as the foundation of our speech extraction model. A TCN enables convolutional neural networks (CNNs) to manage time series modeling, and it can be constructed in various model lengths. Furthermore, we have integrated attention enhancement into the secondary subsystem to provide the speech extraction model with comprehensive and effective target information, which helps to improve the model’s ability to estimate masks. As a result, the quality of the target speaker extraction will be greatly enhanced with a more precise mask.
Keywords: deep learning; target speaker extraction; temporal convolutional network (TCN); convolutional neural network (CNN); automatic speech recognition (ASR) deep learning; target speaker extraction; temporal convolutional network (TCN); convolutional neural network (CNN); automatic speech recognition (ASR)

Share and Cite

MDPI and ACS Style

Wang, J.-H.; Lai, Y.-T.; Tai, T.-C.; Le, P.T.; Pham, T.; Wang, Z.-Y.; Li, Y.-H.; Wang, J.-C.; Chang, P.-C. Target Speaker Extraction Using Attention-Enhanced Temporal Convolutional Network. Electronics 2024, 13, 307. https://doi.org/10.3390/electronics13020307

AMA Style

Wang J-H, Lai Y-T, Tai T-C, Le PT, Pham T, Wang Z-Y, Li Y-H, Wang J-C, Chang P-C. Target Speaker Extraction Using Attention-Enhanced Temporal Convolutional Network. Electronics. 2024; 13(2):307. https://doi.org/10.3390/electronics13020307

Chicago/Turabian Style

Wang, Jian-Hong, Yen-Ting Lai, Tzu-Chiang Tai, Phuong Thi Le, Tuan Pham, Ze-Yu Wang, Yung-Hui Li, Jia-Ching Wang, and Pao-Chi Chang. 2024. "Target Speaker Extraction Using Attention-Enhanced Temporal Convolutional Network" Electronics 13, no. 2: 307. https://doi.org/10.3390/electronics13020307

APA Style

Wang, J.-H., Lai, Y.-T., Tai, T.-C., Le, P. T., Pham, T., Wang, Z.-Y., Li, Y.-H., Wang, J.-C., & Chang, P.-C. (2024). Target Speaker Extraction Using Attention-Enhanced Temporal Convolutional Network. Electronics, 13(2), 307. https://doi.org/10.3390/electronics13020307

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop