Next Article in Journal
FCKDNet: A Feature Condensation Knowledge Distillation Network for Semantic Segmentation
Previous Article in Journal
Unsupervised Anomaly Detection for Intermittent Sequences Based on Multi-Granularity Abnormal Pattern Mining
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Cross-Corpus Speech Emotion Recognition Based on Multi-Task Learning and Subdomain Adaptation

1
College of Information Science and Engineering, Henan University of Technology, Zhengzhou 450001, China
2
Henan Engineering Laboratory of Grain IOT Technology, Henan University of Technology, Zhengzhou 450001, China
3
Key Laboratory of Food Information Processing and Control, Ministry of Education, Henan University of Technology, Zhengzhou 450001, China
*
Author to whom correspondence should be addressed.
Entropy 2023, 25(1), 124; https://doi.org/10.3390/e25010124
Submission received: 26 December 2022 / Revised: 3 January 2023 / Accepted: 4 January 2023 / Published: 7 January 2023
(This article belongs to the Topic Machine and Deep Learning)

Abstract

To solve the problem of feature distribution discrepancy in cross-corpus speech emotion recognition tasks, this paper proposed an emotion recognition model based on multi-task learning and subdomain adaptation, which alleviates the impact on emotion recognition. Existing methods have shortcomings in speech feature representation and cross-corpus feature distribution alignment. The proposed model uses a deep denoising auto-encoder as a shared feature extraction network for multi-task learning, and the fully connected layer and softmax layer are added before each recognition task as task-specific layers. Subsequently, the subdomain adaptation algorithm of emotion and gender features is added to the shared network to obtain the shared emotion features and gender features of the source domain and target domain, respectively. Multi-task learning effectively enhances the representation ability of features, a subdomain adaptive algorithm promotes the migrating ability of features and effectively alleviates the impact of feature distribution differences in emotional features. The average results of six cross-corpus speech emotion recognition experiments show that, compared with other models, the weighted average recall rate is increased by 1.89~10.07%, the experimental results verify the validity of the proposed model.
Keywords: speech emotion recognition; multi-task learning; subdomain adaptation; feature distribution speech emotion recognition; multi-task learning; subdomain adaptation; feature distribution

Share and Cite

MDPI and ACS Style

Fu, H.; Zhuang, Z.; Wang, Y.; Huang, C.; Duan, W. Cross-Corpus Speech Emotion Recognition Based on Multi-Task Learning and Subdomain Adaptation. Entropy 2023, 25, 124. https://doi.org/10.3390/e25010124

AMA Style

Fu H, Zhuang Z, Wang Y, Huang C, Duan W. Cross-Corpus Speech Emotion Recognition Based on Multi-Task Learning and Subdomain Adaptation. Entropy. 2023; 25(1):124. https://doi.org/10.3390/e25010124

Chicago/Turabian Style

Fu, Hongliang, Zhihao Zhuang, Yang Wang, Chen Huang, and Wenzhuo Duan. 2023. "Cross-Corpus Speech Emotion Recognition Based on Multi-Task Learning and Subdomain Adaptation" Entropy 25, no. 1: 124. https://doi.org/10.3390/e25010124

APA Style

Fu, H., Zhuang, Z., Wang, Y., Huang, C., & Duan, W. (2023). Cross-Corpus Speech Emotion Recognition Based on Multi-Task Learning and Subdomain Adaptation. Entropy, 25(1), 124. https://doi.org/10.3390/e25010124

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop