Next Article in Journal
Three-Dimensional Point Cloud Displacement Analysis for Tunnel Deformation Detection Using Mobile Laser Scanning
Previous Article in Journal
Movement of Overlying Strata and Mechanical Responses of Shallow Buried Gas Pipelines in Coal Mining Areas
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Speech Emotion Recognition Model Based on Joint Modeling of Discrete and Dimensional Emotion Representation

by
John Lorenzo Bautista
1,2 and
Hyun Soon Shin
1,2,*
1
Emotion Recognition IoT Research Section, Hyper-Connected Communication Research Laboratory, Electronic and Telecommunications Research Institute (ETRI), Daejeon 34129, Republic of Korea
2
ETRI School of Artificial Intelligence, Korea University of Science and Technology, Daejeon 34113, Republic of Korea
*
Author to whom correspondence should be addressed.
Appl. Sci. 2025, 15(2), 623; https://doi.org/10.3390/app15020623
Submission received: 3 December 2024 / Revised: 6 January 2025 / Accepted: 8 January 2025 / Published: 10 January 2025
(This article belongs to the Section Computing and Artificial Intelligence)

Abstract

This paper introduces a novel joint model architecture for Speech Emotion Recognition (SER) that integrates both discrete and dimensional emotional representations, allowing for the simultaneous training of classification and regression tasks to improve the comprehensiveness and interpretability of emotion recognition. By employing a joint loss function that combines categorical and regression losses, the model ensures balanced optimization across tasks, with experiments exploring various weighting schemes using a tunable parameter to adjust task importance. Two adaptive weight balancing schemes, Dynamic Weighting and Joint Weighting, further enhance performance by dynamically adjusting task weights based on optimization progress and ensuring balanced emotion representation during backpropagation. The architecture employs parallel feature extraction through independent encoders, designed to capture unique features from multiple modalities, including Mel-frequency Cepstral Coefficients (MFCC), Short-term Features (STF), Mel-spectrograms, and raw audio signals. Additionally, pre-trained models such as Wav2Vec 2.0 and HuBERT are integrated to leverage their robust latent features. The inclusion of self-attention and co-attention mechanisms allows the model to capture relationships between input modalities and interdependencies among features, further improving its interpretability and integration capabilities. Experiments conducted on the IEMOCAP dataset using a leave-one-subject-out approach demonstrate the model’s effectiveness, with results showing a 1–2% accuracy improvement over classification-only models. The optimal configuration, incorporating the joint architecture, dynamic weighting, and parallel processing of multimodal features, achieves a weighted accuracy of 72.66%, an unweighted accuracy of 73.22%, and a mean Concordance Correlation Coefficient (CCC) of 0.3717. These results validate the effectiveness of the proposed joint model architecture and adaptive balancing weight schemes in improving SER performance.
Keywords: adaptive weight balancing scheme; affective computing; dimensional emotion representation; discrete emotion representation; joint model architecture; Speech Emotion Recognition (SER) adaptive weight balancing scheme; affective computing; dimensional emotion representation; discrete emotion representation; joint model architecture; Speech Emotion Recognition (SER)

Share and Cite

MDPI and ACS Style

Bautista, J.L.; Shin, H.S. Speech Emotion Recognition Model Based on Joint Modeling of Discrete and Dimensional Emotion Representation. Appl. Sci. 2025, 15, 623. https://doi.org/10.3390/app15020623

AMA Style

Bautista JL, Shin HS. Speech Emotion Recognition Model Based on Joint Modeling of Discrete and Dimensional Emotion Representation. Applied Sciences. 2025; 15(2):623. https://doi.org/10.3390/app15020623

Chicago/Turabian Style

Bautista, John Lorenzo, and Hyun Soon Shin. 2025. "Speech Emotion Recognition Model Based on Joint Modeling of Discrete and Dimensional Emotion Representation" Applied Sciences 15, no. 2: 623. https://doi.org/10.3390/app15020623

APA Style

Bautista, J. L., & Shin, H. S. (2025). Speech Emotion Recognition Model Based on Joint Modeling of Discrete and Dimensional Emotion Representation. Applied Sciences, 15(2), 623. https://doi.org/10.3390/app15020623

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop