Next Article in Journal
Optimal Coordination of Over-Current Relays in Microgrids Using Principal Component Analysis and K-Means
Next Article in Special Issue
Multi-Modal Emotion Recognition Using Speech Features and Text-Embedding
Previous Article in Journal
A Novel Hybrid Deep Learning Model for Detecting COVID-19-Related Rumors on Social Media Based on LSTM and Concatenated Parallel CNNs
Previous Article in Special Issue
Emotion Identification in Movies through Facial Expression Recognition
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Deep Multimodal Emotion Recognition on Human Speech: A Review

by
Panagiotis Koromilas
* and
Theodoros Giannakopoulos
Institute of Informatics and Telecommunications, National Center for Scientific Research—Demokritos, 15310 Athens, Greece
*
Author to whom correspondence should be addressed.
Appl. Sci. 2021, 11(17), 7962; https://doi.org/10.3390/app11177962
Submission received: 26 July 2021 / Revised: 20 August 2021 / Accepted: 26 August 2021 / Published: 28 August 2021
(This article belongs to the Special Issue Pattern Recognition in Multimedia Signal Analysis)

Abstract

This work reviews the state of the art in multimodal speech emotion recognition methodologies, focusing on audio, text and visual information. We provide a new, descriptive categorization of methods, based on the way they handle the inter-modality and intra-modality dynamics in the temporal dimension: (i) non-temporal architectures (NTA), which do not significantly model the temporal dimension in both unimodal and multimodal interaction; (ii) pseudo-temporal architectures (PTA), which also assume an oversimplification of the temporal dimension, although in one of the unimodal or multimodal interactions; and (iii) temporal architectures (TA), which try to capture both unimodal and cross-modal temporal dependencies. In addition, we review the basic feature representation methods for each modality, and we present aggregated evaluation results on the reported methodologies. Finally, we conclude this work with an in-depth analysis of the future challenges related to validation procedures, representation learning and method robustness.
Keywords: multimodal emotion recognition; multimodal temporal learning; multimodal signal processing; affective computing; speech emotion recognition multimodal emotion recognition; multimodal temporal learning; multimodal signal processing; affective computing; speech emotion recognition

Share and Cite

MDPI and ACS Style

Koromilas, P.; Giannakopoulos, T. Deep Multimodal Emotion Recognition on Human Speech: A Review. Appl. Sci. 2021, 11, 7962. https://doi.org/10.3390/app11177962

AMA Style

Koromilas P, Giannakopoulos T. Deep Multimodal Emotion Recognition on Human Speech: A Review. Applied Sciences. 2021; 11(17):7962. https://doi.org/10.3390/app11177962

Chicago/Turabian Style

Koromilas, Panagiotis, and Theodoros Giannakopoulos. 2021. "Deep Multimodal Emotion Recognition on Human Speech: A Review" Applied Sciences 11, no. 17: 7962. https://doi.org/10.3390/app11177962

APA Style

Koromilas, P., & Giannakopoulos, T. (2021). Deep Multimodal Emotion Recognition on Human Speech: A Review. Applied Sciences, 11(17), 7962. https://doi.org/10.3390/app11177962

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop