Next Article in Journal
Identifying the Edges of the Optic Cup and the Optic Disc in Glaucoma Patients by Segmentation
Next Article in Special Issue
Diversifying Emotional Dialogue Generation via Selective Adversarial Training
Previous Article in Journal
NMR Magnetometer Based on Dynamic Nuclear-Polarization for Low-Strength Magnetic Field Measurement
Previous Article in Special Issue
Physiological Synchrony Predict Task Performance and Negative Emotional State during a Three-Member Collaborative Task
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Progressive Learning of a Multimodal Classifier Accounting for Different Modality Combinations

Guardian Robot Project, RIKEN, Seika-cho, Kyoto 619-0288, Japan
*
Author to whom correspondence should be addressed.
Sensors 2023, 23(10), 4666; https://doi.org/10.3390/s23104666
Submission received: 31 March 2023 / Revised: 9 May 2023 / Accepted: 9 May 2023 / Published: 11 May 2023
(This article belongs to the Special Issue Advanced-Sensors-Based Emotion Sensing and Recognition)

Abstract

In classification tasks, such as face recognition and emotion recognition, multimodal information is used for accurate classification. Once a multimodal classification model is trained with a set of modalities, it estimates the class label by using the entire modality set. A trained classifier is typically not formulated to perform classification for various subsets of modalities. Thus, the model would be useful and portable if it could be used for any subset of modalities. We refer to this problem as the multimodal portability problem. Moreover, in the multimodal model, classification accuracy is reduced when one or more modalities are missing. We term this problem the missing modality problem. This article proposes a novel deep learning model, termed KModNet, and a novel learning strategy, termed progressive learning, to simultaneously address missing modality and multimodal portability problems. KModNet, formulated with the transformer, contains multiple branches corresponding to different k-combinations of the modality set S. KModNet is trained using a multi-step progressive learning framework, where the k-th step uses a k-modal model to train different branches up to the k-th combination branch. To address the missing modality problem, the training multimodal data is randomly ablated. The proposed learning framework is formulated and validated using two multimodal classification problems: audio-video-thermal person classification and audio-video emotion classification. The two classification problems are validated using the Speaking Faces, RAVDESS, and SAVEE datasets. The results demonstrate that the progressive learning framework enhances the robustness of multimodal classification, even under the conditions of missing modalities, while being portable to different modality subsets.
Keywords: multimodal learning; person classification; emotion classification; missing modality; multimodal portability; sensor fusion multimodal learning; person classification; emotion classification; missing modality; multimodal portability; sensor fusion

Share and Cite

MDPI and ACS Style

John, V.; Kawanishi, Y. Progressive Learning of a Multimodal Classifier Accounting for Different Modality Combinations. Sensors 2023, 23, 4666. https://doi.org/10.3390/s23104666

AMA Style

John V, Kawanishi Y. Progressive Learning of a Multimodal Classifier Accounting for Different Modality Combinations. Sensors. 2023; 23(10):4666. https://doi.org/10.3390/s23104666

Chicago/Turabian Style

John, Vijay, and Yasutomo Kawanishi. 2023. "Progressive Learning of a Multimodal Classifier Accounting for Different Modality Combinations" Sensors 23, no. 10: 4666. https://doi.org/10.3390/s23104666

APA Style

John, V., & Kawanishi, Y. (2023). Progressive Learning of a Multimodal Classifier Accounting for Different Modality Combinations. Sensors, 23(10), 4666. https://doi.org/10.3390/s23104666

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop