Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (289)

Search Parameters:
Keywords = multi-modal emotion recognition

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
18 pages, 8848 KB  
Article
Multi-Objective Performance-Cost Optimization of Multimodal EEG-Eye Tracking Systems for Emotion Recognition
by Eda Dagdevir
Electronics 2026, 15(18), 4182; https://doi.org/10.3390/electronics15184182 - 15 Sep 2026
Abstract
Multimodal emotion recognition systems based on electroencephalography (EEG) and eye tracking (ET) provide complementary information about neural and visual responses; however, practical deployment requires balancing classification performance and computational cost. This study investigates this trade-off by systematically evaluating 60 feature representation configurations generated [...] Read more.
Multimodal emotion recognition systems based on electroencephalography (EEG) and eye tracking (ET) provide complementary information about neural and visual responses; however, practical deployment requires balancing classification performance and computational cost. This study investigates this trade-off by systematically evaluating 60 feature representation configurations generated from different EEG channel regions, frequency bands, feature types, and signal modalities. Support Vector Machine (SVM), Random Forest (RF), and Artificial Neural Network (ANN) classifiers were evaluated under subject-independent leave-one-subject-out (LOSO) cross-validation, resulting in 180 realizable classifier-feature configuration systems. Median Macro-F1 was used as the primary performance objective, while median classifier inference time was used as the computational-cost objective. A multi-stage selection strategy integrating Δ-based near-optimal filtering and Pareto dominance analysis was applied to identify performance-efficient systems. The highest median Macro-F1 (0.6474) was achieved by an ANN using temporal delta-band PSD features combined with ET information, with a median classifier inference time of 0.0057 s. This system remained the final selected system across Δ values of 0.01, 0.02, and 0.03. An ET-only ANN baseline achieved a median Macro-F1 of 0.6265, indicating a modest improvement when temporal delta-band EEG information was added. These findings demonstrate that feature representation, modality composition, classifier choice, and computational cost should be considered jointly when designing subject-independent multimodal emotion recognition systems. Full article
Show Figures

Figure 1

30 pages, 4184 KB  
Article
Spatial-Frequency Hypergraph Neural Network for EEG-fNIRS Emotion Recognition
by Haifeng Li, Xueying Zhang, Guijun Chen, Yaru Zhou, Ying Sun and Lixia Huang
Brain Sci. 2026, 16(9), 962; https://doi.org/10.3390/brainsci16090962 - 11 Sep 2026
Viewed by 135
Abstract
Background/Objectives: Hybrid EEG-fNIRS emotion recognition aims to accurately identify an individual’s emotional state by analyzing neurophysiological signals and constitutes an important research direction in affective brain–computer interfaces and human–computer interaction. In recent years, EEG-fNIRS emotion recognition has advanced from handcrafted feature extraction and [...] Read more.
Background/Objectives: Hybrid EEG-fNIRS emotion recognition aims to accurately identify an individual’s emotional state by analyzing neurophysiological signals and constitutes an important research direction in affective brain–computer interfaces and human–computer interaction. In recent years, EEG-fNIRS emotion recognition has advanced from handcrafted feature extraction and shallow fusion to deep learning and graph-based modeling. However, most existing methods rely on predefined fixed frequency-band partitioning and second-order graph structures that only support pairwise connections, making it difficult to accommodate inter-subject frequency variability and to characterize high-order brain network relationships such as multi-channel synergistic activation within a frequency band and cross-frequency coupling. Methods: To address these issues, this paper proposes an EEG-fNIRS emotion recognition framework based on a Spatial-Frequency Hypergraph Neural Network (SF-HGNN). First, a Dynamic Frequency Band Decomposition module is designed to achieve adaptive optimization of the EEG and fNIRS frequency bands; second, a Multi-scale Temporal Convolution module extracts temporal features at different time scales; third, a Spatial-Frequency Adaptive Hypergraph Convolution module is constructed to model intra-band cross-channel spatial synergy and channel-wise cross-frequency coupling; and finally, a Cross-Modal Attention Fusion mechanism achieves high-order interaction between the complementary information of the two modalities. The proposed method was validated on the public ENTER dataset comprising 50 participants and four emotion categories (sadness, happiness, fear, and calm). Results: Experimental results show that SF-HGNN achieves accuracies of 82.94% and 68.85% in subject-dependent and subject-independent experiments, respectively; ablation studies and visualization analyses further verify the effectiveness of each module and the interpretability of the model. Conclusions: Future work will focus on validation with larger-scale data and improving cross-subject domain generalization. Full article
(This article belongs to the Section Neurotechnology and Neuroimaging)
Show Figures

Figure 1

25 pages, 15474 KB  
Article
Emotion-Aware Virtual Reality Through Multimodal ECG and Postural Fusion
by Juan Benavides, Mayra Carrión-Toro, Cindy López, David Morales-Martínez, Marco Santórum and Patricia Acosta-Vargas
Sensors 2026, 26(18), 5726; https://doi.org/10.3390/s26185726 - 9 Sep 2026
Viewed by 350
Abstract
Immersive Virtual Reality (VR) environments are increasingly adopted in clinical psychology and stress-management contexts; however, their therapeutic effectiveness depends on the system’s ability to understand and dynamically respond to users’ affective states. Emotional regulation plays a key role in psychological resilience, directly influencing [...] Read more.
Immersive Virtual Reality (VR) environments are increasingly adopted in clinical psychology and stress-management contexts; however, their therapeutic effectiveness depends on the system’s ability to understand and dynamically respond to users’ affective states. Emotional regulation plays a key role in psychological resilience, directly influencing stress coping mechanisms, cognitive performance, and overall mental well-being. Despite recent advances, automatic recognition of scenario-associated affective conditions in VR remains challenging because head-mounted displays occlude facial features. This study proposes a VR-based serious game for affective training and regulation, in which users interact with goal-oriented scenarios targeting fear, anger, and joy. We introduce a multimodal affective computing model to objectively assess users’ emotional responses by integrating electrocardiogram (ECG) signals and posture-based features extracted through computer vision. An early-fusion architecture combined with a Long Short-Term Memory (LSTM) network captures temporal dependencies in synchronized multimodal data. We established a controlled experimental framework using immersive VR scenarios, enabling the collection of synchronized physiological and behavioral data from a cohort of 20 healthy adult participants. The proposed model was evaluated under a strict Leave-One-Subject-Out (LOSO) cross-validation scheme across independent subjects, achieving a robust inter-subject accuracy of 78.94%±10.77% and a global macro F1-score of 0.635, demonstrating strong generalization to entirely unseen users without data leakage. Furthermore, the system maintained an outstanding balance in detecting active emotional states (recall > 80.0% for fear, anger, and joy). Additionally, subjective evaluations using the PANAS and SGU questionnaires confirmed the coherence between detected and perceived emotional states, as well as the system’s high usability. The results suggest the potential viability of combining immersive environments and multimodal affective computing to explore the technical feasibility of adaptive frameworks that could eventually translate into healthcare contexts. This work may contribute to the development of intelligent digital health technologies by providing a foundation for responsive VR systems that can monitor emotional regulation and are fully aligned with sustainable well-being ecosystems (SDG 3: Good Health and Well-being and SDG 10: Reduced Inequalities). Full article
(This article belongs to the Special Issue Advanced Signal Processing for Affective Computing)
Show Figures

Figure 1

51 pages, 4421 KB  
Systematic Review
Affective Computing Approaches in Child–Robot Interaction: A Systematic Review and Taxonomy
by Sandra Cano, Juan Pablo Vásconez, Kiara Villarroel, Juan Carlos Geraldo and Sergio Albiol-Pérez
Sensors 2026, 26(18), 5721; https://doi.org/10.3390/s26185721 - 9 Sep 2026
Viewed by 238
Abstract
Affective computing has become increasingly relevant in child–robot interaction (CRI), particularly in social robotics, emotion recognition, engagement assessment, and autism-related interventions. This systematic review with a critical and integrative synthesis analyzes 105 included studies to examine how affect is sensed, represented, processed, expressed, [...] Read more.
Affective computing has become increasingly relevant in child–robot interaction (CRI), particularly in social robotics, emotion recognition, engagement assessment, and autism-related interventions. This systematic review with a critical and integrative synthesis analyzes 105 included studies to examine how affect is sensed, represented, processed, expressed, and evaluated in CRI. The literature search was conducted in IEEE Xplore, Web of Science, Scopus, and PubMed, following a systematic screening process guided by the review objectives. A descriptive and structured narrative synthesis was conducted considering publication characteristics, robot platform and morphology, target population, sensing modalities and observed affect-relevant features, affective constructs and representation models, computational and control mechanisms, robot affective expression, evaluation strategies, and remaining research gaps. The findings show a strong emphasis on ASD-related contexts, visually observable and behavioral features, facial emotion recognition, body movement analysis, and engagement assessment. The review also identifies important limitations, including reliance on camera-based affect recognition, comparatively limited use of physiological and other complementary sensing modalities, unclear alignment between robot roles and interaction strategies, insufficient reporting of robot emotional expressiveness and control mechanisms, and limited attention to explainability, data governance, and long-term ethical implications. Based on these findings, an integrative taxonomy of affective computing in CRI is proposed, comprising six interconnected dimensions: interaction context; sensing modalities and observed features; affective constructs and representation models; computational and control mechanisms; robot affective expression; and evaluation and adaptation strategies. Rather than treating these dimensions as entirely novel categories, the taxonomy consolidates and extends previously fragmented classifications into a child-centered representation of the affective interaction process. Overall, this review argues that affective CRI should move beyond automatic emotion recognition toward multimodal, embodied, developmentally appropriate, explainable, and ethically grounded robot interaction. Full article
(This article belongs to the Special Issue Sensors and Sensing Technologies for Social Robots)
Show Figures

Figure 1

32 pages, 1400 KB  
Article
KTU-MEDAFE: A Newly Developed Multimodal Dataset for Emotion Recognition Using EEG–Speech Decision-Level Fusion
by Bahar Hatipoglu Yilmaz, Betul Mumcu and Busra Ozkellekci
Sensors 2026, 26(17), 5608; https://doi.org/10.3390/s26175608 - 3 Sep 2026
Viewed by 367
Abstract
One of the central challenges in affective computing is achieving reliable emotion recognition for natural and effective human–computer interaction. In this study, we introduce KTU-MEDAFE (Karadeniz Technical University Multimodal Emotion Dataset using Audio, Facial Images, and EEG), a newly developed multimodal dataset containing [...] Read more.
One of the central challenges in affective computing is achieving reliable emotion recognition for natural and effective human–computer interaction. In this study, we introduce KTU-MEDAFE (Karadeniz Technical University Multimodal Emotion Dataset using Audio, Facial Images, and EEG), a newly developed multimodal dataset containing synchronized EEG signals, speech recordings, and facial videos collected from 40 participants under controlled emotional elicitation conditions. The dataset includes Turkish emotional speech and two recording sessions conducted on separate days, providing a language-specific resource that supports both participant-dependent baseline evaluation and future session-separated analysis. Although KTU-MEDAFE comprises three modalities, the present study focuses on EEG and speech integration. EEG and speech recordings meeting signal quality criteria were transformed into image representations using the Angle–Amplitude Graph (AAG) method and classified using transfer learning with ResNet-50 and GoogLeNet architectures. To exploit complementary information across modalities, multiple decision-level fusion strategies were evaluated. Experimental findings show that multimodal fusion provides higher average classification performance than unimodal EEG and speech models across the evaluated binary emotion pairs, with performance varying according to subject, fusion strategy, and model architecture. Overall, the results support the potential benefit of combining EEG and speech for multimodal emotion recognition while highlighting substantial subject-dependent variability in classification performance. Full article
(This article belongs to the Special Issue EEG Signal Processing Techniques and Applications—3rd Edition)
Show Figures

Graphical abstract

30 pages, 7584 KB  
Article
Parameter-Efficient Audio-Visual Dynamic Facial Expression Recognition with Mamba Fusion Adapters and Frame-Level Feature Arrangement
by Kangbo Ning, Shanshan Gao, Zhaoqiang Xia, Dong Huang and Lei Li
Sensors 2026, 26(17), 5384; https://doi.org/10.3390/s26175384 - 26 Aug 2026
Viewed by 312
Abstract
Dynamic Facial Expression Recognition (DFER) has recently attracted significant interest due to its vital role in enabling empathetic and human-compatible technologies. Developing models that remain robust under in-the-wild variability is a key motivation for DFER research and its practical applications. Improving models through [...] Read more.
Dynamic Facial Expression Recognition (DFER) has recently attracted significant interest due to its vital role in enabling empathetic and human-compatible technologies. Developing models that remain robust under in-the-wild variability is a key motivation for DFER research and its practical applications. Improving models through multimodal learning, which leverages audio and video data for richer, complementary representations, is one promising direction. However, existing methods still rely heavily on modality-specific encoders and coarse-grained content-level alignment, which hinders their ability to capture fine-grained emotional semantics and dynamic cross-modal interactions. To address this, we adopt parameter-efficient fine-tuning (PEFT) to facilitate audio-visual interaction. This strategy offers key advantages: (1) freezing parameters preserves upstream pretrained knowledge, ensuring that the model focuses solely on learning modules for audio-visual interaction and modal fusion; (2) a Mamba Fusion Adapter (MFAdapter) is inserted at each encoder layer to perform causal, audio-conditioned fusion over a frame-aligned token sequence, enabling efficient multi-level cross-modal injection with linear complexity; and (3) a Frame-level Feature Arrangement (FFA) strategy is introduced as a deterministic index-based arrangement scheme that arranges audio tokens at the video-frame rate, providing a frame-indexed temporal prior that supports the causal scan in MFAdapter; FFA reduces, but does not eliminate, coarse segment-level mismatch and is not claimed as verified frame-level synchronization. Notably, our method achieves competitive performance on DFEW and MAFW while updating 2.7% (4.7 million) of the model parameters, with the best WAR of 58.70% on MAFW among all compared methods and 76.62%/65.25% (WAR/UAR) on DFEW under the official five-fold cross-validation protocols. Full article
(This article belongs to the Special Issue Advanced Signal Processing for Affective Computing)
Show Figures

Figure 1

19 pages, 12705 KB  
Article
RMDD: Raspberry Pi-Based Multimodal Dangerous Driving Behavior Detection
by Yunsheng Liang, Haiyan Kang and Huan Zhong
Electronics 2026, 15(17), 3828; https://doi.org/10.3390/electronics15173828 - 26 Aug 2026
Viewed by 283
Abstract
With the continuous growth in motor vehicle ownership, traffic safety risks caused by dangerous driving behaviors remain an important concern. This paper presents RMDD, a Raspberry Pi-based multimodal dangerous driving monitoring feasibility prototype using YOLO26 visual models, personalized facial fatigue estimation, emotion2vec-based speech [...] Read more.
With the continuous growth in motor vehicle ownership, traffic safety risks caused by dangerous driving behaviors remain an important concern. This paper presents RMDD, a Raspberry Pi-based multimodal dangerous driving monitoring feasibility prototype using YOLO26 visual models, personalized facial fatigue estimation, emotion2vec-based speech recognition, and hierarchical reliability-gated BFV fusion (HRG-BFV). HRG-BFV treats object detection as positive-only support for behavior evidence and scales its contribution by development-fold reliability, thereby avoiding hard rejection when the object detector fails. On locked module tests, behavior classification achieved 96.77% Top-1 accuracy, auxiliary detection reached 0.916 mAP@0.5, personalized facial calibration reduced the false-positive rate from 0.877 to 0.211, and speaker-disjoint speech recognition achieved 0.890 accuracy at 0 dB vehicle noise. The historical fixed BFV baseline achieved 0.601 balanced accuracy and 0.506 macro-F1 on P03–P08 (96 clips). In a stricter three-fold leave-one-external-participant-out evaluation on P04/P05/P08 (48 clips), HRG-BFV achieved 0.726 balanced accuracy, 0.562 macro-F1, 0.833 specificity, and 0.677 ROC-AUC, versus 0.655, 0.469, 0.833, and 0.595 for the historical frozen hard gate on the same cohort. Target support reliability was low (0.097–0.194), so the adaptive term appropriately reverted toward behavior evidence rather than producing an unsupported object detection gain. A sealed P08 behavior test exposed substantial cross-subject degradation (0.320 frame accuracy). A cooled 30 min Raspberry Pi 5 run completed without thermal throttling at 0.927 processing windows/s. The evidence supports controlled prototype feasibility but not population-level or on-road generalization. Full article
Show Figures

Figure 1

34 pages, 2453 KB  
Article
Reliability-Aware Cross-Modal Learning Behavior Sensing for Student Cognitive Bias Recognition and Teaching-Oriented Psychological Risk Warning
by Luo Xu, Chenlu Jiang, Moxian Lin and Yan Zhan
Sensors 2026, 26(16), 5286; https://doi.org/10.3390/s26165286 - 20 Aug 2026
Viewed by 361
Abstract
With the development of smart classrooms and digital learning platforms, multimodal learning behavior data provide a new sensing basis for understanding students’ cognitive states and psychological risk warnings. However, existing educational data mining methods mainly focus on performance prediction, dropout warning, or surface-level [...] Read more.
With the development of smart classrooms and digital learning platforms, multimodal learning behavior data provide a new sensing basis for understanding students’ cognitive states and psychological risk warnings. However, existing educational data mining methods mainly focus on performance prediction, dropout warning, or surface-level emotion recognition, while continuous and interpretable modeling of deeper cognitive biases and related psychological risks remains insufficient. To address this issue, we propose MLBS-Net, a multimodal learning behavior sensing network for teaching feedback that jointly models students’ textual expressions, behavioral sequences, classroom interactions, and psychological auxiliary signals. MLBS-Net integrates theory-guided textual cognitive bias encoding, temporal behavioral state modeling, and reliability-aware cross-modal fusion to capture psychologically interpretable cognitive patterns, characterize dynamic learning-state changes, and adaptively integrate multimodal information according to data quality and task contribution while providing interpretable feedback for teachers. Experimental results show that MLBS-Net achieves a Macro-F1 of 0.855 for cognitive bias recognition and an AUC of 0.891 for psychological risk warning, outperforming traditional machine learning, unimodal deep learning, and standard multimodal methods. Ablation results further support the effectiveness of theory-guided semantic encoding, temporal behavioral modeling, reliability estimation, and multitask learning. These findings demonstrate that MLBS-Net can jointly characterize cognitive biases and potential psychological risks from multisource learning behaviors, providing a feasible approach for learning-state sensing, risk warning, and interpretable teaching support in smart education. Full article
(This article belongs to the Special Issue Artificial Intelligence-Driven Sensing)
Show Figures

Figure 1

27 pages, 11115 KB  
Article
Learning Adaptive Cross-Modal Interactions for Multimodal Sentiment Analysis
by Chuhan Cheng, Hangcheng Wu, Junqiao Wang and Yuqi Ouyang
Data 2026, 11(8), 209; https://doi.org/10.3390/data11080209 - 20 Aug 2026
Viewed by 385
Abstract
Multimodal sentiment analysis aims to integrate heterogeneous textual, visual, and acoustic information for effective emotion understanding. However, existing methods often suffer from insufficient cross-modal interaction modeling, limited adaptability in multimodal fusion, and inadequate suppression of modality-specific noise under complex conversational scenarios. To address [...] Read more.
Multimodal sentiment analysis aims to integrate heterogeneous textual, visual, and acoustic information for effective emotion understanding. However, existing methods often suffer from insufficient cross-modal interaction modeling, limited adaptability in multimodal fusion, and inadequate suppression of modality-specific noise under complex conversational scenarios. To address these challenges, this paper proposes a framework for learning adaptive cross-modal interactions for multimodal sentiment analysis. The proposed framework consists of three stages: modality-aware preprocessing, heterogeneous representation learning, and adaptive multimodal fusion. First, a unified preprocessing strategy is designed to improve cross-modal consistency through textual normalization, speaker-aware visual alignment, and utterance-level acoustic representation enhancement. Second, modality-specific encoders are constructed to capture complementary semantic, spatial, and utterance-level acoustic characteristics from textual, visual, and acoustic modalities, respectively. Third, an adaptive fusion framework is introduced to explicitly model cross-modal interactions, dynamically estimate the importance of different modality combinations, and further calibrate discriminative feature channels through channel attention. By jointly performing modality-level interaction learning and channel-wise feature refinement, the proposed framework effectively enhances multimodal representation capability for sentiment classification. Extensive experiments conducted on the CMU-MOSI and MELD benchmark datasets demonstrate that our framework consistently outperforms previous methods. In particular, the proposed model achieves 90.27% accuracy and 90.26% F1-score on CMU-MOSI, together with 66.57% accuracy and 66.21% F1-score on MELD. Additional ablation studies and qualitative analyses further validate the effectiveness of the proposed preprocessing strategy, modality-specific representation learning, and adaptive fusion mechanism. Full article
(This article belongs to the Special Issue Vision-Based AI in the Real World: Data, Robustness and Deployment)
Show Figures

Figure 1

22 pages, 6137 KB  
Article
Intact Neural and Behavioral Processing of Vocal Emotional Expressions in Men with Autism
by Silke Vos, Rowena Van den Broeck, Diego Ruiz Callejo, Olivier Collignon and Bart Boets
Brain Sci. 2026, 16(8), 876; https://doi.org/10.3390/brainsci16080876 - 18 Aug 2026
Viewed by 259
Abstract
Background/Objectives. Human voices convey critical socio-affective information, including emotional states. Although autism has frequently been associated with difficulties in processing vocal emotional cues, findings remain inconsistent, particularly in adults. This study investigated neural and behavioral sensitivity to vocal emotion expressions in autistic adults [...] Read more.
Background/Objectives. Human voices convey critical socio-affective information, including emotional states. Although autism has frequently been associated with difficulties in processing vocal emotional cues, findings remain inconsistent, particularly in adults. This study investigated neural and behavioral sensitivity to vocal emotion expressions in autistic adults using an objective auditory frequency-tagging EEG paradigm. Methods. Twenty-five autistic adult men and 25 age- and IQ-matched non-autistic men completed an auditory frequency-tagging EEG task and an auditory and multimodal emotion-recognition assessment. During EEG recording, neutral vocal utterances were presented at 4 Hz, with emotional utterances (fear, anger, happiness, or sadness) inserted every third stimulus, generating an oddball frequency of 1.333 Hz indexing vocal emotion discrimination. Results. No significant group differences were observed in neural or behavioral measures of emotion processing. Robust oddball EEG responses were present in both groups, indicating automatic discrimination of emotional from neutral vocalizations. Fearful and angry vocalizations elicited the strongest neural responses. On the behavioral task, autistic and non-autistic participants showed comparable performance in the auditory modality as well as in the visual and audiovisual modalities, with auditory emotion recognition being the most challenging condition for both groups. Conclusions. These findings provide converging neural and behavioral evidence for intact vocal emotion processing in autistic adult men and are consistent with the view that socio-affective processing differences may attenuate across development. Auditory frequency-tagging EEG shows promise as a sensitive tool for studying individual differences in socio-affective processing. Full article
Show Figures

Figure 1

35 pages, 9123 KB  
Article
Accurate and Robust Multimodal Emotion Recognition for Human–Robot Interaction via Dynamic Graph Learning with Pairwise Cross-Modal Alignment
by Xinyang Zhou, Jiahao Wu, Hongming Xu, Jinghan Mei, Zeyang Chen, Junxiong Zhang, Yitong Chen, Yanrui Jin, Chengliang Liu and Chenggang Yuan
Big Data Cogn. Comput. 2026, 10(8), 276; https://doi.org/10.3390/bdcc10080276 - 18 Aug 2026
Viewed by 429
Abstract
Multimodal emotion recognition in conversation (MERC) aims to identify the emotions in each utterance by modeling textual, acoustic, and visual evidence. Compared with unimodal emotion recognition in conversation (ERC), MERC can leverage complementary textual, acoustic, and visual information to support more accurate and [...] Read more.
Multimodal emotion recognition in conversation (MERC) aims to identify the emotions in each utterance by modeling textual, acoustic, and visual evidence. Compared with unimodal emotion recognition in conversation (ERC), MERC can leverage complementary textual, acoustic, and visual information to support more accurate and consistent emotion inference. However, coordinating intramodal contextual dependencies, cross-modal alignment, and temporal affective dynamics within a unified framework in MERC is challenging. Existing solutions have advanced MERC through contextual modeling, multimodal fusion, and graph-based reasoning, but they still often rely on static relational assumptions or stage-wise coordination of modalities. This limits their ability to jointly model fine-grained relations, selective cross-modal interactions, and dynamic changes in emotion. To address these issues, we propose DGL-PCA (dynamic graph learning with pairwise cross-modal alignment), a dynamic graph-based framework for MERC. The model combines time-aware relation construction, dynamic time-aware heterogeneous graph modeling, and pairwise cross-modal alignment to improve prediction accuracy. This coordinates temporal affective dynamics, structured dialogue context, and multimodal interaction more explicitly than conventional coarse fusion or static graph formulations. Extensive experiments on IEMOCAP and CMU-MOSEI show that DGL-PCA consistently improves weighted F1 by 1.08–19.93% across all reproduced baselines. It achieves 70.02% and 83.91% weighted F1 on the IEMOCAP 6-way and 4-way settings, respectively, and 44.93% and 84.01% weighted F1 on the CMU-MOSEI 7-way and 2-way settings, respectively. Utterance-masking results demonstrated the robustness of the proposed method under dynamic emotional changes. In a 79-utterance IEMOCAP dialogue with 39 adjacent emotion transitions and up to nine transitions within a 15-utterance window, average weighted F1 slightly decreases by 1.87% in the case of any missing utterance, indicating long-context prediction stability under frequent emotion shifts. The proposed model paves the way for developing human-level emotion understanding capability of robots. Full article
(This article belongs to the Special Issue Multimodal Deep Learning and Its Applications)
Show Figures

Figure 1

18 pages, 1133 KB  
Article
Bimodal Speech Emotion Recognition Using a Hybrid CNN-LSTM Architecture with Sentiment Fusion
by Tze-Syn Yap and Lee-Yeng Ong
Future Internet 2026, 18(8), 421; https://doi.org/10.3390/fi18080421 - 10 Aug 2026
Viewed by 265
Abstract
Speech emotion recognition (SER) is a fundamental task in affective computing; however, traditional unimodal approaches often struggle to capture the complex emotional cues present in spontaneous conversational speech. Bimodal frameworks that integrate acoustic and textual information have therefore emerged to provide complementary semantic [...] Read more.
Speech emotion recognition (SER) is a fundamental task in affective computing; however, traditional unimodal approaches often struggle to capture the complex emotional cues present in spontaneous conversational speech. Bimodal frameworks that integrate acoustic and textual information have therefore emerged to provide complementary semantic and acoustic representations. This study proposes a bimodal SER framework based on a hybrid convolutional neural network–long short-term memory (CNN–LSTM) architecture. Using the Multimodal EmotionLines Dataset (MELD), the framework combines temporal acoustic features, statistical acoustic features, and predicted textual sentiment. Experimental results indicate that the proposed model achieves reliable recognition of majority emotion classes but exhibits limited performance on underrepresented minority classes due to severe class imbalance. To better understand the contribution of each modality, feature sufficiency and feature necessity analyses were conducted. Furthermore, an evaluation of alternative fusion strategies showed that the expressive attention networks did not provide meaningful performance improvements over simple feature concatenation. These findings suggest that class imbalance, rather than fusion complexity, remains the primary limitation in conversational SER, highlighting the importance of addressing data imbalance before pursuing more sophisticated multimodal architectures. Full article
(This article belongs to the Special Issue Artificial Intelligence (AI) and Natural Language Processing (NLP))
Show Figures

Graphical abstract

23 pages, 4458 KB  
Article
MSA-CNN: A Multi-Scale Attention Convolutional Neural Network for fNIRS-Based Emotion Recognition
by Deping Huang, Xiu Zhang, Ye Li, Jingfu Wu and Youzhi Yue
Biosensors 2026, 16(8), 434; https://doi.org/10.3390/bios16080434 - 9 Aug 2026
Viewed by 347
Abstract
Functional near-infrared spectroscopy (fNIRS) has attracted increasing attention in affective brain–computer interface research due to its non-invasive nature, portability, and robustness to motion artifacts. However, substantial inter-subject variability in neural responses remains a major challenge for subject-independent emotion recognition. To address this issue, [...] Read more.
Functional near-infrared spectroscopy (fNIRS) has attracted increasing attention in affective brain–computer interface research due to its non-invasive nature, portability, and robustness to motion artifacts. However, substantial inter-subject variability in neural responses remains a major challenge for subject-independent emotion recognition. To address this issue, this work presents an effective integration of multi-scale temporal convolution and dual-attention mechanisms for subject-independent fNIRS emotion recognition evaluated under the leave-one-subject-out protocol within a single dataset. The proposed framework employs multi-scale temporal convolutions to capture hemodynamic characteristics at different temporal resolutions and incorporates channel and temporal attention mechanisms to adaptively emphasize informative brain regions and critical temporal segments. Experiments were conducted on both a self-collected fNIRS emotion dataset and the publicly available ENTER dataset using the Leave-One-Subject-Out (LOSO) evaluation protocol. On the self-collected dataset, MSA-CNN achieved an accuracy of 65.06 ± 7.10% with an F1-score of 0.605. On the ENTER dataset, the proposed model obtained an accuracy of 68.91% and an F1-score of 0.621, outperforming conventional machine learning approaches and several representative deep learning baselines. Ablation studies further demonstrated the positive contributions of both the multi-scale convolutional structure and the dual-attention mechanism. Experimental results on both the self-collected and ENTER datasets demonstrate that the proposed MSA-CNN achieves competitive emotion recognition performance under the LOSO protocol. Class-wise evaluation using precision, recall, and the F1-score further provides a comprehensive assessment of the model’s classification behavior. These results indicate the effectiveness of the proposed framework for cross-subject fNIRS-based emotion recognition under the current experimental settings. The results indicate that multi-scale temporal feature learning combined with attention mechanisms can effectively enhance fNIRS-based emotion recognition performance and provides a promising framework for within-dataset cross-subject evaluation in fNIRS-based emotion recognition. Future work will focus on expanding the subject population, conducting cross-dataset train–test evaluations, and incorporating multimodal neural signals to further improve robustness and generalization. Full article
(This article belongs to the Special Issue Applications of AI in Non-Invasive Biosensing Technologies)
Show Figures

Figure 1

27 pages, 4742 KB  
Article
PRISM-MTL: Inter-Modal Selective Multi-Task Learning for Assistive Driving Perception
by Minjun Kim and Gyuho Choi
Mathematics 2026, 14(15), 2812; https://doi.org/10.3390/math14152812 - 5 Aug 2026
Viewed by 289
Abstract
Advanced driver assistance systems (ADAS) require a comprehensive understanding of multiple tasks related to the physical and mental states of drivers and traffic situations. Existing ADAS studies perform driver emotion recognition (DER), driver behavior recognition (DBR), traffic context recognition (TCR), and vehicle behavior [...] Read more.
Advanced driver assistance systems (ADAS) require a comprehensive understanding of multiple tasks related to the physical and mental states of drivers and traffic situations. Existing ADAS studies perform driver emotion recognition (DER), driver behavior recognition (DBR), traffic context recognition (TCR), and vehicle behavior recognition (VBR) using models designed based on single-task learning, thereby failing to reflect the interactions among tasks in real driving environments. This paper proposes perception and recognition with inter-modal selective multi-task learning (PRISM-MTL), an integrated multimodal and multi-task learning framework that jointly recognizes DER, DBR, TCR, and VBR. The proposed PRISM-MTL consists of a hierarchical stage-wise attention network (HSA-Net)-based multimodal encoder that extracts spatial features from heterogeneous multimodal inputs and task-specific modality fusion (TSMF), which selectively learns effective modality information for each task. This design addresses negative transfer, a key challenge in multi-task learning. In the multimodal encoder, HSA-Net extracts visual modality tokens that emphasize global structural patterns and key spatial regions from multi-view images, while Token-SE generates joint modality tokens that reflect the spatial configuration of joint data. TSMF generates task-specific fusion features that selectively emphasize the modality cues for each task. The generated task-specific fusion features are summarized through temporal mean pooling, and final predictions of driver states and traffic situations are produced by each task head. Experimental results show that the proposed PRISM-MTL achieves state-of-the-art performance on the public AIDE database, with an mAcc of 86.25% ± 0.35 for multi-task recognition of driver states and traffic situations. Full article
(This article belongs to the Section E1: Mathematics and Computer Science)
Show Figures

Figure 1

23 pages, 5613 KB  
Article
Development and Validation of a Multimodal Emotion Recognition Ability Test Based on the Chinese Cultural Context
by Xiaoming Song, Jinmei Leng, Xia Wu, Sheng Yang and Fang Luo
J. Intell. 2026, 14(8), 169; https://doi.org/10.3390/jintelligence14080169 - 1 Aug 2026
Viewed by 402
Abstract
Emotion recognition is essential for social adaptation and mental health. However, existing measures still rely on limited stimulus materials or simplified task formats, which restricts ecological validity. The present study aimed to construct a localized multimodal emotion expression database in the Chinese context [...] Read more.
Emotion recognition is essential for social adaptation and mental health. However, existing measures still rely on limited stimulus materials or simplified task formats, which restricts ecological validity. The present study aimed to construct a localized multimodal emotion expression database in the Chinese context and to develop a multimodal emotion recognition ability test integrating emotion category and intensity recognition. In Study 1, a multimodal emotion expression database was established through actor performances, expert evaluation, and participant validation. The final database contained 2160 videos representing six basic emotions, five levels of emotional intensity, and two expression modes (vocal and non-vocal). In Study 2, 90 videos were selected under three recognition conditions (vocal-expression, muted vocal-expression, and muted non-vocal-expression). An integrated scoring approach combining emotion category and intensity was used to develop the test. A total of 236 valid participants completed the test. The results indicated that, after 61 video clips were retained, the test demonstrated good reliability and item discrimination. Evidence for criterion-related validity and construct validity was also obtained. Overall, the findings suggest that the present measure is a reliable and valid tool for assessing multimodal emotion recognition ability in the Chinese cultural context and offers the first empirical evidence that integrating category and intensity recognition can better capture the multidimensional nature of this ability. Full article
Show Figures

Figure 1

Back to TopTop