Next Article in Journal
Does Generative Artificial Intelligence Improve Students’ Higher-Order Thinking? A Meta-Analysis Based on 29 Experiments and Quasi-Experiments
Next Article in Special Issue
Determinants of Trust: Evidence from Elementary School Classrooms
Previous Article in Journal / Special Issue
Teachers’ Emotional Commitment: The Emotional Bond That Sustains Teaching
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Bridging Text and Speech for Emotion Understanding: An Explainable Multimodal Transformer Fusion Framework with Unified Audio–Text Attribution

Computer Science Engineering Department, Thapar Institute of Engineering and Technology, Patiala 147001, Punjab, India
*
Author to whom correspondence should be addressed.
J. Intell. 2025, 13(12), 159; https://doi.org/10.3390/jintelligence13120159
Submission received: 15 September 2025 / Revised: 15 November 2025 / Accepted: 28 November 2025 / Published: 3 December 2025
(This article belongs to the Special Issue Social Cognition and Emotions)

Abstract

Conversational interactions, rich in both linguistic and vocal cues, provide a natural context for studying these processes. In this work, we propose an explainable multimodal transformer framework that integrates textual semantics (via RoBERTa) and acoustic prosody (via WavLM) to advance emotion understanding. By projecting both modalities into a shared latent space, our model captures the complementary contributions of language and speech to affective communication, achieving an 0.83 accuracy value across five emotion categories. Crucially, we embed explainable AI (XAI) techniques including Integrated Gradients and Occlusion to attribute predictions to specific linguistic tokens and prosodic patterns, thereby aligning computational mechanisms with human cognitive processes of emotion perception. Beyond performance gains, this work demonstrates how multimodal AI systems can support transparent, human-centered emotion recognition.
Keywords: emotion; multimodal learning; explainable AI; speech; text fusion emotion; multimodal learning; explainable AI; speech; text fusion

Share and Cite

MDPI and ACS Style

Pandey, A.; Singh, J.; Kaur, M. Bridging Text and Speech for Emotion Understanding: An Explainable Multimodal Transformer Fusion Framework with Unified Audio–Text Attribution. J. Intell. 2025, 13, 159. https://doi.org/10.3390/jintelligence13120159

AMA Style

Pandey A, Singh J, Kaur M. Bridging Text and Speech for Emotion Understanding: An Explainable Multimodal Transformer Fusion Framework with Unified Audio–Text Attribution. Journal of Intelligence. 2025; 13(12):159. https://doi.org/10.3390/jintelligence13120159

Chicago/Turabian Style

Pandey, Ashutosh, Jasmeet Singh, and Maninder Kaur. 2025. "Bridging Text and Speech for Emotion Understanding: An Explainable Multimodal Transformer Fusion Framework with Unified Audio–Text Attribution" Journal of Intelligence 13, no. 12: 159. https://doi.org/10.3390/jintelligence13120159

APA Style

Pandey, A., Singh, J., & Kaur, M. (2025). Bridging Text and Speech for Emotion Understanding: An Explainable Multimodal Transformer Fusion Framework with Unified Audio–Text Attribution. Journal of Intelligence, 13(12), 159. https://doi.org/10.3390/jintelligence13120159

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop