Task-Adaptive and Multi-Level Contextual Understanding for Emotion Recognition in Conversations
Abstract
1. Introduction
- (1)
- A task-adaptive representation learning mechanism that enhances PLM’s adaptability to specific emotion tasks through emotion space-constrained prompt learning and contrastive fine-tuning;
- (2)
- A multi-level contextual understanding architecture designed to systematically handle the macroscopic narrative flow and microscopic interactions in conversations;
- (3)
- Extensive experiments on benchmark datasets demonstrate that TAMC-ERC achieves highly competitive performance, with ablation studies confirming a synergistic effect between its core components.
2. Related Work
2.1. Emotion Recognition in Conversations
2.1.1. Sequence-Based Models
2.1.2. Graph-Based Models
2.1.3. Knowledge-Enhanced Models
2.2. Contrastive Learning for ERC
2.3. Task-Adaptive Methods for ERC
3. Problem Definition
4. The TAMC-ERC Model
4.1. Model Framework Overview
4.2. Task-Adaptive Representation Learning
4.2.1. Task-Adaptive Prompt Context Encoding
4.2.2. Task-Adaptive Contrastive Fine-Tuning
4.3. Multi-Level Contextual Understanding
4.3.1. Global Contextual Understanding Module
4.3.2. Local Contextual Understanding Module
- (1)
- Foundational Context: Local Semantic Field
- (2)
- Focused Context: Speaker State Trajectory
- (3)
- Core Context: Direct Interaction Adjacency Pair
4.4. Emotion Recognition and Optimization
5. Experiments
5.1. Datasets
- (1)
- IEMOCAP Dataset [37]: This is a multimodal dataset. Each dialogue segment is performed by two actors based on scripts. Only the text modality is used in this experiment. The data samples in this dataset involve 6 emotion categories: happy, sad, angry, frustrated, excited, and neutral. Considering that there is no validation set in the IEMOCAP dataset, following the practice in DAG-ERC, we split the data into a training set and a validation set, with the last 20 dialogues held out for validation.
- (2)
- MELD Dataset [38]: This is a multi-party, multi-modal dialogue dataset derived from the TV show Friends. Different from IEMOCAP, this dataset encompasses seven emotion categories: anger, disgust, fear, happiness, neutral, sadness, and surprise.
- (3)
- EmoryNLP Dataset [39]: Similar to MELD, EmoryNLP is based on Friends, involving multi-party dialogues and containing 7 categories. However, EmoryNLP differs from MELD in scene selection, and the set of 7 emotion categories it defines differs from that of MELD, namely happy, powerful, peaceful, sad, crazy, fear, and neutral.
5.2. Experimental Setup
5.3. Baseline Models
- (1)
- Graph-based Models
- (2)
- Sequence-based Models
- (3)
- Knowledge-Enhanced Models
6. Experimental Results and Analysis
6.1. Overall Performance Comparison
6.2. Ablation Experiments
6.3. Analysis of Synergistic Effects
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Majumder, N.; Poria, S.; Hazarika, D.; Mihalcea, R.; Gelbukh, A.; Cambria, E. DialogueRNN: An attentive RNN for emotion detection in conversations. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, Honolulu, HI, USA, 27 January–1 February 2019; pp. 6818–6825. [Google Scholar]
- Ghosal, D.; Majumder, N.; Poria, S.; Chhaya, N.; Gelbukh, A. DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, 3–7 November 2019; pp. 154–164. [Google Scholar]
- Joshi, A.; Bhat, A.; Jain, A.; Singh, A.V.; Modi, A. Cogmen: Contextualized gnn based multimodal emotion recognition. arXiv 2022, arXiv:2205.02455. [Google Scholar] [CrossRef] [Scilit]
- Shen, W.; Wu, S.; Yang, Y.; Quan, X. Directed acyclic graph network for conversational emotion recognition. arXiv 2021, arXiv:2105.12907. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Yan, H.; Qiu, X. Contrast and generation make bart a good dialogue emotion recognizer. In Proceedings of the AAAI Conference on Artificial Intelligence, Online, 22 February–1 March 2022; Volume 36, pp. 11002–11010. [Google Scholar]
- Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; Zettlemoyer, L. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 5–10 July 2020; pp. 7871–7880. [Google Scholar]
- Hu, D.; Bao, Y.; Wei, L.; Zhou, W.; Hu, S. Supervised adversarial contrastive learning for emotion recognition in conversations. arXiv 2023, arXiv:2306.01505. [Google Scholar]
- Liu, Y.; Zhao, J.; Hu, J.; Li, R.; Jin, Q. Dialogueein: Emotion interaction network for dialogue affective analysis. In Proceedings of the 29th International Conference on Computational Linguistics, Gyeongju, Republic of Korea, 12–17 October 2022; pp. 684–693. [Google Scholar]
- Yang, H.; Gao, X.; Wu, J.; Gan, T.; Ding, N.; Jiang, F.; Nie, L. Self-adaptive Context and Modal-interaction Modeling For Multimodal Emotion Recognition. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2023, Toronto, ON, Canada, 9–14 July 2023; Association for Computational Linguistics: Stroudsburg, PA, USA, 2023; pp. 6267–6281. [Google Scholar]
- Yang, L.; Shen, Y.; Mao, Y.; Cai, L. Hybrid curriculum learning for emotion recognition in conversation. In Proceedings of the AAAI Conference on Artificial Intelligence, Online, 22 February–1 March 2022; Volume 36, pp. 11595–11603. [Google Scholar]
- Lei, S.; Dong, G.; Wang, X.; Wang, K.; Wang, S. Instructerc: Reforming emotion recognition in conversation with a retrieval multi-task llms framework. arXiv 2023, arXiv:2309.11911. [Google Scholar]
- Zhang, Y.; Wang, M.; Tiwari, P.; Li, Q.; Wang, B.; Qin, J. Dialoguellm:Context and emotion knowledge-tuned llama mod-els for emotion recognition in conversations. arXiv 2023, arXiv:2310.11374. [Google Scholar]
- Lee, J. The emotion is not one-hot en-coding: Learning with grayscale label for emo-tion recognition in conversation. arXiv 2022, arXiv:2206.07359. [Google Scholar]
- Jia, Z.; Shi, Y.; Liu, W.; Huang, Z.; Sun, X. Speaker-aware interactive graph attention network for emotion recognition in conversation. ACM Trans. Asian Low-Resour. Lang. Inf. Process. 2023, 22, 1–18. [Google Scholar] [CrossRef] [Scilit]
- Zhang, D.; Chen, F.; Chen, X. Dualgats: Dual graph attention networks for emotion recognition in conversations. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Toronto, ON, Canada, 9–14 July 2023; Volume 1: Long Papers, pp. 7395–7408. [Google Scholar]
- Hu, J.; Liu, Y.; Zhao, J.; Jin, Q. Mmgcn: Multimodal fusion via deep graph convolution network for emotion recognition in con-versation. arXiv 2021, arXiv:2107.06779. [Google Scholar]
- Tu, G.; Xie, T.; Liang, B.; Wang, H.; Xu, R. Adaptive graph learning for multimodal conversational emotion detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 26–27 February 2024; Volume 38, pp. 19089–19097. [Google Scholar]
- Zhong, P.; Wang, D.; Miao, C. Knowledge-enriched transformer for emotion de-tection in textual conversations. arXiv 2019, arXiv:1909.10681. [Google Scholar]
- Ghosal, D.; Majumder, N.; Gel-bukh, A.; Mihalcea, R.; Poria, S. Cosmic: Commonsense knowledge for emotion identification in conversations. arXiv 2020, arXiv:2010.02795. [Google Scholar]
- Zhu, L.; Pergola, G.; Gui, L.; Zhou, D.; He, Y. Topic-driven and knowledge-aware transformer for dialogue emotion detection. arxiv 2021, arXiv:2106.01071. [Google Scholar]
- Fu, Y.; Okada, S.; Wang, L.; Guo, L.; Dang, J. CONSK-GCN: Conversational semantic-and knowledge-oriented graph convolutional network for multimodal emotion recognition. In Proceedings of the 2021 IEEE International Conference on Multimedia and Expo (ICME), Shenzhen, China, 5–9 July 2021; IEEE Computer Society: Washington, DC, USA, 2021; pp. 1–6. [Google Scholar]
- Li, W.; Zhu, L.; Mao, R.; Cambria, E. SKIER: A symbolic knowledge integrated model for conversational emotion recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Montréal, QC, Canada, 8–10 August 2023; Volume 37, pp. 13121–13129. [Google Scholar]
- Xie, Y.; Yang, K.; Sun, C.; Liu, B.; Ji, Z. Knowledge-interactive network with emotion polarity intensity-aware multi-task learning for emotion recognition in conversations. In Proceedings of the Findings of the Association for Computational Linguistics, EMNLP 2021, Virtual, 16–20 November 2021; pp. 2879–2889. [Google Scholar]
- Tu, G.; Wang, J.; Li, Z.; Chen, S.; Liang, B.; Zeng, X.; Yang, M.; Xu, R. Multiple knowledge-enhanced interactive graph network for multimodal conversational emotion recognition. In Proceedings of the Findings of the Association for Computational Linguistics, EMNLP 2024, Miami, FL, USA, 12–16 November 2024; pp. 3861–3874. [Google Scholar]
- Chen, T.; Kornblith, S.; Norouzi, M.; Hinton, G. A simple framework for contrastive learning of visual representations. In Proceedings of the International Conference on Machine Learning, Virtual, 13–18 July 2020; PMLR: Breckenridge, CO, USA, 2020; pp. 1597–1607. [Google Scholar]
- He, K.; Fan, H.; Wu, Y.; Xie, S.; Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 9729–9738. [Google Scholar]
- Gunel, B.; Du, J.; Conneau, A.; Stoyanov, V. Supervised contrastive learning for pre-trained language model fine-tuning. arXiv 2020, arXiv:2011.01403. [Google Scholar]
- Kang, B.; Li, Y.; Xie, S.; Yuan, Z.; Feng, J. Exploring balanced feature spaces for representation learning. In Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia, 30 April 2020. [Google Scholar]
- Li, T.; Cao, P.; Yuan, Y.; Fan, L.; Yang, Y.; Feris, R.S.; Indyk, P.; Katabi, D. Targeted supervised contrastive learning for long-tailed recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 6918–6928. [Google Scholar]
- Zhu, J.; Wang, Z.; Chen, J.; Chen, Y.-P.P.; Jiang, Y.-G. Balanced contrastive learning for long-tailed visual recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 6908–6917. [Google Scholar]
- Zhang, Z.; Zhao, Y.; Chen, M.; He, X.-A. Label anchored contrastive learn-ingfor language understanding. arXiv 2022, arXiv:2205.10227. [Google Scholar]
- Tu, G.; Liang, B.; Mao, R.; Yang, M.; Xu, R. Context or knowledge is not always necessary: A contrastive learning framework for emotion recognition in conversations. In Proceedings of the Findings of the Association for Computational Linguistics, ACL 2023, Toronto, ON, Canada, 9–14 July 2023; pp. 14054–14067. [Google Scholar]
- Yu, F.; Guo, J.; Wu, Z.; Dai, X. Emotion-anchored contrastive learning framework for emotion recognition in conversation. arxiv 2024, arXiv:2403.20289. [Google Scholar]
- Li, C.; Gao, F.; Bu, J.; Xu, L.; Chen, X.; Gu, Y.; Shao, Z.; Zheng, Q.; Zhang, N.; Wang, Y.; et al. Sentiprompt: Sentiment knowledge enhanced prompt-tuning for aspect-based sentiment analysis. arXiv 2021, arXiv:2109.08306. [Google Scholar]
- Cai, C.; Zhang, K.; Hu, Z.; Lin, X.; Pan, Z. Prompt-based hybrid supervised conterative learning for emotion recognition in conversation. Neurocomputing 2025, 647, 130453. [Google Scholar]
- Yang, K.; Zhang, T.; Alhuzali, H.; Ananiadou, S. Cluster-level contrastive learning for emotion recognition in conversations. IEEE Trans. Affect. Comput. 2023, 14, 3269–3280. [Google Scholar] [CrossRef] [Scilit]
- Busso, C.; Bulut, M.; Lee, C.C.; Kazemzadeh, A.; Mower, E.; Kim, S.; Chang, J.N.; Lee, S.; Narayanan, S.S. IEMOCAP: Interactive emotional dyadic motion capture database. Lang. Resour. Eval. 2008, 42, 335–359. [Google Scholar]
- Poria, S.; Hazarika, D.; Majumder, N.; Naik, G.; Cambria, E.; Mihalcea, R. MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019; pp. 527–536. [Google Scholar]
- Zahiri, S.M.; Choi, J.D. Emotion detection on tv show transcripts with sequence-based convolutional neural networks. In Proceedings of the Workshops at the thirty-Second AAAI Conference on Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018; pp. 44–51. [Google Scholar]
- Ishiwatari, T.; Yasuda, Y.; Miyazaki, T.; Goto, J. Relation-aware graph attention networks with relational position encodings for emotion recognition in conversations. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, 16–20 November 2020; pp. 7360–7370. [Google Scholar]
- Lee, J.; Lee, W. Compm: Context modeling with speaker’s pre-trained memory track-ing for emotion recognition in conversation. arXiv 2021, arXiv:2108.11626. [Google Scholar]
- Zhao, S.; Liu, W.; Chen, J.; Sun, X. Dieu: A dynamic interaction emotion unit for emotion recognition in conversation. ACM Trans. Asian Low-Resour. Lang. Inf. Process. 2023, 22, 1–18. [Google Scholar]
- Zhang, T.; Chen, Z.; Zhong, M.; Qian, T. Mimicking the thinking process for emotion recognition in conversation with prompts and paraphrasing. arXiv 2023, arXiv:2306.06601. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.; Tang, F.; Zhao, M.; Zhu, Y. Emocaps: Emotion capsule based model for conversational emotion recognition. arXiv 2022, arXiv:2203.13504. [Google Scholar] [CrossRef] [Scilit]
- Song, X.; Huang, L.; Xue, H.; Hu, S. Supervised prototypical contrastive learning for emotion recognition in conversation. arXiv 2022, arXiv:2210.08713. [Google Scholar] [CrossRef] [Scilit]
- Zhao, W.; Zhao, Y.; Lu, X.; Wang, S.; Qin, B. Is chat-gpt equipped with emotional dialogue capabilities? arXiv 2023, arXiv:2304.09582. [Google Scholar]
- Wang, S. Real operational labeled data of air handling units from office, auditorium, and hospital buildings. Sci. Data 2025, 12, 1481. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Moon, S.; Eum, I.; Hwang, D.; Kim, J. A text dataset of fire door defects for pre-delivery inspections of apartments during the construction stage. Data Brief 2025, 60, 111536. [Google Scholar] [CrossRef] [Scilit]


| Symbol | Description |
|---|---|
| A conversation, composed of multiple utterances. | |
| The set of speakers in a conversation. | |
| The emotion space of the dataset, for instance, E = {neutral, happy, angry, sad, frustrated, excited} in the IEMOCAP benchmark. | |
| The emotion concept vector corresponding to an emotion category in . | |
| The set of all emotion concept vectors, {, …,}. | |
| A prompt template. | |
| The projection head for contrastive learning. | |
| The representation of an utterance. | |
| The projected utterance representation for contrastive learning. | |
| Temperature hyperparameter. | |
| Cosine similarity function. | |
| A local context window centered on the t-th utterance. | |
| Historical utterances of the target speaker (up to time t). | |
| Adjacent utterance pairs within the local context. |
| Dataset | Total Number of Conversations | Number of Emotion Categories | Average Number of Speakers per Conversation | Average Number of Clauses per Conversation |
| IEMOCAP | 151 | 6 | 2 | 49.2 |
| MELD | 1433 | 7 | 9 | 9.6 |
| EmoryNLP | 897 | 7 | 9 | 14.1 |
| Hyperparameters | IEMOCAP | MELD | EmoryNLP |
| Learning rate | 1 × 10−4 | 1 × 10−4 | 1 × 10−4 |
| Dropout | 0.2 | 0.2 | 0.2 |
| Temperature | 0.2 | 0.3 | 0.2 |
| Maximum length | 256 | 256 | 256 |
| Batch size | 32 | 32 | 32 |
| Epochs | 8 | 6 | 6 |
| Methods | IEMOCAP | MELD | EmoryNLP | Average |
|---|---|---|---|---|
| Graph-based models | ||||
| DialogueGCN (Ghosal et al., 2019) [2] | 64.91 | 63.02 | 38.10 | 55.34 |
| RGAT (Ishiwatari et al., 2020) [40] | 66.36 | 62.80 | 37.89 | 55.68 |
| DAG-ERC (Shen et al., 2021) [4] | 68.03 | 63.65 | 39.02 | 56.90 |
| DAG-ERC+HCL (Yang et al., 2022) [10] | 68.73 | 63.89 | 39.82 | 57.48 |
| SIGAT (Jia et al., 2023) [14] | 70.71 | 66.20 | 39.95 | 58.77 |
| AdaIGN (Tu et al., 2024) [17] | 70.74 | 66.79 | - | - |
| Sequence-based models | ||||
| CKCL (Tu et al., 2023) [32] | 67.16 | 66.21 | 40.23 | 57.87 |
| Cog-BART (Li et al., 2022) [5] | 66.18 | 64.81 | 39.04 | 56.68 |
| DialogueEIN (Liu et al., 2022) [8] | 68.93 | 65.37 | 38.92 | 57.74 |
| CoMPM (Lee and Lee., 2021) [41] | 69.46 | 66.52 | 38.93 | 58.30 |
| SupCon (Gunel et al., 2020) [27] | 68.14 | 65.63 | 39.28 | 57.68 |
| Emocaps (Li et al., 2022) [44] | 69.49 | 63.51 | - | - |
| SPCL+CL (Song et al., 2022) [45] | 67.19 | 65.74 | 39.52 | 57.48 |
| SACL (Hu et al., 2023) [7] | 69.22 | 66.45 | 39.65 | 58.44 |
| SCCL (Yang et al., 2023) [36] | 69.88 | 65.70 | 38.75 | 58.11 |
| DIEU (Zhao et al., 2023a) [42] | 69.90 | 66.43 | 40.12 | 58.11 |
| MPLP (Zhang et al., 2023) [43] | 66.65 | 66.51 | - | - |
| ChatGPT 3-shot (Zhao et al., 2023) [46] | 48.58 | 58.35 | 35.92 | 47.62 |
| EACL (Yu et al., 2024) [33] | 70.41 | 67.12 | 40.24 | 59.26 |
| Knowledge-Enhanced Models | ||||
| KET (Zhong et al., 2019) [18] | 59.56 | 58.18 | 33.95 | 50.56 |
| COSMIC (Ghosal et al., 2020) [19] | 65.25 | 65.21 | 38.11 | 56.19 |
| TODKAT (Zhu et al., 2021) [20] | 61.33 | 65.47 | 43.12 | 56.53 |
| KI-Net (Xie et al., 2021) [23] | 66.98 | 63.24 | - | - |
| TAMC-ERC (ours) | 71.04 | 66.95 | 40.99 | 59.66 |
| Dataset | IEMOCAP | MELD | EmoryNLP |
|---|---|---|---|
| Ours | 71.04 | 66.95 | 40.99 |
| w/o Task-Adaptive Contrastive Learning | 68.77 (2.27 ↓) | 66.2 (0.75 ↓) | 40.7 (0.29 ↓) |
| w/o Prototype-Parameterized | 69.45 (1.59 ↓) | 63.24 (3.71 ↓) | 36.61 (4.38 ↓) |
| w/o Global Contextual Understanding | 70.84 (0.20 ↓) | 66.79 (0.16 ↓) | 40.81 (0.18 ↓) |
| w/o Local Contextual Understanding | 70.8 (0.24 ↓) | 66.82 (0.13 ↓) | 40.77 (0.22 ↓) |
| w/o Global-Local Model | 69.64 (1.4 ↓) | 66.26 (0.69 ↓) | 38.92 (2.07 ↓) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yao, X.; Cao, W.; Xue, Y.; Zhang, H.; Fan, X. Task-Adaptive and Multi-Level Contextual Understanding for Emotion Recognition in Conversations. Appl. Sci. 2026, 16, 1706. https://doi.org/10.3390/app16041706
Yao X, Cao W, Xue Y, Zhang H, Fan X. Task-Adaptive and Multi-Level Contextual Understanding for Emotion Recognition in Conversations. Applied Sciences. 2026; 16(4):1706. https://doi.org/10.3390/app16041706
Chicago/Turabian StyleYao, Xiaomeng, Wei Cao, Yuyang Xue, Haijun Zhang, and Xiaochao Fan. 2026. "Task-Adaptive and Multi-Level Contextual Understanding for Emotion Recognition in Conversations" Applied Sciences 16, no. 4: 1706. https://doi.org/10.3390/app16041706
APA StyleYao, X., Cao, W., Xue, Y., Zhang, H., & Fan, X. (2026). Task-Adaptive and Multi-Level Contextual Understanding for Emotion Recognition in Conversations. Applied Sciences, 16(4), 1706. https://doi.org/10.3390/app16041706

