Attention Bidirectional Gated Fusion Based Multimodal Intent Recognition Under Uncertain Missing Modalities
Abstract
1. Introduction
- How can we enhance the quality of multimodal fusion for MIR under uncertain missing modalities? Currently, some research work has demonstrated that the text modality plays a major role in MIR. However, existing works for MIR treat all the three modalities equally in fusion, and cannot deeply interact with multimodal features in the fusion, thus affecting the quality of multimodal fusion.
- How can we overcome the semantic inconsistency caused by uncertain missing modalities for MIR? Uncertain missing modalities cause the semantics of multimodal samples to be inconsistent with the semantics of complete modalities. Existing models trained with full modalities cannot overcome the semantic inconsistency. Therefore, determining how to handle the problem of semantic inconsistency caused by uncertain missing modalities becomes a challenge for MIR.
- How can we handle the uncertain missing modalities for MIR? Existing works on MIR usually assume that all the modalities are available all the time, failed to consider the uncertain missing modalities. Therefore, determining how to effectively handle the uncertain missing modalities for MIR becomes a new challenge.
- To enhance the quality of multimodal fusion, we propose an attention-based bidirectional gated multimodal feature fusion module. Moreover, inspired by the fact that text modality plays a dominant role in MIR, we propose to reduce the distance from other modalities to the text modality to improve the quality of multimodal fusion.
- To address the problem of semantic inconsistency caused by uncertain missing modalities, we propose learnable Transformer prompts to mitigate the performance degradation caused by missing modalities. In this work, different types of attention level prompts are input according to different scenarios of missing modalities, thus enabling the model toward the target domain through training with fewer parameters, guiding the model to focus on the missing modalities.
- To better cope with uncertain missing modalities, we introduce a pre-trained model (Attention-based Bidirectional Gated Neural Network, AGNN) that is trained with complete modalities and is used to provide effective guidance for the final classification in scenarios of uncertain missing modalities, thereby improving the robustness of ABGFMIR.
2. Related Work
3. Methodology
3.1. Problem Statement
3.2. Model Overview
3.3. Extracting Features of Each Modality with LSTM
3.4. Attention-Based Bidirectional Gated Fusion
3.5. Modality Missing Prompt Based on Transformer
3.6. Intent Recognition
3.7. Generalization Learning Based on Pre-Trained Model
3.8. Model Training
4. Experiments
4.1. Benchmark Datasets
4.2. Experimental Environment and Parameter Settings
4.3. Baseline Models
4.4. Hyperparametric Sensitivity Analysis
4.5. Performances Comparison
4.6. Ablation Experiment
4.7. Comparison of Different Grain Size Classifications
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Data Availability Statement
Conflicts of Interest
References
- Louvan, S.; Magnini, B. Recent neural methods on slot filling and intent classification for task-oriented dialogue systems: A survey. In Proceedings of the 28th International Conference on Computational linguistics, Barcelona, Spain, 8–13 December 2020; pp. 480–496. [Google Scholar]
- Singhal, B.; Gupta, A.; Shivasankaran, V.; Krishna, A. IntenDD: A unified contrastive learning approach for intent detection and discovery. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, 6–10 December 2023; pp. 14204–14216. [Google Scholar]
- Sung, M.; Gung, J.; Mansimov, E.; Pappas, N.; Shu, R.; Romeo, S.; Zhang, Y.; Castelli, V. Pre-training intent-aware encoders for zero-and few-shot intent classification. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, 6–10 December 2023; pp. 10433–10442. [Google Scholar]
- Zhang, F.; Chen, W.; Ding, F.; Gao, M.; Wang, T.; Yao, J.; Zheng, J. From Discrimination to Generation: Low-Resource Intent Detection with Language Model Instruction Tuning. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand, 11–16 August 2024; pp. 10174–10187. [Google Scholar]
- Kim, M.; Choi, J.; Jang, C.; Lee, J. Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models. arXiv 2026, arXiv:2605.12517. [Google Scholar]
- Zhu, Z.; Zhang, F.; Zhang, Y.; Sun, J.; Huang, Z.; Long, Q.; Xing, B.; Wu, X. A Survey on Multi-modal Intent Recognition: Recent Advances and New Frontiers. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, 4–9 November 2025; pp. 15223–15236. [Google Scholar]
- Li, H.; Yang, Q.; Xia, Y.; Lu, L.; Wei, Q. SaliText: A multimodal intent recognition method with saliency and text-guided fusion. Signal Process. 2026, 244, 110537. [Google Scholar] [CrossRef]
- Gui, Q.; Liu, X.; Wang, J.; Ouyang, X.; Huang, W.; Zong, L. Evidence-driven ternary contrastive learning with hierarchical mamba fusion for robust multimodal intent recognition. Neurocomputing 2026, 674, 132866. [Google Scholar] [CrossRef]
- Li, Y.; Wang, T.; Tang, M.; Hu, J. Text-guided frequency-decoupled modeling for multimodal intent recognition. Neurocomputing 2026, 687, 133795. [Google Scholar] [CrossRef]
- Wang, M.; Xie, L.; Li, C.; Wang, X.; Sun, M.; Liu, Z. MGC: A modal mapping coupling and gate-driven contrastive learning approach for multimodal intent recognition. Expert Syst. Appl. 2025, 281, 127631. [Google Scholar] [CrossRef]
- Sun, K.; Xie, Z.; Ye, M.; Zhang, H. Contextual Augmented Global Contrast for Multimodal Intent Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 26953–26963. [Google Scholar]
- Huang, X.; Ma, T.; Jia, L.; Zhang, Y.; Rong, H.; Alnabhan, N. An effective multimodal representation and fusion method for multimodal intent recognition. Neurocomputing 2023, 548, 126373. [Google Scholar] [CrossRef]
- Hu, B.; Zhang, K.; Zhang, Y.; Ye, Y. Adaptive multimodal fusion: Dynamic attention allocation for intent recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025; pp. 17267–17275. [Google Scholar]
- Wang, X.; Zhou, Y.; Huang, B.; Chen, H.; Zhu, W. Multi-modal Generative AI: Multi-modal LLMs, Diffusions and the Unification. IEEE Trans. Circuits Syst. Video Technol. 2025, 36, 5621–5641. [Google Scholar]
- Zhang, C.; Cui, Y.; Han, Z.; Zhou, J.T.; Fu, H.; Hu, Q. Deep partial multi-view learning. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 44, 2402–2415. [Google Scholar]
- Qin, L.; Xie, T.; Che, W.; Liu, T. A survey on spoken language understanding: Recent advances and new frontiers. arXiv 2021, arXiv:2103.03095. [Google Scholar]
- Qin, L.; Liu, T.; Che, W.; Kang, B.; Zhao, S.; Liu, T. A co-interactive transformer for joint slot filling and intent detection. In Proceedings of the ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); IEEE: Piscataway, NJ, USA, 2021; pp. 8193–8197. [Google Scholar]
- Wu, D.; Ding, L.; Lu, F.; Xie, J. SlotRefine: A fast non-autoregressive model for joint intent detection and slot filling. arXiv 2020, arXiv:2010.02693. [Google Scholar]
- Qin, L.; Xu, X.; Che, W.; Liu, T. AGIF: An adaptive graph-interactive framework for joint multiple intent detection and slot filling. arXiv 2020, arXiv:2004.10087. [Google Scholar]
- Hou, Y.; Lai, Y.; Wu, Y.; Che, W.; Liu, T. Few-shot learning for multi-label intent detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 19–21 May 2021; pp. 13036–13044. [Google Scholar] [CrossRef]
- Wu, T.W.; Su, R.; Juang, B. A label-aware BERT attention network for zero-shot multi-intent detection in spoken language understanding. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, 7–11 November 2021; pp. 4884–4896. [Google Scholar]
- Jia, M.; Wu, Z.; Reiter, A.; Cardie, C.; Belongie, S.; Lim, S.N. Intentonomy: A dataset and study towards human intent understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 12986–12996. [Google Scholar]
- Joo, J.; Li, W.; Steen, F.F.; Zhu, S.C. Visual persuasion: Inferring communicative intents of images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 216–223. [Google Scholar]
- Fang, Z.; López, A.M. Intention recognition of pedestrians and cyclists by 2d pose estimation. IEEE Trans. Intell. Transp. Syst. 2019, 21, 4773–4783. [Google Scholar] [CrossRef]
- Yang, B.; Zhan, W.; Wang, P.; Chan, C.; Cai, Y.; Wang, N. Crossing or not? Context-based recognition of pedestrian crossing intention in the urban environment. IEEE Trans. Intell. Transp. Syst. 2021, 23, 5338–5349. [Google Scholar] [CrossRef]
- Xu, B.; Li, J.; Wong, Y.; Zhao, Q.; Kankanhalli, M.S. Interact as you intend: Intention-driven human-object interaction detection. IEEE Trans. Multimed. 2019, 22, 1423–1432. [Google Scholar]
- Zhang, H.; Xu, H.; Wang, X.; Zhou, Q.; Zhao, S.; Teng, J. Mintrec: A new dataset for multimodal intent recognition. In Proceedings of the 30th ACM International Conference on Multimedia, Lisboa, Portugal, 10–14 October 2022; pp. 1688–1697. [Google Scholar]
- Maharana, A.; Tran, Q.H.; Dernoncourt, F.; Yoon, S.; Bui, T.; Chang, W.; Bansal, M. Multimodal intent discovery from livestream videos. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2022, Dublin, Ireland, 22–27 May 2022; pp. 476–489. [Google Scholar]
- Kruk, J.; Lubin, J.; Sikka, K.; Lin, X.; Jurafsky, D.; Divakaran, A. Integrating text and image: Determining multimodal document intent in instagram posts. arXiv 2019, arXiv:1904.09073. [Google Scholar]
- Kiela, D.; Firooz, H.; Mohan, A.; Goswami, V.; Singh, A.; Ringshia, P.; Testuggine, D. The hateful memes challenge: Detecting hate speech in multimodal memes. Adv. Neural Inf. Process. Syst. 2020, 33, 2611–2624. [Google Scholar]
- Liu, H.; Wang, W.; Li, H. Towards multi-modal sarcasm detection via hierarchical congruity modeling with knowledge enhancement. arXiv 2022, arXiv:2210.03501. [Google Scholar]
- Zhang, L.; Shen, J.; Zhang, J.; Xu, J.; Li, Z.; Yao, Y.; Yu, L. Multimodal marketing intent analysis for effective targeted advertising. IEEE Trans. Multimed. 2021, 24, 1830–1843. [Google Scholar] [CrossRef]
- Zeng, J.; Liu, T.; Zhou, J. Tag-assisted multimodal sentiment analysis under uncertain missing modalities. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain, 11–15 July 2022; pp. 1545–1554. [Google Scholar]
- He, Y.; Zheng, S.; Tay, Y.; Gupta, J.; Du, Y.; Aribandi, V.; Zhao, Z.; Li, Y.; Chen, Z.; Metzler, D.; et al. Hyperprompt: Prompt-based task-conditioning of transformers. In Proceedings of the International Conference on Machine Learning, Baltimore, MD, USA, 17–23 July 2022; pp. 8678–8690. [Google Scholar]
- Liu, Z.; Zhou, B.; Chu, D.; Sun, Y.; Meng, L. Modality translation-based multimodal sentiment analysis under uncertain missing modalities. Inf. Fusion 2024, 101, 101973. [Google Scholar] [CrossRef]
- Castro, S.; Hazarika, D.; Pérez-Rosas, V.; Zimmermann, R.; Mihalcea, R.; Poria, S. Towards multimodal sarcasm detection (an _obviously_ perfect paper). arXiv 2019, arXiv:1906.01815. [Google Scholar]
- Tsai, Y.H.H.; Bai, S.; Liang, P.P.; Kolter, J.Z.; Morency, L.P.; Salakhutdinov, R. Multimodal transformer for unaligned multimodal language sequences. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy, 28 July–2 August 2019. [Google Scholar]
- Rahman, W.; Hasan, M.K.; Lee, S.; Zadeh, A.; Mao, C.; Morency, L.P.; Hoque, E. Integrating multimodal information in large Pretrained transformers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online, 5–10 July 2020. [Google Scholar]
- Hazarika, D.; Zimmermann, R.; Poria, S. Misa: Modality-invariant and-specific representations for multimodal sentiment analysis. In Proceedings of the 28th ACM International Conference on Multimedia, Online, 12–16 October 2020; pp. 1122–1131. [Google Scholar]
- Zuo, H.; Liu, R.; Zhao, J.; Gao, G.; Li, H. Exploiting modality-invariant feature for robust multimodal emotion recognition with missing modalities. In Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; pp. 1–5. [Google Scholar]







| Hyperparameter | Symbol | Value |
|---|---|---|
| Batch size | b | 16 |
| Epoch size | e | 100 |
| Discard rate | p | 0.1 |
| Hidden layer size | d | 256 |
| Uncertainty modal missing rate | [0–50%] | |
| Learning rate | 0.0003 | |
| Maximum text length | 30 | |
| Maximum audio length | 230 | |
| Maximum video length | 480 | |
| Training weights | 0.8, 0.5, 0.5 |
| Dataset | Models | 0% | 10% | 20% | 30% | 40% | 50% | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | ||
| MIntRec | MAG_Bert | 72.65 | 68.64 | 68.41 | 66.12 | 65.23 | 63.75 | 63.86 | 62.03 | 61.59 | 60.59 | 60.23 | 56.51 |
| Mult | 72.52 | 69.25 | 67.95 | 64.56 | 64.77 | 62.18 | 63.86 | 61.59 | 60.45 | 60.66 | 58.86 | 57.11 | |
| Misa | 72.29 | 69.32 | 68.41 | 65.65 | 63.64 | 61.67 | 62.95 | 59.44 | 56.82 | 56.33 | 58.86 | 55.75 | |
| IF-MMIN | 67.86 | 66.07 | 64.96 | 63.16 | 62.05 | 60.26 | 61.57 | 59.24 | 59.03 | 55.17 | 58.86 | 53.52 | |
| EMRFM | 72.58 | 70.46 | 68.41 | 66.81 | 62.05 | 65.95 | 63.18 | 63.14 | 60.00 | 59.00 | 57.95 | 58.18 | |
| Ours | 73.86 | 71.09 | 70.00 | 67.68 | 66.14 | 66.20 | 65.91 | 64.48 | 62.73 | 62.27 | 61.82 | 60.12 | |
| EMOTyDA | MAG_Bert | 56.45 | 46.39 | 55.24 | 43.52 | 53.83 | 42.38 | 51.21 | 41.68 | 48.79 | 39.48 | 52.82 | 38.61 |
| Mult | 59.48 | 46.36 | 55.44 | 44.03 | 54.03 | 41.64 | 54.64 | 41.10 | 51.81 | 40.62 | 50.00 | 38.57 | |
| Misa | 59.27 | 45.32 | 57.66 | 44.10 | 56.05 | 42.57 | 55.65 | 41.69 | 46.77 | 40.82 | 57.86 | 38.14 | |
| IF-MMIN | 53.83 | 35.58 | 54.44 | 34.47 | 52.62 | 33.83 | 50.00 | 29.26 | 44.76 | 27.76 | 48.19 | 25.82 | |
| EMRFM | 59.07 | 46.68 | 58.06 | 44.08 | 55.24 | 41.73 | 53.83 | 40.10 | 52.42 | 39.50 | 51.61 | 37.94 | |
| Ours | 59.88 | 47.25 | 58.47 | 44.60 | 55.65 | 43.25 | 55.04 | 42.70 | 53.23 | 40.94 | 53.63 | 40.70 | |
| Models | 0% | 10% | 20% | 30% | 40% | 50% | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | ACC | F1 | |
| (1) T | 70.00 | 66.14 | ||||||||||
| (2) V | 16.82 | 3.29 | ||||||||||
| (3) A | 16.14 | 8.02 | ||||||||||
| (4) T+A | 70.91 | 68.21 | 70.68 | 66.80 | 67.27 | 65.44 | 63.37 | 63.96 | 60.46 | 60.14 | 55.68 | 57.34 |
| (5) T+V | 67.73 | 66.18 | 67.05 | 64.38 | 66.14 | 63.82 | 60.00 | 60.15 | 55.24 | 55.21 | 51.59 | 53.25 |
| (6) A+V | 17.50 | 12.39 | 18.18 | 9.26 | 18.64 | 9.27 | 15.45 | 8.01 | 17.50 | 7.87 | 15.68 | 5.41 |
| (7) T+A+V | 73.86 | 71.09 | 70.00 | 67.68 | 66.14 | 66.20 | 65.91 | 64.48 | 62.73 | 62.27 | 61.82 | 60.12 |
| (8) -Fusion | 73.18 | 69.10 | 69.09 | 66.82 | 65.00 | 64.73 | 62.73 | 61.08 | 61.59 | 59.55 | 60.68 | 58.28 |
| (9) -Pre | 72.27 | 70.23 | 69.55 | 65.93 | 63.86 | 64.59 | 63.64 | 62.43 | 62.05 | 60.12 | 60.23 | 59.88 |
| (10) -Prompt | 71.36 | 69.32 | 67.95 | 65.60 | 64.77 | 63.88 | 63.41 | 62.31 | 62.05 | 60.96 | 59.77 | 59.40 |
| Missing Rate | Fine-Grained | Coarse-Grained | |||
|---|---|---|---|---|---|
| ACC | F1 | ACC | F1 | ||
| (1) | 0% | 73.86 | 71.09 | 90.45 | 90.37 |
| (2) | 10% | 70.00 | 67.68 | 89.55 | 89.37 |
| (3) | 20% | 66.14 | 66.20 | 88.64 | 88.46 |
| (4) | 30% | 65.91 | 64.48 | 87.50 | 87.32 |
| (5) | 40% | 62.73 | 62.27 | 87.27 | 87.01 |
| (6) | 50% | 61.82 | 60.12 | 86.82 | 86.63 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Shang, L.; Liu, Z.; Wu, Y.; Song, X.; Yu, J.; Sheng, Q.Z. Attention Bidirectional Gated Fusion Based Multimodal Intent Recognition Under Uncertain Missing Modalities. Information 2026, 17, 728. https://doi.org/10.3390/info17080728
Shang L, Liu Z, Wu Y, Song X, Yu J, Sheng QZ. Attention Bidirectional Gated Fusion Based Multimodal Intent Recognition Under Uncertain Missing Modalities. Information. 2026; 17(8):728. https://doi.org/10.3390/info17080728
Chicago/Turabian StyleShang, Ling, Zhizhong Liu, Yuxuan Wu, Xiaoyu Song, Jian Yu, and Quan Z. Sheng. 2026. "Attention Bidirectional Gated Fusion Based Multimodal Intent Recognition Under Uncertain Missing Modalities" Information 17, no. 8: 728. https://doi.org/10.3390/info17080728
APA StyleShang, L., Liu, Z., Wu, Y., Song, X., Yu, J., & Sheng, Q. Z. (2026). Attention Bidirectional Gated Fusion Based Multimodal Intent Recognition Under Uncertain Missing Modalities. Information, 17(8), 728. https://doi.org/10.3390/info17080728

