HISF: Hierarchical Interactive Semantic Fusion for Multimodal Prompt Learning
Abstract
1. Introduction
- Hierarchical Semantic Fusion Framework:
- 2.
- Label-Guided Cross-Attention Mechanism:
- 3.
- Semantic Alignment and Label Constraints:
- 4.
- Parameter-Efficient and Robust Adaptation:
2. Related Work
2.1. Vision-Language Pre-Training
2.2. Prompt Learning
2.3. Multimodal Prompt Learning
2.4. Semantic Alignment and Label Embedding
2.5. Summary
3. Method
3.1. Overall Framework
3.2. Prompt Initialization
3.3. Hierarchical Semantic Fusion
3.4. Label Embedding and Semantic Alignment
3.5. Complexity Analysis
3.6. Optimization and Training Strategy
3.7. Discussion
4. Experiments
4.1. Experimental Hyperparameters
4.2. Base-to-Novel Generalization Evaluation
4.3. Cross-Dataset Transfer Evaluation
4.4. Domain Generalization Evaluation
4.5. Ablation Studies
4.5.1. Prompt Insertion Depth
4.5.2. HISF Component Effectiveness
4.5.3. Loss Component Influence
4.6. Summary
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| HISF | Hierarchical Interactive Semantic Fusion |
| PI | Prompt Initialization |
| LE | Label Embedding |
| LA | Label Alignment |
References
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, Virtual Event, 18–24 July 2021; pp. 8748–8763. [Google Scholar]
- Chen, F.; Zhang, D.; Han, M.; Chen, X.; Shi, J.; Xu, S.; Xu, B. VLP: A survey on vision-language pre-training. Mach. Intell. Res. 2023, 20, 38–56. [Google Scholar] [CrossRef]
- Gu, Y.; Han, X.; Liu, Z.; Huang, M. PPT: Pre-trained prompt tuning for few-shot learning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Dublin, Ireland, 22–27 May 2022; pp. 8410–8423. [Google Scholar] [CrossRef]
- Zhou, K.; Yang, J.; Loy, C.C.; Liu, Z. Learning to prompt for vision-language models. Int. J. Comput. Vis. 2022, 130, 2337–2348. [Google Scholar] [CrossRef]
- Zhou, K.; Yang, J.; Loy, C.C.; Liu, Z. Conditional prompt learning for vision-language models. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 16795–16804. [Google Scholar] [CrossRef]
- Khattak, M.U.; Rasheed, H.; Maaz, M.; Khan, S.; Khan, F.S. MaPLe: Multi-modal prompt learning. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 19113–19122. [Google Scholar] [CrossRef]
- Gao, J.; Ruan, J.; Xiang, S.; Yu, Z.; Ji, K.; Xie, M.; Liu, T.; Fu, Y. LAMM: Label alignment for multi-modal prompt learning. Proc. AAAI Conf. Artif. Intell. 2024, 38, 1815–1823. [Google Scholar] [CrossRef]
- Liu, Y.; Deng, Y.; Liu, A.; Liu, Y.; Li, S. Fine-grained multi-modal prompt learning for vision–language models. Neurocomputing 2025, 636, 130028. [Google Scholar] [CrossRef]
- Guo, Y.; Gu, X. MMRL: Multi-modal representation learning for vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 25015–25025. [Google Scholar]
- Chen, S.; Ge, C.; Tong, Z.; Wang, J.; Song, Y.; Wang, J.; Luo, P. AdaptFormer: Adapting vision transformers for scalable visual recognition. In Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; pp. 16664–16678. [Google Scholar]
- Du, Y.; Liu, Z.; Li, J.; Zhao, W.X. A survey of vision-language pre-trained models. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence (IJCAI 2022), Vienna, Austria, 23–29 July 2022; pp. 5436–5443. [Google Scholar] [CrossRef]
- Schuhmann, C.; Vencu, R.; Beaumont, R.; Kaczmarczyk, R.; Mullis, C.; Jitsev, J.; Komatsuzaki, A. LAION-400M: Open dataset of clip-filtered 400 million image-text pairs. In Proceedings of the 35th International Conference on Neural Information Processing Systems (NeurIPS 2021), Virtual Event, 6–14 December 2021; pp. 1–5. [Google Scholar]
- Yang, A.; Pan, J.; Lin, J.; Men, R.; Zhang, J.; Zhou, J.; Zhou, C. Chinese CLIP: Contrastive vision-language pretraining in Chinese. arXiv 2022, arXiv:2211.01335. [Google Scholar] [CrossRef]
- Li, X.; Yin, X.; Li, C.; Zhang, P.; Hu, X.; Zhang, L.; Wang, L.; Hu, H.; Dong, L.; Wei, F.; et al. OSCAR: Object-semantics aligned pre-training for vision-language tasks. In Proceedings of the 16th European Conference on Computer Vision—ECCV 2020, Glasgow, UK, 23–28 August 2020; pp. 121–137. [Google Scholar] [CrossRef]
- Li, J.; Li, D.; Savarese, S.; Hoi, S. BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July, 2023; pp. 19730–19742. [Google Scholar]
- Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual instruction tuning. In Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, LA, USA, 10–16 December 2023; pp. 34892–34916. [Google Scholar]
- Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; Zhang, L. Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 6077–6086. [Google Scholar] [CrossRef]
- Jia, C.; Yang, Y.; Xia, Y.; Chen, Y.; Parekh, Z.; Pham, H.; Le, Q.; Sung, Y.H.; Li, Z.; Duerig, T. Scaling up visual and vision-language representation learning with noisy text supervision. In Proceedings of the 38th International Conference on Machine Learning, Virtual Event, 18–24 July 2021; pp. 4904–4916. [Google Scholar]
- Zhai, X.; Wang, X.; Mustafa, B.; Steiner, A.; Keysers, D.; Kolesnikov, A.; Beyer, L. LIT: Zero-shot transfer with locked-image text tuning. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVP), New Orleans, LA, USA, 18–24 June 2022; pp. 18102–18112. [Google Scholar] [CrossRef]
- Gao, P.; Geng, S.; Zhang, R.; Ma, T.; Fang, R.; Zhang, Y.; Li, H.; Qiao, Y. CLIP-Adapter: Better vision-language models with feature adapters. Int. J. Comput. Vis. 2024, 132, 581–595. [Google Scholar] [CrossRef]
- Cho, J.; Nam, G.; Kim, S.; Yang, H.; Kwak, S. PromptStyler: Prompt-driven style generation for source-free domain generalization. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 15656–15666. [Google Scholar] [CrossRef]
- Lu, Y.; Liu, J.; Zhang, Y.; Liu, Y.; Tian, X. Prompt distribution learning. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 5196–5205. [Google Scholar] [CrossRef]
- Sun, X.; Hu, P.; Saenko, K. DualCoOp: Fast adaptation to multi-label recognition with limited annotations. In Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; pp. 30569–30582. [Google Scholar]
- Xie, J.; Zhang, Y.; Peng, J.; Huang, Z.; Cao, L. TextRefiner: Internal visual feature as efficient refiner for vision-language models prompt tuning. Proc. AAAI Conf. Artif. Intell. 2025, 39, 8718–8726. [Google Scholar] [CrossRef]
- Yao, Y.; Zhang, A.; Zhang, Z.; Liu, Z.; Chua, T.S.; Sun, M. CPT: Colorful prompt tuning for pre-trained vision-language models. AI Open 2024, 5, 30–38. [Google Scholar] [CrossRef]
- Zhang, J.; Wang, S. Text-guided visual prompt learning with semantic prompt generation and feature fusion. Neurocomputing 2025, 654, 131253. [Google Scholar] [CrossRef]
- Bahng, H.; Jahanian, A.; Sankaranarayanan, S.; Isola, P. Exploring visual prompts for adapting large-scale models. arXiv 2022, arXiv:2203.17274. [Google Scholar] [CrossRef]
- Roy, S.; Etemad, A. Consistency-guided prompt learning for vision-language models. arXiv 2023, arXiv:2306.01195. [Google Scholar] [CrossRef]
- Zhu, J.; Ruan, Y.; Chang, J.; Sun, W.; Wan, H.; Long, J.; Luo, C. Deep prompt multi-task network for abuse language detection. In Proceedings of the 27th International Conference on Pattern Recognition, Kolkata, India, 1–5 December 2024; pp. 249–263. [Google Scholar] [CrossRef]
- Wang, Z.; Zhang, Z.; Ebrahimi, S.; Sun, R.; Zhang, H.; Lee, C.Y.; Ren, X.; Su, G.; Perot, V.; Dy, J.; et al. DualPrompt: Complementary prompting for rehearsal-free continual learning. In Proceedings of the 17th European Conference on Computer Vision—ECCV 2022, Tel Aviv, Israel, 23–27 October 2022; pp. 631–648. [Google Scholar] [CrossRef]
- Yin, H.; Zhao, Y. Multi-modal prompt learning with bidirectional layer-wise prompt fusion. Inf. Fusion 2025, 117, 102919. [Google Scholar] [CrossRef]
- Zhao, F.; Zhang, C.; Geng, B. Deep Multimodal Data Fusion. ACM Comput. Surv. 2024, 56, 216. [Google Scholar] [CrossRef]
- Li, Y.; Quan, R.; Zhu, L.; Yang, Y. Efficient multimodal fusion via interactive prompting. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 2604–2613. [Google Scholar] [CrossRef]
- Yang, L.; Zhang, R.; Wang, Y.; Xie, X. MMA: Multi-modal adapter for vision-language models. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 23826–23837. [Google Scholar] [CrossRef]
- Wu, Q.; Yu, W.; Zhou, Y.; Huang, S.; Sun, X.; Ji, R. Parameter and computation efficient transfer learning for vision-language pre-trained models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (NeurIPS 2023), New Orleans, LA, USA, 10–16 December 2023; pp. 41034–41050. [Google Scholar]
- Xin, Y.; Du, J.; Wang, Q.; Yan, K.; Ding, S. MmAP: Multi-modal alignment prompt for cross-domain multi-task learning. Proc. AAAI Conf. Artif. Intell. 2024, 38, 16076–16084. [Google Scholar] [CrossRef]
- Özdemir, Ö.; Akagündüz, E. Enhancing visual question answering through question-driven image captions as prompts. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, 17–18 June 2024; pp. 1562–1571. [Google Scholar] [CrossRef]
- Wang, Y.; Jiang, X.; Cheng, D.; Li, D.; Zhao, C. Learning hierarchical prompt with structured linguistic knowledge for vision-language models. Proc. AAAI Conf. Artif. Intell. 2024, 38, 5749–5757. [Google Scholar] [CrossRef]
- Yang, L.; Zhang, R.; Chen, Q.; Xie, X. Learning with enriched inductive biases for vision-language models. Int. J. Comput. Vis. 2025, 133, 3746–3761. [Google Scholar] [CrossRef]
- He, S.; Wang, S.; Long, S. A slim prompt-averaged consistency prompt learning for vision–language model. Knowl. Based Syst. 2025, 310, 113011. [Google Scholar] [CrossRef]
- Wang, G.; Ge, Y.; Ding, X.; Kankanhalli, M.; Shan, Y. What makes for good visual tokenizers for large language models? arXiv 2023, arXiv:2305.12223. [Google Scholar] [CrossRef]
- Li, Z.; Li, X.; Fu, X.; Zhang, X.; Wang, W.; Chen, S.; Yang, J. PromptKD: Unsupervised prompt distillation for vision-language models. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 26607–26616. [Google Scholar] [CrossRef]
- Zhao, C.; Wang, Y.; Jiang, X.; Shen, Y.; Song, K.; Li, D.; Miao, A. Learning domain invariant prompt for vision-language models. IEEE Trans. Image Process. 2024, 33, 1348–1360. [Google Scholar] [CrossRef] [PubMed]
- Xu, C.; Zhu, Y.; Shen, H.; Chen, B.; Liao, Y.; Chen, X.; Wang, L. Progressive visual prompt learning with contrastive feature re-formation. Int. J. Comput. Vis. 2025, 133, 511–526. [Google Scholar] [CrossRef]
- Yao, H.; Zhang, R.; Xu, C. Visual-language prompt tuning with knowledge-guided context optimization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 6757–6767. [Google Scholar] [CrossRef]
- Khattak, M.U.; Wasim, S.T.; Naseer, M.; Khan, S.; Yang, M.-H.; Khan, F.S. Self-regulating prompts: Foundational model adaptation without forgetting. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 15190–15200. [Google Scholar] [CrossRef]
- Yao, H.; Zhang, R.; Xu, C. TCP: Textual-based class-aware prompt tuning for visual-language model. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 23438–23448. [Google Scholar] [CrossRef]



| Method | Average | ImageNet | Caltech101 | OxfordPets | ||||||||
| Base | Novel | HM | Base | Novel | HM | Base | Novel | HM | Base | Novel | HM | |
| CLIP [1] | 69.34 | 74.22 | 71.70 | 72.43 | 68.14 | 70.22 | 96.84 | 94.00 | 95.40 | 91.17 | 97.26 | 94.12 |
| CoOP [4] | 82.69 | 63.22 | 71.66 | 76.47 | 67.88 | 71.92 | 98.00 | 89.81 | 93.73 | 93.67 | 95.29 | 94.47 |
| CoOpOp [5] | 80.47 | 71.69 | 75.83 | 75.98 | 70.43 | 73.10 | 97.96 | 93.81 | 95.84 | 95.20 | 97.69 | 96.43 |
| ProDA [22] | 81.56 | 72.30 | 76.65 | 75.40 | 70.23 | 72.72 | 98.27 | 93.23 | 95.68 | 95.43 | 97.83 | 96.62 |
| KgCoOp [45] | 80.73 | 73.60 | 77.00 | 75.83 | 69.96 | 72.78 | 97.72 | 94.39 | 96.03 | 94.65 | 97.76 | 96.18 |
| MaPLe [6] | 82.28 | 75.14 | 78.55 | 76.66 | 70.54 | 73.47 | 97.74 | 94.36 | 96.02 | 95.43 | 97.76 | 96.58 |
| PromptSRC [46] | 84.26 | 76.10 | 79.97 | 77.60 | 70.73 | 74.01 | 98.10 | 94.03 | 96.02 | 95.33 | 97.30 | 96.30 |
| ProVP [44] | 85.20 | 73.22 | 78.76 | 75.82 | 69.21 | 72.36 | 98.92 | 94.21 | 96.51 | 95.87 | 97.65 | 96.75 |
| MP [43] | 83.65 | 75.48 | 79.09 | 77.52 | 70.83 | 74.02 | 98.13 | 94.58 | 96.32 | 95.53 | 97.00 | 96.26 |
| TCP [47] | 84.13 | 75.36 | 79.51 | 77.27 | 69.87 | 73.38 | 98.23 | 94.67 | 96.42 | 94.67 | 97.20 | 95.92 |
| MMA [34] | 83.20 | 76.80 | 79.87 | 77.31 | 71.00 | 74.02 | 98.40 | 94.00 | 96.15 | 95.40 | 98.07 | 96.72 |
| HISF | 85.39 | 76.12 | 80.75 | 76.8 | 70.1 | 73.4 | 98.20 | 94.6 | 96.4 | 95.7 | 96.9 | 96.3 |
| Method | StanfordCars | Flower102 | Food101 | FGVCAircraft | ||||||||
| Base | Novel | HM | Base | Novel | HM | Base | Novel | HM | Base | Novel | HM | |
| CLIP [1] | 63.37 | 74.89 | 68.65 | 72.08 | 77.80 | 74.83 | 90.10 | 91.22 | 90.66 | 27.19 | 36.29 | 31.09 |
| CoOP [4] | 78.12 | 60.40 | 68.13 | 97.60 | 59.67 | 74.06 | 88.33 | 82.26 | 85.19 | 40.44 | 22.30 | 28.75 |
| CoOpOp [5] | 70.49 | 73.59 | 72.01 | 94.87 | 71.75 | 81.71 | 90.70 | 91.29 | 90.99 | 33.41 | 23.71 | 27.74 |
| ProDA [22] | 74.70 | 71.20 | 72.91 | 97.70 | 68.68 | 80.66 | 90.30 | 88.57 | 89.43 | 36.90 | 34.13 | 35.46 |
| KgCoOp [45] | 71.76 | 75.04 | 73.36 | 95.00 | 74.73 | 83.65 | 90.50 | 91.70 | 91.09 | 36.21 | 33.55 | 34.83 |
| MaPLe [6] | 72.94 | 74.00 | 73.47 | 95.92 | 72.46 | 82.56 | 90.71 | 92.05 | 91.38 | 37.44 | 35.61 | 36.50 |
| PromptSRC [46] | 78.27 | 74.97 | 76.58 | 98.07 | 76.50 | 85.95 | 90.67 | 91.53 | 91.10 | 42.73 | 37.87 | 40.15 |
| ProVP [44] | 80.43 | 67.96 | 73.67 | 98.42 | 72.06 | 83.20 | 90.32 | 90.91 | 90.61 | 47.08 | 29.87 | 36.55 |
| MP [43] | 76.34 | 75.01 | 75.48 | 97.66 | 74.49 | 84.52 | 90.74 | 91.85 | 91.29 | 40.14 | 36.51 | 38.24 |
| TCP [47] | 80.80 | 74.13 | 77.32 | 97.73 | 75.57 | 85.23 | 90.57 | 91.37 | 90.97 | 41.97 | 34.43 | 37.83 |
| MMA [34] | 78.50 | 73.10 | 75.70 | 97.77 | 75.93 | 85.48 | 90.13 | 91.30 | 90.71 | 40.57 | 36.33 | 38.33 |
| HISF | 81.9 | 74.3 | 78.1 | 98.6 | 74.2 | 86.4 | 89.5 | 90.7 | 90.1 | 47.6 | 33.8 | 40.7 |
| Method | SUN397 | DTD | EuroSAT | UCF101 | ||||||||
| Base | Novel | HM | Base | Novel | HM | Base | Novel | HM | Base | Novel | HM | |
| CLIP [1] | 69.36 | 75.35 | 72.23 | 53.24 | 59.90 | 56.37 | 56.48 | 64.05 | 60.03 | 70.53 | 77.50 | 73.85 |
| CoOP [4] | 80.60 | 65.89 | 72.51 | 79.44 | 41.18 | 54.24 | 92.19 | 54.74 | 68.69 | 84.69 | 56.05 | 67.46 |
| CoOpOp [5] | 79.74 | 76.86 | 78.27 | 77.01 | 56.00 | 64.85 | 87.49 | 60.04 | 71.21 | 82.33 | 73.45 | 77.64 |
| ProDA [22] | 78.67 | 76.93 | 77.79 | 80.67 | 56.48 | 66.44 | 83.90 | 66.00 | 73.88 | 85.23 | 71.97 | 78.04 |
| KgCoOp [45] | 80.29 | 76.53 | 78.36 | 77.55 | 54.99 | 64.35 | 85.64 | 64.34 | 73.48 | 82.89 | 76.67 | 79.65 |
| MaPLe [6] | 80.82 | 78.70 | 79.75 | 80.36 | 59.18 | 68.16 | 94.07 | 73.23 | 82.35 | 83.00 | 78.66 | 80.77 |
| PromptSRC [46] | 82.67 | 78.47 | 80.52 | 83.37 | 62.97 | 71.75 | 92.90 | 73.90 | 82.32 | 87.10 | 78.80 | 82.74 |
| ProVP [44] | 80.67 | 76.11 | 78.32 | 83.95 | 59.06 | 69.34 | 97.12 | 72.91 | 83.29 | 88.56 | 75.55 | 81.54 |
| MP [43] | 82.26 | 79.04 | 80.62 | 83.10 | 58.05 | 68.35 | 93.53 | 75.21 | 83.38 | 85.33 | 77.72 | 81.35 |
| TCP [47] | 82.63 | 78.20 | 80.35 | 82.77 | 58.07 | 68.25 | 91.63 | 74.73 | 82.32 | 87.13 | 80.77 | 83.83 |
| MMA [34] | 82.27 | 78.57 | 80.38 | 83.20 | 65.63 | 73.38 | 85.46 | 82.34 | 83.87 | 86.23 | 80.03 | 82.20 |
| HISF | 83.5 | 79.3 | 81.4 | 83.6 | 64.8 | 74.2 | 97.3 | 77.5 | 87.4 | 86.6 | 81.2 | 83.9 |
| Method | Source | Target | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| ImageNet | Caltech | OxfordPets | Stanford Cars | Flowers 101 | Food101 | FGVC Aircraft | SUN397 | DTD | EuroSAT | UCF101 | |
| CoOp [4] | 71.51 | 93.70 | 89.14 | 64.51 | 68.71 | 85.30 | 18.47 | 68.17 | 41.92 | 46.39 | 66.55 |
| CoOpOp [5] | 71.02 | 94.43 | 90.14 | 65.32 | 71.88 | 86.06 | 22.94 | 67.36 | 45.73 | 45.37 | 68.21 |
| MaPLe [6] | 70.72 | 93.53 | 90.49 | 65.57 | 72.23 | 86.20 | 24.74 | 67.01 | 46.49 | 48.06 | 68.69 |
| PromptSRC [46] | 71.27 | 93.60 | 90.25 | 65.70 | 70.25 | 86.15 | 23.90 | 67.10 | 46.87 | 45.50 | 68.75 |
| TCP [47] | 71.40 | 93.97 | 91.25 | 64.49 | 71.21 | 86.69 | 23.45 | 67.15 | 44.35 | 51.45 | 68.73 |
| MMA [34] | 71.00 | 93.80 | 90.30 | 66.13 | 72.07 | 86.12 | 25.33 | 68.17 | 46.57 | 49.24 | 68.32 |
| HISF | 72.20 | 93.49 | 90.32 | 65.40 | 72.41 | 85.41 | 24.31 | 67.57 | 45.91 | 54.70 | 68.41 |
| Method | Source | Target | |||
|---|---|---|---|---|---|
| ImageNet | −S | −A | −R | −V2 | |
| CLIP [1] | 66.73 | 46.15 | 47.77 | 73.96 | 60.83 |
| CoOp [4] | 71.51 | 47.99 | 49.71 | 75.21 | 64.20 |
| CoOpOp [5] | 71.02 | 48.75 | 50.63 | 76.18 | 64.07 |
| Maple [6] | 70.72 | 49.15 | 50.90 | 76.98 | 64.07 |
| PromptSRC [46] | 71.27 | 49.55 | 50.90 | 77.80 | 64.35 |
| HISF | 72.20 | 49.51 | 51.20 | 78.21 | 64.43 |
| Method | Base | Novel | H.M |
|---|---|---|---|
| HISF (Maple setting) | 82.28 | 75.14 | 78.55 |
| HISF (w/o PI) | 83.76 | 75.57 | 79.66 |
| HISF (w/o SF) | 82.45 | 74.96 | 78.71 |
| HISF (PI + SF) | 85.39 | 76.12 | 80.75 |
| LCE | LS | LLP | H.M |
|---|---|---|---|
| √ | 79.41 | ||
| √ | √ | √ | 80.75 |
| √ | √ | 80.67 | |
| √ | √ | 80.35 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Feng, H.; Li, C. HISF: Hierarchical Interactive Semantic Fusion for Multimodal Prompt Learning. Multimodal Technol. Interact. 2026, 10, 6. https://doi.org/10.3390/mti10010006
Feng H, Li C. HISF: Hierarchical Interactive Semantic Fusion for Multimodal Prompt Learning. Multimodal Technologies and Interaction. 2026; 10(1):6. https://doi.org/10.3390/mti10010006
Chicago/Turabian StyleFeng, Haohan, and Chen Li. 2026. "HISF: Hierarchical Interactive Semantic Fusion for Multimodal Prompt Learning" Multimodal Technologies and Interaction 10, no. 1: 6. https://doi.org/10.3390/mti10010006
APA StyleFeng, H., & Li, C. (2026). HISF: Hierarchical Interactive Semantic Fusion for Multimodal Prompt Learning. Multimodal Technologies and Interaction, 10(1), 6. https://doi.org/10.3390/mti10010006
