IE-MAS: Internal–External Multi-Agent Steering for Controllable Image Captioning
Abstract
1. Introduction
- Internal Multimodal Steering (IMS) for fine-grained sentiment control. IMS extracts sentiment-related steering vectors and injects these vectors into the cross-modal attention layers of the base MLLM, enabling fine-grained modulation of affective tone without retraining.
- External Multi-Agent Collaboration System (EMCS) for multi-constraint controllable captioning. The EMCS balances conflicts from multiple constraints by the novel Contactor, who translates high-level external feedback into internal steering intensity. By integrating IMS and the EMCS, the proposed IE-MAS achieves adaptive coordination between internal representational control and external agent collaboration.
- Comprehensive experimental validation. Extensive experiments under different sentiment types and length settings demonstrate that IE-MAS consistently outperforms strong baselines. Layer analysis further confirms the effectiveness of internal modulation, while ablation studies verify the complementary roles of IMS and the EMCS in achieving balanced and stable Controllable Image Captioning.
2. Related Work
2.1. Controllable Image Captioning
2.2. Steering for Controllable Image Captioning
2.3. Multi-Agent System for Controllable Image Captioning
3. Preliminaries
3.1. Problem Formulation
3.2. Steering
3.3. Multi-Agent System
4. Methodology
4.1. Internal Control: Multimodal Activation Steering for Cic
4.2. External Control: Multi-Agent Collaboration System for Cic
| Algorithm 1 The proposed IE-MAS framework |
|
4.3. Information-Theoretic Interpretation of IE-MAS
5. Experiment
5.1. Experimental Settings
5.2. Analysis of Main Results
5.3. Ablation Study: Component Contributions to Multi-Constraint Control
5.4. Internal Layer Analysis
5.5. Sentiment-Specific SAE Feature Distribution
5.6. Sensitivity Analysis
5.7. Information-Theoretic Analysis
5.8. Qualitative Comparison
6. Conclusions and Future Work
6.1. Conclusions
6.2. Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Farhadi, A.; Hejrati, M.; Sadeghi, M.; Young, P.; Rashtchian, C.; Hockenmaier, J.; Forsyth, D. Every picture tells a story: Generating sentences from images. In Proceedings of the European Conference on Computer Vision, Crete, Greece, 5–11 September 2010; Springer: Berlin/Heidelberg, Germany, 2010; pp. 15–29. [Google Scholar]
- Chen, S.; Jin, Q.; Wang, P.; Wu, Q. Say as you wish: Fine-grained control of image caption generation with abstract scene graphs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 9962–9971. [Google Scholar]
- Wang, N.; Xie, J.; Wu, J.; Jia, M.; Li, L. Controllable image captioning via prompting. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; Volume 37, pp. 2617–2625. [Google Scholar]
- Zeng, Z.; Zhang, H.; Lu, R.; Wang, D.; Chen, B.; Wang, Z. Conzic: Controllable zero-shot image captioning by sampling-based polishing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 23465–23476. [Google Scholar]
- Danescu-Niculescu-Mizil, C.; Gamon, M.; Dumais, S. Mark my words! Linguistic style accommodation in social media. In Proceedings of the 20th International Conference on World Wide Web, Hyderabad, India, 28 March–1 April 2011; pp. 745–754. [Google Scholar]
- Ghandi, T.; Pourreza, H.; Mahyar, H. Deep learning approaches on image captioning: A review. ACM Comput. Surv. 2023, 56, 1–39. [Google Scholar] [CrossRef] [Scilit]
- Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. Qwen2.5-vl technical report. arXiv 2025, arXiv:2502.13923. [Google Scholar] [CrossRef] [Scilit]
- Meta, A. Llama 3.2: Revolutionizing edge ai and vision with open, customizable models. Meta AI Blog. Retrieved Dec. 2024, 20, 2024. [Google Scholar]
- Wang, M.; Xu, Z.; Mao, S.; Deng, S.; Tu, Z.; Chen, H.; Zhang, N. Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 27 July–1 August 2025; Che, W., Nabende, J., Shutova, E., Pilehvar, M.T., Eds.; pp. 23381–23399. [Google Scholar] [CrossRef] [Scilit]
- Deng, C.; Ding, N.; Tan, M.; Wu, Q. Length-controllable image captioning. In Proceedings of the European conference on computer vision, Glasgow, UK, 23–28 August 2020; Springer: Berlin/Heidelberg, Germany, 2020; pp. 712–729. [Google Scholar]
- Wang, N.; Duan, F.; Zhang, Y.; Zhou, W.; Xu, K.; Huang, W.; Fu, J. PositionID: LLMs can Control Lengths, Copy and Paste with Explicit Positional Awareness. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, FL, USA, 12–16 November 2024; Association for Computational Linguistics: Stroudsburg, PA, USA, 2024; pp. 16877–16915. [Google Scholar]
- Butcher, B.; O’Keefe, M.; Titchener, J. Precise length control for large language models. Nat. Lang. Process. J. 2025, 11, 100143. [Google Scholar] [CrossRef] [Scilit]
- Gu, Y.; Wang, W.; Feng, X.; Zhong, W.; Zhu, K.; Huang, L.; Liu, T.; Qin, B. Length Controlled Generation for Black-box LLMs. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, 27 July–1 August 2025; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 16878–16895. [Google Scholar]
- Retkowski, F.; Waibel, A. Zero-Shot Strategies for Length-Controllable Summarization. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, NM, USA, 29 April–4 May 2025; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 551–572. [Google Scholar]
- Sweller, J. Cognitive load during problem solving: Effects on learning. Cogn. Sci. 1988, 12, 257–285. [Google Scholar] [CrossRef]
- Bai, L.; Borah, A.; Ignat, O.; Mihalcea, R. The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Albuquerque, NM, USA, 29 April–4 May 2025; pp. 2970–2993. [Google Scholar]
- Lee, S.; Yoon, S.; Bui, T.; Shi, J.; Yoon, S. Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage. In Proceedings of the Forty-second International Conference on Machine Learning, Vancouver, WC, Canada, 13–19 July 2025. [Google Scholar]
- Yang, P.; Dong, B. Mocoll: Agent-based specific and general model collaboration for image captioning. arXiv 2025, arXiv:2501.01834. [Google Scholar]
- Shen, Y.; Song, K.; Tan, X.; Li, D.; Lu, W.; Zhuang, Y. Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face. Adv. Neural Inf. Process. Syst. 2023, 36, 38154–38180. [Google Scholar]
- Zhang, X.; Dong, X.; Wang, Y.; Zhang, D.; Cao, F. A Survey of Multi-AI Agent Collaboration: Theories, Technologies and Applications; Association for Computing Machinery: New York, NY, USA, 2025. [Google Scholar]
- Yang, J.; Sun, Y.; Liang, J.; Ren, B.; Lai, S.H. Image captioning by incorporating affective concepts learned from both visual and textual components. Neurocomputing 2019, 328, 56–68. [Google Scholar] [CrossRef] [Scilit]
- Shannon, C.E. A mathematical theory of communication. Bell Syst. Tech. J. 1948, 27, 379–423. [Google Scholar] [CrossRef] [Scilit]
- Elements of Information Theory; John Wiley & Sons: Hoboken, NJ, USA, 1999.
- Li, J.; Selvaraju, R.; Gotmare, A.; Joty, S.; Xiong, C.; Hoi, S.C.H. Align before fuse: Vision and language representation learning with momentum distillation. Adv. Neural Inf. Process. Syst. 2021, 34, 9694–9705. [Google Scholar]
- Mathews, A.; Xie, L.; He, X. Senticap: Generating image descriptions with sentiments. In Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AR, USA, 12–17 February 2016; Volume 30. [Google Scholar]
- Gan, C.; Gan, Z.; He, X.; Gao, J.; Deng, L. Stylenet: Generating attractive visual captions with styles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 3137–3146. [Google Scholar]
- Guo, L.; Liu, J.; Yao, P.; Li, J.; Lu, H. MSCap: Multi-Style Image Captioning with Unpaired Stylized Text. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019. [Google Scholar]
- Liu, Z.; Lin, W.; Shi, Y.; Zhao, J. A robustly optimized BERT pre-training approach with post-training. In Proceedings of the China National Conference on Chinese Computational Linguistics, Hohhot, China, 13–15 August 2021; Springer: Berlin/Heidelberg, Germany, 2021; pp. 471–484. [Google Scholar]
- Radford, A.; Kim, J.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, PMLR, Online, 18–24 July 2021; pp. 8748–8763. [Google Scholar]
- Li, J.; Li, D.; Savarese, S.; Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning, PMLR, Honolulu, HI, USA, 23–29 July 2023; pp. 19730–19742. [Google Scholar]
- Tian, J.; Yang, Z.; Shi, S. Unsupervised style control for image captioning. In Proceedings of the International Conference of Pioneering Computer Scientists, Engineers and Educators, Chengdu, China, 19–22 August 2022; Springer: Berlin/Heidelberg, Germany, 2022; pp. 413–424. [Google Scholar]
- He, K.; Fan, H.; Wu, Y.; Xie, S.; Girshick, R. Momentum contrast for unsupervised visual representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 9729–9738. [Google Scholar]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; Riedmiller, M. Playing atari with deep reinforcement learning. arXiv 2013, arXiv:1312.5602. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit]
- Wei, J.; Bosma, M.; Zhao, V.Y.; Guu, K.; Yu, A.W.; Lester, B.; Du, N.; Dai, A.M.; Le, Q.V. Finetuned language models are zero-shot learners. arXiv 2021, arXiv:2109.01652. [Google Scholar]
- Kikuchi, Y.; Neubig, G.; Sasano, R.; Takamura, H.; Okumura, M. Controlling Output Length in Neural Encoder-Decoders. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, Texas, USA, 1–4 November 2016. [Google Scholar]
- Takase, S.; Okazaki, N. Positional Encoding to Control Output Sequence Length. In Proceedings of the 2019 Conference of the North. Association for Computational Linguistics, Minneapolis, MN, USA, 3–5 June 2019. [Google Scholar]
- Liu, P.; Yuan, W.; Fu, J.; Jiang, Z.; Hayashi, H.; Neubig, G. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM Comput. Surv. 2023, 55, 1–35. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Zhang, L.; Zhang, K.; Hu, B.; Xie, H.; Mao, Z. Cascade semantic prompt alignment network for image captioning. IEEE Trans. Circuits Syst. Video Technol. 2023, 34, 5266–5281. [Google Scholar] [CrossRef] [Scilit]
- Rimsky, N.; Gabrieli, N.; Schulz, J.; Tong, M.; Hubinger, E.; Turner, A. Steering Llama 2 via Contrastive Activation Addition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, 11–16 August 2024; pp. 15504–15522. [Google Scholar]
- Subramani, N.; Suresh, N.; Peters, M.E. Extracting Latent Steering Vectors from Pretrained Language Models. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2022, Dublin, Ireland, 22–27 May 2022; pp. 566–581. [Google Scholar]
- Turner, A.M.; Thiergart, L.; Leech, G.; Udell, D.; Vazquez, J.J.; Mini, U.; MacDiarmid, M. Steering language models with activation engineering. arXiv 2023, arXiv:2308.10248. [Google Scholar]
- Soo, S.; Teng, W.; Balaganesh, C.; Tan, G.; Yan, M. Interpretable Steering of Large Language Models with Feature Guided Activation Additions. In Proceedings of the Building Trust Workshop at ICLR 2025, Singapore, 24 April 2025. [Google Scholar]
- Bayat, R.; Rahimi-Kalahroudi, A.; Pezeshki, M.; Chandar, S.; Vincent, P. Steering large language model activations in sparse spaces. arXiv 2025, arXiv:2503.00177. [Google Scholar] [CrossRef] [Scilit]
- Su, J.; Chen, J.; Li, H.; Chen, Y.; Qing, L.; Zhang, Z. Activation steering decoding: Mitigating hallucination in large vision-language models through bidirectional hidden state intervention. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 27 July–1 August 2025; pp. 12964–12974. [Google Scholar]
- Kim, J.; Lee, J.; Choi, H.J.; Hsu, T.Y.; Huang, C.Y.; Kim, S.; Rossi, R.; Yu, T.; Giles, C.L.; Huang, T.H.; et al. Multi-LLM Collaborative Caption Generation in Scientific Documents. In Proceedings of the International Workshop on AI for Transportation, Philadelphia, PA, USA, 25 February–4 March 2025; Springer: Berlin/Heidelberg, Germany, 2025; pp. 142–160. [Google Scholar]
- Jiang, A.; Wang, D.; Peng, C.; Wang, M. Relational Reasoning Image Captioning via Multi-Agent Retrieval-Augmented Generation. Knowl.-Based Syst. 2025, 114977. [Google Scholar] [CrossRef] [Scilit]
- Wu, X.; Li, T. Sentimental visual captioning using multimodal transformer. Int. J. Comput. Vis. 2023, 131, 1073–1090. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 10012–10022. [Google Scholar]
- Elhage, N.; Hume, T.; Olsson, C.; Schiefer, N.; Henighan, T.; Kravec, S.; Hatfield-Dodds, Z.; Lasenby, R.; Drain, D.; Chen, C.; et al. Toy models of superposition. arXiv 2022, arXiv:2209.10652. [Google Scholar] [CrossRef] [Scilit]
- Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Pan, A.; Yin, X.; Mazeika, M.; Dombrowski, A.K.; et al. Representation engineering: A top-down approach to ai transparency. arXiv 2023, arXiv:2310.01405. [Google Scholar] [CrossRef] [Scilit]
- Huben, R.; Cunningham, H.; Riggs Smith, L.; Ewart, A.; Sharkey, L. Sparse Autoencoders Find Highly Interpretable Features in Language Models. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR 2024), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Wooldridge, M. An introduction to Multiagent Systems; John Wiley & Sons: Hoboken, NJ, USA, 2009. [Google Scholar]
- Ferber, J.; Weiss, G. Multi-Agent Systems: An Introduction to Distributed Artificial Intelligence; Addison-Wesley Reading: Boston, UK, 1999; Volume 1. [Google Scholar]
- Tishby, N.; Pereira, F.C.; Bialek, W. The Information Bottleneck Method. In Proceedings of the 37th Annual Allerton Conference on Communication, Control, and Computing, Monticello, IL, USA, 4–6 October 2000; pp. 368–377. [Google Scholar]
- Sanh, V.; Debut, L.; Chaumond, J.; Wolf, T. DistilBERT: A Distilled Version of BERT—Smaller, Faster, Cheaper and Lighter. In Proceedings of the 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing, Vancouver, BC, Canada, 8 December 2019. [Google Scholar]
- Hessel, J.; Holtzman, A.; Forbes, M.; Choi, Y. CLIPScore: A Reference-Free Evaluation Metric for Image Captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP 2021), Online. Punta Cana, Dominican Republic, 7–11 November 2021; pp. 7514–7528. [Google Scholar]
- Papineni, K.; Roukos, S.; Ward, T.; Zhu, W. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, Philadelphia, PA, USA, 6–12 July 2002; pp. 311–318. [Google Scholar]
- Vedantam, R.; Lawrence Zitnick, C.; Parikh, D. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 4566–4575. [Google Scholar]
- Tewel, Y.; Shalev, Y.; Schwartz, I.; Wolf, L. Zerocap: Zero-shot image-to-text generation for visual-semantic arithmetic. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 17918–17928. [Google Scholar]








| Study (Ref.) | Objective | Methodology | Strengths | Limitations |
|---|---|---|---|---|
| 2.1 Controllable Image Captioning (CIC) | ||||
| SentiCap [25], StyleNet [26] | Sentiment as a single constraint. | RNN-based captioner; style embedding module. | Pioneering sentimental CIC; explicit sentiment token control. | RNN-based small-scale models; lack multiple constraints; limited scalability and generalizability. |
| MSCap [27] | Multiple sentimental styles. | Multi-style captioner learning distinct style representations. | Supports affective styles with more diverse outputs. | |
| ConCap [3], ConZIC [4] | Multiple constraints. | Zero-/few-shot generation. | Parameter-free; flexible constraints in zero-/few-shot settings. | Visual grounding is implicit; constraint conflicts are unresolved; scalability to MLLMs is unexplored. |
| PositionID [11] | Length as primary constraint. | Length-aware prompts guiding token allocation. | Strong and interpretable length control | Optimizes length alone; sentiment and grounding are not modeled; fails when applying to sentimental CIC. |
| 2.2 Steering for Controllable Image Captioning | ||||
| ActAdd [40], CAA [42] | Sentiment attribute control in text-only LLMs. | Extracts attribute directions and injects into residual streams. | Parameter-free; interpretable activation directions. | No visual grounding; steering strength fixed; ignore multi-constraint or entropy-aware control. |
| Activation–intervention captioning [45] | Reduce hallucination as the main constraint. | Activation interventions during decoding. | Activation-level steering can lower hallucination. | No sentiment or length control; steering intensity is not coordinated. |
| 2.3 Multi-Agent Systems for Controllable Image Captioning | ||||
| MosAIC [16] | Culturally style captions. | Multi-agent framework. | Improves cultural relevance and diversity. | Prompt-level agent interaction; no multi-constraint trade-offs. |
| CapMAS [17], Planner–critic [18] | Improve factual completeness and reduce hallucination. | Planner–Generator–Critic with iterative refinement and retrieval-augmented verification. | Agent collaboration improves entity coverage and grounding. | Lacks sentimental or length control; practical complexity grows. |
| Parameter | Default Value |
|---|---|
| Min steering strength | |
| Max steering strength | 4 |
| Threshold for sentiment evaluation | |
| Threshold for grounding evaluation | |
| Caption length constraint L | 15 |
| Number of picked atoms for activation selection | 20 |
| Max iterations of refinement | 5 |
| Sampling temperature for decoding |
| Model | Sentiment | Length Control | Sentiment Expression | Image–Text Alignment | |||
|---|---|---|---|---|---|---|---|
| MAE↓ | EM↑ | LC↑ | Cls.↑ | ClipS↑ | RefClipS↑ | ||
| PositionID | positive | 10.203 | 0.048 | 0.204 | 0.562 | 0.681 | 0.670 |
| ConZIC | positive | 1.632 | 0.129 | 0.452 | 0.477 | 0.799 | 0.745 |
| Qwen2.5-VL | positive | 3.500 | 0.062 | 0.208 | 0.933 | 0.763 | 0.754 |
| + IE-MAS | positive | 2.871 | 0.091 | 0.282 | 0.980 | 0.773 | 0.774 |
| LLaMA3.2-V | positive | 1.942 | 0.176 | 0.441 | 0.827 | 0.802 | 0.802 |
| + IE-MAS | positive | 0.785 | 0.337 | 0.901 | 0.958 | 0.812 | 0.807 |
| PositionID | negative | 12.274 | 0.042 | 0.143 | 0.498 | 0.648 | 0.678 |
| ConZIC | negative | 1.636 | 0.131 | 0.443 | 0.722 | 0.801 | 0.745 |
| Qwen2.5-VL | negative | 3.161 | 0.107 | 0.292 | 0.425 | 0.798 | 0.779 |
| + IE-MAS | negative | 3.648 | 0.056 | 0.215 | 0.801 | 0.805 | 0.767 |
| LLaMA3.2-V | negative | 1.932 | 0.141 | 0.465 | 0.917 | 0.792 | 0.794 |
| + IE-MAS | negative | 0.773 | 0.363 | 0.901 | 0.928 | 0.811 | 0.804 |
| Model | Sentiment | Length Control | Sentiment Expression | Image–Text Alignment | |||
|---|---|---|---|---|---|---|---|
| MAE↓ | EM↑ | LC↑ | Cls.↑ | ClipS↑ | RefClipS↑ | ||
| IE-MAS | positive | 0.785 | 0.337 | 0.901 | 0.958 | 0.812 | 0.807 |
| w/o EMCS | positive | 2.039 | 0.139 | 0.455 | 0.829 | 0.798 | 0.799 |
| w/o IMS | positive | 1.942 | 0.176 | 0.441 | 0.827 | 0.802 | 0.802 |
| IE-MAS | negative | 0.773 | 0.363 | 0.901 | 0.928 | 0.811 | 0.804 |
| w/o EMCS | negative | 1.810 | 0.182 | 0.483 | 0.914 | 0.786 | 0.795 |
| w/o IMS | negative | 1.932 | 0.141 | 0.465 | 0.917 | 0.792 | 0.794 |
| Constraints | Length Control | Sentiment Expression | Image–Text Alignment | ||||
|---|---|---|---|---|---|---|---|
| Length | Sentiment | MAE↓ | EM↑ | LC↑ | Cls.↑ | ClipS↑ | RefClipS↑ |
| 15 | Positive | 0.785 | 0.337 | 0.901 | 0.958 | 0.812 | 0.807 |
| 20 | Positive | 2.358 | 0.099 | 0.443 | 0.887 | 0.821 | 0.798 |
| 25 | Positive | 5.520 | 0.012 | 0.090 | 0.882 | 0.831 | 0.792 |
| 15 | Negative | 0.773 | 0.363 | 0.901 | 0.928 | 0.811 | 0.804 |
| 20 | Negative | 2.060 | 0.132 | 0.476 | 0.852 | 0.820 | 0.794 |
| 25 | Negative | 4.952 | 0.028 | 0.098 | 0.791 | 0.831 | 0.792 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Cai, T.; Chen, C.; Lin, S.; Ju, S.; Liao, X. IE-MAS: Internal–External Multi-Agent Steering for Controllable Image Captioning. Entropy 2025, 27, 1237. https://doi.org/10.3390/e27121237
Cai T, Chen C, Lin S, Ju S, Liao X. IE-MAS: Internal–External Multi-Agent Steering for Controllable Image Captioning. Entropy. 2025; 27(12):1237. https://doi.org/10.3390/e27121237
Chicago/Turabian StyleCai, Tiecheng, Chao Chen, Shanshan Lin, Sibo Ju, and Xiangwen Liao. 2025. "IE-MAS: Internal–External Multi-Agent Steering for Controllable Image Captioning" Entropy 27, no. 12: 1237. https://doi.org/10.3390/e27121237
APA StyleCai, T., Chen, C., Lin, S., Ju, S., & Liao, X. (2025). IE-MAS: Internal–External Multi-Agent Steering for Controllable Image Captioning. Entropy, 27(12), 1237. https://doi.org/10.3390/e27121237

