Ensuring Consistency for In-Image Translation
Abstract
1. Introduction
- Translation Consistency: Image information should be integrated into the translation process for improving translation accuracy. While previous studies on multimodal machine translation [4,5,6,7,8] have primarily focused on enhancing text translation quality by leveraging image information during the text translation process, text translation accompanied by images is relatively uncommon in real-life scenarios. In IIT, text is always accompanied by an image, making image information particularly crucial during the translation process.
- Image Generation Consistency: Image generation consistency consists of two dimensions. The first is background coherence, which refers to the smoothness and seamless integration between the generated text and its background. AnyTrans has already achieved strong performance in this aspect. However, in certain scenarios—such as film poster translation and children’s picture book translation—it is essential for the visual style of the target text to remain consistent with that of the source text. Therefore, we propose that the generation of target images should also maintain font style consistency. This form of consistency encompasses elements such as font type and color in text-containing images. Existing methods like AnyTrans infer the style of the generated text image from the surrounding context of the rendered scene. A more suitable approach would be to ensure consistency with the style of the original source text image.
- We underline current inadequacies in maintaining consistency in in-image translation tasks, with a specific emphasis on ensuring consistency in translation and image generation.
- We introduce a simple two-stage framework that effectively addresses prevalent issues in existing tasks of text translation on images, such as inaccuracies in translation, inconsistency in style, and blurriness in backgrounds.
- We modify the existing structural network of diffusion model to enhance its ability to generate image-with-text content in a consistent style. In addition, we curate a few style-consistent pseudo-parallel image pairs for training. Our framework yields satisfactory and reliable results in ensuring consistent style preservation.
2. Related Work
2.1. Text Image Translation
2.2. Text Image Generation
3. Methodology
3.1. High-Consistency In-Image Translation Framework
3.2. Stage 1: TIT Based on MMLLM
3.3. Stage 2: Back-Filling Based on Diffusion Model
3.3.1. Style Latent Module
3.3.2. Glyph Latent Module
3.3.3. Text-Controlled Diffusion Pipeline
4. Experiments
4.1. Datasets
4.2. Baselines
4.3. Settings
4.4. Results and Analysis
4.4.1. Translation Results
4.4.2. Results on Image Generation
4.4.3. Case Study
4.4.4. Human Evaluation and Large Language Model Evaluation
- (1)
- Translation Consistency:
- Point 1: Failure to recognize the text, or the text is recognized but the translation contains significant inaccuracies and is completely irrelevant.
- Point 2: The text can be recognized and the translation is generally relevant; however, partial errors exist, or insufficient integration of image information leads to ambiguous translations.
- Point 3: The text is correctly recognized and the translation is completely accurate, effectively integrating image information. All text is displayed correctly without any issues.
- (2)
- Text Style Consistency:
- Point 1: Low text style consistency. The text style is highly inconsistent, with noticeable variations in color, font, thickness, and other visual attributes.
- Point 2: Medium text style consistency. Partial consistency is observed, where some elements such as color, font, or thickness remain consistent.
- Point 3: High text style consistency. The text style is nearly entirely consistent; color, font, and thickness match well, and relevant text details are highly uniform.
- (3)
- Background Coherence:
- Point 1: Low coherence. Clear boundaries exist between the text and the background, resulting in a strong sense of layering.
- Point 2: Medium coherence. The text blends reasonably well into the background without distinct boundaries.
- Point 3: High coherence. The text and background are seamlessly integrated; the background appears clear and natural with no visible boundaries.
5. Conclusions and Future Work
- While our method makes clear progress in preserving text style, it still fails in some more complex style transfer scenarios, such as cases where a single sentence contains multiple distinct text styles. This limitation is likely due to the constraints of our pseudo-data construction process, which cannot fully cover the wide range of style transfer requirements encountered in real-world images. We leave addressing these more complex style variations as an important direction for future work.
- Improving the detailed content of IIT: IIT is a complex task encompassing multiple modalities and diverse scenarios. While our method has proposed a general framework, there are still intricacies that require resolution within this framework. These include translating vertical text images, translating multi-line text images, collaborative translation of multiple texts, translation of lengthy text images, and optimizing text positioning during translation. Our forthcoming research endeavors will concentrate on refining these aspects.
- End-to-end IIT frameworks: Presently, models that exhibit superior performance are often designed in a cascaded manner. Cascaded approaches may encounter issues such as error accumulation and high latency. Therefore, our upcoming research will investigate end-to-end methods for IIT to overcome these challenges and improve efficiency.
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Ma, C.; Zhang, Y.; Tu, M.; Han, X.; Wu, L.; Zhao, Y.; Zhou, Y. Improving end-to-end text image translation from the auxiliary text translation task. In Proceedings of the 2022 26th International Conference on Pattern Recognition (ICPR), Montreal, QC, Canada, 21–25 August 2022; IEEE: New York, NY, USA, 2022; pp. 1664–1670. [Google Scholar]
- Lan, Z.; Yu, J.; Li, X.; Zhang, W.; Luan, J.; Wang, B.; Huang, D.; Su, J. Exploring Better Text Image Translation with Multimodal Codebook. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, ON, Canada, 9–14 July 2023; pp. 3479–3491. [Google Scholar]
- Qian, Z.; Zhang, P.; Yang, B.; Fan, K.; Ma, Y.; Wong, D.F.; Sun, X.; Ji, R. AnyTrans: Translate AnyText in the Image with Large Scale Models. arXiv 2024, arXiv:2406.11432. [Google Scholar]
- Caglayan, O.; Barrault, L.; Bougares, F. Multimodal attention for neural machine translation. arXiv 2016, arXiv:1609.03976. [Google Scholar] [CrossRef] [Scilit]
- Yao, S.; Wan, X. Multimodal transformer for multimodal machine translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Virtual, 5–10 July 2020; pp. 4346–4350. [Google Scholar]
- Yin, Y.; Meng, F.; Su, J.; Zhou, C.; Yang, Z.; Zhou, J.; Luo, J. A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine Translation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Virtual, 5–10 July 2020; pp. 3025–3035. [Google Scholar]
- Caglayan, O.; Kuyu, M.; Amac, M.S.; Madhyastha, P.S.; Erdem, E.; Erdem, A.; Specia, L. Cross-lingual Visual Pre-training for Multimodal Machine Translation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, Virtual, 19–23 April 2021; pp. 1317–1324. [Google Scholar]
- Li, B.; Lv, C.; Zhou, Z.; Zhou, T.; Xiao, T.; Ma, A.; Zhu, J. On Vision Features in Multimodal Machine Translation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Dublin, Ireland, 22–27 May 2022; pp. 6327–6337. [Google Scholar]
- Watanabe, Y.; Okada, Y.; Kim, Y.B.; Takeda, T. Translation camera. In Proceedings of the Fourteenth International Conference on Pattern Recognition (Cat. No. 98EX170), Brisbane, Australia, 16–20 August 1998; IEEE: New York, NY, USA, 1998; Volume 1, pp. 613–617. [Google Scholar]
- Chen, J.; Cao, H.; Natarajan, P. Integrating natural language processing with image document analysis: What we learned from two real-world applications. Int. J. Doc. Anal. Recognit. (IJDAR) 2015, 18, 235–247. [Google Scholar] [CrossRef] [Scilit]
- Lan, Z.; Niu, L.; Meng, F.; Zhou, J.; Zhang, M.; Su, J. Translatotron-V (ison): An End-to-End Model for In-Image Machine Translation. arXiv 2024, arXiv:2407.02894. [Google Scholar]
- Vinyals, O.; Toshev, A.; Bengio, S.; Erhan, D. Show and tell: A neural image caption generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 3156–3164. [Google Scholar]
- Lu, C.; Krishna, R.; Bernstein, M.; Li, F.-F. Visual relationship detection with language priors. In Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, 11–14 October 2016; Springer: Berlin/Heidelberg, Germany, 2016; pp. 852–869. [Google Scholar]
- Deng, J.; Dong, W.; Socher, R.; Li, L.J.; Li, K.; Li, F.-F. Imagenet: A large-scale hierarchical image database. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, 20–25 June 2009; IEEE: New York, NY, USA, 2009; pp. 248–255. [Google Scholar]
- Tian, C.; Cheng, T.; Peng, Z.; Zuo, W.; Tian, Y.; Zhang, Q.; Wang, F.Y.; Zhang, D. A survey on deep learning fundamentals. Artif. Intell. Rev. 2025, 58, 381. [Google Scholar] [CrossRef] [Scilit]
- Tian, C.; Song, M.; Zuo, W.; Du, B.; Zhang, Y.; Zhang, S. Application of convolutional neural networks in image super-resolution. arXiv 2025, arXiv:2506.02604. [Google Scholar] [CrossRef] [Scilit]
- Ma, J.; Zhao, M.; Chen, C.; Wang, R.; Niu, D.; Lu, H.; Lin, X. Glyphdraw: Learning to draw chinese characters in image synthesis models coherently. arXiv 2023, arXiv:2303.17870. [Google Scholar]
- Yang, Y.; Gui, D.; Yuan, Y.; Liang, W.; Ding, H.; Hu, H.; Chen, K. Glyphcontrol: Glyph conditional control for visual text generation. Adv. Neural Inf. Process. Syst. 2024, 36, 44050–44066. [Google Scholar]
- Chen, J.; Huang, Y.; Lv, T.; Cui, L.; Chen, Q.; Wei, F. Textdiffuser: Diffusion models as text painters. Adv. Neural Inf. Process. Syst. 2024, 36, 9353–9387. [Google Scholar]
- Chen, J.; Huang, Y.; Lv, T.; Cui, L.; Chen, Q.; Wei, F. Textdiffuser-2: Unleashing the power of language models for text rendering. arXiv 2023, arXiv:2311.16465. [Google Scholar] [CrossRef] [Scilit]
- Tuo, Y.; Xiang, W.; He, J.Y.; Geng, Y.; Xie, X. AnyText: Multilingual Visual Text Generation and Editing. In Proceedings of the Twelfth International Conference on Learning Representations, Pensacola, FL, USA, 5–7 December 2023. [Google Scholar]
- Huang, Y.; Feng, X.; Li, B.; Fu, C.; Huo, W.; Liu, T.; Qin, B. Aligning translation-specific understanding to general understanding in large language models. arXiv 2024, arXiv:2401.05072. [Google Scholar] [CrossRef] [Scilit]
- Futeral, M.; Schmid, C.; Laptev, I.; Sagot, B.; Bawden, R. Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive Evaluation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, ON, Canada, 9–14 July 2023; pp. 5394–5413. [Google Scholar]
- Nayef, N.; Patel, Y.; Busta, M.; Chowdhury, P.N.; Karatzas, D.; Khlif, W.; Matas, J.; Pal, U.; Burie, J.C.; Liu, C.L. ICDAR2019 Robust Reading Challenge on Multi-lingual Scene Text Detection and Recognition—RRC-MLT-2019. In Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR), Sydney, Australia, 20–25 September 2019. [Google Scholar]
- Fan, A.; Bhosale, S.; Schwenk, H.; Ma, Z.; El-Kishky, A.; Goyal, S.; Baines, M.; Celebi, O.; Wenzek, G.; Chaudhary, V.; et al. Beyond English-Centric Multilingual Machine Translation. arXiv 2020, arXiv:2010.11125. [Google Scholar] [CrossRef] [Scilit]
- Costa-Jussà, M.R.; Cross, J.; Çelebi, O.; Elbayad, M.; Heafield, K.; Heffernan, K.; Kalbassi, E.; Lam, J.; Licht, D.; Maillard, J.; et al. No language left behind: Scaling human-centered machine translation. arXiv 2022, arXiv:2207.04672. [Google Scholar] [CrossRef] [Scilit]
- Bai, S.; Chen, K.; Liu, X.; Wang, J.; Ge, W.; Song, S.; Dang, K.; Wang, P.; Wang, S.; Tang, J.; et al. Qwen2.5-VL Technical Report. arXiv 2025, arXiv:2502.13923. [Google Scholar] [CrossRef] [Scilit]
- Papineni, K.; Roukos, S.; Ward, T.; Zhu, W.J. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Philadelphia, PA, USA, 6–12 July 2002; pp. 311–318. [Google Scholar]
- Rei, R.; Stewart, C.; Farinha, A.C.; Lavie, A. COMET: A Neural Framework for MT Evaluation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Virtual, 16–20 November 2020; pp. 2685–2702. [Google Scholar]
- Wang, T.; Yang, X.; Xu, K.; Chen, S.; Zhang, Q.; Lau, R.W. Spatial attentive single-image deraining with a high quality real rain dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 12270–12279. [Google Scholar]
- Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; et al. Qwen Technical Report. arXiv 2023, arXiv:2309.16609. [Google Scholar] [CrossRef] [Scilit]






| Language | ZH-EN | ZH-JA | ZH-KO | Average | ||||
|---|---|---|---|---|---|---|---|---|
| Metric | COMET | BLEU | COMET | BLEU | COMET | BLEU | COMET | BLEU |
| NLLB-200 (3.3 b) | 66.3 | 29.7 | 77.2 | 24.9 | 72.5 | 20.1 | 72.0 | 24.9 |
| M2M100 (1.2 b) | 66.1 | 33.1 | 79.6 | 29.4 | 71.1 | 18.1 | 72.3 | 26.9 |
| Qwen1.5-7b | 73.4 | 37.4 | 80.9 | 31.2 | 70.4 | 11.4 | 74.9 | 26.7 |
| Qwenvl2.5-7b | 78.2 | 42.7 | 81.6 | 34.6 | 73.0 | 19.1 | 77.6 | 31.1 |
| Qwenvl2.5-7b + COT | 77.9 | 42.3 | 83.3 | 35.1 | 73.3 | 19.9 | 78.1 | 32.4 |
| Language | EN-ZH | JA-ZH | KO-ZH | Average | ||||
| Metric | COMET | BLEU | COMET | BLEU | COMET | BLEU | COMET | BLEU |
| NLLB-200 (3.3 b) | 73.3 | 21.5 | 61.3 | 7.4 | 65.3 | 9.1 | 66.6 | 12.7 |
| M2M100 (1.2 b) | 76.9 | 24.2 | 74.5 | 24.3 | 67.8 | 14.8 | 73.1 | 21.1 |
| Qwen1.5-7b | 80.7 | 27.6 | 78.7 | 30.0 | 75.7 | 20.9 | 78.4 | 26.2 |
| Qwenvl2.5-7b | 79.8 | 26.5 | 83.9 | 39.6 | 80.5 | 32.6 | 81.4 | 32.9 |
| Qwenvl2.5-7b + COT | 80.1 | 27.5 | 84.1 | 39.8 | 79.9 | 32.2 | 81.4 | 33.2 |
| Language | EN-DE | EN-FR | EN-ZH | Average | ||||
|---|---|---|---|---|---|---|---|---|
| Metric | COMET | BLEU | COMET | BLEU | COMET | BLEU | COMET | BLEU |
| Qwen2.5-VL-7b | 81.01 | 31.05 | 81.15 | 36.46 | 86.44 | 47.31 | 82.9 | 38.3 |
| Qwen2.5-VL-7b + COT | 81.95 | 31.61 | 81.95 | 37.76 | 86.98 | 48.89 | 83.6 | 39.4 |
| Language | Aliyun Trans | Youdao Trans | AnyTrans | HCIIT | |
|---|---|---|---|---|---|
| SSIM | En-Zh | 0.605 | 0.516 | 0.381 | 0.744 |
| Zh-En | 0.657 | 0.622 | 0.439 | 0.685 | |
| Fr-En | - | 0.467 | 0.281 | 0.675 | |
| L1 | En-Zh | 0.436 | 0.561 | 0.984 | 0.402 |
| Zh-En | 0.363 | 0.418 | 0.806 | 0.399 | |
| Fr-En | - | 0.625 | 1.163 | 0.467 |
| Language | En-Zh | Zh-En | ||
|---|---|---|---|---|
| Metric | SSIM | L1 | SSIM | L1 |
| HCIIT | 0.744 | 0.402 | 0.685 | 0.399 |
| w.o. Background Image | 0.712 | 0.411 | 0.642 | 0.412 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Fu, C.; Feng, X.; Huang, Y.; Huo, W.; Li, B.; Xiang, Y.; Wang, H.; Liu, T. Ensuring Consistency for In-Image Translation. Mathematics 2026, 14, 490. https://doi.org/10.3390/math14030490
Fu C, Feng X, Huang Y, Huo W, Li B, Xiang Y, Wang H, Liu T. Ensuring Consistency for In-Image Translation. Mathematics. 2026; 14(3):490. https://doi.org/10.3390/math14030490
Chicago/Turabian StyleFu, Chengpeng, Xiaocheng Feng, Yichong Huang, Wenshuai Huo, Baohang Li, Yang Xiang, Hui Wang, and Ting Liu. 2026. "Ensuring Consistency for In-Image Translation" Mathematics 14, no. 3: 490. https://doi.org/10.3390/math14030490
APA StyleFu, C., Feng, X., Huang, Y., Huo, W., Li, B., Xiang, Y., Wang, H., & Liu, T. (2026). Ensuring Consistency for In-Image Translation. Mathematics, 14(3), 490. https://doi.org/10.3390/math14030490

