Collaborative Control in Diffusion Models for Precise Image Generation: A Survey
Abstract
1. Introduction
1.1. Core Theory and Methodological Evolution of Diffusion Models
1.2. Text-to-Image Conditional Image Generation
1.3. Image-to-Image Conditional Image Generation
2. Condition Extensions and Multi-Condition Fusion for Controlled Generation
2.1. Novel Single-Condition Control Extensions
2.2. Fusion Mechanisms for Multi-Condition Control
2.3. Challenges and Optimization of Multi-Condition Control
3. Collaborative Control Paradigms in Diffusion Models
3.1. Definition and Taxonomy of Collaborative Control Paradigms
3.2. Representative Methods and Technical Pathways
3.3. Zero-Shot/Few-Shot Adaptation for New Condition Extensions
3.4. Computational Complexity and Scalability Considerations
4. Application Scenarios of Collaborative Control Methods
4.1. Daily Creation and Commercial Content Production
4.2. Precision Modeling and Optimization in Professional Domains
4.3. Data Support and Evaluation for Model Innovation
4.4. Application Potential and Future Directions
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AP | Average Precision |
| CFG | Classifier-Free Guidance |
| CLIP | Contrastive Language–Image Pre-Training |
| DDIM | Denoising Diffusion Implicit Model |
| DDPM | Denoising Diffusion Probabilistic Model |
| DINO | Self-Distillation with No Labels |
| DiT | Diffusion Transformer |
| DSG | Davidsonian Scene Graph |
| FID | Fréchet Inception Distance |
| GPU | Graphics Processing Unit |
| HED | Holistically Nested Edge Detection |
| I2I | Image-to-Image |
| LDM | Latent Diffusion Model |
| LLM | Large Language Model |
| LoRA | Low-Rank Adaptation |
| mAP | Mean Average Precision |
| mIoU | Mean Intersection over Union |
| MMDiT | Multimodal Diffusion Transformer |
| MSE | Mean Squared Error |
| NR | Not Reported |
| ODE | Ordinary Differential Equation |
| RMSE | Root Mean Squared Error |
| RPG | Recaption, Plan, and Generate |
| SAR | Synthetic Aperture Radar |
| SDE | Stochastic Differential Equation |
| SDXL | Stable Diffusion XL |
| SLD | Self-Correcting LLM-Controlled Diffusion Models |
| SSIM | Structural Similarity Index Measure |
| T2I | Text-to-Image |
| TIFA | Text-to-Image Faithfulness |
| VAE | Variational Autoencoder |
| VQA | Visual Question Answering |
References
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, 33, 6840–6851. [Google Scholar]
- Dhariwal, P.; Nichol, A. Diffusion models beat gans on image synthesis. Adv. Neural Inf. Process. Syst. 2021, 34, 8780–8794. [Google Scholar]
- Song, Y.; Sohl-Dickstein, J.; Kingma, D.P.; Kumar, A.; Ermon, S.; Poole, B. Score-based generative modeling through stochastic differential equations. arXiv 2020, arXiv:2011.13456. [Google Scholar] [CrossRef]
- Zhang, N.; Tang, H. Text-to-image synthesis: A decade survey. arXiv 2024, arXiv:2411.16164. [Google Scholar] [CrossRef]
- Cao, P.; Zhou, F.; Song, Q.; Yang, L. Controllable Generation with Text-to-Image Diffusion Models: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2026, 48, 4771–4791. [Google Scholar] [CrossRef] [PubMed]
- Zhang, L.; Rao, A.; Agrawala, M. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 3836–3847. [Google Scholar] [CrossRef]
- Mou, C.; Wang, X.; Xie, L.; Wu, Y.; Zhang, J.; Qi, Z.; Shan, Y. T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024; Volume 38, pp. 4296–4304. [Google Scholar] [CrossRef]
- Kingma, D.P.; Welling, M. Auto-encoding variational bayes. arXiv 2013, arXiv:1312.6114. [Google Scholar] [CrossRef]
- LeCun, Y.; Chopra, S.; Hadsell, R.; Ranzato, M.; Huang, F.J. Energy-Based Models. In Predicting Structured Data; Bakir, G., Hofmann, T., Scholkopf, B., Smola, A.J., Taskar, B., Vishwanathan, S.V.N., Eds.; MIT Press: Cambridge, MA, USA, 2007; pp. 191–246. [Google Scholar] [CrossRef]
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial networks. Commun. ACM 2020, 63, 139–144. [Google Scholar] [CrossRef]
- Dinh, L.; Sohl-Dickstein, J.; Bengio, S. Density estimation using real nvp. arXiv 2016, arXiv:1605.08803. [Google Scholar] [CrossRef]
- Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; Ganguli, S. Deep unsupervised learning using nonequilibrium thermodynamics. Proc. Mach. Learn. Res. 2015, 37, 2256–2265. [Google Scholar]
- Lu, C.; Zhou, Y.; Bao, F.; Chen, J.; Li, C.; Zhu, J. Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps. Adv. Neural Inf. Process. Syst. 2022, 35, 5775–5787. [Google Scholar] [CrossRef]
- Karras, T.; Aittala, M.; Aila, T.; Laine, S. Elucidating the design space of diffusion-based generative models. Adv. Neural Inf. Process. Syst. 2022, 35, 26565–26577. [Google Scholar] [CrossRef]
- Lipman, Y.; Chen, R.T.; Ben-Hamu, H.; Nickel, M.; Le, M. Flow matching for generative modeling. arXiv 2022, arXiv:2210.02747. [Google Scholar] [CrossRef]
- Liu, X.; Gong, C.; Liu, Q. Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow. arXiv 2022, arXiv:2209.03003. [Google Scholar] [CrossRef]
- Song, Y.; Ermon, S. Generative modeling by estimating gradients of the data distribution. Adv. Neural Inf. Process. Syst. 2019, 32, 11918–11930. [Google Scholar]
- Jolicoeur-Martineau, A.; Piché-Taillefer, R.; Combes, R.T.d.; Mitliagkas, I. Adversarial score matching and improved sampling for image generation. arXiv 2020, arXiv:2009.05475. [Google Scholar] [CrossRef]
- De Bortoli, V.; Thornton, J.; Heng, J.; Doucet, A. Diffusion schrödinger bridge with applications to score-based generative modeling. Adv. Neural Inf. Process. Syst. 2021, 34, 17695–17709. [Google Scholar] [CrossRef]
- Song, J.; Meng, C.; Ermon, S. Denoising diffusion implicit models. arXiv 2020, arXiv:2010.02502. [Google Scholar] [CrossRef]
- Nichol, A.Q.; Dhariwal, P. Improved denoising diffusion probabilistic models. Proc. Mach. Learn. Res. 2021, 139, 8162–8171. [Google Scholar]
- Salimans, T.; Ho, J. Progressive distillation for fast sampling of diffusion models. arXiv 2022, arXiv:2202.00512. [Google Scholar] [CrossRef]
- Zhao, W.; Bai, L.; Rao, Y.; Zhou, J.; Lu, J. Unipc: A unified predictor-corrector framework for fast sampling of diffusion models. Adv. Neural Inf. Process. Syst. 2023, 36, 49842–49869. [Google Scholar] [CrossRef]
- Xue, S.; Liu, Z.; Chen, F.; Zhang, S.; Hu, T.; Xie, E.; Li, Z. Accelerating diffusion sampling with optimized time steps. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 8292–8301. [Google Scholar] [CrossRef]
- Liu, X.; Park, D.H.; Azadi, S.; Zhang, G.; Chopikyan, A.; Hu, Y.; Shi, H.; Rohrbach, A.; Darrell, T. More control for free! image synthesis with semantic diffusion guidance. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 3–7 January 2023; pp. 289–299. [Google Scholar] [CrossRef]
- Ho, J.; Salimans, T. Classifier-free diffusion guidance. arXiv 2022, arXiv:2207.12598. [Google Scholar] [CrossRef]
- Shen, D.; Song, G.; Xue, Z.; Wang, F.Y.; Liu, Y. Rethinking the spatial inconsistency in classifier-free diffusion guidance. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 9370–9379. [Google Scholar] [CrossRef]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 21–24 June 2022; pp. 10684–10695. [Google Scholar] [CrossRef]
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany, 5–9 October 2015; Springer: New York, NY, USA, 2015; pp. 234–241. [Google Scholar] [CrossRef]
- Peebles, W.; Xie, S. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 4195–4205. [Google Scholar] [CrossRef]
- Bao, F.; Nie, S.; Xue, K.; Cao, Y.; Li, C.; Su, H.; Zhu, J. All are worth words: A vit backbone for diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BS, Canada, 18–22 June 2023; pp. 22669–22679. [Google Scholar] [CrossRef]
- Zheng, H.; Nie, W.; Vahdat, A.; Anandkumar, A. Fast training of diffusion models with masked transformers. arXiv 2023, arXiv:2306.09305. [Google Scholar] [CrossRef]
- Gao, S.; Zhou, P.; Cheng, M.M.; Yan, S. Masked Diffusion Transformer Is a Strong Image Synthesizer. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; IEEE: New York, NY, USA, 2023; pp. 23107–23116. [Google Scholar] [CrossRef]
- Lu, Z.; Wang, Z.; Huang, D.; Wu, C.; Liu, X.; Ouyang, W.; Bai, L. Fit: Flexible vision transformer for diffusion model. arXiv 2024, arXiv:2402.12376. [Google Scholar] [CrossRef]
- Shen, T.; Yu, J.; Zhou, D.; Li, D.; Barsoum, E. E-MMDiT: Revisiting Multimodal Diffusion Transformer Design for Fast Image Synthesis under Limited Resources. arXiv 2025, arXiv:2510.27135. [Google Scholar] [CrossRef]
- Yang, L.; Huang, Z.; Zhang, Z.; Liu, Z.; Hong, S.; Zhang, W.; Yang, W.; Cui, B.; Zhang, L. Graphusion: Latent diffusion for graph generation. IEEE Trans. Knowl. Data Eng. 2024, 36, 6358–6369. [Google Scholar] [CrossRef]
- Jia, W.; Huang, M.; Chen, N.; Zhang, L.; Mao, Z. D2iT: Dynamic Diffusion Transformer for Accurate Image Generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 13–15 June 2025; pp. 12860–12870. [Google Scholar] [CrossRef]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. Proc. Mach. Learn. Res. 2021, 139, 8748–8763. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; Sutskever, I. Zero-shot text-to-image generation. Proc. Mach. Learn. Res. 2021, 139, 8821–8831. [Google Scholar]
- Balaji, Y.; Nah, S.; Huang, X.; Vahdat, A.; Song, J.; Zhang, Q.; Kreis, K.; Aittala, M.; Aila, T.; Laine, S.; et al. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv 2022, arXiv:2211.01324. [Google Scholar] [CrossRef]
- Nichol, A.; Dhariwal, P.; Ramesh, A.; Shyam, P.; Mishkin, P.; McGrew, B.; Sutskever, I.; Chen, M. Glide: Towards photorealistic image generation and editing with text-guided diffusion models. arXiv 2021, arXiv:2112.10741. [Google Scholar] [CrossRef]
- Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E.L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. Photorealistic text-to-image diffusion models with deep language understanding. Adv. Neural Inf. Process. Syst. 2022, 35, 36479–36494. [Google Scholar] [CrossRef]
- Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P.J. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 2020, 21, 1–67. [Google Scholar] [CrossRef]
- Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; Rombach, R. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv 2023, arXiv:2307.01952. [Google Scholar] [CrossRef]
- Betker, J.; Goh, G.; Jing, L.; Brooks, T.; Wang, J.; Li, L.; Ouyang, L.; Zhuang, J.; Lee, J.; Guo, Y.; et al. Improving Image Generation with Better Captions. OpenAI Technical Report. 2023. Available online: https://cdn.openai.com/papers/dall-e-3.pdf (accessed on 21 July 2026).
- Chen, J.; Yu, J.; Ge, C.; Yao, L.; Xie, E.; Wu, Y.; Wang, Z.; Kwok, J.; Luo, P.; Lu, H.; et al. Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis. arXiv 2023, arXiv:2310.00426. [Google Scholar] [CrossRef]
- Chen, J.; Ge, C.; Xie, E.; Wu, Y.; Yao, L.; Ren, X.; Wang, Z.; Luo, P.; Lu, H.; Li, Z. Pixart-σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: New York, NY, USA, 2024; pp. 74–91. [Google Scholar] [CrossRef]
- Li, Z.; Zhang, J.; Lin, Q.; Xiong, J.; Long, Y.; Deng, X.; Zhang, Y.; Liu, X.; Huang, M.; Xiao, Z.; et al. Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understanding. arXiv 2024, arXiv:2405.08748. [Google Scholar] [CrossRef]
- Black Forest Labs. FLUX.1: Official Inference Code and Model Releases. GitHub Repository. 2024. Available online: https://github.com/black-forest-labs/flux (accessed on 13 July 2026).
- Xiao, S.; Wang, Y.; Zhou, J.; Yuan, H.; Xing, X.; Yan, R.; Li, C.; Wang, S.; Huang, T.; Liu, Z. OmniGen: Unified Image Generation. In Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 10–17 June 2025; IEEE: New York, NY, USA, 2025; pp. 13294–13304. [Google Scholar] [CrossRef]
- Tan, Z.; Liu, S.; Yang, X.; Xue, Q.; Wang, X. OminiControl: Minimal and Universal Control for Diffusion Transformer. In Proceedings of the 2025 IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, HI, USA, 19–25 October 2025; IEEE: New York, NY, USA, 2025; pp. 14940–14950. [Google Scholar] [CrossRef]
- Choi, J.; Kim, S.; Jeong, Y.; Gwon, Y.; Yoon, S. ILVR: Conditioning Method for Denoising Diffusion Probabilistic Models. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Virtual, 11–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 14347–14356. [Google Scholar] [CrossRef]
- Isola, P.; Zhu, J.Y.; Zhou, T.; Efros, A.A. Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1125–1134. [Google Scholar] [CrossRef]
- Zhu, J.Y.; Park, T.; Isola, P.; Efros, A.A. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2223–2232. [Google Scholar] [CrossRef]
- Huang, X.; Liu, M.Y.; Belongie, S.; Kautz, J. Multimodal unsupervised image-to-image translation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 179–196. [Google Scholar] [CrossRef] [PubMed]
- Lee, H.Y.; Tseng, H.Y.; Huang, J.B.; Singh, M.; Yang, M.H. Diverse image-to-image translation via disentangled representations. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 36–52. [Google Scholar] [CrossRef]
- Kim, G.; Kwon, T.; Ye, J.C. Diffusionclip: Text-guided diffusion models for robust image manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 2426–2435. [Google Scholar] [CrossRef]
- Saharia, C.; Chan, W.; Chang, H.; Lee, C.; Ho, J.; Salimans, T.; Fleet, D.; Norouzi, M. Palette: Image-to-image diffusion models. In Proceedings of the ACM SIGGRAPH 2022 Conference Proceedings, Vancouver, BC, Canada, 7–11 August 2022; pp. 1–10. [Google Scholar] [CrossRef]
- Saharia, C.; Ho, J.; Chan, W.; Salimans, T.; Fleet, D.J.; Norouzi, M. Image super-resolution via iterative refinement. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 4713–4726. [Google Scholar] [CrossRef] [PubMed]
- Xia, B.; Zhang, Y.; Wang, S.; Wang, Y.; Wu, X.; Tian, Y.; Yang, W.; Van Gool, L. Diffir: Efficient diffusion model for image restoration. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 13095–13105. [Google Scholar] [CrossRef]
- Yue, Z.; Wang, J.; Loy, C.C. Resshift: Efficient diffusion model for image super-resolution by residual shifting. Adv. Neural Inf. Process. Syst. 2023, 36, 13294–13307. [Google Scholar] [CrossRef]
- Lugmayr, A.; Danelljan, M.; Romero, A.; Yu, F.; Timofte, R.; Van Gool, L. Repaint: Inpainting using denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 11461–11471. [Google Scholar] [CrossRef]
- Avrahami, O.; Lischinski, D.; Fried, O. Blended diffusion for text-driven editing of natural images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 18208–18218. [Google Scholar] [CrossRef]
- Perez, E.; Strub, F.; De Vries, H.; Dumoulin, V.; Courville, A. Film: Visual reasoning with a general conditioning layer. Proc. AAAI Conf. Artif. Intell. 2018, 32, 3942–3951. [Google Scholar] [CrossRef]
- Yang, L.; Huang, Z.; Song, Y.; Hong, S.; Li, G.; Zhang, W.; Cui, B.; Ghanem, B.; Yang, M.H. Diffusion-based scene graph to image generation with masked contrastive pre-training. arXiv 2022, arXiv:2211.11138. [Google Scholar] [CrossRef]
- Voynov, A.; Aberman, K.; Cohen-Or, D. Sketch-guided text-to-image diffusion models. In Proceedings of the ACM SIGGRAPH 2023 Conference Proceedings, Los Angeles, CA, USA, 23 July 2023; pp. 1–11. [Google Scholar] [CrossRef]
- Li, Y.; Liu, H.; Wu, Q.; Mu, F.; Yang, J.; Gao, J.; Li, C.; Lee, Y.J. Gligen: Open-set grounded text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 22511–22521. [Google Scholar] [CrossRef]
- Bhat, S.F.; Mitra, N.; Wonka, P. Loosecontrol: Lifting controlnet for generalized depth conditioning. In Proceedings of the ACM SIGGRAPH 2024 Conference Papers, Denver, CO, USA, 28 July–1 August 2024; pp. 1–11. [Google Scholar] [CrossRef]
- He, Q.; Peng, J.; Xu, P.; Jiang, B.; Hu, X.; Luo, D.; Liu, Y.; Wang, Y.; Wang, C.; Li, X.; et al. DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation. arXiv 2024, arXiv:2412.03255. [Google Scholar] [CrossRef]
- Zhao, S.; Chen, D.; Chen, Y.C.; Bao, J.; Hao, S.; Yuan, L.; Wong, K.Y.K. Uni-controlnet: All-in-one control to text-to-image diffusion models. Adv. Neural Inf. Process. Syst. 2023, 36, 11127–11150. [Google Scholar] [CrossRef]
- Gafni, O.; Polyak, A.; Ashual, O.; Sheynin, S.; Parikh, D.; Taigman, Y. Make-a-scene: Scene-based text-to-image generation with human priors. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; Springer: New York, NY, USA, 2022; pp. 89–106. [Google Scholar] [CrossRef]
- Avrahami, O.; Hayes, T.; Gafni, O.; Gupta, S.; Taigman, Y.; Parikh, D.; Lischinski, D.; Fried, O.; Yin, X. Spatext: Spatio-textual representation for controllable image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 18370–18380. [Google Scholar] [CrossRef]
- Liu, N.; Li, S.; Du, Y.; Torralba, A.; Tenenbaum, J.B. Compositional visual generation with composable diffusion models. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; Springer: New York, NY, USA, 2022; pp. 423–439. [Google Scholar] [CrossRef]
- Parmar, G.; Kumar Singh, K.; Zhang, R.; Li, Y.; Lu, J.; Zhu, J.Y. Zero-shot image-to-image translation. In Proceedings of the ACM SIGGRAPH 2023 Conference Proceedings, Los Angeles, CA, USA, 6–10 August 2023; pp. 1–11. [Google Scholar] [CrossRef]
- Bar-Tal, O.; Yariv, L.; Lipman, Y.; Dekel, T. MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation. arXiv 2023, arXiv:2302.08113. [Google Scholar] [CrossRef]
- Huang, L.; Chen, D.; Liu, Y.; Shen, Y.; Zhao, D.; Zhou, J. Composer: Creative and controllable image synthesis with composable conditions. arXiv 2023, arXiv:2302.09778. [Google Scholar] [CrossRef]
- Hu, M.; Zheng, J.; Liu, D.; Zheng, C.; Wang, C.; Tao, D.; Cham, T.J. Cocktail: Mixing multi-modality control for text-conditional image generation. In Proceedings of the Thirty-seventh Conference on Neural Information Processing Systems, New Orleans, LA, USA, 10–16 December 2023; pp. 32424–32444. [Google Scholar] [CrossRef]
- Wang, H.; Peng, J.; He, Q.; Yang, H.; Jin, Y.; Wu, J.; Hu, X.; Pan, Y.; Gan, Z.; Chi, M.; et al. UniCombine: Unified Multi-Conditional Combination with Diffusion Transformer. In Proceedings of the 2025 IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, HI, USA, 19–23 October 2025; IEEE: New York, NY, USA, 2025; pp. 18325–18334. [Google Scholar] [CrossRef]
- Yang, H.; Han, W.; Zhou, Y.; Shen, J. DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion Models. In Proceedings of the 2025 IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, HI, USA, 19–20 October 2025; IEEE: New York, NY, USA, 2025; pp. 1–10. [Google Scholar] [CrossRef]
- Xie, Y.; Feng, F.; Shi, R.; Wang, J.; Rui, Y.; Geng, X. DivControl: Knowledge Diversion for Controllable Image Generation. Proc. AAAI Conf. Artif. Intell. 2026, 40, 27108–27116. [Google Scholar] [CrossRef]
- Chen, H.; Zhang, Y.; Wu, S.; Wang, X.; Duan, X.; Zhou, Y.; Zhu, W. Disenbooth: Identity-preserving disentangled tuning for subject-driven text-to-image generation. arXiv 2023, arXiv:2305.03374. [Google Scholar] [CrossRef]
- Yang, Y.; Wang, W.; Peng, L.; Song, C.; Chen, Y.; Li, H.; Yang, X.; Lu, Q.; Cai, D.; He, X.; et al. Lora-composer: Leveraging low-rank adaptation for multi-concept customization in training-free diffusion models. IEEE Trans. Image Process. 2025, 34, 8145–8158. [Google Scholar] [CrossRef] [PubMed]
- Gandikota, R.; Materzyńska, J.; Zhou, T.; Torralba, A.; Bau, D. Concept sliders: Lora adaptors for precise control in diffusion models. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: New York, NY, USA, 2024; pp. 172–188. [Google Scholar] [CrossRef]
- Zheng, G.; Zhou, X.; Li, X.; Qi, Z.; Shan, Y.; Li, X. Layoutdiffusion: Controllable diffusion model for layout-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 22490–22499. [Google Scholar] [CrossRef]
- Chen, H.; Gao, Y.; Zhou, M.; Wang, P.; Li, X.; Ge, T.; Zheng, B. Enhancing prompt following with visual control through training-free mask-guided diffusion. arXiv 2024, arXiv:2404.14768. [Google Scholar] [CrossRef]
- Li, M.; Yang, T.; Kuang, H.; Wu, J.; Wang, Z.; Xiao, X.; Chen, C. ControlNet++: Improving Conditional Controls with Efficient Consistency Feedback. In Proceedings of the Computer Vision—ECCV 2024, Milan, Italy, 29 September–4 October 2024; Springer Nature: Cham, Switzerland, 2024; pp. 129–147. [Google Scholar] [CrossRef]
- Xu, S.; Ma, Z.; Huang, Y.; Lee, H.; Chai, J. Cyclenet: Rethinking cycle consistency in text-guided diffusion for image manipulation. Adv. Neural Inf. Process. Syst. 2023, 36, 10359–10384. [Google Scholar] [CrossRef]
- Cai, X.; Lai, Q.; Pei, G.; Shu, X.; Yao, Y.; Wang, W. Cycle-consistent learning for joint layout-to-image generation and object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Honolulu, HI, USA, 21–23 October 2025; pp. 6797–6807. [Google Scholar] [CrossRef]
- Li, J.; Li, B.; Tu, Z.; Liu, X.; Guo, Q.; Juefei-Xu, F.; Xu, R.; Yu, H. Light the night: A multi-condition diffusion framework for unpaired low-light enhancement in autonomous driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 15205–15215. [Google Scholar] [CrossRef]
- Yang, J.; Li, A.; Liao, X.; Masouros, C. Speeding-up symbol-level precoding using separable and dual optimizations. IEEE Trans. Commun. 2023, 71, 7056–7071. [Google Scholar] [CrossRef]
- Hu, X.; Wang, R.; Fang, Y.; Fu, B.; Cheng, P.; Yu, G. Ella: Equip diffusion models with llm for enhanced semantic alignment. arXiv 2024, arXiv:2403.05135. [Google Scholar] [CrossRef]
- Wu, T.H.; Lian, L.; Gonzalez, J.E.; Li, B.; Darrell, T. Self-correcting llm-controlled diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 6327–6336. [Google Scholar] [CrossRef]
- Yang, L.; Yu, Z.; Meng, C.; Xu, M.; Ermon, S.; Cui, B. Mastering text-to-image diffusion: Recaptioning, planning, and generating with multimodal LLMs. Proc. Mach. Learn. Res. 2024, 235, 56704–56721. [Google Scholar]
- Qin, C.; Zhang, S.; Yu, N.; Feng, Y.; Yang, X.; Zhou, Y.; Wang, H.; Niebles, J.C.; Xiong, C.; Savarese, S.; et al. UniControl: A Unified Diffusion Model for Controllable Visual Generation in the Wild. In Proceedings of the Advances in Neural Information Processing Systems 36, New Orleans, LA, USA, 10–16 December 2023; pp. 42961–42992. [Google Scholar] [CrossRef]
- Ye, H.; Zhang, J.; Liu, S.; Han, X.; Yang, W. IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models. arXiv 2023, arXiv:2308.06721. [Google Scholar] [CrossRef]
- He, X.; Zheng, J.; Fang, J.Z.; Piramuthu, R.; Bansal, M.; Ordonez, V.; Sigurdsson, G.A.; Peng, N.; Wang, X.E. FlexEControl: Flexible and efficient multimodal control for text-to-image generation. arXiv 2024, arXiv:2405.04834. [Google Scholar] [CrossRef]
- Xie, J.; Li, Y.; Huang, Y.; Liu, H.; Zhang, W.; Zheng, Y.; Shou, M.Z. Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 7452–7461. [Google Scholar] [CrossRef]
- Chefer, H.; Alaluf, Y.; Vinker, Y.; Wolf, L.; Cohen-Or, D. Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models. ACM Trans. Graph. (TOG) 2023, 42, 148. [Google Scholar] [CrossRef]
- Chen, J.; Huang, Y.; Lv, T.; Cui, L.; Chen, Q.; Wei, F. Textdiffuser: Diffusion models as text painters. Adv. Neural Inf. Process. Syst. 2023, 36, 9353–9387. [Google Scholar] [CrossRef]
- Chen, J.; Huang, Y.; Lv, T.; Cui, L.; Chen, Q.; Wei, F. Textdiffuser-2: Unleashing the power of language models for text rendering. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: New York, NY, USA, 2024; pp. 386–402. [Google Scholar] [CrossRef]
- Hertz, A.; Mokady, R.; Tenenbaum, J.; Aberman, K.; Pritch, Y.; Cohen-Or, D. Prompt-to-prompt image editing with cross attention control. arXiv 2022, arXiv:2208.01626. [Google Scholar] [CrossRef]
- Brooks, T.; Holynski, A.; Efros, A.A. Instructpix2pix: Learning to follow image editing instructions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–26 June 2023; pp. 18392–18402. [Google Scholar] [CrossRef]
- Meng, C.; He, Y.; Song, Y.; Song, J.; Wu, J.; Zhu, J.Y.; Ermon, S. Sdedit: Guided image synthesis and editing with stochastic differential equations. arXiv 2021, arXiv:2108.01073. [Google Scholar] [CrossRef]
- Couairon, G.; Verbeek, J.; Schwenk, H.; Cord, M. Diffedit: Diffusion-based semantic image editing with mask guidance. arXiv 2022, arXiv:2210.11427. [Google Scholar] [CrossRef]
- Huang, N.; Tang, F.; Dong, W.; Lee, T.Y.; Xu, C. Region-aware diffusion for zero-shot text-driven image editing. arXiv 2023, arXiv:2302.11797. [Google Scholar] [CrossRef]
- Bashkirova, D.; Lezama, J.; Sohn, K.; Saenko, K.; Essa, I. Masksketch: Unpaired structure-guided masked image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 1879–1889. [Google Scholar] [CrossRef]
- Shi, Y.; Xue, C.; Liew, J.H.; Pan, J.; Yan, H.; Zhang, W.; Tan, V.Y.; Bai, S. Dragdiffusion: Harnessing diffusion models for interactive point-based image editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 8839–8849. [Google Scholar] [CrossRef]
- Mokady, R.; Hertz, A.; Aberman, K.; Pritch, Y.; Cohen-Or, D. Null-text inversion for editing real images using guided diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 6038–6047. [Google Scholar] [CrossRef]
- Kawar, B.; Zada, S.; Lang, O.; Tov, O.; Chang, H.; Dekel, T.; Mosseri, I.; Irani, M. Imagic: Text-based real image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 6007–6017. [Google Scholar] [CrossRef]
- Garibi, D.; Patashnik, O.; Voynov, A.; Averbuch-Elor, H.; Cohen-Or, D. Renoise: Real image inversion through iterative noising. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: New York, NY, USA, 2024; pp. 395–413. [Google Scholar] [CrossRef]
- Miyake, D.; Iohara, A.; Saito, Y.; Tanaka, T. Negative-prompt inversion: Fast image inversion for editing with text-guided diffusion models. In Proceedings of the 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Tucson, AZ, USA, 28 February–4 March 2025; IEEE: New York, NY, USA, 2025; pp. 2063–2072. [Google Scholar] [CrossRef]
- Cao, M.; Wang, X.; Qi, Z.; Shan, Y.; Qie, X.; Zheng, Y. Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 22560–22570. [Google Scholar] [CrossRef]
- Tumanyan, N.; Geyer, M.; Bagon, S.; Dekel, T. Plug-and-play diffusion features for text-driven image-to-image translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 1921–1930. [Google Scholar] [CrossRef]
- Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A.H.; Chechik, G.; Cohen-Or, D. An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion. arXiv 2022, arXiv:2208.01618. [Google Scholar] [CrossRef]
- Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; Aberman, K. DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 22500–22510. [Google Scholar] [CrossRef]
- Kumari, N.; Zhang, B.; Zhang, R.; Shechtman, E.; Zhu, J.Y. Multi-Concept Customization of Text-to-Image Diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 1931–1941. [Google Scholar] [CrossRef]
- Han, L.; Li, Y.; Zhang, H.; Milanfar, P.; Metaxas, D.; Yang, F. Svdiff: Compact parameter space for diffusion fine-tuning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 7323–7334. [Google Scholar] [CrossRef]
- Tewel, Y.; Gal, R.; Chechik, G.; Atzmon, Y. Key-locked rank one editing for text-to-image personalization. In Proceedings of the ACM SIGGRAPH 2023 Conference Proceedings, Los Angeles, CA, USA, 23 July 2023; pp. 1–11. [Google Scholar] [CrossRef]
- Zhou, Y.; Zhou, D.; Cheng, M.M.; Feng, J.; Hou, Q. StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation. In Proceedings of the Advances in Neural Information Processing Systems 37, Vancouver, BC, Canada, 10–15 December 2024; pp. 110315–110340. [Google Scholar] [CrossRef]
- Shi, J.; Xiong, W.; Lin, Z.; Jung, H.J. InstantBooth: Personalized Text-to-Image Generation without Test-Time Finetuning. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; IEEE: New York, NY, USA, 2024; pp. 8543–8552. [Google Scholar] [CrossRef]
- Wang, Q.; Bai, X.; Wang, H.; Qin, Z.; Chen, A.; Li, H.; Tang, X.; Hu, Y. Instantid: Zero-shot identity-preserving generation in seconds. arXiv 2024, arXiv:2401.07519. [Google Scholar] [CrossRef]
- Li, Z.; Cao, M.; Wang, X.; Qi, Z.; Cheng, M.M.; Shan, Y. Photomaker: Customizing realistic human photos via stacked id embedding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 19–21 June 2024; pp. 8640–8650. [Google Scholar] [CrossRef]
- Allmendinger, S.; Zipperling, D.; Struppek, L.; Kühl, N. Collafuse: Collaborative diffusion models. arXiv 2024, arXiv:2406.14429. [Google Scholar] [CrossRef]
- Brack, M.; Friedrich, F.; Hintersdorf, D.; Struppek, L.; Schramowski, P.; Kersting, K. Sega: Instructing text-to-image models using semantic guidance. Adv. Neural Inf. Process. Syst. 2023, 36, 25365–25389. [Google Scholar] [CrossRef]
- Chen, X.; Huang, L.; Liu, Y.; Shen, Y.; Zhao, D.; Zhao, H. Anydoor: Zero-shot object-level image customization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 6593–6602. [Google Scholar] [CrossRef]
- Yang, B.; Gu, S.; Zhang, B.; Zhang, T.; Chen, X.; Sun, X.; Chen, D.; Wen, F. Paint by example: Exemplar-based image editing with diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 18381–18391. [Google Scholar] [CrossRef]
- Yang, B.; Su, H.; Gkanatsios, N.; Ke, T.W.; Jain, A.; Schneider, J.; Fragkiadaki, K. Diffusion-ES: Gradient-free planning with diffusion for autonomous and instruction-guided driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 15342–15353. [Google Scholar] [CrossRef]
- Kondo, K.; Tagliabue, A.; Cai, X.; Tewari, C.; Garcia, O.; Espitia-Alvarez, M.; How, J.P. Cgd: Constraint-guided diffusion policies for uav trajectory planning. arXiv 2024, arXiv:2405.01758. [Google Scholar] [CrossRef]
- Dorjsembe, Z.; Pao, H.K.; Odonchimed, S.; Xiao, F. Conditional diffusion models for semantic 3D brain MRI synthesis. IEEE J. Biomed. Health Inform. 2024, 28, 4084–4093. [Google Scholar] [CrossRef] [PubMed]
- Konz, N.; Chen, Y.; Dong, H.; Mazurowski, M.A. Anatomically-controllable medical image generation with segmentation-guided diffusion models. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Marrakesh, Morocco, 6–10 October 2024; Springer: New York, NY, USA, 2024; pp. 88–98. [Google Scholar] [CrossRef]
- Bai, X.; Pu, X.; Xu, F. Conditional diffusion for SAR to optical image translation. IEEE Geosci. Remote Sens. Lett. 2024, 21, 4000605. [Google Scholar] [CrossRef]
- Bai, X.; Xu, F. Accelerating diffusion for sar-to-optical image translation via adversarial consistency distillation. arXiv 2024, arXiv:2407.06095. [Google Scholar] [CrossRef]
- Huang, J.; Dong, X.; Song, W.; Chong, Z.; Tang, Z.; Zhou, J.; Cheng, Y.; Chen, L.; Li, H.; Yan, Y.; et al. Consistentid: Portrait generation with multimodal fine-grained identity preserving. IEEE Trans. Pattern Anal. Mach. Intell. 2026, 48, 5639–5654. [Google Scholar] [CrossRef] [PubMed]
- Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; Song, S. Diffusion policy: Visuomotor policy learning via action diffusion. Int. J. Robot. Res. 2025, 44, 1684–1704. [Google Scholar] [CrossRef]
- Baldrati, A.; Morelli, D.; Cartella, G.; Cornia, M.; Bertini, M.; Cucchiara, R. Multimodal garment designer: Human-centric latent diffusion models for fashion image editing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 23393–23402. [Google Scholar] [CrossRef]
- Choi, Y.; Kwak, S.; Lee, K.; Choi, H.; Shin, J. Improving diffusion models for authentic virtual try-on in the wild. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; pp. 206–235. [Google Scholar] [CrossRef]
- Wan, S.; Li, Y.; Chen, J.; Pan, Y.; Yao, T.; Cao, Y.; Mei, T. Improving virtual try-on with garment-focused diffusion models. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; pp. 184–199. [Google Scholar] [CrossRef] [PubMed]
- Zeng, J.; Song, D.; Nie, W.; Tian, H.; Wang, T.; Liu, A.A. Cat-dm: Controllable accelerated virtual try-on with diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 8372–8382. [Google Scholar] [CrossRef]
- Sun, Z.; Zhou, Y.; He, H.; Mok, P. Sgdiff: A style guided diffusion model for fashion synthesis. In Proceedings of the 31st ACM International Conference on Multimedia, Ottawa, ON, Canada, 29 October–3 November 2023; pp. 8433–8442. [Google Scholar] [CrossRef]
- Cao, S.; Chai, W.; Hao, S.; Zhang, Y.; Chen, H.; Wang, G. Difffashion: Reference-based fashion design with structure-aware transfer by diffusion models. IEEE Trans. Multimed. 2024, 26, 3962–3975. [Google Scholar] [CrossRef]
- Wang, T.; Ye, M. Texfit: Text-driven fashion image editing with diffusion models. Proc. AAAI Conf. Artif. Intell. 2024, 38, 10198–10206. [Google Scholar] [CrossRef]
- Lampe, A.; Stopar, J.; Jain, D.K.; Omachi, S.; Peer, P.; Štruc, V. Dicti: Diffusion-based clothing designer via text-guided input. In Proceedings of the 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG), Istanbul, Turkey, 27–31 May 2024; IEEE: New York, NY, USA, 2024; pp. 1–9. [Google Scholar] [CrossRef]
- Wu, W.; Zhao, Y.; Shou, M.Z.; Zhou, H.; Shen, C. Diffumask: Synthesizing images with pixel-level annotations for semantic segmentation using diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 1206–1217. [Google Scholar] [CrossRef]
- Chen, S.; Sun, P.; Song, Y.; Luo, P. Diffusiondet: Diffusion model for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 19830–19843. [Google Scholar] [CrossRef]
- Saragih, D.G.; Hibi, A.; Tyrrell, P.N. Using diffusion models to generate synthetic labeled data for medical image segmentation. Int. J. Comput. Assist. Radiol. Surg. 2024, 19, 1615–1625. [Google Scholar] [CrossRef] [PubMed]
- Alemohammad, S.; Humayun, A.I.; Agarwal, S.; Collomosse, J.; Baraniuk, R. Self-improving diffusion models with synthetic data. arXiv 2024, arXiv:2408.16333. [Google Scholar] [CrossRef]
- Lee, T.; Yasunaga, M.; Meng, C.; Mai, Y.; Park, J.S.; Gupta, A.; Zhang, Y.; Narayanan, D.; Teufel, H.; Bellagente, M.; et al. Holistic evaluation of text-to-image models. Adv. Neural Inf. Process. Syst. 2023, 36, 69981–70011. [Google Scholar] [CrossRef]
- Hu, Y.; Liu, B.; Kasai, J.; Wang, Y.; Ostendorf, M.; Krishna, R.; Smith, N.A. Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 20406–20417. [Google Scholar] [CrossRef]
- Sun, K.; Huang, K.; Liu, X.; Wu, Y.; Xu, Z.; Li, Z.; Liu, X. T2v-compbench: A comprehensive benchmark for compositional text-to-video generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 8406–8416. [Google Scholar] [CrossRef]
- Li, B.; Lin, Z.; Pathak, D.; Li, J.; Fei, Y.; Wu, K.; Ling, T.; Xia, X.; Zhang, P.; Neubig, G.; et al. Genai-bench: Evaluating and improving compositional text-to-visual generation. arXiv 2024, arXiv:2406.13743. [Google Scholar] [CrossRef]
- Ghosh, D.; Hajishirzi, H.; Schmidt, L. Geneval: An object-focused framework for evaluating text-to-image alignment. Adv. Neural Inf. Process. Syst. 2023, 36, 52132–52152. [Google Scholar] [CrossRef]
- Cho, J.; Hu, Y.; Garg, R.; Anderson, P.; Krishna, R.; Baldridge, J.; Bansal, M.; Pont-Tuset, J.; Wang, S. Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation. arXiv 2023, arXiv:2310.18235. [Google Scholar] [CrossRef]
- Kamath, A.; Chang, K.W.; Krishna, R.; Zettlemoyer, L.; Hu, Y.; Ghazvininejad, M. GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation. arXiv 2025, arXiv:2512.16853. [Google Scholar] [CrossRef]
- Song, Y.; Dhariwal, P.; Chen, M.; Sutskever, I. Consistency Models. Proc. Mach. Learn. Res. 2023, 202, 32211–32252. [Google Scholar]
- Luo, S.; Tan, Y.; Huang, L.; Li, J.; Zhao, H. Latent consistency models: Synthesizing high-resolution images with few-step inference. arXiv 2023, arXiv:2310.04378. [Google Scholar] [CrossRef]
- Sauer, A.; Lorenz, D.; Blattmann, A.; Rombach, R. Adversarial diffusion distillation. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; pp. 87–103. [Google Scholar] [CrossRef]
- Zhao, Y.; Xu, Y.; Xiao, Z.; Jia, H.; Hou, T. Mobilediffusion: Instant text-to-image generation on mobile devices. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; pp. 225–242. [Google Scholar] [CrossRef]
- Fang, G.; Li, K.; Ma, X.; Wang, X. Tinyfusion: Diffusion transformers learned shallow. In Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA, 11–15 June 2025; pp. 18144–18154. [Google Scholar] [CrossRef]
- Fernandez, P.; Couairon, G.; Jégou, H.; Douze, M.; Furon, T. The stable signature: Rooting watermarks in latent diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Nashville, TN, USA, 13–15 June 2023; pp. 22466–22477. [Google Scholar] [CrossRef]
- Wen, Y.; Kirchenbauer, J.; Geiping, J.; Goldstein, T. Tree-ring watermarks: Fingerprints for diffusion images that are invisible and robust. arXiv 2023, arXiv:2305.20030. [Google Scholar] [CrossRef]
- Gowal, S.; Bunel, R.; Stimberg, F.; Stutz, D.; Ortiz-Jimenez, G.; Kouridi, C.; Vecerik, M.; Hayes, J.; Rebuffi, S.A.; Bernard, P.; et al. SynthID-Image: Image watermarking at internet scale. arXiv 2025, arXiv:2510.09263. [Google Scholar] [CrossRef]


| Method | Conditions | Interface | Fusion/Arbitration | Main Trade-Off |
|---|---|---|---|---|
| ControlNet [6] | text + spatial controls | feature branch | additive branch injection | strong control; limited explicit conflict arbitration |
| SpaText [73] | global text + region text + segmentation | attention | region-wise text guidance | fine-grained semantics; costly with dense regions |
| Composable Diffusion [74] | multiple textual concepts | sampling score | compositional score fusion | flexible without retraining; weak spatial grounding |
| MultiDiffusion [76] | text + regional constraints | trajectory/sampling | coupled regional diffusion paths | strong layout control; higher inference cost |
| Composer [77] | text, depth, sketch, color | feature injection | unified factor composition | broad coverage; inconsistent inputs remain difficult |
| Uni-ControlNet [71] | local + global controls | adapters | unified local/global adapters | scalable interface; variable control strength |
| Cocktail [78] | multimodal regional controls | feature + sampling | ControlNorm and spatial guidance | robust multimodal fusion; added tuning burden |
| Dynamic Control [70] | multiple candidate conditions | scheduling | adaptive selection and weighting | handles redundancy/conflict; extra policy cost |
| Method | Backbone | Task/Dataset | Metrics | Reported Results |
|---|---|---|---|---|
| ControlNet [6] | Stable Diffusion | ADE20K segmentation; sketch study | FID, CLIP, IoU, human rank | FID 15.27; CLIP 0.26; IoU ; human quality/fidelity . |
| T2I-Adapter [7] | SD 1.4 | COCO text + segmentation/sketch | FID, CLIP | Seg.: FID 16.78, CLIP 0.2652; sketch: FID 17.36, CLIP 0.2666. |
| GLIGEN [68] | LDM/SD | COCO grounded generation | FID, YOLO AP | COCO2014D: FID 5.61, AP/AP50/AP75 24.0/42.2/24.1; COCO2017 FID 21.04. |
| MultiDiffusion [76] | Stable Diffusion | text-to-panorama | FID, CLIP, aesthetic | FID ; CLIP 0.27; aesthetic 6.36. |
| DynamicControl [70] | SD 1.5 | MultiGen-20M, ADE20K, COCO-Stuff | F1, SSIM, mAP, RMSE, mIoU | Canny F1 39.26; HED SSIM 0.8376; OpenPose mAP 82.63; depth RMSE 23.21. |
| UniCombine [79] | FLUX.1-schnell | multi-spatial and subject control | FID, F1, MSE, CLIP-I/DINO/T | Multi-spatial: FID 6.82, F1 0.64, CLIP-T 33.45; subject insertion: FID 4.55. |
| Method | Main Failure Addressed | Canny F1 ↑ | HED SSIM ↑ | OpenPose mAP ↑ | Depth RMSE ↓ | ADE20K mIoU ↑ | COCO-Stuff mIoU ↑ |
|---|---|---|---|---|---|---|---|
| T2I-Adapter-SDXL [70] | single-condition structural control with lightweight adaptation | 28.03 | NR | 63.89 | 39.76 | NR | NR |
| T2I-Adapter-SD1.5 [70] | lightweight adapter-based control under Stable Diffusion 1.5 | 23.66 | NR | 60.17 | 48.40 | 12.60 | NR |
| GLIGEN-SD1.4 [68,70] | grounding and spatial anchoring for object-level control | 26.92 | 0.5641 | 69.88 | 38.82 | 23.77 | NR |
| Uni-ControlNet [70] | unified local and global condition composition | 27.31 | 0.6912 | 72.71 | 40.66 | 19.39 | NR |
| UniControl [70] | unified controllable visual generation | 30.83 | 0.7967 | 75.87 | 39.17 | 25.45 | NR |
| ControlNet-SD1.5 [6,70] | strong branch-based structural injection | 34.66 | 0.7622 | NR | NR | 32.56 | 27.47 |
| Cocktail [70] | multimodal control fusion and spatially guided sampling | 35.22 | 0.8152 | 78.82 | 35.90 | 36.55 | 29.68 |
| ControlNet++ [70] | consistency-feedback enhanced controllability | 37.04 | 0.8097 | NR | 28.32 | 43.64 | 34.56 |
| DynamicControl [70] | adaptive condition selection and conflict reduction | 39.26 | 0.8376 | 82.63 | 23.21 | 48.56 | 37.78 |
| Method | Parameter Overhead | Train? | Reported Training Hardware/Time | Latency Complexity Scaling |
|---|---|---|---|---|
| ControlNet [6] | approximately one trainable encoder copy per control | Yes | 1× RTX 3090 Ti (24 GB), 5 days for the reported depth model | for K independently stacked branches |
| T2I-Adapter [7] | 77 M; compressed variants 18 M/5 M | Yes | 4× Tesla V100 32 GB, within 3 days | one backbone pass plus |
| GLIGEN [68] | gated grounding layers; total count NR | Yes | NR | one trajectory; attention cost grows with grounding tokens |
| MultiDiffusion [76] | 0 task-specific parameters | No | Not applicable | for W overlapping denoising windows or paths |
| DynamicControl [70] | condition evaluator plus multi-control adapter | Yes | NR | evaluation over K inputs, then denoising with the selected subset |
| UniCombine [79] | 29 M training-free/44 M training-based for two conditions | Optional | 16× V100, 30,000 steps for training-based mode | one shared trajectory; conditional-attention cost grows with condition tokens |
| Method | Reproducibility Role | Official Repository |
|---|---|---|
| ControlNet [6] | branch-based structural-control baseline | https://github.com/lllyasviel/ControlNet (accessed on 21 July 2026) |
| T2I-Adapter [7] | lightweight composable-adapter baseline | https://github.com/TencentARC/T2I-Adapter (accessed on 21 July 2026) |
| GLIGEN [68] | grounded box and region control | https://github.com/gligen/GLIGEN (accessed on 21 July 2026) |
| MultiDiffusion [76] | multi-region and panorama inference | https://github.com/omerbt/MultiDiffusion (accessed on 21 July 2026) |
| ControlNet++ [87] | consistency-feedback control | https://github.com/liming-ai/ControlNet_Plus_Plus (accessed on 21 July 2026) |
| SLD [93] | LLM-audited closed-loop correction | https://github.com/tsunghan-wu/SLD (accessed on 21 July 2026) |
| RPG [94] | multimodal-LLM regional planning | https://github.com/YangLing0818/RPG-DiffusionMaster (accessed on 21 July 2026) |
| StoryDiffusion [120] | cross-image subject consistency | https://github.com/HVision-NKU/StoryDiffusion (accessed on 21 July 2026) |
| OmniGen [51] | unified generation and editing | https://github.com/VectorSpaceLab/OmniGen (accessed on 21 July 2026) |
| OminiControl [52] | parameter-efficient DiT condition injection | https://github.com/Yuanshi9815/OminiControl (accessed on 21 July 2026) |
| FLUX.1 [50] | rectified-flow Transformer backbone | https://github.com/black-forest-labs/flux (accessed on 21 July 2026) |
| Evaluation Target | Representative Metrics | Benchmark/Paper | Use in Collaborative Control |
|---|---|---|---|
| Holistic T2I capability | alignment, quality, aesthetics, reasoning, fairness, robustness, efficiency | HEIM [148] | broad system-level assessment across 12 aspects and 62 scenarios |
| Text-image faithfulness | VQA-based TIFA score; counting, relation, and attribute checks | TIFA [149] | diagnoses semantic adherence beyond CLIP-style global similarity |
| Object compositionality | co-occurrence, position, count, color accuracy | GenEval [152]; GenEval 2 [154] | measures common failures in multi-object and layout-sensitive prompts |
| Fine-grained semantics | dependency-structured question answering | DSG [153] | improves relation-centric diagnosis for complex scenes |
| Structural control | Canny F1, HED SSIM, OpenPose mAP, depth RMSE, mIoU | DynamicControl [70] | evaluates whether edge, pose, depth, or segmentation conditions are preserved |
| Grounding and layout | YOLO AP/AP50/AP75, GLIP score, FID | GLIGEN [68] | checks object-region alignment together with image quality |
| Subject consistency | CLIP-I, DINO; plus FID/SSIM/CLIP-T | UniCombine [79] | evaluates identity preservation under subject-spatial composition |
| Efficiency | GPU memory, extra parameters, inference cost | UniCombine/T2I-Adapter/ ControlNet [6,7,79] | captures deployment overhead introduced by additional controls |
| Benchmark | Target | Granularity | Scope | Signals | Limitation | ControlAnalysis Use |
|---|---|---|---|---|---|---|
| HEIM [148] | holistic evaluation of text-to-image systems | system/multi-dimensional level | primarily T2I | alignment, image quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, efficiency | not specifically designed for structured control or condition arbitration | useful for global system-level assessment, but less diagnostic for fine-grained control errors |
| TIFA [149] | text-image faithfulness | object/attribute/relation level | primarily T2I | question answering derived from prompts and answered from generated images | depends on the reliability of the underlying VQA pipeline | useful for checking whether generated outputs satisfy explicit semantic conditions |
| T2I-CompBench [150] | composition-focused T2I evaluation | composition level | primarily T2I | attribute binding, object relationships, and complex compositions | less comprehensive outside composition scenarios | useful for analyzing semantic conflicts and compositional failures under collaborative control |
| GenAI-Bench [151] | composition and reasoning evaluation | image/video/reasoning level | text-to-visual generation | human preference, composition prompts, and automatic signals | less focused on explicit structural control interfaces | useful for analyzing whether stronger planning or reasoning modules benefit controllable generation |
| GenEval [152] | object-focused evaluation of generated images | object/relation level | primarily T2I | object co-occurrence, position, count, and color | limited primitive coverage relative to broader real-world control requirements | useful for measuring common precise-generation failures such as counting and layout errors |
| DSG [153] | reliability of fine-grained text-image evaluation | relation/semantic-graph level | primarily T2I | dependency-structured question generation and answering | evaluation still depends on external models | useful for relation-centric diagnosis in complex multi-entity scenes |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Qi, J.; Xu, W.; Zhu, Q.; Chu, X.; Wang, Y. Collaborative Control in Diffusion Models for Precise Image Generation: A Survey. Mathematics 2026, 14, 2737. https://doi.org/10.3390/math14152737
Qi J, Xu W, Zhu Q, Chu X, Wang Y. Collaborative Control in Diffusion Models for Precise Image Generation: A Survey. Mathematics. 2026; 14(15):2737. https://doi.org/10.3390/math14152737
Chicago/Turabian StyleQi, Jingzhong, Wei Xu, Qing Zhu, Xinchen Chu, and Yifan Wang. 2026. "Collaborative Control in Diffusion Models for Precise Image Generation: A Survey" Mathematics 14, no. 15: 2737. https://doi.org/10.3390/math14152737
APA StyleQi, J., Xu, W., Zhu, Q., Chu, X., & Wang, Y. (2026). Collaborative Control in Diffusion Models for Precise Image Generation: A Survey. Mathematics, 14(15), 2737. https://doi.org/10.3390/math14152737

