AI for Garden Design Visualization: Development and Validation of the GardenDiff Model
Abstract
1. Introduction
2. Methodology
2.1. Overall Framework of Methods
2.2. Dataset Construction and Annotation Strategy
2.2.1. Image Collection and Garden Style Selection
2.2.2. Construction of Structured Design Captioning
2.3. Experimental Setup and Training Configuration
2.3.1. Base Model and Training Method Selection
2.3.2. Training Parameters and Phased Implementation
2.4. Evaluation Methods
2.4.1. Evaluation Methods for Parameter Optimization Experiments
2.4.2. Evaluation Methods for Multi-Model Comparison Experiments
2.4.3. Questionnaire Survey and Data Analysis
3. Results
3.1. Training Convergence and Model Selection Results
3.2. Image Generation Strategy for Evaluation
3.3. Effects of Caption Systems and Training Resolution on Generation Quality
3.4. Performance Validation and Multi-Style Analysis of the GardenDiff Model
3.4.1. Comparison of the Performance of GardenDiff Model and Conventional Models
3.4.2. Style-Specific Performance Analysis
3.4.3. Expert–Public Evaluation Comparison
4. Discussion
4.1. Training Parameter Effects on Generation Quality
4.2. The Performance Advantages of GardenDiff
4.3. Style-Specific Performance Variations and Its Influencing Factors
4.4. Limitations
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| BLIP | Bootstrapping Language–Image Pre-training |
| CLIP | Contrastive Language–Image Pre-training |
| ControlNet | ControlNet (proper noun, not an abbreviation) |
| DreamShaper XL | DreamShaper XL, community fine-tuned diffusion model |
| GardenDiff | Garden Diffusion Model (proposed in this study) |
| LAION | Large-scale Artificial Intelligence Open Network |
| LoRA | Low-Rank Adaptation |
| SBE | Scenic Beauty Estimation |
| SDC | Structured Design Captioning |
| SDXL | Stable Diffusion XL |
| WD1.4 | Waifu Diffusion 1.4 Tagger |
Appendix A. Disadvantage Identified in Preliminary Baseline Experiments
| Semantic Misalignment | Detail Rendering Deficiency | |
|---|---|---|
| Description | The models misinterpret design terminology, producing inconsistent styles and ambiguous element layouts. | AI models struggle to render complex design elements accurately. |
| Baseline Generation Overall View (No Control) | ![]() | ![]() |
| Baseline Generation Detail View (No Control) | ![]() | ![]() |
| ControlNet- Assisted Generation Overall View | ![]() | ![]() |
| ControlNet- Assisted Generation Detail View | ![]() | ![]() |
Appendix B. Dataset and Captioning Details
Appendix B.1. Dataset Composition
| Garden Style | Design Category | No. of Images | Repeat Times | Data Source |
|---|---|---|---|---|
| Chinese (250 images) | Architectural Facades | 50 | 20 | https://www.freepik.com/ (CC0 License) https://www.pexels.com/zh-cn/ (CC0 License) |
| Garden Structures | 50 | 20 | ||
| Paving | 40 | 20 | ||
| Planting | 60 | 20 | ||
| Detail Close-ups | 50 | 15 | ||
| Japanese (250 images) | Architectural Facades | 23 | 20 | https://www.freepik.com/ (CC0 License) https://unsplash.com/ (CC0 License) |
| Garden Structures | 66 | 20 | ||
| Paving | 80 | 20 | ||
| Planting | 62 | 20 | ||
| Detail Close-ups | 20 | 15 | ||
| Mediterranean (250 images) | Architectural Facades | 60 | 20 | https://www.freepik.com/ (CC0 License) https://unsplash.com/ (CC0 License)) https://www.pexels.com/zh-cn/ (CC0 License) |
| Garden Structures | 64 | 20 | ||
| Paving | 40 | 20 | ||
| Planting | 52 | 20 | ||
| Detail Close-ups | 34 | 15 | ||
| Nordic (250 images) | Architectural Facades | 65 | 20 | https://www.freepik.com/ (CC0 License) https://unsplash.com/ (CC0 License) |
| Garden Structures | 55 | 20 | ||
| Paving | 40 | 20 | ||
| Planting | 50 | 20 | ||
| Detail Close-ups | 40 | 15 | ||
| English (250 images) | Architectural Facades | 45 | 20 | https://www.freepik.com/ (CC0 License) https://www.pexels.com/zh-cn/ (CC0 License) |
| Garden Structures | 60 | 20 | ||
| Paving | 30 | 20 | ||
| Planting | 75 | 20 | ||
| Detail Close-ups | 40 | 15 |
Appendix B.2. Semantic Annotation Examples
| Image Example | Wd1.4 | Blip | SDC |
|---|---|---|---|
![]() | scenery, reflection, tree, outdoors, water, architecture, east_asian_architecture, no_humans, hood, nature, day, plant, grass, forest, hoodie, building | a pond with lily pads and a building in the background with a pond in front of it and a lily pad in the foreground, Cao Buxing, phuoc quan, a detailed matte painting, cloisonnism | chinese courtyard, peaceful, upturned eaves, white walls, black tile roofs, pavilion, corridor, reflecting pool, stones, water lilies, bamboo, lush greenery, serene atmosphere, summer season, tranquil water reflections |
![]() | no_humans, tree, scenery, outdoors, traditional_media, real_world_location, photo_background, day, sky, building, bare_tree, east_asian_architecture, architecture, realistic, road, | a small building with a tree in the middle of it and a snow covered ground in front of it, Eishōsai Chōki, kyoto studio, a tilt shift photo, mingei | japanese courtyard design, minimalism, wooden structure, tiled roof, stone paving, raked gravel, pine tree, deciduous trees, bamboo, moss, overcast weather, tranquil atmosphere |
![]() | scenery, tree, no_humans, sky, outdoors, plant, water, day, building, blue_sky, statue | a fountain in a formal garden with hedges and potted plants in front of a white building with arches, Enguerrand Quarton, arthouse, a digital rendering, heidelberg school | mediterranean courtyard, elegant style, red clay roof tiles, white facade with arched openings, central stone fountain with lion sculptures, terracotta flower pots with orange trees, trimmed hedges, clear blue sky, formal layout |
![]() | scenery, tree, stairs, outdoors, grass, day, no_humans, sunlight, plant, nature, moss, multiple_girls, | a garden with a staircase and a fountain in the middle of it and a stone staircase leading to the upper level, Enguerrand Quarton, enchanting, a flemish Baroque, arts and crafts movement | English courtyard with lush greenery, brick walls, stone stairs, ornate urns, symmetrical stone path, trimmed boxwood hedges, vine-covered walls, sculptures, bright sunny weather, perspective from ground level, strong visual hierarchy |
![]() | no_humans, scenery, tree, outdoors, house, building, road, window, ground_vehicle, autumn_leaves, door, bench, plant | a small cabin with a deck and lights on it’s side in the woods near a picnic table, Dan Frazier, archdaily, a digital rendering, arts and crafts movement | nordic courtyard, minimalist design, dark wood facade, large glass sliding door, outdoor seating, string lights, wooden deck, natural paving, deciduous trees, autumn leaves, evening setting, warm lighting, open courtyard |
Appendix B.3. Annotation Prompt Template
Appendix C. Evaluation Questionnaire Samples
Appendix C.1. Parameter Optimization Experiment Questionnaire Sample

Appendix C.2. Multi-Model Comparison Experiment Questionnaire Sample

Appendix D. Checkpoint Selection Details

Appendix E. Inference Parameters
| Category | Parameter | Value/Setting | Note |
|---|---|---|---|
| Hardware | GPU | RTX4090(24 GB)/64 GBRAM | CUDA acceleration enabled |
| Training toolkit | Kohya-ss(sd-scripts) | Fine-grained control over training parameters | |
| Inference UI | WebUIForge(f2.0.1) | CoreVersion:gf5330788 | |
| Model training | Base model | SDXL1.0(VAEFix) | Base model version |
| Bucket resolution | 1536 × 1536(Bucketing) | Covers 256–2560 px with multiple aspect ratios | |
| Learning rate | 1 | Compatible with AdaFactor optimizer | |
| Optimizer | AdaFactor | Adaptive learning rate schedule | |
| Training epochs | 20Epochs | Loss monitoring + visual review for checkpoint selection | |
| Inference | Sampling method | DPM++SDEKarras | Suitable for generating fine-detailed textures |
| Sampling steps | 20 Steps | Balances generation speed and detail fidelity | |
| Generation resolution | 1720 × 1280/1328 × 1760 | Matches training resolution density | |
| CFG scale | 4.0–5.0 | Low values to achieve natural color rendering | |
| Seed | −1(Random) | Multiple images generated per prompt; manually selected | |
| ControlNet | Preprocessor | Softedge_teed | Processor resolution: 2048 px |
| Control model (example) | canny-sdxl-V2.0 | Canny model used for edge-based soft control | |
| Control weight | 0.5–0.8(dynamic) | Adjusted based on base map complexity |
Appendix F. Additional Technical Analysis
Appendix F.1. Image Generation Stage Details
| Generation Resolution | 768 × 768 | 1024 × 1024 | 1536 × 1536 |
|---|---|---|---|
| Generation Prompt | chinese style, upturned eaves, black tile roofs, white walls, Rockery artificial rockwork, Hedge, shrub, bamboo, stone paving | ||
| SDXL Base 1.0 | ![]() | ![]() | ![]() |
| 768 × 768 Chinese Mode | ![]() | ![]() | ![]() |
| 1024 × 1024 Chinese Model | ![]() | ![]() | ![]() |
| 1536 × 1536 Chinese Model | ![]() | ![]() | ![]() |
Appendix F.2. Control Plugin Impact Analysis
| Canny V2 | MistoLine | |
|---|---|---|
| Base Model—Chinese | ![]() | ![]() |
| Chinese | ![]() | ![]() |
| Japanese | ![]() | ![]() |
| Mediterranean | ![]() | ![]() |
| Nordic | ![]() | ![]() |
| English | ![]() | ![]() |
References
- Dong, X.; Geng, L. Nature Deficit and Mental Health among Adolescents: A Perspectives of Conservation of Resources Theory. J. Environ. Psychol. 2023, 87, 101995. [Google Scholar] [CrossRef]
- Seastedt, H.; Schuetz, J.; Perkins, A.; Gamble, M.; Sinkkonen, A. Impact of Urban Biodiversity and Climate Change on Children’s Health and Well Being. Pediatr. Res. 2025, 98, 452–457. [Google Scholar] [CrossRef]
- Fu, E.; Zhou, J.; Ren, Y.; Deng, X.; Li, L.; Li, X.; Li, X. Exploring the Influence of Residential Courtyard Space Landscape Elements on People’s Emotional Health in an Immersive Virtual Environment. Front. Public Health 2022, 10, 1017993. [Google Scholar] [CrossRef] [PubMed]
- Soflaei, F.; Shokouhian, M.; Soflaei, A. Traditional Courtyard Houses as a Model for Sustainable Design: A Case Study on BWhs Mesoclimate of Iran. Front. Archit. Res. 2017, 6, 329–345. [Google Scholar] [CrossRef]
- Schroth, O.; Maier, A. Integrating Generative Artificial Intelligence into the Landscape Architecture Design Process. J. Digit. Landsc. Archit. 2025, 10, 665–675. [Google Scholar]
- Wang, Q.; Liang, Y.; Zheng, Y.; Xu, K.; Zhao, J.; Wang, S. Generative AI for Urban Planning: Synthesizing Satellite Imagery via Diffusion Models. Comput. Environ. Urban Syst. 2025, 122, 102339. [Google Scholar] [CrossRef]
- Jang, S.; Roh, H.; Lee, G. Generative AI in Architectural Design: Application, Data, and Evaluation Methods. Autom. Constr. 2025, 174, 106174. [Google Scholar] [CrossRef]
- Yang, L.; Zhang, Z.; Song, Y.; Hong, S.; Xu, R.; Zhao, Y.; Zhang, W.; Cui, B.; Yang, M.-H. Diffusion Models: A Comprehensive Survey of Methods and Applications. ACM Comput. Surv. 2024, 56, 105. [Google Scholar] [CrossRef]
- Zhang, L.; Rao, A.; Agrawala, M. Adding Conditional Control to Text-to-Image Diffusion Models. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2023; pp. 3813–3824. [Google Scholar]
- Zhou, H.; Xiang, S. Applicability Evaluation and Reflection on Artificial Intelligence-Based “Image to Image” Generation of Landscape Architecture Masterplans. Landsc. Archit. Front. 2024, 12, 58–67. [Google Scholar] [CrossRef]
- Chen, R.; Zhao, J.; Yao, X.; Jiang, S.; He, Y.; Bao, B.; Luo, X.; Xu, S.; Wang, C. Generative Design of Outdoor Green Spaces Based on Generative Adversarial Networks. Buildings 2023, 13, 1083. [Google Scholar] [CrossRef]
- Chen, R.; Zhao, J.; Yao, X.; He, Y.; Li, Y.; Lian, Z.; Han, Z.; Yi, X.; Li, H. Enhancing Urban Landscape Design: A GAN-Based Approach for Rapid Color Rendering of Park Sketches. Land 2024, 13, 254. [Google Scholar] [CrossRef]
- Chen, R.; Yi, X.; Zhao, J.; He, Y.; Chen, B.; Liu, F.; Yao, X.; Jiang, X.; Lian, Z.; Li, H. AI for Landscape Planning: Assessing Surrounding Contextual Impact on GAN-Generated Green Land Layouts. Cities 2025, 166, 106181. [Google Scholar] [CrossRef]
- Ye, X.; Huang, T.; Song, Y.; Li, X.; Newman, G.; Wu, D.J.; Zeng, Y. Generating Conceptual Landscape Design via Text-to-Image Generative AI Model. Environ. Plan. B Urban Anal. City Sci. 2025, 52, 1903–1919. [Google Scholar] [CrossRef]
- Chen, F.; Mai, M.; Huang, X.; Li, Y. Enhancing the Sustainability of AI Technology in Architectural Design: Improving the Matching Accuracy of Chinese-Style Buildings. Sustainability 2024, 16, 8414. [Google Scholar] [CrossRef]
- Liang, J. The Application of Artificial Intelligence-Assisted Technology in Cultural and Creative Product Design. Sci. Rep. 2024, 14, 31069. [Google Scholar] [CrossRef]
- Lu, L.; Liu, M. Exploring a Spatial-Experiential Structure within the Chinese Literati Garden: The Master of the Nets Garden as a Case Study. Front. Archit. Res. 2023, 12, 923–946. [Google Scholar] [CrossRef]
- Hoyle, H.E. Climate-Adapted, Traditional or Cottage-Garden Planting? Public Perceptions, Values and Socio-Cultural Drivers in a Designed Garden Setting. Urban For. Urban Green. 2021, 65, 127362. [Google Scholar] [CrossRef]
- Zador, A.; Escola, S.; Richards, B.; Ölveczky, B.; Bengio, Y.; Boahen, K.; Botvinick, M.; Chklovskii, D.; Churchland, A.; Clopath, C.; et al. Catalyzing Next-Generation Artificial Intelligence through NeuroAI. Nat. Commun. 2023, 14, 1597. [Google Scholar] [CrossRef] [PubMed]
- Ananthram, A.; Stengel-Eskin, E.; Bansal, M.; McKeown, K. See It from My Perspective: How Language Affects Cultural Bias in Image Understanding. In Proceedings of the International Conference on Learning Representations 2025 (ICLR 2025), Singapore, 24–28 April 2024. [Google Scholar]
- Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; Müller, J.; Penna, J.; Rombach, R. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis. In Proceedings of the International Conference on Learning Representations 2024 (ICLR 2024), Vienna, Austria, 7–11 May 2024. [Google Scholar]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2022; pp. 10674–10685. [Google Scholar]
- Saharia, C.; Chan, W.; Saxena, S.; Li, L.; Whang, J.; Denton, E.L.; Ghasemipour, K.; Gontijo Lopes, R.; Karagol Ayan, B.; Salimans, T.; et al. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding. In Advances in Neural Information Processing Systems 35; Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; Curran Associates, Inc.: New York, NY, USA, 2022; pp. 36479–36494. [Google Scholar]
- SmilingWolf Wd-v1-4-Vit-Tagger-V2. Available online: https://huggingface.co/SmilingWolf/wd-v1-4-vit-tagger-v2 (accessed on 16 November 2025).
- Lyu, M.; Yang, Y.; Hong, H.; Chen, H.; Jin, X.; He, Y.; Xue, H.; Han, J.; Ding, G. One-Dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing Applications. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024; pp. 7559–7568. [Google Scholar]
- Wu, W.; Zhao, Y.; Chen, H.; Gu, Y.; Zhao, R.; He, Y.; Zhou, H.; Shou, M.Z.; Shen, C. DatasetDM: Synthesizing Data with Perception Annotations Using Diffusion Models. In Proceedings of the Advances in Neural Information Processing Systems 36 (NeurIPS 2023), New Orleans, LA, USA, 10–16 December 2023. [Google Scholar]
- Lykon DreamShaper XL. Available online: https://huggingface.co/Lykon/dreamshaper-xl-1-0 (accessed on 17 November 2025).
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. arXiv 2022, arXiv:2106.09685. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models from Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, Virtual Event, 18–24 July 2021. [Google Scholar]
- Chen, S.; Dewancker, B.J. The Influence of Zen Buddhism and Ink Wash Painting on Japanese Gardens during the Medieval Japan. J. Asian Archit. Build. Eng. 2025, 24, 5024–5036. [Google Scholar] [CrossRef]
- Van Tonder, G.J.; Lyons, M.J.; Ejima, Y. Visual Structure of a Japanese Zen Garden. Nature 2002, 419, 359–360. [Google Scholar] [CrossRef]
- Diz-Mellado, E.; López-Cabeza, V.P.; Rivera-Gómez, C.; Galán-Marín, C. Seasonal Analysis of Thermal Comfort in Mediterranean Social Courtyards: A Comparative Study. J. Build. Eng. 2023, 78, 107756. [Google Scholar] [CrossRef]
- Vogiatzakis, I.N.; Terkenli, T.S.; Trovato, M.G.; Abu-Jaber, N. Landscapes in the Eastern Mediterranean between the Future and the Past. Land 2018, 7, 160. [Google Scholar] [CrossRef]
- Zetterman, A. New Nordic Gardens: Scandinavian Landscape Design; Thames & Hudson: London, UK, 2017. [Google Scholar]
- Li, J.; Li, D.; Xiong, C.; Hoi, S. BLIP: Bootstrapping Language-Image Pre-Training for Unified Vision-Language Understanding and Generation. arXiv 2022, arXiv:2201.12086. [Google Scholar]
- Sohl-Dickstein, J.; Weiss, E.; Maheswaranathan, N.; Ganguli, S. Deep Unsupervised Learning Using Nonequilibrium Thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning, Lille, France, 6–11 July 2015. [Google Scholar]
- Goodfellow, I.J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Nets. In Advances in Neural Information Processing Systems 27; Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N., Weinberger, K.Q., Eds.; Curran Associates, Inc.: New York, NY, USA, 2014. [Google Scholar]
- Ruiz, N.; Li, Y.; Jampani, V.; Pritch, Y.; Rubinstein, M.; Aberman, K. DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 22500–22510. [Google Scholar]
- Gal, R.; Alaluf, Y.; Atzmon, Y.; Patashnik, O.; Bermano, A.H.; Chechik, G.; Cohen-Or, D. An Image Is Worth One Word: Personalizing Text-to-Image Generation Using Textual Inversion. arXiv 2022, arXiv:2208.01618. [Google Scholar]
- Ruiz, N.; Li, Y.; Jampani, V.; Wei, W.; Hou, T.; Pritch, Y.; Wadhwa, N.; Rubinstein, M.; Aberman, K. HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024; pp. 6527–6536. [Google Scholar]
- Kohya-ss Sd-Scripts. Available online: https://github.com/kohya-ss/sd-scripts (accessed on 16 November 2025).
- Shazeer, N.; Stern, M. Adafactor: Adaptive Learning Rates with Sublinear Memory Cost. In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018. [Google Scholar]
- Ahmed, N.; Natarajan, T.; Rao, K.R. Discrete Cosine Transform. IEEE Trans. Comput. 1974, 100, 90–93. [Google Scholar] [CrossRef]
- Duda, R.O.; Hart, P.E. Pattern Classification and Scene Analysis; Wiley: New York, NY, USA, 1973. [Google Scholar]
- Canny, J. A Computational Approach to Edge Detection. IEEE Trans. Pattern Anal. Mach. Intell. 1986, PAMI-8, 679–698. [Google Scholar] [CrossRef]
- Peli, E. Contrast in Complex Images. J. Opt. Soc. Am. A 1990, 7, 2032–2040. [Google Scholar] [CrossRef]
- Pertuz, S.; Puig, D.; Garcia, M.A. Analysis of Focus Measure Operators for Shape-from-Focus. Pattern Recognit. 2013, 46, 1415–1432. [Google Scholar] [CrossRef]
- Daniel, T.C. Measuring Landscape Esthetics: The Scenic Beauty Estimation Method; USDA Forest Service, Rocky Mountain Forest and Range Experiment Station: Fort Collins, CO, USA, 1976. [Google Scholar]
- Lothian, A. Landscape and the Philosophy of Aesthetics: Is Landscape Quality Inherent in the Landscape or in the Eye of the Beholder? Landsc. Urban Plan. 1999, 44, 177–198. [Google Scholar] [CrossRef]
- Jeong, S.; Choi, I.; Yun, Y.; Kim, J. Culture-TRIP: Culturally-Aware Text-to-Image Generation with Iterative Prompt Refinement. arXiv 2025, arXiv:2502.16902. [Google Scholar]
- Chen, R.; Luo, X.; Zhao, J. Research on the Adaptability of Generative Algorithm in Generative Landscape Design. Landsc. Archit. 2024, 31, 12–23. (In Chinese) [Google Scholar] [CrossRef]
- Dee, C. Form and Fabric in Landscape Architecture: A Visual Introduction; Spon Press: Abingdon, UK, 2001. [Google Scholar]
- Lu, W.; Luu, R.K.; Buehler, M.J. Fine-Tuning Large Language Models for Domain Adaptation: Exploration of Training Strategies, Scaling, Model Merging and Synergistic Capabilities. npj Comput. Mater. 2025, 11, 84. [Google Scholar] [CrossRef]
- Huang, R.; Lin, H.; Chen, C.; Zhang, K.; Zeng, W. PlantoGraphy: Incorporating Iterative Design Process into Generative Artificial Intelligence for Landscape Rendering. In CHI ’24: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems; ACM: New York, NY, USA, 2024; pp. 1–19. [Google Scholar]
- Koch, J.; Lucero, A.; Hegemann, L.; Oulasvirta, A. May AI? Design Ideation with Cooperative Contextual Bandits. In CHI ’19: Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems; Association for Computing Machinery: New York, NY, USA, 2019; pp. 1–12. [Google Scholar]
- Weisz, J.D.; Muller, M.; He, J.; Houde, S. Toward General Design Principles for Generative AI Applications. arXiv 2023, arXiv:2301.05578. [Google Scholar] [CrossRef]
- De Peuter, S.; Oulasvirta, A.; Kaski, S. Toward AI Assistants That Let Designers Design. AI Mag. 2023, 44, 85–96. [Google Scholar] [CrossRef]
- Zhuang, F.; Qi, Z.; Duan, K.; Xi, D.; Zhu, Y.; Zhu, H.; Xiong, H.; He, Q. A Comprehensive Survey on Transfer Learning. Proc. IEEE 2021, 109, 43–76. [Google Scholar] [CrossRef]
- Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. LAION-5B: An Open Large-Scale Dataset for Training next Generation Image-Text Models. In Advances in Neural Information Processing Systems 35; Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A., Eds.; Curran Associates, Inc.: New York, NY, USA, 2022; pp. 25278–25294. [Google Scholar]
- Sun, K.; Dredze, M. Amuro & Char: Analyzing the Relationship between Pre-Training and Fine-Tuning of Large Language Models. In Proceedings of the 10th Workshop on Representation Learning for NLP (RepL4NLP-2025); Association for Computational Linguistics: Albuquerque, NM, USA, 2025; pp. 131–151. [Google Scholar]
- Xinsir Xinsir/Controlnet-Canny-Sdxl-1.0. 2024. Available online: https://huggingface.co/xinsir/controlnet-canny-sdxl-1.0 (accessed on 25 May 2026).
- TheMistoAI/MistoLine. 2023. Available online: https://huggingface.co/TheMistoAI/MistoLine (accessed on 25 May 2026).
- Soria, X.; Li, Y.; Rouhani, M.; Sappa, A.D. Tiny and Efficient Model for the Edge Detection Generalization. In 2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW); IEEE: New York, NY, USA, 2023. [Google Scholar]














| Design Category | Chinese | Japanese | Mediterranean | Nordic | English |
|---|---|---|---|---|---|
| Architectural Facades | Upturned eaves, Whitewashed walls, Black tiles | Thatched roofs, Tiled roofs, Shoji windows, Bamboo walls | Terracotta roof tiles, Limestone walls, Arched doorways | Natural timber, Stone, Glass | Brick walls, Brick walls with climbing vines |
| Garden Structures | Pavilions, Corridors, Bridges, Moongates | Stone lanterns, Bamboo fences, Stone basins, Water basins | Terracotta planters, Wrought-iron furniture | Outdoor seating, Fire pits | Sculptures, Bird baths |
| Paving | Stone pavers, Brick, Ceramic tiles | Stone slabs, Gravel, Pebbles | Terracotta tiles, Colored ceramics, Natural stone | Stone, Natural wood | Regularly shaped stone paving |
| Rock & Water Features | Rockwork (artificial rockery), Reflection pools | Japanese rock garden (karesansui), Ponds, Streams | Swimming pools, Fountains | N/A | Fountains |
| Planting | Bamboo, Pine, Plum blossoms | Cherry blossoms, Maple, Moss, Ferns | Olive trees, Lemon trees, Grapevines, Pomegranates, Herbs | Pine, Cedar, Spruce | Roses, Boxwood, Lavender, Iris |
| Atmosphere | Serene, Harmonizing tradition | Restrained, Contemplative | Bright, Vibrant | Minimalist, Functional | Lush, Layered |
| Method | Training Time | Generation Quality | Computational Cost | Storage | Primary Use Case |
|---|---|---|---|---|---|
| LoRA | Short | Medium– High | High | Low | Balanced performance–storage trade-off |
| DreamBooth | Long | High | Medium | High | High-fidelity, resource-intensive tasks |
| HyperNetwork | Medium | Medium | Medium | Medium | Preserving original model characteristics |
| Textual Inversion | Shortest | Medium | High | Very Low | Quick adaptation, low-resource scenarios |
| Experiment Stage | Evaluation Metric | Metric Description | Evaluation Type | Scoring Range & Direction |
|---|---|---|---|---|
| Caption System Experiment | CLIP Score | Semantic alignment between generated images and text descriptions | Objective metrics | 0–1, higher is better |
| Spatial Rationale | Rationality of spatial layout, functional zoning, and element configuration | Subjective evaluation | 1–7, higher is better | |
| Training Resolution Experiment | Comprehensive Image Quality Score | Comprehensive assessment based on frequency domain quality, multi-scale band-limited contrast, and sharpness | Objective metrics | 0–5, higher is better |
| Scale Coherence | Accuracy of proportional relationships among design elements such as buildings, plants, and paving | Subjective evaluation | 1–7, higher is better |
| Experiment Stage | Evaluation Metric | Metric Description | Evaluation Type | Scoring Range & Direction |
|---|---|---|---|---|
| Multi -model comparison experiment | Design Rationale | Integrated assessment of Scale Coherence and Spatial Rationale, comprehensively evaluating spatial layout, proportional relationships, and element configuration | Subjective evaluation | 1–7, higher is better |
| Design Professionalism | Professional competence and technical maturity of the design | |||
| Design Accuracy | Completeness and precision of style characteristics and design elements | |||
| Design Satisfaction | Overall acceptability and practical application potential of the design|Subjective evaluation | |||
| Multi- model comparison (Public) | Scenic Beauty | Visual pleasantness, color harmony, and first-impression aesthetics | ||
| Recreational Appeal | Realism and appeal motivating visiting or recreational use | |||
| Style Recognition | Visual distinctiveness enabling identification of garden style |
| Stage | Variable Group | Evaluation Metric | Generation Method | Control Input | Sample Composition | Total |
|---|---|---|---|---|---|---|
| Stage 1 | A:Captioning Systems | CLIP Score | T2I | None | 5 Styles × 3 Captioning Systems × 20 Prompts | 300 |
| Spatial Rationale | 5 Styles × 3 Captioning Systems × 2 Prompts | 30 | ||||
| B: Training Resolutions | Image Quality | 5 Styles × 3 Training Resolutions × 20 Prompts | 300 | |||
| Scale Coherence | T2I + ControlNet (Softedge) | Test Map (w/Scale Ref.) | 5 Styles × 3 Training Resolutions × 2 Prompts | 30 | ||
| Stage 2 | Comparison | Overall Performance | Test Map | 5 Styles × 3 Models × 2 Prompts | 30 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Sun, X.; Chen, X.; Zhou, C.; Wu, S.; Zhao, H.; Li, K. AI for Garden Design Visualization: Development and Validation of the GardenDiff Model. Buildings 2026, 16, 2195. https://doi.org/10.3390/buildings16112195
Sun X, Chen X, Zhou C, Wu S, Zhao H, Li K. AI for Garden Design Visualization: Development and Validation of the GardenDiff Model. Buildings. 2026; 16(11):2195. https://doi.org/10.3390/buildings16112195
Chicago/Turabian StyleSun, Xiaolong, Xi Chen, Chao Zhou, Shengsha Wu, Hongbo Zhao, and Kun Li. 2026. "AI for Garden Design Visualization: Development and Validation of the GardenDiff Model" Buildings 16, no. 11: 2195. https://doi.org/10.3390/buildings16112195
APA StyleSun, X., Chen, X., Zhou, C., Wu, S., Zhao, H., & Li, K. (2026). AI for Garden Design Visualization: Development and Validation of the GardenDiff Model. Buildings, 16(11), 2195. https://doi.org/10.3390/buildings16112195






































