Preserving Formative Tendencies in AI Image Generation: Toward Architectural AI Typologies Through Iterative Blending
Abstract
1. Introduction
2. Background
2.1. Principles of Generative AI
2.2. Midjourney and Generative AI Tool
2.3. Recent Approaches
3. Materials and Methods
3.1. Research Methods and Process
- (i)
- Minimal Initial Data Usage: User-provided images are used only at the initial stage, which limits the method’s reliance on external datasets.
- (ii)
- Iterative Synthesis Process: Images in subsequent stages are blended by the outcomes of earlier iterations, without adding additional external data.
- (iii)
- Verification of Creativity and Tendency Preservation: The study examines how formative tendencies develop across iterative synthesis, with the analysis tools used as supplementary, exploratory indicators.
- (iv)
- Enhancement of User Agency: The user has full control over the selection of final outputs and the decision to continue iterations. Through this, AI is utilized not as a mere generator but as a dialogical tool that integrates decomposition, recomposition, and selection.
- (i)
- Whether creative variation is achievable with minimal initial data.
- (ii)
- Whether the design maintains its tendency without converging toward the labeled images in the training set during iterative synthesis.
- (iii)
- Whether users can lead the design process through selective control rather than being subjected to unilateral AI outputs.
3.2. Selection Criteria for AI-Generated Results
- (i)
- Layer: Layer refers to the effects that emerge through the overlapping of identical or heterogeneous objects. This includes not only the relative depth produced by protruding or overlapping forms but also the curvatures and intricate geometric relationships formed among nano fragmented or subdivided planes.
- (ii)
- Scale: Scale signifies the relative sense by which humans perceive the size of objects. As architecture increasingly relies on digital fabrication, the morphological effects resulting from the combination and assembly of diverse scales play a vital role in shaping spatial perception.
- (iii)
- Density: Density denotes the visual effect that arises from the complexity of distribution and arrangement of objects within a space. It manifests not only on the two-dimensional plane but also through its three-dimensional relationship with layers. Depending on how material and form are organized and controlled, even identical components can produce entirely different visual and spatial effects.
- (iv)
- Assembly: Assembly refers to how distinct parts come into contact and connect with one another. Grounded in the understanding that architecture presupposes human fabrication, this includes whether each part is composed at a scale manageable by human hands. Moreover, assembly anticipates the potential scale of the resulting space and emphasizes the significance of seams.
3.3. Structural (SSIM), Perceptual (LPIPS), and Semantic (CLIP) Similarities
- (i)
- Structural Similarity (SSIM) measures the luminance, contrast, and structural consistency between two images, x and y [25]. It is formally defined as:where and denote the mean luminance, and represent the variances corresponding to contrast, and indicates the covariance reflecting the structural relationship between images and . Higher SSIM values (approaching 1.0) indicate the preservation of the original structural order, proportional balance, and formal language of the image, whereas lower SSIM values imply a deconstruction of form and the exploration of new structural variations.
- (ii)
- Learned Perceptual Image Patch Similarity (LPIPS) metric is a distance-based perceptual similarity measure calibrated on the large-scale Berkeley-Adobe Perceptual Patch Similarity (BAPPS) dataset [26]. Unlike traditional pixel-level metrics such as PSNR and SSIM, LPIPS evaluates perceptual differences through pre-trained convolutional neural networks (CNN) (e.g., AlexNet, VGG, or SqueezeNet), with the VGG backbone employed in this study.A lower LPIPS score indicates greater perceptual similarity between image pairs, whereas higher values suggest distinct yet perceptually coherent variations. For consistency across metrics, all LPIPS scores were normalized to the range [0, 1] by converting the distance measure into a similarity value using .
- (iii)
- Contrastive Language–Image Pretraining (CLIP) was originally trained on large-scale image–text pairs, thereby constructing a semantic geometry in which visually different images sharing similar meanings are aligned near the same textual representations within the joint latent space [27]. Although CLIP was primarily designed to evaluate cross-modal (image–text) semantic correspondence, the cosine similarity between two image embeddings can naturally serve as a measure of intra-modal semantic proximity [28]. This property emerges as a by-product of its contrastive training objective, which aligns images around shared semantic anchors. Consequently, the cosine similarity within the CLIP latent space constitutes a valid proxy for semantic relatedness. This principle is analogous to CLIP’s zero-shot image retrieval mechanism; where images and texts are compared via cosine proximity; and in the present study, it is symmetrically extended to the comparison between image pairs.
4. Experimental Process
4.1. Preparation of Initial Input Data: Shin Takamatsu
4.2. Initial Input Data Blending Stage
4.3. Iterative Synthesis (Blending) Stage
4.4. User Intervention Stage
5. Analysis and Discussion
5.1. Preservation of Tendencies and Sustainability of Creative Variation
5.2. User Agency and Intervention
5.3. AI Tendencies and Typology
5.4. Research Limitations
6. Conclusions
- (i)
- It was observed that variation is achievable with a minimal set of initial images. During iterative synthesis, the AI deconstructed and recombined the formative tendencies of the input images to generate new forms. Although the features of the initial images were gradually transformed, recognizable formal language and consistent style were maintained, indicating exploratory potential of continuous transformation.
- (ii)
- The study observed that AI could preserve formative tendencies while avoiding convergence into generic forms. Images generated through iterative synthesis and blending did not degenerate into unrelated, common patterns; instead, they exhibited unique variations based on the formative characteristics of the initial input data. This suggests that AI can offer novel transformation possibilities based on user-provided image set. However, these observations should be interpreted cautiously, as the experiment relied on a single case and a proprietary model whose internal mechanisms cannot be independently validated.
- (iii)
- User participation navigates unexpected variation and intentional guidance with outcomes. Through active selection and curation, AI functioned not as a mere generator but as a dialogical design tool, granting the designer a more autonomous and participatory role in guiding the design process. Nonetheless, the degree of control remains partial due to opaque model behavior and the absence of fully reproducible parameters.
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Yoffie, D.B.; Von Bargen, S. Nvidia, Inc. in 2024 and the Future of AI; Harvard Business School Case 725-360; Harvard Business School Soldiers Field: Boston, MA, USA, 2024; Volume 1, pp. 1–28. [Google Scholar]
- Toews, R. The geopolitics of AI chips will define the future of AI. Horiz. J. Int. Relat. Sustain. Dev. 2023, 24, 126–138. [Google Scholar]
- Guo, D.; Yang, D.; Zhang, H.; Song, J.; Zhang, R.; Xu, R.; Zhu, Q.; Ma, S.; Wang, P.; Bi, X.; et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv 2025, arXiv:2501.12948. [Google Scholar]
- Miller, C. Chip War: The Fight for the World’s Most Critical Technology; Scribner, Simon and Schuster: New York, NY, USA, 2022. [Google Scholar]
- Floridi, L. Why the AI hype is another tech bubble. Philos. Technol. 2024, 37, 128. [Google Scholar] [CrossRef]
- Chatterji, A.; Cunningham, T.; Deming, D.; Hitzig, Z.; Ong, C.; Shan, C.Y.; Wadman, K. How people use chatgpt. Natl. Bur. Econ. Res. 2025. [Google Scholar] [CrossRef]
- Rahim, A. Catalytic Formations: Architecture and Digital Design; Taylor & Francis: New York, NY, USA, 2006; pp. 10–29. [Google Scholar]
- Samuelson, P. Generative AI meets copyright. Science 2023, 381, 158–161. [Google Scholar] [CrossRef] [PubMed]
- Lee, D.H.; Ko, S.H. Experiment and Evaluation of Architectural Image Generation through Artificial Intelligence-Based Text Image Generation Tool. KIEAE J. 2023, 23, 13–22. [Google Scholar] [CrossRef]
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef] [PubMed]
- Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F.L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. Gpt-4 technical report. arXiv 2023, arXiv:2303.08774. [Google Scholar] [CrossRef]
- Van Den Oord, A.; Kalchbrenner, N.; Kavukcuoglu, K. Pixel recurrent neural networks. In Proceedings of the International Conference on Machine Learning, New York, NY, USA, 20–22 June 2016; pp. 1747–1756. [Google Scholar]
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial networks. Commun. ACM 2020, 63, 139–144. [Google Scholar] [CrossRef]
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, 33, 6840–6851. [Google Scholar]
- Picon, A. Artificial Intelligence and the Future of Architecture Design Keynote Lecture, ACADIA 2022 Conference, University of Pennsylvania, Philadelphia. 29 October 2022; Unpublished manuscript. [Google Scholar]
- Picon, A. Artificial Intelligence and Architectural Intention. Technol. Archit. Des. 2025, 9, 6–9. [Google Scholar] [CrossRef]
- Salkowitz, R. Midjourney Founder David Holz on the Impact of AI on Art, Imagination and the Creative Economy. Forbes. 2022. Available online: https://www.forbes.com/sites/robsalkowitz/2022/09/16/midjourney-founder-david-holz-on-the-impact-of-ai-on-art-imagination-and-the-creative-economy/ (accessed on 28 December 2025).
- Radhakrishnan, A.M. Is Midjourney-AI a new anti-hero of architectural imagery and creativity. GSJ 2023, 11, 94–104. [Google Scholar]
- Leach, N. Architecture in the Age of Artificial Intelligence: An Introduction to AI for Architects; Bloomsbury Publishing: London, UK, 2025. [Google Scholar]
- del Campo, M. (Ed.) Artificial Intelligence in Architecture; John Wiley & Sons: London, UK, 2024. [Google Scholar]
- Albaghajati, Z.M.; Bettaieb, D.M.; Malek, R.B. Exploring text-to-image application in architectural design: Insights and implications. Archit. Struct. Constr. 2023, 3, 475–497. [Google Scholar] [CrossRef]
- Tan, L.; Luhrs, M. Using Generative AI Midjourney to enhance divergent and convergent thinking in an architect’s creative design process. Des. J. 2024, 27, 677–699. [Google Scholar] [CrossRef]
- Petráková, L.; Šimkovič, V. Architectural alchemy: Leveraging Artificial Intelligence for inspired design–a comprehensive study of creativity, control, and collaboration. Archit. Pap. Fac. Archit. Des. STU 2023, 28, 3–14. [Google Scholar] [CrossRef]
- Shin Takamatsu Architect and Associates; Nacasa & Partners Inc.; Katsuaki Furudate, M. Architectural Works and Photographic Materials. Used with Permission. Available online: https://takamatsu.co.jp/en/project/ (accessed on 28 December 2025).
- Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image quality assessment: From error visibility to structural similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef] [PubMed]
- Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 586–595. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021; pp. 8748–8763. [Google Scholar]
- Jeremy, K. Unlocking OpenAI CLIP. Part 2: Image Similarity. Medium. 2023. Available online: https://medium.com/@jeremy-k/unlocking-openai-clip-part-2-image-similarity-bf0224ab5bb0 (accessed on 28 December 2025).
- Szpakowska-Loranc, E. Architectural narration in Shin Takamatsu’s works. Tech. Trans. 2017, 2017, 39–52. [Google Scholar][Green Version]
- Soulard, L. Shin Takamatsu and Architecture as Symbolic Event. Domus. 2020. Available online: https://www.domusweb.it/en/architecture/gallery/2020/08/31/architecture-as-symbolic-event.html (accessed on 28 December 2025).





| Origin I | SYNTAX | Origin III | ||||||
|---|---|---|---|---|---|---|---|---|
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| ARK | Pharaoh | Kirin Plaza Osaka | Imanishi Motoakasaka | |||||
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| Earthtecture Sub-1 | Kunibiki Messe | Quasar | ||||||
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | |||
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | |||
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | |||
| 1st Iteration Set | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| 2nd Iteration Set | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| 3rd Iteration Set | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
| 4th Iteration Set | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() | ![]() |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Lee, D.-H.; Ko, S.-H. Preserving Formative Tendencies in AI Image Generation: Toward Architectural AI Typologies Through Iterative Blending. Buildings 2026, 16, 183. https://doi.org/10.3390/buildings16010183
Lee D-H, Ko S-H. Preserving Formative Tendencies in AI Image Generation: Toward Architectural AI Typologies Through Iterative Blending. Buildings. 2026; 16(1):183. https://doi.org/10.3390/buildings16010183
Chicago/Turabian StyleLee, Dong-Ho, and Sung-Hak Ko. 2026. "Preserving Formative Tendencies in AI Image Generation: Toward Architectural AI Typologies Through Iterative Blending" Buildings 16, no. 1: 183. https://doi.org/10.3390/buildings16010183
APA StyleLee, D.-H., & Ko, S.-H. (2026). Preserving Formative Tendencies in AI Image Generation: Toward Architectural AI Typologies Through Iterative Blending. Buildings, 16(1), 183. https://doi.org/10.3390/buildings16010183















































































































































































































































































































































































































































































































































































































































































































