Next Article in Journal
High-Precision Point Cloud Registration for Long-Span Bridges Based on Iterative Closest-Surface Method
Next Article in Special Issue
Day–Night and Weekday–Weekend Heterogeneity in Built Environment Impacts on Public Space Vitality: A GWRF Analysis in Yuexiu District
Previous Article in Journal
Mitigating Heat Stress for Pedestrians in Residential Neighborhoods: A Simulation-Based Approach to Enhance Outdoor Thermal Comfort
Previous Article in Special Issue
Research Progress and Frontier Trends in Generative AI in Architectural Design
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Modular AI Workflow for Architectural Facade Style Transfer: A Deep-Style Synergy Approach Based on ComfyUI and Flux Models

College of Architecture and Urban Planning, Qingdao University of Technology, Qingdao 266033, China
*
Author to whom correspondence should be addressed.
Buildings 2026, 16(3), 494; https://doi.org/10.3390/buildings16030494
Submission received: 27 November 2025 / Revised: 16 January 2026 / Accepted: 23 January 2026 / Published: 25 January 2026

Abstract

This study focuses on the transfer of architectural facade styles. Using the node-based visual deep learning platform ComfyUI, the system integrates the Flux Redux and Flux Depth models to establish a modular workflow. This workflow achieved style transfer of building facades guided by deep perception, encompassing key stages such as style feature extraction, depth information extraction, positive prompt input, and style image generation. The core innovation of this study lies in two aspects: Methodologically, a modular low-code visual workflow has been established. Through the coordinated operation of different modules, it ensures the visual stability of architectural forms during style conversion. In response to the novel challenges posed by generative AI in altering architectural forms, the evaluation framework innovatively introduces a “semantic inheritance degree” assessment system. This elevates the evaluation perspective beyond traditional “geometric similarity” to a new level of “semantic and imagery inheritance.” It should be clarified that the framework proposed by this research primarily provides innovative tools for architectural education, early design exploration, and visualization analysis. This workflow introduces an efficient “style-space” cognitive and generative tool for teaching architectural design. Students can use this tool to rapidly conduct comparative experiments to generate multiple stylistic facades, intuitively grasping the intrinsic relationships among different styles and architectural volumes/spatial structures. This approach encourages bold formal exploration and deepens understanding of architectural formal language.

1. Introduction

Architectural facade design embodies cultural, temporal, and spatial aesthetic characteristics. In recent years, artificial intelligence has achieved significant breakthroughs through data-driven approaches [1]. With the emergence and mature application of concepts and technologies such as big data and cloud computing, AI has sparked a profound technological revolution worldwide [2]. In the field of architecture, artificial intelligence is transforming buildings and urban environments, making the design process faster, more efficient, and more sustainable [3].
With the advancement of AI and deep learning, style transfer has emerged as a valuable exploratory tool in architectural design. However, existing methods often rely heavily on programming techniques, pose significant operational barriers, and lack user-friendly design workflows. ComfyUI, an emerging visual AI workflow tool, uses a node-based workflow in which each step serves as a node. Upon completing a task at a node, its output feeds into the subsequent step, thereby executing the entire process [4]. As a node-based visual deep learning platform, it provides designers with a low-code prototyping approach.
In this study on architectural facade style transfer using the ComfyUI platform, the combined application of the Flux Redux and Flux Depth models forms the core technical foundation of the workflow. Flux Redux serves as a powerful foundational generative model, excelling in high-fidelity interpretation and fusion of complex style prompts. It generates images with exceptional consistency and rich detail, ensuring accurate and deep stylistic expression. Flux Depth distinguishes itself by introducing a depth-aware dimension, enabling precise extraction and reconstruction of three-dimensional spatial structural information from input architectural model images. During style transfer, the Flux Depth model constrains the generation process to ensure that transferred stylized textures strictly adhere to the original facade’s geometric and compositional logic. This effectively preserves the architectural entity’s structural authenticity and maintains the stability of spatial order. However, it must be clarified that the structural authenticity and spatial order stability discussed here are both defined at the level of visual representation. They are explicitly manifested in the architectural forms depicted in images taken before and after relocation, rather than in the building’s structural integrity. Furthermore, in subsequent articles, all architectural terminology—such as “architecture,” “structure,” and “space”—refers exclusively to their visual representations in two-dimensional images. The scope of this research is strictly confined to visual perception and image generation, and does not address the physical properties of architectural entities, structural engineering, or actual construction.

2. Theoretical Background

2.1. The Development of AI Image Style Transfer Technology

Style transfer technology began developing after Gatys et al. proposed a convolutional neural network-based algorithm. This approach achieves high-quality artistic style transfer by separating and recombining the content and style representations of an image. However, its iterative optimization process is relatively slow [5]. Subsequently, Johnson et al. introduced perceptual loss functions, enabling feedforward networks to perform real-time style transfer and significantly improving efficiency [6]. Subsequent research by Cai et al. demonstrated that multidimensional neural networks can analyze complex data and uncover nonlinear relationships [7], which forms the very foundation of AI image style transfer technology.
Existing style transfer techniques primarily fall into two categories. The first involves optimization-based methods that iteratively refine input images to match the target style’s feature representation, though these are computationally intensive. The second employs feedforward network-based approaches, utilizing pre-trained loss networks to define perceptual loss for rapid style transfer. For instance, Huang et al.’s Adaptive Instance Normalization (AdaIN) layer achieves real-time transfer of arbitrary styles by aligning feature statistics [8].

2.2. Application of AI Image Style Transfer Technology in the Architectural Field

Style transfer is currently applied in artistic creation, image enhancement, and architectural design. Within architectural design, its applications exhibit considerable diversity. For instance, the multidisciplinary design optimization algorithm (MDOA) and model-driven decision support system (MDSS) proposed by Cai et al. for optimizing architectural spatial performance demonstrate that algorithmic models have become practical tools for solving complex architectural problems such as spatial form optimization and aesthetic style transfer [9]. Lin et al. applied conditional generative adversarial networks (CGAN) in their research. By interpreting image data of historical district facades, they developed a method capable of autonomously designing specific facade ornamentation styles [10]. Meng et al. designed a replicable five-stage AIGC workflow to transfer the artistic style of Song Dynasty boundary painting into interior design, advancing the effective integration of traditional aesthetics with modern design concepts [11]. Chen et al. developed a technical pipeline deeply integrating Stable Diffusion (for texture style transfer) with NeRF (for 3D reconstruction), enabling the complete creation process from 2D stylized images to 3D high-fidelity models, demonstrating AI visual technology’s deep application in architecture—transitioning from expression to construction [12]. Yan et al. employed a pix2pix algorithm based on conditional generative adversarial networks (CGAN) for training. This model demonstrated exceptional accuracy and stability in architectural image style transfer tasks, enabling precise historical style transfer while preserving structural features [13]. Chen et al. proposed a typology-guided fusion framework integrating LoRA and diffusion models. Its core innovation lies in systematically embedding architectural typology theory into generative AI training and inference workflows, transcending traditional visual mimicry to achieve “typological transcoding” of complex historical architectural styles [14]. Concurrently, Duan et al. adapted the Stable Diffusion XL model for domain-specific tasks using LoRA, enabling it to generate high-fidelity facade images consistent with historical architectural contexts. This demonstrated the optimization role of LoRA fine-tuning in architectural style image generation [15].

2.3. Visual AI Platform

Research on existing style transfer techniques reveals that current methods for architectural facade style transfer suffer from complex operations, high computational costs, and a lack of systematic attention to preserving the building’s core form before and after style transfer. Based on this, this study explores solutions using the ComfyUI platform. ComfyUI significantly lowers the technical barrier for users through its intuitive visualization and modular interface. Its core strength lies in its node-based visual programming interface, which provides users with unprecedented fine-grained control and workflow reproducibility [16]. Simultaneously, addressing the issue of altered architectural form caused by generative AI during style transfer, this study proposes a semantic inheritance evaluation method to quantitatively assess the similarity of architectural form before and after style transfer.
The two AI models integrated with ComfyUI in this study significantly streamline the AI image generation workflow. The Flux Depth model employs algorithms to convert complex RGB images into single-channel depth maps. While reducing data dimensions and optimizing data structures for AI recognition, it effectively incorporates spatial information to process depth hierarchies within images. This feature is particularly crucial for architecture-related tasks, enabling architects to represent spatial structures and three-dimensional effects in architectural designs with greater accuracy.
Building on this foundation, the Flux Redux model performs feature extraction and mapping on depth maps and style reference images, based on textual requirements for the output image. Adversarial training optimizes parameters to precisely transfer visual elements, such as texture and color, from the style reference image to the depth map. This preserves the architectural model’s content while imparting artistic style, producing high-quality style transfer renderings that meet diverse architectural design needs.

3. Design Process

This section outlines the overall workflow structure and describes the role of each module.

3.1. System Architecture Design

The AI image generation system workflow developed in this study integrates four modules: style feature extraction, depth information extraction, forward prompt input, and image generation (Figure 1).
Within this workflow, the system first uses the Flux Redux model to extract stylistic image feature vectors, enabling efficient capture of stylistic features. Concurrently, it integrates user-provided textual descriptions to guide subsequent semantic-level generation. Subsequently, the Flux Depth model analyzes the provided model images to interpret three-dimensional structure, generating depth maps to impose spatial geometric constraints. Based on this, multimodal features are encoded into a latent space for fusion, reconstructed into images via a decoder, and optimized through adversarial training to achieve high-quality image generation that balances stylistic consistency, semantic accuracy, and spatial plausibility.

3.2. Workflow Setup and Module Functions

3.2.1. Construction of the Style Feature Extraction Module Group

The construction of this module group first processes image files through the image loading and preprocessing module, converting them into a standardized tensor format. Additionally, this module generates image-related masks and ensures consistency in image format and dimensions.
Subsequently, the CLIP visual encoding module extracts deep semantic features from the images. Its model is dynamically loaded via a dedicated CLIP visual loader. The core innovation of this module group lies in introducing a style model application module. This module uses CLIP visual features as conditional input and, driven by pre-trained models (e.g., flux-redux-dev) loaded by the style model loader, performs stylistic adjustments and feature enhancement on the initial conditional data.
This module group systematically addresses the transformation and extraction of specific stylistic features from raw images (Figure 2).

3.2.2. Construction of the Depth Information Extraction Module Group

This module group adopts a serial architecture, sequentially processing images through loading, directional scaling, depth estimation, and visualization modules to achieve end-to-end processing from raw images to depth maps.
Specifically, the workflow sequentially traverses four specialized modules: The image loading module ensures uniformity of data input; The Image Oriented Scaling module performs geometric normalization of input images through configurable scaling standards and cropping strategies, providing downstream models with uniformly sized inputs. This significantly enhances the stability of subsequent processing and the comparability of results. The core DepthAnythingV2 depth preprocessor module generates high-precision, dense depth maps from preprocessed images. Its performance advantages directly benefit from the standardized data provided by preceding stages.
Additionally, to facilitate process monitoring and result validation, a depth map preview module is integrated at the end of the module suite. This optional component provides instant visual feedback and does not affect the generation of core depth data or subsequent style transfer tasks (Figure 3).

3.2.3. Construction of the Prompt Input Module Group

This module group constitutes a dual-condition fusion prompt input system. In this design, the textual input “Generate a new architectural facade form based on the provided picture style” is converted into a semantic feature vector via the CLIP text encoder. Simultaneously, it receives visual style conditions from the style-feature-extraction module group. The core innovation lies in the Flux Guidance Module. Through its unique dual-input design, it fuses textual semantics and visual style features into unified guidance parameters, achieving dual-condition fusion of textual meaning and visual style.
Ultimately, this fused guidance signal is fed into the generative model to achieve controllable shaping of the generated content. This module group design systematizes the conditional encoding, fusion, and guidance process through clear modular interfaces, significantly enhancing the controllability and flexibility of the generation process (Figure 4).

3.2.4. Construction of the Image Generation Module Group

This module group constructs an end-to-end image generation system based on multi-condition guidance. At its core is the InstructPixToPix conditional module, which innovatively integrates conditional inputs from three independent module groups: favorable conditions from the style model, adverse conditions from the CLIP text encoder, and image inputs from the deep preprocessor module. These multimodal conditions are mapped to a unified latent space via the VAE encoder, forming a collaborative guidance signal.
The system adopts a modular loading architecture that dynamically manages VAE and UNet model resources via dedicated loaders. During generation, the K-sampler performs a controllable denoising process in the latent space using the UNet model and the integrated conditional signal, producing high-quality latent representations. Finally, the VAE decoder converts these latent representations back to the pixel space, with the preserved image module outputting the final result (Figure 5).
This completes the construction of the overall model (Figure 6). Additionally, to address inconsistent results caused by data-quality uncertainties in AI learning, we adopted the following solutions.
First, regarding AI model selection, this research workflow does not train models from scratch but instead integrates top-tier open-source foundational models, such as Flux Redux and Flux Depth. These models have already undergone pre-training on massive, high-quality, rigorously cleaned, and curated internet image-text pair datasets.
Second, in workflow design, we innovatively implemented multi-module collaboration. The deep information extraction module allows users to directly input architectural model images, reflecting the input image’s inherent 3D geometric structure. This significantly reduces the uncertainty inherent in AI-generated images.

4. Experiments and Results

4.1. Experimental Setup

The data used for testing primarily consists of architectural images. All architectural photographs originate from publicly available urban architectural image databases. These databases encompass real-world architectural scenes from diverse cities and styles, effectively reflecting visual characteristics under authentic conditions to ensure the test data’s authenticity and representativeness.
For variable configuration, image complexity was hierarchically categorized through comprehensive evaluation. This involved calculating the information entropy (H), edge density (E), and color complexity (C) for each image. A composite metric (S) was then computed using the formula S = 0.4H + 0.3E + 0.3C. Based on these calculations, complexity levels were divided into three tiers.
The weight distribution S = 0.4H + 0.3E + 0.3C is based on the following logic: In image processing, information entropy (H) is widely regarded as the core metric for measuring the overall information content and uncertainty of an image. It carries the most significant weight in emphasizing that the total information contained in the original image is the primary factor influencing processing difficulty. Second, edge density (E) directly reflects the structural complexity of an image, while color complexity (C) relates to the richness and variation of hues. In architectural facades, the complexity of structure (e.g., decorative lines, window divisions) and the diversity of materials (color and texture) collectively contribute to the visual perception of “complexity.” Assigning equal weight to both dimensions reflects a balanced emphasis on the “form” and “material” aspects of architectural facades.
It is imperative to note that the “complexity” measured in this study specifically refers to the visual computational complexity of architectural facade images. This metric primarily quantifies the technical challenges faced by style transfer models when processing input images with varying information density and levels of detail.
The complexity definition employed in this experiment does not aim to replace the architectural ontology complexity derived from multidimensional constraints, such as function, structure, environment, and context, within architectural theory. This study focuses on the performance of AI models in visual generation tasks. Therefore, using image complexity as a control variable for grouped experiments is technically sound and practical.

4.2. Experimental Results

To ensure data authenticity, all 24 selected architectural images were uniformly cropped to a consistent size before complexity calculations. The final experimental data ranged from 3.6 to 15.23. Based on these results, eight images with values between 9.96 and 15.23 were classified into the high-complexity group, eight images between 8.12 and 9.91 into the medium-complexity group, and eight images between 3.6 and 8.08 into the low-complexity group (Table 1). Style transfer was performed on each group, and similarity measurements were conducted between the transferred images and their originals.
Subsequent image similarity calculations reveal that, unlike conventional architectural facade style transfer, this experiment focuses on transferring buildings with altered core structures, forms, and layouts. Consequently, traditional similarity metrics, such as structural similarity, become inadequate. This necessitates a new evaluation framework—one that shifts from assessing “pixel or structural similarity” to evaluating “the inheritance of architectural semantics and stylistic genes.” The evaluation framework elevates from the “geometric level” to the “semantic and imagery level.” It assesses explicitly four aspects: shape contextual similarity (S1), key-point matching accuracy (K), texture feature similarity (T), and color distribution similarity (C). The overall similarity score is calculated as: Comprehensive Score(S2) = 0.35S1 + 0.25K + 0.25T + 0.15C.
The weighting distribution for the composite score (S2) = 0.35S1 + 0.25K + 0.25T + 0.15C is primarily based on the following logic: Shape Contextual Similarity (S1) is the core of this evaluation system. As demonstrated earlier, generative AI style transfer may alter a building’s fundamental form. Yet, a structure’s “essence” and “recognizability” are most fundamentally embodied in its basic shape, silhouette, and volumetric relationships. Therefore, Shape Contextual Similarity is assigned the highest weight. Key point matching (K) and texture feature similarity (T) evaluate two important yet slightly secondary dimensions—structural stability and surface similarity, respectively. Thus, they are assigned equal weight. Color Distribution Similarity (C) is intentionally assigned a lower weight in the evaluation system. Compared to shape and structure, color alterations have a relatively minor impact on the “identity” of the building itself, and color changes tend to be minimal before and after style migration.
Additionally, due to AI processing uncertainties, transferred architectural images may exhibit excessive background rendering. To ensure consistency in subsequent measurements, transferred images undergo standardized preprocessing to eliminate background interference, providing only the main building structure remains (Table 2). Following the preparatory steps, data were measured, and statistical analysis was conducted (Table 3). Comparing results with the prior complexity grouping, the high-complexity group achieved an average comprehensive semantic inheritance score of 78.7125, while the medium-complexity group scored 78.65. The slight difference (approximately 0.06) indicates relatively similar semantic inheritance after style transfer in medium- and high-complexity buildings. However, the low-complexity group achieved an average comprehensive semantic inheritance score of 82, significantly higher than those of the high- and medium-complexity groups. This suggests that style transfer in low-complexity buildings better preserves semantic and stylistic characteristics, resulting in higher inheritance rates (Table 4).
Based on the findings of this experiment, low-complexity buildings feature simpler structures, fewer details, and more uniform textures. This enables AI models to more readily capture core semantic features (such as basic shapes and color distributions), thereby reducing information loss during transfer. Conversely, high- and medium-complexity buildings contain more intricate details, complex edges, and color variations, increasing the model’s learning difficulty. This may result in incomplete preservation of key semantic information during transfer, particularly when structural alterations are significant. The model may overemphasize local features while neglecting overall style inheritance.
Moreover, the evaluation system in this experiment emphasizes semantic levels. The diversity of high-complexity buildings may make key-point matching and shape context similarity calculations more susceptible to noise interference (e.g., residual background rendering artifacts). In contrast, the uniformity of low-complexity buildings facilitates stable metric computation. Additionally, regarding the two evaluation framework models mentioned above, the weighting data represent reasonable and interpretable initial settings derived from theoretical analysis and preliminary experiments. However, their “optimality” is not absolute. For instance, the keypoint matching score (K) consistently measures 50 across all datasets, indicating that this metric may not fully capture variations in complexity and requires further optimization.

4.3. Ablation Studies

To systematically evaluate the role of each module in the style transfer task, this study conducted targeted controlled experiments, primarily focusing on the deep information extraction module and the positive prompt input module. Specifically, the core function of the deep information extraction module is to parse the geometric layout and spatial structure of the scene from the source image, providing crucial structural constraints for subsequent generation. The core function of the positive prompt input module is to assist in generating the transferred image using textual descriptions.
The experiments will focus on observing and comparing performance differences in multiple aspects of the generated images. Through this controlled experimental design—where all other generation module parameters and conditions remain constant while manipulating only the depth information and forward prompt input modules—we can effectively analyze the relationship between variations in image generation outcomes and the absence of depth information. Alternatively, we can examine the control capabilities of textual guidance over the semantic content and stylistic details of generated images (Table 5), thereby dissecting the specific roles of these two modules in maintaining the structural integrity of generated images.
Experimental results demonstrate that the forward prompt input module is an indispensable component of the Flux Redux model’s generation workflow. Disabling this module results in the loss of conditional information relied upon by the model, ultimately causing the generation process to fail.
Turning off the depth information extraction module entirely causes the workflow’s core functionality to revert to the text-to-image generation mechanism based on the Flux Redux model. Under this configuration, the final generated images become highly dependent on two key input elements. One is the style reference image, whose color patterns, texture features, and compositional elements are analyzed by the algorithm and transferred into the generated image. The other element is the forward prompt. Using semantic encoding techniques, concepts, scenes, attributes, and other textual information are converted into feature vectors processable by the image generation model. This unidirectional relationship implies that when deep information is absent, the spatial structure, perspective relationships, and three-dimensional layering of the generated image are not incorporated into the model’s computational process. Consequently, an image generation paradigm dominated by two-dimensional feature transfer and semantic mapping emerges.

4.4. Experimental Analysis and Discussion

Experimental results clearly demonstrate a negative correlation between the initial complexity of architectural images and the semantic inheritance rate after style transfer. The low-complexity group achieved the highest average inheritance score (82), while the medium- and high-complexity groups scored similarly and significantly lower (78.65 and 78.7125, respectively). This finding indicates that a structurally simple, less detailed architecture enables AI models to capture better and preserve their core stylistic essence. Furthermore, ablation experiments preliminarily validate the critical role of the deep information extraction module in constraining the three-dimensional spatial structure of generated images, as well as the indispensable conditional guidance provided by the forward prompt input module within the current workflow.

4.4.1. Limitations of the Study

Although this study has achieved specific results, several limitations remain:
First, the experiment used only 24 architectural images, which may limit the generalizability of the findings. Furthermore, variables such as building type, regional style, and size were not systematically considered. Future experiments should expand the sample database to cover a broader range of architectural types and styles, thereby enhancing the robustness of the conclusions.
Second, the accuracy of the evaluation system requires further refinement. The assessment framework developed in this study remains preliminary, with specific metrics failing to demonstrate sufficient discriminative power in the experimental data. This indicates their limited effectiveness in the evaluation process, necessitating optimization and further research.
Third, while the experiment observed a negative correlation between complexity and preservation, explanations for the underlying mechanisms remain speculative. For instance, the assertion that “high complexity hinders model learning” lacks a direct analysis of the model’s internal feature representations. Techniques such as feature visualization could be considered to more deeply reveal differences in attention distribution when AI models process architectural images of varying complexity.

4.4.2. Scope and Boundary Conditions of the Study

Experiments revealed that to obtain optimal experimental data, this study faces certain limitations in its applicability and requires specific conditions to be met.
For instance, in terms of applicability, low-complexity architecture, such as minimalist modern building facades, yields the best style transfer results, as their uniform structure facilitates semantic preservation.
Regarding boundary conditions, input images must possess precise geometric contours; blurred or heavily occluded architectural photos may result in failed style extraction. Furthermore, the workflow involves serialized multi-module operations, and ablation experiments indicate high inter-module dependencies—failures in individual modules can cause the entire workflow to collapse.

4.4.3. Applications of Research in Architecture

The research currently presented primarily focuses on visual generation techniques for architectural images, yet it also demonstrates practical applications in architectural design, historic buildings, and architectural education.
Gao et al. highlighted in their study of Jiangmen’s overseas Chinese heritage architecture that digital processing and intelligent analysis of architectural imagery are crucial for cultural heritage preservation and restoration [17]. The AI image style transfer technology examined in this research leverages style learning and transfer to provide robust generative capabilities for virtual restoration of damaged historical structures, thereby expanding the application boundaries of AI technology in architectural heritage.
In prior research, Jin et al. developed an AI-embedded teaching model for a 9-week graduate course [18]. This study clarified AI’s core role in architectural education—as a creative generation tool—providing educational value support for our current technical application. Within architectural design education, this research serves as a cognitive tool for generating spatial style. It enables students to create and compare facades across multiple styles, helping them intuitively grasp the relationships among diverse styles and building volumes. This approach both encourages bold formal exploration and deepens understanding of architectural essence.
Regarding the design process, research by Bagasi et al. explicitly positions AI image generation tools as core technologies for early design stages, directly demonstrating the practicality and efficiency gains of AI image style transfer techniques during architectural creative generation [19]. The workflow generated by this research can serve as a “creative accelerator” for early conceptual development. Designers can input images of volumetric models and specify one or more style reference images to rapidly generate architectural facade possibilities that simultaneously meet defined volumetric and stylistic objectives. Crucially, the workflow’s deep information extraction module parses fundamental spatial hierarchies and front-back relationships from simple block model images. This ensures generated outputs transcend flat texture mapping, instead attempting three-dimensional style mapping that provides a more rational visual foundation for subsequent model refinement.
Furthermore, as Cai et al. prospectively noted in their research, integrating artificial intelligence with 3D printing [20] enables the digital models of architecturally designed solutions—generated through style transfer and embodying both aesthetic innovation and performance optimization—to interface directly with 3D printing construction systems. This facilitates the precise and efficient materialization of virtual styles into physical architecture.

5. Conclusions

This study successfully establishes a new paradigm for evaluating semantic inheritance in architectural style transfer and preliminarily validates that image complexity is a key factor influencing style transfer outcomes. Although improvements are possible in sample size, metric sensitivity, and theoretical depth, the research framework, problem awareness, and preliminary conclusions provide valuable insights for interdisciplinary research at the intersection of architectural design and digital technology. It is important to emphasize that the core contribution of this study lies in the innovation of visualization methods. Its structural relevance is applicable only to specific scenarios such as early-stage conceptual design, educational instruction, or visualization analysis, and is not intended to replace actual structural engineering practices.
Future research will advance in the following directions: First, expand the experimental sample size and systematically introduce building types, architectural styles, and other factors as control variables to conduct multi-factor variance analysis, ensuring that the research findings apply to a broader range of architectural contexts. Second, optimize the evaluation metric system, with the primary task being to refine the “keypoint matching accuracy” metric to enhance the scientific rigor of this evaluation framework. Third, conduct a combined qualitative and quantitative analysis of the image results generated by ablation experiments to assess the role of each module within the workflow rigorously. Fourth, based on the finding that “low complexity facilitates inheritance,” explore applying semantic abstraction or simplification preprocessing to high-complexity architectural images before style transfer. Observe whether this enhances semantic inheritance after style transfer, potentially providing a practical technical pathway for digital style innovation in complex historical and landmark buildings.

Author Contributions

Conceptualization, C.Q.; Methodology, C.Q.; Software, C.Q.; Validation, C.Q.; Formal analysis, C.Q.; Investigation, C.Q.; Data curation, C.Q.; Writing—original draft, C.Q.; Writing—review & editing, C.X.; Visualization, C.Q.; Supervision, C.X. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Gu, S.J.; Wang, R.X.; Wu, Y.F.; Xu, X.Y.; Yan, C.; Gao, T.Y.; Yuan, F. Exploration of AI-Heuristic Architectural Generative Design Process Based on FUGenerator Platform. In Proceedings of the 2023 National Conference on Digital Technology in Architectural Education and Research; Tongji University: Shanghai, China, 2023; pp. 441–444. [Google Scholar]
  2. Huang, X.R.; Wang, Y.D.; White, M.; Zhang, B. Logic and Black Box:Prospecting the Integration of Artificial Intelligence and Computer-Aided Technology in Future Architectureand Urban Design. Urban Archit. 2022, 19, 1–6+18. [Google Scholar]
  3. Komatina, D.; Miletić, M.; Mosurović Ružičić, M. Embracing Artificial Intelligence (AI) in Architectural Education: A Step towards Sustainable Practice? Buildings 2024, 14, 2578. [Google Scholar] [CrossRef]
  4. Zhang, J. An Example of Creating a Dragon Year Promotion ID by Invoking the SVD Animation Model via ComfyUI. Video Prod. 2024, 30, 68–70. [Google Scholar]
  5. Singh, A.; Jaiswal, V.; Joshi, G.; Sanjeeve, A.; Gite, S.; Kotecha, K. Neural style transfer: A critical review. IEEE Access 2021, 9, 131583–131613. [Google Scholar] [CrossRef]
  6. Johnson, J.; Alahi, A.; Fei-Fei, L. Perceptual losses for real-time style transfer and super-resolution. In European Conference on Computer Vision; Springer International Publishing: Cham, Switzerland, 2016; pp. 694–711. [Google Scholar]
  7. Cai, G.; Lou, Y.; Lu, F. AI enhancing prefabricated aesthetics and low carbon coupled with 3D printing in chain hotel buildings from multidimensional neural networks. Sci. Rep. 2025, 15, 13229. [Google Scholar] [CrossRef] [PubMed]
  8. Dumoulin, V.; Shlens, J.; Kudlur, M. A learned representation for artistic style. arXiv 2016, arXiv:1610.07629. [Google Scholar]
  9. Cai, G.; Sun, L.; Liu, D.; Xu, B.; Mo, Z. Potential of indoor room 3D ratio in reducing carbon emissions by prefabricated decoration in chain hotel buildings via multidimensional algorithm models for robot in-situ 3D printing. J. Build. Eng. 2025, 101, 111757. [Google Scholar] [CrossRef]
  10. Lin, H.; Huang, L.; Chen, Y.; Zheng, L.; Huang, M.; Chen, Y. Research on the Application of CGAN in the Design of Historic Building Facades in Urban Renewal—Taking Fujian Putian Historic Districts as an Example. Buildings 2023, 13, 1478. [Google Scholar] [CrossRef]
  11. Meng, J.; Fang, X.; Xu, J.; Zhang, Z. Research on the Innovative Application of Song Dynasty Boundary Painting in Interior Soft Decoration Design Based on AIGC. Buildings 2025, 15, 1067. [Google Scholar] [CrossRef]
  12. Chen, G.; Tong, Y.; Wu, Y.; Wu, Y.; Liu, Z.; Huang, J. Reconstruction of Cultural Heritage in Virtual Space Following Disasters. Buildings 2025, 15, 2040. [Google Scholar] [CrossRef]
  13. Yan, W.; Wang, T.; Zhang, C. Renewal Design of Architectural Facade Features in the Shantou Xiaogongyuan Historic District Based on Deep Learning. Buildings 2025, 15, 4404. [Google Scholar] [CrossRef]
  14. Chen, Z.; Zhang, N.; Xu, C.; Xu, Z.; Han, S.; Jiang, L. Typological Transcoding Through LoRA and Diffusion Models: A Methodological Framework for Stylistic Emulation of Eclectic Facades in Krakow. Buildings 2025, 15, 2292. [Google Scholar] [CrossRef]
  15. Duan, W.; Rao, J.; Zhao, J.; Tao, N.; Chen, J. AI-Based Pre-Renewal Design for Historic Building Facades: An AIGC–LoRA Framework with Collaborative Assessment. Buildings 2025, 15, 4212. [Google Scholar] [CrossRef]
  16. Wang, J.; Shi, Y.; Chen, X.; Lan, Y.; Liu, S. Teaching with Artificial Intelligence in Architecture: Embedding Technical Skills and Ethical Reflection in a Core Design Studio. Buildings 2025, 15, 3069. [Google Scholar] [CrossRef]
  17. Gao, L.; Wu, Y.; Yang, T.; Zhang, X.; Zeng, Z.; Chan, C.K.D.; Chen, W. Research on Image Classification and Retrieval Using Deep Learning with Attention Mechanism on Diaspora Chinese Architectural Heritage in Jiangmen, China. Buildings 2023, 13, 275. [Google Scholar] [CrossRef]
  18. Jin, S.; Tu, H.; Li, J.; Fang, Y.; Qu, Z.; Xu, F.; Liu, K.; Lin, Y. Enhancing Architectural Education through Artificial Intelligence: A Case Study of an AI-Assisted Architectural Programming and Design Course. Buildings 2024, 14, 1613. [Google Scholar] [CrossRef]
  19. Bagasi, O.; Nawari, N.O.; Alsaffar, A. BIM and AI in Early Design Stage: Advancing Architect–Client Communication. Buildings 2025, 15, 1977. [Google Scholar] [CrossRef]
  20. Cai, G.; Liu, D.; Wu, Z. Multidimensional algorithms-based carbon efficiency model of building geometric 3D ratios for prefabricated 3D printing design and construction. npj Clean Energy 2025, 1, 5. [Google Scholar] [CrossRef]
Figure 1. Schematic diagram of module composition. (Source: made by authors.)
Figure 1. Schematic diagram of module composition. (Source: made by authors.)
Buildings 16 00494 g001
Figure 2. Hierarchical data flow and module dependencies within the style feature extraction module. (Source: made by authors).
Figure 2. Hierarchical data flow and module dependencies within the style feature extraction module. (Source: made by authors).
Buildings 16 00494 g002
Figure 3. Serialized processing flow of the depth information extraction module. (Source: made by authors.)
Figure 3. Serialized processing flow of the depth information extraction module. (Source: made by authors.)
Buildings 16 00494 g003
Figure 4. System flowchart of the forward prompt input module. (Source: made by authors).
Figure 4. System flowchart of the forward prompt input module. (Source: made by authors).
Buildings 16 00494 g004
Figure 5. Image generation system flow. (Source: made by authors.)
Figure 5. Image generation system flow. (Source: made by authors.)
Buildings 16 00494 g005
Figure 6. Overall workflow. (Source: made by authors.)
Figure 6. Overall workflow. (Source: made by authors.)
Buildings 16 00494 g006
Table 1. Complexity Metrics for Different Image Datasets.
Table 1. Complexity Metrics for Different Image Datasets.
IDInformation Entropy Complexity (H *) Edge Density (E *)Color Complexity (C *)Comprehensive Score (S *)
17.160.4180.2458.96
27.520.2910.2619.08
37.240.3680.2077.68
47.210.3920.2799.91
57.530.310.2619.14
67.490.4040.2799.96
77.490.3420.28710.01
87.470.4430.2148.12
247.320.3920.28610.12
* H = −∑(p(i) × log2(p(i))). * E = (Number of edge pixels)/(Total number of pixels). * C = (Number of Unique Colors)/(Total Number of Pixels). * S = 40%H + 30%E + 30%C.
Table 2. Comparison of Preprocessed and Unprocessed Images.
Table 2. Comparison of Preprocessed and Unprocessed Images.
Original ImageTransferred ImagePreprocessed Image
Buildings 16 00494 i001Buildings 16 00494 i002Buildings 16 00494 i003
Table 3. Similarity Metrics for Transferred Image.
Table 3. Similarity Metrics for Transferred Image.
IDShape Context Similarity (S1 *)Key-Point Matching (K *)Texture Feature Similarity (T *)Color Gradation Similarity (C *)Comprehensive Semantic Inheritance Score (S2 *)
18850996377.5
28950995376.35
39550997381.45
49250997480.55
58850996677.95
68750997178.35
79050996678.65
89150998181.25
249350997180.45
* S1 = 1 − χ2(P1, P2) = 1 − ½∑(P1(i,j) − P2(i,j))2/(P1(i,j) + P2(i,j)). * K = ∑_(i,j) wij·exp(−d_(ij)2/2σ2). * T = cos(θ) = (V1·V2)/(|V1||V2|). * C = 1 − ½∑|H1(i) − H2(i)|. * S2 = 35%S1 + 25%K + 25%T + 15%C
Table 4. Image Similarity Metrics for Different Complexity Groups.
Table 4. Image Similarity Metrics for Different Complexity Groups.
Complexity GroupData IDComplexity RangeAverage Similarity Score
High Complexity Group18 11 22 10 20 24 7 69.96–15.2378.7125
Medium Complexity Group4 15 17 5 2 1 14 88.12–9.9178.65
Low Complexity Group16 23 3 19 12 9 21 133.6–8.0882
Table 5. Ablation Experiment Results.
Table 5. Ablation Experiment Results.
Experimental VariablesExperimental FlowchartExperimental Results
No VariablesBuildings 16 00494 i004Buildings 16 00494 i005
Only the Deep Information Extraction Module is DisabledBuildings 16 00494 i006Buildings 16 00494 i007
Disable only the positive enhancement word input moduleBuildings 16 00494 i008Buildings 16 00494 i009
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Xu, C.; Qu, C. A Modular AI Workflow for Architectural Facade Style Transfer: A Deep-Style Synergy Approach Based on ComfyUI and Flux Models. Buildings 2026, 16, 494. https://doi.org/10.3390/buildings16030494

AMA Style

Xu C, Qu C. A Modular AI Workflow for Architectural Facade Style Transfer: A Deep-Style Synergy Approach Based on ComfyUI and Flux Models. Buildings. 2026; 16(3):494. https://doi.org/10.3390/buildings16030494

Chicago/Turabian Style

Xu, Chong, and Chongbao Qu. 2026. "A Modular AI Workflow for Architectural Facade Style Transfer: A Deep-Style Synergy Approach Based on ComfyUI and Flux Models" Buildings 16, no. 3: 494. https://doi.org/10.3390/buildings16030494

APA Style

Xu, C., & Qu, C. (2026). A Modular AI Workflow for Architectural Facade Style Transfer: A Deep-Style Synergy Approach Based on ComfyUI and Flux Models. Buildings, 16(3), 494. https://doi.org/10.3390/buildings16030494

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop