1. Introduction
Conceptual design is among the architect’s most generative and least determinate tasks. It is the stage at which briefs are translated into spatial propositions, competing constraints of site, program, regulation, performance, expression, cost, and environmental impact are initially negotiated, and design alternatives remain open to exploration [
1,
2,
3,
4]. Traditionally, conceptual design has rested on tacit knowledge, precedent, iterative sketching, and rapid modeling, with limited tooling for systematic exploration. The recent emergence of generative artificial intelligence (GenAI) has begun to reconfigure these practices, especially in the early production of design narratives, visual options, and computational prompts.
Three technical developments are responsible. LLMs and multimodal foundation models have made natural-language interaction with design-oriented tools increasingly plausible beyond isolated research demonstrations [
5,
6,
7]. T2I diffusion models, including Stable Diffusion and Midjourney, can render visually rich concepts from short textual prompts within seconds [
8,
9]. Multimodal systems increasingly combine text, image, sketch, and three-dimensional input. The combination promises to expand both the speed and the bandwidth of early-stage design exploration. These tools have been rapidly adopted in experimental design practice and architectural education, although the evidence base remains uneven.
The promise has been accompanied by a fragmented evidence base. The literature on AI in architectural conceptual design grew quickly between 2018 and 2026, with contributions originating in both computer science and architecture. Yet the resulting corpus is uneven. Computer-science-oriented studies often evaluate models against generic image-, text- or layout-quality metrics that say little about architectural fitness. Architecture-oriented papers often demonstrate a workflow with a single building type and tool, without comparative evaluation or shared quality criteria. Influential commentaries frequently argue from principle rather than from observed practice. Conceptual coherence across these strands is weak, and the boundaries between conceptual design, schematic design, interior styling, and visualization are frequently elided.
Within the agenda of digital buildings, the relevance of GenAI does not lie only in rapid architectural imagery. Its larger significance depends on whether early-stage generative outputs can enter editable, evaluable, and lifecycle-oriented digital representations, particularly BIM and performance-linked design models [
10,
11]. For this reason, this review treats conceptual design not as an isolated act of visual ideation but as the front end of a digital-building workflow.
This review addresses that fragmentation. It takes GenAI applied to architectural conceptual design as its object, narrows the scope deliberately to the early-stage activity of generating, exploring, and refining design propositions, and asks three questions. RQ1: What are the principal modes in which LLMs, T2I diffusion models, spatial layout generators, and related GenAI systems have been applied to architectural conceptual design? RQ2: What can each technical pathway plausibly support, and where are its limits? RQ3: What unresolved problems should organize future research on GenAI for digital building design?
The paper makes three contributions. First, it proposes a four-domain framework that distinguishes language-driven semantic generation, image-driven conceptual visualization, spatially conditioned layout and 3D scene generation, BIM-AI coupling for concept-to-model translation, and human-AI collaborative workflows. Second, it reports the results of the reviewed literature along consistent dimensions—input, output, evaluation strategy, editability, and limitation—rather than presenting each tool in isolation. Third, it operationalizes four problem domains—controllability, evaluability, translatability, and responsibility—that organize the unresolved questions of the field. We follow PRISMA-ScR principles for reporting transparency but present this as a structured scoping-style review rather than a systematic review or meta-analysis [
12,
13].
The remainder of the paper is organized as follows.
Section 2 defines the scope and operational boundaries of the review.
Section 3 reports the search strategy, screening, appraisal, and coding procedure.
Section 4 reports the results of the scoping review.
Section 5 presents the thematic synthesis along the four-domain framework.
Section 6 discusses the cross-cutting issues of controllability, evaluability, translatability, and responsibility.
Section 7 acknowledges limitations, and
Section 8 concludes.
2. Scope and Conceptual Boundaries
Without clearly stated boundaries, the term “generative AI in architectural design” can refer to anything from autocomplete in CAD scripting to reinforcement-learning-based building control. This section delimits the object of the review.
Conceptual design, in this paper, refers to the stage between brief and schematic design at which initial design propositions are generated, explored, and selected. Outputs at this stage are typically exploratory: massing studies, programmatic diagrams, formal sketches, atmospheric renderings, and verbal narratives. Conceptual design is distinct from schematic design and design development, which presume a chosen proposition and develop it under tighter constraints. The distinction matters because the criteria by which a successful conceptual output is judged—generative breadth, communicative force, suggestive precision, and downstream potential—differ from the criteria for a successful schematic output, where coordination and feasibility dominate [
3,
4].
Generative artificial intelligence refers to machine-learning models trained to produce new content—text, image, three-dimensional geometry, or multimodal combinations of these—rather than to classify or regress over existing data. Within this broad class, the review concentrates on model families most visible in the conceptual-design literature: LLMs, T2I diffusion models, image-to-image or layout-conditional generative models, spatial layout and indoor-scene generators, and the workflows through which these systems are coupled to BIM or parametric representations [
14,
15]. Earlier generative families such as generative adversarial networks (GANs) are treated as relevant background since they shaped particular lines of work, most notably floor-plan generation, but the contemporary state of the field is dominated by transformer-based and diffusion-based models [
5,
8,
9,
15].
Interior-design and indoor-scene-synthesis studies were included when they addressed spatial layout, controllable scene generation, concept-to-visualization workflows, or representation issues transferable to architectural conceptual design. Studies limited to decorative styling without spatial, programmatic, or architectural contribution were treated as contextual references or excluded. This boundary is necessary because several recent studies on interior generation directly address problems of layout coherence, structure conditioning, and text-to-3D scene generation that are also central to architectural conceptual design.
Three adjacent fields are distinguished from the object of the review. Computational and parametric design often relies on explicit rules, constraints, and optimization procedures rather than trained generative models; it is therefore treated as adjacent unless the generator itself is learning-based or directly coupled to GenAI [
16]. AI-driven performance optimization applies machine learning to simulate or optimize energy, daylight, structural or acoustic performance; it is highly relevant to schematic and detail design but is not in itself a generator of conceptual propositions. AI-driven construction automation applies machine learning to fabrication, robotics, and site monitoring, which lies downstream of the design phase under consideration. Studies in these adjacent fields are cited where they illuminate the conceptual-design literature but are not the primary object of analysis.
Building information modeling (BIM) enters the review specifically as a downstream editable representation into which conceptual outputs may eventually be translated. We are interested in BIM not as a coordination platform in general, but as a computable design representation that can connect early concepts to downstream modeling, evaluation, documentation, and lifecycle management [
10,
11]. The question of how language-, image-, layout-, and scene-based generation can be coupled to BIM is therefore included as a bridge between conceptual exploration and digital-building workflows.
3. Review Methodology
The review follows a structured scoping-style approach informed by the PRISMA Extension for Scoping Reviews (PRISMA-ScR), which is better suited than a full systematic-review protocol to a field characterized by methodological heterogeneity and the absence of common quantitative outcomes [
13]. We do not perform meta-analysis. We report the search strategy, screening process, appraisal criteria, and coding procedure with the aim of making the review procedure reconstructible.
3.1. Search Strategy
Four databases were searched: Web of Science Core Collection, Scopus, Dimensions, and CNKI (
Table 1). The formal search covered January 2018 to March 2025, and a supplementary targeted search and citation-chasing stage was conducted in June 2026, covering structure-conditioned diffusion, text-driven interior/architectural visualization, and 3D indoor-scene synthesis. The 2018 lower bound corresponds to the period in which architectural GAN and deep-learning design studies became increasingly visible, while the post-2022 expansion of conversational LLMs and diffusion-based image generators motivated the supplementary stage. Selected contextual references, including foundational AI papers, design-theory sources, and product documentation, were cited where necessary but were not counted as coded studies unless they satisfied the eligibility criteria.
The English search string combined three concept blocks with Boolean AND: (“generative artificial intelligence” OR “generative AI” OR “large language model*” OR “LLM” OR “diffusion model*” OR “text-to-image” OR “text to image” OR “Stable Diffusion” OR “Midjourney” OR “DALL-E” OR “AI-assisted design”) AND (“architectural design” OR “building design” OR “architectural concept*” OR “conceptual design” OR “early-stage design” OR “floor plan generation” OR “building layout generation” OR “architectural visualization” OR “architectural visualization” OR “indoor scene synthesis”) AND (“generation” OR “design exploration” OR “design workflow” OR “human-AI collaboration” OR “BIM” OR “parametric model” OR “layout generation”).
The Chinese search string mirrored this structure: (“生成式人工智能” [generative artificial intelligence] OR “大语言模型” [large language model] OR “扩散模型” [diffusion model] OR “文本生成图像” [text-to-image generation] OR “AI辅助设计” [AI-assisted design] OR “人工智能辅助设计” [artificial-intelligence-assisted design]) AND (“建筑设计” [architectural design] OR “建筑概念设计” [architectural conceptual design] OR “建筑方案生成” [architectural scheme generation] OR “建筑布局生成” [building layout generation] OR “建筑可视化” [architectural visualization] OR “室内场景生成” [indoor scene generation]) AND (“概念设计” [conceptual design] OR “早期设计” [early-stage design] OR “设计生成” [design generation] OR “设计探索” [design exploration] OR “人机协同” [human–machine collaboration] OR “BIM” [building information modeling]). Searches were conducted in title, abstract, and keywords where database functions allowed.
3.2. Language Rationale and Eligibility Criteria
The review was restricted to English-language and Chinese-language records. English and Chinese are among the most widely used languages in current global academic communication, and together they offer broad coverage of international AI, architectural computing, and Chinese architecture-oriented design research. English was included because it is the dominant language of internationally indexed research on AI, architectural computing and human-computer interaction. Chinese-language records indexed in CNKI were included because a substantial body of architecture-oriented discussion on GenAI workflows, design pedagogy, BIM-AI integration, and practice-oriented experimentation is published in Chinese and is often not indexed by Web of Science, Scopus or Dimensions [
17]. The restriction is pragmatic and corpus-oriented rather than a claim that relevant work does not exist in other languages. Japanese, Korean, German, Spanish, Portuguese, and other literatures were outside the authors validated screening capacity and are acknowledged as a limitation.
Studies were included if they (i) described or evaluated a generative AI technique, (ii) addressed an architectural conceptual-design task at the building or building-cluster scale or a spatial design task transferable to conceptual design, and (iii) were written in English or Chinese. Review papers were included when they materially organized the field. Preprints were excluded unless they had subsequently appeared in peer-reviewed venues or were needed as contextual technical references. Technical reports of generic AI methods with no architectural application were excluded from the coded corpus, although foundational technical sources were cited for context where necessary. Civil engineering, construction-management and interior-styling studies without architectural or spatial design contributions were excluded, and duplicate publications were resolved to the version of record.
3.3. Screening
Screening followed the four-stage PRISMA-ScR flow shown in
Figure 1. The PRISMA diagram reports detailed full-text exclusion reasons. The full list of coded studies, contextual references, and coding decisions is provided in
Supplementary Table S1. Because the review combines formal database searching with targeted citation chasing after peer review, the review distinguishes coded evidence sources from contextual references throughout.
3.4. Appraisal and Coding
We used an appraisal framework informed by the Mixed Methods Appraisal Tool (MMAT) and adapted it to the heterogeneity of the corpus [
7]. Each coded study was assessed on five dimensions: clarity of research question, transparency of method and data, verifiability of results, relevance to architectural conceptual design, and discussion of limitations. Appraisal was qualitative and was not used as an exclusion criterion; rather, it informed the interpretive weight assigned to each study in the synthesis. Studies with limited methodological transparency were retained for mapping but discussed more cautiously.
Each study was coded on eight attributes: technical type, study object, design stage, input modality, output modality, evaluation strategy, editability, and primary stated limitation. “Study object” refers to the specific design or research target examined by each source, such as prompt generation, conceptual image production, text-driven interior design, graph-constrained floor-plan generation, indoor scene synthesis, LLM-to-script translation, BIM/parametric model generation, or designer-in-the-loop workflow evaluation. For T2I and diffusion-related studies, the named model or platform was also recorded when explicitly reported; where a source discussed commercial T2I workflows without specifying the engine, the tool route was described more generally rather than attributed to a specific platform. A calibration check was conducted on a subset of 12 studies, representing approximately 21% of the coded corpus. Discrepancies in domain allocation and coding labels were discussed to refine the coding scheme. The remaining studies were coded by one author and checked by another author for consistency; unresolved or ambiguous cases were discussed with the corresponding authors.
4. Results of the Scoping Review
This section reports the results of the review before the thematic synthesis. The aim is to make the evidence base visible: what kinds of studies were found, how they are distributed across domains, how they evaluate outputs, and how Chinese-language and English-language studies differ in emphasis.
4.1. Corpus Profile
The evidence corpus was first profiled to clarify the composition of the review before domain-level synthesis.
Table 2 summarizes the distinction between coded studies and contextual references, the language distribution of the coded corpus, and the balance between workflow-oriented, technical-generation and review-based sources.
4.2. Distribution Across Analytical Domains
The coded studies were then grouped into analytical domains according to their primary technical pathway and representational function in the conceptual-design process.
Table 3 reports the distribution of studies across these domains and summarizes their representative study objects, typical inputs, typical outputs and result patterns.
4.3. Appraisal and Evaluation Patterns
The appraisal results indicate that the field is active but unevenly evaluated (
Table 4). Technical studies usually report clearer model architecture, datasets, or evaluation metrics, whereas workflow papers often provide richer design context but rely on single-case demonstrations, designer reflection, or subjective preference. The review therefore treats strong claims about architectural effectiveness cautiously. Appraisal was used to calibrate interpretive weight, not to exclude studies.
4.4. English- and Chinese-Language Evidence
The Chinese-language corpus was not simply an extension of the English-language corpus. In the coded corpus, English-language studies more often emphasize model architecture, datasets, benchmarking, diffusion control, layout generation, and 3D scene synthesis. Chinese-language studies more often address design workflows, architectural education, BIM-AI integration, prompt-based conceptual exploration, and professional adoption. Across both language groups, however, systematic evaluation of architectural validity and downstream editability remains limited. This comparison justifies the bilingual scope while also showing why the language restriction is a limitation rather than a claim of global completeness.
5. Thematic Synthesis: A Four-Domain Framework
The four-domain framework is organized not only by model family but also by representational function in the conceptual-design workflow: semantic articulation, visual exploration, spatial/layout generation, and translation into editable models within human-AI workflows. This classification is preferable to a tool-based taxonomy because tools change rapidly, whereas representational transitions—text, image, layout, geometry, BIM/parametric model, and workflow—are more stable and architecturally meaningful (
Table 5).
5.1. Large Language Models for Design Semantics and Knowledge Support
The first domain concerns the use of LLMs as semantic engines in conceptual design. Three distinguishable sub-activities recur in the literature and warrant separation, since each makes different demands on the model and on its evaluation. In the coded corpus, LLM-related studies examined three main study objects: prompt generation for downstream image models, semantic structuring of design briefs, and natural-language translation into modeling or scripting instructions. Du et al. investigated the use of LLMs to generate prompts for text-to-image systems [
18]. Jiang et al. explored an interactive architectural design paradigm in which LLMs support Rhino-based design operations [
19]. Knowledge-oriented and regulation-related studies examined the ability of NLP or LLM-based systems to extract, interpret, or structure technical information while also exposing persistent reliability problems in specialized construction and regulatory contexts [
20,
21]. These study objects explain why LLMs are treated here as semantic and procedural support systems rather than as direct architectural form generators. Related advances in vision-language and multimodal foundation models further extend these capabilities by integrating visual and textual reasoning [
51,
52,
53].
The first sub-activity is design semantic generation. Here the LLM functions as a writing assistant that turns a designer’s brief or notes into a structured design narrative, an interpretation of the program, or a prompt for a downstream image model. Tasks of this kind exploit the model’s strengths—language fluency and generalist knowledge—and are relatively forgiving of factual error because the designer can inspect, revise, and redirect the output during iteration [
18,
19].
The second sub-activity is design knowledge support. Here the LLM is asked to retrieve, summarize, or apply codified design knowledge: building regulations, typological precedents, performance heuristics, and precedent-based design strategies. The empirical picture is more mixed. Studies have reported that contemporary LLMs perform reasonably on general code interpretation but exhibit hallucinations on jurisdiction-specific regulations and detailed technical clauses [
21]. Structured prompting strategies can guide LLMs through multi-step problems, but regulatory support still requires retrieval, verification, and professional review [
54].
The third sub-activity is natural-language to modeling-instruction translation. Here the LLM serves as a front end that converts a designer’s verbal description into executable code or parameter settings for downstream environments such as Rhino with Grasshopper, Revit with Dynamo, or game-engine geometry libraries. The technical contribution sits at the boundary between conceptual design and parametric implementation, and its evaluation requires a benchmark more rigorous than image-based quality assessment: whether the resulting parametric model is faithful to the designer’s intent and editable in expected ways.
5.2. Diffusion Models for Atmospheric and Structure-Conditioned Visualization
Since the introduction of denoising and latent diffusion models, diffusion-based architectures have largely displaced GAN-based approaches for high-resolution text-conditioned imagery [
8,
9,
55,
56,
57]. In the architectural conceptual-design literature, their use is concentrated around two different study objects: prompt-based atmospheric ideation and structure-conditioned visual generation. The distinction matters because these two study objects support different kinds of claims. The former mainly concerns visual fluency, speed, and exploratory breadth; the latter concerns whether generated imagery can be controlled by spatial or structural constraints.
The first route comprises commercial or designer-facing T2I workflows, including Midjourney, DALL-E, and other image-generation platforms when explicitly reported. Studies in this route typically examine how designers, students, or researchers translate a brief into prompts, generate alternative visual concepts, and select or refine images as early-stage design stimuli [
22,
23]. The results reported in this strand are strongest at the level of atmospheric communication: the tools can rapidly produce visually rich images, support stylistic exploration, and widen the range of initial visual options. However, these studies also show clear limitations. Prompt and model-version settings are often difficult to document exhaustively, outputs are not fully reproducible, and generated images are rarely evaluated against plan organization, structural feasibility, environmental performance, regulatory compliance, or downstream model editability.
The second route comprises open or configurable diffusion pipelines, especially Stable-Diffusion-related and ControlNet-related studies. Here the study object is not simply whether an attractive image can be produced but whether text, style, layout, boundary, or structure can condition the generated result in a more controlled way. Chen et al. developed a diffusion-model-based method for generating interior design from textual descriptions using a domain-specific interior-style dataset [
24]. Chen et al. further examined text-driven generation of master-style architectural designs, showing that prompts can guide outputs toward recognizable stylistic families while leaving architectural validity largely unresolved [
27]. Other studies focused on domain adaptation, aesthetics, and control: Chen et al. proposed an AI-driven diffusion approach for visually pleasing interior design generation [
25]; Yang et al. developed a controllable Stable Diffusion framework for panoramic interior design generation [
26]; and Chen et al. introduced an improved control network to match generated interior design images with existing indoor structure [
29]. These studies demonstrate that configurable diffusion pipelines can improve stylistic consistency, domain adaptation, and geometric alignment, especially when supported by fine-tuning, panoramic conditioning, structure maps, or ControlNet-like mechanisms [
28,
29,
47,
48,
49].
This body of literature therefore supports a more specific conclusion than the general claim that “diffusion models generate images.” Commercial T2I workflows are most often used for designer-facing ideation and atmospheric visualization, whereas Stable-Diffusion-based studies more often pursue methodological controllability, domain adaptation, and structure-conditioned generation. Across both routes, however, the dominant output remains an image. Even when edge maps, depth maps, indoor boundaries, or structure maps improve visual control, the result does not automatically become an architectural proposition. A generated façade may follow a drawn outline without demonstrating that the building behind that outline is programmatically coherent, technically feasible, or code-compliant.
We therefore characterize many T2I outputs as visually plausible but technically underspecified. They can support inspiration, communication, critique, and early visual exploration, but they do not by themselves demonstrate plan organization, structural feasibility, environmental performance, regulatory compliance, or BIM/parametric editability. This finding explains why the review treats diffusion-based visualization as a powerful conceptual medium but not as a sufficient form of architectural design evidence.
5.3. Spatial Layout and 3D Scene Generation
A distinct strand of literature concerns spatial layout and 3D scene generation. Unlike T2I studies whose immediate study object is usually an image or visual concept, studies in this strand take spatial relations themselves as the object of generation. They examine how room types, adjacency relations, floor-plan boundaries, object relations, semantic graphs, or textual instructions can be translated into layouts and scenes. This strand therefore differs from T2I visualization because it treats spatial organization, adjacency, object relations, and layout constraints as part of the generative problem. Its outputs are often less atmospheric than diffusion renderings but more analytically relevant to architectural conceptual design because they encode spatial relationships.
Graph-constrained floor-plan generation studies such as House-GAN and House-GAN++ take room types and adjacency relations as explicit inputs and generate house layouts or iterative layout refinements [
30,
31]. Their evaluation metrics—realism, diversity, and compatibility with graph constraints—are closer to architectural reasoning than generic image quality metrics. They also show the importance of relational representation: conceptual design requires not only visual style but also relations among rooms, circulation, and programmatic zones.
Indoor scene synthesis further extends the representational scope. ATISS generates plausible indoor object sets given room type and floor plan; LayoutGPT uses LLMs as visual planners that translate text conditions into layouts; DiffuScene models 3D indoor scenes as unordered object attributes; and InstructScene introduces instruction-driven 3D scene synthesis with semantic graph priors [
32,
33,
34,
35]. These studies are not equivalent to building-scale architectural design, but they directly address controllability, semantic structure, and text-to-spatial translation—issues central to early architectural generation.
The significance of this strand is that it occupies an intermediate position between visual concept generation and editable modeling. It is more spatially explicit than T2I rendering but less directly editable than BIM. Future work may use such structured layouts or scene graphs as intermediate representations between prompts, images, and BIM/parametric models.
5.4. BIM-AI Coupling: Bridging Concept to Editable Model
The BIM-AI and parametric-coupling studies in the coded corpus mainly examine whether natural-language descriptions, conceptual parameters, or AI-generated specifications can be translated into editable modeling operations. Their study object is therefore not the final rendered image but the representational bridge between concept, parameter, model, and downstream evaluation. Jiang et al. examined an LLM-supported interactive design paradigm connected to Rhino-based operations [
19], while BIM-AI and computational design studies discuss the integration of artificial intelligence with parametric or BIM environments and the possibility of converting early-stage design intentions into computable model variants [
22,
36,
37,
38,
39,
40]. This strand is important because it addresses the point at which GenAI outputs become usable within architectural production rather than remaining external images or textual descriptions.
A persistent obstacle in AI-assisted architectural design is the representational gap between what current GenAI systems can readily produce—prompts, images, and occasionally coarse geometry—and what downstream architectural workflows require: editable, semantically structured parametric or BIM models [
10,
11,
58,
59,
60,
61,
62]. This domain concerns work that explicitly addresses that gap.
The narrowest and most active strand is LLM-driven BIM scripting. Here an LLM acts as a translator between a designer’s natural-language description and the API of a BIM or parametric environment. Reported pipelines take a designer brief, infer programmatic parameters and generate scripts that instantiate a draft model [
19,
39,
40]. The output is editable in the expected sense: parameters can be adjusted and geometries can be modified through standard tooling. The limitation is that script generation accuracy degrades under cross-component constraints; a model that places single elements reliably may struggle with sequences of structurally interrelated members.
A second strand is AI-augmented early-stage design platforms, including commercial tools such as Autodesk Forma, TestFit, and Maket [
39]. These tools combine site analysis, generative layout exploration, and explicit design constraints, although the peer-reviewed literature on such platforms remains limited. Their value lies less in producing a finished architectural proposition than in exploring feasible configurations under designer-set rules.
For the present review, the salient methodological point is that studies in this domain are more likely than language-only or image-only studies to report functional evaluation: whether the resulting model satisfies a constraint, whether the layout meets an area program, whether a regulatory check passes, or whether parameters remain editable. These functional metrics are precisely what much of the diffusion-imaging literature lacks.
5.5. Human-AI Collaborative Workflows
The fourth domain concerns the integration of the foregoing techniques into design practice. Workflow-oriented studies generally take the design process rather than the model architecture as their study object. They examine how designers formulate prompts, curate image or layout outputs, iterate between language and visualization, and decide when generated material should be accepted, revised, or translated into another design medium [
23,
41,
42]. Other workflow studies and adjacent computational-design sources examine how performance simulation, sketch conditioning, segmentation, parametric modeling, or manual BIM translation can be inserted into the AI-assisted design loop [
43,
44,
45,
63,
64,
65,
66,
67,
68,
69].
The typical workflow described in the literature has three or four stages: requirement articulation, prompt or specification generation, generation by a diffusion model, 3D generator or parametric pipeline, and curation and iteration. The reported value of these workflows lies in accelerating early ideation, broadening the range of visual options and supporting designer reflection. However, the reviewed studies also show that workflow effectiveness is difficult to evaluate when studies rely on single cases, studio exercises, designer reflection, or subjective satisfaction.
The literature on workflow evaluation therefore remains thin. Comparative evaluation across designers, tasks, tools, or non-AI baselines remains rare, and few studies measure whether AI-assisted workflows improve architectural validity, downstream editability, or decision quality. This explains why human-AI workflow studies are valuable for understanding emerging practice but weaker as evidence for the architectural performance of GenAI systems.
A more accurate reading is that AI tools shift the locus of designer attention without abolishing its traditional content: the designer remains the agent who defines the problem, evaluates options against criteria that the tool does not encode, and assumes professional responsibility for the result. The reviewed evidence does not support strong claims that current GenAI systems can autonomously replace architects in conceptual design. The domain-level comparison is summarized in
Table 6.
6. Discussion
The cross-cutting issues in this literature do not reduce to any of the technical domains. We organize the discussion around four problem areas that recurred in the coding and that structure the unresolved questions of the field more usefully than a tool taxonomy alone (
Table 7).
6.1. Controllability
One of the most recurrent limitations identified in the coded corpus is the gap between aesthetic quality and design control. Diffusion models produce images that look like architecture; they do not necessarily produce representations that respect the constraints that distinguish architecture from concept art. ControlNet, graph-constrained generation and semantic-scene models improve controllability, but the field still needs machine-tractable representations of program, regulation, structural feasibility and environmental performance. This conclusion follows directly from the contrast between the T2I/diffusion studies and the layout/scene-generation studies reviewed above: the former show increasingly strong visual and stylistic control, whereas the latter show stronger spatial and relational control but weaker integration with architectural production workflows.
6.2. Evaluability
A second pervasive problem is the lack of shared evaluation criteria. Aesthetic preference, designer satisfaction and visual sophistication dominate; performance, regulation, program, and constructability rarely enter the evaluation loop. A practical evaluation framework should distinguish at least four levels: visual-semantic fit, spatial-programmatic coherence, technical feasibility and downstream editability. Until such frameworks exist, claims about whether a given AI workflow improves on practice cannot be settled. The reviewed studies therefore reveal an evaluability imbalance: image-generation and workflow studies are often evaluated through subjective visual judgement or designer reflection, while layout-generation and BIM/parametric studies more often permit constraint-oriented or functional checks.
6.3. Translatability
A third problem concerns the persistent gap between conceptual representations—text and image—and editable design representations—parametric and BIM models. Image-based generation produces visually rich but non-editable artifacts; BIM-based generation produces editable but often visually and conceptually underdeveloped artifacts. The most consequential research agenda is the closure of this gap, whether through text-to-3D models that mature into text-to-BIM, through structured intermediate representations such as scene graphs and layouts, or through workflows that explicitly couple them via designer-mediated translation. The study objects reviewed in
Section 5.2,
Section 5.3 and
Section 5.4 show why translatability is central: diffusion studies generate images, layout studies generate structured spatial relations, and BIM-AI studies attempt to generate editable models, but these stages remain weakly connected.
6.4. Responsibility
The fourth set of issues is professional rather than purely technical. AI-generated architectural imagery raises questions of authorship, intellectual property, and professional liability that the literature has barely begun to address [
69]. When a model adapted from a recognizable architect’s work generates a building in that architect’s style, the status of the resulting work is unclear in both copyright and disciplinary terms. When an AI workflow contributes substantively to a design, the question of who carries professional liability for code violations or structural failures remains unresolved. These questions cannot be resolved by technical advancement alone.
7. Limitations of the Review
This review has the limitations characteristic of structured scoping reviews. The search is bilingual but not multilingual: potentially relevant Japanese, Korean, German, Spanish, Portuguese, and other literature on computational architecture have been excluded by language. Although English and Chinese provide broad coverage of international and Chinese architecture-oriented research, the bilingual scope should not be interpreted as global completeness.
The review combines database searching with targeted citation. This strategy broadens coverage of fast-moving topics such as spatial layout generation, ControlNet-based interior generation and 3D scene synthesis, but it also means that the evidence base includes coded studies and contextual references that serve different functions.
Supplementary Table S1 therefore records the evidence status of each source explicitly.
A more substantive limitation concerns the field itself. The pace of GenAI development is such that any review will be partly obsolete by publication. Specific tool capabilities may shift rapidly across versions; the broader patterns of the field—the domain structure and the cross-cutting problems of controllability, evaluability, translatability and responsibility—are likely to be more durable.
8. Conclusions
This review has surveyed the application of GenAI to architectural conceptual design across four interlocking domains. The review findings and future research recommendations are separated below.
Review finding 1: The dominant pattern of work in the conceptual phase pairs language models for semantic and prompt-engineering tasks with diffusion models for visualization. This pairing is effective at the level of visual concept generation and is increasingly visible in practice-oriented literature and experimental workflows. It is not, on the current evidence, effective at the level of autonomous architectural proposition: the gap between an attractive rendering and a workable design remains substantial and is presently bridged by the designer rather than by the tools.
Review finding 2: Spatial layout and 3D scene generation form a distinct and increasingly important strand. These studies are more constraint-aware than T2I visualization and more spatially explicit than prompt-based workflows, but they remain only partially connected to building-scale architectural practice and BIM-based downstream work.
Review finding 3: The principal unresolved problem is not visual quality alone but translatability—the gap between conceptual representations that AI generates well and editable design representations that downstream design and construction require. Closing this gap, through BIM-AI coupling, text-to-3D and text-to-BIM pipelines, structured intermediate representations, or designer-mediated translation, is the most consequential research direction for the field.
Future research priority 1: Develop multi-criteria evaluation frameworks that combine visual, functional, regulatory, aesthetic and editability criteria and can be applied across tools and studies. Future research priority 2: Mature text- and image-to-editable-model pipelines, particularly in coupling with BIM platforms embedded in mainstream design practice. Future research priority 3: Conduct empirical studies of designer practice and education under GenAI, including comparisons across designers, tasks, tools and non-AI baselines. The reviewed evidence does not support strong claims that current GenAI systems can autonomously replace architects in conceptual design. The more useful question is what kind of architecture architects produce with it, and what the field needs to ask of itself to make that work rigorous, evaluable and responsible.