Next Article in Journal
Utilisation of Oil-Contaminated Sand in 3D-Printed Concrete: Rheological, Mechanical, and Microstructural Assessment
Previous Article in Journal
Probabilistic Characteristics Study of Tensile Properties of Bamboo Inter-Node Material Based on Random Field Theory
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

TRACE-SC: A Protocol-Based Framework for Mapping Generative AI in Space-Constituting Architectural Design Decisions

1
Architectural Design Computing Graduate Program, Istanbul Technical University, 34485 Sarıyer, Türkiye
2
Department of Architecture, Istanbul Technical University, 34367 Sisli, Türkiye
*
Author to whom correspondence should be addressed.
Buildings 2026, 16(14), 2827; https://doi.org/10.3390/buildings16142827
Submission received: 18 June 2026 / Revised: 10 July 2026 / Accepted: 12 July 2026 / Published: 16 July 2026
(This article belongs to the Topic Architectural Education)

Abstract

Architectural space is produced through decisions about form, material, construction, structure, program, and environment, and design education centers on it. In design studios, generative AI (GenAI) most often enters through visual representation, raising the risk we term the visualization trap: images can appear resolved before their tectonic implications are worked out. This exploratory study examines GenAI participation through protocol analysis of the documented process traces of 18 students in a single, AI-aware bioclimatic design studio, yielding 1107 protocols. The TRACE-SC methodology maps space-constituting components, GenAI use types, design phases, and cognitive breaking points; reliability was examined through a blind expert coding audit and cross-LLM comparison. GenAI appeared in 44.2% of protocols (489/1107), where visual generation and information gathering accounted for 81.2% (397/489). Suggestions were transformed before use in 68.1% (333/489) and adopted verbatim in 0.4% (2/489); interaction was designer-initiated. Cognitive breaking points appeared in 2.2% of protocols (24/1107), with indirect evidence in 70.8% (17/24). GenAI proves more than a visual production tool, but its engagement is limited, episodic, and designer-steered rather than a routine design partnership—partial support for the proposition. The transferable contribution is the TRACE-SC framework and codebook; its patterns describe this studio and invite comparison elsewhere.

1. Introduction

Architectural design proceeds mainly through decisions about form, material, construction, structure, program, and environmental response. These decisions are not contained only in final drawings or images; they are negotiated through moves, revisions, and justifications across the design process. Since 2022, generative artificial intelligence (GenAI) has entered this process through text, image, and multimodal tools [1,2,3]. In architectural design studios, however, its contribution is often read through the images it produces. This raises the central problem of the study: how does GenAI participate in the development of architectural design decisions, whether as an instrument or in some form of partnership, rather than only in the production of visual output?
Much of the scientific literature on GenAI in architecture foregrounds visual generation, ideation, creativity, and student experience [1,2,3,4,5,6]. This focus is valuable, but it obscures a specific risk. A generated image may appear persuasive before the proposal has been tested as a spatial, tectonic, and buildable project. This study refers to that tendency as the visualization trap. It is not operationalized as a measured individual outcome or a causal effect, but used as a sensitizing concept for examining whether documented GenAI engagement shows a corpus-level imbalance toward imageable form—especially the concentration of visual generation around formal decisions and the relative thinness of material, constructional, and structural reasoning in the process record.
The gap is therefore twofold. First, many studies evaluate GenAI through final outputs, student perceptions, or broad pedagogical accounts [6,7,8], while fewer trace how decisions develop while GenAI is being used. Second, process-oriented accounts rarely connect GenAI use to the material, structural, and constructional elements through which architectural space becomes a concrete, buildable proposal. The present study addresses this gap through protocol analysis based on documents and artifacts of an undergraduate design studio.
To address this problem, the study develops a faceted analytical framework, which it names TRACE-SC (Tracing Roles of AI across Components and Episodes in Space-Constituting design), for reading GenAI engagement across architectural design processes. Accordingly, its principal research question asks how GenAI becomes involved in space-constituting design decisions, and through which use-related and cognitive patterns this involvement becomes visible in documented process traces within architectural design education. The framework follows four connected dimensions: (i) space-constituting components (SC), (ii) types of AI use (AU), (iii) design process phases (DP), and (iv) cognitive breaking points (BP). Four subquestions follow, in sequence:
  • SQ1: How does GenAI relate to these space-constituting components?
  • SQ2: Through which AI use types and interaction forms does this engagement appear?
  • SQ3: How is this engagement associated with cognitive breaking points?
  • SQ4: What opportunities and risks follow?
The study does not begin by treating GenAI as a design partner; it asks what the evidence supports. To keep this proposition falsifiable, three evidential outcomes were defined in advance: support, partial support, and failure to find grounding, each read from converging indicators rather than a single numerical threshold. The manuscript first develops the theoretical and methodological basis of the framework, then reports the findings, discusses their implications, and concludes with contributions, limitations, and future work.

2. Theoretical Background

The analytical apparatus of this study is a faceted codebook that reads four dimensions of the design process at once. It draws mainly on design cognition, which treats designing as an observable and decomposable activity. Architectural tectonics and building-science discourse are used more selectively, to define the architectural meaning and boundaries of the space-constituting components. Design cognition therefore provides the theoretical lead; architectural theory clarifies the decision domains. The section develops this grounding in three steps and closes by identifying the research gap.

2.1. Design Cognition and the Design Process

The study’s principal research question, which asks how GenAI becomes involved in space-constituting design decisions, assumes that the development of those decisions can be observed. Yet design is often described as intuitive, tacit, or difficult to inspect from the outside. If it is judged only through final products, the design process remains a black box. To trace GenAI at the level of decisions rather than final imagery, the process must be read through observable and operationally distinguishable moves.
Design cognition makes this reading possible. Donald Schön [9] described designing as reflection-in-action: a reflective conversation in which the designer frames a problem, makes moves, and reads the consequences of those moves. Nigel Cross [10] argued that design has its own designerly ways of knowing and observable practices, helping move the field away from the black-box view. Bryan Lawson [11] similarly treated design thinking as an empirically describable process rather than an inaccessible talent. Designing can be observed because it externalizes itself. It leaves discursive and visual traces that can be read as evidence of cognition. Willemien Visser [12] conceptualized designing as the construction of cognitive artifacts, placing such traces within a systematic frame. This is especially relevant in undergraduate studios, where design work is pedagogically documented and repeatedly externalized [13].
Once designing is treated as observable, the next question is how the process is structured. Early design-methods research modeled design as a sequence of analysis, synthesis, and evaluation [14]. Later design-cognition research showed a more cyclical and opportunistic pattern: designers move between activities, work on problem and solution together, and return to earlier decisions [10,11]. This study uses exploration, production, and evaluation to describe phases of thinking and action. The triad is grounded in the design-methods tradition [14], but it is not treated as a fixed ladder. Following Cross [10], the phases are read as behavioral descriptions within a cyclical process. Reviews and interim presentations are therefore treated as moments inside the process, not as a separate terminal stage.
The cyclical character of design follows from the structure of design problems. Herbert Simon [15] described design problems as ill-structured: they lack a definitive formulation, admit no single correct solution, and often leave goals indeterminate. Horst Rittel and Melvin Webber [16] described planning and design problems as wicked, with no exhaustive definition or clear stopping rule. Architectural design problems share these qualities. Problem and solution co-evolve [17], and designers reflect both in action and on action [9]. This justifies a cyclical reading of the studio process and the decision not to assign sessions to fixed phases in advance.
Schön [9] places framing and problem setting at the center of reflective practice, and Kees Dorst [18] treats frame creation as a defining act of design. Cross [19] discusses the creative leap, a term he traces to L. Bruce Archer [20]; backtracking, in turn, has been tracked empirically as part of design activity [21]. On this basis, the study distinguishes four types of cognitive breaking points: reframing, direction change, creative leap, and backtracking. The source of a breaking point, whether the designer, GenAI, or another actor, is treated separately to avoid causal claims after the fact.

2.2. Space-Constituting Components (SCs) and Their Evaluation

The traceability and cyclical structure of design lead to the study’s central analytical unit: SC. The constitution of architectural space is treated not as one holistic act, but as a composite of distinguishable decisions. Design cognition supports this decomposition because moves and cognitive actions can be traced through protocol data [22,23,24]. Architectural thought gives these decision domains their disciplinary meaning. Without decomposition, constituting space remains too broad to analyze; once decomposed, it becomes possible to examine where GenAI becomes visible and with what intensity.
Seven decision types are defined within this study: conceptual, material, constructional, formal, structural, programmatic, and environmental. The set is not imported from a single existing model. It is developed for this study at the intersection of design-cognition decomposability and architectural decision domains. Existing schemes, such as the coding of cognitive actions by Masaki Suwa, Terry Purcell, and John Gero [24] or Gero’s [22] Function–Behavior–Structure framework, support the premise that design can be traced through decision moves, but they do not directly yield these seven categories. A domain is included only if it leaves a discursive, visual, or numerical trace; corresponds to a distinguishable architectural decision domain; and can be separated operationally from neighboring domains.
The classification deliberately separates domains that some architectural traditions discuss together. Material selection and construction logic are often linked in the classical tectonic tradition [25,26,27], but they are traced separately here because they appear through different indicators in protocol data. The logic of joining parts and the logic of the load-bearing system are also separated, following Eduard Sekler’s [28] distinction among structure, construction, and tectonics [29]. The structural component is narrowly defined as load-bearing system and load-transfer logic, supported by structures discourse [30,31,32]. The conceptual component names the intentional orientation that frames the design’s semantic content and interpretation of context. The programmatic component is grounded in the function term of the Function–Behavior–Structure framework [22]. The environmental component is grounded in bioclimatic design discourse [33,34,35] and traces the translation of context into spatial decisions, not raw context itself.
Because several of these domains are discussed together in architectural theory, each pair of neighboring domains is separated in coding by an explicit discriminating question (Table 1); form, for instance, is coded independently of both its underlying concept and its load-bearing logic. The full operational definitions, per-component boundary notes, construct-validity origins, and primary/secondary boundary heuristics with worked examples are provided in Supplementary S2.
These components are not equally easy to trace. Material, structural, programmatic, and environmental decisions often leave relatively concrete lexical or numerical indicators. Conceptual, constructional, and formal decisions more often require interpretation of discourse and visual evidence. This difference is treated as a property of the data rather than a defect. The value of the scheme lies in making these domains separately assessable, so that the sensitivity and limits of each measure remain visible instead of being absorbed into a single holistic category.

2.3. Research Gap

Once design is made observable through traces, the next issue is how those traces are systematized. Protocol analysis offers one established answer: discursive, visual, and behavioral traces of cognitive activity are segmented into analytic units and coded with explicit categories. Its contribution here is methodological. It shows that design processes can be analyzed through categorical coding schemes. Gero and McNeill [36] segmented design protocols into coded units; Suwa, Purcell, and Gero [24] proposed a multidimensional scheme for coding cognitive actions at physical, perceptual, functional, and conceptual levels; and Gabriela Goldschmidt [23] coded design moves and links to reveal patterns in the process. Categorical coding is therefore a well-established tradition in design-process research [37].
The framework developed in this study follows this tradition, but it changes the content of the coding problem. Earlier schemes often code levels of cognitive action, Function-Behavior–Structure relations, or general moves and links. TRACE-SC instead treats space-constituting decisions as an architectural axis and adds GenAI use as a separate axis. It also relies on process traces based on documents and artifacts rather than concurrent recording. This makes the approach suitable for a studio corpus composed of reports, visual material, and embedded GenAI logs.
The GenAI axis must also be placed within the wider debate. GenAI has entered architectural design rapidly, and much of the literature frames its use through visual production, inspiration, or creativity triggering. This framing tends to locate interaction at the formal and representational level of design [4,38]. At the same time, studies note the limits of these tools as design advances, including loss of scale, weak abstraction, and difficulties of control [39,40]. This concern relates to design fixation, where strong and prematurely formed solution images can narrow divergent search [41]. Beyond architecture, recent studies of GenAI ideation report fixation on first examples, reduced diversity and originality [42], and increased individual output alongside decreased collective diversity [43]. The persuasive force of a generated image may therefore draw attention too early toward a finished-looking result and weaken the link to tectonic, structural, and buildability concerns.
Read together, this literature reveals the gap addressed by this study. Conceptually, GenAI research in architecture has not sufficiently distinguished how the tool relates to the different decision domains through which space becomes a concrete proposal. Methodologically, studio and design-education studies rely mainly on outputs, surveys, and experience-based accounts, while protocol-analytic approaches remain rare in relation to GenAI. Cognitively, debates about GenAI’s expansive and restrictive potentials often focus on creative outputs and idea diversity, leaving its relation to transformations inside the design process less theorized. The missing object is the developmental logic through which space-constituting decisions mature, and the way GenAI becomes visible within that logic.
The four axes of TRACE-SC are designed to make this relation traceable. The SC axis identifies the architectural decision at stake. The AU axis identifies what role the tool plays. The DP axis locates the move in the developmental arc. The BP axis asks whether the move reorients the design. Their value lies in cross-readings, such as AI use by component or AI use by breaking point. The operational form of this framework is set out in Section 3.

3. Materials and Methods

3.1. Research Design

The empirical setting was an undergraduate architectural design studio (MIM 401) run over a 13-week summer term, in which 18 students each developed an individual project and their routine studio work was followed across the full term. The studio ran alongside a parallel AI-awareness course (MIM 405) in the same program, which introduced generative text- and image-based tools through a critical conceptual frame. Because the two courses ran together, the study examines a primed studio context rather than unframed GenAI use, a condition carried into the interpretation of the findings. The corpus is therefore useful for examining visible GenAI engagement, but not for estimating ordinary or unframed GenAI use in architectural education. The corpus analyzed here is the studio’s routine output (weekly reports, visual material, and embedded GenAI logs) produced throughout the semester.
Within this study, the TRACE-SC framework is proposed; it draws on protocol analysis based on documents and artifacts within an interpretive paradigm [44]. It traces how GenAI participates in the development of architectural decisions, without treating that participation as an isolated causal effect. The design is qualitative at its core. A structured metadata layer is embedded within the qualitative protocol record and is used descriptively, which corresponds to an embedded mixed-methods design [45]. This choice fits the research questions: the study needs to describe the distribution of GenAI engagement, the pattern of designer–GenAI interaction, and the interpretive texture of design discourse, including shifts in cognitive orientation.
Each protocol record combines two layers generated from the same analytic unit. The first is qualitative: written discourse, visual production, GenAI interaction traces, and contextual notes. The second is structured metadata composed of categorical fields. These nominal fields are quantitized into frequency, distribution, and cross-tabulation summaries [46]. The numerical layer remains descriptive. It maps tendencies in the corpus; it does not support inferential testing or statistical generalization.
The study is organized around a working proposition rather than a confirmatory hypothesis. The proposition offers a directional frame that can be supported, bounded, or left unsupported by the evidence. Its testable core is that GenAI does not serve only as a visualization tool, but appears across cognitive phases and components of designing.
The method differs from concurrent verbal-protocol research. It does not use think-aloud protocols or follow the verbal-protocol model of K. Anders Ericsson and Herbert Simon [47], where cognition is accessed through real-time verbalization under laboratory conditions. Instead, it reads the cognitive artifacts [12] routinely produced in the studio: weekly reports, GenAI dialogue logs, presentation sheets, sketches, models, and renders. This gives the study naturalistic, multi-source records, but also means that it analyzes documented traces rather than the unfolding process itself. Figure 1 summarizes the workflow from research design to data processing, codebook construction, trustworthiness procedures, and reporting.

3.2. Study Context and Participants

The studio brief asked students to design bioclimatic mobile stations for nature exploration across four extreme climate regions selected to maximize contextual variation: Tropical (Manaus, Brazil), Desert (Siwa Oasis, Egypt), Monsoon (Cherrapunji, India), and Cold (Svalbard, Norway). Students were randomly assigned to four five-person climate groups, and each student developed an individual design. The climate regions are used as a means of contextual variation in the studio, not as statistical comparison groups.
During the study period, the GenAI landscape included ChatGPT, with OpenAI’s GPT-4o becoming available during the term; DALL·E 3 (OpenAI); Midjourney v6, which was the platform’s default model for most of the studio period; Adobe Firefly, including the Image 3 model then available in beta alongside Image 2; and Leonardo.Ai, including Phoenix, its first foundational model. These platforms allowed users to select among current or legacy models; therefore, students chose tools and versions freely, and model versions were documented where visible but were not standardized or controlled.
Of the 20 students continuing in the studio, 18 agreed to participate, producing a corpus of 1107 design protocols throughout the semester. Because all consenting students were included, no sampling strategy was applied. Students are identified as S1–S18. Because the researcher’s insider position [48] could, in principle, introduce pedagogical authority, recognition effects, or performative reporting, these possibilities were taken into account through the reflexivity procedures described in Section 3.5 and the studio’s multi-actor structure.

3.3. Unit of Analysis and Segmentation

The corpus is read at three nested levels: raw deliverable material; protocol units extracted from that material; and protocol records, where each unit is standardized in a markdown template with quotation, explanation, evidence, codes, and metadata. Data accumulated through routine studio work during the term. The material was submitted progressively across the semester, as the studio advanced, rather than assembled as a single end-of-term retrospective account. No research-specific deliverables or laboratory conditions were added. Segmentation, coding, and metadata assignment were then conducted after the term on the completed dataset. These records nonetheless remain curated deliverables prepared for studio assessment rather than live process recordings, so the corpus is treated as a temporally staged documentary record and as a documented floor of visible process activity.
A protocol unit is defined as a traceable design event. A design event becomes a unit when it falls within one of the four codebook facets (Section 3.4) and shows a move such as a decision, rationale, alternative, evaluation, revision, GenAI interaction, or spatial/formal proposition. This category-guided threshold ties segmentation to the codebook while acknowledging that boundary decisions remain interpretive. Borderline cases are recorded through bias notes and the audit trail. To avoid inflated counts, the same move repeated across multiple sources in one week is recorded as one protocol record with multiple evidence items. A recurring unchanged visual creates a new unit only when accompanied by morphological revision, new annotation, or a new contextual position. For GenAI dialogues, the unit is the coherent interaction moment, defined by the designer’s cognitive move rather than by the number of conversational turns.
Because protocol-unit counts form the denominators of normalized rates, the threshold is applied consistently across students and weeks. Unresolved boundaries are first marked provisional and checked against chronological consistency, using only information available at that point; this review verifies the continuity of unit boundaries rather than inferring motives or causes retrospectively. Cases that remain unresolved are reported in the unattributable set rather than silently dropped. These procedures reduce, but do not eliminate, the interpretive judgment involved in retrospective segmentation. Full segmentation rules—upper and lower unit boundaries, multi-source handling, repeated visual material, AI-dialogue units, and boundary-case resolution—are provided in Supplementary S1.

3.4. The Faceted Codebook

The codebook follows a faceted classification logic in an established faceted-classification tradition [49]. Each protocol unit is scanned across four independent analytical dimensions and coded on any facet for which evidence is present; not every unit is coded on every facet. Table 2 gives the facets, their purpose, subcodes, and permitted coding density. This structure supports the cross-readings required by the research questions, such as AI use by space-constituting component or AI use by breaking point.
The SC facet identifies the architectural decision domain; AU records how GenAI is used; DP locates the move within the design process; and BP marks whether the move reorients the design. Output type, interaction mode, interaction direction, and triggering source are recorded as metadata rather than as facet criteria. This separation keeps the analytical dimensions distinct while allowing them to be compared in cross-tabulations. Before full coding, one student’s process was used as a pilot to check whether the codebook produced coherent and usable classifications. This pilot did not alter the codebook structure; the same codebook was then applied to all 18 students.
Construct validity was examined by classifying each code as a priori, mixed, or emergent. A priori codes are named directly in the source literature; mixed codes are guided by the literature but operationalized through data and researcher judgment; emergent codes are data-driven and later aligned with the literature. Most codes are mixed-origin, which is appropriate for qualitative work and avoids presenting interpretive categories as falsely deductive.
Coding uses two forms of multiple coding. Vertical multiple coding across facets follows from the faceted structure. Horizontal multiple coding within a facet is used only when one move independently evidences more than one subcode, within the limits shown in Table 2. A primary code is then selected by three boundary rules applied in order: emphasis, logical sequence, and chronological context. Secondary codes require written justification. Twelve structured metadata fields and one optional audit field capture documentation type, AI tool, interaction mode and direction, suggestion incorporation, discourse-visual correspondence, and related dimensions.

3.5. Trustworthiness and Coding Reliability

The reliability architecture is based on Lincoln and Guba’s trustworthiness framework [50]: credibility, dependability, confirmability, and transferability. Five mechanisms operationalize this framework: M1 data triangulation across written, visual, and GenAI traces; M2 negative-case analysis; M3 an audit trail for segmentation, coding, and interpretation decisions; M4 reflexivity through researcher-position disclosure and protocol-embedded bias notes; and M5 thick description.
A chance-corrected coefficient was not used as the primary reliability index. Category-guided segmentation links unit and code boundaries, so independently segmented comparisons could mix unitizing disagreement with coding disagreement. In addition, breaking points are rare, appearing in 24 of 1107 protocols (2.2%); at this base rate, a small number of disagreements can strongly affect chance-corrected estimates [51,52,53].
Within this frame, an independent blind expert audit was conducted to examine coding reproducibility. The auditor was an institution-internal expert peer, a faculty member directing a different studio who did not share the MIM 401/405 priming, did not teach the student cohort, and did not participate in codebook development. The auditor coded blind, using only the raw protocol unit and the codebook, without seeing the canonical codes generated with LLM support and finalized by the researcher. The material coded by both reviewers included S10’s full term (92 protocol units) as a deep single case and a stratified scattered sample of 194 units across 16 students, selected with a seeded blind audit script across students and AI activity levels. Together, these produced about 286 double coded units, or 25.8% of the corpus, with coverage across students, climate groups, and AI activity levels.
Because both coders worked on the same pre-segmented units, the audit isolates code-assignment disagreement from unit-boundary uncertainty. Raw agreement is therefore reported as a transparent reproducibility check, while chance-corrected coefficients are treated as diagnostic rather than primary evidence.
Agreement is reported at two levels. Primary code agreement asks whether both coders selected the same dominant subcode. Overlap in the two leading codes asks whether any code was shared, regardless of order. This distinction matters because faceted coding often identifies co-present components. Evidence comes from the deep single case, the scattered sample, and a comparison across LLMs. Table 3 reports raw percent agreement. Supplementary S4 also reports Krippendorff’s alpha values [54] on the blind audit subsample, interpreted alongside the trustworthiness framework because category-guided segmentation, multi-label interpretive coding, and skewed category distributions limit what such coefficients can establish.
Two caveats apply to these values: the AU figure is partly elevated by protocols both coders marked as no GenAI use, and the BP presence value in the comparison across LLMs is inflated by shared absence of a rare category. Even in the deep single case, where primary agreement is lowest, the two leading codes overlap in 89.1% (SC) and 91.3% (DP) of units, so disagreement there is largely about which co-present code is primary rather than whether a component was detected.
Breaking-point reliability was not estimated through external coding. We initially explored whether an independent coder could reproduce breaking points, but two exploratory attempts during the blind expert evaluation showed that this is not feasible within the available retrospective documentary corpus and audit design, for a rare, high-inference category. In a blind chronological reading of three students’ full terms, the expert detected only 1 of 9 canonical breaking points; and a context-windowed exercise, in which the design steps immediately before and after each breaking point were shown, imposed a high interpretive load and did not yield reliable judgments. We therefore could not develop a separate breaking-point reliability procedure, and we state this as a limitation. Rather than a coding deficiency, this outcome reflects the rarity of breaking points and the difficulty of recovering a cognitive reorientation from documentary before-and-after traces. Reliability for this facet is instead addressed through confirmability mechanisms: audit trail, evidence-level reporting, and disciplined thresholds. Segmentation reliability was also not externally validated against independently determined unit boundaries, and this is stated as a limitation.
A further, independent observation came from outside the coding pipeline. Five senior experts in architectural and computational design, external to the studio team, evaluated the final projects using a structured rubric with five points based on the same space-constituting components. Students did not know the rubric or criteria, so the ratings function as an unobtrusive measure. This external evaluation is reported as an independent, outcome-level observation and contextual background only, not as a validation of the study’s coding (Section 5.2). The rubric, rating distribution, and within-rater consistency are reported in Supplementary S5.

3.6. Candidate Coding Assisted by LLMs

An LLM produced the first consistent pass of segmentation and coding candidates for all 1107 protocol units. For each unit, it generated reasoned candidate codes and metadata. The first author then read every unit with its candidates, assessed them against the codebook, and endorsed the final coding while retaining authority over every code. The LLM is therefore treated as a reasoned candidate generator, not as a coder or a statistical classifier. All material processed with external LLM tools was anonymized beforehand: student identifiers were replaced with codes (S1–S18), and names, signatures, and other direct identifying marks were removed from both documents and images, so that no direct personal identifiers were transmitted to the external tools. Claude (Anthropic; Claude Max plan) and ChatGPT (OpenAI; ChatGPT Pro plan) were accessed within Visual Studio Code 1.110 (Microsoft) through the Claude Code and Codex extensions, respectively; before processing, the account-level settings permitting use of sessions for model improvement were disabled. A companion script parsed the researcher-approved markdown fields into a relational database for cross-tabulation; it made no natural-language processing or coding decisions. During piloting, multi-model triangulation was attempted. One pilot model was removed from the primary and comparison roles because it produced frequency inflation, breaking point thresholding after the fact, and source position hallucination. A single model was retained as the primary candidate generator.
Because canonical codes begin from LLM-generated candidates and the first author’s review largely concurred with them, anchoring risk remains. The independent blind expert audit described in Section 3.5, conducted without access to the canonical codes, is the main safeguard for the SC, AU, and DP facets. For the rare BP facet, which was not externally coded, the remaining risk is addressed through the confirmability mechanisms described above.

4. Results

This section reports the patterns identified in the 1107 design protocols obtained from the students. It addresses the descriptive parts of the study: the distribution of GenAI engagement across space-constituting components (SQ1), the AI use types and interaction forms through which this engagement appears (SQ2), and its association with shifts in cognitive orientation (SQ3). Variation by student, climate group distributions, negative cases, and anchor cases are then used to qualify the pattern. SQ4 concerns opportunity and risk and is therefore addressed in the Discussion. Four denominators are kept separate throughout: the full corpus (1107 protocols), protocols with a primary space-constituting component (1099), GenAI-related protocols (489), and GenAI-related protocols with a primary component (482).

4.1. Space-Constituting Engagement: Component Profile

SQ1 asks which space-constituting components (SCs) are associated with GenAI, and at what intensity. GenAI is present in 489 of 1107 protocols (44.2%). This is a documented minimum because the analysis captures only GenAI use that left an observable trace in the studio record.
In the full corpus, the component profile is uneven (Figure 2). The formal component (SC4) is the largest category, followed by programmatic, conceptual, and environmental components. Material, constructional, and structural components are the least frequent. Architecturally, this suggests that recorded design attention concentrates on formal and representational decisions, while buildability-related registers remain thinner. In this study, structure refers to load transfer and the load-bearing system, not to the abstract structure term in the Function–Behavior–Structure framework [22].
The GenAI-related subset with a primary component (n = 482) shows a different mix. Conceptual and material components rise, while programmatic and environmental components fall. The formal component remains the largest category. Figure 2 summarizes this compositional difference between the GenAI-related subset and the corpus as a whole; they do not show that GenAI caused any component to appear. In the full corpus, 8 protocols carry no primary space-constituting component; among the 489 GenAI-related protocols, 7 carry none.
A final pattern concerns focus and accompaniment. The primary code marks the focal component of a protocol, while secondary codes record components present in the same move. Formal engagement is high in both layers, with 321 primary and 300 secondary codings. Environmental and conceptual components appear more often as accompanying context than as the focal decision. Environmental is coded 138 times as primary and 468 times as secondary; conceptual is coded 170 times as primary and 295 times as secondary. The same asymmetry is visible within the GenAI-related protocols.

4.2. AI Use and Interaction in Relation to Components

SQ2 asks which AI use types and interaction forms constitute GenAI participation. Among the 489 GenAI-related protocols, primary AI use codes are concentrated in visual generation (AU2; 227 protocols, 46.4%) and information gathering (AU1; 170 protocols, 34.8%). Together, they account for more than four-fifths of primary AI use codes. Proposing new possibilities (AU6), critical dialogue (AU4), design testing (AU3), and scenario building (AU5) appear less often as primary AI use codes. The secondary layer adds nuance: AU6 appears 39 times as a primary AI use code but 134 times as a secondary AI use code, suggesting that it often accompanies another activity rather than leading the protocol. Text and visual outputs are nearly balanced, so GenAI engagement cannot be described as only visual.
The relation between AI use and spatial component can be read directly from the co-occurrence heat map in Figure 3, in which each cell counts protocols where the row use type and the column component are both primary, within the 489 GenAI-related protocols.
Three readings stand out (Figure 3). First, AU2 is strongly tied to formal decisions: 125 of 147 formal GenAI-related protocols (85.0%) are AU2 cases. Second, AU1 is more widely distributed across conceptual, programmatic, material, environmental, and constructional registers. Third, the more metacognitive uses sit closer to conceptual and programmatic decisions. AU6 co-occurs most with programmatic and conceptual components, while AU4 is most often associated with the conceptual component. In short, output production concentrates on represented form, whereas uses oriented toward process appear nearer to the conceptual core of the project.
Suggestion uptake is the most distinctive SQ2 pattern. Of the 489 GenAI-related protocols, 333 (68.1%) transform a GenAI suggestion before incorporating it, while only 2 (0.4%) adopt one verbatim; the remainder are ambiguous, carry no incorporable suggestion, or reject the suggestion. Among the 335 protocols in which a suggestion enters the design, uptake is therefore almost entirely transformative. This pattern is visible across use types and components, especially where AU2 meets formal decisions and where AU1 supports conceptual or programmatic decisions.
Interaction form and direction reinforce this reading. Among the 479 protocols with a coded interaction form, iterative exchange is dominant (308 protocols, 64.3%). Single shot and cyclical interactions follow. Direction is also weighted toward the designer: 370 of 489 GenAI-related protocols (75.7%) move from designer to GenAI, 92 (18.8%) are bidirectional, and 25 (5.1%) move from GenAI to designer; 2 protocols have no coded direction. Overall, students usually initiate and steer the exchange, and GenAI output is generally reworked rather than transferred literally.
Figure 4 grounds these patterns in a single deep case: twelve GenAI-engaged protocols by one student (S10), traced from the student’s own sketches and study models through successive GenAI iterations of a fractured-ice glacier-crevasse station. Across the semester this student’s GenAI use ranges over conceptual, formal, material, structural, programmatic, and interior decisions and over both designer-led and bidirectional interaction, with tools including Leonardo.Ai, Adobe Firefly, DALL·E, and ChatGPT. The most developed exchange is S10-D24, where a Leonardo.Ai model is conditioned on the student’s own study-model photographs (Training & Datasets) and DALL·E renders the crack pattern in stages—an instance of the designer steering, rather than only consuming, the generative tool. The cohort-wide set of GenAI-engaged protocols, one row per coded protocol across all students, is provided in Supplementary Figure S1.

4.3. Cognitive Breaking Points (BPs) Around Spatial Decisions

SQ3 asks how GenAI engagement is associated with shifts in the designer’s cognitive orientation. The empirical marker is the cognitive breaking point, defined as a clear departure from a previous frame, direction, or decision. Coding was conservative: a move had to show both a clear departure and enough evidence to support the classification.
Breaking points occur in 24 of 1107 protocols (2.2%). Reframing is the most common type (BP1, 13 protocols), followed by direction change (BP2, 10 protocols). Creative leap appears once, and backtracking does not appear. The 24 breaking points are distributed across 11 students, while 7 students have none (Figure 5). This is treated as a finding rather than a gap, because some trajectories remain linear or never cross the threshold for a breaking point. Temporally, breaking points cluster in the early and middle weeks, when project frames are still being formed, rather than at the studio jury sessions.
The relation between GenAI and breaking points is layered. By trigger coding, 9 of 24 breaking points are GenAI-related, 11 are linked to the designer’s own thinking, 2 are mixed, and 2 are ambiguous. The evidence is mostly indirect: only 5 breaking points include an explicit designer statement, while 17 rely on indirect inference and 2 remain ambiguous. A protocol can also contain a GenAI use code without being coded as having a GenAI-related trigger. Of the 24 breaking points, 14 carry a primary GenAI use code, 9 are trigger-coded as GenAI-related, and 7 fall in both groups. No breaking point was attributed to an external actor such as a jury member or peer; when the source could not be traced, as later illustrated by S2’s interim-jury direction change, the trigger was left ambiguous rather than inferred. The breaking points cluster mainly around conceptual and formal reframing, not around material, constructional, or structural decisions.
Two additional readings situate these shifts in the process. A supplementary reading assisted by an LLM classifies 20 of the 24 breaking points as systemic, meaning that they reshape later decisions; 2 are local and 2 are ambiguous. This breadth does not differ clearly by trigger type. At the process level, the corpus is weighted toward production: 748 protocols carry production as their primary phase code, compared with 271 as exploration and 88 as evaluation. A separate, overlapping view of production sub-types, counted wherever production appears as a primary or secondary code and therefore reported as multi-coded counts rather than as subdivisions of the 748, shows development and iteration (564 protocols) exceeding new idea generation (249 protocols). This suggests that much of the studio record concerns elaboration of an existing direction rather than the start of a new one.

4.4. Variation Across Students and Climate Contexts

Distributions by student are heterogeneous. Protocol counts range from 15 (S14) to 101 (S1), with a mean of 61.5. The GenAI-related share per student ranges from 27% (S14) to 60% (S17), close to the corpus rate weighted by protocol count (44%). Figure 6 summarizes this distribution. These shares depend on each student’s denominator; when a student has few protocols, a single unit carries more weight. Breaking points are also uneven. Eleven students have at least one, with a maximum of 4 for S4. Intensive GenAI use and breaking points do not always coincide: S13 has a high GenAI share but no coded breaking point, while S4 combines a high GenAI share with the largest number of breaking points. The unweighted student mean for GenAI share is 43.4%, and the median is 44.0%, indicating that the corpus rate is not driven only by the most prolific documenters.
Climate group distributions also vary descriptively. The Desert group has the highest GenAI-related share (52%) and the most breaking points (10), while the Tropical group has the lowest values on both measures (38%, 2 breaking points). Monsoon and Cold fall between these values. These differences are not interpreted as climate effects. The groups are small after the two consent-based exclusions (Cold 5, Desert 5, Monsoon 4, Tropical 4), and group patterns may reflect individual students. For example, the Desert group’s high GenAI share is largely attributable to S8 and S13, while its breaking-point count is mainly carried by S2 and S8.

4.5. Negative Cases and Robustness

Negative case analysis used metadata filters for rejected suggestions, low discourse and visual correspondence, and superficial single-use AU2. Researcher review then grouped the candidates into three types. Type A includes 10 protocols in which designers deliberately rejected a GenAI suggestion on grounds such as contextual fit, source checking, material criteria, or weak inspiration. The 10 Type A cases are the clearly reasoned subset of the 14 protocols metadata-coded as rejected suggestions in Section 4.2. Type B includes 4 protocols in which the tool could not produce the designer’s concept, or the discourse and visual output did not align enough to enter the design. Type C is an aggregate pattern rather than a single case: 50 protocols, about 10% of GenAI-related protocols, where AU2 appears as a one-off use without another GenAI use code. These cases show that GenAI engagement can be rejected, stall, or remain superficial. The analysis still covers the full corpus: the 618 protocols with no GenAI relation are coded as such, not excluded.

4.6. Anchor Cases

The distributions above are grounded in individual design moments through thirteen anchor cases. The cases were selected to test the working proposition from more than one direction: 5 support a GenAI-related shift in cognitive orientation, 5 refute or qualify that relation, and 3 remain ambiguous. Together, they cover 10 students and all four climate groups. One representative move from each set is summarized below; the full set of thirteen cases is reported in Supplementary S3.
One supporting case is S7 (Monsoon climate, Week 3). The designer pasted an entire project report into a text tool and asked for open conceptualization. The tool proposed the metaphor “Nebula” and linked it to stellar clouds as “the birth and death places of stars.” The designer did not copy this output. Instead, the metaphor was connected to the cloud-rich site and became a new design problem: a structure where visitors “can interact with the clouds” and move “in harmony with the rhythm of the clouds.” A latent cloud motif from the previous week became an explicit conceptual axis (BP1; trigger = GenAI-related, indirect inference; suggestion transformed and incorporated).
A refuting case is S11 (Tropical climate, Week 3). When the tool claimed that floating city culture emerged as adaptation to natural dynamics, the designer challenged the claim through independent research and argued that floating cities arose for economic reasons. The designer asked for sources, found them inadequate, and had the tool retract the claim. No frame shift occurred, and the suggestion was not incorporated.
An ambiguous case is S2 (Desert climate, Week 10), the only breaking point in the corpus with no GenAI evidence. A linear route schema that had remained stable from D08 to D16 changed in an interim jury deck into a centered radial form: a circular palm garden and water core surrounded by concentric program bands. The project frame remained stable, but the solution path changed. Because no GenAI was present and the source of the change could not be traced, the trigger was coded as ambiguous rather than attributed after the fact (BP2; trigger = ambiguous).
Together, the anchor cases show that GenAI engagement is not uniform. In the supporting cases, GenAI-related conceptualization or dialogue appears alongside reframing. In the refuting cases, designers reject the suggestion or judge the tool insufficient. In the ambiguous cases, the evidence does not permit attribution. These cases provide the qualitative basis for Section 5 assessment of the working proposition.

4.7. Confirmability Summary

Because no second canonical coder was used, confirmability was tracked through protocol-level bias notes functioning as a procedural reflexive journal [50]. These notes helped check for file name leakage, assistance anchoring, frequency inflation, speculation about unverifiable tools, and post hoc attribution of breaking points. The resulting pattern is conservative: GenAI is not counted as a design decision merely because it appears in the record, and ambiguous evidence is left ambiguous. The descriptive findings assembled here, including the concentration on formal decisions, the thinness of material, constructional, and structural registers, the dominance of transformed uptake, the direction of interaction weighted toward the designer, and the rarity of evidence for breaking points, provide the empirical basis for Section 5.

5. Discussion

Section 4 reported the empirical pattern: GenAI appears in 44.2% of the documented protocols (489/1107), but its presence is uneven across components, use types, students, and moments in the process. This section interprets that pattern in relation to the working proposition. The argument remains associational rather than causal. The corpus records process traces prepared for evaluation, not the lived design process itself; therefore, the percentages below should be read as documented tendencies, not as a complete measure of use.

5.1. Adjudicating the Working Proposition

The working proposition is partly supported. The evidence does not sustain the view that GenAI is merely a visual production tool: its documented use includes information gathering, conceptual prompting, critical dialogue, design testing, and the proposal of new possibilities. Suggestions are usually transformed before entering the design, and a small set of breaking points appears alongside activity related to GenAI. At the same time, the stronger claim that GenAI routinely operates as a design partner is not supported. Most engagement is led by the designer, oriented toward production or information, and only occasionally dialogic in a stronger sense. This pattern also resonates with recent accounts of digital tools in design education that frame such tools as supportive rather than autonomous actors [55]. The most defensible conclusion is therefore partial support: GenAI can move beyond visual production, but this engagement remains limited, episodic, and strongly shaped by designer judgment. In these terms, what the evidence supports is designer-steered dialogic engagement, not partnership.

5.2. Beyond the Visual Tool Assumption

The first subquestion is answered with qualification. GenAI engagement is broader than the formal surface alone, because information gathering appears across conceptual, programmatic, material, environmental, and constructional registers. Yet breadth should not be confused with depth. Much of this contact appears in the secondary layer, while the formal component remains the largest primary category. The subset related to GenAI with a primary component (n = 482) shows a modest shift: conceptual and material components rise, while programmatic and environmental components fall. AU2 paired with the formal component alone (AU2 × SC4) accounts for 125 of 227 AU2 protocols, or 55.1%. The result is therefore not a simple rejection of the visual tool assumption. GenAI touches multiple components, but its strongest center of gravity remains representational. Material, constructional, and structural components remain comparatively thin. This pattern does not demonstrate that individual students fell into a visualization trap, nor does it establish any effect on their design thinking, constructional reasoning, or structural judgment; it is a corpus-level imbalance in which documented GenAI engagement is more visible around imageable form than around tectonic reasoning.
The external expert evaluation offers an independent, outcome-level observation. Five senior experts, external to the studio team, evaluated the final projects using a structured rubric; students did not know the criteria. The evaluation was designed to assess project outcomes, not to validate the process coding or the distribution of AI use. On this rubric, structural and constructional registers were rated weakest and the conceptual register strongest. Because there is no direct correspondence between final-project quality and process-level AI use, this outcome-level pattern is reported only as indirect contextual background; it is not used to validate the protocol analysis, and no inference is drawn from final-project quality to the AI use types recorded in the process.

5.3. From Tool Use to Designer-Steered Dialogic Engagement

The second subquestion concerns whether GenAI is used as a tool, a partner, or something between the two. The balance of evidence is weighted toward tool use. AU2 (227 protocols, 46.4%) and AU1 (170 protocols, 34.8%) account for more than four-fifths of primary GenAI use codes. More dialogic uses, such as AU4 (32 protocols, 6.5%) and AU6 (39 protocols, 8.0%), remain marginal as primary AI use codes, though they appear more often in the secondary layer. Suggestion uptake is strongly transformative: 333 protocols rework a suggestion before incorporation, while only 2 adopt one verbatim. This shows active designer judgment, but it also limits the partner reading because the decisive transformation occurs on the designer’s side. Directionality points the same way: 370 of 489 protocols related to GenAI (75.7%) move from designer to GenAI, compared with 25 protocols (5.1%) moving from GenAI to designer. In this study, a design-partner reading does not imply autonomous agency or authorship by GenAI. It is reserved for protocols in which bidirectional exchange, transformative uptake, and a dialogic or propositional AI use, such as AU4 or AU6, coincide. Under this definition, the broadest version appears in 52 protocols (10.6%); a stricter version requiring the dialogic or propositional AI use to be primary appears in 27 protocols (5.5%). When tied to cognitive reorientation, the intersection narrows to 7 protocols (1.4%). These indicators—bidirectional exchange, transformative uptake, and dialogic or propositional AI use (AU4/AU6)—are treated as minimum evidential signs, not as sufficient evidence of partnership; individually they may also reflect routine filtering, revision, and reuse of tool outputs. We therefore do not describe the observed episodes as partnership in a symmetrical or agentic sense. The stronger pattern is better described as designer-steered dialogic engagement: GenAI may provoke, extend, or challenge a design move, but the designer initiates, selects, interprets, and transforms the output before it becomes part of the project. This is more than technical tool use, since the output is reworked rather than merely executed, but it falls short of co-creation or autonomous design agency, which would require stronger evidence of shared authorship, autonomous initiative, or design agency.

5.4. Cognitive Shifts and Evidentiary Limits

The third subquestion requires the most cautious interpretation. Cognitive breaking points are rare, appearing in 24 of 1107 protocols (2.2%). Reframing is the dominant type (13 protocols), followed by direction change (10), while creative leap appears once and backtracking does not appear. This rarity is not a weakness in itself; it reflects the conservative threshold used for identifying a genuine shift in orientation. The GenAI relation is also limited. Nine breaking points have a trigger related to GenAI, 14 carry a primary GenAI use code, and only 7 fall in both groups. Most evidence is indirect: only 5 breaking points include an explicit student statement, while 17 rely on inference from sequence, content overlap, and context. These figures rule out causal language. The anchor cases show that GenAI can accompany conceptual turns, as in S7, S8, and S4 (Supplementary S3), but other cases show rejection, withdrawal, or no traceable relation. The study therefore establishes a documented possibility, not a prevalent pattern. The breaking-point facet is therefore treated as the study’s most exploratory and least externally verifiable dimension. It informs SQ3 as evidence of documented possibility, but it is not used as a stand-alone basis for adjudicating the working proposition.

5.5. Opportunity, Risk, and Contextual Variation

The fourth subquestion is best answered as a balance rather than a verdict. The opportunities are visible where GenAI supports broad information gathering, selective transformation of suggestions, and early conceptual provocation. The risks appear where AU2 concentrates around formal components, where single shot visual use remains superficial, and where the interaction is so strongly led by the designer that the tool rarely challenges the initial frame. These tendencies are not uniform. GenAI association ranges from 27% to 60% across students, and 7 of 18 students show no breaking points. Climate group differences, such as the Desert group having the highest share related to GenAI and the Tropical group the lowest, are descriptive only. With four or five students per group, they should not be read as climate effects. The balance between opportunity and risk is therefore not a property of the tool alone; it depends on student practice, documentation, and the local studio context.

5.6. Dialogue with the Literature

These findings fit within design cognition theory, but they do not justify strong claims about mechanism. Episodes of reframing can be read through reflection-in-action [9], where a design situation talks back and prompts reconsideration. Here, GenAI sometimes helps make that response visible, but it does not replace the designer’s judgment. The 2.2% breaking point rate is lower than the 10% to 12% rate of critical moves reported in some linkographic studies [56], yet the comparison is not direct. This study uses retrospective documents, a higher threshold, and a different unit of analysis. The finding is therefore not that design movement is absent, but that only a small number of documented moves meet the stricter definition of a cognitive breaking point. The results also align with accounts of design as the redefinition of ill-defined problems [15], with GenAI occasionally participating in that redefinition.
Recent scholarship on design studios often approaches GenAI through outcome quality [6], student perception [7], pedagogical feedback [38], or proposed collaboration models [8]. This study makes a different contribution by following GenAI engagement at protocol level and across space-constituting components. That lens makes it possible to see not only where GenAI is active, but also where it remains weak: material, constructional, and structural reasoning are less visible than formal and conceptual work. The contribution is therefore both methodological and architectural. It offers a way to study GenAI in design education without reducing its role to either image production or generalized creativity.

5.7. Limitations

The study’s claims are bounded by the nature of the data. The corpus consists of curated evaluation artifacts, including reports, jury sheets, and partial GenAI logs selected by students themselves. It records traces of process rather than the process itself, and it may be shaped by post hoc rationalization, selective documentation, and survivorship. The reliability architecture supports procedural trustworthiness, but it does not establish direct access to lived cognition. The 44.2% GenAI association rate is therefore a documented floor, not a measure of total use. The visualization trap is used here as an interpretive, sensitizing concept rather than a measured variable: the study documents a corpus-level imbalance consistent with this risk but does not establish it as an effect on individual students.
Reliability is also uneven across facets. Space-constituting components show strong overlap for the two leading codes, so disagreement usually concerns which component is primary rather than whether a component is present. The AI use facet is more interpretive. GenAI presence was detected concordantly in 92% of audited units, but agreement on its use type was lower among units where GenAI was present: about 66% for the primary code and 78% for the two leading codes. The comparison across LLMs is useful as a further check, but it is weaker than human audit evidence because language models may share training priors. For this reason, AI use distributions should be read as broad tendencies rather than exact proportions.
Four further limits bound generalization. First, the data come from a single, primed studio in which GenAI use had been explicitly framed by the researcher. Second, the weekly documentary record leaves activity between deliverables largely invisible. Third, interpretation passes through a double mediation: candidate coding assisted by LLMs followed by researchers’ judgment, leaving residual anchoring risk even though the blind expert audit mitigates it. Fourth, segmentation was retrospective and category-guided. As described in Section 3.6, the protocol units were produced as an LLM-generated first pass and then reviewed and endorsed against explicit segmentation rules (Section 3.3); they were not independently re-segmented by a second coder, so unit boundaries were not subjected to a separate unitizing-reliability test. Because these counts form the denominator for the reported rates, the frequencies should be read as descriptive tendencies within a disciplined, rule-based segmentation, while the principal patterns—for example the concentration around formal decisions and the predominance of transformed over verbatim uptake—do not depend on exact unit counts. These limits confine the findings to an analytic, not statistical, generalization. What may transfer beyond this studio is the TRACE-SC framework and its codebook logic; the specific distributions reported here require testing in other studios and with richer process data. Finally, although only de-identified material was processed by external tools, future studies should include consent language that explicitly names third-party, AI-assisted processing.

6. Conclusions

This study examined how GenAI appears in the documented process traces of an undergraduate architectural design studio. It analyzed 1107 protocols from 18 consenting students through the TRACE-SC framework, which links GenAI activity to space-constituting components, AI use types, design process phases, and cognitive breaking points. The aim was not to measure all AI use or to infer causality, but to identify how documented GenAI engagement becomes visible in relation to architectural decisions.
The main conclusion is one of partial support. GenAI cannot be described only as a visual production tool: it appears in nearly half of the documented protocols, supports more than image generation, and is usually transformed before entering the design rather than copied directly. At the same time, the evidence does not support a strong claim of routine design partnership with AI. Most interaction is initiated and steered by students, and the clearest link between GenAI and cognitive reorientation remains rare and mostly indirect: breaking points appeared in 24 of 1107 protocols (2.2%), and 17 of the 24 cases relied on inference rather than explicit student statements. In this studio corpus, GenAI appeared less as an autonomous partner than as a situated design resource whose influence depended on how students questioned, interpreted, and reworked its outputs.
The architectural significance of the findings lies in their unevenness. GenAI engagement reaches several space-constituting components, but its strongest concentration remains around formal and representational decisions. Material, constructional, and structural registers are comparatively thin in the protocol analysis. This is a corpus-level imbalance in where documented GenAI engagement is most visible; it does not, on its own, establish an effect on students’ design thinking, constructional reasoning, or structural judgment. These findings raise a pedagogical question for AI-aware studio settings: not only whether students use GenAI, but which parts of architectural thinking the tool appears to support and which remain underdeveloped.
The study contributes a reusable analytical framework for examining design processes assisted by GenAI without reducing them to final outputs, student perceptions, or generic creativity claims. Its main contribution is methodological and architectural: it shows how GenAI activity can be traced across components of space, AI use types, and moments of cognitive shift. The findings are bounded by the single, primed studio context and by the documentary nature of the data. The reported percentages should be read as documented tendencies and as a floor of visible use, not as complete measures of student activity.
Future research should test the framework in studios with richer process evidence, including live interaction logs, observation, stimulated recall, and member checking. Comparative studies across institutions, cohorts, design briefs, and model ecosystems would clarify which patterns are local and which travel across contexts. In this sense, the study offers not a closed account of GenAI in architectural education, but a disciplined way to examine where it participates in design thinking, where it remains superficial, and where pedagogical attention is most needed.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/buildings16142827/s1. S1: the segmentation rules and protocol-unit definition. S2: the complete faceted codebook (SC1-SC7, AU1-AU6, DP1-DP3, BP1-BP4) with subcode definitions and construct-validity origins. S3: detailed records of the thirteen anchor cases. S4: reliability evidence (the independent blind audit, cross-LLM comparison, the AI-use agreement decomposition, and diagnostic Krippendorff’s alpha values). S5: the independent jury quality evaluation (rubric, rating distribution, and within-rater consistency). S6 (Figure S1): the cohort-wide visual corpus of GenAI-engaged protocols, one row per coded protocol across all students.

Author Contributions

Conceptualization, N.E. and D.G.O.; methodology, N.E. and D.G.O.; software, N.E.; validation, N.E. and D.G.O.; formal analysis, N.E.; investigation, N.E.; resources, N.E.; data curation, N.E.; writing, original draft preparation, N.E.; writing, review and editing, N.E. and D.G.O.; visualization, N.E.; supervision, D.G.O.; project administration, D.G.O. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Ethical approval for this study was granted by the Human Research Ethics Committee of TOBB University of Economics and Technology (decision no. E-27393295-100-56585, 12 March 2024).

Informed Consent Statement

Written informed consent was obtained from all 18 students whose data were included in the study. Two registered students did not provide written consent and were excluded from all analyses and reporting.

Data Availability Statement

The data presented in this study are available on request from the corresponding author. The data are not publicly available because they consist of student design records subject to consent and anonymization constraints. Students are identified throughout only by anonymized codes (S1–S18).

Acknowledgments

The authors thank TOBB University of Economics and Technology for granting permission to conduct this research within the Architectural Design Studio VII context, the students who consented to participate, the jury members and external experts who contributed to the project evaluation, and Murat Sönmez for conducting the independent blind coding audit. This paper is submitted in partial fulfilment of the requirements for the PhD degree at Istanbul Technical University. This study examines students’ use of generative AI (GenAI) in the design studio. Separately, large language models (LLMs) were used as research instruments at three points in the method. (1) For candidate coding (Section 3.6), Claude Opus 4.8 (Anthropic) generated reasoned candidate codes and metadata; the first author read these outputs, assessed them against the codebook, endorsed them where appropriate, and retained final authority over every code. (2) For cross-model reliability triangulation (Section 3.5), a second model, GPT-5.5 (OpenAI), independently coded all 18 students for comparison; an additional model evaluated in the pilot was excluded for documented reasons. (3) During manuscript preparation, an LLM was used for language editing. The authors reviewed and edited all outputs and take full responsibility for the content, coding decisions, and conclusions of this publication. The transfer of anonymized study material to external LLM tools during coding preparation is explicitly declared here.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Chaillou, S. Artificial Intelligence and Architecture: From Research to Practice; Birkhäuser: Basel, Switzerland, 2022. [Google Scholar]
  2. Bolojan, D.; Vermisso, E.; Yousif, S. Is language all we need? A query into architectural semantics using a multimodal generative workflow. In Proceedings of the 27th CAADRIA Conference; CAADRIA: Hong Kong, China, 2022; Volume 1, pp. 353–362. [Google Scholar]
  3. Leach, N. Architecture in the Age of Artificial Intelligence: An Introduction to AI for Architects; Bloomsbury Visual Arts: London, UK, 2022. [Google Scholar]
  4. Bank Stigsen, M.; Moisi, A.; Rasoulzadeh, S.; Schinegger, K.; Rutzinger, S. AI diffusion as design vocabulary—Investigating the use of AI image generation in early architectural design and education. In Proceedings of eCAADe 2023: Digital Design Reconsidered; eCAADe: Graz, Austria, 2023; Volume 2, pp. 587–596. [Google Scholar] [CrossRef] [Scilit]
  5. Tong, H.; Ülken, G.; Türel, A.; Şenkal, H.; Yağcı Ergün, F.; Güzelci, O.Z.; Alaçam, S. An attempt to integrate AI-based techniques into first year design representation course. In Cumulus Conference Proceedings: Connectivity and Creativity in Times of Conflict; Academia Press: Ghent, Belgium, 2023. [Google Scholar]
  6. Dong, Q.; He, J.; Li, N.; Wang, B.; Lu, H.; Yang, Y. Exploring the cognitive reconstruction mechanism of generative AI in outcome-based design education: A study on load optimization and performance impact based on dual-path teaching. Buildings 2025, 15, 2864. [Google Scholar] [CrossRef] [Scilit]
  7. Karadağ, D.; Ozar, B. A new frontier in design studio: AI and human collaboration in conceptual design. Front. Archit. Res. 2025, 14, 1536–1550. [Google Scholar] [CrossRef] [Scilit]
  8. Seyman Güray, T.; Uyan, B. Collaboration with generative artificial intelligence in the early stages of design studio: A model proposal. J. Archit. Sci. Appl. 2025, 10, 434–450. [Google Scholar] [CrossRef] [Scilit]
  9. Schön, D.A. The Reflective Practitioner: How Professionals Think in Action; Basic Books: New York, NY, USA, 1983. [Google Scholar]
  10. Cross, N. Designerly Ways of Knowing; Springer: London, UK, 2006. [Google Scholar]
  11. Lawson, B. How Designers Think: The Design Process Demystified, 4th ed.; Architectural Press: Oxford, UK, 2005. [Google Scholar]
  12. Visser, W. The Cognitive Artifacts of Designing; Lawrence Erlbaum Associates: Mahwah, NJ, USA, 2006. [Google Scholar]
  13. Eastman, C.M.; McCracken, W.M.; Newstetter, W.C. (Eds.) Design Knowing and Learning: Cognition in Design Education; Elsevier: Amsterdam, The Netherlands, 2001. [Google Scholar]
  14. Asimow, M. Introduction to Design; Prentice-Hall: Englewood Cliffs, NJ, USA, 1962. [Google Scholar]
  15. Simon, H.A. The structure of ill-structured problems. Artif. Intell. 1973, 4, 181–201. [Google Scholar] [CrossRef] [Scilit]
  16. Rittel, H.W.J.; Webber, M.M. Dilemmas in a general theory of planning. Policy Sci. 1973, 4, 155–169. [Google Scholar] [CrossRef] [Scilit]
  17. Dorst, K.; Cross, N. Creativity in the design process: Co-evolution of problem–solution. Des. Stud. 2001, 22, 425–437. [Google Scholar] [CrossRef] [Scilit]
  18. Dorst, K. The core of “design thinking” and its application. Des. Stud. 2011, 32, 521–532. [Google Scholar] [CrossRef] [Scilit]
  19. Cross, N. Descriptive models of creative design: Application to an example. Des. Stud. 1997, 18, 427–440. [Google Scholar] [CrossRef] [Scilit]
  20. Archer, L.B. Systematic Method for Designers; Council of Industrial Design: London, UK, 1965. [Google Scholar]
  21. Adams, R.S.; Atman, C.J. Characterizing engineering student design processes: An illustration of iteration. In Proceedings of the 2000 ASEE Annual Conference & Exposition; American Society for Engineering Education: St. Louis, MO, USA, 2000; Session 2330. [Google Scholar] [CrossRef] [Scilit]
  22. Gero, J.S. Design prototypes: A knowledge representation schema for design. AI Mag. 1990, 11, 26–36. [Google Scholar]
  23. Goldschmidt, G. The dialectics of sketching. Creat. Res. J. 1991, 4, 123–143. [Google Scholar] [CrossRef] [Scilit]
  24. Suwa, M.; Purcell, T.; Gero, J. Macroscopic analysis of design processes based on a scheme for coding designers’ cognitive actions. Des. Stud. 1998, 19, 455–483. [Google Scholar] [CrossRef] [Scilit]
  25. Semper, G. The Four Elements of Architecture and Other Writings; Mallgrave, H.F.; Herrmann, W., Translators; Cambridge University Press: Cambridge, UK, 1989. [Google Scholar]
  26. Frampton, K. Studies in Tectonic Culture: The Poetics of Construction in Nineteenth and Twentieth Century Architecture; Cava, J., Ed.; MIT Press: Cambridge, MA, USA, 1995. [Google Scholar]
  27. Frascari, M. The tell-the-tale detail. VIA 1984, 7, 23–37. [Google Scholar]
  28. Sekler, E.F. Structure, construction, tectonics. In Structure in Art and in Science; Kepes, G., Ed.; George Braziller: New York, NY, USA, 1965; pp. 89–95. [Google Scholar]
  29. Sönmez, M. Technique and tectonic concepts as theoretical tools in object and space production: An experimental approach to Building Technologies I and II courses. Buildings 2024, 14, 2866. [Google Scholar] [CrossRef] [Scilit]
  30. Engel, H. Structure Systems, 2nd ed.; Hatje Cantz: Ostfildern, Germany, 1997. [Google Scholar]
  31. Salvadori, M. Why Buildings Stand Up: The Strength of Architecture; W.W. Norton & Company: New York, NY, USA, 1980. [Google Scholar]
  32. Fahmy, A. A review on structural literacy in architectural education. Buildings 2025, 15, 4312. [Google Scholar] [CrossRef] [Scilit]
  33. Olgyay, V. Design with Climate: Bioclimatic Approach to Architectural Regionalism; Princeton University Press: Princeton, NJ, USA, 1963. [Google Scholar]
  34. Banham, R. The Architecture of the Well-Tempered Environment; Architectural Press: London, UK, 1969. [Google Scholar]
  35. Fathy, H. Natural Energy and Vernacular Architecture: Principles and Examples with Reference to Hot Arid Climates; University of Chicago Press: Chicago, IL, USA, 1986. [Google Scholar]
  36. Gero, J.S.; McNeill, T. An approach to the analysis of design protocols. Des. Stud. 1998, 19, 21–61. [Google Scholar] [CrossRef] [Scilit]
  37. Hay, L.; Duffy, A.H.B.; McTeague, C.; Pidgeon, L.M.; Vuletic, T.; Grealy, M. A systematic review of protocol studies on conceptual design cognition: Design as search and exploration. Des. Sci. 2017, 3, e10. [Google Scholar] [CrossRef] [Scilit]
  38. Wang, J.; Shi, Y.; Chen, X.; Lan, Y.; Liu, S. Teaching with artificial intelligence in architecture: Embedding technical skills and ethical reflection in a core design studio. Buildings 2025, 15, 3069. [Google Scholar] [CrossRef] [Scilit]
  39. Dortheimer, J.; Schubert, G.; Dalach, A.; Brenner, L.J.; Martelaro, N. Think AI-side the box! Exploring the usability of text-to-image generators for architecture students. In Proceedings of eCAADe 2023: Digital Design Reconsidered, 41st Conference on Education and Research in Computer Aided Architectural Design in Europe, Graz, Austria, 20–22 September 2023; Dokonal, W., Hirschberg, U., Wurzer, G., Eds.; eCAADe: Graz, Austria, 2023; Volume 2, pp. 567–576. [Google Scholar] [CrossRef] [Scilit]
  40. Iranmanesh, A.; Lotfabadi, P. Critical questions on the emergence of text-to-image artificial intelligence in architectural design pedagogy. AI Soc. 2025, 40, 3557–3571. [Google Scholar] [CrossRef] [Scilit]
  41. Purcell, A.T.; Gero, J.S. Design and other types of fixation. Des. Stud. 1996, 17, 363–383. [Google Scholar] [CrossRef] [Scilit]
  42. Wadinambiarachchi, S.; Kelly, R.M.; Pareek, S.; Zhou, Q.; Velloso, E. The effects of generative AI on design fixation and divergent thinking. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems; ACM: New York, NY, USA, 2024; pp. 1–18. [Google Scholar] [CrossRef] [Scilit]
  43. Doshi, A.R.; Hauser, O.P. Generative AI enhances individual creativity but reduces the collective diversity of novel content. Sci. Adv. 2024, 10, eadn5290. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Guba, E.G.; Lincoln, Y.S. Competing paradigms in qualitative research. In Handbook of Qualitative Research; Denzin, N.K., Lincoln, Y.S., Eds.; SAGE Publications: Thousand Oaks, CA, USA, 1994; pp. 105–117. [Google Scholar]
  45. Creswell, J.W.; Plano Clark, V.L. Designing and Conducting Mixed Methods Research, 3rd ed.; SAGE Publications: Thousand Oaks, CA, USA, 2018. [Google Scholar]
  46. Tashakkori, A.; Teddlie, C. (Eds.) SAGE Handbook of Mixed Methods in Social & Behavioral Research, 2nd ed.; SAGE Publications: Thousand Oaks, CA, USA, 2010. [Google Scholar]
  47. Ericsson, K.A.; Simon, H.A. Protocol Analysis: Verbal Reports as Data, Revised ed.; MIT Press: Cambridge, MA, USA, 1993. [Google Scholar]
  48. Mercer, J. The challenges of insider research in educational institutions: Wielding a double-edged sword and resolving delicate dilemmas. Oxford Rev. Educ. 2007, 33, 1–17. [Google Scholar] [CrossRef] [Scilit]
  49. Ranganathan, S.R. Colon Classification; Madras Library Association: Madras, India, 1933. [Google Scholar]
  50. Lincoln, Y.S.; Guba, E.G. Naturalistic Inquiry; SAGE Publications: Beverly Hills, CA, USA, 1985. [Google Scholar]
  51. Hammersley, M. What’s Wrong with Ethnography? Methodological Explorations; Routledge: London, UK, 1992. [Google Scholar]
  52. McDonald, N.; Schoenebeck, S.; Forte, A. Reliability and inter-rater reliability in qualitative research: Norms and guidelines for CSCW and HCI practice. Proc. ACM Hum.-Comput. Interact. 2019, 3, 72. [Google Scholar] [CrossRef] [Scilit]
  53. Bakeman, R.; Gottman, J.M. Observing Interaction: An Introduction to Sequential Analysis, 2nd ed.; Cambridge University Press: Cambridge, UK, 1997. [Google Scholar] [CrossRef] [Scilit]
  54. Krippendorff, K. Content Analysis: An Introduction to Its Methodology, 2nd ed.; SAGE Publications: Thousand Oaks, CA, USA, 2004. [Google Scholar]
  55. Kaylak, E.; Kurt, S.; Saymanlıer, A.M. The role of VR in supporting body-centered phenomenology in interior design education. Buildings 2026, 16, 250. [Google Scholar] [CrossRef] [Scilit]
  56. Goldschmidt, G. Linkography: Unfolding the Design Process; MIT Press: Cambridge, MA, USA, 2014. [Google Scholar]
Figure 1. Methodological workflow of the study (TRACE-SC).
Figure 1. Methodological workflow of the study (TRACE-SC).
Buildings 16 02827 g001
Figure 2. Distribution of primary space-constituting component codes across the full corpus (n = 1099) and the GenAI-related subset (n = 482). Bar labels show counts (n); the y-axis shows the share of primary codes (%).
Figure 2. Distribution of primary space-constituting component codes across the full corpus (n = 1099) and the GenAI-related subset (n = 482). Bar labels show counts (n); the y-axis shows the share of primary codes (%).
Buildings 16 02827 g002
Figure 3. Co-occurrence heat map of GenAI use (AU) by space-constituting component (SC), primary codes, within the GenAI-related subset. The seven GenAI-related protocols with no primary SC are omitted, so the cells sum to 482 of 489 protocols. The outlined cell marks the largest concentration, AU2 × SC4 (visual generation on the formal component; 125 protocols).
Figure 3. Co-occurrence heat map of GenAI use (AU) by space-constituting component (SC), primary codes, within the GenAI-related subset. The seven GenAI-related protocols with no primary SC are omitted, so the cells sum to 482 of 489 protocols. The outlined cell marks the largest concentration, AU2 × SC4 (visual generation on the formal component; 125 protocols).
Buildings 16 02827 g003
Figure 4. Worked example: twelve GenAI-engaged protocols by one student (S10), the study’s deep case. Each row is one coded protocol; panels run left to right from the student’s own sketch or study model through successive GenAI iterations. Labels give the protocol code (S10-D##; e.g., S10-D24) and primary coding.
Figure 4. Worked example: twelve GenAI-engaged protocols by one student (S10), the study’s deep case. Each row is one coded protocol; panels run left to right from the student’s own sketch or study model through successive GenAI iterations. Labels give the protocol code (S10-D##; e.g., S10-D24) and primary coding.
Buildings 16 02827 g004
Figure 5. Breaking points across the 18-student census and studio sessions. Cell shading shows the number of protocols a student produced in each studio session; markers locate the 24 breaking points by type, with adjacent markers indicating more than one breaking point in the same session.
Figure 5. Breaking points across the 18-student census and studio sessions. Cell shading shows the number of protocols a student produced in each studio session; markers locate the 24 breaking points by type, with adjacent markers indicating more than one breaking point in the same session.
Buildings 16 02827 g005
Figure 6. GenAI-related protocol share for each student across the census of 18 students, sorted by share. The dashed line shows the corpus mean weighted by protocol count (44%). Bar shading reflects the same share; a darker shade indicates a higher GenAI-related share.
Figure 6. GenAI-related protocol share for each student across the census of 18 students, sorted by share. The dashed line shows the corpus mean weighted by protocol count (44%). Bar shading reflects the same share; a darker shade indicates a higher GenAI-related share.
Buildings 16 02827 g006
Table 1. Discriminating questions used to separate neighboring space-constituting components (SCs) during coding. Full per-component definitions and boundary notes are given in Supplementary S2.
Table 1. Discriminating questions used to separate neighboring space-constituting components (SCs) during coding. Full per-component definitions and boundary notes are given in Supplementary S2.
BoundaryDiscriminating Question
Conceptual vs. FormalIs the unit about semantic intent, or geometric/volumetric realization?
Formal vs. StructuralIs the unit about shape and massing, or load transfer and load-bearing logic?
Constructional vs. StructuralIs the unit about joining and making, or load paths and spanning logic?
Material vs. ConstructionalIs the unit about what the project is made of, or how parts are assembled?
Environmental vs. Conceptual/MaterialIs external climate/site data translated into spatial response, or is it symbolic meaning or material selection?
Programmatic vs. FormalIs the unit about what happens where, or the geometric configuration of space?
Table 2. Facets, subcodes, and codes per unit of the codebook; codes per unit is the number of subcodes a coder may assign to a single protocol on that facet (the permitted range).
Table 2. Facets, subcodes, and codes per unit of the codebook; codes per unit is the number of subcodes a coder may assign to a single protocol on that facet (the permitted range).
FacetWhat It CapturesSubcodesCodes per Unit
SCSpace-constituting
decision domain
SC1 Conceptual; SC2 Material; SC3 Construction; SC4 Formal; SC5 Structural; SC6 Programmatic; SC7 Environmental1–3
AUGenAI use in the moveAU1 Information gathering; AU2 Visual generation;
AU3 Design testing; AU4 Critical dialogue;
AU5 Scenario building; AU6 Proposing new possibilities
0–3
DPDesign process phaseDP1 Exploration; DP2 Production; DP3 Evaluation1–2
BPCognitive breaking pointBP1 Reframing; BP2 Direction change;
BP3 Creative leap; BP4 Backtracking
0–1
Table 3. Coding reliability evidence from the independent blind expert audit and the cross-LLM comparison (raw percent agreement, not chance-corrected).
Table 3. Coding reliability evidence from the independent blind expert audit and the cross-LLM comparison (raw percent agreement, not chance-corrected).
SourceSC
Primary %
AU
Primary %
DP
Primary %
BP
Presence %
SC
Top-Two %
AU
Top-Two %
DP
Top-Two %
Deep single case
(S10, 92 units)
63.070.768.5n/a89.175.091.3
Scattered sample
(194 units, 16 students)
70.184.572.2n/a93.890.792.3
LLM to LLM comparison
(18 students, 1013 protocols)
71747996n/an/an/a
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Eyce, N.; Ozer, D.G. TRACE-SC: A Protocol-Based Framework for Mapping Generative AI in Space-Constituting Architectural Design Decisions. Buildings 2026, 16, 2827. https://doi.org/10.3390/buildings16142827

AMA Style

Eyce N, Ozer DG. TRACE-SC: A Protocol-Based Framework for Mapping Generative AI in Space-Constituting Architectural Design Decisions. Buildings. 2026; 16(14):2827. https://doi.org/10.3390/buildings16142827

Chicago/Turabian Style

Eyce, Nihat, and Derya Gulec Ozer. 2026. "TRACE-SC: A Protocol-Based Framework for Mapping Generative AI in Space-Constituting Architectural Design Decisions" Buildings 16, no. 14: 2827. https://doi.org/10.3390/buildings16142827

APA Style

Eyce, N., & Ozer, D. G. (2026). TRACE-SC: A Protocol-Based Framework for Mapping Generative AI in Space-Constituting Architectural Design Decisions. Buildings, 16(14), 2827. https://doi.org/10.3390/buildings16142827

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop