Next Article in Journal
Research on Evolutionary Game and Implementation Strategies for Promoting Near-Zero Energy Building Technologies
Next Article in Special Issue
The Role of Artificial Intelligence in Architecture and Interior Design
Previous Article in Journal
Analysis of Temperature Field Characteristics of Highway Tunnels During Fire
Previous Article in Special Issue
Human–AI Collaborative Design in Architectural Studios: Evaluating Paradigm Shifts Across the Six Stages of the Design Process
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Leveraging Generative AI for High-Fidelity 360° Spatial Images: Methodological Validation for Use as Experimental Stimuli

Department of Interior Architecture and Built Environment, Yonsei University, Seoul 03722, Republic of Korea
*
Author to whom correspondence should be addressed.
Buildings 2026, 16(9), 1679; https://doi.org/10.3390/buildings16091679
Submission received: 23 March 2026 / Revised: 17 April 2026 / Accepted: 22 April 2026 / Published: 24 April 2026
(This article belongs to the Special Issue Artificial Intelligence in Architecture and Interior Design)

Abstract

Despite its efficiency, the structural integrity and geometric accuracy of artificial intelligence (AI)-generated imagery used in environmental psychology experiments have not been sufficiently validated. This study investigated the methodological validity and substitutability of generative AI-generated 360° images as experimental stimuli for indoor environmental research. Using a three-stage framework, we generated base panoramas with controlled structural parameters, integrated greenery via AI-based inpainting, and conducted multifaceted validation through objective quality metrics and expert assessments. Quantitative results confirmed high technical integrity, indicating that structural distortions at panoramic stitching points were effectively minimized. Furthermore, the AI-generated stimuli maintained stable visual quality across varying greenery densities. Expert evaluations confirmed that the AI-driven approach significantly outperforms conventional 3D modeling, particularly in terms of presence and realism. By achieving high usability and spatial integrity scores, we established a novel standard for employing generative AI to create high-fidelity virtual environments for architectural and psychological research.

1. Introduction

In contemporary indoor environmental research, space is recognized not merely as a physical backdrop but as a complex environmental stimulus—defined in this study as a controlled visual setting designed to evoke specific psychological and physiological responses [1,2]. While natural elements like greenery are essential variables for emotional restoration, their application in virtual reality (VR) experiments has been frequently hindered by the high technical costs and extensive time required for conventional 3D rendering [3,4]. This study addresses the methodological constraints inherent in the creation of conventional stimuli by proposing a generative AI-based workflow that ensures both efficiency and scientific rigor.
Currently, 360° panoramic imagery is a standard format for providing immersive spatial stimuli [5]. However, traditional production processes—involving sequential modeling, texturing, and lighting—limit the diversity of experimental conditions and make it difficult to maintain visual consistency while rapidly altering specific environmental variables [6]. Although several studies have attempted to utilize AI in this field [7,8,9,10,11], empirical validation of whether AI-generated panoramas can maintain the architectural rigor and consistency required for scientific experiments, compared with conventional 3D rendering, remains insufficient.
Building on this objective, although previous research has primarily applied generative AI for visual ideation and design exploration, its application as a controlled experimental tool in spatial research remains insufficiently examined. We argue that, in experimental contexts, spatial images are not merely visual representations but function as stimuli that require high levels of reliability, reproducibility, and perceptual validity. In contrast to exploratory design contexts where visual inspiration is prioritized, experimental settings demand that every visual element be strictly controlled to ensure that participant responses are triggered by intended variables rather than unintended AI artifacts. However, the extent to which AI-generated 360° panoramic images can satisfy these requirements has not been thoroughly validated.
To bridge this gap, the primary objective of this research is to validate the effectiveness of AI-generated 360° panoramic images as experimental stimuli for indoor environmental psychology. We focus on three key dimensions: (1) evaluating the precision of text-to-image (T2I) models for architectural environments, (2) examining the capacity for systematic greenery manipulation across multiple levels and configurations, and (3) establishing methodological validity through a multi-layered verification framework. In this framework, diffusion-based T2I models are adopted as the most effective technical instruments to achieve high fidelity and precise variable control; concurrently, their performance is rigorously verified beyond technical Image Quality Assessment (IQA) metrics through structured perceptual evaluations by architecture and design experts. By adopting the igroup Presence Questionnaire (IPQ) and a multi-dimensional quality assessment framework derived from generative AI literature, we verify whether these AI-generated 360° environments possess the architectural integrity, greenery fidelity, and overall usability required for immersive psychological experiments.
The academic contribution of this study lies in its shift from using AI for exploratory design ideation to systematic stimulus production for controlled experiments. We present a methodology that drastically reduces the time and labor typically required for manual 3D modeling while maintaining the scientific rigor—specifically reliability, reproducibility, and strict variable control—essential for empirical research. Crucially, the absence of such rigor leads to unreliable experimental comparisons and misattribution of perceptual responses, compromising the credibility of spatial research. By overcoming these limitations and the inefficiencies of traditional methods, this study validates the potential for researchers to develop diverse, high-quality 360° experimental environments with significantly improved efficiency.

2. Literature Review

2.1. Generative AI for Spatial Image Generation and Experimental Visual Stimuli

Recent studies on architectural and interior designs have increasingly explored the use of generative AI as a tool for producing spatial images during the early stages of design development, particularly for concept exploration and visual ideation [12]. Unlike conventional visualization workflows that rely on manual modeling and rendering, generative AI enables rapid image synthesis based on abstract inputs, allowing designers to externalize, develop, and refine spatial concepts at an early conceptual level without the need for labor-intensive manual processes [13].
Text-to-image generation models have been widely adopted in architectural research owing to their ability to translate linguistic descriptions into visual representations of space. Existing studies have demonstrated that prompt-based interactions enable designers to articulate abstract spatial attributes, including atmosphere, materiality, lighting conditions, and stylistic references, through natural language, which are then synthesized into spatial images using diffusion-based models [12,14]. This approach supports exploratory design processes by enabling rapid iteration and comparison of alternative spatial expressions without requiring predefined geometric models. From an experimental perspective, this property also allows the explicit specification and controlled variation in spatial attributes through structured linguistic inputs.
By contrast, image-to-image (img2img) generation approaches produce spatial variations by transforming existing reference images while partially preserving their underlying structure. Horvath and Pouliou [13] report that img2img workflows, such as textures, lighting artifacts, and compositional biases, tend to retain unintended visual elements embedded in the source image. These inherited visual features may hinder experimental control and limit the clarity of spatial manipulation, particularly when systematic variation in visual conditions is required; this limitation becomes critical in experimental scenarios, where the isolation of specific variables requires precise and independent control over visual attributes.
Several studies have emphasized that generative AI-based spatial images are shaped by prompt content and the selection of visual descriptors related to material surfaces, volumetric composition, scale perception, and stylistic references. Jeong et al. [15] demonstrated that fine-tuned diffusion models can learn domain-specific visual vocabularies, including interior materials, color palettes, and furniture styles, through keyword-based training, thereby enabling more consistent control over spatial appearance. Similarly, Paananen, Oppenlaender and Visuri [12] reported that architecture students often employ biophilic and nature-inspired descriptors, such as plants, wood, and organic forms, in text prompts to support imaginative exploration and creative ideation during early-stage design tasks.
From a technical perspective, diffusion-based generative models have become the dominant framework for spatial image synthesis owing to their stability and ability to produce coherent global structures [7,16]. Unlike earlier generative models, diffusion models iteratively refine images through probabilistic denoising processes that recover structured signals by progressively removing Gaussian noise, thereby improving structural consistency and reducing visual artifacts [17]. This characteristic has expanded the applicability of generative AI from artistic image production to spatially structured architectural representations [16]. While diffusion models improve overall image quality and structural coherence, the level of controllability in experimental situations is determined not just by the model itself, but also by how inputs are structured and limited. In this context, “text-based generation” refers to the input modality that provides spatial descriptions, whereas the diffusion model refers to the generative architecture that synthesizes the images. These two dimensions are conceptually distinct: one governs how spatial intent is specified, and the other governs how images are computationally generated.
Despite these advancements, existing architectural applications of generative AI primarily focus on ideation and visual communication rather than on controlled experimental usage. Previous works have focused on creative support and stylistic exploration, where outputs are typically judged in terms of visual plausibility instead of experimental rigor. In such contexts, spatial images must maintain consistency across conditions while enabling precise manipulation of target variables. Accordingly, the extent to which AI-generated spatial images can ensure reliability, reproducibility, and perceptual validity under these conditions has not been thoroughly investigated, despite the fact that these are critical requirements for valid experimental comparison.

2.2. 360° Panorama Images and the Use of Generative AI

360° panoramic images provide continuous visual coverage of spatial environments and have been widely employed in spatial and environmental research as immersive and ecologically valid experimental stimuli [5]. By enabling users to perceive an environment from all directions, panoramic imagery offers a closer approximation of real-world spatial experiences than conventional perspective images.
Previous studies on virtual and immersive environments have demonstrated that panoramic visual stimuli significantly influence the visual attention, arousal responses, and perceptual engagement of users. Kim and Lee [3] showed that 360° virtual retail environments elicit gaze patterns and physiological responses comparable to those observed in physical spaces. Similarly, Lee et al. [4] reported that overlaying 3D models onto panoramic images within augmented reality (AR) environments facilitates effective comparison of architectural design alternatives, supporting the validity of panoramic imagery as a spatial evaluation medium.
Despite these advantages, the production of panoramic images has conventionally relied on resource-intensive workflows, including 3D modeling, 360° camera capture, panoramic stitching, and manual post-processing [4,10,18]. These processes require specialized expertise and substantial labor, limiting scalability and constraining the systematic manipulation of visual variables while maintaining consistent spatial structure across experimental conditions, which is significant for controlled experimental research.
To address these limitations, generative AI has recently been employed for the automated generation and modification of 360° spatial images. Diffusion-based generative approaches have demonstrated the capacity to synthesize panoramic imagery with improved global spatial coherence, enabling the generation of multiple visual conditions within a shared spatial framework [17]. This capacity indicates the potential of generative AI to aid in achieving both efficiency and experimental control in spatial stimulus generation.
In parallel, recent improvements in computer vision have led to the design of text-based panoramic image generation techniques which deal with spatial continuity and geometric distortion phenomena of 360° images. Recent studies have proposed generation strategies that enhance consistency across different viewpoints, enabling panoramic images to maintain coherent spatial structures throughout the full field of view. For instance, Kalischek et al. [9] proposed a cube map-based generation approach that improves the continuity between adjacent views and reduces visible seams and distortions. Zhang et al. [8] further demonstrated that integrating the global spatial context with localized perspective information supports more stable and spatially coherent panoramic environments. In addition, Ni et al. [11] demonstrated that parameter-efficient fine-tuning methods enable the generation of high-quality panoramic images with limited model adaptation, indicating the feasibility of producing controlled spatial stimuli without extensive computational resources.
Despite these technical advances, most studies have primarily evaluated generative performance in terms of visual realism, distortion reduction, and structural consistency. However, these criteria do not directly address the requirements of experimental research, where spatial stimuli must maintain consistency across conditions while enabling precise manipulation of target variables. As a result, the empirical validation of the AI-generated panoramic images as such controlled experimental stimuli, especially compared with traditional panoramas or 3D-rendered surroundings, remains insufficiently examined. In contrast, this study investigates whether AI-generated panoramic photos can support controlled modification of environmental variables while retaining spatial consistency between conditions, allowing them to be used as legitimate experimental stimuli.

3. Method: Generation of 360° Panoramic Spatial Images Using Generative AI

3.1. Terminology and Definitions

To ensure methodological rigor and address potential ambiguities, this study establishes clear operational definitions for its core concepts. Table 1 summarizes the distinction between the conceptual environment, the generated visual artifacts, and the technical processes involved. By separating the ‘environmental stimulus’ (the conceptual intent) from the ‘spatial stimulus’ (the visual output), this framework provides a structured approach to validating generative AI as a reliable tool for indoor environmental research.

3.2. Tool Selection for 360° Panoramic Spatial Images

Skybox AI (Model 4, Blockade Labs, Indianapolis, IN, USA), a generative AI service that produces 360° panoramic images using text prompts, was selected as the primary stimulus generation tool. This tool was selected due to its capability to directly generate 360° equirectangular panoramic images with stable global coherence, which is essential for maintaining spatial continuity across experimental conditions. It has been increasingly adopted as a background generation tool in virtual production pipelines and extended reality content development [19]. While text-based panoramic generation offers superior visual coherence despite potential spatial ambiguities, img2img-based workflows are unsuitable for controlled VR environments due to persistent source image artifacts; thus, this study exclusively employed text-driven generation, supplemented by post-processing to ensure structural consistency and minimize noise.
Although Skybox AI stably generates spatial structures, its ability to place discrete objects—like plants—while maintaining consistent lighting and scale remains limited. Specifically, the “Remix” function does not support fixed random seeds, rendering it challenging to precisely control object placement across experimental conditions. To address this, vegetation elements were incorporated via post-processing. The Gemini 3 Flash Image model (codenamed Nano Banana 2; Google, Mountain View, CA, USA), accessed via the Envato platform (Envato, Melbourne, VIC, Australia), was employed to insert elements into the generated panoramas. This tool supports visually coherent object insertion within 360° environments while preserving lighting consistency and spatial realism, ensuring suitability for experimental stimulus construction.

3.3. Condition Setting for 360° Panoramic Spatial Image Generation

3.3.1. Derivation of Architectural Characteristics of Public Areas in Mixed-Use Facilities

Before stimulus production, architectural characteristics of domestic mixed-use facilities were examined to establish a neutral spatial typology. Specifically, we conducted an extensive analysis involving field surveys, photographic documentation, spatial analysis, and a comprehensive review of online spatial data focused on public common areas within domestic mixed-use facilities. These characteristics support experimental neutrality and generalizability, rather than association with specific buildings. The derived characteristics and analytical categories are summarized in Table A1.

3.3.2. Construction of Controlled Stimulus Environment

To establish a baseline environment, we developed three versions by varying material conditions while maintaining consistent spatial structures and lighting: white, light wood, and dark wood (Table A2). The selection of these three specific material-tone conditions was directly informed by the results of our prior field analysis, which identified them as the most representative and prevalent finishes in the common areas of domestic mixed-use facilities. The white version uses light and white-toned finishes common in domestic mixed-use facilities, with neutral configurations to avoid specific program interpretation. The light wood version incorporates warm wood finishes while restricting color range and application to prevent wood from becoming the dominant visual element. The dark wood version reflects a premium atmosphere, while the material hierarchy and furniture quantity are controlled to avoid interpretation as a hotel lounge. Spatial structure, circulation, and lighting were controlled across all stimuli (Table 2). This strategy minimized unintended AI-generated scene transitions and suppressed non-essential visual differences.

3.4. Prompt Writing Strategy and Optimization

Trial generations were conducted to examine Skybox AI’s language processing characteristics. Skybox AI assigns substantial weights to initial prompt conditions and operates with strict input length constraints. Prompts were constructed using ~1200 characters, placing core architectural elements—void configuration, lighting, and slab geometry—at the beginning, followed by material and atmospheric descriptions. When large public spaces were specified, the system automatically generated sky elements; thus, constraints were repeated throughout prompts to reinforce exclusion. A photorealistic rendering configuration (realistic, m3 advanced) was selected to approximate actual visual perception.
Negative prompts were limited to 480 characters. Applying identical negative text across all stimuli was insufficient, as uniform criteria either suppressed material characteristics or failed to block unintended interpretations. Therefore, negative prompts used a shared base with material-specific adjustments. This process was refined through repeated visual inspection, adding exclusion items only when necessary and removing items affecting spatial coherence.
To ensure methodological clarity and reproducibility, a representative prompt used for generating the base panoramic environments is presented below. The prompt design follows a hierarchical structure: it begins with spatial configuration and geometric constraints to stabilize the global structure, followed by material specifications and atmospheric details to refine the visual characteristics.
  • Positive Prompt: “A large central atrium inside a contemporary mixed-use commercial complex. The viewpoint is at human eye level on the basement floor. A central vertical void defines the atrium, with two to three upper levels visible around the perimeter. Upper floors form continuous horizontal slabs with consistent thickness, clearly supported by aligned columns. Materials follow a refined beige and light-brown commercial palette. Railings are continuous clear glass with slim light-toned metal handrails. The atrium is fully enclosed below ground with no skylights, exterior views, or visible sky. Lighting is entirely artificial. The space is empty, with no people, vegetation, or decorative installations”.
  • Negative Prompt: “red, blue, green vivid colored paintings, paintings, people, crowds, silhouettes, plants, greenery, trees, sky, skylight, translucent ceiling, glass roof, exterior view, windows to outside, black doors…”.

3.5. Spatial Image Generation and Validation

3.5.1. Evaluation Metrics

To evaluate the quality of the AI-generated 360° indoor panoramas, we employed five metrics covering semantic alignment, geometric integrity, and structural continuity at the seams. To account for the omnidirectional nature of panoramic images, each panorama was divided into 20 normal field-of-view (NFoV) images. Metrics were calculated for each view, and the average score was used as the final indicator.
The suitability of the generated images was evaluated using Contrastive Language-Image Pre-training (CLIP) scores to quantitatively evaluate the semantic alignment between prompts and generated images using a contrastive CLIP model [20]. Higher scores indicated better alignment with the user’s intent. The structural integrity of the 360° panoramic images was evaluated using the straight-line deviation (SLD) metric, which quantifies geometric distortion by calculating the distance between detected edge pixels and their corresponding idealized straight lines. Based on the “straight-line constraint” principle [21], we measured the deviation of edge pixels from idealized Hough-transformed lines. Lower values indicated superior structural preservation. Continuity of the generated panoramas was assessed using seam mean squared error (MSE), which measures the average squared difference between pixel intensities at the junction, peak signal-to-noise ratio (PSNR) to assess the ratio between the maximum possible power of a signal and the power of corrupting noise, and structural similarity index (SSIM), which evaluates the degradation of structural information by modeling the perceived change in luminance and contrast [22]. Seam MSE and PSNR assess pixel-level discontinuity at the interface [23], where high PSNRs and low MSEs signify seamless transition and high-intensity consistency. SSIM evaluates perceptual continuity by modeling human visual sensitivity to luminance, contrast, and structural patterns [24]. Unlike pixel-based error metrics, SSIM captures the preservation of local structural information across junctions; values approaching 1.0 indicate the visual integrity of the 360° representation.

3.5.2. Results of Spatial Images Generation and Evaluation

The white, light wood, and dark wood stimuli were generated as base environments reflecting only architectural and material conditions. All stimuli were 360° panoramic images with similar spatial types, structural principles, lighting, circulation, and viewpoints. The quantitative results are presented in Table 3.
The dark wood variant exhibited the highest geometric precision (SLD: 0.040%), whereas the light wood variant showed higher distortion (0.093%). Although light wood (0.359) exhibited superior semantic fidelity, all results exceeded the 0.3 threshold, indicating high semantic alignment. This threshold is established as a benchmark where CLIP scores demonstrate a strong correlation with human perceptual judgment, confirming that the generated imagery accurately reflects the input prompts [25]. However, seam metrics revealed significant performance gaps. The white variant showed acceptable geometric stability (SLD: 0.050%) but high seam discontinuity (MSE: 328.190). Conversely, light wood exhibited the most refined stitching quality (MSE: 104.820, PSNR: 27.930). Based on SSIM, dark wood (0.847) exhibited the highest visual and structural continuity. The dark wood version was selected for the final experimental control. Despite the light wood variant’s superior semantic alignment and surface continuity, the dark wood version achieved exceptional geometric precision (SLD: 0.040%), which is vital for accurate spatial reconstruction and minimizing visual distortion variables. With high perceptual integrity (SSIM: 0.847), the dark wood variant provides the most stable, mathematically consistent foundation for spatial analysis and subsequent studies.

4. Systematic Manipulation of Greenery on 360° Panoramic Spatial Images Using Generative AI

4.1. Criteria for Greenery Insertion

After generating 360° interior scenes, indoor greenery was introduced using a post-processing workflow. Viewpoints, camera positions, equirectangular geometry, architectural layouts, materials, lighting, and spatial proportions remained constant across all versions. Vegetation consisted of low-salience foliage species, primarily Epipremnum aureum, with limited use of Dracaena. Greenery was integrated exclusively into existing architectural surfaces, such as walls, columns, and balustrades, while freestanding, floor-based, and ceiling-mounted elements were excluded. Due to generative AI inpainting constraints, the final version was rendered at 4K resolution.

4.2. Greenery Manipulation Strategies for Comparative Analysis

The five configuration types were derived based on primary architectural surfaces (walls, columns, and balustrades), which differ in spatial prominence and visual exposure within the panoramic field of view. These configurations were designed to examine how the placement of greenery across different architectural elements influences spatial perception while preserving a consistent underlying spatial structure. They represent distinct spatial distribution patterns of greenery, enabling comparison between localized and distributed interventions, while the three density levels allow incremental assessment of perceptual impact without altering the underlying spatial structure. The three greenery density levels (low, medium, and high) represent incremental increases in visual exposure, enabling the assessment of perceptual responses across a controlled range of greenery intensity under uniform spatial conditions. The overall experimental framework, including the fixed spatial conditions and controlled greenery manipulations, is summarized in Figure 1.
Accordingly, five configurations were defined based on primary surface location: wall only (V1), column only (V2), balustrade only (V3), column–balustrade (V4), and wall–column–balustrade systems (V5). Independent of the version-based greenery distribution stimuli, a separate set of panoramic images was generated to examine the perceptual effects of greenery intensity. Greenery intensity was operationalized at low, medium, and high levels, with application locations limited to column surfaces and inner balustrade edges. Differences between levels were achieved solely by adjusting vertical reach, continuity, and perceptual prominence. Low levels represent minimal greenery, characterized by limited application on selected columns’ lower sections and a single thin continuous linear band along balustrades. Medium levels increase vertical reach to mid-column height and introduce continuous linear bands across visible balustrade segments. Finally, high levels further intensify repetition across columns and balustrades without full spatial enclosure, extending vertically toward the ceiling (while stopping short of it) and increasing balustrade continuity. Greenery density levels were defined as follows: low (limited vertical application and minimal continuity), medium (moderate vertical extension and continuous distribution), and high (extended vertical reach with repeated distribution across surfaces).

4.3. Greenery Integration and Validation

4.3.1. Evaluation Framework and Methodology

To evaluate the quality of the image outputs modified by generative AI, a comprehensive assessment framework was established using both full-reference (FR) metrics that compare the output to a ground-truth image and no-reference (NR) IQA metrics that evaluate quality without a baseline. Structural integrity of the original image was primarily assessed using the SSIM, while learned perceptual image patch similarity (LPIPS) was simultaneously used to quantify perceptual distance in accordance with human visual judgment [26]. This evaluation was further extended to assess artifacts and naturalness through the analysis of Blind/Referenceless Image Spatial Quality Evaluator (BRISQUE) and Natural Image Quality Evaluator (NIQE), which quantify the absence of artificial distortions. This was complemented by the quantification of technical and aesthetic excellence via Multi-scale Image Quality transformer (MUSIQ) and aesthetic scores. CLIP–IQA, an IQA method based on the CLIP model, was integrated into the pipeline to cross-verify visual quality in terms of semantic consistency and high-level conceptual alignment between the source and generated outputs.

4.3.2. Greenery Insertion Results

The quantitative performance of the generative modifications across various experimental configurations is presented in Table 4. The data present a comparative analysis of structural fidelity, measured by SSIM and LPIPS, along with technical quality and aesthetic assessments. These metrics serve as an empirical foundation for identifying the optimal generative parameters within each experimental category.

4.3.3. Selection of Optimal Results by Experimental Group

By analyzing the trade-offs between structural fidelity and generative quality across all experimental groups (Versions 1–5, and Low, Medium, and High), the following optimal candidates were identified for each category:
In the quantitative assessment for determining the optimal generative outcomes across diverse experimental configurations, each specimen was systematically identified based on its capacity to balance structural fidelity with generative quality.
Specifically, Ver1-1 demonstrated exceptional fidelity, yielding a group-leading SSIM of 0.808, whereas Ver2-3 was prioritized for its superior perceptual alignment (LPIPS = 0.100) relative to the same series. Most significantly, Ver3-3 emerged as the overall optimal output across the entire dataset, attaining a peak SSIM of 0.824 and a minimum LPIPS of 0.075, indicating the near-optimal preservation of the semantic essence of the source. Furthermore, Ver4-2 was recognized for maintaining a stable SSIM of 0.776 despite the increased complexity of generative modification. Ver5-2 was selected owing to its superior structural preservation (SSIM = 0.784) and performance on non-reference quality metrics. Collectively, these results signify the highest level of adherence to established quantitative benchmarks by effectively suppressing generative artifacts while maintaining the integrity of the original visual information.
Within the density iterations, High-3 was designated as the representative for its category because it achieved a more harmonized distribution across multiple quality metrics than other candidates in the high-variance group. Concurrently, Low-2 and Med-2 exhibited minimal perceptual distortion (LPIPS = 0.093) and optimal synergy between robust structural similarity (SSIM = 0.800) and technical aesthetic excellence (Aesthetic = 4.513), respectively. Table 5 presents the representative optimal images selected from each experimental group. To provide clear visual evidence of the AI’s precision, each entry includes the final output and a Difference Map. The Difference Map highlights the specific pixels modified from the baseline environment, demonstrating the capability to maintain architectural integrity while precisely manipulating independent variables (i.e., greenery).

4.3.4. Perceptual Verification via VR-Based Assessment

In AI-generated 360° indoor imagery, conventional IQA metrics alone cannot fully capture the nuanced perceptual characteristics of omnidirectional environments [26,27,28]. While automated algorithms provide a technical baseline, determining the structural integrity and spatial plausibility remains a high-level cognitive task. Therefore, this study employed an expert-based evaluation as a final perceptual verification stage to complement the technical data.
To bridge the gap between automated metrics and human perception, a structured perceptual verification was conducted with seven senior experts (four architectural designers and three researchers) with over ten years of professional experience. For a controlled comparison, conventional 3D reference stimuli were developed using SketchUp and Enscape, applying identical environmental parameters—including spatial scale, lighting intensity, and material properties—to match the AI-generated environments (Table A3). The generated images were presented via Meta Quest 3 HMDs to ensure an immersive, high-fidelity spatial experience.
Experts evaluated both environments using a seven-point Likert scale, ensuring methodological consistency with the IPQ [29]. The evaluation of Presence revealed a substantial performance gap: AI-generated stimuli achieved a significantly higher mean score (6.143) compared to conventional modeling (4.397) (Table 6). To determine statistical significance, a Wilcoxon signed-rank test was performed on the aggregated presence scores. The results confirmed that AI-generated stimuli evoked a statistically significant higher level of spatial presence (Z = −2.3664, p = 0.018), proving that the proposed workflow achieves superior perceptual realism that surpasses traditional manual modeling standards.
Beyond Presence, to evaluate the specific performance of the generative AI workflow, we established a multi-dimensional framework derived from established generative AI literature [7,30,31,32,33]. This framework comprised three domains: (1) Architectural Quality, (2) Greenery Quality, and (3) Usability as an experimental stimulus (detailed in Table A4). The evaluation functioned as a technical audit; a high level of descriptive consensus was observed among the experts, confirming that the AI-generated stimuli meet the rigorous standards required for empirical research. Based on this framework, the AI-generated environments were further evaluated across specific quality metrics to verify their technical validity (Table 7). Within the Architectural Quality dimension, ‘Space function’ achieved the highest score (6.429), indicating that the AI accurately reflects the intended architectural purpose. Both ‘Text-to-image alignment’ and ‘Visual consistency’ scored 6.286, demonstrating the reliability of the system in translating descriptive prompts into coherent 360° visuals. In the Greenery Quality dimension, while ‘Design details’ (5.571) and ‘Greenery details’ (5.429) received slightly lower ratings, they still exceeded the scale’s median, confirming high overall fidelity. Finally, the Usability score of 6.000 serves as a critical benchmark, confirming that the experts considered these virtual spaces sufficiently robust for use as empirical research stimuli.

5. Discussion and Conclusions

The empirical results of this study support the methodological soundness and technical validity of generative AI-driven 360° images as experimental stimuli. Specifically, the ability to maintain environmental consistency while manipulating greenery elements via inpainting addresses a major challenge in spatial perception research.
The primary strength was revealed through the evaluation of quantitative metrics across different materials and greenery densities. In the first phase of space generation, the seam SSIM and PSNR values reached 0.847 and 27.930, respectively. These results are particularly significant as the seam area in equirectangular projections typically suffers from the highest distortion and lower SSIM values compared to the global average, yet our score of 0.847 surpasses the low 0.8 range identified as the critical threshold for eliminating perceptible ghosting artifacts and structural discontinuities [34]. The PSNR of 27.930 exceeds typical acceptable thresholds for generative outputs, where complex textures often lower numerical scores, confirming high signal precision. This indicates that the generated images maintain stable pixel-level correspondence and minimal noise at stitching boundaries. Such results are competitive with recent studies [34,35], verifying that the proposed approach delivers robust reconstruction quality consistent with current 360° image synthesis standards.
Furthermore, the second phase involving greenery integration demonstrated remarkable stability. The stability of SSIM (0.739–0.824) and Aesthetic scores (consistently above 4.4 on a scale of 1–10) across different greenery densities confirms that AI-based inpainting enables precise variable manipulation without reducing the underlying spatial quality. This level of aesthetic integrity is consistent with results reported in recent 360-degree panoramic studies [35,36], demonstrating that it ensures the visual consistency required for immersive environments. Even under varying greenery densities and surfaces, the stability of the MUSIQ (55.577–57.807) index demonstrates that AI can maintain high-fidelity visual standards regardless of experimental complexity. Notably, these scores are highly meaningful as they fall within the stable mid-to-high range reported in the original study [37], indicating professional-grade visual integrity despite inherent spherical distortions. The minimal variance in this metric further confirms the robustness of the generation process, ensuring uniform aesthetic quality across diverse spatial conditions.
The reliability of this approach was further supported by expert assessment, resulting in a significant increase in general presence (6.429) and visual realism (6.714) compared with the 4.286 and 4.000 scores of 3D conventional modeling, respectively. This discrepancy shows that AI-generated stimuli can more effectively overcome the limitations of visual artificiality by synthesizing complex lighting, subtle shadows, and realistic material textures, which typically require extensive computational resources and time for conventional modeling. Unlike previous studies that relied on static 2D AI images, this research demonstrates that 360° panoramic AI stimuli can maintain high spatial presence (6.429), enabling users to develop the psychological state of being “inside” the space.
This study offers a validated AI-generative approach that significantly reduces technical barriers to creating high-fidelity virtual stimuli, moving beyond the costly and rigid nature of conventional 3D modeling. The high text-to-greenery alignment (5.911) indicates that nuanced experiments on biophilic design can be conducted by controlling specific elements with high accuracy. For practitioners, the high usability (6.000) and space function (6.429) scores imply that AI can be used as a pre-occupancy evaluation tool to simulate user reactions to public space designs before actual construction. This convergence of subjective expert satisfaction and objective quality metrics confirms that the proposed three-stage framework offers a scientifically rigorous alternative to high-cost physical photography or complex 3D simulations, particularly for studies requiring precise control over environmental variables. Despite these strengths, this study has certain limitations. Although the visual consistency (6.286) was rated high, the current AI workflow may still produce minor artifacts in complex architectural geometries. In addition, the reliance on static 360° images limits the evaluation of the movement of dynamic user within the space. Furthermore, the generative process may be subject to biases from prompt engineering and a reliance on specific AI tools, which presents challenges for reproducibility. As results can vary across different generative models or versions, future research should explore the generalizability of this workflow across diverse platforms to mitigate model-specific uncertainties.
In conclusion, this study validates the methodological feasibility of using generative AI for creating high-fidelity 360° environments, confirming the technical and psychological validity of generative AI-driven 360° stimuli. The AI-generated stimuli achieved higher scores in professional evaluations, markedly surpassing the scores recorded by conventional modeling methods. By integrating objective quality metrics such as SSIM, PSNR, and CLIP-based assessments with subjective expert feedback, this study confirmed that AI can produce experimental stimuli that are both aesthetically superior and scientifically rigorous. The key findings highlight that AI-driven methodologies enable high visual consistency and precise text-to-greenery alignment, thereby enabling the manipulation of environmental variables with unprecedented flexibility. The high scores in the space function and presence dimensions reinforce the potential of AI to generate psychologically valid virtual representations of complex public spaces. This study offers a high-efficiency, low-cost paradigm for spatial perception studies, paving the way for future indoor environmental research utilizing AI-generated VR stimuli for more diverse and scalable experimental designs.
Future studies would focus on integrating these AI-generated 360° environments into real-time VR engines to facilitate interactive navigation. Additionally, testing the framework across a broader range of complex public facilities, such as transit hubs and healthcare environments, will further validate its generalizability. Finally, a comparison of psychological responses related to AI-generated stimuli and actual physical environments will provide definitive proof of the substitutability of this innovative method.

Author Contributions

Conceptualization, Y.H.; methodology, Y.H.; software, Y.H. and J.J.; validation, Y.H.; formal analysis, Y.H.; investigation, Y.H.; resources, Y.H.; data curation, Y.H. and J.J.; writing—original draft preparation, Y.H. and J.J.; writing—review and editing, Y.H.; visualization, Y.H. and J.J.; supervision, Y.H.; project administration, Y.H.; funding acquisition, Y.H. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korean government (MSIT) (No. RS-2024-00360680).

Data Availability Statement

The data that support the findings of this study are available from the corresponding author, Y.H., upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Table A1. Architectural characteristics of public common areas in domestic mixed-use facilities.
Table A1. Architectural characteristics of public common areas in domestic mixed-use facilities.
Analysis CategorySummary of Key Characteristics
Spatial TypePublic common areas in domestic mixed-use facilities
(e.g., department stores, large commercial complexes)
Method of Spatial AnalysisField surveys, photographic documentation, spatial analysis, and online data review
Spatial StructureCentral vertical void penetrating multiple floors
Spatial FormPredominantly rectilinear configuration
Horizontal ElementsRepetitive slabs of uniform thickness
Structural SystemRegular slab–column structural configuration
Balustrade ConfigurationTransparent glass balustrades with light-toned metal frames
Daylighting ConditionFully enclosed interior space with exclusion of external views
Lighting EnvironmentUniform artificial lighting using indirect illumination
Material CharacteristicsNeutral-toned interior finish materials
Table A2. Summary of material control strategies by stimulus version.
Table A2. Summary of material control strategies by stimulus version.
CategoryWhiteLight WoodDark Wood
Material
Basis
White plaster, light-toned stoneLight beige and light woodDark-toned wood combined with stone and metal
Referenced
Spatial
Characteristics
Neutral finishes commonly observed in public areas of mixed-use facilitiesWarm material impression typical of underground common areasPremium atmosphere of high-end mixed-use facilities
Risk of AI Auto-
Interpretation
LowRisk of being interpreted as retail or café space if wood use is excessiveRisk of being interpreted as hotel or lounge program
Control
Strategy
Restriction of color contrast and surface textureLimitation of color range and proportion of wood applicationMaintenance of material hierarchy and minimization of furniture
Target Spatial
Perception
Neutral public atriumWarm but non-program-specific public spaceHigh-end yet clearly commercial public space
Table A3. AI-generated and conventional modeling-based 360° panoramic images.
Table A3. AI-generated and conventional modeling-based 360° panoramic images.
AI-Generated ImageConventional Modeling-Based Image
Buildings 16 01679 i021Buildings 16 01679 i022
Table A4. Expert perceptual assessment items.
Table A4. Expert perceptual assessment items.
DimensionsItemsDescription
PresenceGeneral presenceThe virtual space felt like a real place rather than an image or simulation.
Spatial presenceI felt as if I was really inside the space shown in the 360° environment.
InvolvementI had a strong sense of “being there” in the environment.
I was fully engaged with the environment during the experience.
I paid more attention to the environment than to the real world around me.
I felt deeply immersed in the virtual environment.
RealismThe environment looked visually realistic.
The spatial layout and depth felt similar to a real indoor space.
The lighting and materials appeared natural and believable.
Architectural
quality
Visual qualityThe environment is aesthetically pleasing and free of obvious errors.
Space functionThe design meets its intended purpose.
Text-to-image
alignment
The 360-panoramic image matches the text description (prompt).
Design detailsThe design elements are thoughtfully designed, with a reasonable layout and composition.
Visual consistencyThe 360-panoramic images maintain a high level of visual consistency.
Greenery
quality
Text-to-greenery
alignment
The greenery of the space aligns with the provided text description (prompt).
Greenery detailsThe arrangement and composition of the greenery elements reflect a high level of intricacy and thoughtfulness in their design.
UsabilityUsabilityThe quality of this virtual space is sufficient to be used as actual research stimuli.

References

  1. Ulrich, R.S. View through a window may influence recovery from surgery. Science 1984, 224, 420–421. [Google Scholar] [CrossRef] [PubMed]
  2. Kaplan, R.; Kaplan, S. The Experience of Nature: A Psychological Perspective; Cambridge University Press: New York, NY, USA, 1989. [Google Scholar]
  3. Kim, N.; Lee, H. Assessing consumer attention and arousal using eye-tracking technology in virtual retail environment. Front. Psychol. 2021, 12, 665658. [Google Scholar] [CrossRef] [PubMed]
  4. Lee, J.K.; Lee, S.; Kim, Y.C.; Kim, S.; Hong, S.W. Augmented virtual reality and 360 spatial visualization for supporting user-engaged design. J. Comput. Des. Eng. 2023, 10, 1047–1059. [Google Scholar] [CrossRef]
  5. Shinde, Y.; Lee, K.; Kiper, B.; Hasanzadeh, S. A Systematic Literature Review on 360° Panoramic Applications in the Architecture, Engineering, and Construction (AEC) Industry. J. Inf. Technol. Constr. 2023, 28, 405–437. [Google Scholar] [CrossRef]
  6. Zhu, Y.; Fukuda, T.; Yabuki, N. A mixed reality design system for interior renovation: Inpainting with 360-degree live streaming and generative adversarial networks after removal. Technologies 2024, 12, 9. [Google Scholar] [CrossRef]
  7. Yang, W.; Wang, C.; Liu, L.; Dong, S.; Zhao, Y. Advancing interior design with AI: Controllable Stable Diffusion for panoramic image generation. Buildings 2025, 15, 1391. [Google Scholar] [CrossRef]
  8. Zhang, C.; Wu, Q.; Gambardella, C.C.; Huang, X.; Phung, D.; Ouyang, W.; Cai, J. Taming Stable Diffusion for text to 360° panorama image generation. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024), Seattle, WA, USA, 17–21 June 2024; pp. 6347–6357. [Google Scholar]
  9. Kalischek, N.; Oechsle, M.; Manhardt, F.; Henzler, P.; Schindler, K.; Tombari, F. Cubediff: Repurposing diffusion-based image models for panorama generation. In Proceedings of the Thirteenth International Conference on Learning Representations (ICLR 2025), Singapore, 24–28 April 2025. [Google Scholar]
  10. Lee, J.K.; Kim, Y.; Shin, E.; Choo, S.; Cha, S.H. An AI-assisted approach for creating and archiving interior design references using 360-degree panoramic images. Archit. Sci. Rev. 2025, 68, 506–518. [Google Scholar] [CrossRef]
  11. Ni, J.; Zhang, C.B.; Zhang, Q.; Zhang, J. What makes for text to 360-degree panorama generation with Stable Diffusion? In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2025), Honolulu, HI, USA, 19–23 October 2025; pp. 16555–16564. [Google Scholar]
  12. Paananen, V.; Oppenlaender, J.; Visuri, A. Using text-to-image generation for architectural design ideation. Int. J. Archit. Comput. 2023, 22, 458–474. [Google Scholar] [CrossRef]
  13. Horvath, A.-S.; Pouliou, P. AI for conceptual architecture: Reflections on designing with text-to-text, text-to-image, and image-to-image generators. Front. Archit. Res. 2024, 13, 593–612. [Google Scholar] [CrossRef]
  14. Akyildiz, E.C. An analysis of text-to-image generative models as creativity support tools. Int. J. Arts Technol. 2025, 15, 257–282. [Google Scholar] [CrossRef]
  15. Jeong, H.; Kim, Y.; Yoo, Y.; Cha, S.; Lee, J.-K. Gen AI and interior design representation: Applying design styles using fine-tuned models. In Proceedings of the 23rd International Conference on Construction Applications of Virtual Reality (CONVR 2023), Florence, Italy, 13–16 November 2023; pp. 950–957. [Google Scholar] [CrossRef]
  16. Li, C.; Zhang, T.; Du, X.; Zhang, Y.; Xie, H. Generative AI models for different steps in architectural design: A literature review. Front. Archit. Res. 2025, 14, 759–783. [Google Scholar] [CrossRef]
  17. Chen, J.; Zheng, X.; Shao, Z.; Ruan, M.; Li, H.; Zheng, D.; Liang, Y. Creative interior design matching the indoor structure generated through diffusion model with an improved control network. Front. Archit. Res. 2025, 14, 614–629. [Google Scholar] [CrossRef]
  18. Pramalystianto, A. Study of image resolution in virtual reality panorama Technology 360 in comfort of interior visualization. Int. J. Archit. Urban. 2023, 7, 244–250. [Google Scholar] [CrossRef]
  19. Song, J.; Guo, H.; Zhang, L.; Wang, Z.; Yip, D. Expanding virtual production frontiers: AI-driven workflows for enhanced cinematic creation. In Proceedings of the 18th International Symposium on Visual Information Communication and Interaction (VINCI ’25), Linz, Austria, 1–3 December 2025; pp. 1–5. [Google Scholar] [CrossRef]
  20. Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, Online, 18–24 July 2021; Volume 139, pp. 8748–8763. Available online: https://proceedings.mlr.press/v139/radford21a.html (accessed on 10 January 2026).
  21. Devernay, F.; Faugeras, O. Straight lines have to be straight. Mach. Vis. Appl. 2001, 13, 14–24. [Google Scholar] [CrossRef]
  22. Szeliski, R. Image alignment and stitching: A tutorial. Found. Trends Comput. Graph. Vis. 2007, 2, 1–104. [Google Scholar] [CrossRef]
  23. Wang, Z.; Bovik, A.C. Mean squared error: Love it or leave it? A new look at signal fidelity measures. IEEE Signal Process. Mag. 2009, 26, 98–117. [Google Scholar] [CrossRef]
  24. Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image quality assessment: From error visibility to structural similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef]
  25. Hessel, J.; Holtzman, A.; Forbes, M.; Le Bras, R.; Choi, Y. Clipscore: A Reference-Free Evaluation Metric for Image Captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online and Punta Cana, Dominican Republic, 7–11 November, 2021; pp. 7514–7528. [Google Scholar]
  26. Zhang, R.; Isola, P.; Efros, A.A.; Shechtman, E.; Wang, O. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 586–595. [Google Scholar] [CrossRef]
  27. Tran, H.T.; Ngoc, N.P.; Pham, C.T.; Jung, Y.J.; Thang, T.C. A subjective study on QoE of 360 video for VR communication. In Proceedings of the IEEE 19th International Workshop on Multimedia Signal Processing (MMSP), Luton Bedfordshire, UK, 16–18 October 2017; pp. 1–6. [Google Scholar] [CrossRef]
  28. Hai Uyen, T.T.; Kwon, O.J.; Choi, S.; Hussain, I. Subjective assessment of 360° image projection formats. IEEE Access 2020, 8, 33588–33599. [Google Scholar] [CrossRef]
  29. Schubert, T.; Friedmann, F.; Regenbrecht, H. The experience of presence: Factor analytic insights. Presence Teleoper. Virtual Environ. 2001, 10, 266–281. [Google Scholar] [CrossRef]
  30. Shao, Z.; Chen, J.; Zeng, H.; Hu, W.; Xu, Q.; Zhang, Y. A new approach to interior design: Generating creative interior design videos of various design styles from indoor texture-free 3D models. Buildings 2024, 14, 1528. [Google Scholar] [CrossRef]
  31. Shi, M.; Seo, J.; Cha, S.H.; Xiao, B.; Chi, H.L. Generative AI-powered architectural exterior conceptual design based on the design intent. J. Comput. Des. Eng. 2024, 11, 125–142. [Google Scholar] [CrossRef]
  32. Jiang, J. Enhancing interior design and space planning via human–machine intelligent interaction for artistic cognition. Sci. Rep. 2025, 15, 32344. [Google Scholar] [CrossRef]
  33. Kuang, Z.; Zhang, J.; Li, Y.; Fukuda, T. Preserving architectural heritage in urban renewal: A Stable Diffusion model framework for automated historical facade generation. npj Herit. Sci. 2025, 13, 256. [Google Scholar] [CrossRef]
  34. Li, J.; Yu, K.; Zhao, Y.; Zhang, Y.; Xu, L. Cross-reference stitching quality assessment for 360 omnidirectional images. In Proceedings of the 27th ACM International Conference on Multimedia, Nice, France, 21–25 October 2019; pp. 2360–2368. [Google Scholar]
  35. Wu, Z.; Watson, D.; Tagliasacchi, A.; Fleet, D.J.; Brubaker, M.A.; Saxena, S. 360Anything: Geometry-Free Lifting of Images and Videos to 360°. arXiv 2026, arXiv:2601.16192. [Google Scholar] [CrossRef]
  36. Luo, R.; Wallingford, M.; Farhadi, A.; Snavely, N.; Ma, W.C. Beyond the Frame: Generating 360° panoramic videos from perspective videos. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2025), Honolulu, HI, USA, 19–23 October 2025; pp. 14336–14345. [Google Scholar]
  37. Ke, J.; Wang, Q.; Wang, Y.; Milanfar, P.; Yang, F. MUSIQ: Multi-scale image quality transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2021), Online, 11–17 October 2021; pp. 5148–5157. [Google Scholar]
Figure 1. Experimental framework illustrating the fixed base environment and the systematic manipulation of greenery placement (V1–V5) and density (low-high) for perceptual evaluation.
Figure 1. Experimental framework illustrating the fixed base environment and the systematic manipulation of greenery placement (V1–V5) and density (low-high) for perceptual evaluation.
Buildings 16 01679 g001
Table 1. Operational definitions of key terminology.
Table 1. Operational definitions of key terminology.
CategoryTermOperational Definition in This Study
ConceptEnvironmental
Stimulus
The physical and sensory attributes (e.g., wood tones, greenery density) designed to elicit psychological or physiological responses in indoor environmental research.
ArtifactSpatial
Stimulus
The specific 360° panoramic image generated by AI and presented to participants as the visual proxy for the virtual environment.
ProcessStimulus
Production
The systematic technical workflow encompassing prompt engineering, diffusion-based image generation, and post-optimization.
MechanismText-to-Image (T2I)A generative synthesis approach that creates visual content strictly based on textual descriptions, ensuring precise control over experimental variables.
EvaluationSemantic
Alignment
The degree of correspondence between the intended research variables (input prompts) and the resulting visual content (output images), quantified via CLIP scores.
Table 2. Common environmental control conditions applied to all stimuli.
Table 2. Common environmental control conditions applied to all stimuli.
CategoryItemControl Specification
Spatial
conditions
Spatial structureCentral vertical void combined with repetitive horizontal slabs
Floor configurationSimultaneous perception of two to three floors
CirculationCentral open atrium with peripheral circulation
BalustradesTransparent glass balustrades with light-toned metal frames
LightingFully artificial lighting with uniform illumination
External elementsExclusion of exterior views, natural daylight, and skylights
People & DecorationsComplete exclusion of people, greenery, and installations
Camera
parameters
Camera positionFixed (eye-level)
Field of view±40° horizontal field of view
Image
specifications
Projection type360° equirectangular (2:1)
Resolution8K
Table 3. Quantitative evaluation results.
Table 3. Quantitative evaluation results.
ALT.ImagesScores
WhiteBuildings 16 01679 i001CLIP Score0.309
Relative SLD (%)0.050
Seam MSE328.190
Seam PSNR22.970
Seam SSIM0.752
Light woodBuildings 16 01679 i002CLIP Score0.359
Relative SLD (%)0.093
Seam MSE104.820
Seam PSNR27.930
Seam SSIM0.840
Dark woodBuildings 16 01679 i003CLIP Score0.319
Relative SLD (%)0.040
Seam MSE305.160
Seam PSNR23.290
Seam SSIM0.847
Table 4. Evaluation results on set of generated panoramic images.
Table 4. Evaluation results on set of generated panoramic images.
GroupSSIM
(↑)
LPIPS
(↓)
BRISQUE
(↓)
NIQE
(↓)
MUSIQ
(↑)
Aesthetic
(↑)
CLIP-IQA
(↑)
MMMSDMSDMSDMSDMSD
Orig.--34.3623.1368.5051.06463.2857.5594.9710.5940.6690.076
Ver110.8080.09436.1404.1797.6861.48856.6696.7914.5500.5760.4440.048
20.7910.13433.2856.0167.5691.53656.3176.7314.5400.5670.4430.054
30.7910.13234.4844.9657.7021.53156.9197.1104.5300.6010.4400.056
Ver210.7790.12334.8004.2007.5881.41657.6176.8504.4690.5510.4570.058
20.7870.09536.7683.4918.0341.49657.3817.2674.4610.5450.4760.058
30.7940.10035.6084.1077.6951.50357.3437.1864.4760.5520.4570.057
Ver310.7950.13931.7157.1927.3991.67756.6656.1664.5370.5630.4550.060
20.8200.07537.2144.5277.9591.48956.9057.1614.4700.5410.4580.073
30.8240.07536.7274.8697.9161.31357.0087.0824.4850.5870.4730.079
Ver410.7550.17833.6255.1287.6121.59056.6166.2844.5090.5570.4340.059
20.7760.12534.4654.7817.4531.44956.4276.7214.4560.5270.4380.061
30.7750.14031.3946.8907.5471.56656.2366.4034.4450.4860.4260.065
Ver510.7390.19829.6138.1407.5611.58157.4016.2994.4770.5220.4300.051
20.7840.14433.8995.0607.6221.39055.5776.4464.4580.4900.4260.060
30.7580.18132.5795.9787.5641.70656.6026.7234.4760.5690.4320.061
Low10.7790.12334.8004.2007.5881.41657.6176.8504.4690.5510.4570.058
20.7980.09336.7534.2927.6221.16656.3556.4124.4210.5180.4450.080
30.7940.10335.5974.2127.6821.45556.1116.7194.4650.5250.4420.072
Med10.7550.17833.6255.1287.6121.59056.6166.2844.5090.5570.4340.059
20.8000.07937.2063.3007.8261.42957.8077.3234.5130.5500.4610.061
30.7710.14132.1856.2887.2641.61355.8956.2314.4370.4880.4380.067
High10.7390.19829.6138.1407.5611.58157.4016.2994.4770.5220.4300.051
20.7780.12432.1226.0107.5351.35456.7066.7504.4450.5120.4310.063
30.7560.15731.1946.6517.5391.55856.9487.3394.4600.5260.4500.067
Table 5. Selected images by experimental group.
Table 5. Selected images by experimental group.
Group ImagesDifference Map
Ver1Buildings 16 01679 i004Buildings 16 01679 i005
2Buildings 16 01679 i006Buildings 16 01679 i007
3Buildings 16 01679 i008Buildings 16 01679 i009
4Buildings 16 01679 i010Buildings 16 01679 i011
5Buildings 16 01679 i012Buildings 16 01679 i013
DensityLowBuildings 16 01679 i014Buildings 16 01679 i015
MedBuildings 16 01679 i016Buildings 16 01679 i017
HighBuildings 16 01679 i018Buildings 16 01679 i019
Buildings 16 01679 i020
Table 6. Comparative analysis of mean presence scores: AI-generated vs. traditional modeling.
Table 6. Comparative analysis of mean presence scores: AI-generated vs. traditional modeling.
DimensionsSub-DimensionsAI-Generated (M)Modeling (M)Difference
PresenceGeneral Presence6.4294.286+2.143
Spatial Presence6.4294.714+1.715
Involvement-15.8574.571+1.286
Involvement-25.7144.000+1.714
Involvement-36.2864.857+1.429
Involvement-46.0004.143+1.857
Realism-16.7144.000+2.714
Realism-26.2864.857+1.429
Realism-35.5714.143+1.428
Total 6.1434.397+1.746
Table 7. Expert assessment scores for AI-generated 360° panoramic images.
Table 7. Expert assessment scores for AI-generated 360° panoramic images.
DimensionsItemsScore (M)
Architectural qualityVisual quality5.429
Space function6.429
Text-to-image alignment6.286
Design details5.571
Visual consistency6.286
Greenery qualityText-to-greenery alignment5.911
Greenery details5.429
UsabilityUsability6.000
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Han, Y.; Jeong, J. Leveraging Generative AI for High-Fidelity 360° Spatial Images: Methodological Validation for Use as Experimental Stimuli. Buildings 2026, 16, 1679. https://doi.org/10.3390/buildings16091679

AMA Style

Han Y, Jeong J. Leveraging Generative AI for High-Fidelity 360° Spatial Images: Methodological Validation for Use as Experimental Stimuli. Buildings. 2026; 16(9):1679. https://doi.org/10.3390/buildings16091679

Chicago/Turabian Style

Han, Yoojin, and Joowon Jeong. 2026. "Leveraging Generative AI for High-Fidelity 360° Spatial Images: Methodological Validation for Use as Experimental Stimuli" Buildings 16, no. 9: 1679. https://doi.org/10.3390/buildings16091679

APA Style

Han, Y., & Jeong, J. (2026). Leveraging Generative AI for High-Fidelity 360° Spatial Images: Methodological Validation for Use as Experimental Stimuli. Buildings, 16(9), 1679. https://doi.org/10.3390/buildings16091679

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop