Next Article in Journal
Children’s Perception of Urban Outdoor Spaces and Playground Design: A Sensory Walk Study in Zagreb, Croatia
Previous Article in Journal
Flexible Futures: Designing High Streets for the Crises to Come
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Urban Comfort Perception Under Induced Emotional Conditions: A Multi-Method Analysis of Architectural and Streetscape Imagery Using Fractal Analysis, Self-Report, and Eye-Tracking

by
Satrio Agung Perwira
1,*,
Bart Julien Dewancker
2 and
Dimas Herjuno
3
1
Graduate School of Environmental Engineering, The University of Kitakyushu, 1-1 Hibikino, Wakamatsu Ward, Kitakyushu 808-0135, Fukuoka, Japan
2
Department of Architecture, The University of Kitakyushu, 1-1 Hibikino, Wakamatsu Ward, Kitakyushu 808-0135, Fukuoka, Japan
3
Graduate School of Life Science and Systems Engineering, Kyushu Institute of Technology, 2-4 Hibikino, Wakamatsu Ward, Kitakyushu 808-0196, Fukuoka, Japan
*
Author to whom correspondence should be addressed.
Architecture 2026, 6(2), 91; https://doi.org/10.3390/architecture6020091
Submission received: 31 March 2026 / Revised: 13 May 2026 / Accepted: 3 June 2026 / Published: 8 June 2026

Abstract

This pilot study examines how experimentally induced emotional states interact with the visual properties of urban environments to shape comfort perception. A controlled laboratory experiment was conducted with 17 participants assigned to one of four emotional conditions (Fear, Anger, Sad, Happy) through audio-visual induction. Participants evaluated 73 building façade and 42 pedestrian streetscape stimuli from three urban areas in Kitakyushu, Japan (Wakamatsu, Tobata, Mojiko) using a multi-method framework combining fractal analysis (D, Λ), six pedestrian visual metrics, webcam-based eye-tracking (Visual Attention Score, VAS), and self-reported comfort votes. Emotion induction was effective for Fear and Anger groups and partial for Sad and Happy groups, with the latter attributable to experimental fatigue. Cross-method correlation analysis revealed that fractal dimension D significantly predicted comfort vote consensus (Spearman r = 0.369, p = 0.013), while VAS showed no significant relationship with comfort votes (r = 0.097, ns) or with fractal dimension (r = 0.015, ns), confirming that visual attention and comfort preference are independent dimensions. For building façades, the ‘Complex but Organized’ fractal profile (D ≥ 1.70, Λ < 0.60) was the consistent comfort driver across all emotion groups. For pedestrian streetscapes, low spatial enclosure and spatially integrated tree canopy were the primary comfort predictors. Multi-method synthesis identified five empirical paradoxes and three design principles: (1) target D ≥ 1.70 with Λ < 0.60; (2) prioritize spatially integrated canopy over visible greenery quantity; and (3) leverage civic legibility as an independent comfort pathway. These findings support the development of emotion-independent frameworks for urban comfort evaluation. Replication with larger, more diverse samples is recommended.

1. Introduction

1.1. Background

Urban environments influence not only how people move and function within a space but also how they feel and perceive it. The visual characteristics of the built environment, including architectural buildings, streetscapes, and pedestrian-level views play a significant role in shaping comfort perception, spatial preference, and emotional response. Research in environmental psychology has shown that the physical features of urban settings can directly affect human emotional and cognitive evaluations, often beyond conscious awareness [1,2]. A key question in this field is whether these perceptual responses are stable across different emotional states, or whether a person’s current emotional condition shapes how they evaluate urban environments.
At the pedestrian scale, individuals experience cities through continuous visual interaction with their surroundings. Elements such as building form, façade composition, spatial enclosure, and street complexity contribute to how people judge the overall quality of an urban area. Studies have demonstrated that visually coherent and well-structured environments tend to be associated with more positive emotional responses, while disordered or overly complex environments may reduce perceived comfort [3,4]. Importantly, these judgments are not always explicitly articulated, meaning that individuals may perceive a space as comfortable or uncomfortable without clearly identifying the underlying reasons.
This study focuses on selected urban areas in Kitakyushu, where different urban typologies coexist, including residential, commercial, industrial, and tourism-oriented districts. These areas present varying spatial characteristics, ranging from organically developed environments to more planned and structured urban forms. Such variation provides a valuable context for examining how differences in architectural and streetscape composition influence emotional perception and comfort evaluation [5].
To better understand these relationships, this research adopts a multimodal approach combining self-reported emotional data with eye-tracking observation. Self-reporting methods capture participants’ subjective evaluations of comfort and emotional experience, while eye-tracking provides objective insights into visual attention and gaze behavior toward specific environmental features [6,7]. The integration of these methods allows for a more comprehensive understanding of how urban imagery is perceived and how different components contribute to overall emotional responses.
Based on this approach, the objectives of this study are as follows:
(1)
To examine how architectural and streetscape visual properties influence participants’ comfort perception under experimentally induced emotional conditions.
(2)
To identify visual attention patterns toward specific urban elements using eye-tracking data, and assess whether these patterns vary across emotional conditions.
(3)
To analyze the relationship between induced emotional state and the selectivity of self-reported comfort responses.
(4)
To explore how different urban typologies in Kitakyushu contribute to variations in perceived comfort across emotional conditions.
Through these objectives, this study aims to contribute to the development of an emotionscape-based analytical framework for understanding human responses to urban environments, supporting more perceptually informed approaches to architectural and urban design.

1.2. Aim and Contribution

This study aims to improve understandings of urban comfort perception by examining how experimentally induced emotional states interact with the visual properties of architectural and streetscape environments. Rather than treating comfort as a fixed quality of a space, this research investigates whether and how a person’s current emotional condition shapes their comfort evaluation of the same visual environment. Previous studies have established that built environment features can influence emotional responses, preferences, and overall wellbeing [8]. Building on this foundation, the present study uses controlled emotional induction combined with self-reported comfort evaluations and eye-tracking observations to examine how visual attention, environmental features, and perceived comfort relate to one another across different emotional states [6,7].
The primary contribution of this research is empirical evidence that urban comfort perception is driven predominantly by the quantifiable visual properties of the environment, specifically fractal complexity and spatial enclosure rather than by the observer’s transient emotional state. This finding supports the development of more reliable, perception-based methods for urban design evaluation, where comfort assessments can be treated as relatively stable across diverse user emotional profiles. This is consistent with prior research in environmental psychology emphasizing the role of physical environmental qualities in shaping user experience [3].
A further contribution is the structured analytical framework connecting visual metrics, fractal analysis, and eye-tracking data with self-reported comfort responses under controlled emotional conditions. This framework provides a replicable and evidence-based approach to evaluating architectural and streetscape quality from a human perception standpoint. The findings also establish a foundation for future research incorporating more immersive methods, such as virtual reality environments or physiological measurements including EEG, to examine whether the stability of comfort perception observed here holds under conditions that more closely replicate lived urban experience.

2. Literature Review

Recent studies on visual comfort and urban perception have increasingly applied quantitative and computational approaches to analyze the relationship between environmental features and human evaluation. Image-based methods using street-view data and machine learning have been widely used to extract visual elements such as buildings, greenery, and sky visibility, and to relate these features to perceived qualities such as safety, attractiveness, and preference [8,9]. These approaches enable large-scale analysis and allow researchers to quantify visual characteristics through measurable indicators.
One commonly used technique in this field is the extraction of environmental indices from images, such as the Green View Index (GVI), which represents the proportion of visible greenery within a scene [10]. Similar segmentation-based methods have been applied to quantify other visual components, including buildings, roads, and open sky, providing a structured way to describe the visual composition of urban environments. These quantified variables are often used as independent variables to predict human perception outcomes.
In parallel, perception-based studies have relied on self-report methods to evaluate how individuals respond to urban environments. Participants are typically asked to rate scenes based on perceived qualities such as comfort, safety, or preference, allowing researchers to capture subjective evaluation directly [1,3]. These responses are then analyzed in relation to environmental variables, establishing links between physical characteristics and perceived experience.
More recent studies have introduced behavioral observation techniques, such as eye-tracking, to complement self-reported data. Eye-tracking provides detailed information about visual attention, including fixation points, gaze paths, and viewing duration, can indicate how individuals interact with specific elements in a scene [6,7]. By combining gaze data with perception ratings, researchers are able to better understand how visual attention relates to environmental evaluation.
Overall, existing studies demonstrate three main approaches in the analysis of urban visual comfort: (1) computational methods that quantify visual elements, (2) self-report methods that capture subjective perception, and (3) behavioral observation techniques that record visual attention. These approaches provide complementary insights into how urban environments are experienced and evaluated.

2.1. Impact of Architectural Building on Visual Comfort

Architectural building characteristics have been widely shown to influence visual comfort and environmental perception. Elements such as façade articulation, building height, and spatial enclosure affect how people visually experience urban spaces. Environments with balanced complexity and coherent façade composition tend to produce higher visual preference and perceived comfort [1,3,11], while excessive visual disorder or monotony may reduce environmental quality.
Spatial form also plays a key role in shaping emotional and perceptual responses. Variations in enclosure, openness, and architectural configuration influence how spaces are evaluated and experienced [2,12]. These spatial conditions contribute to feelings of comfort, interest, or stress depending on their visual organization.
In addition, building geometry and façade characteristics affect preference. Curved and visually softer forms are generally perceived more positively than rigid or sharp forms [13]. Architectural variation can enhance visual interest, while maintaining coherence supports clarity and ease of perception [14,15].
At the pedestrian scale, building proportion and human-scale design are critical for visual comfort. Environments that align with human scale and provide appropriate spatial proportions are more likely to be perceived as comfortable and walkable [16,17]. In contrast, overly dense or visually overwhelming configurations can lead to discomfort and negative perception [18].

2.2. Impact of Pedestrian Landscape and Hardscape on Visual Comfort

Pedestrian landscape and hardscape elements significantly influence visual comfort and environmental perception. Features such as vegetation, ground surfaces, and street layout affect how individuals visually experience urban spaces. Studies have shown that the presence of greenery, including trees and planting elements, is associated with higher visual preference and improved psychological responses [19,20,21].
Hardscape components also contribute to visual evaluation. The quality and organization of pavements, walkways, and street elements can support a sense of clarity and usability in pedestrian environments. Well-structured pedestrian spaces tend to enhance comfort and perceived environmental quality, while disordered or cluttered settings may reduce visual satisfaction [22,23].
Spatial configuration at the pedestrian level further affects perception. The relationship between openness, visibility, and enclosure influences how comfortable and legible a space feels. Environments that provide clear visual access and defined spatial structure are generally associated with better orientation and positive user experience [24,25].
Material and surface characteristics also play a role in shaping visual comfort. Variations in texture, color, and paving patterns can enhance visual interest when applied consistently, contributing to a more engaging and readable environment [16]. In contrast, inconsistent or overly complex surface treatments may reduce clarity and affect overall perception.

2.3. Measuring Perceived Visual Comfort Through Personal Perception-Based Experimental Observation

Perceived visual comfort is commonly evaluated through personal perception using experimental observation methods. This approach captures how individuals directly assess visual environments based on their own experience, which is essential for understanding subjective responses. Foundational studies in environmental psychology highlight that emotional and perceptual evaluations can be reliably measured through self-report scales [26,27], and these approaches continue to be widely used in recent urban perception research [28].
Experimental observation is often conducted using visual stimuli such as photographs or videos representing urban environments. This method allows controlled comparison between different scenes while maintaining consistent evaluation conditions. Earlier work demonstrated that image-based assessments can effectively represent real-world environmental perception [29], while more recent studies confirm their validity in large-scale perception analysis and urban evaluation [4,30].
In addition, perception-based evaluation is increasingly supported by behavioral measurement techniques. Eye-tracking methods provide objective data on visual attention, including fixation duration and gaze patterns, which help explain how individuals interact with environmental elements during evaluation [6,7]. Recent studies combining eye-tracking with perception data show that visual attention is closely related to preference and comfort, strengthening the interpretation of subjective responses [28,31].

3. Data and Methodology

3.1. Study Area

Three case study areas were selected: Wakamatsu, Tobata, and Mojiko. These areas represent different urban functions, including residential, commercial, industrial, and special district (tourist) environments. This selection ensures diversity in urban typologies while maintaining a consistent perceptual context (Figure 1).
From each area, one specific route was selected based on locations commonly experienced by pedestrians. The selected routes include movement corridors such as from a train station to another station, or in this case, from the station to the port connecting Wakamatsu and Tobata. In Mojiko, the selected area covers the route from the station to the northeast zone, which is accessible by walking (Figure 2).
In terms of characteristics, the selected area in Wakamatsu represents a mix of commercial and residential buildings. Tobata mainly represents residential environments, while Mojiko includes commercial, industrial, and landmark buildings typically found in special district areas, such as tourist destinations. The architectural features across these three areas vary, ranging from simple forms and minimal ornamentation to more unique and complex decorative elements.

3.2. Participants and Setting

A total of 22 participants were recruited, comprising undergraduate to postgraduate students from the University of Kitakyushu, Kyushu Institute of Technology, and Waseda University, with one member from a professional engineering background. Of these, 17 datasets were retained for analysis following data quality screening; five were excluded due to incomplete questionnaire responses or eye-tracking data loss.
Each participant completed the experiment individually in a controlled laboratory session of approximately 60 min after providing written informed consent (Figure 3). Given the exploratory nature of this study and the practical constraints of laboratory-based recruitment, the sample represents a pilot-scale investigation.
Given the exploratory and pilot-scale nature of this study, between-group statistical comparisons should be treated as descriptive and hypothesis-generating rather than confirmatory. With approximately four participants per emotion group, the sample is insufficient to reliably detect small-to-medium effect sizes in pairwise comparisons; this is consistent with field norms for multi-method studies combining eye-tracking and physiological measures, where samples of n = 12–20 are typical due to equipment cost and protocol duration, given the previous examples ref. [5] with n = 20, ref. [21] with n = 14, and ref. [32] with n = 16. Null results should not be interpreted as evidence of no effect. The primary analytical value of between-group comparison lies in identifying directional patterns to inform adequately powered future studies.
Additionally, emotion-induced effectiveness was uneven across conditions: induced was confirmed for the Fear and Anger groups but only partial for the Sad and Happy groups, attributable to experimental fatigue. This asymmetry further constrains the interpretive validity of cross-group comparisons involving those two groups, and directional patterns from Sad and Happy data should be treated with additional caution.
The equipment used in the experiment is shown in Figure 4 and comprised five items such as, a webcam-based eye-tracking system using Google Media Pipe Face Mesh (version 0.10.21) running on Tensor Flow Lite (version 2.14.0) for real-time gaze estimation; a laptop for running the questionnaire and recording data; a 4K monitor (3840 × 2160 pixels) for displaying all stimulus content to the participant; a camera on a tripod for the observational recording of each session; and a portable speaker for continuous audio playback.

3.3. Experimental Procedure

Each session lasted approximately 60 min and followed a fixed sequence of five stages, as illustrated in Figure 5.
In Stage A (preparation), participants were briefed and seated facing the large monitor, with the eye-tracking camera positioned toward them. Before any architectural stimuli were shown, emotion-inducing video clips were played to establish a specific target emotional state in the participant, corresponding to one of four categories: Fear, Anger, Sad, or Happy (Figure 6). Each video clip contained both visual and audio content selected to reliably evoke the target emotion, as shown in Figure 5. The use of film-based audio-visual stimuli for laboratory emotion induction is well established in experimental psychology research [33,34].
In Stage B (eye-tracker calibration), the webcam-based eye-tracking system was calibrated for each participant using a green dot-grid sequence displayed on the monitor before the observation session began.
Participants’ self-reported emotional states were recorded immediately before Stage A (baseline) and again after Stage E (post-experiment) as a manipulation check to assess induction effectiveness. Reponses were classified using the circumplex model of effect [26] into Positive, Neutral, or Negative valence categories; low-arousal states such as tired and sleepy were classified as Neutral rather than Negative, as they reflect fatigue rather than target emotional induction.
To address the potential decay of the induced emotional state across 115 stimuli presented over approximately 50 min, continuous emotion-congruent audio stimuli were played via the portable speaker throughout both observation sessions (Stages C and D). The audio tracks were selected to reinforce the emotional tone of the induction video without introducing new narrative content, providing sustained affective priming across the full viewing period. This design follows the recommendation that emotion maintenance in extended laboratory protocols requires continuous or periodic re-exposure to emotion-congruent stimuli [35].
EEG data were also recorded throughout the experiment as an independent physiological measure; these will be reported in a separate publication focused on neurophysiological responses. The present paper focuses exclusively on vision-based and self-reported measures.
In stage C (first observation session), 73 architectural façade figures were displayed sequentially on the monitor, each shown for 10 s, while gaze data were recorded continuously. A single Area of Interest (AOI) was defined per image as a bounding box enclosing the entire building façade surface, and the proportion of each participant’s gaze falling within this AOI was measured [6]. After this session, participants took a short resting break. In Stage D (second observation session), the same procedure was repeated for 42 pedestrian streetscape figures at 10 s per figure with simultaneous eye-tracking, followed by a final rest with tea and chocolate.
In Stage E (post-experiment questionnaire), participants completed a structured questionnaire covering: their post-exposure emotional state; their perceptual comfort ratings of the observed figures; and a preference vote identifying the most comforting building facade and streetscape images. Demographic information and baseline emotional state had been collected prior to Stage A.

3.4. Analytical Methods

A.
Fractal dimension analysis.
Fractal dimension (D) and lacunarity (Λ) were computed for a targeted set of architectural façade images using the box-counting method via the FracLac plugin for ImageJ (version 1.54, 64-bit, bundled with Java 8). The lacunarity formula applied:
λ ϵ , g =   C V ϵ , g 2 =   σ ϵ , g μ ϵ , g 2
Measurement of spatial heterogeneity and translational invariance (Lacunarity) across scale ε. where σ is the standard deviation and μ is the mean of pixel mass distribution across box sizes [36].
Analysis was applied to all comfort-voted figures and a matched set on non-voted figures for comparison. Figures receiving only one or two votes were classified under the standard typology framework but are not foregrounded as key preference cases [36,37].
Fractal dimension measures self-similar visual complexity across spatial scales (higher D = more visually rich), while lacunarity measures spatial irregularity (lower Lac = more uniform arrangement). In the box-counting method, a grid of boxes of side length ε is superimposed over the image, and the number of boxes N(ε) containing any part of the pattern is counted. This procedure is repeated at multiple scales, and D is derived from the scaling relationship:
D = lim ε 0 log N ε log ε
Fractal dimension D via box-counting, where N(ε) = number of boxes of size ε containing the pattern.
Research has shown that human observers consistently prefer environments with moderate-to-high fractal complexity, and that fractal dimension is a reliable predictor of visual preference in both natural and built environments [38,39].
Each facade image was classified into one of four complexity levels based on its D value: Low complexity (D < 1.3), Moderate (1.3 ≤ D < 1.5), High complexity (1.5 ≤ D < 1.7), and Very high complexity (D ≥ 1.7).
The spatial distribution of facade elements was classified using Lacunarity (Λ) into four levels: Very uniform (Λ < 0.3), Moderate variation (0.3 ≤ Λ < 0.6), Irregular (0.6 ≤ Λ < 1.0), and Highly irregular (Λ ≥ 1.0).
The combined interpretation of D and Λ together produces one of six facade typology classes: Simple and uniform (D < 1.5, Λ < 0.3); Simple but irregular (D < 1.5, Λ ≥ 0.3); Moderately complex and balanced (1.5 ≤ D < 1.7, Λ < 0.6); Moderately complex but irregular (1.5 ≤ D < 1.7, Λ ≥ 0.6); Complex but Organized (D ≥ 1.7, Λ < 0.6); and Complex and chaotic (D ≥ 1.7, Λ ≥ 0.6).
To aid the interpretation of these metrics, Figure 7 provides a visual reference using actual building façade stimuli from this study. The top strip shows five images arranged from lowest (D = 1.428) to highest (D = 1.838) fractal dimension, illustrating how increasing D corresponds to progressively richer visual articulation, from plain rendered surfaces to ornately layered façades. The middle strip shows five images arranged from lowest (Λ = 0.203) to highest (Λ = 1.200) lacunarity, illustrating how increasing Λ corresponds to progressively more irregular spatial distribution of visual elements, from rhythmically organized façades to those with unpredictable void-mass patterns. The bottom panel plots all 45 images with fractal data in the D × Λ space, with nine representative examples shown at key positions across the typology range. Green borders indicate images voted comfortable by two or more emotion groups; orange borders indicate images not voted comfortable.
The eye-tracking system was a webcam-based gaze estimation tool built on Google Media Pipe Face Mesh with a Tensor Flow Lite XNNPACK CPU backend, detecting 468 facial landmarks including iris and eyelid positions for real-time gaze estimation. Stimuli were displayed on a 4K monitor at 3840 × 2160 pixels, each shown for 10 s. Calibration was performed individually per participant using a green dot-grid sequence. Gaze data were recorded at approximately 12–22 samples per second as CSV files (timestamp, x, y, elapsed time).
The composite Visual Attention Score (VAS) reported in this study combines three complementary dimensions of visual attention into a single normalized score: dwell time captures the total duration of gaze engagement; fixation count reflects the frequency of discrete attentional events and indicates active scanning; and gaze entropy measures the spatial distribution of fixations across the AOI, with higher entropy indicating more dispersed exploration.
Each metric was min–max normalized to [0, 1] before averaging, ensuring equal contribution regardless of absolute scale. This composite approach follows the rationale that no single metric fully characterizes visual attention. Equal weighting was applied in the absence of an established domain-specific rationale for differential weighting; sensitivity analysis showed that figure rankings were largely insensitive to moderate changes in metric weights. VAS = (normalized dwell time + normalized fixation count + normalized gaze entropy) ÷ 3.
Inter-rater reliability was assessed by computing pairwise Pearson r between participants within each group across all stimulus figures. Dwell time showed moderate-to-acceptable reliability (mean r = 0.47–0.68). Fixation count (mean r = 0.07–0.14) and gaze entropy (mean r = 0.09–0.28) showed weaker agreement, reflecting individual differences in scanning strategy (Figure 8).
Within-stimulus consistency (CV = SD/Mean) confirmed dwell time was highly stable (CV < 0.01) and entropy showed acceptable consistency (CV ≈ 0.12–0.16, below the 0.20 threshold) (Figure 9). These findings support the composite VAS, which averages across metrics to reduce individual-level variability.
As a webcam-based system, spatial precision is lower than commercial infrared trackers and is sensitive to lighting and head movement; gaze may drop out if the face leaves the camera frame. Calibration reliability was assessed via R2: Excellent (R2 > 0.98), Good (0.95 < R2 ≤ 0.98), Acceptable (0.90 < R2 ≤ 0.95), or Low reliability (R2 ≤ 0.90).
The webcam-based system (MediaPipe Face Mesh, TensorFlow Lite) achieves an estimated gaze accuracy of ±2–3° visual angle, consistent with validated webcam-based tracking benchmarks [40], yet lower than professional IR systems like the Tobii Pro Nano (0.30°). While insufficient for granular within-image region analysis, this accuracy is adequate for the full-image AOI design employed here. At a 60–70 cm viewing distance, a ±3° error corresponds to ~3 cm on screen, well within the boundaries of a full facade AOI spanning the 4K display (Figure 10).
A further limitation concerns the use of flat, single-view point photographs rather than 360° imagery. Standard photographs capture only a frontal subset of the visual that a person at the same location would perceive: the sky, peripheral zones, and overhead elements are systematically underrepresented.
Accordingly, visual metric proportions—particularly Sky Ratio and Canopy Ratio—may differ from values obtained under fully immersive conditions, and findings should not be directly extrapolated to real-world perception without acknowledging this constraint.
B.
Pedestrian streetscape visual metric analysis
Six visual metrics were extracted for each of the 42 pedestrian streetscape figures through semantic image segmentation performed manually in ImageJ [37] and tabulated in Microsoft Excel.
The metrics are: Sky Ratio visible open sky); Green Ratio (all vegetation: trees, shrubs, ground planting); Canopy Ratio (tree crown foliage only, a subset of Green Ratio); Hardscape Ratio (paved or built surfaces); Enclosure Ratio (vertical surfaces such as facades and walls); and Occupied Space (areas with visible people).
As Canopy Ratio is a subset of Green Ratio, metric sums can exceed 100% in vegetated scenes; this is expected and reflects intentional measurement design, not segmentation error. Each metric is expressed as a ratio of the pixel area of the target element to the total scene area:
metric   ratio i   = A i A t o t a l   × 100 %
Visual metric ratio, where Ai = pixel area of element i and Atotal = total scene pixel area.
Manual segmentation was adopted in preference to automated semantic segmentation because the pedestrian streetscape images in this study were captured from fixed viewpoints at specific historical locations in Kitakyushu, where automated models trained on large-scale street-view datasets (e.g., Cityscapes, ADE20K) systematically misclassify heritage building facades, port infrastructure, and seasonal foliage.
Manual delineation in ImageJ ensured accurate category assignment for each specific scene. Measurements were performed by a single trained researcher and cross-checked against photographic ground truth; the absence of formal inter-rater reliability assessment is acknowledged as a limitation.
Further work could address depth-plane confounding by applying single-view depth estimation—a deep learning technique that classifies each image pixel as foreground, middle ground, or background—enabling separate analysis of canopy and greenery contributions at different spatial distances from the viewer.
A normalized Combined Score was then computed per figure as the mean of all six normalized metric values, providing a single composite indicator of overall streetscape visual quality:
CS = 1 n i = 1 n M i norm
Combined Score (CS), where n = 6 metrics and M i n o r m = min–max normalized value of metric i, scaled to [0, 1].
Pearson correlation analysis was applied to examine statistical relationships between the six visual metrics and participant emotional responses (Comfortable/Neutral/Uncomfortable). The Pearson correlation coefficient r is calculated as shown in:
r = i = 1 n X i X ¯ Y i Y ¯ i = 1 n X i X ¯ 2   ·   i = 1 n Y i Y ¯ 2
Pearson correlation coefficient r, where Xi = visual metric value, Yi = emotion score, X ¯ and Y ¯ = respective means. Values range from −1 (perfect negative) to +1 (perfect positive correlation).

4. Results

This section presents findings across three analytical dimensions: eye-tracking visual attention, fractal complexity and comfort preference voting, and pedestrian streetscape visual metrics. Comfort voting results are reported within the fractal subsection as the two are directly linked.

4.1. Eye-Tracking Results (Visual Attention Score/VAS)

As presented in Table 1, for the buildings’ facades, the Sad group produced the highest mean VAS (0.686, SD = 0.038), while the Fear group scored lowest (0.636, SD = 0.048). The same pattern held for pedestrian streetscapes (Sad: 0.682; Fear: 0.631). A Kruskal–Wallis test confirmed statistically significant differences across groups for both buildings (H = 43.06, p < 0.001) and pedestrian figures (H = 38.80, p < 0.001).
The absolute mean differences of approximately 0.05 points suggests that the Sad emotional condition may have induced more focused visual scanning of environmental stimuli, while Fear produced slightly more scattered attention. Despite these group level differences, the ranking of individual figures by VAS remained broadly stable across all four emotional conditions, confirming that visual attention patterns are primarily driven by the visual properties of the stimuli rather than the participant’s emotional state.
For building façades, VAS ranged from 0.532 to 0.785 across all four emotion groups. Mojiko buildings produced the widest spread and highest peak values, with Fig. 66 recording the highest single score in the buildings’ dataset (0.785, Sad group), and consistently appearing among the top-ranked figures across emotion conditions. Wakamatsu and Tobata buildings showed narrower, lower-clustered scores, consistent with their more uniform architectural character (Figure 11).
For pedestrian streetscapes, scores ranged from 0.533 to 0.784. The highest single score in the entire pedestrian dataset is Tobata Fig. 13 under Anger (0.784); however, Tobata pedestrian figures (Figs. 13–18, n = 6) are excluded from the per-figure most/least ranking in Table 2, as with only 6 figures, selecting the top 3 and bottom 3 would cover the entire subgroup (Figure 12).
Among Wakamatsu and Mojiko pedestrian figures, the highest scores are Mojiko Fig. 42 under Sad (0.753), Fig. 40 under Happy (0.740), and Fig. 21 under Sad (0.739). The lowest are Mojiko figures under Fear, Fig. 37 (0.533), Fig. 32 (0.559), and Fig. 39 (0.565). Across both stimulus types, the ranking of individual figures by VAS remained broadly stable across emotion groups, confirming that visual attention is primarily driven by stimulus properties rather than emotional state.
Table 2 reveals consistent patterns across emotion groups and study areas. For building façades, Mojiko figures dominate the most-attracted rankings across all four emotion conditions, with Fig. 66 recording the highest peak VAS in the entire buildings dataset (0.785 under Sad) and consistently appearing among the top-ranked figures. Wakamatsu figures recur most frequently among the least attracted, reflecting the area’s lower architectural complexity and more uniform façade character. Tobata shows intermediate performance, with selected most-attracted figures concentrated among its more ornamented heritage structures.

4.2. Fractal Dimension Analysis and Comfort Preferences Results

An important methodological note: the stimulus-level analysis in this section—comparing the fractal profiles of comfort-voted versus non-voted figures across all 73 stimuli rated by all 17 participants—has substantially greater statistical power than the between-group emotional comparisons (n = 4 per group). The Spearman correlation (r = 0.369, p = 0.013) draws on 73 data points and represents a statistically reliable finding that does not depend on the between-group design. Between-group comparisons in this section should be interpreted as descriptive patterns.
Across all four emotion groups, comfort-voted facades are dominated by Very High Complexity (D ≥ 1.70) profiles classified as Complex but Organized, high D combined with low to moderate Λ. This pattern holds regardless of the induced emotion, confirming that high fractal complexity with controlled spatial regularity is a stable driver of comfort preference. All R2 values exceeded 0.98, confirming measurement reliability throughout.
A.
Comfort votes and fractal profiles by emotion group
Table 3 presents the fractal profiles of all comfort-voted facades. Across all four groups, 22 of 23 voted figure instances are classified as Very High Complexity and fall within the Complex but Organized category. The only exception is Fig. 64 (D = 1.517, Λ = 1.200), the sole High Complexity and Highly Irregular facade among the selected figures. Notably, Fig. 64 appears across the Fear, Anger, and Happy groups, indicating a recurring pattern that cannot be explained by fractal complexity alone.
The Fear group (n = 4, mean D = 1.706) produced the lowest number of comfort selections. Most selected figures clustered within the Complex but Organized category (Λ = 0.357–0.456). However, Fig. 64 (D = 1.517, Λ = 1.200) appeared as a unique outlier—the only High Complexity and Mod. Complex but Irregular façade to receive comfort votes—suggesting that under Fear conditions, comfort may be influenced by factors beyond structural complexity.
The Anger group (n = 8, mean D = 1.735) records the highest number of comfort selections and the greatest proportion of Very High Complexity Figures (7 of 8, 88%). Fig. 61 (D = 1.828) represents the most complex comfort-selected façade across all groups. Fig. 64 again appears as the sole irregular outlier.
The Sad group (n = 5, mean D = 1.763, mean Λ = 0.371) records the highest mean complexity (D) and lowest mean lacunarity (Λ) among all groups. Four of the five figures are classified as Very High Complexity. Fig. 58 (D = 1.684, Λ = 0.514) appears as a minor deviation, placing it in the Moderately Complex and balanced category rather than Complex but Organized. Fig. 64 again appears as an irregular outlier, consistent with its presence across all four emotion groups.
The Happy group (n = 6, mean D = 1.724) follows the same dominant pattern, with five figures classified as Very High Complexity and Complex but Organized. Fig. 64 appears for the fourth time across all groups, confirming that its comfort appeal operates independently of its fractal characteristics.
Across all four groups, Fig. 60 (D = 1.760, Λ = 0.379) and Fig. 64 (D = 1.517, Λ = 1.200) are the only two façades selected as comfortable by all four emotion groups. Fig. 60 represents the archetypal preferred profile—Very High Complexity with low lacunarity—while Fig. 64 represents a structural paradox, receiving universal comfort votes despite having the lowest D and highest Λ among all voted figures. Figs. 46 and 56 appear across three groups, both characterized by Very High Complexity and low lacunarity, reinforcing the dominant comfort profile observed throughout.
B.
Most voted and non-voted data comparison
The most-voted group (n = 16, mean D = 1.736) clusters tightly at high complexity (D) with low to moderate lacunarity (Λ), with 81% of figures classified as Very High Complexity. In contrast, the not-voted group (n = 29, mean D = 1.664) shows a more dispersed distribution, with only 45% Very High Complexity and lacunarity values extending up to 0.787. The observed difference in mean D (0.072) and the substantially higher proportion of Very High Complexity figures indicate that higher fractal complexity combined with lower spatial irregularity is associated with preferred facades (Figure 13).
C.
Fractal analysis by study area
Across the three study areas, a consistent gradient is observed between architectural complexity and comfort consensus. Wakamatsu (mean D = 1.602) produced zero comfort votes across all four emotion groups. Tobata (mean D = 1.653) produced three comfort-voted figures, with the three selected buildings (Figs. 24, 26 and 36) separating clearly at higher D and lower Λ from non-voted figures; Fig. 26 (D = 1.838, Λ = 0.203) is the most complex and spatially uniform façade in the entire dataset. Mojiko (mean D = 1.710) produced the strongest concentration, with 19 of 28 façades receiving comfort selections from at least one emotion group, the majority clustering at D > 1.70 and Λ < 0.50. This area-level gradient—from zero selections in Wakamatsu to the highest consensus in Mojiko—indicates that higher architectural complexity is associated with increased comfort consensus across emotion groups.
This gradient is further confirmed in Figure 14, which shows the distribution of fractal dimension D and lacunarity Λ by comfort vote consensus level (0 = not voted to 4 = all four groups voted). Figures receiving a higher consensus tend to cluster above D = 1.70 and below Λ = 0.60, supporting the positive relationship between complexity and comfort consensus (Spearman r = 0.369, p = 0.013).
For building façades, eight images received comfort votes from all four emotion groups: Fig. 36 (Tobata) and Figs. 46, 50, 56, 60, 61, 62, and 64 (all Mojiko). Among these, six images—Fig. 36, 46, 50, 56, 60, and 61—are classified as Very High Complexity with low lacunarity (Λ < 0.46), and their fractal profiles follow the expected pattern of organized visual richness. Fig. 60 (D = 1.760, Λ = 0.379) is taken as the archetypal preferred profile, combining Very High Complexity with low lacunarity and receiving consistent votes across all emotion conditions. Fig. 62 (D = 1.627, Λ = 0.441) is classified as High Complexity with a balanced spatial distribution and sits close to the threshold between the two groups. Fig. 64 (D = 1.517, Λ = 1.200) is the structural paradox of the dataset: it is the only image classified as High Complexity with irregular spatial distribution, yet it received comfort votes from all four emotion groups. This case is discussed separately as the civic legibility paradox (Figure 15).
At the group level, the Anger group recorded the highest number of comfort selections for building façades (n = 35 images), while the Fear group recorded the lowest (n = 10 images). This wide gap suggests that the Anger condition produces a broad and inclusive comfort response across many façade types, whereas the Fear condition produces a highly selective pattern, concentrating comfort votes on images with strong visual structure. The Sad and Happy groups fall between these extremes (Sad: n = 23; Happy: n = 14).

4.3. Pedestrian Streetscape Visual Metric Results

Three main findings emerge. First, Enclosure Ratio is the only metric that consistently differentiates comfortable from uncomfortable streetscapes across all three areas. Second, vegetation is associated with higher comfort in Wakamatsu and Mojiko, but shows an opposite pattern in Tobata. Third, no single metric alone significantly predicts emotion perception.
A.
Summary heatmap
None of the six Tobata figures (Figs. 13–18) received a Comfortable label; all are classified as Neutral or Uncomfortable, with Combined Scores ranging from 0.22 to 0.37, the lowest range across all three areas. In Wakamatsu, only Figs. 3 and 5 are Comfortable, both showing markedly higher Green Ratio (30.4% and 43.9%) and Canopy Ratio (29.9% and 37.1%) than surrounding Uncomfortable figures. Figure 12 records the lowest Combined Score in the entire dataset (CS = 0.16), corresponding to near-zero vegetation cover. In Mojiko, the three Comfortable figures—Figs. 29, 32, and 41 (CS = 0.37–0.39)—consistently show high Tree Canopy and low Enclosure Ratio, a combination absent from all Uncomfortable-labeled Mojiko figures (Figure 16).
Figure 15. Distribution of Fractal Dimension (D) and Lacunarity (Λ) for Comfortable (≥2 group votes, n = 14 images) and Uncomfortable (≤1 vote, n = 31 images) building façades (total n = 45). Thresholds at D = 1.702 and Λ = 0.436 (group mean midpoints); Mann–Whitney D p = 0.029 *, Λ p = 0.093 (ns). Image 64 is a structural paradox.
Figure 15. Distribution of Fractal Dimension (D) and Lacunarity (Λ) for Comfortable (≥2 group votes, n = 14 images) and Uncomfortable (≤1 vote, n = 31 images) building façades (total n = 45). Thresholds at D = 1.702 and Λ = 0.436 (group mean midpoints); Mann–Whitney D p = 0.029 *, Λ p = 0.093 (ns). Image 64 is a structural paradox.
Architecture 06 00091 g015
B.
Mean metric comparison by emotion group and area
Enclosure Ratio is consistently lower in Comfortable figures across all areas without exception: Wakamatsu (2.0% vs. 7.1%), Tobata (9.6% vs. 16.3%), and Mojiko (5.1% vs. 17.2%). Sky Ratio shows a substantial difference in Wakamatsu (49.7% vs. 22.1%), suggesting that open sky may compensate for low vegetation; however, this difference becomes less pronounced in Tobata and Mojiko. Green and Canopy Ratios are higher in Comfortable figures in Wakamatsu and Mojiko, but show a reversed pattern in Tobata, where Uncomfortable figures exhibit higher vegetation values. This contrast appears to be specific to Tobata’s streetscape configuration and should be interpreted with caution due to the small sample size (n = 6). In contrast, Hardscape Ratio does not show a consistent directional pattern across areas or emotion groups (Figure 17).
C.
Pearson correlation analysis
A near-perfect correlation between Canopy and Green Ratio (r = 0.98, p < 0.001) reflects the compositional relationship between these metrics: Canopy Ratio is a subset of Green Ratio, measuring only tree crown foliage within the broader vegetation category. Streets with abundant tree cover consistently produce high values for both, confirming that mature street trees are the dominant vegetation form in the study areas.
The two metrics capture complementary dimensions of the same environmental quality rather than independent variables. Enclosure Ratio shows strong negative correlations with green (r = −0.64, p < 0.001) and Canopy (r = −0.62, p < 0.001), suggesting that more enclosed streets tend to lack vegetation, while openness and greenery co-occur. Sky Ratio is moderately negatively correlated with Canopy (r = −0.47, p < 0.01), reflecting a visual trade-off between shaded and open sky conditions. Importantly, no individual metric demonstrates a significant correlation with emotion perception (all |r| < 0.40; highest: Canopy r = 0.24, Enclosure r = −0.23), indicating that perceived comfort in pedestrian environments is associated with the combined spatial configuration of enclosure, vegetation, and openness rather than any single visual factor (Figure 18).

4.4. Cross-Method Results Synthesis

Cross-method correlation analysis examined the relationships between the three independent measurement approaches—fractal analysis, eye-tracking (VAS), and self-reported comfort votes—for the 73 building façade stimuli (Figure 19).
Step 1. Correlation 1—D vs. comfort vote consensus (Spearman r = 0.369, p = 0.013, significant): Fractal complexity significantly predicted comfort vote consensus across all emotion groups. The D = 1.70 threshold visually separates most-voted from non-voted figures in the histogram distribution, confirming that fractal dimension is a statistically reliable, emotion-independent predictor of comfort preference.
Step 2. Correlation 2—Λ vs. comfort votes (Spearman r = −0.216, p = 0.154, not significant): Lacunarity did not significantly predict comfort vote consensus. The distributions of comfortable and uncomfortable figures overlap substantially across the Λ range, and the exploratory threshold at Λ = 0.436 (midpoint between group means) does not produce a clean separation. Lacunarity therefore functions as a secondary descriptive classifier rather than an independent predictor and is retained as a companion measure to D rather than a standalone comfort indicator.
Step 3. Correlation 3—VAS vs. comfort votes (r = 0.097, p = 0.526, not significant): No relationship was found between visual attention and comfort vote consensus. Participants attended visually to all stimuli at similarly high levels regardless of whether they voted them as comfortable. Attention is necessary but not sufficient for comfort preference; eye-tracking and self-report capture fundamentally different dimensions of the perception process.
Step 4. Correlation 4—D vs. VAS (r = 0.015, p = 0.829, not significant): No relationship was found between fractal complexity and visual attention. Participants did not systematically attend more to high-D façades. Fractal complexity predicts comfort preference through a pathway independent of attentional capture; the reflective comfort judgment, not the pre-reflective gaze, is sensitive to fractal organization.
Step 5. The per-figure method agreement analysis shows: six figures with full positive agreement (comfort voted + high VAS + D ≥ 1.70), all located in Mojiko; eight with full negative agreement (not voted + low VAS + D < 1.70), predominantly in Wakamatsu; and 59 mixed. The most frequent mixed pattern is ‘Vote + D agree (VAS independent)’, confirming that attentional engagement does not predict the direction of comfort preference, only its occurrence (Figure 19).

4.5. Most Comfortable Figures (What Makes Them Comfortable)

Figure 20 presents the top 12 building façades and top 10 pedestrian streetscapes by comfort consensus, with their metric profiles, self-reported votes per emotion group, and VAS per emotion group shown side by side.
For building façades, the metric profile reveals that all top figures except Fig. 62 and Fig. 64 exceed the D = 1.70 threshold with Λ < 0.60. Across all top buildings, vote distributions are broadly consistent across emotion groups; the Anger group produces the highest counts and Fear the lowest, but the same figures are selected by all groups. VAS is uniformly high across all top figures and all emotion groups, confirming strong visual engagement regardless of emotion.
For pedestrian streetscapes, two distinct comfort pathways are visible: Wakamatsu figures achieve high comfort through high Canopy Ratio (37–49%) with Very Low Enclosure; Mojiko waterfront figures (Ped 39, 42) achieve comfort through high Sky Ratio (29–39%) and spatial openness with minimal canopy. This confirms that comfort in pedestrian environments can be achieved through either the vegetation–canopy pathway or the openness–sky pathway depending on the spatial context (Figure 20).

5. Discussion

5.1. Participant Profile and Emotion Induction Validity

The sample (n = 17; 15 male, 2 female; predominantly architecture-trained) introduces two specific limitations. First, architecture-trained observers may demonstrate heightened sensitivity to façade complexity and ornamental detail relative to general populations; the findings should therefore be understood as an upper-bound sensitivity estimate. If even this sensitized group shows largely emotion-independent comfort responses, the pattern is likely to be at least as robust in general populations. Second, gender differences in emotional reactivity to environmental stimuli have been documented [32], and the near-absence of female participants means gender as a moderating variable cannot be assessed.
The partial effectiveness of emotion induction in Sad and Happy groups—where fatigue dominated the post-experiment self-report—means that between-group comparisons involving these groups should be treated as exploratory. A further limitation concerns the maintenance of induced emotional state during the observation sessions. Continuous emotion-congruent audio was played throughout (Section 3.3), but without mid-session physiological monitoring (e.g., skin conductance, heart rate), it cannot be confirmed that the induced state was sustained across all 115 stimuli, particularly toward the end of the 50 min viewing period [35]. Future studies should include either periodic stimulus re-exposure or continuous physiological monitoring. An online supplementary study with larger samples would substantially increase statistical power (following the approach of [41]).
Replication with a larger, more demographically diverse sample is recommended before generalizing these findings beyond this specific pilot-scale context.

5.2. Building Façade Comfort: Fractal Complexity and Spatial Organization

The fractal analysis and comfort selection results indicate that façade visual complexity, measured by fractal dimension (D) and lacunarity (Λ), is a consistent predictor of comfort preference across all four emotional conditions. This finding aligns with established research showing that human visual systems are naturally attuned to fractal patterns, with moderate-to-high fractal dimensions (D ≈ 1.3–1.8) being most preferred [38,42].
In architectural contexts, façades with higher D provide layered visual information across multiple spatial scales, allowing viewers to engage with both overall form and fine detail without cognitive overload. At the same time, low lacunarity (Λ) reflects a more uniform and organized spatial distribution of elements, which corresponds to the concept of coherence in environmental preference theory [12].
Together, these properties define a “Complex but Organized” visual profile, façades that are visually rich yet spatially legible. This combination is consistently associated with higher comfort ratings across all emotional groups, suggesting that it functions as a stimulus-driven predictor of comfort, largely independent of the observer’s emotional state.
However, the variation observed in the most and least preferred façades indicates that additional factors—such as architectural style, historical identity, perceived function, and environmental context—also contribute to comfort perception and are not fully captured by fractal metrics alone.
A.
Most comforting facades (Shared visual qualities)
The façades that received comfort selections across all four emotion groups include Fig. 60 (4/4 groups), Fig. 46 (4/4 groups), and Fig. 56 (4/4 groups), which share a consistent fractal profile, characterized by Very High Complexity (D ≥ 1.70) and low to moderate lacunarity, corresponding to the Complex but Organized category. Beyond these quantitative properties, these facades also exhibit a set of shared architectural qualities that are not fully captured by fractal metrics alone [Figure 21].
Fig. 60 is a Meiji-era red brick heritage building in Mojiko, characterized by layered Tudor-style half-timbering, conical corner towers, decorative chimneys, arched window surrounds, and a highly articulated roofline. This richness of surface detail corresponds to its high fractal dimension (D = 1.760), while the organized repetition of these elements maintains relatively low lacunarity (Λ = 0.379). Participants showed divided interpretations of its function, with equal classification as a Commercial and Special District building (three votes each), suggesting that functional ambiguity does not limit comfort response. Instead, visual richness appears to outweigh functional legibility. Its selection across all four emotion groups further indicates that the combination of ornamental richness and spatial coherence is associated with a consistent, emotion-independent perception of comfort.
Fig. 46 (D = 1.826, Λ = 0.253), the former Moji Railway Station, is a large Taisho-era public building characterized by a symmetrical Neo-Baroque facade, copper-domed roofline, monumental columns, and decorative balustrades. Participants predominantly identified it as a Special District building (four votes), correctly recognizing its civic and heritage landmark status. This clear functional legibility as a public monument, combined with Very High Complexity and the lowest lacunarity among the key comfort facades, suggests that participants associate it with both visual richness and a sense of social safety, potentially reinforced by its identity as a recognized civic space.
Fig. 56 (D = 1.718, Λ = 0.388) presents a distinct case: all participants identified it as Residential (7/7 votes), making it the only comfort-selected facade not perceived as a public or heritage landmark. Visually, it is a Showa-era structure combining half-timbered upper sections with traditional Japanese tiled roofing and dense hedge planting at street level. This hybrid architectural composition comprising Western structural articulation integrated with Japanese roof forms and domestic vegetation suggests that comfort can also emerge from a sense of homeliness and familiarity rather than civic monumentality. The presence of vegetation at street level further aligns this facade with the pedestrian findings, where tree canopy and greenery are consistently associated with higher comfort perception.
B.
Anomaly (Fig. 64) comfort beyond complexity
Fig. 64 (“Minato House”) represents a key outlier in the dataset. With D = 1.517 and Λ = 1.200, it is the only High Complexity, Highly Irregular facade to receive comfort selections, appearing in three of the four emotion groups (Fear, Anger, Sad, and Happy). Its fractal profile contrasts sharply with the other preferred facades, exhibiting lower complexity and higher irregularity, and is classified as Moderately Complex but Irregular [Figure 22].
Visually, Fig. 64 is a contemporary minimalist structure, defined by a plain white gabled volume, a single circular window, and a wide arched entrance which contrasts sharply with the ornamental richness of Figs. 60 and 46. However, design perception data provides a key explanatory layer: participants predominantly identified it as a Special District building (five votes), correctly recognizing its role as the Minato House visitor and information center within the Mojiko Retro heritage area.
This civic and publicly accessible identity, supported by contextual cues such as the open plaza setting, the presence of urban furniture, and visible pedestrian activity, appears to evoke comfort through associative recognition and familiarity rather than visual complexity. This interpretation aligns with theories of environmental perception suggesting that cognitive mapping and place identity influence the emotional evaluation of urban environments [24], as well as research showing that familiar and legible environments can reduce stress and increase perceived safety [43].
Accordingly, this case suggests that typological identity and contextual social meaning can override the complexity to comfort relationship, particularly under emotionally activated states such as Fear and Anger, where familiarity and civic legibility may provide reassurance.
C.
Least comforting facades
Buildings receiving zero comfort selections across all four emotion groups were predominantly located in Wakamatsu, which has the lowest mean fractal dimension (D = 1.602). Design perception data provides an additional interpretive layer: Fig. 4 was unanimously identified as Industrial (7/7 votes), while Fig. 10 was primarily perceived as Residential (four votes). Despite these differing functional identities, neither received any comfort selections. This pattern suggests that the absence of visual complexity, rather than perceived function alone, is associated with non-comfort responses (Figure 23).
Fig. 4 (D = 1.628) is a low-rise building characterized by a flat facade, repetitive horizontal striping in yellow and gray, and minimal window articulation, lacking ornamentation, textural variation, or hierarchical differentiation between structural and decorative elements. Fig. 10 (D = 1.428) is a narrow three-story building with a plain rendered surface and uniformly repeated rectangular windows across all levels. Both are classified at or below Moderately Complex.
In contrast to the comfort-selected facades in Mojiko, this comparison reveals a consistent visual pattern: comfort preference is associated with facades that provide layered visual information across multiple viewing scales, including overall massing and silhouette at a distance, and surface detail, texture, and material transitions at closer range. By comparison, Wakamatsu facades offer limited visual information at both scales.
D.
Design perception and its relationship to comfort
Design perception results reveal a clear pattern: buildings identified as Special District are typically associated with heritage, civic, or tourism functions which account for the majority of comfort-selected facades across all emotion groups. In contrast, buildings unanimously perceived as Industrial (Fig. 4: 7/7 votes) received no comfort selections in any group, representing the strongest categorical distinction in the dataset.
A key comparative case is Fig. 56, which was unanimously perceived as Residential (7/7) yet received comfort selections in three of four groups, indicating that domestic familiarity combined with high visual complexity (D = 1.718) can generate strong comfort responses. Conversely, Fig. 10, also perceived as Residential, received no comfort selections, suggesting that residential identity alone is insufficient when expressed with low visual complexity. This contrast indicates that fractal complexity (D) remains a primary factor in comfort perception even when functional interpretation is held constant.

5.3. What Makes a Pedestrian Streetscape Comforting

The pedestrian analysis identifies three interrelated qualities that distinguish comfortable from uncomfortable streetscapes: spatial openness, tree canopy, and the absence of dominant vehicular infrastructure. However, these factors do not consistently co-occur. Notably, the most-selected pedestrian figure, Fig. 42, achieves a high level of comfort primarily through spatial openness, despite limited vegetation.
A.
Most comforting pedestrian streetscapes
Fig. 42 received the highest total comfort votes in the pedestrian dataset (11 votes across all four emotion groups), making it the most strongly and consistently preferred pedestrian figure. It is one of 16 figures (38% of all pedestrian stimuli) voted by all four emotion groups but stands out for the breadth and strength of its votes across all conditions (Figure 24).
This pattern suggests that its perceived comfort is primarily associated with spatial openness, including expansive views of water, harbor activity, and sky. The case demonstrates that openness alone, even in the absence of significant green elements, can function as a key driver of comfort in pedestrian environments.
Seven additional figures appear in the comfort selections of three out of four emotion groups, all combining tree canopy with open spatial configurations. Fig. 38, 29, 41, and 39 (Mojiko), and Figs. 4, 5, and 6 (Wakamatsu) exhibit consistently low Enclosure Ratios (1.0–8.0%) alongside high Green and Canopy values (Green: 36–57%, Canopy: 19–49%), indicating that vegetation and openness operate together to support comfort perception.
Fig. 38 is a wide paved promenade in Mojiko, lined with autumn-colored trees and active pedestrian use. Its combination of moderate tree canopy (13.2%) and low enclosure (10.4%) creates a spatial condition that is sheltered yet not confined. Fig. 29 represents the Mojiko station plaza, an active civic space with market activity, street trees, and a strong pedestrian presence, characterized by the lowest Enclosure Ratio (1.3%) and Substantial Canopy (31.3%). Fig. 41 is an open, tree-lined promenade with distant mountain views and seasonal foliage, combining low enclosure (6.1%) with a relatively high Canopy Ratio (33.7%). Together, these examples illustrate how the integration of tree canopy and spatial openness can produce comfortable pedestrian environments that balance exposure and enclosure (Figure 25).
Among the Wakamatsu Comfortable figures, Fig. 5 records the highest Combined Score in the pedestrian dataset (CS = 0.50) and the highest Sky Ratio among all comfort-selected figures (80.5%), alongside Substantial Canopy (37.1%) and Very Low Enclosure (2.3%). It represents a wide, tree-lined road opening toward the waterfront, with the Wakamatsu suspension bridge visible, creating a highly open spatial condition (Figure 26).
This combination of expansive sky, mature tree canopy, and minimal enclosure forms a balanced spatial profile associated with strong comfort perception. Figs. 4 and 6 in Wakamatsu exhibit similar characteristics with tree-lined paths with open water access and low enclosure, and both appear in comfort selections in three of four emotion groups, indicating that this environmental configuration is consistently associated with comfort across emotional states.
B.
Least comforting pedestrian streetscapes
Unlike the building façade dataset, no pedestrian figure received zero comfort votes, all 42 figures were selected by at least one emotion group. However, figures with the lowest vote counts and lowest consensus represent the least-preferred environments. The lowest-performing figures are those voted by only one or two emotion groups: Ped 1 (Wakamatsu, 2/4 groups), Ped 3 (Wakamatsu, 2/4 groups), Ped 13 (Tobata, 2/4 groups), and Ped 20 (Mojiko, 1/4 groups). Notably, Ped 16 (Tobata) was voted by all four groups despite being a relatively enclosed streetscape, suggesting that Tobata figures as a whole received more consistent comfort responses than the original analysis indicated.
Fig. 8 (Wakamatsu) represents the lowest-consensus case among comfortable figures: while selected by 3/4 emotion groups (Fear, Anger, Sad), it received the lowest total vote count in the dataset. It depicts a bare road junction with minimal vegetation (Canopy = 0.3%, Green = 1.2%), relatively high enclosure (16.1%), and a visual field dominated by road surfaces, traffic signals, and port-related infrastructure. In contrast, Figs. 4–6 within the same Wakamatsu area—also road-based environments—are lined with mature trees and received comfort selections from all four emotion groups. This comparison confirms that street tree canopy is the critical differentiating factor between high-consensus and low-consensus comfort environments within the same spatial context.
This pattern suggests that the presence of vegetation alone is insufficient to generate comfort. Instead, the spatial integration of greenery with pedestrian pathways, particularly through continuous canopy coverage, appears to be a critical factor. Isolated vegetation within vehicle-dominated environments does not produce the same perceptual effect as tree-lined, walkable spaces (Figure 27).
C.
Pedestrian Fig. 3 metric (Vote divergence)
Fig. 3 of data stimuli (Wakamatsu) presents a particularly revealing case. Although classified as Comfortable in the heatmap (CS = 0.34; Enclosure = 1.8%; Green = 30.4%; Canopy = 29.9%), it received no comfort selections across all four emotion groups. The scene depicts a crosswalk at a road intersection viewed from the pedestrian waiting position. While trees and the waterfront are visible in the background, the immediate perceptual field is dominated by road surface and traffic infrastructure, with no overhead canopy within the walking zone.
This discrepancy highlights a limitation of pixel-based visual metrics: they capture visible elements within the scene but do not account for their spatial relationship to the observer. In this case, vegetation is present but not experientially accessible to the pedestrian. The result suggests that comfort in pedestrian environments depends not only on visual content, but also on the spatial integration between the individual and surrounding elements (Figure 28).
This divergence between metric-based classification and participant responses is also evident in Fig. 42, the most frequently selected comfort figure. Despite being classified as Neutral in the heatmap (Combined Score = 0.24), due to its limited vegetation and moderate sky values it was consistently identified as comforting across all emotion groups. Together with the case of Fig. 3, this contrast indicates that metric-based evaluation and participant responses capture distinct yet complementary dimensions of comfort. While the former reflects quantifiable visual attributes, the latter incorporates experiential and perceptual factors. Neither approach alone appears sufficient to fully explain comfort perception in pedestrian environments.

5.4. The Role of Induced Emotion in Comfort Perception

Comfort preferences were broadly stable across emotional conditions, consistent with the quantitative findings reported in Section 4.2 and Section 4.3. Induced emotion appears to influence the selectivity rather than the direction of comfort responses.
For building façades, the Anger group showed the broadest comfort selections while Sad and Fear groups showed narrower, more structured preferences, a pattern consistent with the fractal profiles reported in Section 4.2.
For pedestrian streetscapes, the Fear group was the most selective (n = 3), focusing on open waterfront environments, while the Sad group showed the widest range of selections (n = 27) across all areas. This may indicate different patterns of environmental preference under varying emotional states.
These findings are broadly consistent with theories suggesting that negative emotional states can increase attention to restorative environmental features [43,44]. At the same time, the overall consistency of selected figures suggests the presence of relatively stable perceptual preferences, particularly for organized visual complexity in facades and spatial openness in streetscapes.
Although emotion induction was partially effective in the Sad and Happy groups (see Section 5.1), their comfort selection patterns were comparable to the Fear and Anger groups. Key figures (Figs. 60 and 42) were consistently selected across conditions. This indicates that environmental characteristics likely played a central role in shaping comfort perception, alongside any effects of induced emotional state.

5.5. Design Implications

Three evidence-based design principles emerge from the convergent multi-method findings:
Principle 1. Target the ‘Complex but Organized’ façade profile: Design interventions on building façades should aim for D ≥ 1.70 combined with Λ < 0.60. In practice, this corresponds to façades with layered ornamental detail that is spatially organized: heritage-style articulation, rhythmic window patterns, material transitions, and vertical differentiation. Simple rendered surfaces and monotonous repetition consistently failed to generate comfort responses regardless of emotional state.
Principle 2. Prioritize the spatial integration of greenery over total vegetation quantity: Pedestrian comfort is determined by whether tree canopy is spatially accessible—along the walking path, with overhead coverage—not by how much vegetation is visible in the scene. Design guidelines should specify continuous canopy along pedestrian corridors, with a target Enclosure Ratio below 15% and Canopy Ratio above 20% within the immediate walking zone.
Principle 3. Leverage civic legibility as an independent comfort pathway: Buildings functioning as legible civic anchors—publicly accessible, socially activated, contextually identifiable—can generate comfort responses even when their physical complexity falls below the fractal threshold. Urban comfort evaluation frameworks should integrate typological identity and social activation alongside quantitative visual measures, particularly for contemporary buildings inserted into heritage contexts.

6. Conclusions

This pilot study examined how experimentally induced emotional states (Fear, Anger, Sad, Happy) interact with the visual properties of urban environments to shape comfort perception across three areas of Kitakyushu, Japan (Wakamatsu, Tobata, Mojiko). Using a multi-method approach combining fractal analysis, pedestrian visual metrics, webcam-based eye-tracking, and self-reported comfort votes, four research objectives were addressed:
Regarding Objective 1 (Visual properties and comfort): The ‘Complex but Organized’ fractal profile (D ≥ 1.70, Λ < 0.60) was a consistent, emotion-independent predictor of building façade comfort selection (Spearman r = 0.369, p = 0.013). For pedestrian streetscapes, low Enclosure Ratio was the most consistent comfort predictor, with spatially integrated tree canopy as a secondary predictor. Induced emotion modulated the breadth of comfort selections (Anger: most selections; Fear: fewest) but did not determine which environments were preferred.
Regarding Objective 2 (Visual attention patterns): Eye-tracking confirmed that visual attention (VAS) was driven by stimulus properties rather than emotional state. VAS showed no significant relationship with comfort votes (r = 0.097, ns) or with fractal dimension (r = 0.015, ns), establishing that visual attention and comfort preference are independent dimensions that capture fundamentally different aspects of the perception process.
Regarding Objective 3 (Emotional state and comfort selectivity): Emotion induction was effective for Fear and Anger and partial for Sad and Happy (fatigue-dominant). Despite this variability, comfort vote patterns were consistent across all four groups, confirming that emotional context operates as a moderating rather than determining variable. Between-group comparisons are descriptive given the pilot-scale sample (n ≈ 4 per group).
Regarding Objective 4 (Urban typology): The three study areas exhibited distinct and persistent comfort profiles: Wakamatsu (mean D = 1.602, no 4/4 group façade selections) then Tobata (1 figure at 4/4), then Mojiko (7 figures at 4/4, 19 of 28 façades voted). For pedestrian streetscapes, all 42 figures received votes; the 16 figures with 4/4 consensus include Mojiko waterfront, Tobata promenades, and Wakamatsu tree-lined roads.
Beyond these objectives, multi-method synthesis identified five empirical paradoxes—cases where single-method predictions fail—and three design principles that follow from the convergent evidence: target the Complex but Organized façade profile, prioritize spatially integrated canopy over visible greenery quantity, and leverage civic legibility as an independent comfort pathway.
Future investigations should also consider depth-stratified visual analysis using monocular depth estimation networks (e.g., MiDaS, DPT), which could enable fine-grained differentiation between foreground and background vegetation, addressing a current limitation of pixel-based segmentation metrics.
As a pilot investigation, findings are directional rather than confirmatory; replication with a larger, more demographically diverse sample across additional urban contexts is necessary before the design principles can be generalized.

Author Contributions

S.A.P.: Writer, Conceptualization, Methodology, Experiment, Investigation, Data curation. B.J.D.: Reviewer, Supervisor. D.H.: Eye-tracking preparation, Experiment assistant, Reviewer. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Ethical review and approval were waived for this study in accordance with the Act on the Protection of Personal Information (APPI, Act No.57 of 2003, amended 2022) of Japan and the Ethical Guidelines for Life Science and Medical Research Involving Human Subjects (MEXT; MHLW; METI, 2021), which exempt observational studies using anonymized data and non-invasive behavioral measurement from requiring ethics committee review. All participant data were fully anonymized prior to analysis, and all participants provided voluntary informed consent to participate and agreed to the use of their data for research and publication purposes.

Informed Consent Statement

Informed consent was obtained from all participants involved in the study. Each participant was briefed on the research objectives, experimental procedure, and intended use of their data before providing written informed consent. All collected data were fully anonymized prior to analysis.

Data Availability Statement

The data supporting the findings of this study are not publicly available due to privacy and ethical restrictions concerning human participant data. Anonymized data may be made available upon reasonable request to the corresponding author.

Conflicts of Interest

The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Satrio Agung Perwira reports financial support was provided by the Government of Japan’s Ministry of Education, Culture Sports, Science, and Technology. If there are other authors, they declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

References

  1. Ewing, R.; Handy, S. Measuring the Unmeasurable: Urban Design Qualities Related to Walkability. J. Urban Des. 2009, 14, 65–84. [Google Scholar] [CrossRef]
  2. Vartanian, O.; Navarrete, G.; Chatterjee, A.; Fich, L.B.; Gonzalez-Mora, J.L.; Leder, H.; Modroño, C.; Nadal, M.; Rostrup, N.; Skov, M. Architectural design and the brain: Effects of ceiling height and perceived enclosure on beauty judgments and approach-avoidance decisions. J. Environ. Psychol. 2015, 41, 10–18. [Google Scholar] [CrossRef]
  3. Nasar, J. Urban Design Aesthetics: The Evaluative Qualities of Building Exteriors. Environ. Behav. 1994, 26, 377–401. [Google Scholar]
  4. Zhang, Z.; Zhuo, K.; Wei, W.; Li, F.; Yin, J.; Xu, L. Emotional Responses to the Visual Patterns of Urban Streets: Evidence from Physiological and Subjective Indicators. Int. J. Environ. Res. Public Health 2021, 18, 9677. [Google Scholar] [CrossRef]
  5. Higuera-Trujillo, J.L.; Maldonado, J.L.-T.; Llinares Millán, M.C. Psychological and physiological human responses to simulated and real environments: A comparison between Photographs, 360° Panoramas, and Virtual Reality. Appl. Ergon. 2017, 65, 398–409. [Google Scholar] [CrossRef] [PubMed]
  6. Holmqvist, K.; Nyström, M.; Andersson, R.; Dewhurst, R.; Jarodzka, H.; van de Weijer, J. Eye Tracking: A Comprehensive Guide to Methods and Measures; Oxford University Press (OUP): Oxford, UK, 2011. [Google Scholar]
  7. Duchowski, A. Eye Tracking Methodology: Theory and Practice; Springer International Publishing: Cham, Switzerland, 2017. [Google Scholar]
  8. Naik, N.; Philipoom, J.; Raskar, R.; Hidalgo, C. Streetscore—Predicting the Perceived Safety of One Million Streetscapes. In IEEE Conference on Computer Vision and Pattern Recognition Workshops; Institute of Electrical and Electronics Engineers (IEEE): New York, NY, USA, 2014; pp. 793–799. [Google Scholar] [CrossRef]
  9. Dubey, A.; Naik, N.; Parikh, D.; Raskar, R.; Hidalgo, C. Deep Learning the City: Quantifying Urban Perception at a Global Scale. In European Conference on Computer Vision (ECCV); Springer International Publishing: Cham, Switzerland, 2016; pp. 196–212. [Google Scholar] [CrossRef]
  10. Li, X.; Zhang, C.; Li, W.; Ricard, R.; Meng, Q.; Zhang, W. Assessing street-level urban greenery using Google Street View and a modified green view index. Urban For. Urban Green. 2015, 14, 675–685. [Google Scholar] [CrossRef]
  11. Stamps, A., III. Psychology and the Aesthetics of the Built Environment; Kluwer Academic Publishers: Boston, MA, USA, 2000. [Google Scholar] [CrossRef]
  12. Kaplan, R.; Kaplan, S. The Experience of Nature: A Psychological Perspective; Press Syndicate of the University of Cambridge: Cambridge, UK, 1989. [Google Scholar]
  13. Bar, M.; Neta, M. Visual elements of subjective preferences: The case of curvature. Psychol. Sci. 2006, 17, 645–648. [Google Scholar] [CrossRef]
  14. Alexander, C. The Nature of Order: An Essay on the Art of Building and the Nature of the Universe, Book One: The Phenomenon of Life; Center for Environmental Structure (Later distributed by Taylor & Francis/Routledge): Berkeley, CA, USA, 2002. [Google Scholar]
  15. Rapoport, A. History and Precedent in Environmental Design; Plenum Press: New York, NY, USA, 1990. [Google Scholar] [CrossRef]
  16. Gehl, J. Cities for People; Island Press: London, UK, 2010. [Google Scholar]
  17. Carmona, M. Contemporary Public Space: Critique and Classification, Part One: Critique. J. Urban Des. 2010, 15, 123–148. [Google Scholar] [CrossRef]
  18. Ewing, R.; Clemente, O. Measuring Urban Design: Metrics for Livable Places; Island Press: Washington, WA, USA, 2013. [Google Scholar] [CrossRef]
  19. Wolf, K. Public Response to the Urban Forest in Inner-City Business Districts. Arboric. Urban For. 2003, 29, 117–126. [Google Scholar] [CrossRef]
  20. Kardan, O.; Gozdyra, P.; Misic, B.; Moola, F.; Palmer, L.; Paus, T.; Berman, M. Neighborhood greenspace and health in a large urban center. Sci. Rep. 2015, 5, 11610. [Google Scholar] [CrossRef]
  21. Sussman, A.; Hollander, J.B. Cognitive Architecture: Designing for How We Respond to the Built Environment; Routledge: New York, NY, USA, 2015. [Google Scholar]
  22. Southworth, M. Designing the Walkable City. J. Urban Plan. Dev. 2005, 131, 246–257. [Google Scholar] [CrossRef]
  23. Whyte, W.H. The Social Life of Small Urban Spaces; Conservation Foundation: Washington, WA, USA, 1980. [Google Scholar]
  24. Lynch, K. The Image of the City; The MIT Press: Cambridge, MA, USA, 1960. [Google Scholar]
  25. Carmona, M.; Heath, T.; Oc, T.; Tiesdell, S. Public Places, Urban Spaces: The Dimensions of Urban Design, 2nd ed.; Architectural Press: Oxford, UK, 2010. [Google Scholar]
  26. Russell, J.A. A circumplex Model of Affect. J. Personal. Soc. Psychol. 1980, 39, 1161–1178. [Google Scholar] [CrossRef]
  27. Mehrabian, A.; Russell, J. An Approach to Environmental Psychology; MIT Press: Cambridge, MA, USA, 1974. [Google Scholar]
  28. Wang, Z.; Yao, Y.; Liu, Y.; Helbich, M. The relationship between urban greenness and mental health: A nationwide study in China. Landsc. Urban Plan. 2023, 238, 104830. [Google Scholar] [CrossRef]
  29. Stamps, A., III. Use of Photographs to Simulate Environments: A Meta-Analysis. Percept. Mot. Ski. 1990, 71, 907–913. [Google Scholar] [CrossRef]
  30. Ito, K.; Ogawa, Y.; Miyata, Y.; Shibusawa, H. Evaluating the subjective perceptions of streetscapes using street-view imagery: A deep learning framework. Landsc. Urban Plan. 2024, 247, 105073. [Google Scholar] [CrossRef]
  31. Evans, D.; Chamberlain, B. In Pursuit of Eye Tracking for Visual Landscape Assessments. Land 2024, 13, 1184. [Google Scholar] [CrossRef]
  32. Vartanian, O.; Navarrete, G.; Chatterjee, A.; Brorson Fich, L.; Leder, H.; Modroño, C.; Nadal, M.; Rostrup, N.; Skov, M. Impact of Contour on Aesthetic Judgments and Approach-Avoidance Decisions in Architecture. Proc. Natl. Acad. Sci. USA 2013, 110, 10446–10453. [Google Scholar] [CrossRef]
  33. Eerola, T.; Vuoskoski, J.K. A Comparison of the Discrete and Dimensional Models of Emotion in Music. Psychol. Music 2011, 39, 18–49. [Google Scholar] [CrossRef]
  34. Zentner, M.; Grandjean, D.; Scherer, K.R. Emotions Evoked by the Sound of Music: Characterization, Classification, and Measurement. Emotion 2008, 8, 494–521. [Google Scholar] [CrossRef]
  35. Weibel, R.; Grübel, J.; Zhao, H.; Thrash, T.; Meloni, D.; Hölscher, C.; Schinazi, V. Virtual reality experiments with physiological measures. J. Vis. Exp. 2018, 138, 58318. [Google Scholar] [CrossRef]
  36. Karperien, A. FracLac for ImageJ. 1999–2013. Available online: https://imagej.net/ij/plugins/fraclac/FLHelp/Introduction.htm (accessed on 13 May 2026).
  37. Schneider, C.A.; Rasband, W.S.; Eliceiri, K.W. NIH Image to ImageJ: 25 years of image analysis. Nat. Methods 2012, 9, 671–675. [Google Scholar] [CrossRef]
  38. Hagerhall, C.M.; Purcell, T.; Taylor, R. Fractal dimension of landscape silhouette outlines as a predictor of landscape preference. J. Environ. Psychol. 2004, 24, 247–255. [Google Scholar] [CrossRef]
  39. Abboushi, B.; Elzeyadi, I.; Taylor, R.; Sereno, M. Fractals in architecture: The visual interest, preference, and mood response to projected fractal light patterns in interior spaces. J. Environ. Psychol. 2019, 61, 57–70. [Google Scholar] [CrossRef]
  40. Papoutsaki, A.; Sangkloy, P.; Laskey, J.; Daskalova, N.; Huang, J.; Hays, J. WebGazer: Scalable Webcam Gaze Tracking Using User Interactions. In Proceedings of the 25th International Joint Conference on Artificial Intelligence (IJCAI); IJCAI/AAAI: Menlo Park, CA, USA, 2016; pp. 3839–3845. [Google Scholar]
  41. Mavros, P.; Ngoi, Z.L.; Kirk, S.; Te, A.S.; Grubel, J.; Aguilar, L.; Olszewska-Guizzo, A.; Makowski, D. Attenuating Subjective Crowding Through Beauty: An Online Study on the Interaction Between Environment Aesthetics, Typology and Crowdedness; PsyArXiv: Charlottesville, VA, USA, 2023; pp. 1–32. [Google Scholar]
  42. Taylor, R.P.; Spehar, B.; Wise, J.A.; Clifford, C.W.G.; Newell, B.R.; Hagerhall, C.M.; Purcell, T.; Martin, T.P. Perceptual and Physiological Responses to the Visual Complexity of Fractal Patterns. J. Non-Linear Dyn. Psychol. Life Sci. 2005, 9, 89–114. [Google Scholar]
  43. Ulrich, R.S. View Through a Window may Influence Recovery from Surgery. Science 1984, 224, 420–421. [Google Scholar] [CrossRef] [PubMed]
  44. Kaplan, S. The restorative benefits of nature: Toward an integrative framework. J. Environ. Psychol. 1995, 15, 169–182. [Google Scholar] [CrossRef]
Figure 1. Case study area location. A total of 73 building facades and 42 pedestrian streetscapes were distributed across three wards: Wakamatsu, Tobata, and Mojiko.
Figure 1. Case study area location. A total of 73 building facades and 42 pedestrian streetscapes were distributed across three wards: Wakamatsu, Tobata, and Mojiko.
Architecture 06 00091 g001
Figure 2. Selected study area stimuli. Wakamatsu (Figs. 1–23 buildings; Figs. 1–12 pedestrian), Tobata (Figs. 24–45; Figs. 13–18), and Mojiko (Figs. 46–73; Figs. 19–42).
Figure 2. Selected study area stimuli. Wakamatsu (Figs. 1–23 buildings; Figs. 1–12 pedestrian), Tobata (Figs. 24–45; Figs. 13–18), and Mojiko (Figs. 46–73; Figs. 19–42).
Architecture 06 00091 g002
Figure 3. Laboratory layout and experiment room. (Left): floor plan showing the spatial arrangement of equipment and zones, the observing desk (researcher), experiment area (participant), and questionnaire station. (Right): photograph of the actual experiment room.
Figure 3. Laboratory layout and experiment room. (Left): floor plan showing the spatial arrangement of equipment and zones, the observing desk (researcher), experiment area (participant), and questionnaire station. (Right): photograph of the actual experiment room.
Architecture 06 00091 g003
Figure 4. Equipment used in the experiment: eye-tracking system, laptop, monitor, tripod camera, and speaker.
Figure 4. Equipment used in the experiment: eye-tracking system, laptop, monitor, tripod camera, and speaker.
Architecture 06 00091 g004
Figure 5. Experimental procedure: (A) preparation and emotion induction via video; (B) eye-tracker calibration; (C) first observation session with 73 architectural facade figures at 10 s each; (D) second observation session with 42 pedestrian streetscape figures at 10 s each; (E) post-experiment questionnaire. Total duration: approximately 60 min.
Figure 5. Experimental procedure: (A) preparation and emotion induction via video; (B) eye-tracker calibration; (C) first observation session with 73 architectural facade figures at 10 s each; (D) second observation session with 42 pedestrian streetscape figures at 10 s each; (E) post-experiment questionnaire. Total duration: approximately 60 min.
Architecture 06 00091 g005
Figure 6. Emotion induction setting: participants observed emotion-inducing video clips (Fear, Anger, Sad, and Happy) on the large monitor prior to viewing the architectural and streetscape stimuli.
Figure 6. Emotion induction setting: participants observed emotion-inducing video clips (Fear, Anger, Sad, and Happy) on the large monitor prior to viewing the architectural and streetscape stimuli.
Architecture 06 00091 g006
Figure 7. Visual reference for Fractal Dimension (D) and Lacunarity (Λ) using building façade stimuli from this study.
Figure 7. Visual reference for Fractal Dimension (D) and Lacunarity (Λ) using building façade stimuli from this study.
Architecture 06 00091 g007
Figure 8. Inter-rater reliability—Pearson r between participants within each emotion group. Bar = mean r across all pairwise combinations; error bar = range. Green dashed line = r = 0.70 (acceptable threshold); orange dotted = r = 0.40 (moderate). Buildings (left) and Pedestrian (right).
Figure 8. Inter-rater reliability—Pearson r between participants within each emotion group. Bar = mean r across all pairwise combinations; error bar = range. Green dashed line = r = 0.70 (acceptable threshold); orange dotted = r = 0.40 (moderate). Buildings (left) and Pedestrian (right).
Architecture 06 00091 g008
Figure 9. Within-stimulus consistency—Coefficient of Variation (CV = SD/Mean) per figure across participants. Lower CV indicates higher consistency. Box = IQR; line = median; dots = individual stimuli. Green dashed = 0.20 threshold; orange dotted = 0.50.
Figure 9. Within-stimulus consistency—Coefficient of Variation (CV = SD/Mean) per figure across participants. Lower CV indicates higher consistency. Box = IQR; line = median; dots = individual stimuli. Green dashed = 0.20 threshold; orange dotted = 0.50.
Architecture 06 00091 g009
Figure 10. Eye-tracking system accuracy comparison (left) and AOI coverage schematic (right). The webcam-based MediaPipe system (estimated ±2–3°) is consistent with comparable webcam systems and adequate for the full-image AOI design used in this study.
Figure 10. Eye-tracking system accuracy comparison (left) and AOI coverage schematic (right). The webcam-based MediaPipe system (estimated ±2–3°) is consistent with comparable webcam systems and adequate for the full-image AOI design used in this study.
Architecture 06 00091 g010
Figure 11. Comfort vote stability across emotion groups—Building Façades (73 figures). Each dot = one emotion group’s vote count; vertical bar = inter-group spread. Figures sorted high to low by grand mean comfort vote. Purple labels indicate emotion-dependent divergence (spread > 1).
Figure 11. Comfort vote stability across emotion groups—Building Façades (73 figures). Each dot = one emotion group’s vote count; vertical bar = inter-group spread. Figures sorted high to low by grand mean comfort vote. Purple labels indicate emotion-dependent divergence (spread > 1).
Architecture 06 00091 g011
Figure 12. Comfort vote stability across emotion groups—Pedestrian Streetscapes (42 figures). Each dot = one emotion group’s vote count; vertical bar = inter-group spread. Figures sorted high to low by grand mean comfort vote. Purple labels indicate emotion-dependent divergence (spread > 1).
Figure 12. Comfort vote stability across emotion groups—Pedestrian Streetscapes (42 figures). Each dot = one emotion group’s vote count; vertical bar = inter-group spread. Figures sorted high to low by grand mean comfort vote. Purple labels indicate emotion-dependent divergence (spread > 1).
Architecture 06 00091 g012
Figure 13. Fractal Dimension (D) vs. Lacunarity (Λ)—Most Voted (n = 16) vs. Not Voted (n = 29) Building Façades. Density Distributions Shown on Marginal Axes.
Figure 13. Fractal Dimension (D) vs. Lacunarity (Λ)—Most Voted (n = 16) vs. Not Voted (n = 29) Building Façades. Density Distributions Shown on Marginal Axes.
Architecture 06 00091 g013
Figure 14. Fractal analysis by study area: Most Voted vs. Not Voted building façades. Wakamatsu (n = 6), Tobata (n = 17), Mojiko (n = 22). Dashed lines = area mean D and Λ.
Figure 14. Fractal analysis by study area: Most Voted vs. Not Voted building façades. Wakamatsu (n = 6), Tobata (n = 17), Mojiko (n = 22). Dashed lines = area mean D and Λ.
Architecture 06 00091 g014
Figure 16. Visual metric heatmap for 42 pedestrian streetscape figures by study area. Cell color: viridis scale (dark = low, bright = high); values are raw percentages. Architecture 06 00091 i001 Comfortable · Architecture 06 00091 i002 Neutral · Architecture 06 00091 i003 Uncomfortable. Combined Score = mean of normalized metrics (0–1).
Figure 16. Visual metric heatmap for 42 pedestrian streetscape figures by study area. Cell color: viridis scale (dark = low, bright = high); values are raw percentages. Architecture 06 00091 i001 Comfortable · Architecture 06 00091 i002 Neutral · Architecture 06 00091 i003 Uncomfortable. Combined Score = mean of normalized metrics (0–1).
Architecture 06 00091 g016
Figure 17. Mean Street Metric Values by Emotion Group and Area. Each panel shows one metric (rows) across emotion groups (Comfortable, Neutral, and Uncomfortable) for each area (columns). Black border indicates the highest value per panel.
Figure 17. Mean Street Metric Values by Emotion Group and Area. Each panel shows one metric (rows) across emotion groups (Comfortable, Neutral, and Uncomfortable) for each area (columns). Black border indicates the highest value per panel.
Architecture 06 00091 g017
Figure 18. Pearson Correlation Analysis of Pedestrian Street Metrics and Emotion Perception. Green = positive; red = negative; white = near zero. Significant findings (|r| ≥ 0.40, p < 0.05): Canopy × Green (r = 0.98 ***); Enclosure × Green (r = −0.64 ***); Enclosure × Canopy (r = −0.62 ***); Sky × Canopy (r = −0.47 **); Sky × Green (r = −0.45 **). * p < 0.05, ** p < 0.01, *** p < 0.001.
Figure 18. Pearson Correlation Analysis of Pedestrian Street Metrics and Emotion Perception. Green = positive; red = negative; white = near zero. Significant findings (|r| ≥ 0.40, p < 0.05): Canopy × Green (r = 0.98 ***); Enclosure × Green (r = −0.64 ***); Enclosure × Canopy (r = −0.62 ***); Sky × Canopy (r = −0.47 **); Sky × Green (r = −0.45 **). * p < 0.05, ** p < 0.01, *** p < 0.001.
Architecture 06 00091 g018
Figure 19. Cross-method synthesis—Building Façades (73 figures). Steps 1–5: D vs. comfort votes; Λ vs. comfort votes; VAS vs. comfort votes; preference by D and Λ group; emotion stability.
Figure 19. Cross-method synthesis—Building Façades (73 figures). Steps 1–5: D vs. comfort votes; Λ vs. comfort votes; VAS vs. comfort votes; preference by D and Λ group; emotion stability.
Architecture 06 00091 g019
Figure 20. Top comfortable figures: metric profile, comfort votes, and VAS per emotion group. Buildings (top 12); Pedestrian (top 10). Dot color = group consensus (4/4, 3/4, 2/4).
Figure 20. Top comfortable figures: metric profile, comfort votes, and VAS per emotion group. Buildings (top 12); Pedestrian (top 10). Dot color = group consensus (4/4, 3/4, 2/4).
Architecture 06 00091 g020
Figure 21. Most comfortable building façades—Figs. 60, 46, and 56 (all Mojiko, 4/4 groups).
Figure 21. Most comfortable building façades—Figs. 60, 46, and 56 (all Mojiko, 4/4 groups).
Architecture 06 00091 g021
Figure 22. Fig. 64 (Minato House, Mojiko)—structural paradox. D = 1.517, Λ = 1.200; voted comfortable by all 4 groups despite low D and high Λ.
Figure 22. Fig. 64 (Minato House, Mojiko)—structural paradox. D = 1.517, Λ = 1.200; voted comfortable by all 4 groups despite low D and high Λ.
Architecture 06 00091 g022
Figure 23. Least comfortable building façades—Fig. 4 and Fig. 10 (Wakamatsu). Zero comfort votes across all emotion groups.
Figure 23. Least comfortable building façades—Fig. 4 and Fig. 10 (Wakamatsu). Zero comfort votes across all emotion groups.
Architecture 06 00091 g023
Figure 24. Fig. 42 (Mojiko waterfront)—highest comfort vote consensus in the pedestrian dataset (4/4 groups, 11 total votes).
Figure 24. Fig. 42 (Mojiko waterfront)—highest comfort vote consensus in the pedestrian dataset (4/4 groups, 11 total votes).
Architecture 06 00091 g024
Figure 25. Most comfortable pedestrian streetscapes voted by 3/4 groups—Fig. 38, 29, and 41 (all Mojiko). Characterized by tree canopy and low enclosure.
Figure 25. Most comfortable pedestrian streetscapes voted by 3/4 groups—Fig. 38, 29, and 41 (all Mojiko). Characterized by tree canopy and low enclosure.
Architecture 06 00091 g025
Figure 26. Fig. 5 (Wakamatsu)—highest Combined Score in the pedestrian dataset (CS = 0.50). Sky Ratio 80.5%, Canopy 37.1%, Enclosure 2.3%.
Figure 26. Fig. 5 (Wakamatsu)—highest Combined Score in the pedestrian dataset (CS = 0.50). Sky Ratio 80.5%, Canopy 37.1%, Enclosure 2.3%.
Architecture 06 00091 g026
Figure 27. Lowest comfort consensus pedestrian streetscapes—Fig. 8 (Wakamatsu) and Fig. 18 (Tobata). Characterized by road dominance and absent or inaccessible vegetation.
Figure 27. Lowest comfort consensus pedestrian streetscapes—Fig. 8 (Wakamatsu) and Fig. 18 (Tobata). Characterized by road dominance and absent or inaccessible vegetation.
Architecture 06 00091 g027
Figure 28. Fig. 3 (Wakamatsu)—metric-vote discrepancy case. Classified as Comfortable by metrics (CS = 0.34) but voted by only 2/4 groups. Vegetation is visible but outside the pedestrian zone.
Figure 28. Fig. 3 (Wakamatsu)—metric-vote discrepancy case. Classified as Comfortable by metrics (CS = 0.34) but voted by only 2/4 groups. Vegetation is visible but outside the pedestrian zone.
Architecture 06 00091 g028
Table 1. Mean Visual Attention Score (VAS) by emotion group for building facades (n = 73) and pedestrian streetscapes (n = 42). VAS (normalized dwell time + normalized fixation count + normalized gaze entropy) divided by 3. Score range from 0 (lo attention) to 1 (high attention).
Table 1. Mean Visual Attention Score (VAS) by emotion group for building facades (n = 73) and pedestrian streetscapes (n = 42). VAS (normalized dwell time + normalized fixation count + normalized gaze entropy) divided by 3. Score range from 0 (lo attention) to 1 (high attention).
(a) Building Facades of 73 Figures
Emotion GroupnMean VASStandard DeviationMinMax
Fear730.6360.0480.5350.728
Anger730.6550.0410.5320.734
Sad730.6860.0380.6030.785
Happy730.6670.0390.5420.753
Kruskal–Wallis: H = 43.063, p = 0.0; Sad group scored highest (mean = 0.686); Fear group lowest (mean = 0.636). Absolute difference of 0.05 suggests Sad condition induced more focused scanning.
(b) Pedestrian Streetscape of 42 Figures
Emotion GroupnMean VASStandard DeviationMinMax
Fear420.6310.0360.5330.686
Anger420.6650.0380.5750.784
Sad420.6820.040.6010.757
Happy420.680.030.610.74
Kruskal–Wallis: H = 38.799, p = 0.0; Sad group scored highest (mean = 0.682); Fear group lowest (mean = 0.631). Pattern consistent with buildings across all four emotion conditions.
Table 2. Top 3 most and least visually attracted figures by emotion group, stimulus type, and area. VAS shown in parentheses.
Table 2. Top 3 most and least visually attracted figures by emotion group, stimulus type, and area. VAS shown in parentheses.
EmotionStimulusArea Top 3 Most Attracted (VAS)Top 3 Least Attracted (VAS)
Fear Buildings WakamatsuFig. 12 (0.721), Fig. 11 (0.717), Fig. 7 (0.717)Fig. 1 (0.559), Fig. 8 (0.563), Fig. 9 (0.571)
Tobata Fig. 42 (0.709), Fig. 43 (0.692), Fig. 34 (0.660)Fig. 41 (0.535), Fig. 28 (0.557), Fig. 40 (0.561)
Mojiko Fig. 58 (0.728), Fig. 68 (0.721), Fig. 69 (0.698)Fig. 65 (0.571), Fig. 67 (0.579), Fig. 49 (0.582)
Pedestrian WakamatsuFig. 8 (0.675), Fig. 11 (0.649), Fig. 6 (0.648)Fig. 3 (0.569), Fig. 5 (0.599), Fig. 10 (0.609)
MojikoFig. 30 (0.686), Fig. 28 (0.683), Fig. 38 (0.674)Fig. 37 (0.533), Fig. 32 (0.559), Fig. 39 (0.565)
AngerBuildingsWakamatsuFig. 16 (0.707), Fig. 4 (0.692), Fig. 15 (0.691)Fig. 9 (0.532), Fig. 12 (0.575), Fig. 10 (0.583)
TobataFig. 33 (0.714), Fig. 32 (0.705), Fig. 37 (0.687)Fig. 45 (0.570), Fig. 36 (0.575), Fig. 29 (0.592)
MojikoFig. 52 (0.735), Fig. 57 (0.734), Fig. 62 (0.730)Fig. 67 (0.594), Fig. 69 (0.614), Fig. 64 (0.624)
PedestrianWakamatsuFig. 10 (0.734), Fig. 6 (0.715), Fig. 4 (0.713)Fig. 3 (0.626), Fig. 11 (0.635), Fig. 9 (0.663)
MojikoFig. 20 (0.713), Fig. 31 (0.703), Fig. 19 (0.696)Fig. 21 (0.575), Fig. 42 (0.586), Fig. 26 (0.626)
SadBuildingsWakamatsuFig. 19 (0.744), Fig. 11 (0.733), Fig. 18 (0.718)Fig. 1 (0.621), Fig. 20 (0.639), Fig. 16 (0.648)
TobataFig. 38 (0.746), Fig. 28 (0.741), Fig. 37 (0.734)Fig. 36 (0.603), Fig. 32 (0.609), Fig. 30 (0.620)
MojikoFig. 66 (0.785), Fig. 64 (0.758), Fig. 48 (0.757)Fig. 53 (0.618), Fig. 63 (0.626), Fig. 67 (0.637)
PedestrianWakamatsuFig. 9 (0.737), Fig. 8 (0.726), Fig. 10 (0.710)Fig. 11 (0.622), Fig. 1 (0.627), Fig. 12 (0.637)
MojikoFig. 42 (0.753), Fig. 21 (0.739), Fig. 23 (0.732)Fig. 28 (0.629), Fig. 40 (0.630), Fig. 33 (0.631)
HappyBuildingsWakamatsuFig. 7 (0.753), Fig. 23 (0.712), Fig. 3 (0.705)Fig. 20 (0.548), Fig. 12 (0.581), Fig. 8 (0.614)
TobataFig. 36 (0.735), Fig. 34 (0.726), Fig. 37 (0.722)Fig. 31 (0.619), Fig. 30 (0.621), Fig. 45 (0.622)
MojikoFig. 58 (0.724), Fig. 50 (0.707), Fig. 46 (0.703)Fig. 52 (0.542), Fig. 53 (0.594), Fig. 61 (0.600)
PedestrianWakamatsuFig. 4 (0.722), Fig. 3 (0.709), Fig. 11 (0.696)Fig. 6 (0.630), Fig. 5 (0.641), Fig. 1 (0.652)
MojikoFig. 40 (0.740), Fig. 33 (0.737), Fig. 21 (0.716)Fig. 35 (0.610), Fig. 32 (0.612), Fig. 26 (0.630)
Note: Tobata pedestrian (Figs. 13–18, n = 6) omitted from most/least ranking due to small subgroup size.
Table 3. Fractal dimension (D) and lacunarity (Λ) of all comfort-voted facade figures per emotion group. All R2 values are Excellent (>0.98). Fig. 64 appears in three of four groups as a consistent outlier.
Table 3. Fractal dimension (D) and lacunarity (Λ) of all comfort-voted facade figures per emotion group. All R2 values are Excellent (>0.98). Fig. 64 appears in three of four groups as a consistent outlier.
Group Figure (D) Fractal Dimension(Λ) LacunarityComplexity ClassInterpretation
FearFig. 361.7700.456Very HighComplex but Organized
Fig. 501.7760.357Very HighComplex but Organized
Fig. 601.7600.379Very HighComplex but Organized
Fig. 641.5171.200HighMod. complex but irregular
AngerFig. 461.8260.253Very HighComplex but Organized
Fig. 471.7260.279Very HighComplex but Organized
Fig. 541.7590.396Very HighComplex but Organized
Fig. 551.7460.473Very HighComplex but Organized
Fig. 561.7180.388Very HighComplex but Organized
Fig. 601.7600.379Very HighComplex but Organized
Fig. 611.8280.320Very HighComplex but Organized
Fig. 641.5171.200HighMod. complex but irregular
SadFig. 461.8260.253Very HighComplex but Organized
Fig. 561.7180.388Very HighComplex but Organized
Fig. 581.6840.514HighMod. complex and balanced
Fig. 601.7600.379Very HighComplex but Organized
Fig. 611.8280.320Very HighComplex but Organized
Fig. 641.5171.200HighMod. complex but irregular
HappyFig. 461.8260.253Very HighComplex but Organized
Fig. 501.7760.357Very HighComplex but Organized
Fig. 531.7460.364Very HighComplex but Organized
Fig. 561.7180.388Very HighComplex but Organized
Fig. 601.7600.379Very HighComplex but Organized
Fig. 641.5171.200HighMod. complex but irregular
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Perwira, S.A.; Dewancker, B.J.; Herjuno, D. Urban Comfort Perception Under Induced Emotional Conditions: A Multi-Method Analysis of Architectural and Streetscape Imagery Using Fractal Analysis, Self-Report, and Eye-Tracking. Architecture 2026, 6, 91. https://doi.org/10.3390/architecture6020091

AMA Style

Perwira SA, Dewancker BJ, Herjuno D. Urban Comfort Perception Under Induced Emotional Conditions: A Multi-Method Analysis of Architectural and Streetscape Imagery Using Fractal Analysis, Self-Report, and Eye-Tracking. Architecture. 2026; 6(2):91. https://doi.org/10.3390/architecture6020091

Chicago/Turabian Style

Perwira, Satrio Agung, Bart Julien Dewancker, and Dimas Herjuno. 2026. "Urban Comfort Perception Under Induced Emotional Conditions: A Multi-Method Analysis of Architectural and Streetscape Imagery Using Fractal Analysis, Self-Report, and Eye-Tracking" Architecture 6, no. 2: 91. https://doi.org/10.3390/architecture6020091

APA Style

Perwira, S. A., Dewancker, B. J., & Herjuno, D. (2026). Urban Comfort Perception Under Induced Emotional Conditions: A Multi-Method Analysis of Architectural and Streetscape Imagery Using Fractal Analysis, Self-Report, and Eye-Tracking. Architecture, 6(2), 91. https://doi.org/10.3390/architecture6020091

Article Metrics

Back to TopTop