Next Article in Journal
Machine Learning-Based Assessment of Land-Use Change, Forest Recovery, and Landscape Connectivity in Islamabad
Previous Article in Journal
Tangible and Intangible Landscape Elements in the Adaptive Reuse of Historic Railway Corridors: Landscape Perception and Preference Along the Taichung Green Corridor 1908, Taiwan
Previous Article in Special Issue
Freight Big Data-Based Dual-Scale Study of Economic Spatial Organization and Planning Responses in Hubei Province
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Multidimensional Framework for Diagnosing Streetscape Perception in Historic-District Renewal Using Street-View Imagery and Deep Learning

School of Architecture and Design, China University of Mining and Technology, Xuzhou 221116, China
*
Author to whom correspondence should be addressed.
Land 2026, 15(9), 1640; https://doi.org/10.3390/land15091640
Submission received: 20 July 2026 / Revised: 22 August 2026 / Accepted: 1 September 2026 / Published: 3 September 2026
(This article belongs to the Special Issue Big Data-Driven Urban Spatial Perception)

Abstract

The renewal of historic districts needs to address human-centered issues, such as cultural expression, spatial experience, and visual comfort, at the street scale. However, existing assessment methods predominantly rely on field surveys, expert judgment, or individual visual indicators, making it difficult to produce reproducible and spatially explicit diagnostic evidence across extensive street networks. This study proposes a multidimensional streetscape perception diagnostic framework that evaluates three dimensions: cultural character recognition (CCR), spatial order (SO), and visual comfort (VC). Taking the Pengcheng Qili historic district in Xuzhou, China, as a case study, the framework integrates street-view imagery, subjective pairwise comparisons, the Bradley–Terry model, and ResNet50-based deep learning prediction. Based on 2243 sampling points and 8693 street-view images, three perception–prediction models were developed and evaluated using five-fold stratified cross-validation, followed by independent external validation using additional historic-district images. The outputs were subsequently mapped onto the street network and examined using spatial statistical analysis and multidimensional profile classification. The results show that the three perception dimensions exhibit distinct spatial patterns and significant spatial clustering. CCR forms localized clusters around historical nodes and heritage-rich areas, whereas SO and VC show clearer corridor-like and network-like patterns. The proposed framework organizes streetscape perception predictions into interpretable multidimensional spatial profiles, thereby providing spatial evidence for conservation-oriented renewal and fine-scale governance of historic districts.

1. Introduction

Historic districts (HDs) are not merely physical repositories of urban heritage; they are also places where cultural memory, spatial experience, and public life are continuously perceived, practiced, and reproduced [1,2,3,4]. Accordingly, the emphasis in HD renewal has shifted from the restoration of individual buildings and the upgrading of physical infrastructure to the overall streetscape and its associated cultural and human-scale experiences [5,6]. From a human-scale perspective, street-level perception is particularly important in HDs since it mediates how cultural heritage is recognized and how visual qualities are experienced in everyday life [7,8]. Previous studies have mainly focused on several dimensions relevant to HD renewal, such as inherent cultural character, perceived spatial organization, and visual comfort [9,10]. However, these dimensions tend to be examined separately, with limited effort to integrate them within a diagnostic framework.
This study proposes a multidimensional framework for diagnosing streetscape perception in HD renewal, incorporating three related but non-interchangeable dimensions: cultural character recognition (CCR), spatial order (SO), and visual comfort (VC). The rationale for integrating these dimensions is due to the hypothesis that satisfactory performance in one dimension does not necessarily compensate for deficiencies in another. For example, although physical upgrading may improve the spatial order and visual comfort of a historic streetscape, the extensive use of modern materials and façades may weaken the recognizability of its cultural character and associated historical atmosphere. An integrated assessment is therefore required to identify dimension-specific and overlapping deficiencies that may be overlooked when the three dimensions are examined separately in existing studies. Moreover, such an assessment can provide a replicable and fine-scale basis for identifying perceptual deficiencies and supporting evidence-based renewal decisions in HDs [11,12].
Traditional assessments of HDs mainly rely on field surveys, expert scoring, and questionnaires [3,6,13]. These methods are indispensable for subjectively interpreting the cultural value of an HD, yet they are often labor-intensive and costly, resulting in limited spatial coverage [14,15,16,17]. Recent advances in street-view imagery (SVI), computer vision (CV), and urban artificial intelligence (AI) have provided new opportunities to assess the built environment [18,19,20]. In particular, SVI has become an important data source for human-scale urban studies as it records the urban environment at the street level [12,19,21,22]. Biljecki and Ito [14] reviewed its applications in urban analytics and GISscience, including vegetation assessment, transportation analysis, health studies, and environmental-quality assessment. Compared with remote-sensing imagery, SVI captures visible elements that can be directly perceived by pedestrians, including building façades, sidewalks, vegetation, sky visibility and signage [3,8,23,24]. CV methods further extend the analytical capacity of SVI. Representative semantic segmentation models, such as DeepLab, fully convolutional networks (FCNs), and U-Net, have been extensively used to extract pixel-level visual elements in the literature [13,20,25,26]. Image classification tasks have been performed to infer subjective perceptions of streetscapes, including beauty, safety, vitality, wealth, depression, and boredom [3,9,27,28,29]. These methods enable large-scale street-view images to be transformed into reproducible indicators and spatial maps, making them particularly suitable for streetscape assessments that require broader spatial coverage [15,16,19,30].
Existing studies of streetscape perception can be broadly summarized into three groups relevant to HD renewal [7,8,31,32,33,34]. The first focuses on cultural character recognition (CCR). These studies examine how visible heritage attributes—including historic façades, traditional materials and colors, architectural details, local symbols, and streetscape form—influence the recognition of heritage value, place identity, and historic character [3,4,7,8,13,17].
The second focuses on spatial order (SO) perception and its influence on pedestrian experience. Spatial attributes such as openness, enclosure, street curvature, building density, height-to-width ratio and interface continuity have been examined in relation to spatial legibility, coherence, and pedestrian mobility [6,15,23,35]. Studies of HDs have shown that the balance between openness and enclosure affects how street space is perceived, while excessive enclosure may produce a sense of oppression and psychological stress [1,6,9]. In historic alleyway settings, spatial openness measured by the depth-to-height ratio, building density, and street curvature has been found to affect perceived spatial harmony and legibility, suggesting that SO depends on the combined effects of openness, enclosure, and street-scale continuity [6].
The third group examines people’s evaluative responses to overall visual comfort (VC). It has been suggested that sky visibility and the height-to-width ratio are associated with human comfort [10,27,36,37,38]. VC is further shaped by greenery, building interfaces, road surfaces, vehicles, pedestrians, water bodies, and visually disordered elements [39,40,41,42,43,44,45]. Recent street-view-based studies have also shown that green-view-related indicators, together with walkability, enclosure, imageability, and visual complexity, are useful for explaining human-scale visual environmental quality [30,46]. However, it should be noted that the effects of these elements on cultural and esthetic perception, emotional responses, and environmental stress may vary across road types and spatial contexts [9,16,23,45]. Therefore, streetscape assessment in HDs should not be reduced to a single score but should integrate CCR, SO, and VC within a unified framework.
In summary, three gaps remain in the literature. First, although SVI has been widely applied in urban studies, its use in integrated streetscape assessments specifically designed for HDs remains limited. Second, existing studies often examine individual dimensions such as esthetics, comfort, safety, vitality, or historic character, but pay insufficient attention to the relationships and overlapping deficiencies among CCR, SO, and VC [46,47,48,49,50]. Third, how streetscape perception results can inform spatial diagnosis and targeted renewal decision-making has received limited attention in the literature [6,16].
To address these gaps, this study proposes a multidimensional framework for diagnosing streetscape perception using SVI, subjective perception evaluation, and deep-learning techniques. Taking the Pengcheng Qili historic district in Xuzhou as a case study, the framework assesses CCR, SO, and VC. First, subjective perception judgments were collected through pairwise comparisons of sampled street-view images, and the Bradley–Terry model [25,42] was used to construct perception scores and grade labels. Second, three ResNet50 classification models [9,39] were trained to extend the subjective evaluations from the annotated samples to street-view images throughout the study area. Finally, point-level aggregation, spatial mapping, spatial statistical analysis, and multidimensional profile classification were applied to identify the spatial distribution of single- and multidimensional perceptual deficiencies and provide evidence for differentiated renewal strategies. By linking subjective evaluation with deep learning prediction and spatial diagnosis, the study provides a replicable approach for supporting perception-oriented renewal decisions in HDs. It should be noted that this study focuses on internal perceptual differentiation within a historic-district context, rather than distinguishing historic districts from ordinary urban areas.

2. Materials and Methods

2.1. Study Area

This study selected the Pengcheng Qili historic district in Xuzhou, China, as the case study area (Figure 1), as the area contains numerous cultural heritage resources (Figure 2) and is currently undergoing active urban renewal. Pengcheng Qili refers to an approximately 3.5 km long historic and cultural urban axis. This axis connects important historical and cultural nodes such as the Underground City Site Museum (E8 in Figure 2), Huilongwo (B9 in Figure 2), Kuaizaiting Park (B10 in Figure 2), Hubushan (A1 in Figure 2), and Huanglou Park (B1 in Figure 2). In addition, the Huanghe Gudao forms a riverine corridor that partially encircles the study area and constitutes an important landscape and spatial boundary of Pengcheng Qili. The area contains 97 heritage assets and 235 historical remains, making it central to Xuzhou’s historical memory and urban identity. As a key urban renewal project, the city aims to integrate historic conservation, cultural tourism, commercial activity, and everyday residential life. The street environments within the area are diverse, including traditional alleys, historic street frontages, commercial streets, and neighborhood streets in residential areas. These characteristics make Pengcheng Qili a suitable case for evaluating CCR, SO, and VC using SVI and applying the proposed multidimensional framework.

2.2. Multidimensional Perception Framework

This study develops a multidimensional streetscape perception diagnostic workflow for HD renewal (Figure 3). Based on SVI and subjective perception evaluations, the workflow uses deep learning models to estimate CCR, SO, and VC across the study area. It further integrates GIS-based mapping and spatial statistical analysis methods to convert image-level predictions into spatial diagnostic results that can support evidence-based renewal decision-making. The overall workflow includes six main steps.

2.3. Data Sources

During the data acquisition stage, this study used SVI data, subjective perception evaluations, and geographic coordinates to construct the dataset for multidimensional streetscape perception diagnosis in the Pengcheng Qili historic district [4,6,7]. First, street networks were obtained from OpenStreetMap (OSM) and processed using QGIS 3.40.14 [51] to generate sampling points at 50 m intervals along the road network (Figure 1). Based on these sampling points, SVI was collected in four directions from Baidu Street View (https://api.map.baidu.com/panorama/v2, (accessed on 14 May 2026), with the images mainly acquired in 2022. In total, 2243 sampling points and 8693 valid street-view images were obtained, covering the principal streets within the study area [17,46]. The data sources and sample construction statistics are shown in Table 1.

2.4. Perception Label Construction

This study conceptualized street-level perception using three dimensions: cultural character recognition (CCR), spatial order (SO), and visual comfort (VC) (Table 2). Although interrelated, they are not interchangeable aspects of streetscape perception in HDs. The subsequent perception–prediction models were designed to classify SVI according to each of these three dimensions. To select representative images for subjective evaluation, the original images were divided into nine sequential groups based on the sampling-point numbers, with up to 1000 images in each group. A histogram-difference method was then used to select images with relatively distinct visual characteristics. Low-quality, duplicate, and severely occluded images, as well as images containing limited streetscape information, were further removed through manual screening. Finally, 910 representative street-view images were retained for subjective perception evaluation.
To better capture relative perceptual differences among SVI, this study adopted pairwise comparison rather than direct scoring, and data were collected through a local annotation system built using Streamlit 1.40.1 (Figure 4). Three questions were designed to correspond to CCR, SO, and VC, respectively:
Q1 Cultural Character Recognition (CCR): Which image better reflects the historic and cultural character of the district?
Q2 Spatial Order (SO): Which image presents a more orderly and coherent street space?
Q3 Visual Comfort (VC): Which image provides a more comfortable visual experience?
Figure 4. Screenshot of the Streamlit-based pairwise comparison interface used for the subjective evaluation of cultural character recognition, spatial order, and visual comfort.
Figure 4. Screenshot of the Streamlit-based pairwise comparison interface used for the subjective evaluation of cultural character recognition, spatial order, and visual comfort.
Land 15 01640 g004
Participants selected one of the two images according to the corresponding question or chose “difficult to judge” when necessary. Each participant was asked to complete 100 comparisons for each dimension, resulting in a target of 300 comparisons per participant. Each image was included in up to 15 pairwise comparisons for each dimension. In total, 71 volunteers, including 34 women and 37 men, participated in the evaluation. The participants were primarily architecture students, whose disciplinary training provided familiarity with spatial and visual assessment. The sample also included local participants who were familiar with the study area. A total of 20,475 valid pairwise-comparison records were obtained. Responses marked as “difficult to judge” were excluded from the win–loss matrix and were not treated as ties in the BT estimation, because they reflected ambiguous perceptual differences between image pairs. The proportions of “difficult to judge” responses were 53.74% for CCR, 27.87% for SO, and 22.85% for VC (Table A1). The higher uncertainty observed for CCR resulted in fewer images meeting the screening criteria, contributing to its lower retention rate compared with SO and VC. This screening helped retain samples with clearer relative preferences and more stable BT scores for model training.
The pairwise-comparison records obtained from the subjective evaluation were converted into latent perception scores using the Bradley–Terry (BT) model [25,42]. Because the questionnaire recorded pairwise preferences between images rather than absolute ratings of individual images, the BT model provided a statistical basis for estimating the perceptual quality of each SVI. After the estimated scores were screened according to the number of valid win–loss comparisons, the proportion of “difficult to judge” responses, and score stability, 501 images were retained for CCR, 735 for SO, and 768 for VC. The detailed statistics of the BT-based screening process, including the number of retained samples, valid comparisons, and score uncertainty indicators for each dimension, are provided in Appendix A Table A1. Although the number of valid images varied among the three dimensions, the retained sample sizes were broadly consistent with street-view perception studies, supporting their use for subsequent model training [27]. The Jenks natural breaks [52] method was then applied separately to each dimension to divide the retained BT scores into high, medium, and low perception grades. This method minimizes variation within classes and maximizes differences between classes without requiring equally sized groups. To examine the sensitivity of this classification choice, Jenks natural breaks were compared with quantile and equal-interval classification using the retained BT scores for each dimension. Classification performance was evaluated using the within-class sum of squares (WCSS) and goodness of variance fit (GVF), with lower WCSS and higher GVF indicating greater within-class homogeneity and a better fit to the empirical score distribution. The resulting classification labels were used to train three separate ResNet50 models for CCR, SO, and VC. In addition, the spatial coordinates of the sampling points were used to map the model-predicted perception scores to their corresponding geographic locations and generate spatial distribution maps for the three dimensions.

2.5. ResNet50-Based Perception Prediction

During the perception–prediction stage, three separate ResNet50 models with ImageNet-pretrained weights were adopted to identify high, medium, and low perception grades for CCR, SO, and VC. ResNet50 was selected because its residual architecture, combined with ImageNet-based transfer learning, is suitable for extracting robust visual features from complex SVI in the present studies [9,39]. The final fully connected layer of each model was replaced with a classification head containing a dropout layer and a three-class linear output layer to adapt the model to the three-class classification task. A two-stage transfer-learning strategy was used during training. In the first stage, the ResNet50 backbone was frozen, and only the newly added classification layer was trained. In the second stage, the high-level feature block (“layer4”) and the classification layer of ResNet50 were unfrozen and fine-tuned using a smaller learning rate. Only layer4 and the classification head were unfrozen during the second stage to adapt the high-level semantic representations to the streetscape perception task while retaining the general low-level visual features learned from ImageNet and reducing the risk of overfitting under the limited annotated sample size. The learning rates were set to 1 × 10−3 for the first stage and 1 × 10−5 for the fine-tuning stage (Table A2). The loss function was cross-entropy loss with label smoothing, which was used to reduce model overconfidence and sensitivity to uncertainty in the subjective perception labels. Model performance was evaluated using five-fold stratified cross-validation to ensure relatively balanced proportions of low-, medium-, and high-grade samples in each fold. The evaluation metrics included accuracy, Macro-F1, precision, recall, and confusion matrices, with Macro-F1 used as the primary indicator for model evaluation and selection since it assigns equal importance to the three classes. To improve model training with limited local samples, supplementary street-view images from other historic districts were included to enrich the representation of typical historic-district features. These images were used only for feature learning, while the final prediction, spatial analysis, and profile interpretation were conducted for the Pengcheng Qili historic district. To further examine model transferability, an independent external validation was subsequently conducted using 100 different collected images from historic districts in Jiangsu Province that were not used for model training, cross-validation, or model selection. A separate pairwise-comparison experiment was conducted for the three perception dimensions, with each participant completing 50 comparisons per dimension and each image appearing approximately seven times in aggregate across all participants. Independent perceptual scores were estimated using the BT model. Because BT scores represent relative preferences within the external sample, Spearman rank correlation between the independent BT assessments and model-predicted scores was used as the primary external-validation measure. Detailed training parameters, including the optimizer, batch size, learning rates, data-augmentation operations, random seed, learning-rate scheduler, and model-selection criterion, are provided in Table A2.
After model validation, the trained ResNet50 models were applied to all SVI in the Pengcheng Qili historic district. For each image, each model produced three estimated class probabilities, namely Plow, Pmedium, and Phigh. To convert the classification output into a probability-weighted ordinal index, this study calculated the perception score using the following formula:
S = 0 × Plow + 0.5 × Pmedium + 1 × Phigh
which can be simplified as follows:
S = 0.5 × Pmedium + Phigh
The value of S ranges from 0 to 1, with higher values indicating higher predicted perceptual quality and lower values indicating lower predicted perceptual quality. The predicted scores for CCR, SO, and VC were then linked to the corresponding geographic coordinates of the sampling points and were aggregated at the point level, as described in Section 2.6. The resulting point-level scores were used to generate spatial distribution maps for the three dimensions in ArcGIS 10.7.

2.6. Point-Level Aggregation and Spatial Mapping

The spatial aggregation stage was used to convert the model-prediction results into point-level spatial units for HD renewal. The perception–prediction models first output image-level perception scores for CCR, SO, and VC. Each image-level score was then linked to the geographic coordinates of its corresponding sampling point. Because a single sampling point usually contained SVI from four directions, this study averaged the perception scores from the four directions to obtain the final point-level score for the corresponding dimension. This procedure reduces the influence of view-specific occlusion, local lighting conditions, or temporary traffic factors and provides a more balanced representation of the overall streetscape perception at that location.
Si,d = (Si,d,0° + Si,d,90° + Si,d,180° + Si,d,270°)/4
Here, Si,d denotes the point-level perception score at sampling point (i) for dimension d, where d represents CCR, SO, or VC. This aggregation was intended to characterize the overall streetscape perception at each sampling location rather than orientation-specific perceptual differences. Although averaging may attenuate directional variation, retaining a consistent four-view aggregation reduces the influence of view-specific occlusion, lighting, and temporary traffic conditions and provides a comparable point-level measure for subsequent spatial analysis. After obtaining the point-level scores for the three dimensions, the results were visualized in GIS. For descriptive visualization, point-level perception scores were interpolated using the Kriging method in ArcGIS 10.7 with the software’s default parameter settings to generate continuous perception surfaces. The interpolation was used only for visualization, whereas the subsequent spatial autocorrelation and hot-spot analyses were conducted using the original point-level perception scores. This allowed the visualization of areas with relatively high, medium, and low predicted scores for each perception dimension. To provide a supplementary overview, the three dimension-specific perception scores were further aggregated into a comprehensive perception score:
Soverall = (Sculture + Sspatial + Svisual)/3
Here, Sculture, Sspatial, and Svisual denote the scores for CCR, SO, and VC, respectively.
Note that equal weights were assigned because there was no established basis for prioritizing one dimension over another. To examine the sensitivity of the comprehensive score to this assumption, three alternative weighting scenarios were additionally tested by assigning a weight of 0.50 to CCR, SO, or VC in turn, and 0.25 to each of the remaining two dimensions. Ranking stability relative to the equal-weighted baseline was evaluated using Spearman’s rank correlation, mean absolute percentile-rank change, and the overlap of sampling points within the top and bottom 10% of the rankings. These measures were used to assess overall rank consistency, the magnitude of point-level rank shifts, and the stability of the highest- and lowest-ranked locations, respectively. The comprehensive score was used only as a supplementary overview, while the main diagnostic interpretation relied on the separate CCR, SO, and VC scores and their multidimensional profiles.

2.7. Spatial Autocorrelation and Hot/Cold-Spot Analysis

To examine whether the multidimensional perception predictions exhibited spatial clustering, this study used global Moran’s I to assess the global spatial autocorrelation [51] of the CCR, SO, VC, and comprehensive perception scores. Global Moran’s I indicates whether a variable exhibits a clustered, dispersed, or random spatial distribution across the study area. A significantly positive Moran’s I indicates that similar values tend to be spatially adjacent. In other words, high-value points tend to occur near other high-value points, while low-value points tend to occur near other low-value points. A significantly negative Moran’s I indicates spatial dispersion, whereas a value close to its expected value indicates an approximately random spatial pattern.
Using the sampling points as the spatial units of analysis, this study used the point-level CCR, SO, VC, and comprehensive scores as input variables and constructed an inverse-distance spatial-weight matrix based on the Euclidean distance. This method statistically assesses whether the multidimensional perception scores predicted by the models exhibit global spatial dependence. Significant positive spatial autocorrelation would indicate that high and low perception scores were spatially organized rather than randomly distributed, thereby providing statistical support for interpreting the spatial distribution maps and subsequent diagnostic results.
Based on the global spatial autocorrelation analysis, this study further used the Getis–Ord General G statistic to determine whether the overall spatial pattern was characterized primarily by high- or low-value clustering [51]. If the observed General G value is higher than its expected value and the z-score is significantly positive, high values are significantly clustered. Conversely, if the observed value is lower than its expected value and the z-score is significantly negative, low values are significantly clustered. General G identifies the overall tendency towards high- or low-value clustering but does not locate the individual clusters.
To locate the specific spatial positions of high- and low-value clusters, this study used local Getis–Ord Gi* analysis to identify hot and cold spots at different confidence levels [51]. Hot spots indicate significant local clustering of high perception scores, whereas cold spots indicate significant local clustering of low perception scores. These results provide spatial statistical evidence for interpreting the local concentration of perception scores and supporting the subsequent diagnostic analysis. Global Moran’s I, General G, and Getis–Ord Gi* statistics were calculated in ArcGIS using the default spatial relationship and distance settings of the corresponding tools; no distance threshold was manually specified or optimized. To evaluate the sensitivity of the spatial-autocorrelation results to alternative definitions of spatial relationships, Global Moran’s I was additionally recalculated using a 100 m fixed-distance band and eight-nearest-neighbor (K = 8) weights. In addition, Pearson correlation analysis was conducted using the point-level prediction scores of CCR, SO, and VC to further examine the relationships among the three perception dimensions, with Spearman correlation used as a robustness check. To further examine whether the pairwise associations among the three perception dimensions persisted independently of the remaining dimension, partial Spearman correlations were calculated for each pair while controlling for the third dimension.

2.8. Classification of Multidimensional Perception Profiles

To integrate the three perception dimensions while retaining their distinct information, the point-level CCR, SO, and VC scores were independently classified into low, medium, and high levels using the Jenks natural breaks method (Table A3). The classification was conducted separately for each dimension because their score distributions differed. Each sampling point was subsequently represented by a three-dimensional perception profile comprising its respective CCR, SO, and VC levels. The Jenks classification was used as an exploratory and cartographic tool for internal relative classification. The resulting low, medium, and high levels should not be interpreted as universal perceptual thresholds, and continuous scores remained the primary measures for spatial autocorrelation and correlation analyses.
Based on the number and combination of low-value dimensions, the profiles were grouped into five general types: uniformly high profiles, mixed medium-to-high profiles, single-low profiles, dual-low profiles, and uniformly low profiles. Single- and dual-low profiles were further distinguished according to the dimensions in which low values occurred. This classification was used to interpret the multidimensional perceptual conditions of the sampling points, which may inform corresponding spatial strategies for urban renewal. Although unsupervised clustering was not applied, the profile classification preserved different combinations of CCR, SO, and VC, thereby avoiding the loss of dimension-specific information that may occur when relying only on a single comprehensive score.

3. Results

3.1. Performance of the SVI-Based Perception–Prediction Models

Before model training, sensitivity analysis of the BT-score classification showed that the three classification methods produced different class distributions (Figure 5). Jenks natural breaks consistently achieved the lowest WCSS and highest GVF across CCR, SO, and VC (Table A4), indicating greater within-class homogeneity and a better fit to the empirical BT-score distributions than quantile and equal-interval classification.
The model-training results show that the three ResNet50 models were able to distinguish among low, medium, and high perception grades and support subsequent prediction across the study area (Figure 6). The learning curves tended to stabilize after approximately 40 epochs, and the final models were trained for 80 epochs to ensure stable convergence. The five-fold cross-validation results showed that the CCR model achieved the highest performance, with a mean accuracy of 0.8638 and a mean Macro-F1 of 0.8606. The SO model achieved a mean accuracy of 0.8125 and a mean Macro-F1 of 0.7808, while the VC model achieved a mean accuracy of 0.8398 and a mean Macro-F1 of 0.7945. Overall, all three models demonstrated reasonable classification performance, although their predictive performance varied across the three dimensions. The full performance statistics, including accuracy, Macro-F1, macro precision, and macro recall, are provided in Table A5.
The stronger performance of the CCR model may reflect the greater visual salience of cues associated with historic character, such as traditional eaves, historic façades, and cultural symbols in SVI. By contrast, SO and VC depend on multiple interacting streetscape features, including openness, enclosure, interface continuity, greenery, traffic, and visual clutter, which may make these perceptions more difficult to classify. Overall, the mean accuracy of the three models exceeded 80% (Figure 6), while the Macro-F1 values varied across dimensions. This indicates that the models achieved a robust performance level comparable to that reported in recent street-view image classification studies [27]. The comparatively lower class-level performance for SO and VC may also reflect greater perceptual ambiguity in these dimensions, particularly for intermediate cases where spatial or visual characteristics are less clearly differentiated.
The independent external validation further showed significant agreement between human-derived perceptual rankings and model predictions. After BT screening, 68, 47, and 62 images were retained for CCR, SO, and VC, respectively. Significant positive rank correlations were observed for all three dimensions, with a Spearman’s ρ of 0.721 for CCR (95% CI: 0.590–0.806), 0.671 for SO (95% CI: 0.486–0.775), and 0.565 for VC (95% CI: 0.365–0.711), all p < 0.001 (Figure 7). These results indicate that the trained models generally preserved the relative perceptual ordering of previously unseen historic-district images. The detailed BT screening results, model predictions, and rank-based comparison data used for the independent external validation are provided in the Supplementary Materials.

3.2. Spatial Patterns of Multidimensional Streetscape Perception

The study-area-wide prediction results show that the three perception dimensions exhibited distinct spatial patterns (Figure 8). CCR displayed a pronounced node-centered pattern, with high-value areas mainly concentrated around Hubushan (A1 in Figure 2), Huanglou Park (B1 in Figure 2), Huilongwo (B9 in Figure 2), and other historical and cultural nodes. These areas contain more recognizable cultural symbols, traditional architectural elements, and historic street frontages. In contrast, low-value CCR areas were mainly located in residential neighborhoods, commercial streets, and locations where elements of historic character were fragmented or less visually prominent (Figure 9).
The predicted SO scores exhibited more evident corridor- and network-like patterns. High-value areas were distributed along major roads characterized by greater openness and spatial continuity, such as Huaihai Road, Zhongshan Road, and Jianguo Road, as well as around public parks, including Yunlong Park (E14 in Figure 2) and Kuaizaiting Park (B10 in Figure 2). These patterns indicate that higher SO scores were associated with street openness, spatial continuity, and the organization of street interfaces (Figure 9). Low-value SO areas were mainly distributed in narrow alleys, enclosed neighborhood streets, and older residential areas. These locations were characterized by limited visibility and stronger enclosure, corresponding to lower perceived spatial order (Figure 9).
Although the SO and VC maps show similar corridor-like tendencies, their spatial patterns are not identical. The predicted VC scores showed the coexistence of high-value corridors and localized low-value areas, including a high-VC corridor along the Huanghe Gudao (Figure 8d). Areas with moderately open views, orderly street interfaces, and higher green-view ratios were generally associated with higher VC scores (Figure 9). Conversely, low-value areas were commonly characterized by strong enclosure, fragmented façades, high wall proportions, traffic interference, and visual clutter, which partially overlap with low SO areas (Figure 9). Compared with CCR, VC was less dependent on specific historical nodes and exhibited a broader association with the organization of street interfaces, greenery, traffic, and visual clutter.
The comprehensive perception score integrates the spatial patterns of CCR, SO, and VC. High-value areas exhibited a network-like distribution centered on Hubushan (A1 in Figure 2) and extended along major roads and their intersections, while low-value areas were mainly distributed in the northwestern and southeastern parts of the study area. It can be observed that areas with high comprehensive scores generally exhibited higher values across the three perception dimensions, whereas areas with low comprehensive scores generally reflected weaker performance across multiple dimensions. However, it should be noted that the high-VC corridor along the Huanghe Gudao (Figure 8d) became less prominent in the comprehensive map because its CCR and SO scores were comparatively lower. The weighting-sensitivity analysis further showed that the point-level rankings were highly stable under the three alternative weighting scenarios. Spearman’s ρ values between the equal-weighted score and the CCR-, SO-, and VC-emphasis scenarios were 0.989, 0.997, and 0.994, respectively (all p < 0.001; Table A6). The corresponding mean absolute percentile-rank changes were only 2.91%, 1.49%, and 2.27%. Moreover, 88.1–94.0% of the sampling points in the top 10% and 94.5–97.2% of those in the bottom 10% were retained across the alternative weighting scenarios. These results indicate that the overall point-level ranking, as well as the identification of extreme high- and low-ranked locations, remained largely stable under the tested variations in dimensional weights.

3.3. Global Spatial Autocorrelation and High-Value Clustering

The global Moran’s I analysis shows that the CCR, SO, VC, and comprehensive scores all exhibit significant positive spatial autocorrelation (Table 3). This indicates that the predicted streetscape perception scores in the Pengcheng Qili historic district were not randomly distributed but formed spatially clustered patterns. Specifically, the Moran’s I values for CCR, SO, VC, and the comprehensive score were 0.5926, 0.5891, 0.6049, and 0.6058, respectively, with all p-values below 0.001 (Table 3).
Although the comprehensive score had the highest numerical Moran’s I, the differences among the four values were small. Overall, the results indicate that the predicted perception scores displayed substantial spatial continuity, with high-score points adjacent to other high-score points and low-score points near other low-score points.
The Getis–Ord General G statistic was used to determine whether the global spatial patterns were characterized by significant high- or low-value clustering. The observed General G values for CCR, SO, VC, and the comprehensive score were all higher than their expected values. The z-scores for CCR, SO, VC, and the comprehensive perception score were 19.6542, 26.2878, 26.5851, and 25.3876, respectively, with all p-values below 0.001 (Table 4). These results indicate that all four score sets exhibited significant global high-value clustering, meaning that sampling points with high perception scores tended to be spatially concentrated. As General G is a global statistic, these results establish an overall high-value clustering tendency but do not identify the locations of individual clusters. These locations are examined using local Gi analysis in Section 3.4. The sensitivity analysis showed consistently positive and significant spatial autocorrelation under both alternative spatial-weight specifications, with Moran’s I ranging from 0.526 to 0.587 across the three perception dimensions (all p < 0.001; Table 5). This consistency indicates that the principal spatial-clustering conclusions were robust to alternative definitions of spatial relationships.

3.4. Local Hot- and Cold-Spot Patterns

The local Gi* analysis confirmed that the spatial patterns identified in Section 3.2 included statistically significant hot and cold spots (Figure 10). CCR exhibited localized and node-centered clustering, with significant clusters concentrated along the central historic axis of the district. By contrast, SO and VC exhibited similar corridor-like clustering patterns. The comprehensive score map visually resembled the SO and VC patterns more closely than the localized CCR pattern. Thus, the Gi* analysis complements the spatial distribution maps (Figure 8) by distinguishing statistically significant local concentrations within the predicted scores. Since the comprehensive score is an equal-weight average, its hot and cold spots do not necessarily represent locations where all three dimensions are simultaneously high or low.

3.5. Inter-Dimensional Correlations of Perception Scores

The relationships among the three perception dimensions were further examined using correlation analysis. As shown in Figure 11, CCR, SO, and VC were all positively correlated, indicating that the three dimensions are not isolated components of streetscape perception. The strongest correlation was observed between SO and VC, suggesting that perceived spatial organization is closely associated with visual comfort. In contrast, CCR showed moderate correlations with SO and VC, indicating that cultural character recognition is related to spatial and visual conditions but cannot be fully substituted by either of them. These results further support the necessity of integrating CCR, SO, and VC within a multidimensional perception framework while retaining their distinct diagnostic meanings. Spearman correlation analysis produced similar positive associations among CCR, SO, VC, and the comprehensive perception score, confirming the robustness of the Pearson correlation results. The Spearman correlation matrix, with coefficients, is shown in Figure A1. Partial correlation analysis further revealed differentiated relationships among the three dimensions (Table 6). After controlling for VC, CCR remained moderately associated with SO (partial ρ = 0.480, p < 0.001), while the strong SO–VC association persisted after controlling for CCR (partial ρ = 0.867, p < 0.001). In contrast, the CCR–VC association became weakly negative after controlling for SO (partial ρ = −0.171, p < 0.001). These results further indicate that the three dimensions are interrelated but retain distinct perceptual information.

3.6. Multidimensional Perception Profiles and Renewal Priorities

Overlaying the classified CCR, SO, and VC levels produced five general types of multidimensional perception profiles: uniformly high, mixed medium-to-high, single-low, dual-low, and uniformly low profiles (Figure 12). The results show that low perception scores were not confined to a single dimension but occurred in both dimension-specific and overlapping forms.
Uniformly high profiles (Figure 12a) were mainly concentrated around Hubushan (A1 in Figure 2) and Huanglou Park (B1 in Figure 2) and along several major roads, whereas uniformly low profiles (Figure 12b) occurred primarily in residential areas in the northwestern and southeastern parts of the study area. Mixed medium-to-high profiles (Figure 12c) formed a broad network across the district. Among the single-low profiles, CCR-low (Figure 12i) points displayed the most extensive distribution, including a noticeable band along the Huanghe Gudao (Figure 8d). SO-low (Figure 12h) and VC-low (Figure 12g) profiles were more localized, while the dual-low profiles (Figure 12) were scattered across the northwestern area and residential neighborhoods south of Hubushan (A1 in Figure 2).
Table 7 summarizes the number and proportion of points in each multidimensional perception profile. Among the 2176 classified sampling points, mixed medium-to-high profiles (Figure 12c) accounted for the largest proportion (46.00%). Uniformly low profiles represented 18.66%, compared with 8.23% for uniformly high profiles. Single-low profiles accounted for 19.16%, of which CCR-low was the most frequent subtype (14.61%). Dual-low profiles represented a further 7.95%. These results indicate that weak CCR was the most common dimension-specific condition, while a considerable proportion of the study area exhibited simultaneously low levels across all three dimensions.

4. Discussion

4.1. Contributions to Perception-Oriented Diagnosis of Historic Districts

The main contributions of this study can be summarized in three aspects. First, this paper extends streetscape perception in historic districts from general visual-quality evaluation to a multidimensional diagnostic framework composed of CCR, SO, and VC. Existing SVI studies generally emphasize individual outcomes such as esthetics, safety, comfort, or vitality [9,10]. However, perceptual quality in historic districts depends not only on whether the visual environment is pleasant but also on whether cultural character is recognizable and street space is coherently organized. The distinct spatial patterns of CCR, SO, and VC, together with the occurrence of single- and multiple-low profiles, support their treatment as separate but complementary dimensions [4,10]. Their integration can therefore reveal dimension-specific and overlapping conditions that may be obscured by a single perception score.
Second, this paper establishes a technical pathway from limited subjective evaluation to study-area-wide spatial diagnosis. Through pairwise comparison and the Bradley–Terry model, the study converts relative preferences among SVI into perception labels for model training. ResNet50 was then used to predict the three perception dimensions across the study area and map the results geographically. This workflow does not aim to replace traditional expert surveys and field assessments but provides a reproducible and spatially explicit supplementary method [1]. Its particular value lies in identifying perception patterns in neighborhood streets, residential areas, and non-landmark spaces that may be overlooked as heritage assets in conventional studies [4,53].
Third, the study strengthens the spatial interpretation of the predicted scores through global Moran’s I, Getis–Ord General G, and local Gi* analysis. Moran’s I confirmed significant global spatial autocorrelation, General G identified a global tendency towards high-value clustering, and local Gi* helped identify statistically significant hot and cold spots. Combined with the multidimensional perception profiles, they provide stronger spatial evidence for subsequent diagnostic interpretation.

4.2. Spatial Differentiation of Cultural Character Recognition, Spatial Order and Visual Comfort

The results show that CCR, SO, and VC exhibit distinct but complementary spatial patterns in the Pengcheng Qili historic district. CCR is characterized by localized and node-centered clustering around heritage-rich areas. These locations contained visible heritage cues, including traditional eaves, cultural symbols, and historic street frontages. This pattern is consistent with the importance of the visibility and continuity of cultural elements in shaping the recognition of historic character [8]. However, the localized distribution of CCR also indicates that recognizable cultural character is less evident beyond major heritage nodes.
SO showed a stronger orientation towards road corridors and public open spaces. Open and continuous roads and areas surrounding parks were more likely to exhibit high SO scores, whereas narrow alleys and enclosed residential streets more frequently exhibited low values. These findings indicate that SO in historic districts depends on the relationship between openness, enclosure, continuity, and traditional street scale rather than on maximizing openness alone [17,54]. Moderate enclosure can help retain the spatial character of traditional streets, whereas excessive enclosure may reduce visibility and pedestrian legibility. VC also displayed corridor-like patterns but was more closely associated with the overall condition of street interfaces, greenery, parking, facilities, and visual clutter. The partial similarity between SO and VC, together with their difference from the node-centered CCR pattern, supports the distinction among the three dimensions.
Furthermore, the correlation results also show that the three dimensions are not interchangeable (Figure 11). CCR showed only moderate correlations with SO and VC, indicating that cultural character recognition is related to spatial and visual conditions but cannot be fully replaced by them. In other words, a street may have relatively good spatial order or visual comfort while still lacking recognizable historical character. Conversely, areas around heritage nodes may achieve high CCR values without necessarily showing equally high SO or VC values. The Huanghe Gudao (Figure 8d) corridor provides a useful example: it exhibited relatively high VC, but this advantage became less prominent in the comprehensive result because its CCR and SO scores were comparatively lower.
Therefore, the value of multidimensional diagnosis lies not in replacing CCR, SO, and VC with a single arithmetic mean, but in identifying where strengths in one dimension coexist with weaknesses in others. This distinction is important for historic-district renewal because different perceptual deficiencies require different planning responses. The combined interpretation of spatial patterns and inter-dimensional correlations therefore supports a more differentiated and diagnosis-oriented renewal approach.

4.3. Implications for Differentiated Renewal Strategies

The multidimensional perception profiles and hot- and cold-spot results provide a basis for differentiated renewal strategies. Rather than treating the historic district as a homogeneous renewal object, the profile-based diagnosis distinguishes areas with multidimensional advantages, balanced medium-level conditions, single-dimensional deficits, dual-dimensional deficits, and uniformly low perception levels.
Uniformly high profiles (Figure 12a), which accounted for 8.23% of the classified points, are more suitable for protective enhancement than intensive intervention. Renewal in these areas should avoid excessive commercialization, replacement of historic materials, and homogenization of visual character. Mixed medium-to-high profiles (Figure 12c) accounted for the largest proportion, reaching 46.00%, and therefore require incremental improvement and routine monitoring rather than comprehensive renewal.
Single-low profiles accounted for 19.16%, with CCR-low (Figure 12i) as the most frequent subtype. These profiles suggest dimension-specific improvement needs. CCR-low (Figure 12i) areas may benefit from enhancing the visibility, continuity, and interpretation of authentic cultural elements. SO-low (Figure 12h) areas indicate potential needs for improving pedestrian continuity, spatial legibility, interface permeability, and the balance between openness and enclosure. VC-low (Figure 12g) areas point to possible improvements in façade quality, signage order, parking management, street facilities, greenery, traffic exposure, and visual clutter.
Dual-low profiles accounted for 7.95% and indicate overlapping perceptual deficiencies. CCR–SO low (Figure 12e) areas suggest combined attention to cultural character and spatial organization; CCR–VC low (Figure 12d) areas indicate the need to consider both cultural visibility and visual environmental quality; and SO–VC low (Figure 12f) areas may benefit from the coordinated improvement of spatial structure and visual comfort. For dual-low profiles, interventions should accordingly combine the corresponding dimension-specific responses: CCR–SO deficiencies call for coordinated reinforcement of historical character and spatial-interface organization; CCR–VC deficiencies require cultural-character enhancement together with visual-clutter and environmental-quality management; and SO–VC deficiencies emphasize coordinated improvements in spatial organization and visual comfort.
Uniformly low profiles (Figure 12b) represented 18.66% of the classified points and indicate simultaneous weaknesses in CCR, SO, and VC. These areas may warrant first-priority renewal attention, particularly when they overlap with statistically significant cold spots. However, renewal priority should not be determined by perception scores alone; it should also be verified based on heritage significance, field conditions, residents’ needs, conservation requirements, and implementation feasibility. Overall, the framework supports a shift from generalized streetscape beautification toward profile-based and context-sensitive renewal.

4.4. Limitations and Future Research

This study has several limitations. First, although pairwise comparison reduces some of the instability associated with direct scoring, the number of annotated images and the composition of the participant sample may affect the resulting perception labels. Because the participants were primarily architecture students, the evaluations may reflect professionally informed perceptions more strongly than those of the general public. Future research should compare local residents, tourists, professionals, and other user groups to assess the transferability of the three perception dimensions. Second, SVI primarily records static visual information and cannot fully represent sound, smell, seasonal variation, temporal activity patterns, and social interaction [21,40]. Future studies could integrate multi-temporal SVI, on-site behavioral observation, walking trajectories, and environmental-sensor data to provide a more comprehensive understanding of streetscape perception. Third, the present study identifies perception patterns but does not quantitatively explain how individual physical elements contribute to CCR, SO, or VC. The representative images provide only qualitative interpretation. Future research could combine semantic segmentation and explainable modeling to assess the contributions and potential nonlinear effects of heritage elements, spatial configuration, greenery, traffic, and visual clutter. In addition, although the sensitivity analyses indicated that the comprehensive score rankings remained largely stable under the tested alternative weighting scenarios and that Jenks provided a better fit to the empirical BT-score distributions than the tested classification alternatives, both procedures remain data-dependent analytical choices. Future studies could further examine data-driven or stakeholder-informed weighting schemes and validate classification thresholds across different historic-district contexts. Finally, although the independent external validation provided additional evidence of model transferability to previously unseen historic-district images, the validation sample was limited to historic districts within Jiangsu Province. Further cross-case validation across different regions, heritage types, participant groups, and SVI sources is therefore still required.

5. Conclusions

Taking the Pengcheng Qili historic district in Xuzhou, China, as a case study, this paper developed a multidimensional streetscape perception diagnostic framework based on street-view imagery, subjective pairwise comparison, the Bradley–Terry model, and ResNet50-based prediction. The framework integrated three related but non-interchangeable dimensions—cultural character recognition (CCR), spatial order (SO), and visual comfort (VC)—and translated subjective perception judgments into spatially explicit diagnostic evidence through point-level aggregation, GIS mapping, and spatial statistical analysis.
The results show that the three perception dimensions exhibited distinct but related spatial patterns. High CCR scores were mainly concentrated around historical and cultural nodes and historic street frontages, while high SO scores were more strongly associated with open and continuous roads and public spaces. VC displayed corridor-like patterns related to street-interface quality, greenery, traffic exposure, and visual order. Correlation analysis further showed that CCR, SO, and VC were significantly and positively associated, with the strongest relationship between SO and VC and more moderate relationships between CCR and the other two dimensions. These findings indicate that the three dimensions are interrelated components of streetscape perception, but they retain distinct diagnostic meanings.
Global Moran’s I, Getis–Ord General G, and local Gi* analyses confirmed significant spatial clustering in the predicted perception scores. The multidimensional profile classification further revealed that perceptual deficiencies were not limited to single dimensions. Mixed medium-to-high profiles (Figure 12c) accounted for the largest proportion of classified points, while uniformly low profiles represented a considerable share. Among the single-low profiles, CCR-low was the most frequent subtype, indicating that weakened cultural character recognition is a prominent issue in the study area.
Overall, the proposed framework provides a reproducible and spatially explicit approach for diagnosing streetscape perception in historic districts. Independent external validation further showed significant correspondence between model predictions and human-derived perceptual rankings across all three dimensions. Future studies could further compare ResNet50 with alternative architectures, such as Vision Transformers, particularly when larger annotated datasets become available, to examine whether different feature-learning mechanisms improve the discrimination of perceptually ambiguous classes. By linking deep learning prediction, spatial statistical analysis, inter-dimensional correlation analysis, and multidimensional profile interpretation, the framework supports a shift from generalized streetscape beautification toward conservation-oriented, profile-based, and context-sensitive renewal.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/land15091640/s1.

Author Contributions

Conceptualization, P.L. and L.S.; methodology, P.L., J.L. and H.C.; software, P.L. and J.L.; validation, P.L., R.L. and H.C.; formal analysis, P.L.; investigation, P.L. and J.L.; resources, L.S.; data curation, P.L.; writing—original draft preparation, P.L. and J.L.; writing—review and editing, H.C. and L.S.; visualization, P.L. and J.L.; supervision, L.S.; project administration, R.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Fundamental Research Funds for the Central Universities, grant number 2024QN11058, and the Graduate Innovation Program of China University of Mining and Technology, grant number 2026WLJCRCZL278.

Institutional Review Board Statement

This study involved human participants through a non-sensitive, questionnaire-based pairwise comparison of street-view images. The study posed minimal risk to participants and did not involve medical interventions, the collection of biological samples, or sensitive personal information. According to Article 32 of the Measures for Ethical Review of Life Science and Medical Research Involving Human Subjects (jointly issued by the National Health Commission, the Ministry of Education, the Ministry of Science and Technology, and the National Administration of Traditional Chinese Medicine of China on 18 February 2023; Guowei Kejiao Fa [2023] No. 4), research using anonymized information or data that does not cause harm to human participants or involve sensitive personal information or commercial interests may be exempt from ethical review. The present study met these conditions.

Informed Consent Statement

All participants were informed of the purpose and procedures of the study, participated voluntarily, and provided informed consent prior to participation. Participant responses were anonymized, and privacy and confidentiality were protected throughout data collection, analysis, and reporting.

Data Availability Statement

The datasets generated and/or analyzed during the current study are not publicly available due to third-party restrictions associated with street-view images and field-survey data, but are available from the corresponding author on reasonable request.

Acknowledgments

The authors would like to thank the participants who contributed to the pairwise comparison experiment and the reviewers for their constructive comments.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
BTBradley–Terry
CCRCultural Character Recognition
CVComputer Vision
FCNFully Convolutional Network
GISGeographic Information System
Gi*Getis–Ord Gi*
HDHistoric District
OSMOpenStreetMap
QGISQGIS Geographic Information System
ResNet50Residual Network-50
SOSpatial Order
SVIStreet-View Imagery
VCVisual Comfort

Appendix A

Table A1. Statistics of Bradley–Terry score screening and uncertainty assessment.
Table A1. Statistics of Bradley–Terry score screening and uncertainty assessment.
DimensionInitial SamplesRetained SamplesMinimum ComparisonsMean Valid ComparisonsDifficult-to-Judge Responses (%)Retention Rate
CCR91050168.5753.7455.12%
SO910735811.5627.8780.77%
VC910768812.0322.8584.40%
Table A2. Training parameters of the ResNet50 perception–prediction models.
Table A2. Training parameters of the ResNet50 perception–prediction models.
ParameterSetting
Model architectureResNet50
Pretrained weightsImageNet pretrained weights
Classification taskThree-class classification: low, medium, and high
Classifier headDropout layer followed by a linear classification layer
Dropout rate0.30
Training strategyTwo-stage transfer learning
Stage 1Frozen ResNet50 backbone; only the classification head was trained
Stage 2The layer4 block and classification head were unfrozen and fine-tuned
Cross-validationFive-fold stratified cross-validation
OptimizerAdamW
Initial learning rate1 × 10−3 for the frozen-backbone stage
Fine-tuning learning rate1 × 10−5 for the fine-tuning stage
Weight decay1 × 10−4
Loss functionCross-entropy loss with label smoothing
Label smoothing0.05
Batch size16
Epochs30 frozen-backbone epochs and 50 fine-tuning epochs, 80 epochs in total
Learning-rate schedulerCosine annealing scheduler during the fine-tuning stage
Data augmentationResize to 256 × 256, random resized crop to 224 × 224, random horizontal flip, color jitter, and ImageNet normalization
Validation preprocessingResize to 256 × 256, center crop to 224 × 224, and ImageNet normalization
Random seed42
Model-selection criterionBest validation Macro-F1
Evaluation metricsAccuracy, Macro-F1, weighted F1, macro precision, macro recall, classification report, and confusion matrix
Saved outputsFold-level training logs, best epoch information, fold-level and combined confusion matrices, all-fold validation predictions, and training-time summary
Table A3 reports the Jenks natural break intervals used to classify the continuous, point-level perception scores into low, medium, and high levels. The classification was conducted separately for CCR, SO, VC, and comprehensive perception because the score distributions differed across dimensions. These intervals were used to support the diagnostic coding process described in Section 2.8. In the diagnostic coding system, the categories were primarily derived from the three single-dimensional levels of CCR, SO, and VC, while the comprehensive perception level was used as supplementary information for overall spatial interpretation.
Table A3. Jenks natural breaks intervals for classifying point-level perception scores.
Table A3. Jenks natural breaks intervals for classifying point-level perception scores.
DimensionLowMediumHigh
CCR0.0232–0.22190.2222–0.37740.3780–0.9175
SO0.0118–0.30640.3065–0.50650.5072–0.8403
VC0.0141–0.33990.3402–0.55730.5582–0.9332
Comprehensive perception0.0181–0.29250.2926–0.46070.4614–0.7965
Table A4. Sensitivity comparison of classification methods for the retained BT scores used in model-label construction.
Table A4. Sensitivity comparison of classification methods for the retained BT scores used in model-label construction.
DimensionMethodLow (n)Medium (n)High (n)WCSSGVF
CCRJenks25319256195.4510.845
Quantile167167167339.3470.731
Equal interval9135456303.2780.760
SOJenks279284172228.0260.831
Quantile245245245239.1540.823
Equal interval154444137317.8040.765
VCJenks290290188271.0240.839
Quantile256256256291.7710.826
Equal interval154473141423.9430.748
WCSS, within-class sum of squares; GVF, goodness of variance fit. Lower WCSS and higher GVF indicate greater within-class homogeneity and better representation of the empirical score distribution.
Table A5. Five-fold cross-validation performance of the ResNet50 perception–prediction models.
Table A5. Five-fold cross-validation performance of the ResNet50 perception–prediction models.
DimensionAccuracyMacro-F1Macro PrecisionMacro Recall
CCR0.8638 ± 0.02900.8606 ± 0.02530.8828 ± 0.02010.8452 ± 0.0343
SO0.8125 ± 0.01870.7808 ± 0.02080.7897 ± 0.02040.7752 ± 0.0240
VC0.8398 ± 0.02310.7945 ± 0.03650.8200 ± 0.02390.7791 ± 0.0471
Table A6. Sensitivity of the comprehensive perception scores to alternative weighting scenarios.
Table A6. Sensitivity of the comprehensive perception scores to alternative weighting scenarios.
ComparisonSpearman’s ρMean Absolute Percentile-Rank Change (%)Top 10% Overlap (%)Bottom 10% Overlap (%)p-Value
Equal weighting vs. CCR-emphasis0.9892.9188.194.5<0.001
Equal weighting vs. SO-emphasis0.9971.4994.096.8<0.001
Equal weighting vs. VC-emphasis0.9942.2791.397.2<0.001
The equal-weighted scenario (CCR = SO = VC = 1/3) was used as the reference. In each alternative scenario, the focal dimension was assigned a weight of 0.50 and the remaining two dimensions were assigned weights of 0.25. Mean absolute percentile-rank change represents the average absolute shift in point-level percentile ranking relative to the equal-weighted scenario. Top and bottom 10% overlap represent the proportion of sampling points retained within the corresponding extreme-ranking group.
Figure A1. Spearman correlation matrix among cultural character recognition (CCR), spatial order (SO), visual comfort (VC), and comprehensive perception scores.
Figure A1. Spearman correlation matrix among cultural character recognition (CCR), spatial order (SO), visual comfort (VC), and comprehensive perception scores.
Land 15 01640 g0a1

References

  1. Tang, J.; Long, Y. Measuring Visual Quality of Street Space and Its Temporal Variation: Methodology and Its Application in the Hutong Area in Beijing. Landsc. Urban Plan. 2019, 191, 103436. [Google Scholar] [CrossRef] [Scilit]
  2. Ren, L.; Xiong, W.; Li, J.; Namaiti, A.; Zhuang, J. Revealing the Nonlinear Effects of Traditional Rural Streetscape Features on Visual Quality Based on the XGBoost-SHAP Model. Environ. Impact Assess. Rev. 2026, 118, 108282. [Google Scholar] [CrossRef] [Scilit]
  3. Xu, H.; Yang, T.; Guo, Z. Exploring Nonlinear Effects of Visual Elements on Perceived Landscape Quality in Historical and Cultural Districts: A Deep Learning Case from Wuhan, China. Buildings 2025, 15, 4338. [Google Scholar] [CrossRef] [Scilit]
  4. Tang, N.; Wang, S.; Lyu, M. Urban Style and Features’ Visual Quality and Influencing Factors: A Case Study of Fangcheng Historical and Cultural District in Shenyang, China. Buildings 2026, 16, 1455. [Google Scholar] [CrossRef] [Scilit]
  5. Li, P.; Xu, Y.; Liu, Z.; Jiang, H.; Liu, A. Evaluation and Optimization of Urban Street Spatial Quality Based on Street View Images and Machine Learning: A Case Study of the Jinan Old City. Buildings 2025, 15, 1408. [Google Scholar] [CrossRef] [Scilit]
  6. Ma, Y.; Wang, L.; Zhang, J. Quantifying Spatial Openness and Visual Perception in Historic Urban Environments. Buildings 2025, 15, 3295. [Google Scholar] [CrossRef] [Scilit]
  7. Gao, X.; Wang, H.; Zhao, J.; Wang, Y.; Li, C.; Gong, C. Visual Comfort Impact Assessment for Walking Spaces of Urban Historic District in China Based on Semantic Segmentation Algorithm. Environ. Impact Assess. Rev. 2025, 114, 107917. [Google Scholar] [CrossRef] [Scilit]
  8. Shen, Z.; Xie, D.; Su, C. Street View Image-Based Method for Assessing the Thermal Environment in Urban Historic Districts: A Case Study of Guangzhou’s Arcade Streets. Build. Environ. 2026, 288, 114022. [Google Scholar] [CrossRef] [Scilit]
  9. Yang, X.; Shen, J. Examining Streetscape Visuals and Emotional Responses through Social Media and Street View Image Analysis. Int. J. Health Geogr. 2025, 25, 15–37. [Google Scholar] [CrossRef] [Scilit]
  10. Wang, Z.; Zhang, W.; Huang, Y. Nonlinear Perceptual Thresholds and Trade-Offs of Visual Environment in Historic Districts: Evidence from Street View Images in Shanghai. Sustainability 2025, 17, 11075. [Google Scholar] [CrossRef] [Scilit]
  11. Li, M.; Fan, Z. Constructing High-Quality Livable Cities: A Comprehensive Evaluation of Urban Street Livability Using an Approach Based on Human Needs Theory, Street View Images, and Deep Learning. Land 2025, 14, 1095. [Google Scholar] [CrossRef] [Scilit]
  12. Han, X.; Zhu, Y.; Wang, L.; Guo, Z. Enhancing the Understanding of Urban Street Perception with LLM s and Street View Imagery. Trans. GIS 2026, 30, e70280. [Google Scholar] [CrossRef] [Scilit]
  13. Sun, M.; Zhang, F.; Duarte, F.; Ratti, C. Understanding Architecture Age and Style through Deep Learning. Cities 2022, 128, 103787. [Google Scholar] [CrossRef] [Scilit]
  14. Biljecki, F.; Ito, K. Street View Imagery in Urban Analytics and GIS: A Review. Landsc. Urban Plan. 2021, 215, 104217. [Google Scholar] [CrossRef] [Scilit]
  15. Jiang, Z.; Wu, C.; Chung, H. The 15-Minute Community Life Circle for Older People: Walkability Measurement Based on Service Accessibility and Street-Level Built Environment—A Case Study of Suzhou, China. Cities 2025, 157, 105587. [Google Scholar] [CrossRef] [Scilit]
  16. Wang, L.; Han, X.; He, J.; Jung, T. Measuring Residents’ Perceptions of City Streets to Inform Better Street Planning through Deep Learning and Space Syntax. ISPRS J. Photogramm. Remote Sens. 2022, 190, 215–230. [Google Scholar] [CrossRef] [Scilit]
  17. Xu, J.; Dai, Y.; Cai, J.; Qian, H.; Peng, Z.; Zhong, T. Evaluation of Urban Street Historical Appearance Integrity Based on Street View Images and Transfer Learning. ISPRS Int. J. Geo-Inf. 2025, 14, 266. [Google Scholar] [CrossRef] [Scilit]
  18. Zhao, T.; Liang, X.; Tu, W.; Huang, Z.; Biljecki, F. Sensing Urban Soundscapes from Street View Imagery. Comput. Environ. Urban Syst. 2023, 99, 101915. [Google Scholar] [CrossRef] [Scilit]
  19. Hou, Y.; Quintana, M.; Khomiakov, M.; Yap, W.; Ouyang, J.; Ito, K.; Wang, Z.; Zhao, T.; Biljecki, F. Global Streetscapes—A Comprehensive Dataset of 10 Million Street-Level Images across 688 Cities for Urban Science and Analytics. ISPRS J. Photogramm. Remote Sens. 2024, 215, 216–238. [Google Scholar] [CrossRef] [Scilit]
  20. Zhang, L.; Huang, X.; Zhong, H. Measurement of Street Greenness and Interface Permeability Based on Street View Image Analysis. Trait. Signal 2024, 41, 1679–1688. [Google Scholar] [CrossRef] [Scilit]
  21. Qiu, W.; Zhang, Z.; Liu, X.; Li, W.; Li, X.; Xu, X.; Huang, X. Subjective or Objective Measures of Street Environment, Which Are More Effective in Explaining Housing Prices? Landsc. Urban Plan. 2022, 221, 104358. [Google Scholar] [CrossRef] [Scilit]
  22. Zhang, P.; Liu, Y.; Huang, Y. Dynamic Assessment of Street Environmental Quality Using Time-Series Street View Imagery within Daily Intervals. Land 2025, 14, 1544. [Google Scholar] [CrossRef] [Scilit]
  23. Bai, Z.; Mao, Y.; Gao, F.; Liu, S. Pedestrian and Cycling Vitality in Scenic and Downtown Contexts: Decoding Nonlinear Associations with the Built Environment via Geo-XAI. Cities 2026, 172, 106906. [Google Scholar] [CrossRef] [Scilit]
  24. Knoblauch, S.; Li, H.; Biljecki, F.; Li, W.; Zipf, A. Urban AI for a Sustainable Built Environment: Progress and Future Directions. Environ. Plan. B Urban Anal. City Sci. 2026, 53, 255–262. [Google Scholar] [CrossRef] [Scilit]
  25. Ogawa, Y.; Oki, T.; Zhao, C.; Sekimoto, Y.; Shimizu, C. Evaluating the Subjective Perceptions of Streetscapes Using Street-View Images. Landsc. Urban Plan. 2024, 247, 105073. [Google Scholar] [CrossRef] [Scilit]
  26. Yao, Y.; Liang, Z.; Yuan, Z.; Liu, P.; Bie, Y.; Zhang, J.; Wang, R.; Wang, J.; Guan, Q. A Human-Machine Adversarial Scoring Framework for Urban Perception Assessment Using Street-View Images. Int. J. Geogr. Inf. Sci. 2019, 33, 2363–2384. [Google Scholar] [CrossRef] [Scilit]
  27. Ma, H.; Li, J.; Ye, X. Deep Learning Meets Urban Design: Assessing Streetscape Aesthetic and Design Quality through AI and Cluster Analysis. Cities 2025, 162, 105939. [Google Scholar] [CrossRef] [Scilit]
  28. Yang, H.; Zhang, Q.; Helbich, M.; Lu, Y.; He, D.; Ettema, D.; Chen, L. Examining Non-Linear Associations between Built Environments around Workplace and Adults’ Walking Behaviour in Shanghai, China. Transp. Res. Part A Policy Pract. 2022, 155, 234–246. [Google Scholar] [CrossRef] [Scilit]
  29. Xia, Z.; Zhang, X.; Zhai, G.; Zhang, Y. Integrating Visual Spatial Vulnerability to Quantify Fire-Prone Neighborhoods in Cities: A Case Study of Nanjing, China. Int. J. Disaster Risk Reduct. 2025, 128, 105758. [Google Scholar] [CrossRef] [Scilit]
  30. Wei, Z.; Cao, K.; Kwan, M.-P.; Jiang, Y.; Feng, Q. Measuring the Age-Friendliness of Streets’ Walking Environment Using Multi-Source Big Data: A Case Study in Shanghai, China. Cities 2024, 148, 104829. [Google Scholar] [CrossRef] [Scilit]
  31. Li, W.; Li, Y.; Zhang, L.; Gao, J.; Xie, S.; Feng, Y. Road-Type-Specific Streetscape Renewal Effects on Urban Beauty Perception: A Spatiotemporal SHAP Analysis Using Historical Street Views. Buildings 2026, 16, 653. [Google Scholar] [CrossRef] [Scilit]
  32. Liu, Z.; Feng, X.; Yang, H.; Chang, F.; Zhu, D. Impact of Visual Features and Urban Form on Street Vitality Using Explainable AI: A Study of Historic Districts in Shanghai. All Earth 2026, 38, 2647240. [Google Scholar] [CrossRef] [Scilit]
  33. Xiong, X.; Wu, Y.; Ma, M.; Yang, S.; Zhang, J.; Zhang, Q.; Ye, H.; Hu, Y. Exploring the Multidimensional Visual Perception of Urban Riverfront Street Environments: A Framework Using Street View Images, Deep Learning and Eye-Tracking. Land 2025, 14, 2039. [Google Scholar] [CrossRef] [Scilit]
  34. Lu, K.; Zhao, X.; Li, M.; Wang, Z.; Zhou, Y.; Wang, J. StreetSenser: A Novel Approach to Sensing Street View via a Fine-Tuned Multimodal Large Language Model. Int. J. Geogr. Inf. Sci. 2026, 40, 1547–1575. [Google Scholar] [CrossRef] [Scilit]
  35. Ma, S.; Wang, B.; Liu, W.; Zhou, H.; Wang, Y.; Li, S. Assessment of Street Space Quality and Subjective Well-Being Mismatch and Its Impact, Using Multi-Source Big Data. Cities 2024, 147, 104797. [Google Scholar] [CrossRef] [Scilit]
  36. Klemm, W.; Heusinkveld, B.G.; Lenzholzer, S.; Van Hove, B. Street Greenery and Its Physical and Psychological Impact on Thermal Comfort. Landsc. Urban Plan. 2015, 138, 87–98. [Google Scholar] [CrossRef] [Scilit]
  37. Wang, R.; Lu, Y.; Zhang, J.; Liu, P.; Yao, Y.; Liu, Y. The Relationship between Visual Enclosure for Neighbourhood Street Walkability and Elders’ Mental Health in China: Using Street View Images. J. Transp. Health 2019, 13, 90–102. [Google Scholar] [CrossRef] [Scilit]
  38. Kuang, B.; Yang, H.; Jung, T. The Impact of Visual Elements in Street View on Street Quality: A Quantitative Study Based on Deep Learning, Elastic Net Regression, and SHapley Additive exPlanations (SHAP). Sustainability 2025, 17, 3454. [Google Scholar] [CrossRef] [Scilit]
  39. Lin, Y.; Liu, W.; Sun, X. Explainable Vision Analytics for Adaptive Campus Design: Diagnosing Multi-Dimensional Perceptual Differences. Buildings 2026, 16, 1623. [Google Scholar] [CrossRef] [Scilit]
  40. Lu, Y. Using Google Street View to Investigate the Association between Street Greenery and Physical Activity. Landsc. Urban Plan. 2019, 191, 103435. [Google Scholar] [CrossRef] [Scilit]
  41. Wang, R.; Feng, Z.; Pearce, J.; Zhou, S.; Zhang, L.; Liu, Y. Dynamic Greenspace Exposure and Residents’ Mental Health in Guangzhou, China: From over-Head to Eye-Level Perspective, from Quantity to Quality. Landsc. Urban Plan. 2021, 215, 104230. [Google Scholar] [CrossRef] [Scilit]
  42. Zhang, P.; Ghosh, D.; Park, S. Spatial Measures and Methods in Sustainable Urban Morphology: A Systematic Review. Landsc. Urban Plan. 2023, 237, 104776. [Google Scholar] [CrossRef] [Scilit]
  43. Kang, Y.; Chen, J.; Liu, L.; Sharma, K.; Mazzarello, M.; Mora, S.; Duarte, F.; Ratti, C. Decoding Human Safety Perception with Eye-Tracking Systems, Street View Images, and Explainable AI. Comput. Environ. Urban Syst. 2026, 123, 102356. [Google Scholar] [CrossRef] [Scilit]
  44. He, D.; Miao, J.; Lu, Y.; Song, Y.; Chen, L.; Liu, Y. Urban Greenery Mitigates the Negative Effect of Urban Density on Older Adults’ Life Satisfaction: Evidence from Shanghai, China. Cities 2022, 124, 103607. [Google Scholar] [CrossRef] [Scilit]
  45. Chen, L.; Lu, Y.; Ye, Y.; Xiao, Y.; Yang, L. Examining the Association between the Built Environment and Pedestrian Volume Using Street View Images. Cities 2022, 127, 103734. [Google Scholar] [CrossRef] [Scilit]
  46. He, J.; Zhang, J.; Yao, Y.; Li, X. Extracting Human Perceptions from Street View Images for Better Assessing Urban Renewal Potential. Cities 2023, 134, 104189. [Google Scholar] [CrossRef] [Scilit]
  47. Jin, A.; Ge, Y.; Zhang, S. Spatial Characteristics of Multidimensional Urban Vitality and Its Impact Mechanisms by the Built Environment. Land 2024, 13, 991. [Google Scholar] [CrossRef] [Scilit]
  48. Yu, X.; Ma, J.; Tang, Y.; Yang, T.; Jiang, F. Can We Trust Our Eyes? Interpreting the Misperception of Road Safety from Street View Images and Deep Learning. Accid. Anal. Prev. 2024, 197, 107455. [Google Scholar] [CrossRef] [Scilit]
  49. Wu, Y.; Liu, Q.; Hang, T.; Yang, Y.; Wang, Y.; Cao, L. Integrating Restorative Perception into Urban Street Planning: A Framework Using Street View Images, Deep Learning, and Space Syntax. Cities 2024, 147, 104791. [Google Scholar] [CrossRef] [Scilit]
  50. Qin, J.; Feng, Y.; Sheng, Y.; Huang, Y.; Zhang, F.; Zhang, K. Evaluation of Pedestrian-Perceived Comfort on Urban Streets Using Multi-Source Data: A Case Study in Nanjing, China. ISPRS Int. J. Geo-Inf. 2025, 14, 63. [Google Scholar] [CrossRef] [Scilit]
  51. Gu, N.; Liao, P.; Yu, R. The Mathematics of Historic Chinese-Built Heritage: Syntactical Properties and Implications. In Handbook of the Mathematics of the Arts and Sciences; Sriraman, B., Ed.; Springer Nature Switzerland: Cham, Switzerland, 2026; pp. 1–14. ISBN 978-3-319-70658-0. [Google Scholar]
  52. Xu, D.; Wu, H.; Yao, Q.; Song, F.; Su, F. Spatiotemporal Evolution and Driving Factors of Desertification Sensitivity during Urbanization: A Case Study of the Beijing–Tianjin–Hebei Core Region. Land 2025, 14, 858. [Google Scholar] [CrossRef] [Scilit]
  53. Zhang, H.; Jiang, J.; Guo, X. HSVI-Net: A Deep Learning Network for Scene Understanding Based on Panoramic Street-View Images within Historical Districts. npj Herit. Sci. 2025, 13, 383–398. [Google Scholar] [CrossRef] [Scilit]
  54. Liao, P.; Li, J.; Chen, H.; Sun, L. Decoding Space Genes in Traditional Village Morphology: A Quantitative Study of Xishan Island, Suzhou. npj Herit. Sci. 2026, 1–40. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Location of the study area and street-view image sampling in the Pengcheng Qili historic district, China.
Figure 1. Location of the study area and street-view image sampling in the Pengcheng Qili historic district, China.
Land 15 01640 g001
Figure 2. Spatial distribution of historical and cultural resources in the Pengcheng Qili Historic District.
Figure 2. Spatial distribution of historical and cultural resources in the Pengcheng Qili Historic District.
Land 15 01640 g002
Figure 3. Research framework.
Figure 3. Research framework.
Land 15 01640 g003
Figure 5. Comparison of BT-score distributions under Jenks natural breaks, quantile, and equal-interval classifications.
Figure 5. Comparison of BT-score distributions under Jenks natural breaks, quantile, and equal-interval classifications.
Land 15 01640 g005
Figure 6. Performance of the ResNet50 perception–prediction models and confusion matrices.
Figure 6. Performance of the ResNet50 perception–prediction models and confusion matrices.
Land 15 01640 g006
Figure 7. Rank-based independent external validation of the ResNet50 perception–prediction models.
Figure 7. Rank-based independent external validation of the ResNet50 perception–prediction models.
Land 15 01640 g007
Figure 8. Spatial distributions of the CCR, SO, VC, and comprehensive perception scores.
Figure 8. Spatial distributions of the CCR, SO, VC, and comprehensive perception scores.
Land 15 01640 g008
Figure 9. Locations and representative street-view examples of high and low predicted levels of cultural character recognition (CCR), spatial order (SO), and visual comfort (VC).
Figure 9. Locations and representative street-view examples of high and low predicted levels of cultural character recognition (CCR), spatial order (SO), and visual comfort (VC).
Land 15 01640 g009
Figure 10. Hot- and cold-spot analysis of multidimensional streetscape perception scores.
Figure 10. Hot- and cold-spot analysis of multidimensional streetscape perception scores.
Land 15 01640 g010
Figure 11. Pearson correlation matrix among cultural character recognition (CCR), spatial order (SO), visual comfort (VC), and comprehensive perception scores.
Figure 11. Pearson correlation matrix among cultural character recognition (CCR), spatial order (SO), visual comfort (VC), and comprehensive perception scores.
Land 15 01640 g011
Figure 12. Spatial distribution of multidimensional perception profiles based on the classified CCR, SO, and VC levels.
Figure 12. Spatial distribution of multidimensional perception profiles based on the classified CCR, SO, and VC levels.
Land 15 01640 g012
Table 1. Data sources and sample construction.
Table 1. Data sources and sample construction.
ItemValue
Sampling interval50 m
Sampling points2243
Street-view directions0°, 90°, 180°, 270°
Valid street-view images8693
Representative images for subjective evaluation910
Participants71
Pairwise comparison records20,475
Valid CCR perception samples501
Valid SO perception samples735
Valid VC perception samples768
Table 2. Definitions of the three perception dimensions.
Table 2. Definitions of the three perception dimensions.
DimensionDiagnostic Meaning
CCRRefers to the perceived recognizability of historic character, traditional architectural features, local cultural symbols, and place-specific cultural atmosphere within the streetscape.
SORefers to the perceived organization and coherence of street space at the pedestrian scale, including openness and enclosure, street scale, interface continuity, and spatial legibility.
VCRefers to the perceived comfort and pleasantness of the overall visual environment, as influenced by façade condition, greenery, sky visibility, traffic, street facilities, and visual clutter.
Table 3. Global Moran’s I results for multidimensional streetscape perception scores.
Table 3. Global Moran’s I results for multidimensional streetscape perception scores.
DimensionMoran’s IExpected Iz-Scorep-ValueSpatial Pattern
CCR0.5926−0.000553.7063<0.001Significant clustering
SO0.5891−0.000553.3525<0.001Significant clustering
VC0.6049−0.000554.7757<0.001Significant clustering
Comprehensive score0.6058−0.000554.8650<0.001Significant clustering
Table 4. Getis–Ord General G results for high/low clustering.
Table 4. Getis–Ord General G results for high/low clustering.
DimensionObserved General GExpected General Gz-Scorep-ValueClustering Type
CCR0.0041420.00344219.654249<0.001High-value clustering
SO0.0042520.00344226.287766<0.001High-value clustering
VC0.0042970.00344226.585145<0.001High-value clustering
Comprehensive perception0.0041880.00344225.387634<0.001High-value clustering
Table 5. Sensitivity analysis of Global Moran’s I under alternative spatial-weight specifications.
Table 5. Sensitivity analysis of Global Moran’s I under alternative spatial-weight specifications.
DimensionSpatial-Weight Specificationz-Scorep-ValueMoran’s I
CCRFixed-distance band (100 m)59.12<0.0010.544
CCR8-nearest neighbors57.14<0.0010.574
SOFixed-distance band (100 m)58.33<0.0010.526
SO8-nearest neighbors57.42<0.0010.560
VCFixed-distance band (100 m)60.06<0.0010.552
VC8-nearest neighbors58.65<0.0010.587
Table 6. Partial Spearman correlations.
Table 6. Partial Spearman correlations.
RelationshipControlled DimensionPartial Spearman ρp-Value
CCR–SOVC0.480<0.001
CCR–VCSO−0.171<0.001
SO–VCCCR0.867<0.001
Table 7. Distribution of multidimensional perception profiles.
Table 7. Distribution of multidimensional perception profiles.
General ProfileProfile SubtypeCountPercentage (%)Renewal Implication
Uniformly highUniformly high1798.23Protective Enhancement Area
Uniformly lowUniformly low40618.66First-Priority Renewal Area
Mixed medium-to-highMixed medium-to-high100146.00General Coordination Area
Dual-lowCCR–VC low642.94Second-Priority Renewal Area
Dual-lowCCR–SO low281.29
Dual-lowSO–VC low813.72
Single-lowCCR-low31814.61Third-Priority Optimization Area
Single-lowVC-low833.81
Single-lowSO-low160.74
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Li, J.; Chen, H.; Liao, P.; Sun, L.; Liu, R. A Multidimensional Framework for Diagnosing Streetscape Perception in Historic-District Renewal Using Street-View Imagery and Deep Learning. Land 2026, 15, 1640. https://doi.org/10.3390/land15091640

AMA Style

Li J, Chen H, Liao P, Sun L, Liu R. A Multidimensional Framework for Diagnosing Streetscape Perception in Historic-District Renewal Using Street-View Imagery and Deep Learning. Land. 2026; 15(9):1640. https://doi.org/10.3390/land15091640

Chicago/Turabian Style

Li, Jilong, Hongyang Chen, Pan Liao, Liang Sun, and Ruxin Liu. 2026. "A Multidimensional Framework for Diagnosing Streetscape Perception in Historic-District Renewal Using Street-View Imagery and Deep Learning" Land 15, no. 9: 1640. https://doi.org/10.3390/land15091640

APA Style

Li, J., Chen, H., Liao, P., Sun, L., & Liu, R. (2026). A Multidimensional Framework for Diagnosing Streetscape Perception in Historic-District Renewal Using Street-View Imagery and Deep Learning. Land, 15(9), 1640. https://doi.org/10.3390/land15091640

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop