Next Article in Journal
The Landscape–Hydrological Organization of the East Kazakhstan Region
Previous Article in Journal
The Correlations of Blight, Crime, and Litter: A Path Analysis with Urban Morphology
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Analysis and Considerations to Lessen Visual Clutter: Multidimensional Associations Between Streetscape Visual Clutter and Urban Perceptions in Shanghai

1
Department of Landscape Architecture, College of Architecture and Urban Planning, Tongji University, Shanghai 200092, China
2
Center of Ecological Planning and Environment Effects Research, Shanghai Joint Laboratory of Ecological Urban Design, Tongji University, Shanghai 200092, China
*
Author to whom correspondence should be addressed.
Land 2026, 15(8), 1402; https://doi.org/10.3390/land15081402
Submission received: 3 July 2026 / Revised: 28 July 2026 / Accepted: 29 July 2026 / Published: 4 August 2026
(This article belongs to the Section Land Planning and Landscape Architecture)

Abstract

Visual clutter can affect attention, cognitive load, and environmental evaluation, yet streetscape studies often treat visual complexity as a broad construct without distinguishing visual richness from competition among visual features. Shanghai was selected because its dense, heterogeneous urban fabric and central–peripheral variation make it well suited to city-scale analysis. Using 69,244 street-view images, this study combined computer vision, deep learning, and modeling to quantify edge, color, semantic, and depth clutter and estimate beauty, wealth, liveliness, safety, boredom, and depressiveness. After controlling for built-environment characteristics, 23 of 24 linear associations were significant, although their directions differed. Edge and color clutter were associated with higher positive perceptions and lower negative perceptions. Semantic clutter was positively associated with wealth, liveliness, and safety but negatively associated with beauty; depth clutter was negatively associated with beauty, wealth, and liveliness and positively associated with boredom and depressiveness. Stable inverted U-shaped patterns occurred in only four relationships. Streetscape management in Shanghai should avoid uniformly increasing or reducing visual information. Planning and design may retain architectural detail, vegetation edges, coordinated color variation, and semantic richness while better coordinating facades, signage, street furniture, vehicles, and foreground elements, improving sightline continuity, and reducing obstruction and abrupt spatial discontinuity.

1. Introduction

Urban perception refers to the subjective evaluations that people form about the safety, attractiveness, liveliness, social characteristics, and emotional atmosphere of urban places based on environmental cues. Early studies commonly used pairwise comparisons of street-view images to quantify perceptions of spatial attributes such as safety, social class, and uniqueness [1]. With the development of street-view imagery, deep learning, and large-scale perception datasets, recent studies have expanded this framework to multiple dimensions, including beauty, wealth, liveliness, safety, boredom, and depressiveness [2,3]. Previous research has also shown that urban perception is associated with health-related outcomes, psychological distress, and urban well-being [4,5]. Urban perception can therefore provide an important evidence base for health-oriented urban planning and management.
Vision is one of the primary ways through which people perceive urban environments [6]. Visual features in street spaces can shape cognition, attention allocation, and environmental evaluation [7,8]. Accordingly, a growing body of research has examined visual characteristics of the built environment, including streetscape composition [9], architectural form [10], color [11], visual complexity [12], and visual coherence and order [13]. In visual psychology, visual clutter refers to the spatial congestion of visual information, in which visual features or object properties are densely packed and change repeatedly within a limited area [14]. Such frequent local transitions can generate competing visual signals, thereby increasing the difficulty of visual search, object recognition, and information processing under the limited capacity of human visual attention [15]. Visual complexity, by contrast, refers more broadly to the overall amount, richness, diversity, or degree of stimulation provided by visual information. The two concepts are related but not equivalent. Streetscapes containing similar amounts or types of visual information may differ substantially in clutter: visual elements organized into large and continuous regions produce relatively few local transitions, whereas the same elements divided into small, closely interspersed regions produce denser transitions and greater visual congestion. Accordingly, the conceptual distinction adopted in this study lies in the spatial density of feature transitions rather than in the overall amount or diversity of visual information.
Existing studies in visual psychology and computer vision have developed various approaches to measuring visual clutter. These approaches examine factors such as feature congestion, edge density, color transitions, and semantic fragmentation [16]. In the present study, streetscape visual clutter is operationalized as the spatial density of feature transitions in street-view images. For edge, color, semantic, and depth clutter, respectively, the indicators quantify the density of structural contours and of boundaries between adjacent dominant color regions, semantic object categories, and relative spatial-depth layers. Across all four dimensions, denser boundaries indicate that the corresponding visual feature changes more frequently within the image and is therefore more spatially congested. The proposed indicators therefore characterize the spatial congestion of feature transitions rather than the number or diversity of visual elements per se.
This operationalization is consistent with previous studies showing that clutter arises from the combined effects of multiple visual features, which can intensify information competition, distract attention, and increase the difficulty of visual recognition [14,17,18]. However, most of these concepts and methods have been applied to general images, interface design, or controlled experimental settings. They have not yet been fully translated into a multidimensional framework for studying perceptions of urban streetscapes and open spaces. Existing studies of streetscapes and open spaces have widely examined visual complexity and the amount of visual information, but have paid less attention to the distinct perceptual effects of different types of visual information. Moreover, the association between visual clutter and urban perception may not be strictly linear [8,19,20,21]. Previous research has also focused largely on small neighborhoods or individual sites, leaving citywide evidence limited. It is therefore necessary to examine both the differentiated and nonlinear associations between multiple types of streetscape visual clutter and multidimensional urban perceptions at the city scale.
The development of street-view imagery and deep learning has created new opportunities for large-scale analysis of urban visual environments. Compared with traditional field surveys and questionnaires, street-view imagery records urban spaces from a human-centered perspective and supports continuous visual assessment over large areas [22]. In recent years, deep learning methods have been widely used to identify streetscape elements, estimate urban perceptions, and quantify visual characteristics [2,9,22,23]. In particular, deep learning models trained on the Place Pulse 2.0 dataset have shown a reliable capacity to predict perceptual dimensions such as beauty, safety, liveliness, and depressiveness in urban environments [24]. These developments provide a methodological basis for translating the concept of visual clutter from visual psychology into city-scale streetscape indicators and for examining its associations with urban perceptions.
Building on previous research, this study addresses the limited city-scale evidence on whether different dimensions of streetscape visual clutter show distinct associations with specific urban perceptions and whether these associations exhibit stable nonlinear patterns. Shanghai was selected because its dense and heterogeneous urban fabric, extensive street network, and substantial variation across central, newly developed, and peripheral areas provide a suitable context for examining multiple dimensions of streetscape visual clutter at the city scale. Using street-view imagery from Shanghai, the study combines deep learning and computer vision methods to operationalize streetscape visual clutter through four dimensions—edge, color, semantic, and depth clutter—and examines their associations with six model-estimated urban perception dimensions: beauty, wealth, liveliness, safety, boredom, and depressiveness. The study aims not only to develop a multidimensional measurement framework but also to provide diagnostic evidence for municipal and district-level planning authorities, urban designers, landscape architects, and agencies responsible for streetscape renewal, signage, street furniture, planting, and public-space management. Rather than recommending a uniform increase or reduction in visual information, the findings are intended to support more targeted identification and management of different forms of streetscape visual clutter. The resulting evidence is intended to support targeted diagnosis and design review in visually congested settings, including major intersections, mixed-use streets, and peripheral town centers, where buildings, signs, vegetation, vehicles, and public facilities must be coordinated within limited street space.
Specifically, this study addresses three questions and one practical objective: (1) Does streetscape visual clutter remain associated with urban perceptions after controlling for built-environment characteristics? (2) Do different types of visual clutter show distinct associations with different dimensions of urban perception? (3) Are there stable nonlinear relationships between streetscape visual clutter and urban perceptions? With the questions addressed, (4) the study suggests considerations to lessen visual clutter.
The overall research framework is shown in Figure 1.

2. Study Area and Methods

2.1. Study Area and Street-View Data

Shanghai is a high-density megacity characterized by substantial spatial heterogeneity. Its urban landscape includes dense central districts, peripheral new towns, suburban built-up areas, and urban–rural transition zones. These areas differ markedly in building form, street frontage, vegetation, transport conditions, and levels of human activity. Such diversity provides sufficient spatial variation for examining the associations between streetscape visual clutter and urban perceptions [3].
In Shanghai, the potential application of these findings involves multiple municipal, district-level, and local actors. Planning and natural-resources authorities, urban-design and urban-renewal agencies, and housing and urban–rural development departments may use such evidence to inform street-interface design, public-space layout, sightline organization, spatial continuity, and the coordination of buildings with surrounding streetscapes. Landscaping and city-appearance authorities are involved in the management of outdoor signs, landscape lighting, street greenery, and the visual quality of public environments, while transport and road-management authorities, district governments, and subdistrict offices participate in street renewal and the provision and maintenance of public facilities. Property owners, designers, and facility operators also influence facades, signs, lighting, planting, and other visible street elements. The findings may therefore support urban-design guidance, streetscape-renewal schemes, public-space design and review, signage and lighting management, street-furniture coordination, and planting design.
Road-network data were obtained from OpenStreetMap, and sampling points were generated at 100 m intervals along the road network. Based on the coordinates of these points, panoramic street-view images captured in the winter of 2022 were collected from Baidu Maps. A total of 34,622 sampling locations were obtained. To reduce redundancy between views along the road direction, only the left- and right-facing images were retained from each panorama, resulting in 69,244 directional images with a resolution of 512 × 512 pixels. The images were subsequently processed using the PSShadowHighlight function in Python 3.10 to reduce uneven illumination caused by shadows and highlights.
The data were organized at three analytical levels. The 69,244 directional-image observations were used for the primary OLS and nonlinear regression analyses. Measurements from the left- and right-facing images at each sampling location were averaged to generate 34,622 sampling-location records, which were used for principal component analysis, the sampling-location-level sensitivity analysis, and residual spatial autocorrelation diagnostics. For spatial visualization, the sampling-location values were further aggregated into 250 m × 250 m grid cells by calculating the mean value of all valid sampling locations within each cell.

2.2. Measurement of Streetscape Visual Clutter and Estimation of Urban Perception Scores

2.2.1. Measurement of Streetscape Visual Clutter

Streetscape visual clutter was measured from four complementary dimensions: edge, color, semantic, and depth clutter (Table 1). These dimensions represent different forms of visual variation within streetscape images, including structural contours and textures, transitions among dominant colors, fragmentation among semantic object categories, and changes between relative spatial-depth layers.
Edge clutter was used to characterize the density of lines, object contours, textures, and spatial boundaries within the streetscape. Each directional street-view image was first converted to grayscale, after which Canny edge detection was applied using lower and upper thresholds of 0.11 × 255 and 0.27 × 255, respectively. The resulting binary map represented structural changes generated by elements such as building facades, road edges, vegetation, signs, windows, and street facilities.
Color clutter was used to describe the frequency and spatial distribution of transitions among dominant colors. Each image was converted from the RGB color space to the HSV color space, which separates hue, saturation, and brightness information. K-means clustering was then applied with the number of clusters set to five to identify the dominant color categories within each image. Boundaries between adjacent color categories were subsequently extracted to generate a color-boundary map, with denser boundaries indicating more frequent local color changes.
Semantic clutter was used to represent the spatial fragmentation and interaction of different object categories within the streetscape. Semantic segmentation was performed using the publicly available OneFormer Swin-L model pretrained on the MS COCO Panoptic dataset [25]. The pixel-level classes generated by the model were regrouped into streetscape-related categories, including roads, buildings, walls and fences, vegetation, sky, pedestrians, vehicles, and other visible elements. The detailed correspondence between the original model classes and the regrouped categories is provided in Supplementary Table S1. Boundaries between adjacent semantic categories were then extracted to produce semantic-boundary maps. A higher density of such boundaries indicates more frequent alternation and greater fragmentation among different streetscape elements.
Depth clutter was used to characterize the frequency of transitions among relative depth layers and variations in foreground–background organization. Monocular relative-depth maps were generated using the publicly available DPT-Hybrid model [26,27]. The predicted depth values were normalized to the range of 0–255 and divided into four relative depth layers using K-means clustering. Boundaries between adjacent depth layers were then extracted to produce depth-boundary maps. The resulting indicator represents an image-based measure of relative depth transitions rather than a direct measurement of physical enclosure or occlusion.
After the four feature maps had been obtained, they were processed using a consistent procedure. Each map was converted into a binary boundary map and smoothed using a Gaussian kernel of 31 × 31 pixels with a standard deviation of 7. Gaussian smoothing was used to represent the local concentration of visual changes and to reduce isolated pixel-level responses. The mean value of each smoothed map was then calculated as the corresponding image-level clutter score. Higher values indicate a greater overall density of structural, color, semantic, or depth-related boundaries in the streetscape image.
This treatment is conceptually consistent with the spatial characteristics of human vision, in which visual resolution is highest in central vision and gradually decreases toward the periphery. Visual clutter was therefore interpreted as the aggregation of visual changes within a local neighborhood rather than merely as a count of isolated boundary pixels. Because the normalized Gaussian kernel largely preserved the global mean, smoothing did not materially alter the relative scores among images but improved the representation and visualization of clutter patterns.
To assess the degree of overlap among the four clutter indicators, pairwise Pearson correlations were calculated using the 69,244 directional-image observations. The resulting correlation matrix was used to determine whether the indicators were highly redundant or captured related but distinct dimensions of streetscape visual clutter.

2.2.2. Estimation of Urban Perception Scores

Place Pulse-based computational perception modeling has become a well-established approach in large-scale urban research, enabling consistent estimation of subjective streetscape attributes from extensive street-view image collections. Models trained on Place Pulse 2.0 have been applied across multiple urban contexts and spatial scales [2,9,11].
Six urban perception dimensions were estimated using pretrained models developed from the Place Pulse 2.0 dataset. Place Pulse 2.0 contains more than 1.1 million pairwise comparisons of urban images and records public evaluations across six perceptual dimensions: beauty, wealth, liveliness, safety, boredom, and depressiveness [2,9]. These pairwise comparisons provide the training basis for extending subjective urban perception assessment to large-scale street-view imagery.
This study used the publicly available human-perception-place-pulse models released by Ouyang, which have also been applied in the Global Streetscapes project to extract human-perception attributes from street-view images [28]. The models use ViT-B/16 as the backbone and consist of six independent classifiers corresponding to beauty, wealth, liveliness, safety, boredom, and depressiveness. Each directional street-view image was independently processed by the six models, and the predicted probability of the positive class was used to represent the intensity of the corresponding perception. To facilitate interpretation and comparison, the positive-class probability was multiplied by 10 to obtain a perception score ranging from 0 to 10. According to the validation results reported by the model developer, the classification accuracies were 76.9% for beauty, 72.9% for wealth, 77.1% for liveliness, 76.7% for safety, 61.6% for boredom, and 67.2% for depressiveness.

2.3. Statistical Analysis

2.3.1. Construction of the Composite Streetscape Visual Clutter Index

Principal component analysis (PCA) was conducted using the four standardized clutter indicators at the sampling-location level. The first principal component was retained to construct the composite streetscape visual clutter index. This index was used to characterize and visualize the overall spatial distribution of visual clutter across the study area. The PCA diagnostics, eigenvalues, explained variance, and other detailed results are provided in the Supplementary Materials. The composite index was not included in the subsequent regression analyses; instead, edge, color, semantic, and depth clutter were retained as separate explanatory variables to identify their dimension-specific associations with urban perceptions.

2.3.2. Linear Association Analysis

Ordinary least squares (OLS) regression was used to examine the linear associations between streetscape visual clutter and the six model-estimated urban perception scores. Separate models were estimated for beauty, wealth, liveliness, safety, boredom, and depressiveness. In each model, edge, color, semantic, and depth clutter were included simultaneously as the main explanatory variables. The proportions of vegetation, buildings, roads, walls, vehicles, and pedestrians, derived from the semantic segmentation results, were included as built-environment control variables.
These controls were included because visible physical elements may independently influence how streetscapes are perceived. Roads, vehicles, and pedestrian presence can indicate accessibility, mobility, and street activity and may therefore contribute to perceived liveliness. Vegetation is commonly associated with naturalness and environmental quality and can enhance perceptions such as beauty and wealth. By contrast, walls may interrupt visual continuity, reduce openness, and negatively affect safety and other environmental evaluations. Building proportion can also reflect development intensity, spatial enclosure, and the continuity of street frontages. Controlling for these elements allowed the associations of visual clutter with urban perceptions to be estimated beyond the effects of streetscape composition itself [9,11,29].
The regression model was specified as follows:
Y i = β 0 + β 1 E d g e i + β 2 C o l o r i + β 3 S e m a n t i c i + β 4 D e p t h i + j = 1 6 γ j C j i + ε i
where ( Y i ) denotes the score of observation ( i ) for a specific urban perception dimension. ( E d g e i ), ( C o l o r i ), ( S e m a n t i c i ), and ( D e p t h i ) represent edge, color, semantic, and depth clutter, respectively. ( C j i ) denotes the control variables, including vegetation, roads, walls, buildings, vehicles, and pedestrians. ( β 0 ) is the intercept; ( β 1 ) to ( β 4 ) are the regression coefficients of the four core independent variables; ( γ j ) denotes the coefficient of control variable ( j ); and ( ε i ) is the random error term. Separate regression models were estimated for each of the six dependent variables.
Before model estimation, infinite values were treated as missing, and observations with missing values in any model variable were excluded. Variables were transformed where necessary to reduce distributional skewness and were subsequently standardized using Z-scores, allowing the magnitudes of the regression coefficients to be compared across indicators. HC3 heteroskedasticity-consistent standard errors were used for statistical inference. Multicollinearity was assessed using the variance inflation factor (VIF), and all statistical tests were two-sided, with p   <   0.05 regarded as statistically significant. The detailed VIF results are provided in Supplementary Table S2.
To assess the robustness of the primary OLS results and diagnose potential spatial dependence, two complementary analyses were conducted at the sampling-location level. First, the four clutter indicators, six model-estimated urban perception scores, and six built-environment control variables derived from the left- and right-facing images were averaged within each of the 34,622 sampling locations. The OLS models were then re-estimated using the same variable transformations, standardization procedures, control variables, and HC3 heteroskedasticity-robust standard errors as in the primary analysis. Second, residual spatial autocorrelation was assessed for the six clutter-only and six controlled models using Global Moran’s I . A row-standardized eight-nearest-neighbor spatial weights matrix ( k = 8 ) was constructed for all 34,622 sampling locations. The mean distance to the eight nearest neighbors was 272.45 m, and statistical significance was assessed using 499 permutations. Detailed results are provided in Supplementary Tables S3 and S4.

2.3.3. Nonlinear Association Analysis

In addition to the OLS models, nonlinear analyses were conducted to determine whether the associations between visual clutter and urban perceptions changed across different levels of clutter. The analyses covered all 24 combinations of the four clutter indicators and six perception dimensions. The same directional-image dataset and built-environment controls used in the linear models were retained to maintain comparability between the linear and nonlinear results.
Three complementary methods were applied. First, locally weighted scatterplot smoothing (LOWESS) was used to examine the overall functional form of each relationship without imposing a predetermined curve. Second, segmented regression was used to estimate a potential breakpoint and the slopes below and above that breakpoint. Third, quadratic regression was used to test whether the relationship showed statistically significant curvature and to determine the direction of the quadratic term. Applying the three methods jointly reduced the risk of classifying a relationship as nonlinear on the basis of a single model or statistical test.
To account for multiple testing across the 24 relationships, statistical significance was adjusted using the false discovery rate (FDR). A relationship was classified as a stable inverted U-shape only when LOWESS indicated an initial increase followed by a decrease, segmented regression identified the same directional change around a statistically supported breakpoint, the quadratic term was significantly negative after FDR correction, and the overall pattern remained stable in the sensitivity analyses. Detailed model specifications, breakpoint estimation procedures, and sensitivity tests are reported in the Supplementary Materials.

3. Results

3.1. Streetscape Visual Clutter Features in Shanghai

Using the street-view imagery, we derived four dimensions of streetscape visual clutter across Shanghai: edge, color, semantic, and depth clutter. A composite clutter indicator was also constructed from the first principal component of the PCA.
Pairwise Pearson correlations were examined to assess the degree of overlap among the four dimension-specific indicators. As shown in Figure 2, the correlation coefficients ranged from 0.079 to 0.463. The strongest correlation was observed between color clutter and edge clutter ( r = 0.463 ), whereas semantic clutter showed only weak correlations with color clutter and edge clutter (both r = 0.079 ). The correlations between depth clutter and the other three indicators ranged from 0.234 to 0.298. Overall, no pair of indicators showed a high correlation, indicating that the four measures were related but not highly redundant and captured distinct dimensions of streetscape visual clutter.
To further examine their spatial patterns, the point-level clutter values were mapped onto Shanghai’s street network and aggregated into 250 m × 250 m grid cells. The resulting values were classified using the Jenks natural breaks method. This method determines class boundaries based on clusters inherent in the data, allowing the spatial variation in each clutter indicator to be displayed more clearly.
Figure 3 presents the spatial distributions of the streetscape visual clutter indicators at the 250 m grid-cell level. The grid-cell values were classified using the Jenks natural breaks method for visualization.
Edge clutter exhibits a clear concentration in central Shanghai. High-value areas are distributed across Huangpu, Jing’an, Hongkou, Yangpu, Putuo, Changning, and Xuhui, forming a relatively dense pattern within the central street network. Outside the urban core, localized high-value clusters are also present around built-up nodes in Minhang, Songjiang, Jiading, northern Jinshan, and Pudong. Some elevated values extend along selected road corridors, particularly in the southwestern and southern parts of the study area. Lower values are more common in western Qingpu and parts of eastern and southern Pudong.
Color clutter shows a broadly similar center–periphery pattern but has a more continuous high-value concentration in the urban core. Elevated values extend across the central districts, while peripheral high-value areas are generally smaller and more fragmented. These localized clusters occur mainly around Minhang, Songjiang, northern Jinshan, Fengxian, and parts of Pudong. By contrast, lower values are more widely distributed in Qingpu, peripheral Jiading, and eastern Pudong. Compared with edge clutter, color clutter displays a clearer contrast between the continuous central cluster and the discontinuous peripheral distribution.
Semantic clutter is also relatively high in central Shanghai, although its high-value distribution extends more widely beyond the urban core than that of color clutter. In addition to the central districts, elevated values occur around peripheral built-up nodes, major road intersections, and town centers across Shanghai, particularly in Minhang, Songjiang, Jinshan, Fengxian, and Pudong. Outside these areas, high values generally appear as scattered clusters or short linear segments rather than continuous zones. Lower values are concentrated mainly in western Qingpu, eastern Pudong, and the peripheral parts of Jiading and Fengxian.
Depth clutter exhibits a markedly different spatial pattern. High-value cells are relatively sparse and widely dispersed, with no extensive continuous concentration in the central city. Most central and peripheral roads are dominated by low or medium–low values, while localized high-value segments are scattered across Huangpu, Jing’an, Yangpu, Minhang, Songjiang, Jinshan, Fengxian, and Pudong. Compared with edge, color, and semantic clutter, depth clutter shows weaker central concentration and a more fragmented distribution across both central and peripheral areas.
Composite streetscape visual clutter is generally higher in the urban core, where high and medium-high values form an interconnected pattern across the street network. Outside the central area, elevated values occur around several peripheral built-up nodes and selected road corridors in Baoshan, Minhang, Songjiang, Jinshan, Fengxian, and Pudong. Lower values are more common in western Qingpu, peripheral Jiading, eastern and southern Pudong, and parts of the southern urban fringe. Overall, the composite index shows a spatial pattern characterized by central concentration, localized peripheral clustering, and corridor-like extensions rather than a simple continuous decline from the urban center to the periphery. Edge, color, and semantic clutter share this general pattern but differ in spatial continuity and extent, whereas depth clutter remains substantially more dispersed and fragmented.

3.2. Model-Estimated Urban Perceptions in Shanghai

This study further mapped six model-estimated urban perception scores across Shanghai’s street network: beauty, wealth, liveliness, safety, boredom, and depressiveness. Figure 4 presents the spatial distributions of these perception scores. All six dimensions show clear spatial variation across the city.
Figure 4 presents the spatial distributions of the six model-estimated urban perception dimensions across Shanghai. The grid-cell values were classified using the Jenks natural breaks method for visualization.
Beauty scores are generally higher in central Shanghai, with additional high-value clusters around several peripheral built-up nodes and selected road corridors in Minhang, Songjiang, Jinshan, Fengxian, and Pudong. Lower values are more common in western Qingpu, eastern and southern Pudong, and parts of the southern urban fringe.
Wealth shows a strong and relatively continuous high-value concentration across the central urban area. Smaller high-value clusters also occur around peripheral nodes in Minhang, Songjiang, Baoshan, Jiading, Fengxian, and Pudong, whereas lower values are concentrated mainly in western Qingpu, eastern Pudong, and southern Fengxian.
Liveliness exhibits a spatial pattern broadly similar to that of wealth, with a clear central concentration and smaller, more fragmented high-value clusters in peripheral areas. Among the positive perception dimensions, wealth and liveliness show the strongest center–periphery contrast.
Safety is also generally higher in central Shanghai, but its high-value distribution is less continuous than those of wealth and liveliness. Elevated values occur in both the urban core and several peripheral nodes, while lower values are more widely distributed in Qingpu, peripheral Jiading, eastern Pudong, and southern Fengxian.
Boredom displays a broadly inverse pattern, with lower values concentrated in the central districts and higher values distributed across much of the urban periphery. Some high-value areas also extend along road corridors and around peripheral built-up nodes.
Depressiveness shows a similar center–periphery contrast, but its high-value areas are more fragmented and dispersed than those of boredom. Overall, positive perceptions are generally higher in the urban core, whereas boredom and depressiveness tend to be higher toward the periphery. Wealth and liveliness form the most continuous central clusters, while beauty and safety show more dispersed peripheral high-value nodes.

3.3. The Relationship Between Streetscape Clutter Features and Urban Perceptions

3.3.1. Linear Relationships

To examine the relationships between different types of streetscape visual clutter and model-estimated urban perception scores, this study estimated the associations of edge, color, semantic, and depth clutter with six perception dimensions: beauty, wealth, liveliness, safety, boredom, and depressiveness. The models controlled for built-environment characteristics, including vegetation, buildings, roads, walls, pedestrians, and vehicles.
Table 2 presents the OLS regression results for the associations between the four streetscape visual clutter indicators and the six urban perception dimensions. Background shading indicates the direction and statistical significance of the associations: red represents a significant positive association, blue represents a significant negative association, and gray indicates a nonsignificant association.
Overall, the streetscape visual clutter indicators were significantly associated with most urban perception dimensions. Among the 24 associations between the four clutter indicators and the six perception dimensions, 23 reached statistical significance. The only nonsignificant result was the association between depth clutter and perceived safety. These findings indicate that streetscape visual clutter is not a uniformly negative visual characteristic. Instead, different forms of clutter show distinct directions and strengths of association with urban perceptions.
Edge and color clutter showed broadly consistent associations with the six urban perception dimensions. Edge clutter was positively associated with beauty, liveliness, safety, and wealth, with t-values of 31.465, 30.918, 26.424, and 17.292, respectively. It was negatively associated with boredom and depressiveness, with t-values of −2.014 and −7.053.
Color clutter showed a similar pattern. It was positively associated with beauty, liveliness, safety, and wealth, with t-values of 24.565, 16.479, 13.985, and 14.155, respectively. It was negatively associated with boredom and depressiveness, with t-values of −14.683 and −3.783.
Semantic clutter showed a more differentiated pattern of association. It was positively associated with liveliness, safety, and wealth, with t-values of 31.788, 10.042, and 38.506, respectively. It was negatively associated with boredom and depressiveness, with t-values of −23.165 and −14.044. However, semantic clutter was also negatively associated with beauty ( t = 10.476 ).
Depth clutter displayed a pattern distinct from those of the other three clutter indicators. It was negatively associated with beauty, liveliness, and wealth, with t-values of −15.165, −6.214, and −20.181, respectively. By contrast, it was positively associated with boredom and depressiveness, with t-values of 7.464 and 10.946. Its association with safety was not statistically significant ( t = 1.365 ).
A comparison of the absolute t-values further identifies the clutter indicator most strongly associated with each perception dimension. Edge clutter had the largest absolute t-values for beauty and safety, reaching 31.465 and 26.424, respectively. Semantic clutter had the largest absolute t-values for boredom, depressiveness, liveliness, and wealth, at −23.165, −14.044, 31.788, and 38.506, respectively. These results indicate that edge clutter is more strongly associated with evaluations of beauty and safety, whereas semantic clutter is more strongly associated with boredom, depressiveness, liveliness, and wealth.
The sampling-location-level estimates were broadly consistent with those from the primary directional-image-level models. The directions of all 24 associations between the four clutter indicators and six urban perception dimensions remained unchanged, and 23 of the 24 associations were statistically significant at p < 0.05 . Relative to the primary models, the weak negative association between edge clutter and boredom was no longer statistically significant ( t = 1.302 , 95% CI [−0.019, 0.004]), whereas the positive association between depth clutter and safety became statistically significant ( t = 3.398 , 95% CI [0.008, 0.028]). All other associations retained their significance status. Overall, the main linear patterns were robust to aggregation at the sampling-location level (Supplementary Table S3).
Residual spatial autocorrelation diagnostics showed that all 12 sets of model residuals exhibited significant positive spatial autocorrelation (two-sided permutation p = 0.004 ). Among the controlled models, Moran’s I ranged from 0.110 for depressiveness to 0.211 for wealth and was lower than that of the corresponding clutter-only model for all six perception dimensions. These results indicate that the built-environment controls reduced but did not eliminate residual spatial dependence. Therefore, although the sampling-location-level analysis supported the stability of the main coefficient patterns, statistical inference from the OLS models should be interpreted cautiously because residual spatial dependence remained (Supplementary Table S4).
Regarding model fit, explanatory power varied across perception dimensions. After built-environment controls were included, adjusted R 2 increased for all six models, ranging from 0.074 for depressiveness to 0.351 for liveliness. Although the models explained only part of the variation in urban perceptions, particularly for depressiveness and boredom, the broadly consistent directions and statistical support of the clutter coefficients indicate that streetscape visual clutter remained systematically associated with urban perceptions. These results therefore suggest stable associations, while the explanatory contribution varied across perception dimensions (Supplementary Table S5).

3.3.2. Nonlinear Relationships

As shown in Table 3, four of the 24 relationships met all criteria for a strict inverted U-shape: edge clutter–boredom, color clutter–depressiveness, depth clutter–wealth, and semantic clutter–wealth. For all four relationships, the slopes on both sides of the breakpoint, the hinge term in the segmented regression, and the quadratic term remained statistically significant after FDR correction (all adjusted ( p < 0.001 ) ). The patterns identified by LOWESS and segmented regression also remained stable across the different sample ranges examined in the sensitivity analyses.
A significant inverted U-shaped relationship was identified between edge clutter and boredom. The segmented regression estimated a breakpoint of 0.486. The slope below the breakpoint was significantly positive ( ( β = 0.026 ), adjusted ( p < 0.001 )), whereas the slope above the breakpoint was significantly negative ( ( β = 0.128 ) , adjusted ( p < 0.001 ) ). The hinge term also remained significant after FDR correction (adjusted ( p < 0.001 ) ). Consistent with these results, the coefficient of the quadratic term was significantly negative (( β = 0.024 ) , adjusted ( p < 0.001 ) ). Thus, boredom initially increased with edge clutter but began to decrease after edge clutter exceeded the estimated breakpoint.
Color clutter showed a significant inverted U-shaped relationship with depressiveness. The segmented regression estimated a breakpoint of −0.264. The slope below the breakpoint was significantly positive ( β = 0.100 ) , ( p < 0.001 ) , whereas the slope above the breakpoint was significantly negative ( β = 0.097 ) , ( p < 0.001 ) . The hinge term remained significant after FDR correction ( p < 0.001 ). The quadratic term was also significantly negative ( β = 0.054 ) , ( p < 0.001 ) . Thus, depressiveness increased with color clutter below the breakpoint and decreased once color clutter exceeded the breakpoint.
Depth clutter also showed a significant inverted U-shaped relationship with wealth. The estimated breakpoint was −0.558. The slope below the breakpoint was significantly positive ( β = 0.043 ) , ( p < 0.001 ) , whereas the slope above the breakpoint was significantly negative ( β = 0.138 ) , ( p < 0.001 ) . The hinge term remained significant after FDR correction ( p < 0.001 ), and the quadratic term was significantly negative ( β = 0.030 ) , ( p < 0.001 ) . Wealth therefore increased with depth clutter below the breakpoint but declined after the breakpoint was exceeded.
A significant inverted U-shaped relationship was likewise identified between semantic clutter and wealth. The segmented regression estimated a breakpoint of 1.165. The slope below the breakpoint was significantly positive ( β = 0.184 ) , ( p < 0.001 ) , whereas the slope above the breakpoint was significantly negative ( β = 0.162 ) , ( p < 0.001 ) . The hinge term remained significant after FDR correction ( p < 0.001 ), and the quadratic term was significantly negative ( β = 0.021 ) , ( p < 0.001 ) . Among the four strict inverted U-shaped relationships, the semantic clutter–wealth relationship showed relatively large absolute slopes on both sides of the breakpoint.
The remaining 20 relationships between visual clutter and urban perception did not satisfy all criteria for a strict inverted U-shape. Although some relationships showed statistically significant hinge terms in the segmented regressions or significant quadratic terms, the LOWESS curve, the directions and significance of the slopes on either side of the breakpoint, the sign of the quadratic term, or the sensitivity-analysis results were not fully consistent. These relationships were therefore not classified as stable inverted U-shaped relationships.

4. Discussion

This study measured streetscape visual clutter across four dimensions—edge, color, semantic, and depth clutter—and examined their associations with six model-estimated urban perception scores. After controlling for built-environment characteristics, including vegetation, buildings, roads, walls, vehicles, and pedestrians, the four clutter dimensions remained associated with most perception outcomes. However, both the directions and magnitudes of these associations varied across clutter types and perception dimensions. The variation in explanatory power across perception dimensions further suggests that visual clutter represents one explanatory factor rather than a comprehensive account of perceptual variation. Streetscape visual clutter should therefore not be treated as a single or inherently negative environmental attribute. Its perceptual implications depend on the type of visual information involved and how that information is organized within the streetscape.
Methodologically, this study decomposed streetscape visual clutter into four distinct dimensions rather than representing it with a single overall measure. Compared with aggregate indicators such as image entropy, fractal dimension, the Shannon diversity index, or composite complexity scores, this framework provides greater insight into the specific sources of visual variation. The low-to-moderate correlations among the four indicators, together with their distinct spatial distributions and differentiated associations with urban perceptions, suggest that they capture related but non-redundant aspects of streetscape visual clutter. By distinguishing variations in structural boundaries, color composition, semantic elements, and relative depth transitions, the proposed framework improves the interpretability of streetscape visual measurement. It also provides a basis for linking visual clutter to specific planning and design elements, including building interfaces, urban color, visible functional elements, and foreground–background organization.
At the city scale, the spatial distributions of the perception scores further show that the implications of visual clutter depend on both the broader urban context and the specific type of street space. The continuous concentrations of model-estimated wealth and liveliness in central Shanghai may be related to active ground-floor frontages, commercial and service functions, pedestrian activity, architectural articulation, and dense but recognizable visual cues. By contrast, higher boredom and depressiveness in many peripheral areas may be associated with inactive or discontinuous street interfaces, fewer visible activities and facilities, and fragmented foreground–background relationships. These spatial differences indicate that the perceptual implications of visual clutter vary across urban contexts and street-space types.
Accordingly, the associations between streetscape visual clutter and model-estimated urban perceptions require dimension-specific interpretation. Edge and color clutter were generally positively associated with beauty, liveliness, safety, and wealth, and negatively associated with the two negative perception dimensions. Although both indicators capture changes in visual boundaries, they also reflect aspects of visual richness, such as architectural detail, vegetation contours, material variation, and color contrast. Their positive associations may therefore reflect the contribution of environmental detail and visual stimulation to lower perceived monotony and greater streetscape legibility and attractiveness [30,31,32].
Semantic clutter showed a more differentiated pattern across perception dimensions. It was positively associated with liveliness, wealth, and safety, but negatively associated with beauty. Diverse building frontages, commercial facilities, pedestrians, and vehicles can convey functional diversity, social activity, and economic vitality, thereby strengthening perceptions of liveliness and wealth [33,34]. However, the frequent intermingling of different object categories may also disrupt compositional coherence and the continuity of street interfaces, reducing aesthetic harmony. This finding suggests that the same type of visual information may support functional evaluations while weakening aesthetic evaluations.
Depth clutter displayed a markedly different pattern from the other three indicators. It was negatively associated with beauty, liveliness, and wealth, and positively associated with boredom and depressiveness. Higher depth clutter reflects more frequent changes in relative depth structure and stronger variation in foreground–background organization, which may make spatial boundaries, road orientation, and relative depth relationships more difficult to interpret quickly. Research in visual psychology has shown that cluttered scenes can slow visual search, increase cognitive demands, and impair object recognition and situational awareness [17,35,36,37]. Depth clutter may therefore correspond more closely than edge, color, or semantic clutter to the concepts of information competition and cognitive load in visual psychology.
The nonlinear analyses identified stable inverted U-shaped patterns in only four of the 24 clutter–perception relationships. These patterns should therefore be regarded as relationship-specific supplementary findings rather than evidence of a universal nonlinear pattern or a generally optimal intermediate level of visual clutter.
Both depth and semantic clutter showed inverted U-shaped relationships with perceived wealth. At relatively low levels, limited variation in relative depth structure and semantic composition may correspond to insufficient visible facilities, low functional diversity, or weak street activity. As spatial layering, visible facilities, and semantic diversity increase, streetscapes may convey stronger cues of functional diversity, developed infrastructure, and economic activity, thereby strengthening model-estimated wealth scores. Beyond a certain level, however, greater occlusion, object intermingling, and frequent interface transitions may cause visual richness to be perceived as spatial congestion, visual disorder, or declining environmental quality, which may weaken perceived wealth.
Inverted U-shaped relationships were also found between edge clutter and boredom and between color clutter and depressiveness. When edge or color clutter increases from a low to a moderate level, the additional visual information may consist of scattered, discontinuous, or weakly organized local variations. Such changes may increase perceptions of fragmentation, boredom, or depressiveness. At higher levels, however, increased edge and color information may reflect richer architectural details, vegetation, commercial frontages, and street activity. These elements may provide stronger place cues and visual stimulation, thereby reducing negative perceptions.

5. Limitations

This study has several limitations. First, the cross-sectional research design identifies statistical associations rather than causal effects. The mechanisms proposed in the Discussion should therefore be interpreted as plausible explanations of the observed patterns rather than as confirmed causal processes.
Second, the six perception scores were estimated using pretrained Place Pulse-based models rather than obtained directly from surveys of Shanghai residents. Although this approach enables city-scale assessment, the resulting scores may not fully capture locally specific cultural, demographic, and contextual differences in environmental evaluation. Moreover, the clutter indicators, semantic control variables, and perception scores were all derived from the same street-view images using computer vision or deep learning procedures. This shared image source and partial algorithmic coupling may introduce correlated measurement errors or amplify visual characteristics that are simultaneously salient to multiple models.
Third, the primary regression analyses used 69,244 directional images, including paired left- and right-facing views from 34,622 sampling locations distributed at 100 m intervals along the road network. Although the sampling-location-level sensitivity analysis reduced the dependence arising from paired views at the same location, the Global Moran’s I diagnostics indicated significant positive residual spatial autocorrelation in all six controlled models. The inclusion of built-environment controls reduced but did not eliminate this spatial dependence, suggesting that neighboring sampling locations may share unobserved spatial characteristics not fully captured by the current explanatory variables. Consequently, the statistical significance of the OLS estimates should be interpreted with caution.
Finally, depth clutter should be interpreted with caution. As operationalized in this study, it represents image-based variations in relative depth structure and foreground–background organization rather than a direct measurement of physical occlusion, spatial enclosure, or three-dimensional spatial hierarchy. Its physical interpretation should therefore remain limited to relative depth transitions observed in street-view images.

6. Conclusions

Using street-view imagery from Shanghai, this study developed a multidimensional framework for measuring streetscape visual clutter through four indicators: edge, color, semantic, and depth clutter. The framework distinguishes different sources of visual variation rather than relying solely on a single aggregate measure and was used to examine associations with six model-estimated urban perception scores: beauty, wealth, liveliness, safety, boredom, and depressiveness.
The results revealed differentiated associations across clutter types and perception dimensions. Edge and color clutter were generally associated with more favorable perception scores, whereas semantic clutter showed contrasting associations across functional and aesthetic evaluations. Depth clutter exhibited a distinct pattern and was more frequently associated with less favorable perception scores. These findings indicate that streetscape visual clutter is neither a uniform nor inherently negative environmental attribute; its perceptual implications depend on the type and organization of visual information.
Stable inverted U-shaped patterns were identified in only four of the 24 clutter–perception relationships. These patterns should therefore be interpreted as relationship-specific supplementary findings rather than evidence of a universal nonlinear rule or a generally preferable intermediate level of visual clutter.

7. Practical Application and Future Research

To facilitate practical application, the four clutter dimensions can be translated into distinct planning and design considerations. High edge clutter may reflect either beneficial architectural articulation and vegetation contours or excessive accumulations of small projections, overlapping details, and fragmented street interfaces. Beneficial details may be retained, whereas unnecessary fragmentation can be reduced through clearer facade organization and the consolidation of minor visual elements. Color clutter should be managed by coordinating the hue, saturation, scale, and placement of facades, signs, lighting, and street facilities rather than by imposing visual uniformity. For semantic clutter, functional diversity and active street uses should be preserved, while signs, benches, waste bins, utility boxes, parked vehicles, and temporary facilities may be grouped, aligned, or incorporated into designated frontage and street-furniture zones. Depth clutter requires particular attention to foreground obstruction and spatial continuity. Decorative elements, signs, utilities, temporary installations, and dense planting should be kept outside pedestrian routes, major sightlines, intersections, crossings, and other visibility-critical locations.
These dimension-specific considerations can be applied differently across spatial settings. At intersections and road crossings, interventions may prioritize sightline continuity, crossing visibility, the consolidation of redundant signs, and the management of curbside parking and temporary installations. Along mixed-use and commercial streets, active frontages and functional diversity may be retained while signage, facade lighting, awnings, street furniture, and service facilities are organized within a clearer visual hierarchy. In peripheral town centers and road corridors with lower model-estimated wealth or liveliness scores, improvement should focus on continuous and well-maintained street interfaces, visible public and community facilities, active ground-floor uses, connected pedestrian spaces, vegetation, lighting, and coordinated material and color variation rather than further visual simplification. Citywide clutter and perception maps may support initial screening, followed by field audits and site-specific design review by planning, urban renewal, landscaping, transport, and district-level authorities. Because the present study identifies statistical associations rather than causal effects, the resulting recommendations should be regarded as evidence-informed planning priorities and tested through context-specific pilot interventions and before-and-after evaluations before wider implementation.
Future research should further test and refine these planning considerations. Multi-temporal street-view imagery, natural experiments, quasi-experimental designs, and before-and-after studies could be used to examine whether changes in specific clutter dimensions lead to corresponding changes in urban perceptions. Model-estimated perception scores should also be validated against ratings from local residents and supplemented with field surveys, behavioral observations, eye-tracking experiments, traffic-safety records, GIS data, and manually verified environmental indicators. Spatial econometric models, network-based spatial weights, spatially clustered standard errors, and grid-level analyses could be applied to address residual spatial dependence more directly. In addition, LiDAR, three-dimensional point clouds, stereo imagery, and independently measured urban-form indicators could help validate the physical meaning of depth clutter. Comparative studies across cities, seasons, and cultural contexts would further assess the generalizability of the proposed framework and identify where context-specific adaptation is required.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/land15081402/s1, Table S1: Mapping of COCO panoptic categories to the custom streetscape semantic taxonomy; Table S2: Variance inflation factors for the explanatory and control variables; Table S3: Sampling-location-level OLS regression results for streetscape visual clutter and model-estimated urban perception scores; Table S4: Global Moran’s I statistics for the residuals of the sampling-location-level OLS models; Table S5: Model fit before and after the inclusion of built-environment controls and standardized clutter coefficients from the controlled models.

Author Contributions

Conceptualization, F.L., J.S. and Y.W.; methodology, F.L., J.S. and Y.W.; software, F.L.; validation, F.L. and J.S.; formal analysis, F.L.; investigation, F.L.; data curation, F.L.; writing—original draft preparation, F.L.; writing—review and editing, F.L., X.L., J.S. and Y.W.; visualization, F.L.; supervision, J.S. and Y.W.; project administration, J.S. and Y.W.; funding acquisition, J.S. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the National Natural Science Foundation of China (General Program, No. 52578091).

Data Availability Statement

The original contributions presented in the study are included in the article, further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Salesses, P.; Schechtner, K.; Hidalgo, C.A. The Collaborative Image of the City: Mapping the Inequality of Urban Perception. PLoS ONE 2013, 8, e68400. [Google Scholar] [CrossRef] [PubMed]
  2. Dubey, A.; Naik, N.; Parikh, D.; Raskar, R.; Hidalgo, C.A. Deep Learning the City: Quantifying Urban Perception at a Global Scale. In European Conference on Computer Vision, Part I; Leibe, B., Matas, J., Sebe, N., Welling, M., Eds.; Springer International Publishing Ag: Cham, Switzerland, 2016; Volume 9905, pp. 196–212. [Google Scholar]
  3. Wei, J.; Yue, W.; Li, M.; Gao, J. Mapping Human Perception of Urban Landscape from Street-View Images: A Deep-Learning Approach. Int. J. Appl. Earth Obs. Geoinf. 2022, 112, 102886. [Google Scholar] [CrossRef]
  4. Won, J.; Lee, C.; Forjuoh, S.N.; Ory, M.G. Neighborhood Safety Factors Associated with Older Adults’ Health-Related Outcomes: A Systematic Literature Review. Soc. Sci. Med. 2016, 165, 177–186. [Google Scholar] [CrossRef] [PubMed]
  5. Gong, Y.; Palmer, S.; Gallacher, J.; Marsden, T.; Fone, D. A Systematic Review of the Relationship between Objective Measurements of the Urban Environment and Psychological Distress. Environ. Int. 2016, 96, 48–57. [Google Scholar] [CrossRef] [PubMed]
  6. Lynch, K. The Image of the City; MIT Press: Cambridge, MA, USA, 1964; ISBN 978-0-262-62001-7. [Google Scholar]
  7. Ewing, R.; Handy, S. Measuring the Unmeasurable: Urban Design Qualities Related to Walkability. J. Urban Des. 2009, 14, 65–84. [Google Scholar] [CrossRef]
  8. Nasar, J. Urban Design Aesthetics—The Evaluative Qualities of Building Exteriors. Environ. Behav. 1994, 26, 377–401. [Google Scholar] [CrossRef]
  9. Zhang, F.; Zhou, B.; Liu, L.; Liu, Y.; Fung, H.H.; Lin, H.; Ratti, C. Measuring Human Perceptions of a Large-Scale Urban Region Using Machine Learning. Landsc. Urban Plan. 2018, 180, 148–160. [Google Scholar] [CrossRef]
  10. Liang, X.; Chang, J.H.; Gao, S.; Zhao, T.; Biljecki, F. Evaluating Human Perception of Building Exteriors Using Street View Imagery. Build. Environ. 2024, 263, 111875. [Google Scholar] [CrossRef]
  11. Song, M.; Xiao, Y. Does Streetscape Color Matter for Urban Perceptions? A Deep Learning Approach to Street View Images. Land Use Policy 2025, 155, 107581. [Google Scholar] [CrossRef]
  12. Ma, L.; Guo, Z.; Lu, M.; He, S.; Wang, M. Developing an Urban Streetscape Indexing Based on Visual Complexity and Self-Organizing Map. Build. Environ. 2023, 242, 110549. [Google Scholar] [CrossRef]
  13. Weber, R.; Schnier, J.; Jacobsen, T. Aesthetics of Streetscapes: Influence of Fundamental Properties on Aesthetic Judgments of Urban Space. Percept. Mot. Ski. 2008, 106, 128–146. [Google Scholar] [CrossRef] [PubMed]
  14. Rosenholtz, R.; Li, Y.; Nakano, L. Measuring Visual Clutter. J. Vis. 2007, 7, 17. [Google Scholar] [CrossRef] [PubMed]
  15. Kaplan, S. The Restorative Benefits of Nature: Toward an Integrative Framework. J. Environ. Psychol. 1995, 15, 169–182. [Google Scholar] [CrossRef]
  16. Moacdieh, N.; Sarter, N. Display Clutter: A Review of Definitions and Measurement Techniques. Hum. Factors 2015, 57, 61–100. [Google Scholar] [CrossRef] [PubMed]
  17. Kim, S.-H.; Kaber, D.B. Assessing the Effects of Conformal Terrain Features in Advanced Head-up Displays on Pilot Performance. Proc. Hum. Factors Ergon. Soc. Annu. Meet. 2009, 53, 36–40. [Google Scholar] [CrossRef]
  18. Westerbeek, H.G.W.; Maes, A. Referential Scope and Visual Clutter in Navigation Tasks. In Bridging the Gap Between Computational, Empirical, & Theoretical Approaches to Reference (PRE-CogSci); Tilburg Centre for Cognition and Communication, Tilburg University: Tilburg, The Netherlands, 2011. [Google Scholar]
  19. Akalin, A.; Yildirim, K.; Wilson, C.; Kilicoglu, O. Architecture and engineering students’ evaluations of house facades: Preference, complexity and impressiveness. J. Environ. Psychol. 2009, 29, 124–132. [Google Scholar] [CrossRef]
  20. Kawshalya, L.W.G.; Weerasinghe, U.G.D.; Chandrasekara, D.P. The Impact of Visual Complexity on Perceived Safety and Comfort of the Users: A Study on Urban Streetscape of Sri Lanka. PLoS ONE 2022, 17, e0272074. [Google Scholar] [CrossRef] [PubMed]
  21. Yu, M.; Zheng, X.; Qin, P.; Cui, W.; Ji, Q. Urban Color Perception and Sentiment Analysis Based on Deep Learning and Street View Big Data. Appl. Sci. 2024, 14, 9521. [Google Scholar] [CrossRef]
  22. Chen, S.; Biljecki, F. Automatic Assessment of Public Open Spaces Using Street View Imagery. Cities 2023, 137, 104329. [Google Scholar] [CrossRef]
  23. Zhao, J.; Suo, W. Research on the Construction and Application of a SVM-Based Quantification Model for Streetscape Visual Complexity. Land 2024, 13, 1953. [Google Scholar] [CrossRef]
  24. Porzi, L.; Rota Bulò, S.; Lepri, B.; Ricci, E. Predicting and Understanding Urban Perception with Convolutional Neural Networks. In Proceedings of the 23rd ACM International Conference on Multimedia; ACM: Brisbane, Australia, 2015; pp. 139–148. [Google Scholar]
  25. Jain, J.; Li, J.; Chiu, M.T.; Hassani, A.; Orlov, N.; Shi, H. OneFormer: One Transformer to Rule Universal Image Segmentation. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); Association for Computing Machinery: New York, NY, USA, 2023; pp. 2989–2998. [Google Scholar]
  26. Ranftl, R.; Bochkovskiy, A.; Koltun, V. Vision Transformers for Dense Prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2021; pp. 12159–12168. [Google Scholar] [CrossRef]
  27. Intel Labs. MiDaS: Code for Robust Monocular Depth Estimation, Version 3.0. Available online: https://github.com/isl-org/MiDaS/tree/v3 (accessed on 2 July 2026).
  28. Hou, Y.; Quintana, M.; Khomiakov, M.; Yap, W.; Ouyang, J.; Ito, K.; Wang, Z.; Zhao, T.; Biljecki, F. Global Streetscapes—A Comprehensive Dataset of 10 Million Street-Level Images across 688 Cities for Urban Science and Analytics. ISPRS J. Photogramm. Remote Sens. 2024, 215, 216–238. [Google Scholar] [CrossRef]
  29. Quercia, D.; O’Hare, N.; Cramer, H. Aesthetic capital: What makes London look beautiful, quiet, and happy? In Proceedings of the 17th ACM Conference on Computer Supported Cooperative Work & Social Computing; Association for Computing Machinery: New York, NY, USA, 2014; pp. 945–955. [Google Scholar]
  30. Kardan, O.; Demiralp, E.; Hout, M.C.; Hunter, M.R.; Karimi, H.; Hanayik, T.; Yourganov, G.; Jonides, J.; Berman, M.G. Is the Preference of Natural versus Man-Made Scenes Driven by Bottom–up Processing of the Visual Features of Nature? Front. Psychol. 2015, 6, 471. [Google Scholar] [CrossRef] [PubMed]
  31. Heath, T.; Smith, S.G.; Lim, B. Tall Buildings and the Urban Skyline: The Effect of Visual Complexity on Preferences. Environ. Behav. 2000, 32, 541–556. [Google Scholar] [CrossRef]
  32. Zhou, H.; He, S.; Cai, Y.; Wang, M.; Su, S. Social Inequalities in Neighborhood Visual Walkability: Using Street View Imagery and Deep Learning Technologies to Facilitate Healthy City Planning. Sustain. Cities Soc. 2019, 50, 101605. [Google Scholar] [CrossRef]
  33. Pazhouhanfar, M.; Mustafa Kamal, M.S. Effect of Predictors of Visual Preference as Characteristics of Urban Natural Landscapes in Increasing Perceived Restorative Potential. Urban For. Urban Green. 2014, 13, 145–151. [Google Scholar] [CrossRef]
  34. Xiao, Y.; Song, M. How Are Urban Design Qualities Associated with Perceived Walkability? An AI Approach Using Street View Images and Deep Learning. Int. J. Urban Sci. 2025, 29, 66–91. [Google Scholar] [CrossRef]
  35. Baldassi, S.; Megna, N.; Burr, D.C. Visual Clutter Causes High-Magnitude Errors. PLoS Biol. 2006, 4, e56. [Google Scholar] [CrossRef] [PubMed]
  36. Ewing, G.J.; Woodruff, C.J.; Vickers, D. Effects of “local” Clutter on Human Target Detection. Spat. Vis. 2006, 19, 37–60. [Google Scholar] [CrossRef] [PubMed]
  37. Roster, C.A.; Ferrari, J.R.; Jurkat, M.P. The Dark Side of Home: Assessing Possession “clutter” on Subjective Well-Being. J. Environ. Psychol. 2016, 46, 32–41. [Google Scholar] [CrossRef]
Figure 1. Research framework diagram.
Figure 1. Research framework diagram.
Land 15 01402 g001
Figure 2. Pearson correlation matrix of the four streetscape visual clutter indicators.
Figure 2. Pearson correlation matrix of the four streetscape visual clutter indicators.
Land 15 01402 g002
Figure 3. Spatial distributions of visual clutter.
Figure 3. Spatial distributions of visual clutter.
Land 15 01402 g003
Figure 4. Spatial distributions of the six model-estimated urban perception scores.
Figure 4. Spatial distributions of the six model-estimated urban perception scores.
Land 15 01402 g004
Table 1. Measurement of the four streetscape visual clutter dimensions.
Table 1. Measurement of the four streetscape visual clutter dimensions.
Clutter
Dimension
Preprocessing and Feature ExtractionBoundary ExtractionOutput
Edge
clutter
Grayscale conversionCanny edge detectionStructural-edge map
Color
clutter
HSV conversion and K-means clustering ( k = 5 )Boundaries between dominant color categoriesColor-boundary map
Semantic
clutter
OneFormer segmentation and category regroupingBoundaries between semantic categoriesSemantic-boundary map
Depth
clutter
DPT-Hybrid estimation and K-means clustering ( k = 4 )Boundaries between relative depth layersDepth-boundary map
Table 2. OLS regression results for streetscape visual clutter and model-estimated urban perceptions with built-environment controls.
Table 2. OLS regression results for streetscape visual clutter and model-estimated urban perceptions with built-environment controls.
Explanatory
Variables
BeautyWealthLivelinessSafetyBoredomDepressiveness
t-Value95%CIt-Value95%CIt-Value95%CIt-Value95%CIt-Value95%CIt-Value95%CI
edge
clutter
31.465
***
[0.114,
0.130]
17.292
***
[0.062,
0.078]
30.918
***
[0.106,
0.120]
26.424
***
[0.098,
0.114]
−2.014
*
[−0.017,
−0.000]
−7.053
***
[−0.038,
−0.022]
color
clutter
24.565
***
[0.112,
0.131]
14.155
***
[0.062,
0.082]
16.479
***
[0.067,
0.085]
13.985
***
[0.060,
0.079]
−14.683
***
[−0.089,
−0.068]
−3.783
***
[−0.032,
−0.010]
semantic
clutter
−10.476
***
[−0.047,
−0.032]
38.506
***
[0.140,
0.155]
31.788
***
[0.102,
0.116]
10.042
***
[0.030,
0.044]
−23.165
***
[−0.098,
−0.083]
−14.044
***
[−0.064,
−0.048]
depth
clutter
−15.165
***
[−0.064,
−0.050]
−20.181
***
[−0.084,
−0.069]
−6.214
***
[−0.028,
−0.015]
1.365[−0.002,
0.013]
7.464
***
[0.022,
0.037]
10.946
***
[0.036,
0.052]
vegetation
proportion
37.959
***
[0.197,
0.219]
54.718
***
[0.305,
0.328]
55.975
***
[0.276,
0.296]
60.407
***
[0.324,
0.345]
−33.139
***
[−0.212,
−0.188]
−38.230
***
[−0.238,
−0.215]
building
proportion
−40.488
***
[−0.181,
−0.164]
61.274
***
[0.263,
0.280]
95.935
***
[0.365,
0.380]
32.960
***
[0.133,
0.150]
−47.553
***
[−0.226,
−0.208]
6.237
***
[0.020,
0.037]
road
proportion
25.687
***
[0.095,
0.111]
14.846
***
[0.052,
0.068]
55.916
***
[0.200,
0.214]
27.198
***
[0.103,
0.118]
−28.725
***
[−0.131,
−0.115]
−28.189
***
[−0.133,
−0.116]
wall
proportion
−37.336
***
[−0.143,
−0.129]
−1.849[−0.014,
0.000]
−17.109
***
[−0.065,
−0.052]
−27.545
***
[−0.108,
−0.094]
−2.898
**
[−0.019,
−0.004]
13.354
***
[0.045,
0.061]
vehicle
proportion
−21.897
***
[−0.086,
−0.072]
−19.619
***
[−0.077,
−0.063]
29.754
***
[0.090,
0.103]
−10.552
***
[−0.046,
−0.032]
6.044
***
[0.016,
0.031]
−1.455[−0.014,
0.002]
pedestrian
proportion
32.595
***
[0.114,
0.129]
−4.513
***
[−0.024,
−0.010]
35.911
***
[0.115,
0.128]
−3.050
**
[−0.019,
−0.004]
−23.019
***
[−0.099,
−0.083]
−8.743
***
[−0.043,
−0.027]
N69,24469,24469,24469,24469,24469,244
Adjusted R20.2290.2010.3510.2020.1250.074
F-statistic 1892.104
***
1644.872
***
3355.414
***
1683.919
***
928.996
***
531.666
***
AIC178,496.873180,949.105166,544.286180,934.607187,265.691191,165.997
BIC178,597.472181,049.704166,644.886181,035.206187,366.291191,266.596
Notes: HC3 heteroskedasticity-robust t-values and 95% confidence intervals for the regression coefficients are reported. The model-level F-statistics are HC3 robust. AIC and BIC are based on the fitted OLS models. Background shading indicates the direction and statistical significance of the associations: red indicates a significant positive association, blue indicates a significant negative association, and gray indicates a nonsignificant association. * p < 0.05, ** p < 0.01, *** p < 0.001.
Table 3. Summary of nonlinear relationship tests between streetscape visual clutter and urban perceptions.
Table 3. Summary of nonlinear relationship tests between streetscape visual clutter and urban perceptions.
PerceptionClutterLOWESSSegmented RegressionQuadratic RegressionSensitivityStrict
Inverted-U
ShapeShape
(Knot)
Left
Slope
Right
Slope
Hinge pShapeb2L/P
beautyedge clutterNo clear patternNo clear pattern
(knot = −0.759)
0.241
***
0.079
***
3.26 × 10−28
***
Inverted-U-shaped−0.027
***
0/0No
color clutterInverted-U-shapedInverted-U-shaped
(knot = 0.802)
0.141
***
−4.39 × 10−2
ns
4.09 × 10−12
***
No clear pattern2.47 × 10−2
ns
3/2No
semantic clutterNo clear patternNo clear pattern
(knot = 1.165)
−7.58 × 10−3
ns
−0.311
***
2.82 × 10−49
***
Inverted-U-shaped−0.022
***
0/1No
depth clutterNo clear patternInverted-U-shaped
(knot = −1.223)
3.69 × 10−3
ns
−0.067
***
1.89 × 10−4
***
No clear pattern−4.11 × 10−3
ns
0/1No
wealthedge clutterNo clear patternInverted-U-shaped
(knot = −0.552)
0.208
***
−1.22 × 10−4
ns
6.39 × 10−52
***
Inverted-U-shaped−0.031
***
0/3No
color clutterNo clear patternU-shaped
(knot = −0.859)
−1.50 × 10−3
ns
0.085
***
4.09 × 10−4
***
U-shaped0.011
*
0/0No
semantic clutterInverted-U-shapedInverted-U-shaped
(knot = 1.165)
0.184
***
−0.162
***
4.12 × 10−61
***
Inverted-U-shaped−0.021
***
3/3Yes
depth clutterInverted-U-shapedInverted-U-shaped
(knot = −0.558)
0.043
***
−0.138
***
9.19 × 10−42
***
Inverted-U-shaped−0.030
***
3/3Yes
livelinessedge clutterNo clear patternNo clear pattern
(knot = −0.617)
0.208
***
0.070
***
1.38 × 10−27
***
Inverted-U-shaped−0.019
***
0/0No
color clutterNo clear patternNo clear pattern
(knot = −0.700)
1.98 × 10−2
ns
0.090
***
2.34 × 10−4
***
U-shaped0.015
***
0/0No
semantic clutterNo clear patternU-shaped
(knot = −1.158)
−0.056
***
0.133
***
7.48 × 10−31
***
U-shaped0.012
***
0/0No
depth clutterNo clear patternNo clear pattern
(knot = −0.894)
−0.057
***
−0.011
*
8.67 × 10−4
***
U-shaped0.007
**
0/0No
safetyedge clutterNo clear patternNo clear pattern
(knot = −0.617)
0.323
***
7.35 × 10−3
ns
4.72 × 10−112
***
Inverted-U-shaped−0.050
***
0/0No
color clutterNo clear patternInverted-U-shaped
(knot = 0.382)
0.102
***
−2.36 × 10−2
ns
1.11 × 10−11
***
Inverted-U-shaped−0.025
***
0/3No
semantic clutterNo clear patternInverted-U-shaped
(knot = 0.907)
0.087
***
−0.222
***
4.93 × 10−74
***
Inverted-U-shaped−0.033
***
1/3No
depth clutterNo clear patternNo clear pattern
(knot = −1.223)
0.036
*
−1.20 × 10−4
ns
5.62 × 10−2
ns
No clear pattern−1.84 × 10−3
ns
0/0No
boredomedge clutterInverted-U-shapedInverted-U-shaped
(knot = 0.486)
0.026
***
−0.128
***
2.71 × 10−21
***
Inverted-U-shaped−0.024
***
3/3Yes
color clutterNo clear patternNo clear pattern
(knot = 0.496)
−0.028
***
−0.271
***
1.01 × 10−30
***
Inverted-U-shaped−0.050
***
0/0No
semantic clutterNo clear patternInverted-U-shaped
(knot = −1.158)
2.48 × 10−2
ns
−0.108
***
3.72 × 10−12
***
No clear pattern−3.98 × 10−3
ns
0/1No
depth clutterU-shapedU-shaped
(knot = −1.223)
−0.050
**
0.043
***
3.91 × 10−6
***
No clear pattern2.20 × 10−3
ns
0/1No
depressive-
ness
edge clutterNo clear patternNo clear pattern
(knot = −0.492)
−0.054
***
−0.017
*
0.010
*
No clear pattern5.25 × 10−4
ns
0/0No
color clutterInverted-U-shapedInverted-U-shaped
(knot = −0.264)
0.100
***
−0.097
***
1.50 × 10−28
***
Inverted-U-shaped−0.054
***
3/3Yes
semantic clutterNo clear patternInverted-U-shaped
(knot = −1.158)
0.066
***
−0.074
***
7.60 × 10−13
***
Inverted-U-shaped−0.016
***
0/3No
depth clutterInverted-U-shapedNo clear pattern
(knot = −0.262)
0.023
**
0.062
***
0.003
**
U-shaped0.006
*
2/0No
Notes: N = 69,244. Slopes and quadratic coefficients (b2) are followed by significance levels based on Benjamini–Hochberg FDR-adjusted p-values. Hinge-test significance is also FDR-adjusted. Sensitivity values report the number of inverted-U classifications across three LOWESS tests and three segmented-regression tests, respectively. Strict inverted-U = Yes only when LOWESS, segmented regression, and quadratic regression are mutually consistent and the result remains stable in sensitivity tests. Background colors indicate the classified nonlinear patterns: light blue indicates no clear pattern, light red indicates an inverted-U-shaped relationship, and light yellow indicates a U-shaped relationship. In the sensitivity column, light orange highlights results for which all three LOWESS tests identified an inverted-U-shaped relationship; in the “Strict Inverted-U” column, light red indicates that the strict inverted-U criteria were satisfied. Gray font indicates statistically nonsignificant results. * p < 0.05, ** p < 0.01, *** p < 0.001; ns, not significant.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Li, F.; Lu, X.; Shen, J.; Wang, Y. Analysis and Considerations to Lessen Visual Clutter: Multidimensional Associations Between Streetscape Visual Clutter and Urban Perceptions in Shanghai. Land 2026, 15, 1402. https://doi.org/10.3390/land15081402

AMA Style

Li F, Lu X, Shen J, Wang Y. Analysis and Considerations to Lessen Visual Clutter: Multidimensional Associations Between Streetscape Visual Clutter and Urban Perceptions in Shanghai. Land. 2026; 15(8):1402. https://doi.org/10.3390/land15081402

Chicago/Turabian Style

Li, Fujun, Xinghao Lu, Jiake Shen, and Yuncai Wang. 2026. "Analysis and Considerations to Lessen Visual Clutter: Multidimensional Associations Between Streetscape Visual Clutter and Urban Perceptions in Shanghai" Land 15, no. 8: 1402. https://doi.org/10.3390/land15081402

APA Style

Li, F., Lu, X., Shen, J., & Wang, Y. (2026). Analysis and Considerations to Lessen Visual Clutter: Multidimensional Associations Between Streetscape Visual Clutter and Urban Perceptions in Shanghai. Land, 15(8), 1402. https://doi.org/10.3390/land15081402

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop