Skip to Content
BuildingsBuildings
  • Article
  • Open Access

19 July 2026

When Visual Conservation Meets Auditory Homogenization: Multimodal Evidence from Ximen Street Historic District, Qujing, China

,
and
1
Faculty of Architecture and City Planning, Kunming University of Science and Technology, Kunming 650500, China
2
Institute of Urban and Sustainable Development, City University of Macau, Avenida Padre Tomás Pereira, Taipa, Macau 999078, China
*
Authors to whom correspondence should be addressed.

Abstract

The conservation of historic districts is often evaluated on the basis of visible material features, yet commercialized renewal may weaken visual and auditory cues of locality while preserving building façades and street textures. We established 340 sampling points at 20 m intervals along Ximen Street Historic District in Qujing, Yunnan Province, and constructed a streetscape–audio matched sample dataset. Through multimodal model-assisted recognition, K-means clustering, XGBoost, and SHAP methods, we identified the degree of visual and auditory homogenization and analyzed the nonlinear effects and interactions of visual and auditory factors on the perception of homogenization. The results show that (1) visual homogenization presents a continuous distribution, while soundscape homogenization presents a patch-like distribution and is higher overall; (2) commercial symbolism, generic decoration, and open interfaces lead to visual substitutability, while commercial broadcasting and mechanical sounds intensify soundscape standardization; dialects, residents’ conversations, and diverse sound events enhance the perception of locality; and (3) visual commercialization and soundscape standardization have a clear interaction effect. This case study provides insights into the living conservation and renewal of historic districts in similar contexts and may offer a fine-grained analytical reference for collaborative governance that considers both visible townscape features and daily soundscape locality.

1. Introduction

1.1. Realistic Dilemmas in the Conservation of Historic Districts and the Continuation of Local Identity

Historic districts serve as important spatial carriers of urban historical memory and local cultural identity [1]. Against the backdrop of ongoing urban regeneration, they have become key spatial units for balancing heritage conservation, spatial vitality, and public value [2]. At the international level, the 1987 Washington Charter expanded the scope of conservation from individual buildings to historic towns, urban areas, and their natural and human-made environments. It also emphasized that the conservation of historic urban areas should be integrated into urban planning while accommodating residents’ daily lives and broader social development [3]. In 2011, UNESCO’s Recommendation on the Historic Urban Landscape further proposed that the built environment, spatial patterns, natural conditions, sociocultural practices, and intangible values should be addressed in an integrated manner within a broader urban context. This marked a shift in historic urban conservation from the preservation of physical remains toward a more holistic and living approach [4]. Accordingly, the conservation of historic districts should extend beyond the restoration and preservation of physical heritage. Greater attention should also be given to the role of residents’ everyday lives, social interactions, and community participation in sustaining local values [4,5].
In China, the Regulations on the Protection of Famous Historical and Cultural Cities, Towns, and Villages emphasize the integrated protection of traditional spatial patterns, historic character, and urban scale at the national level. They also require that the authenticity and integrity of historical and cultural heritage be maintained [6]. At the provincial level, the Regulations of Yunnan Province on the Protection of Famous Historical and Cultural Cities, Towns, Villages, and Streets provide a legal basis for the conservation, management, and appropriate use of historic and cultural districts in Yunnan [7]. Achieving a balance among the continuity of historical culture, the improvement of living environments, and the enhancement of local vitality has therefore become a critical issue for the sustainable development of historic districts.
Visible material elements constitute an important basis for assessing conservation effectiveness [6,7]. However, the continuity of locality is also shaped by the social connections, collective memories, and daily experiences formed between residents and places [8]. Existing approaches to the identification, evaluation, and management of historic district conservation and renewal are usually based on material-spatial elements, such as historic buildings and their façades [9], materials and construction characteristics [10], and district patterns and their relationships with open spaces [3]. They also maintain overall character through the coordinated control of townscape elements, including architectural colors, shop signs, decorations, and street interfaces [11]. By contrast, relatively limited attention has been paid to perceptual elements, such as sound and daily activities, in the evaluation of authenticity in historic districts [12].
In the renewal process, the introduction of commercial functions, tourism consumption, and streetscape improvement may bring financial investment and spatial vitality, but they may also reshape residents’ livelihoods, community life, and the structure of daily activities [13]. The continued introduction of replicated visual elements, such as standardized commercial colors, signs, and decorations, may weaken the original visual characteristics and local recognizability of historic districts [14]. When the shaping of commercial space becomes further detached from residents’ daily lives, residents’ sense of place, sense of belonging, and local identity may also be affected [15]. Accordingly, the homogenization of historic districts examined in this paper refers to the process in which place-specific spatial forms, environmental characteristics, and daily life cues are weakened or replaced under the continued intervention of standardized and replicable renewal elements. Whether the conservation of visible townscape features alone can adequately identify the continuity of perceived locality in historic districts remains to be further examined.

1.2. Insufficient Synergy in Visual–Auditory Perception Assessment for Historic District Conservation

Streetscape imagery has become an important data source for research on the urban built environment [16]. It has produced significant achievements in identifying façade style characteristics [17], street interface forms [18], intensity of commercial signage [19], distribution of architectural colors [20], and degree of spatial openness [21].
Compared with traditional field surveys, streetscape imagery can assess the street environment from a perspective close to that of pedestrians, thereby effectively improving the spatial precision and comparability of urban environmental measurements [22,23]. With the development of computer vision and multimodal artificial intelligence methods, streetscape studies have expanded from the identification of physical elements, such as buildings, roads, and vegetation, to the evaluation of semantic dimensions related to human-centered experience, including street quality [24], visual order [25], and environmental atmosphere [26]. These methods can provide important technical support for identifying visual perception in historic districts.
In recent years, soundscape research in historical and heritage contexts has primarily focused on the acoustic characteristics of historic urban areas [27], soundscape perception and visitor experience [28,29], and the cultural meanings and place-based representations of distinctive sound sources [30]. Acoustic communication theory suggests that particular sounds can acquire place-specific meanings through long-term social interaction and become important auditory elements representing local identity [30]. However, this line of research has largely emphasized the interpretation of the cultural meanings of sound and has rarely translated locally distinctive sounds into spatially explicit and measurable conservation approaches.
Previous studies have used sound source surveys, sound pressure level measurements, and time–frequency analyses to identify the composition of acoustic environments in historic urban areas. They have also examined the conservation value of characteristic sounds with local significance [27]. Nevertheless, these studies have mainly concentrated on acoustic properties and the spatial distribution of sound sources, making it difficult to capture how different sounds influence the perception of historical atmosphere and sense of place. Building on this work, other studies have investigated the effects of soundscape perception on visitor experience and the ways in which different spatial contexts within historic districts shape soundscape evaluations [28,29]. However, most of this research has centered on tourist experiences, with limited attention paid to residents’ everyday soundscapes.
At the regulatory level, the Law of the People’s Republic of China on Noise Pollution Prevention and Control and the Environmental Quality Standard for Noise address harmful noise through emission control, acoustic environment functional zoning, and environmental quality limits [31,32]. However, clear provisions and operational evaluation criteria for conserving soundscape environments with local cultural significance remain lacking. Overall, existing research has yet to establish an integrated framework that combines residents’ everyday soundscapes, place-based perception, and spatial measurement. It has also not fully revealed the mechanisms through which sound and the visible environment jointly shape historical atmosphere and local identity. However, human environmental perception is characterized by multisensory integration. Vision and hearing do not act independently in environmental evaluation; rather, they mutually regulate one another in processes of attention, emotion, and judgment of meaning, jointly shaping people’s comprehensive experience of environmental quality, atmosphere, and behavioral suitability [33,34]. Vision primarily provides visible information such as architectural form, material texture, street scale, interface order, and commercial landscape, while sound conveys dynamic information related to crowd activities, modes of social interaction, functional use, and environmental rhythm [35,36,37]. The interaction between the two further influences individuals’ judgments of place perception. Therefore, a fragmented evaluation of visual and auditory dimensions may equate the preservation of visible townscape features with the continuity of locality, while overlooking situations in which traditional forms are retained but daily soundscapes change. It is therefore necessary to integrate both dimensions within the same spatial unit for evaluation.

1.3. Research Objectives

Previous studies have provided an important foundation for the conservation and renewal of historic districts, yet several limitations remain. First, current evaluations of historic district conservation and renewal largely rely on visible spatial elements to assess conservation effectiveness, while auditory cues are considered less frequently. As a result, they are unable to fully identify the extent to which locality is sustained in daily perception. Second, visual and auditory studies are usually conducted separately, with insufficient attention paid to their coordinated representation and spatial correspondence within the same place-based context. Moreover, nonlinear and interaction relationships may exist between visual and auditory elements in historic districts, whereas existing single-dimensional or linear analyses are still limited in revealing their combined mechanisms of influence on the perception of homogenization.
We collected streetscape images and field recordings from Ximen Street Historic District in Qujing, Yunnan Province, to use as a case study. After recognition using LLMs and VLMs and validation through manual perceptual evaluation, a visual–soundscape matched dataset was established at the sampling-point scale. An indicator system was then constructed from three dimensions: heritage authenticity, landscape standardization, and soundscape locality. By further integrating spatial distribution analysis, K-means clustering, the XGBoost model, and SHAP-based interpretation, we identified the distribution and types of visual and auditory homogenization, and revealed the mechanisms through which different visual and auditory variables affect the perception of sensory homogenization. The study addresses the following questions:
Q1: What are the degrees of visual and auditory homogenization in the Ximen Street Historic District? What spatial differences do they exhibit, and how are they associated with street hierarchy and spatial functions?
Q2: What types of multisensory homogenization can be identified in the Ximen Street Historic District? How do these types differ in terms of visual townscape, commercial interface, and soundscape locality?
Q3: Which visual and soundscape variables are important predictors of multisensory homogenization? Do these variables exhibit meaningful thresholds or interaction effects?
This study contributes to the literature in three main ways. Theoretically, this study extends the discussion of homogenization in historic districts from the traditional focus on visual townscape standardization to the integrated perceptual dimensions of vision and hearing. It emphasizes that the acoustic environment and visual townscape jointly participate in the construction of heritage identity, local authenticity, and daily sense of place. Methodologically, the study provides an operational pathway for measuring perceived homogenization within historic districts. Practically, the findings offer a reference for refined governance in the conservation and renewal of historic districts of a similar type.

2. Study Area, Datasets, and Research Methods

2.1. Study Area

The Ximen Street Historic District is located in Qilin District, Qujing City (25.49° N, 103.80° E). It extends north to Wenchang Street, south to Shengfeng Road, west to Liaokuo South Road, and east to Qilin South Road, covering a total area of approximately 24 ha. It is not only the core space of the spatial development of the ancient city of Qujing, but also the only well-preserved historic district remaining in Qilin District. The district has formed a relatively stable spatial framework around Qujing. It has largely maintained the nested “street–lane–courtyard” spatial pattern, and preserved the traditional residential style of the “Yikeyin” courtyard house. It is therefore a historical epitome of urban development in eastern Yunnan and an important material carrier of the cultural heritage of Cuan culture (Figure 1).
Figure 1. Location and current conditions of the Ximen Street Historic District, Qujing City: (a) location of Qujing City in Yunnan Province; (b) location and geographical setting of Qilin District within Qujing City; (c) location of the Ximen Street Historic District within Qilin District; and (d) current conditions of the Ximen Street Historic District (Source: Compiled by the authors based on materials provided by the Qujing Planning Bureau).
Compared with more intensively developed and commercially oriented historic urban areas in Yunnan Province, such as the Old Town of Lijiang and the Ancient City of Dali, the Ximen Street Historic District in Qujing retains stronger characteristics of everyday life and neighborhood interaction. It therefore represents a type of historic district in Yunnan’s small- and medium-sized cities that has not undergone intensive tourism development (Table 1).
Table 1. Comparison of different types of historic districts in Yunnan.
Since the launch of the urban regeneration “micro-renewal” initiative in 2022, the Ximen Street Historic District in Qujing has gradually implemented a series of interventions, including heritage conservation, infrastructure improvement, public space enhancement, commercial revitalization, and the upgrading of street interfaces. The district therefore exhibits a representative pattern in which heritage conservation, everyday life, and incremental renewal coexist and interact. Within the broader context of urban regeneration in China, this type of historic district is highly representative of current conservation and renewal practices. It provides an appropriate empirical setting for examining the relationships among heritage conservation, commercial activities, and everyday environments, as well as for exploring the assessment of multisensory homogenization in historic districts.

2.2. Datasets

2.2.1. Streetscape Data Collection

Before streetscape data collection, the road network of the study area was first extracted from OSM, and the study boundary and existing street–lane network were examined in GIS. Main streets, secondary streets, internal lanes, and daily living spaces within the Ximen Street Historic District were then identified. Sampling points were arranged along the centerlines of streets and lanes at intervals of approximately 20 m, resulting in a total of 340 sampling points.
Streetscape images were collected from 9:00 to 18:30 on 9–10 December 2025, using an iPhone 15 Pro Max. During the collection period, the weather was clear, with a light breeze of approximately Level 1–2 (Figure 2). The sampling route covered urban arterial roads, secondary roads, core internal streets, and narrow lanes within the district to reflect differences among spatial types and street interfaces. During image acquisition, the shooting height was controlled at approximately 1.7 m to simulate pedestrians’ daily visual perception. To ensure sample comparability, approximately 180° panoramic images were collected at each sampling point along the main axis of the street or lane and in the opposite direction, together covering an approximately 360° field of view. For intersections and open spaces, the shooting orientation was determined according to the main directions of pedestrian activity and the principal visible street interfaces.
Figure 2. Sampling design, streetscape image examples and spatial structure of the Ximen Street Historic Area: (a) distribution of 340 sampling points and workflow of field-collected streetscape image acquisition; (b) representative streetscape images of different spatial types; and (c) spatial classification of streets, lanes, public spaces and green spaces (Source: Drawn and photographed by the authors).
After field collection, a total of 701 original streetscape images were obtained. All images were manually screened for quality, and images with obvious occlusion, blurring, severe exposure problems, duplicated content, or insufficient representation of street-space characteristics were removed or replaced. For each sampling point, two images from adjacent viewing directions were selected and stitched to form a 360° streetscape image covering the street environment. Finally, 340 valid field streetscape images were obtained (Figure 2).

2.2.2. Sound Data Collection

Soundscape data were collected simultaneously with streetscape images to achieve one-to-one image–audio matching at the same sampling points. Recordings were mainly made using the built-in microphone of an iPhone 15 Pro Max, supplemented by an Apple Watch Series 10 to observe changes in the on-site environmental sound level. All recordings were conducted using the same device, recording mode, and operating procedure, and the original audio files were saved in M4A format. During data processing, the original audio files were uniformly converted to wav format with a sampling rate of 48 kHz and 16-bit PCM encoding. Before acoustic feature extraction, the audio files were further converted to mono to ensure comparability among samples.
Soundscape sampling took place at the same 340 sampling points used for streetscape image collection, with an interval of approximately 20 m between points. Each audio segment corresponded to one valid streetscape image at the same location, forming an image–audio matched dataset. During the field survey, the “Two Steps” app was used to record route and time information in order to verify the sampling route, collection sequence, and point–location correspondence.
The researchers walked throughout the sound recording survey and avoided using bicycles, electric bicycles, or other transport modes that could introduce non-environmental sound interference. After arriving at each sampling point, the researchers stopped walking, kept the device stable, and recorded toward the main axis of the street or lane to capture an acoustic environment close to daily pedestrian experience. Each recording lasted approximately 20–30 s.
After field collection, 383 audio records were obtained. Through manual screening, samples with obvious sudden interference were removed, and 340 valid soundscape samples were ultimately retained and matched one-to-one with the 340 valid streetscape images. Daily traffic sounds, residents’ conversations, commercial activity sounds, and tourist-related sounds were retained as components of the district soundscape, while only occasional, sudden, and unrepresentative interfering sounds were removed.

2.3. Methods

2.3.1. Research Framework

We constructed a homogenization assessment framework for historic districts that integrates multisource data and multimodal models. The framework consists of five main steps. First, streetscape images, soundscape audio, and building–road network vector data corresponding spatial-location data were collected to construct a multimodal base database. After data cleaning and manual screening, 340 valid streetscape images and 340 matched audio samples were retained to construct the multimodal database. Second, images and audio were spatially matched, and basic features were extracted. At the visual level, semantic elements such as buildings, sky, and vegetation were quantified. At the auditory level, acoustic features such as pitch, zero-crossing rate, and MFCCs were extracted. Third, visual–language models (VLMs) and audio–language models (ALMs) were used to construct high-dimensional homogenization indicators. These models were applied to identify audiovisual features, including façade historical traces, material locality, signage standardization, dialect intensity, and the proportion of commercial broadcasting. The reliability of the scores was then validated through a manual perceptual experiment. Fourth, spatial differentiation analysis, multisensory homogenization type identification, and SHAP-based interpretation were conducted to reveal the spatial patterns, dominant drivers, and nonlinear effects of perceived homogenization within the district. Finally, the analytical findings were translated into differentiated conservation and management recommendations, including type-specific responses, control of commercial and technical interventions, maintenance of everyday soundscapes, and preservation of local identity (Figure 3).
Figure 3. Research framework (Source: authors).

2.3.2. Construction of Independent and Dependent Variables

To explain why perceptions of homogenization emerge in historic districts, we organized the independent variables into two modal layers, visual and auditory, based on the logic of “locality retention–standardized replacement.” These variables are further divided into four secondary dimensions: weakening of visual authenticity, standardization of commercial landscapes, weakening of soundscape locality, and intrusion of commercial–technical sound sources. The indicator system includes eight visual variables and six auditory variables, comprising a total of 14 independent variables (Table 2). To improve the discrimination of subtle audiovisual differences among samples and reduce tied scores, we adopted a 0–10 rating scale for use in the AI-assisted assessment. The original 0–10 scores were used for spatial analysis, cluster analysis, and machine-learning modeling. They were converted to a 0–5 scale during the human–AI agreement analysis, so as to ensure consistency with the scale used in the human evaluation.
Table 2. Independent variable framework and 0–10 scoring criteria for multisensory homogenization in historic districts (Source: Compiled by the authors).
The weakening of visual authenticity was used to characterize the loss of place-specific material cues in historic buildings and street interfaces. It comprised four indicators: traces of historical development on façades, local distinctiveness of materials, complexity of architectural details, and organic character of street interfaces. Commercial landscape standardization was used to capture the replacement of everyday-life settings by consumption-oriented symbols. It was assessed using four indicators: standardization of storefront signage, prominence of chain brands, generic decoration, and openness of commercial interfaces.
For the auditory dimension, the weakening of soundscape locality was represented by the intensity of local dialect use and the diversity of sound events. Commercial and technological sound intrusion was assessed using four indicators: intensity of commercial broadcasting, influence of tourism activities, proportion of technological sound sources, and perceived quietness.

2.3.3. LLM- and ALM-Assisted Scoring

A total of 16 AI-assisted assessment items were established, comprising eight primary visual perception indicators, six primary auditory perception indicators, one overall visual homogenization indicator, and one overall auditory homogenization indicator. The visual and auditory indicators were evaluated using a visual language model (VLM) and an audio language model (ALM), respectively. The models initially generated scores on a 0–10 scale, which were subsequently converted to a 0–5 scale to ensure consistency with the human perception assessment and subsequent statistical analyses. Following the predefined scoring criteria, the models made judgments solely on the basis of information visible in the images or audible in the recordings. They were not permitted to infer information beyond the provided samples. The complete prompts and model parameter settings are presented in Table A1, Table A2, Table A3 and Table A4.
The street-view image and audio recording corresponding to each sampling point were treated as independent units of analysis. Both visual and auditory assessments were conducted using the GPT-5.4 model, which was accessed on 7 April 2026. All street-view images were stored in RGB JPEG format. Their widths were standardized to 1024 pixels while preserving the original aspect ratios. No further resizing or compression was performed during the scoring process; instead, the processed images were directly encoded and submitted to the model. The audio recordings were stored in M4A/AAC format with a sampling rate of 48 kHz and a single audio channel.
Both visual and auditory samples were processed through separate model calls for each individual sample. For the visual model, the temperature was set to 0.1, the maximum output length to 3000 tokens, the top-p value to 0.8, the top-k value to 40, and the request timeout to 90 s. The program automatically retried a request up to five times in cases of empty responses, missing fields, out-of-range scores, JSON-parsing failures, request timeouts, or other application programming interface errors. The type of error and the number of retry attempts were recorded. Samples that still failed to produce valid results after five automatic retries were manually checked by the researchers for problems with the original files or request outputs and were then resubmitted to the model. Invalid outputs and placeholder values were excluded from subsequent statistical analyses.
The pilot experiment showed that processing multiple sampling points within a single request frequently resulted in missing fields, null values, or JSON-parsing failures. The failure rate was substantially higher for images than for audio recordings. After adopting a sample-by-sample processing strategy, 337 of the 340 candidate samples produced complete and parsable structured outputs on the first attempt, corresponding to an initial valid generation rate of 99.14%. The remaining three samples were successfully processed after manual inspection and resubmission, resulting in a final valid generation rate of 100%. Following image quality control and one-to-one image–audio matching, valid recognition data were obtained for all 340 samples.
Finally, the model returned the sample identifier, individual indicator scores, and a brief justification for each judgment in strict JSON format. The program then automatically performed JSON parsing, required-field verification, score-range validation, and CSV export. The complete implementation code is available from the corresponding author upon reasonable request.

2.3.4. Human Perception Validation and Agreement Assessment

A human perception validation experiment was conducted using matched image–audio samples. Its sole purpose was to assess the agreement between AI-assisted scores and human perceptual judgments. Thirty matched image–audio sample sets were selected from the 340 sampling points. These samples represented spatial locations with pronounced differences across the district, thereby ensuring that the validation set captured representative audiovisual variation (Figure 4).
Figure 4. Human validation survey workflow for visual and soundscape homogenization perception (Source: authors).
Given the exploratory purpose of the validation, 30 adult participants were recruited to complete the human assessment. Participant characteristics are reported in Table A5. A total of 900 valid evaluation records were obtained. The participant group provided basic coverage in terms of gender, age, disciplinary background, residential experience, and familiarity with the district. Each participant was required to evaluate all 30 matched image–audio sample sets, with each set involving separate visual and auditory perceptual judgments. Because this procedure imposed a substantial assessment burden, the study adopted a design comprising 30 participants and 30 validation samples to balance evaluation feasibility, perceptual diversity, and rating stability (Table A5 and Table A6).
During the experiment, participants completed the assessment in a relatively quiet environment while wearing headphones. They viewed each street-view image or listened to a 20–30 s audio recording before assigning a score. The samples were presented in randomized order, and all assessment items were rated using a 0–5 Likert scale (Figure 4).
The validation results indicated a high level of agreement between the AI-assisted scores and the mean human ratings. The overall Pearson correlation coefficient was r = 0.79 (p < 0.01). The corresponding coefficients were r = 0.82 for the visual dimension (p < 0.01), r = 0.77 for the auditory dimension (p < 0.01), and r = 0.80 for the overall homogenization assessment (p < 0.01). Pearson’s correlation coefficient primarily reflects the extent to which AI-generated scores and human perceptual judgments follow similar patterns of variation.
Further agreement analysis produced an overall intraclass correlation coefficient (ICC) of 0.81. The ICC values were 0.84 for the visual dimension and 0.78 for the auditory dimension. These results indicate good agreement between the AI-generated scores and the aggregated human ratings, with stronger agreement for the visual dimension than for the auditory dimension. Cronbach’s alpha was 0.87, indicating good internal consistency among the human raters. The AI-assisted scores were therefore used as a source of structured perceptual data in the subsequent spatial analyses and statistical modeling.
These validation results support the use of AI-assisted scores as structured approximations of human perception in subsequent spatial analyses and statistical modeling. However, such scores should not be regarded as objective substitutes for heritage value assessments, residents’ actual perceptions, or expert judgments. AI-generated scores may be influenced by the model’s training data, prompt design, image and audio quality, indicator definitions, and the local context of the case study area. Human evaluations may likewise be affected by participants’ age, disciplinary background, familiarity with the local area, ability to understand local dialects, and personal esthetic experience.

2.3.5. Model Design

First, we compared eight regression models: multiple linear regression (MLR), Ridge regression, support vector regression (SVR), Random Forest, Extra Trees, XGBoost, LightGBM, and CatBoost. All models were evaluated using the same 80/20 training–test split and five-fold cross-validation. Their predictive performance was assessed using R2, RMSE, and MAE. The results showed that the nonlinear models generally outperformed the linear models. Although CatBoost achieved the best overall predictive metrics, XGBoost also performed well on both the held-out test set and cross-validation, and provided a favorable balance between predictive accuracy and generalizability. Because the primary objective of this study was to identify variable importance, nonlinear response patterns, and audiovisual interaction effects, XGBoost was ultimately selected as the principal explanatory model. Its compatibility with SHAP analysis also enabled the development of a unified and reproducible model interpretation framework (Table 3 and Table 4).
Table 3. Training and test performance and generalization gaps of eight regression models.
Table 4. Comparative performance of eight regression models on the independent test set and five-fold cross-validation.
Before model fitting, multicollinearity among the 14 predictor variables was assessed using variance inflation factors (VIFs) and tolerance statistics. The VIF values ranged from 1.434 to 4.396, all below the commonly adopted threshold of 5, indicating no serious multicollinearity among the predictors (Table A8).
To identify the combined effects of visual and auditory environmental factors on perceived multisensory homogenization in the historic district, we developed an XGBoost-based machine-learning framework integrating predictive modeling, hyperparameter optimization, and interpretability analysis [38,39]. The model used 340 image–audio matched samples as analytical units, with each sample corresponding to one street-view image, one environmental audio clip, and one spatial location.
The independent variables consisted of 14 multisensory environmental indicators, including eight visual variables and six auditory variables. The visual variables included facade patina, material vernacularity, architectural complexity, street interface organic-ness, signage standardization, chain brand salience, generic decoration, and commercial interface openness. The auditory variables included local dialect intensity, commercial broadcasting intensity, sound event diversity, technophony ratio, tourist activity impact, and perceived calmness. To ensure directional consistency, all variables were processed so that higher values indicated a greater risk of multisensory homogenization. The original scoring scales were retained because XGBoost is insensitive to feature scaling and because preserving the original scale facilitates the identification of practically interpretable thresholds.
The dependent variable was the integrated multisensory homogenization perception score of each image–audio matched sample. This score reflected participants’ overall judgment of visual commercialization, soundscape standardization, and environmental interchangeability. It was measured on a 1–5 scale, with higher values indicating stronger perceived multisensory homogenization. Visual homogenization scores, soundscape homogenization scores, and their difference were only used for subsequent spatial mismatch analysis and were not included as model inputs, thereby avoiding dependent variable information leakage.
XGBoost was employed because it can capture nonlinear relationships and interaction effects among multisensory environmental variables while incorporating regularization mechanisms to reduce overfitting. This is particularly suitable for examining the potentially complex relationships among visual authenticity loss, commercial landscape standardization, change in soundscape locality, and integrated homogenization perception.
All samples were randomly divided into training and testing sets at an 8:2 ratio, yielding 272 training samples and 68 testing samples. Hyperparameter tuning was conducted only on the training set. Given the relatively limited sample size of this single-case study, five-fold cross-validation was implemented within the training set to evaluate model stability and identify the optimal parameter combination. The independent testing set was subsequently used to assess the generalization performance of the final model. Model performance was evaluated using the coefficient of determination (R2), root mean square error (RMSE), and mean absolute error (MAE). Higher R2 values and lower RMSE and MAE values indicate better predictive performance (Equations (A1) and (A2)).
Hyperparameter optimization was conducted using grid search combined with five-fold cross-validation. The search ranges were 100–800 for ‘n_estimators’, 0.01–0.10 for ‘learning_rate’, 2–6 for ‘max_depth’, 1–5 for ‘min_child_weight’, 0.6–1.0 for both ‘subsample’ and ‘colsample_bytree’, 0–0.5 for ‘gamma’, 0–1 for ‘reg_alpha’, and 0.5–5 for ‘reg_lambda’. During model training, the validation subset within each cross-validation fold was used as the early-stopping monitoring set, and training was stopped when no performance improvement occurred for 50 consecutive rounds. This procedure reduced the risk of overfitting caused by excessive tree growth.
The final XGBoost model achieved an R2 of 0.886, an RMSE of 0.345, and an MAE of 0.228 on the independent testing set, indicating good predictive performance for integrated multisensory homogenization perception.

3. Results

3.1. Spatial Differentiation of Audio-Visual Indicators

3.1.1. Spatial Differentiation of Visual Homogenization Indicators

The semantic composition of the district streetscape is dominated by buildings, vegetation, and roads, while sky and sidewalks account for relatively low proportions, and the remaining elements make up only small shares overall (Figure 5a). The indicators of visual homogenization show clear spatial differences among sampling points (Figure 5(b1–b4)).
Figure 5. Streetscape semantic composition and bivariate spatial patterns of visual homogenization indicators: (a) distribution of the top 25 semantic segmentation features; (b1) façade patina and material vernacularity; (b2) architectural complexity and interface organic-ness; (b3) signage standardization and chain brand salience; and (b4) generic decoration and commercial frontage openness (Source: authors).
Façade aging traces (FP) and material locality (MV) are generally at relatively high levels. High-value combinations are mainly concentrated in the southeastern part of the district and in traditional streets and lanes such as Ximen Street and Nanmen Street, where the original architectural townscape and local materials are relatively well preserved (Figure 5(b1)). High-value combinations of architectural complexity (AC) and interface organicity (IO) are mainly distributed in the southern part of the site and around Tianchi Park, where the street interface morphology and visual organization are relatively organic (Figure 5(b2)). Shop sign standardization (SS) and chain brand prominence (CS) are generally at low levels, with only localized high values appearing along some street-front commercial interfaces. Existing commercial activities are mainly composed of small local shops and traditional business types, while the intervention of chain based and standardized commercial signage remains relatively limited (Figure 5(b3)). High-value combinations of generic decoration (GD) and commercial interface openness (CO) are mainly concentrated along the periphery of the district, especially along urban arterial roads and at their intersections. In these areas, uniformly installed shop signs and open commercial interfaces are relatively concentrated (Figure 5(b4)).

3.1.2. Spatial Differentiation of Auditory Homogenization Indicators

The auditory indicators of the district show a distribution pattern characterized by the “dispersed retention of living sounds, the clustering of commercial and technical sound sources along main streets, and the localized embedding of quiet spaces.” Living sounds are discretely retained in internal streets and lanes, while commercial and technical sound sources are concentrated along main streets and peripheral roads. Quiet spaces appear only locally in Tianchi Park on the southern side and in adjacent slow traffic areas, and the overall influence of tourist activities is weak (Figure 6).
Figure 6. Spatial distribution of audio perception indicators: (a) local dialect intensity; (b) commercial broadcasting; (c) sound event diversity; (d) technophony ratio; (e) tourist activity impact; and (f) perceived calmness (Source: authors).
Local dialect intensity (LDI) is generally high, but its spatial distribution is relatively dispersed. High values mainly occur in internal living-oriented streets and lanes and in areas where residents’ activities are concentrated, whereas values are relatively low along peripheral roads and continuous commercial interfaces (Figure 6a). Commercial broadcasting intensity (CB) shows continuous high values along street sections with concentrated commercial shops, such as Ximen Street, Xueyuan Street, Qilin South Road, and Wenchang Street, while remaining generally low in internal branch lanes (Figure 6b). Sound event diversity (SED) is relatively high around schools, in internal living-oriented streets and lanes, and at some public activity nodes, indicating that the composition of sound sources in these areas is relatively diverse. The remaining street sections are dominated by medium and low values (Figure 6c).
The proportion of technical sound sources (TR) is generally high and is mainly distributed along peripheral roads, main streets, and traffic connection spaces. Localized high values also appear near cotton-fluffing workshops, construction sites, and interfaces where equipment is concentrated (Figure 6d). The influence of tourist activities (TAI) remains low overall and increases only locally around a few open spaces and commercial nodes. This indicates that changes in the district’s soundscape are mainly affected by daily commercial activities and technical sound sources, rather than by intensive tourism activities (Figure 6e). Perceived quietness (PC) is generally low, with high values mainly concentrated inside Tianchi Park on the southern side and in some adjacent slow traffic spaces. It is relatively low along peripheral roads, commercial main streets, and street sections with stronger technical sound sources (Figure 6f).

3.1.3. Distributional and Correlation Characteristics of Visual and Auditory Indicators

Relatively clear internal associations exist among traditional visual elements, among commercial landscape elements, and among some soundscape elements, while weak correlations are observed between visual and auditory indicators (Figure 7a).
Figure 7. Correlation structure and distributional characteristics of visual and audio homogenization indicators: (a) correlation matrix of visual and audio homogenization indicators; (b) distribution of visual indicators; (c) distribution of audio indicators (Source: authors).
Further comparison of Pearson and Spearman correlation coefficients revealed several moderate-to-strong associations among indicators within the same dimension. The strongest positive association was observed between façade patina and material vernacularity (r = 0.755r = 0.755r = 0.755, ρ = 0.756\rho = 0.756ρ = 0.756), whereas technophony ratio was moderately and negatively associated with perceived calmness (r = −0.641r = −0.641r = −0.641, ρ = −0.635\rho = −0.635ρ = −0.635). Moderate positive associations were also found between signage standardization and generic decoration, chain brand salience and commercial openness, commercial broadcasting and tourist activity impact, and local dialect intensity and sound event diversity. Differences between Pearson and Spearman coefficients for some variable pairs suggest potentially monotonic but not strictly linear relationships. Detailed results are reported in Table A9.
In terms of indicator distributions, the degree of dispersion differs markedly among visual indicators. Façade aging traces, material locality, and interface organicity show relatively wide value ranges, whereas architectural complexity, shop sign standardization, and chain brand prominence are more concentrated in lower value ranges. Generic decoration and commercial interface openness show more evident differences among sampling points (Figure 7b). Among the auditory indicators, local dialect intensity and the proportion of technical sound sources have relatively wide distribution ranges, while sound event diversity is mainly concentrated in the medium range. Commercial broadcasting and the influence of tourist activities are dominated by low-value samples, whereas perceived quietness shows a more distinct dispersed distribution (Figure 7c).

3.1.4. Spatial Autocorrelation of Visual and Auditory Homogenization Indices

To examine the spatial distribution patterns of the homogenization indices, Global Moran’s I was calculated for both visual and auditory homogenization using a fixed-distance-band spatial weight matrix, Euclidean distance, and row standardization. Based on the sampling interval and the spatial scale of the district, 50 m was selected as the primary distance threshold, while 40 m and 60 m were used for sensitivity testing.
The results showed that both the visual and auditory homogenization indices exhibited significant positive spatial autocorrelation across all distance thresholds (p < 0.001). For visual homogenization, Moran’s I ranged from 0.365670 to 0.448925. At the 50 m threshold, Moran’s I was 0.384152, with a z-score of 11.947115. For auditory homogenization, Moran’s I ranged from 0.128422 to 0.212447. At the 50 m threshold, Moran’s I was 0.166666, with a z-score of 5.234960. These findings indicate significant spatial clustering in both dimensions, with visual homogenization exhibiting stronger spatial autocorrelation than auditory homogenization (Table 5).
Table 5. Global Moran’s I results for visual and auditory homogenization indices under different distance thresholds (Source: Compiled by the authors).

3.1.5. Spatial Mismatch Between Visual and Auditory Homogenization Indicators

There is a clear perceptual mismatch between visual and auditory homogenization within the district, and soundscape homogenization is stronger overall than visual homogenization (Figure 8b). High-value sampling points are mainly distributed along peripheral roads such as Wenchang Street, Qilin South Road, Liaokuo South Road, and Shengfeng Road, and are clustered around the commercial interfaces of main streets such as Ximen Street and Nanmen Street, as well as near nodes where these streets connect with peripheral roads. By contrast, internal living-oriented streets and lanes, such as Zhuge Lane, Wufu Lane, and Fensi Lane, show relatively low levels of homogenization (Figure 8a).
Figure 8. Spatial statistical evidence and interpolated surfaces of visual–audio homogenization mismatch: (a) spatial distribution of visual–audio mismatch; (b) joint distribution of visual and audio homogenization; (c) comparison between visual and audio homogenization indices; (d) mean-mismatch plot of visual–audio homogenization; (e) distribution of visual–audio mismatch values; (f1) interpolated surface of visual homogenization; (f2) interpolated surface of audio homogenization; and (f3) interpolated surface of visual–audio mismatch (Source: authors).
At the overall level, both the median and mean values of soundscape homogenization in the district are higher than those of visual homogenization (Figure 8c). To further determine whether this difference exists at the level of individual sampling points, we used the Visual–Audio mismatch to represent the difference between visual homogenization and soundscape homogenization. The results show that many sampling points are located below the zero line, indicating that soundscape homogenization is higher than visual homogenization at most sampling points. A small number of sampling points fall within the positive-value range, suggesting that visual homogenization is stronger in some local street sections (Figure 8d). The distribution of mismatch values further confirms this tendency. The mean value of the Visual–Audio mismatch is −1.10, and the median value is −1.29, with the overall distribution skewed toward negative values. This demonstrates that “soundscape homogenization being stronger than visual homogenization” is not an isolated phenomenon at individual sampling points, but a dominant trend across the district as a whole (Figure 8e). Moreover, visual homogenization generally shows relatively continuous spatial variation, whereas soundscape homogenization is more characterized by localized patches and nodal clustering (Figure 8(f1–f3)).

3.2. Identification of Multi-Sensory Homogeneity Types

To identify different visual–soundscape combination characteristics within Ximen Street, we conducted K-means clustering analysis based on 14 standardized visual and auditory indicators. We further compared the silhouette coefficient, within-cluster sum of squares (WCSS), Calinski–Harabasz index, Davies–Bouldin index, and type interpretability under different K values.
The number of clusters was determined by jointly considering the silhouette coefficient, within-cluster sum of squares (WCSS), Calinski–Harabasz index, Davies–Bouldin index, and the interpretive value of the resulting typology. The results showed that K = 2 yielded the highest silhouette coefficient (0.1857) and Calinski–Harabasz index (63.67), indicating relatively strong overall cluster separation. This solution was therefore suitable for capturing the broad distinction between high and low levels of multisensory homogenization. However, it tended to obscure differences among specific visual–auditory combinations.
As K increased, WCSS declined continuously and the Davies–Bouldin index generally decreased, indicating improved cluster compactness. Beyond K = 6, however, the reduction in WCSS became less pronounced, while the Calinski–Harabasz index generally declined. This suggests that further increases in the number of clusters provided only limited marginal interpretive gains. Accordingly, K = 2 was retained as a reference for broad statistical classification, whereas K = 6 was adopted as the more detailed typological solution that balanced statistical performance with interpretability for planning and management (Table 6).
Table 6. Comparison of alternative K-means cluster solutions. (Source: Compiled by the authors).
For the K = 6 solution, the silhouette coefficient was 0.1578, the WCSS was 2791.21, the Calinski–Harabasz index was 47.12, and the Davies–Bouldin index was 1.7230. While maintaining an acceptable level of statistical performance, this solution distinguished six types characterized by different visual–auditory configurations and management needs: low commercial disturbance, commercially intensified soundscape, traditional townscape with technological sound disturbance, low visual homogenization with soundscape convergence, high commercial homogenization, and preserved traditional townscape with a mixed soundscape.
However, the K = 6 solution was not regarded as the uniquely optimal statistical solution. Rather, it was adopted as an interpretive classification designed to support the analysis of visual–soundscape mismatches and discussions of differentiated management strategies (Figure 9).
Figure 9. K-means clustering results and cluster-specific distributions of multisensory homogenization features: (a) silhouette scores under different numbers of clusters, with K = 6 selected for typological interpretation; (b) t-SNE projection of the six cluster types; (c) cluster-specific distributions of visual, audio and overall homogenization indicators visualized using parallel coordinates (Source: authors). (FP = Facade Patina; MV = Material Vernacularity; AC = Architectural Complexity; IO = Interface Organic-ness; SS = Signage Standardization; CS = Chain brand Salience; GD = Generic Decoration; CO = Commercial Openness; LDI = Local Dialect Intensity; CB = Commercial Broadcasting; SED = Sound Event Diversity; TR = Technophony Ratio; TAI = Tourist Activity Impact; PC = Perceived Calmness; VH = Visual Homogenization; AH = Audio Homogenization).

3.3. Key Drivers and Nonlinear Effects Based on SHAP

3.3.1. Model Performance and Spatial Cross-Validation

Under the random training–test split, the XGBoost model achieved a coefficient of determination (R2) of 0.886, a root mean square error (RMSE) of 0.345, and a mean absolute error (MAE) of 0.228 on the test set. The five-fold spatial cross-validation yielded a mean R2 of 0.868 ± 0.057, an RMSE of 0.344 ± 0.074, and an MAE of 0.246 ± 0.038. Compared with the random-split results, spatial cross-validation produced a slightly lower mean R2 and a modestly higher MAE, while the RMSE remained nearly unchanged. These findings indicate that the model maintained broadly comparable predictive performance under the more stringent condition of spatially separated validation (Table 7).
Table 7. Model performance under random train–test split and five-fold spatial cross-validation.
Model performance varied to some extent across the spatial folds, with R2 values ranging from 0.793 to 0.924. Fold 5 produced the lowest R2 and the highest RMSE and MAE, indicating relatively weaker generalizability in that spatial area. Nevertheless, R2 remained positive in every spatial fold, demonstrating that the model retained predictive capacity across all spatial partitions.
The observed values and out-of-fold predictions obtained from the five-fold spatial cross-validation were generally distributed close to the 1:1 reference line. When the out-of-fold predictions from all five spatial folds were pooled, the overall R2 was 0.898 (Figure 10a). This value was calculated from the combined set of all out-of-fold predictions. By contrast, the value of 0.868 ± 0.057 reported in Table 3 represents the mean and standard deviation of the R2 values calculated separately for the five spatial folds. Because these two statistics were derived using different calculation procedures, they should not be compared directly.
Figure 10. Spatial cross-validation and residual diagnosis of the XGBoost model: (a) observed versus predicted values under five-fold spatial cross-validation; (b) spatial folds used for spatial cross-validation; and (c) spatial distribution of prediction residuals.
The spatial-fold configuration showed that the 340 sampling points were divided into relatively contiguous spatial partitions, thereby reducing the likelihood that spatially adjacent samples were simultaneously assigned to the training and test sets (Figure 10b). The spatial distribution of the out-of-fold residuals showed no clear evidence of extensive and continuous clustering of model errors, although some local sampling points exhibited residual variation (Figure 10c). Overall, model performance under spatial cross-validation was broadly comparable to that obtained from the random split, suggesting that the random-split results were not subject to substantial performance inflation. Nevertheless, potential spatial dependence near the boundaries of the spatial folds cannot be entirely ruled out.

3.3.2. Global Importance of Visual and Audio Indicators

The XGBoost model demonstrates good predictive performance for multisensory homogenization. The test-set results show that the model achieved an R2 of 0.886, an RMSE of 0.345, and an MAE of 0.228. The predicted values are generally close to the observed values, and the residuals are mainly distributed around zero, indicating that the model can characterize the level of multisensory homogenization in a relatively stable manner.
In terms of global feature importance, generic decoration (0.276) makes the highest contribution, indicating that replicated, template-based, and landscape-oriented decoration is the most critical visual factor affecting homogenization prediction. Commercial broadcasting (0.227) and the proportion of mechanical sounds (0.186) rank second and third, respectively, suggesting that standardized sound sources, such as outdoor shop broadcasting, promotional announcements, vehicle sounds, and equipment sounds, have an important influence on auditory homogenization. Material locality (0.172), local dialect intensity (0.170), and perceived quietness (0.134) also make relatively high contributions (Figure 11a).
Figure 11. SHAP global feature importance and sample-level heatmap for multisensory homogenization: (a) global contribution of visual and audio indicators to the model output; (b) distribution of SHAP. (Source: authors).
Their relatively high importance indicates strong explanatory power in distinguishing between high- and low-homogenization samples. Combined analysis of the SHAP scatter plot and SHAP heatmap further shows that, as the model output value f(x) increases, generic decoration, commercial broadcasting, and the proportion of mechanical sounds make continuous positive contributions across multiple samples. This suggests that highly homogenized samples are often jointly driven by visual commercialization and sound standardization. In contrast, façade weathering traces, material locality, local dialect intensity, and perceived quietness show negative contributions in some samples, reflecting the inhibitory effects of traditional visual elements and local soundscapes on homogenization prediction (Figure 11b).

3.3.3. Main and Interaction Effects of Multisensory Factors

The results of the SHAP main effects and interaction effects indicate that multisensory homogenization is jointly influenced by visual commercialization, soundscape standardization, and traditional local elements. Generic decoration has the highest node importance and shows strong associations with variables such as commercial broadcasting, the proportion of technical sound sources, façade weathering traces, and material locality. This suggests that visual commercialization interacts with commercial sound sources, technical sound sources, and traditional townscape elements in shaping perceived homogenization (Figure 12a).
Figure 12. SHAP-based main and interaction effects of factors influencing multisensory homogenization: (a) feature importance and interaction network of visual and audio indicators; (b) comparison of main and interaction effects across indicators. (Source: authors).
A further comparison of main and interaction effects shows that most indicators are still dominated by main effects. Among them, generic decoration has the highest main effect (0.308), making it the primary factor influencing the prediction of multisensory homogenization. This indicates that replicated, template-based, and landscape-oriented decorations have the strongest explanatory power for perceived homogenization. Commercial broadcasting (0.227) and the proportion of technical sound sources (0.192) rank next, suggesting that standardized sound sources, such as outdoor shop broadcasting, promotional announcements, vehicle sounds, and equipment sounds, are also important drivers of homogenization. At the same time, indicators such as façade weathering traces, material locality, commercial broadcasting, and the proportion of technical sound sources still make certain interaction contributions, reflecting the complex relationships among traditional visual townscape, commercial symbols, and technologically mediated soundscapes (Figure 12b).

3.3.4. Nonlinear Effects and Threshold Responses

The two-dimensional PDP results reveal the nonlinear effects of key variable combinations on the prediction of multisensory homogenization. When GD reaches a high value, the predicted level of homogenization increases markedly, indicating that template-based decoration is an important factor weakening visual locality. Second, GD × CB and GD × TR show audiovisual superposition effects; that is, when generic decoration, commercial broadcasting, and technical sound sources increase simultaneously, the predicted value of multisensory homogenization further rises. Finally, GD × LDI reflects the tension between commercialized visual interfaces and local sounds. When the intensity of local dialects is relatively high, the tendency toward homogenization is weakened to some extent (Figure A1).
The SHAP dependence plots show that visual indicators have clear nonlinear and threshold effects on the prediction of multisensory homogenization (Figure 13).
Figure 13. SHAP dependence plots and threshold effects of factors influencing multisensory homogenization. Light blue histograms indicate the distribution of feature values, blue points represent SHAP values of samples, pink curves show the smoothed fitting results, shaded bands indicate the 95% confidence intervals, and labeled points denote zero-crossing thresholds of SHAP values (Source: authors).
(1) Inverted U-shaped relationships. FP, MV, AC, and IO all exhibit inverted U-shaped or weak inverted U-shaped relationships, characterized by an initial increase followed by a subsequent decline. FP turns into a positive contributor after approximately 1.7, reaches its peak at a moderate level, and becomes a negative contributor after approximately 6.0. MV, AC, and IO change from positive to negative contributors at approximately 4.7, 3.9, and 5.6, respectively. These results indicate that historical traces, local materials, architectural complexity, and organic interfaces are not sufficient to weaken perceived homogenization when they remain at low or moderate levels. Only when they reach relatively high levels do they clearly reduce the predicted value of multisensory homogenization (Figure 13a–d).
(2) Declining relationships. LDI, SED, and PC exhibit declining relationships and become negative contributors after approximately 5.0, 4.8, and 3.6, respectively. This suggests that local dialects, rich daily sound events, and a higher sense of quietness can weaken perceived soundscape homogenization once they reach a certain level (Figure 13i,k,n).
(3) Increasing relationships. SS, CS, GD, and CO generally show increasing relationships. SS, GD, and CO become positive contributors after approximately 3.9, 4.4, and 3.7, respectively. CS shows a positive effect after approximately 1.1 and continues to increase the model output in the medium- and high-value ranges. This indicates that once shop sign standardization, chain brand presence, generic decoration, and open commercial interfaces exceed certain levels, they significantly strengthen perceived visual homogenization (Figure 13e–h).
To further assess whether these breakpoints had the potential to generalize beyond the observed response patterns, formal segmented regression analyses were conducted. The results provided statistical support for case-specific breakpoints in six visual indicators: traces of historical development on façades, local distinctiveness of materials, architectural complexity, organic character of street interfaces, standardization of storefront signage, and openness of commercial interfaces. For the first four indicators, the slope changed from positive to slightly negative beyond the breakpoint. Storefront signage standardization and commercial interface openness remained positively associated with homogenization beyond their respective breakpoints, but their marginal effects weakened substantially.
None of the auditory indicators exhibited statistically supported breakpoints after false discovery rate (FDR) correction. The identified values should therefore be interpreted as empirical breakpoints specific to the Ximen Street case rather than universally applicable thresholds (Figure A2 and Figure A3 and Table A7).

4. Discussion

This study discusses the findings in relation to the three research questions proposed in the Introduction. First, visual and auditory homogenization in Ximen Street were not synchronized. Soundscape homogenization was generally more pronounced and exhibited stronger patch-like and node-based clustering, indicating that the continuity of the visible townscape does not necessarily ensure the continuity of place-based perception. Second, the six multisensory types revealed distinct combinations of visual character, commercial interfaces, and soundscape locality across different street segments. Third, multisensory homogenization was jointly shaped by commercial visual elements, technological sound sources, the continuity of the historic environment, and the everyday sounds of residents. These influences also exhibited nonlinear and cross-modal interaction effects. The following discussion addresses these findings from three perspectives: audiovisual mismatch, typology-based management, and compound driving mechanisms.

4.1. Mismatch Between Visual Conservation and Soundscape Locality

The current conservation and renewal processes in Ximen Street Historic District reveal a degree of imbalance between the preservation of visual character and the maintenance of soundscape locality. Although elements of the visible historic townscape have been retained in some parts of the district, place-specific auditory characteristics have weakened to a certain extent. This mismatch is particularly evident along commercial streets and peripheral traffic corridors, where preserved façades coexist with commercial broadcasting, traffic noise, and mechanical sounds. The conservation of the visible physical environment is undoubtedly important for sustaining historical continuity. However, if locally distinctive sounds are gradually replaced by standardized commercial soundscapes, technological sound sources, or traffic noise, the preservation of architectural character and spatial form alone is unlikely to sustain the locality of historic districts [35,36,37]. Conservation management should therefore move beyond static townscape improvement toward a form of “dynamic conservation” that integrates both visual and auditory dimensions.
Environmental perception is jointly shaped by visual and auditory information, while the intensity and distribution of different sound categories vary with rhythms of human activity and environmental conditions [32,33,34]. Visual homogenization in Ximen Street exhibited relatively continuous spatial variation. By contrast, soundscape homogenization was more strongly concentrated along commercial streets, peripheral traffic corridors, and activity nodes, and its overall level was higher than that of visual homogenization. Wenchang Street, Qilin South Road, Liaokuo South Road, and Shengfeng Road showed stronger exposure to traffic and technical sounds, whereas internal residential lanes and Tianchi Park retained more localized or relatively calm soundscapes. This divergence reflects the different mechanisms through which the two environmental dimensions change. Building form is constrained by conservation regulations and construction cycles and therefore usually changes slowly. Soundscapes, in contrast, can change rapidly in response to commercial operations, traffic volume, equipment use, and patterns of human activity.
Conservation policies for historic districts should therefore adopt differentiated measures based on the spatial distribution, temporal dynamics, and underlying mechanisms of visual and auditory change. At the visual level, long-term control and incremental repair should be strengthened for façades, materials, and street interfaces. Along the commercial frontages of Ximen Street and Nanmen Street, continuous standardized signage and generic decoration that conceal local materials or historic details should be avoided.
At the auditory level, time-specific monitoring and sound source management should be introduced to regulate commercial broadcasting, mechanical equipment, and traffic disturbance, while maintaining locally distinctive sounds associated with dialect use, neighborhood interaction, and everyday commercial activities. At the audiovisual level, an integrated assessment and renewal review mechanism should be established. Townscape improvement, commercial functions, equipment installation, traffic organization, and public activities should be incorporated into a unified management framework.
The findings further indicate that respondents’ identities and place-based experiences influenced their judgments of homogenizing renewal. For long-term residents, the continuity of place memory was important, but so were improvements in environmental sanitation, housing safety, accessibility, and neighborhood vitality. Street front business operators and some respondents who had previously studied or worked in the district also tended to view the overall environmental improvements positively. Compared with the former conditions of deteriorating facilities, disorderly surroundings, and poor housing quality, they considered a certain degree of visual or soundscape homogenization acceptable when accompanied by broader improvements to the living environment.
These findings suggest that assessments of homogenization are also shaped by lived experience, practical needs, and stakeholder interests. Conservation and renewal policies should therefore balance the continuity of local identity with residents’ legitimate demands for greater safety, convenience, and quality of life (Table A6).

4.2. Multi-Sensory Type Identification and Governance Implications

The identification of multisensory types provides a practical basis for translating research findings into differentiated governance strategies. The Ximen Street Historic District cannot simply be described as either “well preserved” or “highly homogenized”. Some street sections retain traditional visual townscape features but are disturbed by technical sound sources; some show low levels of visual homogenization but have already experienced soundscape standardization; and others simultaneously exhibit high-intensity commercial interfaces, visual homogenization, and soundscape homogenization. These differences indicate that a single façade control strategy is insufficient. For street sections with relatively intact visual heritage but high soundscape homogenization, particularly peripheral roads and major junctions, priority should be given to controlling outdoor broadcasting, technical equipment, and traffic disturbances. Outward-facing loudspeakers should be restricted, broadcasting periods should be regulated, and mechanical equipment should be screened or acoustically treated. For commercial street sections with high levels of both visual and auditory homogenization, especially the commercial frontages of Ximen Street and Nanmen Street, shop signs, generic decorations, open commercial interfaces, and sound-producing equipment should be managed in a coordinated manner. By contrast, residential lanes with low commercial disturbance should be protected as environments in which daily activities, social interactions, and local sound events remain perceptible.
This type-based approach to supporting hierarchical and differentiated governance is consistent with existing studies on historic districts that allocate different renewal strategies according to street type classification [2,40]. It is also grounded in the material form, spatial organization, sociocultural practices, and economic processes of specific places, rather than treating classification results as fixed zones detached from local contexts [4,8,41,42,43]. Therefore, this method can be understood as a governance-oriented typological tool for identifying intervention priorities in different spatial units and for providing a basis for refined renewal at the level of street sections and street front interfaces, rather than as a fixed or universally applicable classification system (Table 8).
Table 8. Governance priorities and feasibility considerations for different multisensory homogenization types (Source: Compiled by the authors).

4.3. Composite Driving Mechanism of Visual and Auditory Homogeneity Perception

In the Ximen Street Historic District, multisensory homogenization primarily results from the cumulative effects of dispersed interventions, including generic decoration, standardized storefront signage, commercial broadcasting, and noise from mechanical equipment. The influence of any single instance of similar decoration, amplified advertising, or equipment operation may be limited. However, when such elements repeatedly occur or become concentrated along particular street segments, they may gradually weaken the district’s original visual and auditory distinctiveness.
At the same time, traces of historical development on façades, local materials, architectural complexity, and the organic character of street interfaces appear to reduce perceived homogenization only when they form a sufficiently continuous environmental pattern. The isolated preservation of traditional elements is unlikely to offset the cumulative effects of commercial standardization. The combined presence of generic decoration, commercial broadcasting, and technological sound sources further intensifies perceived homogenization. By contrast, everyday sounds such as local dialects and residents’ conversations are associated with a lower tendency toward homogenization. The rankings of variable importance, effect magnitudes, and empirical response ranges identified in this study are specific to the Ximen Street sample and should not be directly generalized to other heritage sites.
Nevertheless, the underlying mechanisms may have some transferability to other heritage contexts. Commercial activities and anthropogenic sound sources can gradually reshape the environmental experience and local atmosphere of historic districts through long-term accumulation [30,31]. The maintenance of locality does not depend merely on the presence of individual traditional elements. Rather, it relies on continuous relationships among architectural form, local materials, historical traces, street interfaces, and everyday life [3,4,8].
Visual and auditory information may also jointly contribute to the formation of place perception [28,29,32,33,34]. When repetitive commercial landscapes coincide with broadcasting, traffic, or equipment noise, the perceived standardization of the environment may be intensified. Conversely, local dialects, neighborhood interactions, and other everyday sounds can convey sociocultural information that the physical environment alone cannot fully express. The resulting mechanism—characterized by the accumulation of dispersed interventions, the continuity of heritage cues, and audiovisual synergy—may provide an analytical reference for other historic districts that retain active residential life while facing pressures from commercial renewal.

4.4. Limitations and Future Work

In this study, we integrated street-view imagery, soundscape recordings, AI-assisted scoring, human perception validation, spatial analysis, and explainable machine learning to establish an analytical framework for multisensory homogenization in historic districts. It also identified spatial differences between visual townscape character and soundscape locality, together with their composite driving mechanisms. Nevertheless, the study has several limitations related to the scope of the case study, data conditions, and methodological applicability.
First, the analysis was limited to the Ximen Street Historic District. Spatial cross-validation assessed model stability only across spatially separated samples within the district and cannot substitute for external validation across different regions. The spatial patterns, variable importance rankings, cluster types, and empirical breakpoints identified here are dependent on local conditions, including street morphology, residential activities, commercial structure, the extent of tourism involvement, and soundscape composition. They should therefore not be directly generalized. By comparison, the analytical procedures involving matched street-view and audio sampling, audiovisual mismatch identification, typological classification, and nonlinear interpretation may have greater transferability. Their external validity should be further examined through multi-case comparisons under standardized sampling conditions, cross-regional model testing, and leave-one-region-out cross-validation.
Second, the study primarily reflects daytime perceptual conditions during winter and does not fully capture seasonal variation, day–night transitions, or differences between weekdays and holidays. Soundscapes are highly dynamic and may change rapidly with commercial operations, traffic flows, school schedules, pedestrian activity, and temporary events. Visual environments may likewise be influenced by lighting, weather, commercial displays, and festival decorations. The present findings should therefore not be interpreted as representing a stable long-term perceptual condition. Future research should establish repeated sampling protocols across multiple seasons, time periods, and dates. Continuous acoustic monitoring could also be incorporated to assess the temporal stability of audiovisual homogenization patterns.
Third, the conditions under which street-view images and audio recordings were collected may still affect comparability across sampling points. Although we standardized the equipment, camera height, survey period, and basic weather conditions, automatic exposure, lighting direction, partial occlusion, and unexpected sound events may have influenced some indicators. Future studies could use fixed exposure settings and standardized color charts. Longer recording durations, repeated sampling, and the annotation of anomalous sound events could further improve the consistency and stability of multimodal data.
Fourth, AI-assisted scoring should be regarded only as a structured approximation of perception. It cannot replace residents’ actual perceptions, expert judgments, or heritage value assessments. Pearson correlation primarily reflects correspondence in the patterns of variation between AI-generated and aggregated human ratings, while the ICC reflects their degree of agreement. Neither measure is equivalent to objective accuracy. The scoring results may still be influenced by model training data, prompt design, indicator definitions, model version, input quality, and local cultural context. AI models may also reproduce existing human cognitive biases. Future studies should enhance transparency, reproducibility, and robustness through cross-model evaluation, expert review, larger human validation samples, and the disclosure of non-sensitive prompts, parameters, and code.
Fifth, the samples used for human perception validation and interviews were limited in size. The 30 participants were recruited primarily for an exploratory assessment of human–AI agreement and cannot represent all residents and users of the district. Their evaluations may have been affected by age, disciplinary background, familiarity with the district, ability to understand local dialects, and place attachment. The interviews also relied on a small purposive sample. Respondents’ lived experiences, commercial interests, and recall biases, as well as the researchers’ familiarity with the local context, may have influenced the interpretation of the findings. Future research should include broader representation of long-term residents, older adults, business operators, tourists, and heritage professionals. Soundwalks, focus groups, and participatory mapping could also be used for triangulation.
Sixth, uncertainty remains in the model interpretations, clustering results, and nonlinear breakpoints. SHAP values and partial dependence plots reflect predictive associations rather than causal relationships. The empirical breakpoints are applicable only to the Ximen Street sample, while the selection of K = 6 represents a compromise between statistical performance and management interpretability. Future studies should use larger samples, spatially blocked cross-validation, independent test sets, repeated clustering, and uncertainty analysis to improve result robustness .
Finally, the proposed management recommendations are context-dependent. Their implementation is also constrained by funding, administrative authority, business needs, resident acceptance, and relevant institutional arrangements. Future research could conduct pilot interventions in representative street segments and compare audiovisual environments, resident perceptions, commercial effects, and management costs before and after implementation. Such evaluations would help determine the practical feasibility of the multisensory governance framework.

5. Conclusions

This study used the Ximen Street Historic District in Qujing, Yunnan Province, as a case study and constructed a framework for identifying and interpreting multisensory homogenization in historic districts based on streetscape image–environmental audio matched samples. The framework integrates visual and auditory indicators, manual perceptual validation, spatial analysis, K-means clustering, XGBoost, and SHAP-based interpretation. Compared with traditional townscape conservation, this framework incorporates local sounds into the same evaluation system and identifies six types of homogenized street sections within the district: low commercial disturbance, intensified commercial soundscape, traditional townscape with technical sound disturbance, low visual homogenization with soundscape convergence, high commercial homogenization, and retained traditional townscape with mixed soundscape.
Within the district, visual homogenization shows a relatively continuous distribution, whereas soundscape homogenization is more characterized by patch-like patterns and nodal clustering, with a higher overall level. Visual and auditory cues are preserved, replaced, and superimposed in an interwoven and complex manner within the district. Moreover, multisensory homogenization is not determined by a single visual or auditory factor, but is jointly shaped by the standardization of commercial landscapes, the intrusion of technical sound sources, and the degree to which local cues are retained. Generic decoration, standardized shop signs, and open commercial interfaces increase the replicability of the visual environment, while commercial broadcasting, vehicle sounds, and equipment sounds cover or reorganize residents’ conversations, local dialects, and sounds of daily activities. When visual commercialization and soundscape standardization are superimposed within the same spatial unit, they exert a synergistic amplifying effect on the perception of overall homogenization. Conversely, historical traces, local materials, interface organicity, local dialects, sound event diversity, and perceived quietness jointly constitute an important basis for inhibiting the perception of homogenization.
This paper extends the issue of homogenization in historic districts from the traditional level of visual townscape to the weakening of multisensory locality. It reveals the hidden risk that visual conservation does not necessarily equate to the continuity of perceived locality. For the practice of historic district conservation, future renewal governance should not only focus on townscape conservation, but also emphasize the coordinated management of soundscape order. Future research may further extend to historic districts in different cities, with different degrees of tourism development and at different spatial scales. It may also incorporate long-term monitoring across seasons and day–night periods to enhance the refined conservation of historic districts.

Author Contributions

Conceptualization, Y.Q., Y.H. and D.Y.; methodology, Y.Q.; software, Y.Q.; validation, Y.Q. and D.Y.; formal analysis, Y.Q., D.Y. and Y.H.; investigation, Y.Q. and D.Y.; resources, Y.Q. and D.Y.; data curation, Y.Q.; writing—original draft preparation, Y.Q.; writing—review and editing, Y.Q., D.Y. and Y.H.; visualization, Y.Q.; supervision, D.Y. and Y.H.; project administration, Y.Q. and D.Y.; funding acquisition, D.Y.; translation and language polishing, Y.H. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China, grant number 52478018.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Kunming University of Science and Technology Medical Ethics Committee (approval number: KMUST-MEC-2026-102; date of approval: 23 June 2026).

Data Availability Statement

The data presented in this study are available from the corresponding author upon reasonable request. The data are not publicly available due to privacy restrictions stipulated in the informed consent obtained from the participants.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Table A1. Definitions, measurement dimensions, and VLM prompt logic for visual heritage authenticity indicators (Source: Compiled by the authors).
Table A2. Definitions, measurement dimensions, and VLM prompt logic for commercial landscape standardization indicators (Source: Compiled by the authors).
Table A3. Definitions, measurement dimensions, and ALM prompt logic for auditory homogenization indicators (Source: Compiled by the authors).
Table A4. Model parameters and request settings for visual and auditory scoring.
Table A5. Participant characteristics and place familiarity in the human perception. (Source: Compiled by the authors).
Table A6. Characteristics of interview participants, length of place experience, and summary of interview content (Source: Compiled by the authors).

Appendix B

XGBoost Objective Function

At the t-th iteration, XGBoost adds a new regression tree f t ( x ) to minimize the sum of the loss function and model complexity. The objective function is expressed as
L ( t ) = i = 1 n l y i , y ^ i ( t 1 ) + f t x i + Ω f t Ω f t = γ T + 1 2 λ j = 1 T w j 2
where y i denotes the integrated multisensory homogenization perception score of sample i , y ^ i ( t 1 ) is the predicted value after the ((t − 1))-th iteration, f t x i is the prediction of the newly added tree, T is the number of leaf nodes, w j is the weight of the j -th leaf node, and γ and λ control model complexity and leaf node weights, thereby improving model generalization.
In the model interpretation stage, SHAP was used to quantify the marginal contribution of each variable to the prediction of an individual sample [44]. For feature j , the SHAP value is defined as
ϕ j = S F \ j | S | ! ( M | S | 1 ) ! M ! f S j x S j f S x S
where F is the full feature set, S is any feature subset that does not include feature j , (M) is the total number of features, and f S x S denotes the model prediction using only subset S . A positive SHAP value indicates that the variable increases the predicted multisensory homogenization value, whereas a negative SHAP value indicates that the variable helps reduce perceived homogenization. Based on SHAP global importance, dependence plots, and interaction values, the study further identified the main effects, nonlinear thresholds, and cross-modal interactions among visual commercialization, traditional streetscape preservation, and soundscape locality.

Appendix C

Figure A1. Two-dimensional partial dependence plots of key visual–audio interaction effects (Source: authors).
Figure A2. Segmented regression results and bootstrap confidence intervals for empirical breakpoints of the 14 visual and auditory indicators: (a) façade patina; (b) material vernacularity; (c) architectural complexity; (d) interface organic-ness; (e) signage standardization; (f) chain brand salience; (g) generic decoration; (h) commercial openness; (i) local dialect intensity; (j) commercial broadcasting; (k) sound event diversity; (l) technophony ratio; (m) tourist activity impact; and (n) perceived calmness. Blue points represent sample observations, pink lines indicate segmented regression fits, dashed lines denote estimated breakpoints, and shaded areas show the 95% bootstrap confidence intervals (Source: Authors).
Table A7. Formal Segmented-Regression Tests and Spatial-Block Bootstrap Uncertainty for Empirical Breakpoints of the 14 Visual and Auditory Indicators.
Table A8. Multicollinearity diagnostics of the 14 visual and auditory indicators. (Source: Compiled by the authors).
Figure A3. Spatial-Block Bootstrap Estimates and 95% Confidence Intervals of Empirical Breakpoints for the 14 Visual and Auditory Indicators(Source: authors).
Table A9. Moderate and high pairwise correlations among the 14 indicators. (Source: Compiled by the authors).

References

  1. Graham, B.J. Senses of place, senses of time and heritage. In Senses of Place: Senses of Time; Ashworth, G.J., Graham, B., Eds.; Ashgate Publishing: Aldershot, UK; Burlington, VT, USA, 2005; pp. 3–14. [Google Scholar]
  2. Zhang, R.; Martí Casanovas, M.; Bosch González, M.; Sun, S. Revitalizing heritage: The role of urban morphology in creating public value in China’s historic districts. Land 2024, 13, 1919. [Google Scholar] [CrossRef] [Scilit]
  3. International Council on Monuments and Sites. Charter for the Conservation of Historic Towns and Urban Areas (Washington Charter); ICOMOS: Paris, France, 1987; Available online: https://www.icomos.org/images/DOCUMENTS/Charters/towns_e.pdf (accessed on 12 July 2026).
  4. United Nations Educational, Scientific and Cultural Organization. Recommendation on the Historic Urban Landscape, Including a Glossary of Definitions; UNESCO: Paris, France, 2011; Available online: https://www.unesco.org/en/legal-affairs/recommendation-historic-urban-landscape-including-glossary-definitions (accessed on 12 July 2026).
  5. Pulles, K.; Conti, I.A.M.; de Kleijn, M.B.; Kusters, B.; Rous, T.; Havinga, L.C.; Ikiz Kaya, D. Emerging strategies for regeneration of historic urban sites: A systematic literature review. City Cult. Soc. 2023, 35, 100539. [Google Scholar] [CrossRef] [Scilit]
  6. State Council of the People’s Republic of China. Regulations on the Protection of Famous Historical and Cultural Cities, Towns and Villages; Order No. 524 of the State Council of the People’s Republic of China; State Council of the People’s Republic of China: Beijing, China, 2008; Revised 2017. Available online: https://xzfg.moj.gov.cn/front/law/detail?LawID=212 (accessed on 12 July 2026). (In Chinese)
  7. Standing Committee of the Yunnan Provincial People’s Congress. Regulations of Yunnan Province on the Protection of Famous Historical and Cultural Cities, Towns, Villages and Streets; Announcement No. 65 of the Standing Committee of the Tenth Yunnan Provincial People’s Congress; Standing Committee of the Yunnan Provincial People’s Congress: Kunming, China, 2007; Amended 2012. Available online: https://www.cxs.gov.cn/info/1562/16685.htm (accessed on 12 July 2026). (In Chinese)
  8. Chen, Y.; Wang, Y.-W. Approaches to sustaining people–place bonds in conservation planning: From value-based, living heritage, to the glocal community. Built Herit. 2024, 8, 10. [Google Scholar] [CrossRef] [Scilit]
  9. Kuang, Z.; Zhang, J.; Li, Y.; Fukuda, T. Preserving architectural heritage in urban renewal: A stable diffusion model framework for automated historical facade generation. npj Herit. Sci. 2025, 13, 256. [Google Scholar] [CrossRef] [Scilit]
  10. Xie, K.; Xiong, R.; Bai, Y.; Zhang, M.; Zhang, Y.; Han, W. Traditional architectural heritage conservation and green renovation with eco materials: Design strategy and field practice in cultural Tibetan town. Sustainability 2024, 16, 6834. [Google Scholar] [CrossRef] [Scilit]
  11. Yao, W.; Miao, M.; Ding, Z.; Wu, Y.; Zhan, M. Commercial color impact on traditional heritage features in Suzhou Shiquan historic district. npj Herit. Sci. 2025, 13, 395. [Google Scholar] [CrossRef] [Scilit]
  12. Hu, Y.; Meng, Q.; Li, M.; Yang, D. Enhancing authenticity in historic districts via soundscape design. Herit. Sci. 2024, 12, 396. [Google Scholar] [CrossRef] [Scilit]
  13. Liu, X. The effects of commercialisation on urban heritage in Tianjin: A study of citizens’ livelihood in the Five Avenues (Wudadao) historical district. Built Herit. 2024, 8, 42. [Google Scholar] [CrossRef] [Scilit]
  14. Portella, A. Evaluating the effects of commercial signs on the appearance of historic streetscapes in different countries. J. Res. Archit. Plan. 2007, 6, 10–20. [Google Scholar]
  15. Zhu, Y.; González Martínez, P. Heritage, values and gentrification: The redevelopment of historic areas in China. Int. J. Herit. Stud. 2022, 28, 476–494. [Google Scholar] [CrossRef] [Scilit]
  16. Biljecki, F.; Ito, K. Street view imagery in urban analytics and GIS: A review. Landsc. Urban Plan. 2021, 215, 104217. [Google Scholar] [CrossRef] [Scilit]
  17. Xu, H.; Sun, H.; Wang, L.; Yu, X.; Li, T. Urban architectural style recognition and dataset construction method under deep learning of street view images: A case study of Wuhan. ISPRS Int. J. Geo-Inf. 2023, 12, 264. [Google Scholar] [CrossRef] [Scilit]
  18. Huang, K.; Kang, P.; Zhao, Y. Quantitative research of street interface morphology in urban historic districts: A case study of West Street Historic District, Quanzhou. Herit. Sci. 2024, 12, 226. [Google Scholar] [CrossRef] [Scilit]
  19. Huang, G.; Yu, Y.; Lyu, M.; Sun, D.; Dewancker, B.; Gao, W. Impact of physical features on visual walkability perception in urban commercial streets by using street-view images and deep learning. Buildings 2025, 15, 113. [Google Scholar] [CrossRef] [Scilit]
  20. Zhong, T.; Ye, C.; Wang, Z.; Tang, G.; Zhang, W.; Ye, Y. City-scale mapping of urban façade color using street-view imagery. Remote Sens. 2021, 13, 1591. [Google Scholar] [CrossRef] [Scilit]
  21. Yin, L.; Wang, Z. Measuring visual enclosure for street walkability: Using machine learning algorithms and Google Street View imagery. Appl. Geogr. 2016, 76, 147–153. [Google Scholar] [CrossRef] [Scilit]
  22. Kelly, C.M.; Wilson, J.S.; Baker, E.A.; Miller, D.K.; Schootman, M. Using Google Street View to audit the built environment: Inter-rater reliability results. Ann. Behav. Med. 2013, 45, S108–S112. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Li, Y.; Peng, L.; Wu, C.; Zhang, J. Street View Imagery (SVI) in the built environment: A theoretical and systematic review. Buildings 2022, 12, 1167. [Google Scholar] [CrossRef] [Scilit]
  24. Ye, Y.; Zeng, W.; Shen, Q.; Zhang, X.; Lu, Y. The visual quality of streets: A human-centred continuous measurement based on machine learning algorithms and street view images. Environ. Plan. B Urban Anal. City Sci. 2019, 46, 1439–1457. [Google Scholar] [CrossRef] [Scilit]
  25. Wang, Z.; Zhang, W.; Huang, Y. Nonlinear perceptual thresholds and trade-offs of visual environment in historic districts: Evidence from street view images in Shanghai. Sustainability 2025, 17, 11075. [Google Scholar] [CrossRef] [Scilit]
  26. Yu, M.; Chen, X.; Zheng, X.; Cui, W.; Ji, Q.; Xing, H. Evaluation of spatial visual perception of streets based on deep learning and spatial syntax. Sci. Rep. 2025, 15, 18439. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Huang, L.; Kang, J. The sound environment and soundscape preservation in historic city centres—The case study of Lhasa. Environ. Plan. B Plan. Des. 2015, 42, 652–674. [Google Scholar] [CrossRef] [Scilit]
  28. Liu, J.; Yang, L.; Xiong, Y.; Yang, Y. Effects of soundscape perception on visiting experience in a renovated historical block. Build. Environ. 2019, 165, 106375. [Google Scholar] [CrossRef] [Scilit]
  29. Ye, J.; Li, S.; Chen, Y.; Ma, Y.; Chen, L.; He, T.; Zheng, Y. A study of the effects of historical block context on soundscape perception. Buildings 2024, 14, 621. [Google Scholar] [CrossRef] [Scilit]
  30. Truax, B. Acoustic Communication, 2nd ed.; Ablex Publishing: Westport, CT, USA, 2001. [Google Scholar]
  31. Standing Committee of the National People’s Congress of the People’s Republic of China. Law of the People’s Republic of China on Noise Pollution Prevention; Order No. 104 of the President of the People’s Republic of China; Standing Committee of the National People’s Congress of the People’s Republic of China: Beijing, China, 2021; Effective 5 June 2022. (In Chinese)
  32. GB 3096-2008; Environmental Quality Standard for Noise. China Environmental Science Press: Beijing, China, 2008. (In Chinese)
  33. Gan, Y.; Luo, T.; Breitung, W.; Kang, J.; Zhang, T. Multi-sensory landscape assessment: The contribution of acoustic perception to landscape evaluation. J. Acoust. Soc. Am. 2014, 136, 3200–3210. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Jeon, J.Y.; Jo, H.I. Effects of audio-visual interactions on soundscape and landscape perception and their influence on satisfaction with the urban environment. Build. Environ. 2020, 169, 106544. [Google Scholar] [CrossRef] [Scilit]
  35. Relph, E. Place and Placelessness; Pion: London, UK, 1976. [Google Scholar]
  36. Norberg-Schulz, C. Genius Loci: Towards a Phenomenology of Architecture; Rizzoli: New York, NY, USA, 1980. [Google Scholar]
  37. Baudrillard, J. The Consumer Society: Myths and Structures; Sage: London, UK, 1998. [Google Scholar]
  38. Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef] [Scilit]
  39. Chen, T.; Guestrin, C. XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar] [CrossRef] [Scilit]
  40. Yan, Y.; Xu, X.; Huang, Y.; Zhang, Q. Micro-scale land-use functional patterns and driving mechanisms in historic districts: A multi-source GIS and interpretable machine learning approach. Front. Environ. Sci. 2026, 14, 1784087. [Google Scholar] [CrossRef] [Scilit]
  41. Huang, Y.; Shi, Z.; Chen, Y.; Zhou, K.; Ying, Z.; Zheng, L.; Yang, S. Urban form evolution and transformation of traditional water towns based on space syntax and GIS: Evidence from ancient Wenzhou city (16th to the 20th century). Front. Earth Sci. 2025, 13, 1520643. [Google Scholar] [CrossRef] [Scilit]
  42. Huang, Y.; Huang, Y.; Chen, Y.; Song, J.; Yang, S.; Huang, L.; Zheng, L.; Gao, Y. The evolution and construction of Shan-shui cities: Evidence from the ancient city of Hangzhou from the sixth to the twenty-first century via geographical information systems and space syntax. Front. Earth Sci. 2025, 13, 1551117. [Google Scholar] [CrossRef] [Scilit]
  43. Zhu, Y.; Huang, Y. HGIS-based analysis of urban morphological evolution in historic Kaifeng. npj Herit. Sci. 2026, 14, 32. [Google Scholar] [CrossRef] [Scilit]
  44. Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30; Guyon, I., von Luxburg, U., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2017; pp. 4765–4774. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.