Next Article in Journal
Uniaxial Compressive Behavior and Constitutive Modeling of Fiber-Reinforced Self-Compacting Concrete with Granite Powder and Expansive Agent: An Experimental Study with Acoustic Emission Monitoring
Next Article in Special Issue
Sustainable Revitalization of a Heritage Site: Community-Based Conservation and Adaptive Reuse of Wuxi’s Sanliqiao Catholic Church
Previous Article in Journal
Early-Stage (10-Cycle) Freeze–Thaw Damage Sensitivity and Multi-Metric Conservation Assessment of Historic Blue Bricks from Beijing
Previous Article in Special Issue
Spatiotemporal Distribution Characteristics and Influencing Factors of Historic Buildings in the Mount Tai Region: Implications for Tourism Planning
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

When Visual Conservation Meets Auditory Homogenization: Multimodal Evidence from Ximen Street Historic District, Qujing, China

1
Faculty of Architecture and City Planning, Kunming University of Science and Technology, Kunming 650500, China
2
Institute of Urban and Sustainable Development, City University of Macau, Avenida Padre Tomás Pereira, Taipa, Macau 999078, China
*
Authors to whom correspondence should be addressed.
Buildings 2026, 16(14), 2871; https://doi.org/10.3390/buildings16142871
Submission received: 25 June 2026 / Revised: 16 July 2026 / Accepted: 17 July 2026 / Published: 19 July 2026
(This article belongs to the Special Issue Built Heritage Conservation in the Twenty-First Century: 3rd Edition)

Abstract

The conservation of historic districts is often evaluated on the basis of visible material features, yet commercialized renewal may weaken visual and auditory cues of locality while preserving building façades and street textures. We established 340 sampling points at 20 m intervals along Ximen Street Historic District in Qujing, Yunnan Province, and constructed a streetscape–audio matched sample dataset. Through multimodal model-assisted recognition, K-means clustering, XGBoost, and SHAP methods, we identified the degree of visual and auditory homogenization and analyzed the nonlinear effects and interactions of visual and auditory factors on the perception of homogenization. The results show that (1) visual homogenization presents a continuous distribution, while soundscape homogenization presents a patch-like distribution and is higher overall; (2) commercial symbolism, generic decoration, and open interfaces lead to visual substitutability, while commercial broadcasting and mechanical sounds intensify soundscape standardization; dialects, residents’ conversations, and diverse sound events enhance the perception of locality; and (3) visual commercialization and soundscape standardization have a clear interaction effect. This case study provides insights into the living conservation and renewal of historic districts in similar contexts and may offer a fine-grained analytical reference for collaborative governance that considers both visible townscape features and daily soundscape locality.

1. Introduction

1.1. Realistic Dilemmas in the Conservation of Historic Districts and the Continuation of Local Identity

Historic districts serve as important spatial carriers of urban historical memory and local cultural identity [1]. Against the backdrop of ongoing urban regeneration, they have become key spatial units for balancing heritage conservation, spatial vitality, and public value [2]. At the international level, the 1987 Washington Charter expanded the scope of conservation from individual buildings to historic towns, urban areas, and their natural and human-made environments. It also emphasized that the conservation of historic urban areas should be integrated into urban planning while accommodating residents’ daily lives and broader social development [3]. In 2011, UNESCO’s Recommendation on the Historic Urban Landscape further proposed that the built environment, spatial patterns, natural conditions, sociocultural practices, and intangible values should be addressed in an integrated manner within a broader urban context. This marked a shift in historic urban conservation from the preservation of physical remains toward a more holistic and living approach [4]. Accordingly, the conservation of historic districts should extend beyond the restoration and preservation of physical heritage. Greater attention should also be given to the role of residents’ everyday lives, social interactions, and community participation in sustaining local values [4,5].
In China, the Regulations on the Protection of Famous Historical and Cultural Cities, Towns, and Villages emphasize the integrated protection of traditional spatial patterns, historic character, and urban scale at the national level. They also require that the authenticity and integrity of historical and cultural heritage be maintained [6]. At the provincial level, the Regulations of Yunnan Province on the Protection of Famous Historical and Cultural Cities, Towns, Villages, and Streets provide a legal basis for the conservation, management, and appropriate use of historic and cultural districts in Yunnan [7]. Achieving a balance among the continuity of historical culture, the improvement of living environments, and the enhancement of local vitality has therefore become a critical issue for the sustainable development of historic districts.
Visible material elements constitute an important basis for assessing conservation effectiveness [6,7]. However, the continuity of locality is also shaped by the social connections, collective memories, and daily experiences formed between residents and places [8]. Existing approaches to the identification, evaluation, and management of historic district conservation and renewal are usually based on material-spatial elements, such as historic buildings and their façades [9], materials and construction characteristics [10], and district patterns and their relationships with open spaces [3]. They also maintain overall character through the coordinated control of townscape elements, including architectural colors, shop signs, decorations, and street interfaces [11]. By contrast, relatively limited attention has been paid to perceptual elements, such as sound and daily activities, in the evaluation of authenticity in historic districts [12].
In the renewal process, the introduction of commercial functions, tourism consumption, and streetscape improvement may bring financial investment and spatial vitality, but they may also reshape residents’ livelihoods, community life, and the structure of daily activities [13]. The continued introduction of replicated visual elements, such as standardized commercial colors, signs, and decorations, may weaken the original visual characteristics and local recognizability of historic districts [14]. When the shaping of commercial space becomes further detached from residents’ daily lives, residents’ sense of place, sense of belonging, and local identity may also be affected [15]. Accordingly, the homogenization of historic districts examined in this paper refers to the process in which place-specific spatial forms, environmental characteristics, and daily life cues are weakened or replaced under the continued intervention of standardized and replicable renewal elements. Whether the conservation of visible townscape features alone can adequately identify the continuity of perceived locality in historic districts remains to be further examined.

1.2. Insufficient Synergy in Visual–Auditory Perception Assessment for Historic District Conservation

Streetscape imagery has become an important data source for research on the urban built environment [16]. It has produced significant achievements in identifying façade style characteristics [17], street interface forms [18], intensity of commercial signage [19], distribution of architectural colors [20], and degree of spatial openness [21].
Compared with traditional field surveys, streetscape imagery can assess the street environment from a perspective close to that of pedestrians, thereby effectively improving the spatial precision and comparability of urban environmental measurements [22,23]. With the development of computer vision and multimodal artificial intelligence methods, streetscape studies have expanded from the identification of physical elements, such as buildings, roads, and vegetation, to the evaluation of semantic dimensions related to human-centered experience, including street quality [24], visual order [25], and environmental atmosphere [26]. These methods can provide important technical support for identifying visual perception in historic districts.
In recent years, soundscape research in historical and heritage contexts has primarily focused on the acoustic characteristics of historic urban areas [27], soundscape perception and visitor experience [28,29], and the cultural meanings and place-based representations of distinctive sound sources [30]. Acoustic communication theory suggests that particular sounds can acquire place-specific meanings through long-term social interaction and become important auditory elements representing local identity [30]. However, this line of research has largely emphasized the interpretation of the cultural meanings of sound and has rarely translated locally distinctive sounds into spatially explicit and measurable conservation approaches.
Previous studies have used sound source surveys, sound pressure level measurements, and time–frequency analyses to identify the composition of acoustic environments in historic urban areas. They have also examined the conservation value of characteristic sounds with local significance [27]. Nevertheless, these studies have mainly concentrated on acoustic properties and the spatial distribution of sound sources, making it difficult to capture how different sounds influence the perception of historical atmosphere and sense of place. Building on this work, other studies have investigated the effects of soundscape perception on visitor experience and the ways in which different spatial contexts within historic districts shape soundscape evaluations [28,29]. However, most of this research has centered on tourist experiences, with limited attention paid to residents’ everyday soundscapes.
At the regulatory level, the Law of the People’s Republic of China on Noise Pollution Prevention and Control and the Environmental Quality Standard for Noise address harmful noise through emission control, acoustic environment functional zoning, and environmental quality limits [31,32]. However, clear provisions and operational evaluation criteria for conserving soundscape environments with local cultural significance remain lacking. Overall, existing research has yet to establish an integrated framework that combines residents’ everyday soundscapes, place-based perception, and spatial measurement. It has also not fully revealed the mechanisms through which sound and the visible environment jointly shape historical atmosphere and local identity. However, human environmental perception is characterized by multisensory integration. Vision and hearing do not act independently in environmental evaluation; rather, they mutually regulate one another in processes of attention, emotion, and judgment of meaning, jointly shaping people’s comprehensive experience of environmental quality, atmosphere, and behavioral suitability [33,34]. Vision primarily provides visible information such as architectural form, material texture, street scale, interface order, and commercial landscape, while sound conveys dynamic information related to crowd activities, modes of social interaction, functional use, and environmental rhythm [35,36,37]. The interaction between the two further influences individuals’ judgments of place perception. Therefore, a fragmented evaluation of visual and auditory dimensions may equate the preservation of visible townscape features with the continuity of locality, while overlooking situations in which traditional forms are retained but daily soundscapes change. It is therefore necessary to integrate both dimensions within the same spatial unit for evaluation.

1.3. Research Objectives

Previous studies have provided an important foundation for the conservation and renewal of historic districts, yet several limitations remain. First, current evaluations of historic district conservation and renewal largely rely on visible spatial elements to assess conservation effectiveness, while auditory cues are considered less frequently. As a result, they are unable to fully identify the extent to which locality is sustained in daily perception. Second, visual and auditory studies are usually conducted separately, with insufficient attention paid to their coordinated representation and spatial correspondence within the same place-based context. Moreover, nonlinear and interaction relationships may exist between visual and auditory elements in historic districts, whereas existing single-dimensional or linear analyses are still limited in revealing their combined mechanisms of influence on the perception of homogenization.
We collected streetscape images and field recordings from Ximen Street Historic District in Qujing, Yunnan Province, to use as a case study. After recognition using LLMs and VLMs and validation through manual perceptual evaluation, a visual–soundscape matched dataset was established at the sampling-point scale. An indicator system was then constructed from three dimensions: heritage authenticity, landscape standardization, and soundscape locality. By further integrating spatial distribution analysis, K-means clustering, the XGBoost model, and SHAP-based interpretation, we identified the distribution and types of visual and auditory homogenization, and revealed the mechanisms through which different visual and auditory variables affect the perception of sensory homogenization. The study addresses the following questions:
Q1: What are the degrees of visual and auditory homogenization in the Ximen Street Historic District? What spatial differences do they exhibit, and how are they associated with street hierarchy and spatial functions?
Q2: What types of multisensory homogenization can be identified in the Ximen Street Historic District? How do these types differ in terms of visual townscape, commercial interface, and soundscape locality?
Q3: Which visual and soundscape variables are important predictors of multisensory homogenization? Do these variables exhibit meaningful thresholds or interaction effects?
This study contributes to the literature in three main ways. Theoretically, this study extends the discussion of homogenization in historic districts from the traditional focus on visual townscape standardization to the integrated perceptual dimensions of vision and hearing. It emphasizes that the acoustic environment and visual townscape jointly participate in the construction of heritage identity, local authenticity, and daily sense of place. Methodologically, the study provides an operational pathway for measuring perceived homogenization within historic districts. Practically, the findings offer a reference for refined governance in the conservation and renewal of historic districts of a similar type.

2. Study Area, Datasets, and Research Methods

2.1. Study Area

The Ximen Street Historic District is located in Qilin District, Qujing City (25.49° N, 103.80° E). It extends north to Wenchang Street, south to Shengfeng Road, west to Liaokuo South Road, and east to Qilin South Road, covering a total area of approximately 24 ha. It is not only the core space of the spatial development of the ancient city of Qujing, but also the only well-preserved historic district remaining in Qilin District. The district has formed a relatively stable spatial framework around Qujing. It has largely maintained the nested “street–lane–courtyard” spatial pattern, and preserved the traditional residential style of the “Yikeyin” courtyard house. It is therefore a historical epitome of urban development in eastern Yunnan and an important material carrier of the cultural heritage of Cuan culture (Figure 1).
Compared with more intensively developed and commercially oriented historic urban areas in Yunnan Province, such as the Old Town of Lijiang and the Ancient City of Dali, the Ximen Street Historic District in Qujing retains stronger characteristics of everyday life and neighborhood interaction. It therefore represents a type of historic district in Yunnan’s small- and medium-sized cities that has not undergone intensive tourism development (Table 1).
Since the launch of the urban regeneration “micro-renewal” initiative in 2022, the Ximen Street Historic District in Qujing has gradually implemented a series of interventions, including heritage conservation, infrastructure improvement, public space enhancement, commercial revitalization, and the upgrading of street interfaces. The district therefore exhibits a representative pattern in which heritage conservation, everyday life, and incremental renewal coexist and interact. Within the broader context of urban regeneration in China, this type of historic district is highly representative of current conservation and renewal practices. It provides an appropriate empirical setting for examining the relationships among heritage conservation, commercial activities, and everyday environments, as well as for exploring the assessment of multisensory homogenization in historic districts.

2.2. Datasets

2.2.1. Streetscape Data Collection

Before streetscape data collection, the road network of the study area was first extracted from OSM, and the study boundary and existing street–lane network were examined in GIS. Main streets, secondary streets, internal lanes, and daily living spaces within the Ximen Street Historic District were then identified. Sampling points were arranged along the centerlines of streets and lanes at intervals of approximately 20 m, resulting in a total of 340 sampling points.
Streetscape images were collected from 9:00 to 18:30 on 9–10 December 2025, using an iPhone 15 Pro Max. During the collection period, the weather was clear, with a light breeze of approximately Level 1–2 (Figure 2). The sampling route covered urban arterial roads, secondary roads, core internal streets, and narrow lanes within the district to reflect differences among spatial types and street interfaces. During image acquisition, the shooting height was controlled at approximately 1.7 m to simulate pedestrians’ daily visual perception. To ensure sample comparability, approximately 180° panoramic images were collected at each sampling point along the main axis of the street or lane and in the opposite direction, together covering an approximately 360° field of view. For intersections and open spaces, the shooting orientation was determined according to the main directions of pedestrian activity and the principal visible street interfaces.
After field collection, a total of 701 original streetscape images were obtained. All images were manually screened for quality, and images with obvious occlusion, blurring, severe exposure problems, duplicated content, or insufficient representation of street-space characteristics were removed or replaced. For each sampling point, two images from adjacent viewing directions were selected and stitched to form a 360° streetscape image covering the street environment. Finally, 340 valid field streetscape images were obtained (Figure 2).

2.2.2. Sound Data Collection

Soundscape data were collected simultaneously with streetscape images to achieve one-to-one image–audio matching at the same sampling points. Recordings were mainly made using the built-in microphone of an iPhone 15 Pro Max, supplemented by an Apple Watch Series 10 to observe changes in the on-site environmental sound level. All recordings were conducted using the same device, recording mode, and operating procedure, and the original audio files were saved in M4A format. During data processing, the original audio files were uniformly converted to wav format with a sampling rate of 48 kHz and 16-bit PCM encoding. Before acoustic feature extraction, the audio files were further converted to mono to ensure comparability among samples.
Soundscape sampling took place at the same 340 sampling points used for streetscape image collection, with an interval of approximately 20 m between points. Each audio segment corresponded to one valid streetscape image at the same location, forming an image–audio matched dataset. During the field survey, the “Two Steps” app was used to record route and time information in order to verify the sampling route, collection sequence, and point–location correspondence.
The researchers walked throughout the sound recording survey and avoided using bicycles, electric bicycles, or other transport modes that could introduce non-environmental sound interference. After arriving at each sampling point, the researchers stopped walking, kept the device stable, and recorded toward the main axis of the street or lane to capture an acoustic environment close to daily pedestrian experience. Each recording lasted approximately 20–30 s.
After field collection, 383 audio records were obtained. Through manual screening, samples with obvious sudden interference were removed, and 340 valid soundscape samples were ultimately retained and matched one-to-one with the 340 valid streetscape images. Daily traffic sounds, residents’ conversations, commercial activity sounds, and tourist-related sounds were retained as components of the district soundscape, while only occasional, sudden, and unrepresentative interfering sounds were removed.

2.3. Methods

2.3.1. Research Framework

We constructed a homogenization assessment framework for historic districts that integrates multisource data and multimodal models. The framework consists of five main steps. First, streetscape images, soundscape audio, and building–road network vector data corresponding spatial-location data were collected to construct a multimodal base database. After data cleaning and manual screening, 340 valid streetscape images and 340 matched audio samples were retained to construct the multimodal database. Second, images and audio were spatially matched, and basic features were extracted. At the visual level, semantic elements such as buildings, sky, and vegetation were quantified. At the auditory level, acoustic features such as pitch, zero-crossing rate, and MFCCs were extracted. Third, visual–language models (VLMs) and audio–language models (ALMs) were used to construct high-dimensional homogenization indicators. These models were applied to identify audiovisual features, including façade historical traces, material locality, signage standardization, dialect intensity, and the proportion of commercial broadcasting. The reliability of the scores was then validated through a manual perceptual experiment. Fourth, spatial differentiation analysis, multisensory homogenization type identification, and SHAP-based interpretation were conducted to reveal the spatial patterns, dominant drivers, and nonlinear effects of perceived homogenization within the district. Finally, the analytical findings were translated into differentiated conservation and management recommendations, including type-specific responses, control of commercial and technical interventions, maintenance of everyday soundscapes, and preservation of local identity (Figure 3).

2.3.2. Construction of Independent and Dependent Variables

To explain why perceptions of homogenization emerge in historic districts, we organized the independent variables into two modal layers, visual and auditory, based on the logic of “locality retention–standardized replacement.” These variables are further divided into four secondary dimensions: weakening of visual authenticity, standardization of commercial landscapes, weakening of soundscape locality, and intrusion of commercial–technical sound sources. The indicator system includes eight visual variables and six auditory variables, comprising a total of 14 independent variables (Table 2). To improve the discrimination of subtle audiovisual differences among samples and reduce tied scores, we adopted a 0–10 rating scale for use in the AI-assisted assessment. The original 0–10 scores were used for spatial analysis, cluster analysis, and machine-learning modeling. They were converted to a 0–5 scale during the human–AI agreement analysis, so as to ensure consistency with the scale used in the human evaluation.
The weakening of visual authenticity was used to characterize the loss of place-specific material cues in historic buildings and street interfaces. It comprised four indicators: traces of historical development on façades, local distinctiveness of materials, complexity of architectural details, and organic character of street interfaces. Commercial landscape standardization was used to capture the replacement of everyday-life settings by consumption-oriented symbols. It was assessed using four indicators: standardization of storefront signage, prominence of chain brands, generic decoration, and openness of commercial interfaces.
For the auditory dimension, the weakening of soundscape locality was represented by the intensity of local dialect use and the diversity of sound events. Commercial and technological sound intrusion was assessed using four indicators: intensity of commercial broadcasting, influence of tourism activities, proportion of technological sound sources, and perceived quietness.

2.3.3. LLM- and ALM-Assisted Scoring

A total of 16 AI-assisted assessment items were established, comprising eight primary visual perception indicators, six primary auditory perception indicators, one overall visual homogenization indicator, and one overall auditory homogenization indicator. The visual and auditory indicators were evaluated using a visual language model (VLM) and an audio language model (ALM), respectively. The models initially generated scores on a 0–10 scale, which were subsequently converted to a 0–5 scale to ensure consistency with the human perception assessment and subsequent statistical analyses. Following the predefined scoring criteria, the models made judgments solely on the basis of information visible in the images or audible in the recordings. They were not permitted to infer information beyond the provided samples. The complete prompts and model parameter settings are presented in Table A1, Table A2, Table A3 and Table A4.
The street-view image and audio recording corresponding to each sampling point were treated as independent units of analysis. Both visual and auditory assessments were conducted using the GPT-5.4 model, which was accessed on 7 April 2026. All street-view images were stored in RGB JPEG format. Their widths were standardized to 1024 pixels while preserving the original aspect ratios. No further resizing or compression was performed during the scoring process; instead, the processed images were directly encoded and submitted to the model. The audio recordings were stored in M4A/AAC format with a sampling rate of 48 kHz and a single audio channel.
Both visual and auditory samples were processed through separate model calls for each individual sample. For the visual model, the temperature was set to 0.1, the maximum output length to 3000 tokens, the top-p value to 0.8, the top-k value to 40, and the request timeout to 90 s. The program automatically retried a request up to five times in cases of empty responses, missing fields, out-of-range scores, JSON-parsing failures, request timeouts, or other application programming interface errors. The type of error and the number of retry attempts were recorded. Samples that still failed to produce valid results after five automatic retries were manually checked by the researchers for problems with the original files or request outputs and were then resubmitted to the model. Invalid outputs and placeholder values were excluded from subsequent statistical analyses.
The pilot experiment showed that processing multiple sampling points within a single request frequently resulted in missing fields, null values, or JSON-parsing failures. The failure rate was substantially higher for images than for audio recordings. After adopting a sample-by-sample processing strategy, 337 of the 340 candidate samples produced complete and parsable structured outputs on the first attempt, corresponding to an initial valid generation rate of 99.14%. The remaining three samples were successfully processed after manual inspection and resubmission, resulting in a final valid generation rate of 100%. Following image quality control and one-to-one image–audio matching, valid recognition data were obtained for all 340 samples.
Finally, the model returned the sample identifier, individual indicator scores, and a brief justification for each judgment in strict JSON format. The program then automatically performed JSON parsing, required-field verification, score-range validation, and CSV export. The complete implementation code is available from the corresponding author upon reasonable request.

2.3.4. Human Perception Validation and Agreement Assessment

A human perception validation experiment was conducted using matched image–audio samples. Its sole purpose was to assess the agreement between AI-assisted scores and human perceptual judgments. Thirty matched image–audio sample sets were selected from the 340 sampling points. These samples represented spatial locations with pronounced differences across the district, thereby ensuring that the validation set captured representative audiovisual variation (Figure 4).
Given the exploratory purpose of the validation, 30 adult participants were recruited to complete the human assessment. Participant characteristics are reported in Table A5. A total of 900 valid evaluation records were obtained. The participant group provided basic coverage in terms of gender, age, disciplinary background, residential experience, and familiarity with the district. Each participant was required to evaluate all 30 matched image–audio sample sets, with each set involving separate visual and auditory perceptual judgments. Because this procedure imposed a substantial assessment burden, the study adopted a design comprising 30 participants and 30 validation samples to balance evaluation feasibility, perceptual diversity, and rating stability (Table A5 and Table A6).
During the experiment, participants completed the assessment in a relatively quiet environment while wearing headphones. They viewed each street-view image or listened to a 20–30 s audio recording before assigning a score. The samples were presented in randomized order, and all assessment items were rated using a 0–5 Likert scale (Figure 4).
The validation results indicated a high level of agreement between the AI-assisted scores and the mean human ratings. The overall Pearson correlation coefficient was r = 0.79 (p < 0.01). The corresponding coefficients were r = 0.82 for the visual dimension (p < 0.01), r = 0.77 for the auditory dimension (p < 0.01), and r = 0.80 for the overall homogenization assessment (p < 0.01). Pearson’s correlation coefficient primarily reflects the extent to which AI-generated scores and human perceptual judgments follow similar patterns of variation.
Further agreement analysis produced an overall intraclass correlation coefficient (ICC) of 0.81. The ICC values were 0.84 for the visual dimension and 0.78 for the auditory dimension. These results indicate good agreement between the AI-generated scores and the aggregated human ratings, with stronger agreement for the visual dimension than for the auditory dimension. Cronbach’s alpha was 0.87, indicating good internal consistency among the human raters. The AI-assisted scores were therefore used as a source of structured perceptual data in the subsequent spatial analyses and statistical modeling.
These validation results support the use of AI-assisted scores as structured approximations of human perception in subsequent spatial analyses and statistical modeling. However, such scores should not be regarded as objective substitutes for heritage value assessments, residents’ actual perceptions, or expert judgments. AI-generated scores may be influenced by the model’s training data, prompt design, image and audio quality, indicator definitions, and the local context of the case study area. Human evaluations may likewise be affected by participants’ age, disciplinary background, familiarity with the local area, ability to understand local dialects, and personal esthetic experience.

2.3.5. Model Design

First, we compared eight regression models: multiple linear regression (MLR), Ridge regression, support vector regression (SVR), Random Forest, Extra Trees, XGBoost, LightGBM, and CatBoost. All models were evaluated using the same 80/20 training–test split and five-fold cross-validation. Their predictive performance was assessed using R2, RMSE, and MAE. The results showed that the nonlinear models generally outperformed the linear models. Although CatBoost achieved the best overall predictive metrics, XGBoost also performed well on both the held-out test set and cross-validation, and provided a favorable balance between predictive accuracy and generalizability. Because the primary objective of this study was to identify variable importance, nonlinear response patterns, and audiovisual interaction effects, XGBoost was ultimately selected as the principal explanatory model. Its compatibility with SHAP analysis also enabled the development of a unified and reproducible model interpretation framework (Table 3 and Table 4).
Before model fitting, multicollinearity among the 14 predictor variables was assessed using variance inflation factors (VIFs) and tolerance statistics. The VIF values ranged from 1.434 to 4.396, all below the commonly adopted threshold of 5, indicating no serious multicollinearity among the predictors (Table A8).
To identify the combined effects of visual and auditory environmental factors on perceived multisensory homogenization in the historic district, we developed an XGBoost-based machine-learning framework integrating predictive modeling, hyperparameter optimization, and interpretability analysis [38,39]. The model used 340 image–audio matched samples as analytical units, with each sample corresponding to one street-view image, one environmental audio clip, and one spatial location.
The independent variables consisted of 14 multisensory environmental indicators, including eight visual variables and six auditory variables. The visual variables included facade patina, material vernacularity, architectural complexity, street interface organic-ness, signage standardization, chain brand salience, generic decoration, and commercial interface openness. The auditory variables included local dialect intensity, commercial broadcasting intensity, sound event diversity, technophony ratio, tourist activity impact, and perceived calmness. To ensure directional consistency, all variables were processed so that higher values indicated a greater risk of multisensory homogenization. The original scoring scales were retained because XGBoost is insensitive to feature scaling and because preserving the original scale facilitates the identification of practically interpretable thresholds.
The dependent variable was the integrated multisensory homogenization perception score of each image–audio matched sample. This score reflected participants’ overall judgment of visual commercialization, soundscape standardization, and environmental interchangeability. It was measured on a 1–5 scale, with higher values indicating stronger perceived multisensory homogenization. Visual homogenization scores, soundscape homogenization scores, and their difference were only used for subsequent spatial mismatch analysis and were not included as model inputs, thereby avoiding dependent variable information leakage.
XGBoost was employed because it can capture nonlinear relationships and interaction effects among multisensory environmental variables while incorporating regularization mechanisms to reduce overfitting. This is particularly suitable for examining the potentially complex relationships among visual authenticity loss, commercial landscape standardization, change in soundscape locality, and integrated homogenization perception.
All samples were randomly divided into training and testing sets at an 8:2 ratio, yielding 272 training samples and 68 testing samples. Hyperparameter tuning was conducted only on the training set. Given the relatively limited sample size of this single-case study, five-fold cross-validation was implemented within the training set to evaluate model stability and identify the optimal parameter combination. The independent testing set was subsequently used to assess the generalization performance of the final model. Model performance was evaluated using the coefficient of determination (R2), root mean square error (RMSE), and mean absolute error (MAE). Higher R2 values and lower RMSE and MAE values indicate better predictive performance (Equations (A1) and (A2)).
Hyperparameter optimization was conducted using grid search combined with five-fold cross-validation. The search ranges were 100–800 for ‘n_estimators’, 0.01–0.10 for ‘learning_rate’, 2–6 for ‘max_depth’, 1–5 for ‘min_child_weight’, 0.6–1.0 for both ‘subsample’ and ‘colsample_bytree’, 0–0.5 for ‘gamma’, 0–1 for ‘reg_alpha’, and 0.5–5 for ‘reg_lambda’. During model training, the validation subset within each cross-validation fold was used as the early-stopping monitoring set, and training was stopped when no performance improvement occurred for 50 consecutive rounds. This procedure reduced the risk of overfitting caused by excessive tree growth.
The final XGBoost model achieved an R2 of 0.886, an RMSE of 0.345, and an MAE of 0.228 on the independent testing set, indicating good predictive performance for integrated multisensory homogenization perception.

3. Results

3.1. Spatial Differentiation of Audio-Visual Indicators

3.1.1. Spatial Differentiation of Visual Homogenization Indicators

The semantic composition of the district streetscape is dominated by buildings, vegetation, and roads, while sky and sidewalks account for relatively low proportions, and the remaining elements make up only small shares overall (Figure 5a). The indicators of visual homogenization show clear spatial differences among sampling points (Figure 5(b1–b4)).
Façade aging traces (FP) and material locality (MV) are generally at relatively high levels. High-value combinations are mainly concentrated in the southeastern part of the district and in traditional streets and lanes such as Ximen Street and Nanmen Street, where the original architectural townscape and local materials are relatively well preserved (Figure 5(b1)). High-value combinations of architectural complexity (AC) and interface organicity (IO) are mainly distributed in the southern part of the site and around Tianchi Park, where the street interface morphology and visual organization are relatively organic (Figure 5(b2)). Shop sign standardization (SS) and chain brand prominence (CS) are generally at low levels, with only localized high values appearing along some street-front commercial interfaces. Existing commercial activities are mainly composed of small local shops and traditional business types, while the intervention of chain based and standardized commercial signage remains relatively limited (Figure 5(b3)). High-value combinations of generic decoration (GD) and commercial interface openness (CO) are mainly concentrated along the periphery of the district, especially along urban arterial roads and at their intersections. In these areas, uniformly installed shop signs and open commercial interfaces are relatively concentrated (Figure 5(b4)).

3.1.2. Spatial Differentiation of Auditory Homogenization Indicators

The auditory indicators of the district show a distribution pattern characterized by the “dispersed retention of living sounds, the clustering of commercial and technical sound sources along main streets, and the localized embedding of quiet spaces.” Living sounds are discretely retained in internal streets and lanes, while commercial and technical sound sources are concentrated along main streets and peripheral roads. Quiet spaces appear only locally in Tianchi Park on the southern side and in adjacent slow traffic areas, and the overall influence of tourist activities is weak (Figure 6).
Local dialect intensity (LDI) is generally high, but its spatial distribution is relatively dispersed. High values mainly occur in internal living-oriented streets and lanes and in areas where residents’ activities are concentrated, whereas values are relatively low along peripheral roads and continuous commercial interfaces (Figure 6a). Commercial broadcasting intensity (CB) shows continuous high values along street sections with concentrated commercial shops, such as Ximen Street, Xueyuan Street, Qilin South Road, and Wenchang Street, while remaining generally low in internal branch lanes (Figure 6b). Sound event diversity (SED) is relatively high around schools, in internal living-oriented streets and lanes, and at some public activity nodes, indicating that the composition of sound sources in these areas is relatively diverse. The remaining street sections are dominated by medium and low values (Figure 6c).
The proportion of technical sound sources (TR) is generally high and is mainly distributed along peripheral roads, main streets, and traffic connection spaces. Localized high values also appear near cotton-fluffing workshops, construction sites, and interfaces where equipment is concentrated (Figure 6d). The influence of tourist activities (TAI) remains low overall and increases only locally around a few open spaces and commercial nodes. This indicates that changes in the district’s soundscape are mainly affected by daily commercial activities and technical sound sources, rather than by intensive tourism activities (Figure 6e). Perceived quietness (PC) is generally low, with high values mainly concentrated inside Tianchi Park on the southern side and in some adjacent slow traffic spaces. It is relatively low along peripheral roads, commercial main streets, and street sections with stronger technical sound sources (Figure 6f).

3.1.3. Distributional and Correlation Characteristics of Visual and Auditory Indicators

Relatively clear internal associations exist among traditional visual elements, among commercial landscape elements, and among some soundscape elements, while weak correlations are observed between visual and auditory indicators (Figure 7a).
Further comparison of Pearson and Spearman correlation coefficients revealed several moderate-to-strong associations among indicators within the same dimension. The strongest positive association was observed between façade patina and material vernacularity (r = 0.755r = 0.755r = 0.755, ρ = 0.756\rho = 0.756ρ = 0.756), whereas technophony ratio was moderately and negatively associated with perceived calmness (r = −0.641r = −0.641r = −0.641, ρ = −0.635\rho = −0.635ρ = −0.635). Moderate positive associations were also found between signage standardization and generic decoration, chain brand salience and commercial openness, commercial broadcasting and tourist activity impact, and local dialect intensity and sound event diversity. Differences between Pearson and Spearman coefficients for some variable pairs suggest potentially monotonic but not strictly linear relationships. Detailed results are reported in Table A9.
In terms of indicator distributions, the degree of dispersion differs markedly among visual indicators. Façade aging traces, material locality, and interface organicity show relatively wide value ranges, whereas architectural complexity, shop sign standardization, and chain brand prominence are more concentrated in lower value ranges. Generic decoration and commercial interface openness show more evident differences among sampling points (Figure 7b). Among the auditory indicators, local dialect intensity and the proportion of technical sound sources have relatively wide distribution ranges, while sound event diversity is mainly concentrated in the medium range. Commercial broadcasting and the influence of tourist activities are dominated by low-value samples, whereas perceived quietness shows a more distinct dispersed distribution (Figure 7c).

3.1.4. Spatial Autocorrelation of Visual and Auditory Homogenization Indices

To examine the spatial distribution patterns of the homogenization indices, Global Moran’s I was calculated for both visual and auditory homogenization using a fixed-distance-band spatial weight matrix, Euclidean distance, and row standardization. Based on the sampling interval and the spatial scale of the district, 50 m was selected as the primary distance threshold, while 40 m and 60 m were used for sensitivity testing.
The results showed that both the visual and auditory homogenization indices exhibited significant positive spatial autocorrelation across all distance thresholds (p < 0.001). For visual homogenization, Moran’s I ranged from 0.365670 to 0.448925. At the 50 m threshold, Moran’s I was 0.384152, with a z-score of 11.947115. For auditory homogenization, Moran’s I ranged from 0.128422 to 0.212447. At the 50 m threshold, Moran’s I was 0.166666, with a z-score of 5.234960. These findings indicate significant spatial clustering in both dimensions, with visual homogenization exhibiting stronger spatial autocorrelation than auditory homogenization (Table 5).

3.1.5. Spatial Mismatch Between Visual and Auditory Homogenization Indicators

There is a clear perceptual mismatch between visual and auditory homogenization within the district, and soundscape homogenization is stronger overall than visual homogenization (Figure 8b). High-value sampling points are mainly distributed along peripheral roads such as Wenchang Street, Qilin South Road, Liaokuo South Road, and Shengfeng Road, and are clustered around the commercial interfaces of main streets such as Ximen Street and Nanmen Street, as well as near nodes where these streets connect with peripheral roads. By contrast, internal living-oriented streets and lanes, such as Zhuge Lane, Wufu Lane, and Fensi Lane, show relatively low levels of homogenization (Figure 8a).
At the overall level, both the median and mean values of soundscape homogenization in the district are higher than those of visual homogenization (Figure 8c). To further determine whether this difference exists at the level of individual sampling points, we used the Visual–Audio mismatch to represent the difference between visual homogenization and soundscape homogenization. The results show that many sampling points are located below the zero line, indicating that soundscape homogenization is higher than visual homogenization at most sampling points. A small number of sampling points fall within the positive-value range, suggesting that visual homogenization is stronger in some local street sections (Figure 8d). The distribution of mismatch values further confirms this tendency. The mean value of the Visual–Audio mismatch is −1.10, and the median value is −1.29, with the overall distribution skewed toward negative values. This demonstrates that “soundscape homogenization being stronger than visual homogenization” is not an isolated phenomenon at individual sampling points, but a dominant trend across the district as a whole (Figure 8e). Moreover, visual homogenization generally shows relatively continuous spatial variation, whereas soundscape homogenization is more characterized by localized patches and nodal clustering (Figure 8(f1–f3)).

3.2. Identification of Multi-Sensory Homogeneity Types

To identify different visual–soundscape combination characteristics within Ximen Street, we conducted K-means clustering analysis based on 14 standardized visual and auditory indicators. We further compared the silhouette coefficient, within-cluster sum of squares (WCSS), Calinski–Harabasz index, Davies–Bouldin index, and type interpretability under different K values.
The number of clusters was determined by jointly considering the silhouette coefficient, within-cluster sum of squares (WCSS), Calinski–Harabasz index, Davies–Bouldin index, and the interpretive value of the resulting typology. The results showed that K = 2 yielded the highest silhouette coefficient (0.1857) and Calinski–Harabasz index (63.67), indicating relatively strong overall cluster separation. This solution was therefore suitable for capturing the broad distinction between high and low levels of multisensory homogenization. However, it tended to obscure differences among specific visual–auditory combinations.
As K increased, WCSS declined continuously and the Davies–Bouldin index generally decreased, indicating improved cluster compactness. Beyond K = 6, however, the reduction in WCSS became less pronounced, while the Calinski–Harabasz index generally declined. This suggests that further increases in the number of clusters provided only limited marginal interpretive gains. Accordingly, K = 2 was retained as a reference for broad statistical classification, whereas K = 6 was adopted as the more detailed typological solution that balanced statistical performance with interpretability for planning and management (Table 6).
For the K = 6 solution, the silhouette coefficient was 0.1578, the WCSS was 2791.21, the Calinski–Harabasz index was 47.12, and the Davies–Bouldin index was 1.7230. While maintaining an acceptable level of statistical performance, this solution distinguished six types characterized by different visual–auditory configurations and management needs: low commercial disturbance, commercially intensified soundscape, traditional townscape with technological sound disturbance, low visual homogenization with soundscape convergence, high commercial homogenization, and preserved traditional townscape with a mixed soundscape.
However, the K = 6 solution was not regarded as the uniquely optimal statistical solution. Rather, it was adopted as an interpretive classification designed to support the analysis of visual–soundscape mismatches and discussions of differentiated management strategies (Figure 9).

3.3. Key Drivers and Nonlinear Effects Based on SHAP

3.3.1. Model Performance and Spatial Cross-Validation

Under the random training–test split, the XGBoost model achieved a coefficient of determination (R2) of 0.886, a root mean square error (RMSE) of 0.345, and a mean absolute error (MAE) of 0.228 on the test set. The five-fold spatial cross-validation yielded a mean R2 of 0.868 ± 0.057, an RMSE of 0.344 ± 0.074, and an MAE of 0.246 ± 0.038. Compared with the random-split results, spatial cross-validation produced a slightly lower mean R2 and a modestly higher MAE, while the RMSE remained nearly unchanged. These findings indicate that the model maintained broadly comparable predictive performance under the more stringent condition of spatially separated validation (Table 7).
Model performance varied to some extent across the spatial folds, with R2 values ranging from 0.793 to 0.924. Fold 5 produced the lowest R2 and the highest RMSE and MAE, indicating relatively weaker generalizability in that spatial area. Nevertheless, R2 remained positive in every spatial fold, demonstrating that the model retained predictive capacity across all spatial partitions.
The observed values and out-of-fold predictions obtained from the five-fold spatial cross-validation were generally distributed close to the 1:1 reference line. When the out-of-fold predictions from all five spatial folds were pooled, the overall R2 was 0.898 (Figure 10a). This value was calculated from the combined set of all out-of-fold predictions. By contrast, the value of 0.868 ± 0.057 reported in Table 3 represents the mean and standard deviation of the R2 values calculated separately for the five spatial folds. Because these two statistics were derived using different calculation procedures, they should not be compared directly.
The spatial-fold configuration showed that the 340 sampling points were divided into relatively contiguous spatial partitions, thereby reducing the likelihood that spatially adjacent samples were simultaneously assigned to the training and test sets (Figure 10b). The spatial distribution of the out-of-fold residuals showed no clear evidence of extensive and continuous clustering of model errors, although some local sampling points exhibited residual variation (Figure 10c). Overall, model performance under spatial cross-validation was broadly comparable to that obtained from the random split, suggesting that the random-split results were not subject to substantial performance inflation. Nevertheless, potential spatial dependence near the boundaries of the spatial folds cannot be entirely ruled out.

3.3.2. Global Importance of Visual and Audio Indicators

The XGBoost model demonstrates good predictive performance for multisensory homogenization. The test-set results show that the model achieved an R2 of 0.886, an RMSE of 0.345, and an MAE of 0.228. The predicted values are generally close to the observed values, and the residuals are mainly distributed around zero, indicating that the model can characterize the level of multisensory homogenization in a relatively stable manner.
In terms of global feature importance, generic decoration (0.276) makes the highest contribution, indicating that replicated, template-based, and landscape-oriented decoration is the most critical visual factor affecting homogenization prediction. Commercial broadcasting (0.227) and the proportion of mechanical sounds (0.186) rank second and third, respectively, suggesting that standardized sound sources, such as outdoor shop broadcasting, promotional announcements, vehicle sounds, and equipment sounds, have an important influence on auditory homogenization. Material locality (0.172), local dialect intensity (0.170), and perceived quietness (0.134) also make relatively high contributions (Figure 11a).
Their relatively high importance indicates strong explanatory power in distinguishing between high- and low-homogenization samples. Combined analysis of the SHAP scatter plot and SHAP heatmap further shows that, as the model output value f(x) increases, generic decoration, commercial broadcasting, and the proportion of mechanical sounds make continuous positive contributions across multiple samples. This suggests that highly homogenized samples are often jointly driven by visual commercialization and sound standardization. In contrast, façade weathering traces, material locality, local dialect intensity, and perceived quietness show negative contributions in some samples, reflecting the inhibitory effects of traditional visual elements and local soundscapes on homogenization prediction (Figure 11b).

3.3.3. Main and Interaction Effects of Multisensory Factors

The results of the SHAP main effects and interaction effects indicate that multisensory homogenization is jointly influenced by visual commercialization, soundscape standardization, and traditional local elements. Generic decoration has the highest node importance and shows strong associations with variables such as commercial broadcasting, the proportion of technical sound sources, façade weathering traces, and material locality. This suggests that visual commercialization interacts with commercial sound sources, technical sound sources, and traditional townscape elements in shaping perceived homogenization (Figure 12a).
A further comparison of main and interaction effects shows that most indicators are still dominated by main effects. Among them, generic decoration has the highest main effect (0.308), making it the primary factor influencing the prediction of multisensory homogenization. This indicates that replicated, template-based, and landscape-oriented decorations have the strongest explanatory power for perceived homogenization. Commercial broadcasting (0.227) and the proportion of technical sound sources (0.192) rank next, suggesting that standardized sound sources, such as outdoor shop broadcasting, promotional announcements, vehicle sounds, and equipment sounds, are also important drivers of homogenization. At the same time, indicators such as façade weathering traces, material locality, commercial broadcasting, and the proportion of technical sound sources still make certain interaction contributions, reflecting the complex relationships among traditional visual townscape, commercial symbols, and technologically mediated soundscapes (Figure 12b).

3.3.4. Nonlinear Effects and Threshold Responses

The two-dimensional PDP results reveal the nonlinear effects of key variable combinations on the prediction of multisensory homogenization. When GD reaches a high value, the predicted level of homogenization increases markedly, indicating that template-based decoration is an important factor weakening visual locality. Second, GD × CB and GD × TR show audiovisual superposition effects; that is, when generic decoration, commercial broadcasting, and technical sound sources increase simultaneously, the predicted value of multisensory homogenization further rises. Finally, GD × LDI reflects the tension between commercialized visual interfaces and local sounds. When the intensity of local dialects is relatively high, the tendency toward homogenization is weakened to some extent (Figure A1).
The SHAP dependence plots show that visual indicators have clear nonlinear and threshold effects on the prediction of multisensory homogenization (Figure 13).
(1) Inverted U-shaped relationships. FP, MV, AC, and IO all exhibit inverted U-shaped or weak inverted U-shaped relationships, characterized by an initial increase followed by a subsequent decline. FP turns into a positive contributor after approximately 1.7, reaches its peak at a moderate level, and becomes a negative contributor after approximately 6.0. MV, AC, and IO change from positive to negative contributors at approximately 4.7, 3.9, and 5.6, respectively. These results indicate that historical traces, local materials, architectural complexity, and organic interfaces are not sufficient to weaken perceived homogenization when they remain at low or moderate levels. Only when they reach relatively high levels do they clearly reduce the predicted value of multisensory homogenization (Figure 13a–d).
(2) Declining relationships. LDI, SED, and PC exhibit declining relationships and become negative contributors after approximately 5.0, 4.8, and 3.6, respectively. This suggests that local dialects, rich daily sound events, and a higher sense of quietness can weaken perceived soundscape homogenization once they reach a certain level (Figure 13i,k,n).
(3) Increasing relationships. SS, CS, GD, and CO generally show increasing relationships. SS, GD, and CO become positive contributors after approximately 3.9, 4.4, and 3.7, respectively. CS shows a positive effect after approximately 1.1 and continues to increase the model output in the medium- and high-value ranges. This indicates that once shop sign standardization, chain brand presence, generic decoration, and open commercial interfaces exceed certain levels, they significantly strengthen perceived visual homogenization (Figure 13e–h).
To further assess whether these breakpoints had the potential to generalize beyond the observed response patterns, formal segmented regression analyses were conducted. The results provided statistical support for case-specific breakpoints in six visual indicators: traces of historical development on façades, local distinctiveness of materials, architectural complexity, organic character of street interfaces, standardization of storefront signage, and openness of commercial interfaces. For the first four indicators, the slope changed from positive to slightly negative beyond the breakpoint. Storefront signage standardization and commercial interface openness remained positively associated with homogenization beyond their respective breakpoints, but their marginal effects weakened substantially.
None of the auditory indicators exhibited statistically supported breakpoints after false discovery rate (FDR) correction. The identified values should therefore be interpreted as empirical breakpoints specific to the Ximen Street case rather than universally applicable thresholds (Figure A2 and Figure A3 and Table A7).

4. Discussion

This study discusses the findings in relation to the three research questions proposed in the Introduction. First, visual and auditory homogenization in Ximen Street were not synchronized. Soundscape homogenization was generally more pronounced and exhibited stronger patch-like and node-based clustering, indicating that the continuity of the visible townscape does not necessarily ensure the continuity of place-based perception. Second, the six multisensory types revealed distinct combinations of visual character, commercial interfaces, and soundscape locality across different street segments. Third, multisensory homogenization was jointly shaped by commercial visual elements, technological sound sources, the continuity of the historic environment, and the everyday sounds of residents. These influences also exhibited nonlinear and cross-modal interaction effects. The following discussion addresses these findings from three perspectives: audiovisual mismatch, typology-based management, and compound driving mechanisms.

4.1. Mismatch Between Visual Conservation and Soundscape Locality

The current conservation and renewal processes in Ximen Street Historic District reveal a degree of imbalance between the preservation of visual character and the maintenance of soundscape locality. Although elements of the visible historic townscape have been retained in some parts of the district, place-specific auditory characteristics have weakened to a certain extent. This mismatch is particularly evident along commercial streets and peripheral traffic corridors, where preserved façades coexist with commercial broadcasting, traffic noise, and mechanical sounds. The conservation of the visible physical environment is undoubtedly important for sustaining historical continuity. However, if locally distinctive sounds are gradually replaced by standardized commercial soundscapes, technological sound sources, or traffic noise, the preservation of architectural character and spatial form alone is unlikely to sustain the locality of historic districts [35,36,37]. Conservation management should therefore move beyond static townscape improvement toward a form of “dynamic conservation” that integrates both visual and auditory dimensions.
Environmental perception is jointly shaped by visual and auditory information, while the intensity and distribution of different sound categories vary with rhythms of human activity and environmental conditions [32,33,34]. Visual homogenization in Ximen Street exhibited relatively continuous spatial variation. By contrast, soundscape homogenization was more strongly concentrated along commercial streets, peripheral traffic corridors, and activity nodes, and its overall level was higher than that of visual homogenization. Wenchang Street, Qilin South Road, Liaokuo South Road, and Shengfeng Road showed stronger exposure to traffic and technical sounds, whereas internal residential lanes and Tianchi Park retained more localized or relatively calm soundscapes. This divergence reflects the different mechanisms through which the two environmental dimensions change. Building form is constrained by conservation regulations and construction cycles and therefore usually changes slowly. Soundscapes, in contrast, can change rapidly in response to commercial operations, traffic volume, equipment use, and patterns of human activity.
Conservation policies for historic districts should therefore adopt differentiated measures based on the spatial distribution, temporal dynamics, and underlying mechanisms of visual and auditory change. At the visual level, long-term control and incremental repair should be strengthened for façades, materials, and street interfaces. Along the commercial frontages of Ximen Street and Nanmen Street, continuous standardized signage and generic decoration that conceal local materials or historic details should be avoided.
At the auditory level, time-specific monitoring and sound source management should be introduced to regulate commercial broadcasting, mechanical equipment, and traffic disturbance, while maintaining locally distinctive sounds associated with dialect use, neighborhood interaction, and everyday commercial activities. At the audiovisual level, an integrated assessment and renewal review mechanism should be established. Townscape improvement, commercial functions, equipment installation, traffic organization, and public activities should be incorporated into a unified management framework.
The findings further indicate that respondents’ identities and place-based experiences influenced their judgments of homogenizing renewal. For long-term residents, the continuity of place memory was important, but so were improvements in environmental sanitation, housing safety, accessibility, and neighborhood vitality. Street front business operators and some respondents who had previously studied or worked in the district also tended to view the overall environmental improvements positively. Compared with the former conditions of deteriorating facilities, disorderly surroundings, and poor housing quality, they considered a certain degree of visual or soundscape homogenization acceptable when accompanied by broader improvements to the living environment.
These findings suggest that assessments of homogenization are also shaped by lived experience, practical needs, and stakeholder interests. Conservation and renewal policies should therefore balance the continuity of local identity with residents’ legitimate demands for greater safety, convenience, and quality of life (Table A6).

4.2. Multi-Sensory Type Identification and Governance Implications

The identification of multisensory types provides a practical basis for translating research findings into differentiated governance strategies. The Ximen Street Historic District cannot simply be described as either “well preserved” or “highly homogenized”. Some street sections retain traditional visual townscape features but are disturbed by technical sound sources; some show low levels of visual homogenization but have already experienced soundscape standardization; and others simultaneously exhibit high-intensity commercial interfaces, visual homogenization, and soundscape homogenization. These differences indicate that a single façade control strategy is insufficient. For street sections with relatively intact visual heritage but high soundscape homogenization, particularly peripheral roads and major junctions, priority should be given to controlling outdoor broadcasting, technical equipment, and traffic disturbances. Outward-facing loudspeakers should be restricted, broadcasting periods should be regulated, and mechanical equipment should be screened or acoustically treated. For commercial street sections with high levels of both visual and auditory homogenization, especially the commercial frontages of Ximen Street and Nanmen Street, shop signs, generic decorations, open commercial interfaces, and sound-producing equipment should be managed in a coordinated manner. By contrast, residential lanes with low commercial disturbance should be protected as environments in which daily activities, social interactions, and local sound events remain perceptible.
This type-based approach to supporting hierarchical and differentiated governance is consistent with existing studies on historic districts that allocate different renewal strategies according to street type classification [2,40]. It is also grounded in the material form, spatial organization, sociocultural practices, and economic processes of specific places, rather than treating classification results as fixed zones detached from local contexts [4,8,41,42,43]. Therefore, this method can be understood as a governance-oriented typological tool for identifying intervention priorities in different spatial units and for providing a basis for refined renewal at the level of street sections and street front interfaces, rather than as a fixed or universally applicable classification system (Table 8).

4.3. Composite Driving Mechanism of Visual and Auditory Homogeneity Perception

In the Ximen Street Historic District, multisensory homogenization primarily results from the cumulative effects of dispersed interventions, including generic decoration, standardized storefront signage, commercial broadcasting, and noise from mechanical equipment. The influence of any single instance of similar decoration, amplified advertising, or equipment operation may be limited. However, when such elements repeatedly occur or become concentrated along particular street segments, they may gradually weaken the district’s original visual and auditory distinctiveness.
At the same time, traces of historical development on façades, local materials, architectural complexity, and the organic character of street interfaces appear to reduce perceived homogenization only when they form a sufficiently continuous environmental pattern. The isolated preservation of traditional elements is unlikely to offset the cumulative effects of commercial standardization. The combined presence of generic decoration, commercial broadcasting, and technological sound sources further intensifies perceived homogenization. By contrast, everyday sounds such as local dialects and residents’ conversations are associated with a lower tendency toward homogenization. The rankings of variable importance, effect magnitudes, and empirical response ranges identified in this study are specific to the Ximen Street sample and should not be directly generalized to other heritage sites.
Nevertheless, the underlying mechanisms may have some transferability to other heritage contexts. Commercial activities and anthropogenic sound sources can gradually reshape the environmental experience and local atmosphere of historic districts through long-term accumulation [30,31]. The maintenance of locality does not depend merely on the presence of individual traditional elements. Rather, it relies on continuous relationships among architectural form, local materials, historical traces, street interfaces, and everyday life [3,4,8].
Visual and auditory information may also jointly contribute to the formation of place perception [28,29,32,33,34]. When repetitive commercial landscapes coincide with broadcasting, traffic, or equipment noise, the perceived standardization of the environment may be intensified. Conversely, local dialects, neighborhood interactions, and other everyday sounds can convey sociocultural information that the physical environment alone cannot fully express. The resulting mechanism—characterized by the accumulation of dispersed interventions, the continuity of heritage cues, and audiovisual synergy—may provide an analytical reference for other historic districts that retain active residential life while facing pressures from commercial renewal.

4.4. Limitations and Future Work

In this study, we integrated street-view imagery, soundscape recordings, AI-assisted scoring, human perception validation, spatial analysis, and explainable machine learning to establish an analytical framework for multisensory homogenization in historic districts. It also identified spatial differences between visual townscape character and soundscape locality, together with their composite driving mechanisms. Nevertheless, the study has several limitations related to the scope of the case study, data conditions, and methodological applicability.
First, the analysis was limited to the Ximen Street Historic District. Spatial cross-validation assessed model stability only across spatially separated samples within the district and cannot substitute for external validation across different regions. The spatial patterns, variable importance rankings, cluster types, and empirical breakpoints identified here are dependent on local conditions, including street morphology, residential activities, commercial structure, the extent of tourism involvement, and soundscape composition. They should therefore not be directly generalized. By comparison, the analytical procedures involving matched street-view and audio sampling, audiovisual mismatch identification, typological classification, and nonlinear interpretation may have greater transferability. Their external validity should be further examined through multi-case comparisons under standardized sampling conditions, cross-regional model testing, and leave-one-region-out cross-validation.
Second, the study primarily reflects daytime perceptual conditions during winter and does not fully capture seasonal variation, day–night transitions, or differences between weekdays and holidays. Soundscapes are highly dynamic and may change rapidly with commercial operations, traffic flows, school schedules, pedestrian activity, and temporary events. Visual environments may likewise be influenced by lighting, weather, commercial displays, and festival decorations. The present findings should therefore not be interpreted as representing a stable long-term perceptual condition. Future research should establish repeated sampling protocols across multiple seasons, time periods, and dates. Continuous acoustic monitoring could also be incorporated to assess the temporal stability of audiovisual homogenization patterns.
Third, the conditions under which street-view images and audio recordings were collected may still affect comparability across sampling points. Although we standardized the equipment, camera height, survey period, and basic weather conditions, automatic exposure, lighting direction, partial occlusion, and unexpected sound events may have influenced some indicators. Future studies could use fixed exposure settings and standardized color charts. Longer recording durations, repeated sampling, and the annotation of anomalous sound events could further improve the consistency and stability of multimodal data.
Fourth, AI-assisted scoring should be regarded only as a structured approximation of perception. It cannot replace residents’ actual perceptions, expert judgments, or heritage value assessments. Pearson correlation primarily reflects correspondence in the patterns of variation between AI-generated and aggregated human ratings, while the ICC reflects their degree of agreement. Neither measure is equivalent to objective accuracy. The scoring results may still be influenced by model training data, prompt design, indicator definitions, model version, input quality, and local cultural context. AI models may also reproduce existing human cognitive biases. Future studies should enhance transparency, reproducibility, and robustness through cross-model evaluation, expert review, larger human validation samples, and the disclosure of non-sensitive prompts, parameters, and code.
Fifth, the samples used for human perception validation and interviews were limited in size. The 30 participants were recruited primarily for an exploratory assessment of human–AI agreement and cannot represent all residents and users of the district. Their evaluations may have been affected by age, disciplinary background, familiarity with the district, ability to understand local dialects, and place attachment. The interviews also relied on a small purposive sample. Respondents’ lived experiences, commercial interests, and recall biases, as well as the researchers’ familiarity with the local context, may have influenced the interpretation of the findings. Future research should include broader representation of long-term residents, older adults, business operators, tourists, and heritage professionals. Soundwalks, focus groups, and participatory mapping could also be used for triangulation.
Sixth, uncertainty remains in the model interpretations, clustering results, and nonlinear breakpoints. SHAP values and partial dependence plots reflect predictive associations rather than causal relationships. The empirical breakpoints are applicable only to the Ximen Street sample, while the selection of K = 6 represents a compromise between statistical performance and management interpretability. Future studies should use larger samples, spatially blocked cross-validation, independent test sets, repeated clustering, and uncertainty analysis to improve result robustness .
Finally, the proposed management recommendations are context-dependent. Their implementation is also constrained by funding, administrative authority, business needs, resident acceptance, and relevant institutional arrangements. Future research could conduct pilot interventions in representative street segments and compare audiovisual environments, resident perceptions, commercial effects, and management costs before and after implementation. Such evaluations would help determine the practical feasibility of the multisensory governance framework.

5. Conclusions

This study used the Ximen Street Historic District in Qujing, Yunnan Province, as a case study and constructed a framework for identifying and interpreting multisensory homogenization in historic districts based on streetscape image–environmental audio matched samples. The framework integrates visual and auditory indicators, manual perceptual validation, spatial analysis, K-means clustering, XGBoost, and SHAP-based interpretation. Compared with traditional townscape conservation, this framework incorporates local sounds into the same evaluation system and identifies six types of homogenized street sections within the district: low commercial disturbance, intensified commercial soundscape, traditional townscape with technical sound disturbance, low visual homogenization with soundscape convergence, high commercial homogenization, and retained traditional townscape with mixed soundscape.
Within the district, visual homogenization shows a relatively continuous distribution, whereas soundscape homogenization is more characterized by patch-like patterns and nodal clustering, with a higher overall level. Visual and auditory cues are preserved, replaced, and superimposed in an interwoven and complex manner within the district. Moreover, multisensory homogenization is not determined by a single visual or auditory factor, but is jointly shaped by the standardization of commercial landscapes, the intrusion of technical sound sources, and the degree to which local cues are retained. Generic decoration, standardized shop signs, and open commercial interfaces increase the replicability of the visual environment, while commercial broadcasting, vehicle sounds, and equipment sounds cover or reorganize residents’ conversations, local dialects, and sounds of daily activities. When visual commercialization and soundscape standardization are superimposed within the same spatial unit, they exert a synergistic amplifying effect on the perception of overall homogenization. Conversely, historical traces, local materials, interface organicity, local dialects, sound event diversity, and perceived quietness jointly constitute an important basis for inhibiting the perception of homogenization.
This paper extends the issue of homogenization in historic districts from the traditional level of visual townscape to the weakening of multisensory locality. It reveals the hidden risk that visual conservation does not necessarily equate to the continuity of perceived locality. For the practice of historic district conservation, future renewal governance should not only focus on townscape conservation, but also emphasize the coordinated management of soundscape order. Future research may further extend to historic districts in different cities, with different degrees of tourism development and at different spatial scales. It may also incorporate long-term monitoring across seasons and day–night periods to enhance the refined conservation of historic districts.

Author Contributions

Conceptualization, Y.Q., Y.H. and D.Y.; methodology, Y.Q.; software, Y.Q.; validation, Y.Q. and D.Y.; formal analysis, Y.Q., D.Y. and Y.H.; investigation, Y.Q. and D.Y.; resources, Y.Q. and D.Y.; data curation, Y.Q.; writing—original draft preparation, Y.Q.; writing—review and editing, Y.Q., D.Y. and Y.H.; visualization, Y.Q.; supervision, D.Y. and Y.H.; project administration, Y.Q. and D.Y.; funding acquisition, D.Y.; translation and language polishing, Y.H. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China, grant number 52478018.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Kunming University of Science and Technology Medical Ethics Committee (approval number: KMUST-MEC-2026-102; date of approval: 23 June 2026).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study. All participants were adults and participated voluntarily and anonymously. No personally identifiable information was collected.

Data Availability Statement

The data presented in this study are available from the corresponding author upon reasonable request. The data are not publicly available due to privacy restrictions stipulated in the informed consent obtained from the participants.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A

Table A1. Definitions, measurement dimensions, and VLM prompt logic for visual heritage authenticity indicators (Source: Compiled by the authors).
Table A1. Definitions, measurement dimensions, and VLM prompt logic for visual heritage authenticity indicators (Source: Compiled by the authors).
VariableDefinition and Measurement DimensionVLM Prompt Logic
Façade PatinaMeasures the degree of historical weathering and age traces visible on building façades. In more homogenized streetscapes, façades are often excessively renovated and appear clean, polished, and newly finished.“Rate the weathering and age traces on the building surfaces.”
Material VernacularityAssesses the proportion of traditional local materials (e.g., blue brick, timber, stone) relative to modern industrial materials (e.g., paint, ceramic tiles, aluminum composite panels).“Identify the dominant building materials. Estimate the proportion of traditional local materials versus modern industrial materials.”
Architectural ComplexityEvaluates the richness of façade details, such as carvings, lattice windows, eaves, and decorative components. Homogenization often leads to simplified details or superficial symbolic replication.“Analyze the visual complexity of architectural details.”
Interface Organic-nessMeasures whether the street interface is naturally developed and spatially varied, or instead aligned, standardized, and uniformly planned.“Evaluate the rhythm of the street wall. Is it organic and varied or rigidly uniform and standardized ?”
Table A2. Definitions, measurement dimensions, and VLM prompt logic for commercial landscape standardization indicators (Source: Compiled by the authors).
Table A2. Definitions, measurement dimensions, and VLM prompt logic for commercial landscape standardization indicators (Source: Compiled by the authors).
VariableDefinition and Measurement DimensionVLM Prompt Logic
Signage StandardizationMeasures whether shop signs are standardized in font, color, size, material, placement, and template design. This is a typical visual feature of homogenization in Chinese historic streets. Higher scores indicate stronger standardization of commercial signage.“Look at the shop signs in the image. Are they standardized in font, color, size, material, placement, or template design? “
Chain Brand SalienceMeasures the visibility and dominance of well-known chain brands, including global, national, or regional chain stores, within the street-view scene. Higher scores indicate stronger dominance of chain brand commercial landscapes and weaker presence of local independent businesses.“Detect visible logos, storefronts, or signs of major chain brands. Estimate the dominance of global, national, or regional chain stores in the image.”
Generic DecorationEvaluates whether the scene contains generic, replicable, or tourism-oriented decorative elements, such as lantern arrays, umbrella canopies, uniform planter boxes, flags, artificial retro-style props, themed photo spots, or Instagrammable installations. Higher scores indicate stronger use of generic tourist decorations and weaker local specificity.“Identify decorative elements such as lanterns, umbrellas, flags, uniform planter boxes, artificial retro-style props, themed photo spots, or tourism-oriented installations. Are these decorations generic and replicable, or culturally specific to the local context? “
Standardized Commercial OpennessMeasures the extent to which the ground-floor street interface is occupied by open, consumption-oriented, and standardized commercial frontages, such as glass storefronts, product displays, outdoor seating, standardized shopfront layouts, and continuous commercial openings. Higher scores indicate stronger replacement of residential or life-oriented interfaces by standardized commercial spaces.“Estimate the proportion of the ground-floor façade dedicated to commercial display, open storefronts, product display, outdoor seating, standardized shopfront layouts, or continuous commercial openings.”
Table A3. Definitions, measurement dimensions, and ALM prompt logic for auditory homogenization indicators (Source: Compiled by the authors).
Table A3. Definitions, measurement dimensions, and ALM prompt logic for auditory homogenization indicators (Source: Compiled by the authors).
VariableDefinition and Measurement DimensionALM Prompt Logic
Local Dialect IntensityMeasures the weakening of local dialects or local accents in the recording. Higher scores indicate that speech is dominated by standard Mandarin, non-local tourist accents, or foreign languages, while local dialects are weak or absent.Use ASR to transcribe spoken content and identify language or accent features. The model can be prompted: “Analyze the spoken language and accent characteristics in the audio. Is the speech primarily local dialect, mixed speech, standard Mandarin, non-local accents, or foreign languages?”
Sound Event DiversityEvaluates the reduction in diverse and place-specific sound events, such as local conversations, footsteps, cooking sounds, handcraft sounds, birdsong, water sounds, mahjong sounds, street vending sounds, and community activity sounds. Higher scores indicate fewer sound event types and a more simplified soundscape.Count or estimate the number of distinct recognizable sound event categories. The model can be prompted: “Identify the recognizable sound event categories in the audio, such as birdsong, footsteps, cooking sounds, local conversations, handcraft sounds, mahjong sounds, traffic sounds, commercial broadcasting, or community activity sounds.”
Commercial Broadcasting IntensityMeasures the dominance of commercial sounds, including shop music, promotional broadcasting, looped advertisements, amplified sales calls, and commercial loudspeakers. Higher scores indicate stronger commercial occupation of the acoustic environment.Detect repetitive, amplified, and promotional audio content. The model can be prompted: “Detect the presence and dominance of amplified sales pitches, looped advertisements, promotional broadcasting, shop music, or commercial loudspeakers.”
Tourist Activity ImpactMeasures the intensity and intrusion of tourist-related sounds, including tour guide explanations, loudspeaker use, group chatter, crowd noise, photo-taking activities, tourism-oriented performances, and organized tour groups. Higher scores indicate stronger tourist occupation of the acoustic foreground.Identify tourist-related speech, group noise, and loudspeaker characteristics. The model can detect keywords or expressions such as “everyone look here,” “follow me,” or amplified guide commentary, as well as the acoustic features of loudspeakers and tour groups. The model can be prompted: “Assess the intensity of tourist-related sounds in the audio.”
Technophony RatioMeasures the proportion and dominance of mechanical, technological, or traffic-related sounds, such as motor vehicles, electric bikes, air-conditioning units, construction equipment, generators, loudspeakers, and other equipment noise. Higher scores indicate stronger intrusion of non-local mechanical sound sources.Estimate the proportion of time dominated by mechanical, technological, construction, or traffic-related sounds. The model can be prompted: “Estimate the proportion of the audio duration dominated by mechanical, technological, construction, or traffic-related sounds, such as vehicles, electric bikes, air-conditioning units, construction equipment, generators, or equipment noise.”
Perceived CalmnessEvaluates whether the expected calmness and acoustic order of the historic street are masked by noisy background sounds. It reflects perceived tranquility, acoustic comfort, signal clarity, and overall soundscape order. Higher scores indicate stronger acoustic disturbance, lower tranquility, and poorer signal-to-noise quality.Assess the perceived calmness, acoustic order, and signal-to-noise quality of the recording. The model can be prompted: “Evaluate the overall acoustic comfort, calmness, and disturbance level of the environment.”
Table A4. Model parameters and request settings for visual and auditory scoring.
Table A4. Model parameters and request settings for visual and auditory scoring.
ItemVisual ScoringAuditory Scoring
Score scale0.00–10.00, two decimal places0.00–10.00, two decimal places
Temperature0.10.1
Top-p0.80.8
Top-k4040
Maximum output length3000 tokens3000 tokens
Request timeout90 s90 s
Maximum retries55
Table A5. Participant characteristics and place familiarity in the human perception. (Source: Compiled by the authors).
Table A5. Participant characteristics and place familiarity in the human perception. (Source: Compiled by the authors).
CategoryItemn (%)
GenderMale15 (50.0%)
Female15 (50.0%)
Age18–29 years12 (40.0%)
30–44 years10 (33.3%)
45 years and above8 (26.7%)
Professional backgroundArchitecture, urban planning, landscape architecture, heritage conservation, or acoustics-related fields10 (33.3%)
Other disciplines or non-related occupations20 (66.7%)
Residence or life experienceResidents or street-front business operators in Ximen Street10 (33.3%)
Residents of other areas in Qujing10 (33.3%)
Non-residents of Qujing10 (33.3%)
Familiarity with Ximen StreetNever visited or visited only briefly6 (20.0%)
Occasional visitors8 (26.7%)
Frequent visitors6 (20.0%)
Long-term residents, workers, or business operators10 (33.3%)
Exposure to the Qujing local dialectAble to understand or use it proficiently10 (33.3%)
Able to understand it partially12 (40.0%)
Unable to understand it8 (26.7%)
Sensory conditionsNormal or corrected-to-normal vision30 (100.0%)
Self-reported normal hearing30 (100.0%)
Table A6. Characteristics of interview participants, length of place experience, and summary of interview content (Source: Compiled by the authors).
Table A6. Characteristics of interview participants, length of place experience, and summary of interview content (Source: Compiled by the authors).
IDTypePlace Experience (Years)Summary of Interview Content
R01Long-term resident and local worker40The participant first became familiar with the area around 1985 and began teaching at a local school around 1988. The former school, established in 1958, once contained relatively well-preserved and esthetically distinctive traditional buildings, which were later demolished. The participant expressed regret over their loss. In the participant’s view, the main section of Ximen Street has changed relatively little. Visiting the same long-established breakfast shops over several decades has become part of the participant’s daily routine. The participant also acknowledged the visible improvements resulting from recent renewal.
R02Long-term resident35The participant has lived in the area for several decades. In the participant’s view, recent policy support and renewal projects have improved both the physical environment and the vitality of the district. During holidays, students and visitors frequently come to the food street, creating a lively atmosphere resembling that of a small tourist attraction.
R03Street-front business operator2–3The participant reported that business was very good and that customers often had to queue during holidays. Many nearby shops had also opened only in recent years, making the area increasingly resemble a small tourist attraction.
R04Local worker36The participant considered Ximen Street to possess considerable historical and cultural significance and argued that its heritage should be further developed and presented on the basis of appropriate conservation. The old urban area was perceived as still being relatively disordered and insufficiently regulated, requiring a balance between modernization and the preservation of historic cultural characteristics. The participant positively evaluated the renovation of traditional urban courtyards and lanes, while also identifying aging electrical systems, unsafe buildings, potential safety hazards, and poor accessibility in some deep lanes.
R05Long-term resident and former local student25The participant observed frequent turnover among shop operators, with businesses repeatedly entering and leaving the district. The participant also retained memories of traditional daily and craft-related activities and sounds, particularly cotton fluffing.
R06Student or short-term district user3–4The participant noted that public activities are frequently organized during holidays. During some events, purchases are supported by government subsidies, creating an atmosphere similar to that of a lively traditional market.
R07Long-term resident and volunteer or sanitation worker35The participant has long been involved in, or has closely observed, cleaning and environmental maintenance in the district. The area was previously perceived as relatively dirty, with frequent littering. Following recent renewal and management measures, the district has become noticeably cleaner, and the burden of daily maintenance has been reduced.
R08Local worker20–30The participant retained strong memories of the former school entrance, the Qilin campus, and local foods such as erkuai. These school spaces and food-related activities constituted important components of the participant’s place identity and collective memory of the district.
R09Street-front business operator or long-term district user15The participant considered that Ximen Street had changed substantially in recent years, while its overall development was viewed positively. School life, class reunions, and long-established local food businesses serving chicken-soup rice, Dafugui specialties, and cold rice noodles constituted the participant’s most prominent memories of the district.
R10Long-term resident20The participant mentioned daily neighborhood and leisure activities such as mahjong playing. The participant also recalled that some buildings in the district had previously been of poor quality, including earth wall structures and unsafe houses, indicating the practical need for renewal and improvement of residential conditions.
R11Street-front business operator or long-term district user15The participant retained strong memories of long-established local businesses such as Sanmei Soy Milk, where customers often had to queue. Clear differences were perceived between Ximen Street and the historic towns of Dali and Huize. Dali and Huize were considered to have more complete historic architecture and stronger tourism-oriented development, whereas Ximen Street remained primarily characterized by the daily lives of long-term residents, including food markets, local snack shops, roasted chestnut vendors, and street cries. In auditory terms, school lessons, kindergarten and primary school activities, breakfast preparation, residential activities, and street vending were regarded as distinctive local sound cues. By contrast, sounds associated with bars, cafés, and clubs were perceived as more similar to the commercial soundscapes of tourism-oriented historic towns such as Dali. Overall, the participant positively evaluated the substantial improvements resulting from recent conservation and renewal.

Appendix B

XGBoost Objective Function

At the t-th iteration, XGBoost adds a new regression tree f t ( x ) to minimize the sum of the loss function and model complexity. The objective function is expressed as
L ( t ) = i = 1 n l y i , y ^ i ( t 1 ) + f t x i + Ω f t Ω f t = γ T + 1 2 λ j = 1 T w j 2
where y i denotes the integrated multisensory homogenization perception score of sample i , y ^ i ( t 1 ) is the predicted value after the ((t − 1))-th iteration, f t x i is the prediction of the newly added tree, T is the number of leaf nodes, w j is the weight of the j -th leaf node, and γ and λ control model complexity and leaf node weights, thereby improving model generalization.
In the model interpretation stage, SHAP was used to quantify the marginal contribution of each variable to the prediction of an individual sample [44]. For feature j , the SHAP value is defined as
ϕ j = S F \ j | S | ! ( M | S | 1 ) ! M ! f S j x S j f S x S
where F is the full feature set, S is any feature subset that does not include feature j , (M) is the total number of features, and f S x S denotes the model prediction using only subset S . A positive SHAP value indicates that the variable increases the predicted multisensory homogenization value, whereas a negative SHAP value indicates that the variable helps reduce perceived homogenization. Based on SHAP global importance, dependence plots, and interaction values, the study further identified the main effects, nonlinear thresholds, and cross-modal interactions among visual commercialization, traditional streetscape preservation, and soundscape locality.

Appendix C

Figure A1. Two-dimensional partial dependence plots of key visual–audio interaction effects (Source: authors).
Figure A1. Two-dimensional partial dependence plots of key visual–audio interaction effects (Source: authors).
Buildings 16 02871 g0a1
Figure A2. Segmented regression results and bootstrap confidence intervals for empirical breakpoints of the 14 visual and auditory indicators: (a) façade patina; (b) material vernacularity; (c) architectural complexity; (d) interface organic-ness; (e) signage standardization; (f) chain brand salience; (g) generic decoration; (h) commercial openness; (i) local dialect intensity; (j) commercial broadcasting; (k) sound event diversity; (l) technophony ratio; (m) tourist activity impact; and (n) perceived calmness. Blue points represent sample observations, pink lines indicate segmented regression fits, dashed lines denote estimated breakpoints, and shaded areas show the 95% bootstrap confidence intervals (Source: Authors).
Figure A2. Segmented regression results and bootstrap confidence intervals for empirical breakpoints of the 14 visual and auditory indicators: (a) façade patina; (b) material vernacularity; (c) architectural complexity; (d) interface organic-ness; (e) signage standardization; (f) chain brand salience; (g) generic decoration; (h) commercial openness; (i) local dialect intensity; (j) commercial broadcasting; (k) sound event diversity; (l) technophony ratio; (m) tourist activity impact; and (n) perceived calmness. Blue points represent sample observations, pink lines indicate segmented regression fits, dashed lines denote estimated breakpoints, and shaded areas show the 95% bootstrap confidence intervals (Source: Authors).
Buildings 16 02871 g0a2aBuildings 16 02871 g0a2b
Table A7. Formal Segmented-Regression Tests and Spatial-Block Bootstrap Uncertainty for Empirical Breakpoints of the 14 Visual and Auditory Indicators.
Table A7. Formal Segmented-Regression Tests and Spatial-Block Bootstrap Uncertainty for Empirical Breakpoints of the 14 Visual and Auditory Indicators.
VariableBreakpoint (0–10)95% Spatial-Block Bootstrap CICI WidthSlope BeforeSlope AfterNonlinearity FDR qBreakpoint FDR qΔAICConclusion
Façade Patina (FP)2.252.25–2.750.500.871−0.0470.00350.0035219.64Supported case-specific breakpoint
Material Vernacularity (MV)1.751.75–1.750.001.211−0.0860.00350.0035210.39Supported case-specific breakpoint
Architectural Complexity (AC)2.252.25–2.250.000.739−0.1060.00350.0035142.61Supported case-specific breakpoint
Interface Organicness (IO)2.252.25–2.750.500.622−0.1220.00350.0035111.21Supported case-specific breakpoint
Signage Standardization (SS)1.250.75–3.753.000.5150.1130.00700.008419.78Supported, but relatively uncertain
Commercial Openness (CO)0.750.75–3.252.500.4770.0210.00700.012010.01Supported, but relatively uncertain
Chain brand Salience (CS)0.750.25–0.750.500.3760.0410.19600.01177.54Not formally supported; nonlinearity nonsignificant
Generic Decoration (GD)1.251.25–6.755.500.2280.1250.69440.7257−0.69Not supported
Local Dialect Intensity (LDI)1.251.25–8.257.00−0.046−0.0850.92800.9640−1.71Not supported
Commercial Broadcasting (CB)0.750.10–7.257.15−0.0380.0910.78910.7165−0.40Not supported
Sound event Diversity (SED)2.252.25–5.253.000.042−0.0820.92800.74850.01Not supported
Technophony Ratio (TR)3.751.75–7.906.150.0560.1350.47080.44331.98Not supported
Tourist activity Impact (TAI)0.400.15–1.751.60−0.1610.1030.62220.40950.38Not supported
Perceived Calmness (PC)6.252.25–6.254.00−0.090−0.0170.91350.7165−0.50Not supported
Note: CI = confidence interval; FDR = false discovery rate; ΔAIC = AIC of the linear model minus AIC of the segmented model. Positive ΔAIC values indicate that the segmented model performed better than the corresponding linear model. A breakpoint was considered formally supported only when both the nonlinear effect and the breakpoint test remained significant after FDR correction (q < 0.05) and the segmented model improved model fit (ΔAIC > 2). Breakpoints were estimated on the original 0–10 scale. The zero-width intervals for MV and AC reflect the discrete breakpoint-search grid and repeated selection of the same candidate value across bootstrap samples; they should not be interpreted as indicating absolute measurement precision.
Table A8. Multicollinearity diagnostics of the 14 visual and auditory indicators. (Source: Compiled by the authors).
Table A8. Multicollinearity diagnostics of the 14 visual and auditory indicators. (Source: Compiled by the authors).
Indicator TypeVariableVIFToleranceMost Related VariablePearson rSpearman ρ
Visualmaterial_vernacularity4.3960.227facade_patina0.7550.756
Visualfacade_patina3.5320.283material_vernacularity0.7550.756
Visualinterface_organicness2.5920.386facade_patina0.6880.598
Auditoryvoice_perceived_calmness2.2770.439voice_technophony_ratio−0.641−0.635
Visualcommercial_openness2.0190.495chain_brand_salience0.4230.633
Visualarchitectural_complexity2.0000.500material_vernacularity0.5370.526
Visualsignage_standardization1.9770.506generic_decoration0.6060.610
Auditoryvoice_technophony_ratio1.9300.518voice_perceived_calmness−0.641−0.635
Visualgeneric_decoration1.9100.524signage_standardization0.6060.610
Auditoryvoice_sound_event_diversity1.7230.580voice_local_dialect_intensity0.5670.599
Auditoryvoice_local_dialect_intensity1.6900.592voice_sound_event_diversity0.5670.599
Auditoryvoice_commercial_broadcasting1.6680.600voice_tourist_activity_impact0.4750.605
Visualchain_brand_salience1.4900.671commercial_openness0.4230.633
Auditoryvoice_tourist_activity_impact1.4340.697voice_commercial_broadcasting0.4750.605
Note: VIF = variance inflation factor. Tolerance = 1/VIF. A VIF value below 5 indicates that severe multicollinearity is unlikely.
Figure A3. Spatial-Block Bootstrap Estimates and 95% Confidence Intervals of Empirical Breakpoints for the 14 Visual and Auditory Indicators(Source: authors).
Figure A3. Spatial-Block Bootstrap Estimates and 95% Confidence Intervals of Empirical Breakpoints for the 14 Visual and Auditory Indicators(Source: authors).
Buildings 16 02871 g0a3
Table A9. Moderate and high pairwise correlations among the 14 indicators. (Source: Compiled by the authors).
Table A9. Moderate and high pairwise correlations among the 14 indicators. (Source: Compiled by the authors).
Variable 1Variable 2Pearson rSpearman ρCorrelation Level
Facade patinamaterial_vernacularity0.7550.756High
Voice Technophony ratiovoice_perceived_calmness−0.641−0.635Moderate negative
Chain brand saliencecommercial_openness0.4230.633Moderate
Signage standardizationgeneric_decoration0.6060.610Moderate
commercial broadcastingvoice_tourist_activity_impact0.4750.605Moderate
local dialect intensityvoice_sound_event_diversity0.5670.599Moderate
facade patinainterface_organicness0.6880.598Moderate
material vernacularityinterface_organicness0.6130.581Moderate
material vernacularityarchitectural_complexity0.5370.526Moderate
local dialect intensityvoice_tourist_activity_impact0.2660.503Moderate in rank correlation

References

  1. Graham, B.J. Senses of place, senses of time and heritage. In Senses of Place: Senses of Time; Ashworth, G.J., Graham, B., Eds.; Ashgate Publishing: Aldershot, UK; Burlington, VT, USA, 2005; pp. 3–14. [Google Scholar]
  2. Zhang, R.; Martí Casanovas, M.; Bosch González, M.; Sun, S. Revitalizing heritage: The role of urban morphology in creating public value in China’s historic districts. Land 2024, 13, 1919. [Google Scholar] [CrossRef] [Scilit]
  3. International Council on Monuments and Sites. Charter for the Conservation of Historic Towns and Urban Areas (Washington Charter); ICOMOS: Paris, France, 1987; Available online: https://www.icomos.org/images/DOCUMENTS/Charters/towns_e.pdf (accessed on 12 July 2026).
  4. United Nations Educational, Scientific and Cultural Organization. Recommendation on the Historic Urban Landscape, Including a Glossary of Definitions; UNESCO: Paris, France, 2011; Available online: https://www.unesco.org/en/legal-affairs/recommendation-historic-urban-landscape-including-glossary-definitions (accessed on 12 July 2026).
  5. Pulles, K.; Conti, I.A.M.; de Kleijn, M.B.; Kusters, B.; Rous, T.; Havinga, L.C.; Ikiz Kaya, D. Emerging strategies for regeneration of historic urban sites: A systematic literature review. City Cult. Soc. 2023, 35, 100539. [Google Scholar] [CrossRef] [Scilit]
  6. State Council of the People’s Republic of China. Regulations on the Protection of Famous Historical and Cultural Cities, Towns and Villages; Order No. 524 of the State Council of the People’s Republic of China; State Council of the People’s Republic of China: Beijing, China, 2008; Revised 2017. Available online: https://xzfg.moj.gov.cn/front/law/detail?LawID=212 (accessed on 12 July 2026). (In Chinese)
  7. Standing Committee of the Yunnan Provincial People’s Congress. Regulations of Yunnan Province on the Protection of Famous Historical and Cultural Cities, Towns, Villages and Streets; Announcement No. 65 of the Standing Committee of the Tenth Yunnan Provincial People’s Congress; Standing Committee of the Yunnan Provincial People’s Congress: Kunming, China, 2007; Amended 2012. Available online: https://www.cxs.gov.cn/info/1562/16685.htm (accessed on 12 July 2026). (In Chinese)
  8. Chen, Y.; Wang, Y.-W. Approaches to sustaining people–place bonds in conservation planning: From value-based, living heritage, to the glocal community. Built Herit. 2024, 8, 10. [Google Scholar] [CrossRef] [Scilit]
  9. Kuang, Z.; Zhang, J.; Li, Y.; Fukuda, T. Preserving architectural heritage in urban renewal: A stable diffusion model framework for automated historical facade generation. npj Herit. Sci. 2025, 13, 256. [Google Scholar] [CrossRef] [Scilit]
  10. Xie, K.; Xiong, R.; Bai, Y.; Zhang, M.; Zhang, Y.; Han, W. Traditional architectural heritage conservation and green renovation with eco materials: Design strategy and field practice in cultural Tibetan town. Sustainability 2024, 16, 6834. [Google Scholar] [CrossRef] [Scilit]
  11. Yao, W.; Miao, M.; Ding, Z.; Wu, Y.; Zhan, M. Commercial color impact on traditional heritage features in Suzhou Shiquan historic district. npj Herit. Sci. 2025, 13, 395. [Google Scholar] [CrossRef] [Scilit]
  12. Hu, Y.; Meng, Q.; Li, M.; Yang, D. Enhancing authenticity in historic districts via soundscape design. Herit. Sci. 2024, 12, 396. [Google Scholar] [CrossRef] [Scilit]
  13. Liu, X. The effects of commercialisation on urban heritage in Tianjin: A study of citizens’ livelihood in the Five Avenues (Wudadao) historical district. Built Herit. 2024, 8, 42. [Google Scholar] [CrossRef] [Scilit]
  14. Portella, A. Evaluating the effects of commercial signs on the appearance of historic streetscapes in different countries. J. Res. Archit. Plan. 2007, 6, 10–20. [Google Scholar]
  15. Zhu, Y.; González Martínez, P. Heritage, values and gentrification: The redevelopment of historic areas in China. Int. J. Herit. Stud. 2022, 28, 476–494. [Google Scholar] [CrossRef] [Scilit]
  16. Biljecki, F.; Ito, K. Street view imagery in urban analytics and GIS: A review. Landsc. Urban Plan. 2021, 215, 104217. [Google Scholar] [CrossRef] [Scilit]
  17. Xu, H.; Sun, H.; Wang, L.; Yu, X.; Li, T. Urban architectural style recognition and dataset construction method under deep learning of street view images: A case study of Wuhan. ISPRS Int. J. Geo-Inf. 2023, 12, 264. [Google Scholar] [CrossRef] [Scilit]
  18. Huang, K.; Kang, P.; Zhao, Y. Quantitative research of street interface morphology in urban historic districts: A case study of West Street Historic District, Quanzhou. Herit. Sci. 2024, 12, 226. [Google Scholar] [CrossRef] [Scilit]
  19. Huang, G.; Yu, Y.; Lyu, M.; Sun, D.; Dewancker, B.; Gao, W. Impact of physical features on visual walkability perception in urban commercial streets by using street-view images and deep learning. Buildings 2025, 15, 113. [Google Scholar] [CrossRef] [Scilit]
  20. Zhong, T.; Ye, C.; Wang, Z.; Tang, G.; Zhang, W.; Ye, Y. City-scale mapping of urban façade color using street-view imagery. Remote Sens. 2021, 13, 1591. [Google Scholar] [CrossRef] [Scilit]
  21. Yin, L.; Wang, Z. Measuring visual enclosure for street walkability: Using machine learning algorithms and Google Street View imagery. Appl. Geogr. 2016, 76, 147–153. [Google Scholar] [CrossRef] [Scilit]
  22. Kelly, C.M.; Wilson, J.S.; Baker, E.A.; Miller, D.K.; Schootman, M. Using Google Street View to audit the built environment: Inter-rater reliability results. Ann. Behav. Med. 2013, 45, S108–S112. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Li, Y.; Peng, L.; Wu, C.; Zhang, J. Street View Imagery (SVI) in the built environment: A theoretical and systematic review. Buildings 2022, 12, 1167. [Google Scholar] [CrossRef] [Scilit]
  24. Ye, Y.; Zeng, W.; Shen, Q.; Zhang, X.; Lu, Y. The visual quality of streets: A human-centred continuous measurement based on machine learning algorithms and street view images. Environ. Plan. B Urban Anal. City Sci. 2019, 46, 1439–1457. [Google Scholar] [CrossRef] [Scilit]
  25. Wang, Z.; Zhang, W.; Huang, Y. Nonlinear perceptual thresholds and trade-offs of visual environment in historic districts: Evidence from street view images in Shanghai. Sustainability 2025, 17, 11075. [Google Scholar] [CrossRef] [Scilit]
  26. Yu, M.; Chen, X.; Zheng, X.; Cui, W.; Ji, Q.; Xing, H. Evaluation of spatial visual perception of streets based on deep learning and spatial syntax. Sci. Rep. 2025, 15, 18439. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Huang, L.; Kang, J. The sound environment and soundscape preservation in historic city centres—The case study of Lhasa. Environ. Plan. B Plan. Des. 2015, 42, 652–674. [Google Scholar] [CrossRef] [Scilit]
  28. Liu, J.; Yang, L.; Xiong, Y.; Yang, Y. Effects of soundscape perception on visiting experience in a renovated historical block. Build. Environ. 2019, 165, 106375. [Google Scholar] [CrossRef] [Scilit]
  29. Ye, J.; Li, S.; Chen, Y.; Ma, Y.; Chen, L.; He, T.; Zheng, Y. A study of the effects of historical block context on soundscape perception. Buildings 2024, 14, 621. [Google Scholar] [CrossRef] [Scilit]
  30. Truax, B. Acoustic Communication, 2nd ed.; Ablex Publishing: Westport, CT, USA, 2001. [Google Scholar]
  31. Standing Committee of the National People’s Congress of the People’s Republic of China. Law of the People’s Republic of China on Noise Pollution Prevention; Order No. 104 of the President of the People’s Republic of China; Standing Committee of the National People’s Congress of the People’s Republic of China: Beijing, China, 2021; Effective 5 June 2022. (In Chinese)
  32. GB 3096-2008; Environmental Quality Standard for Noise. China Environmental Science Press: Beijing, China, 2008. (In Chinese)
  33. Gan, Y.; Luo, T.; Breitung, W.; Kang, J.; Zhang, T. Multi-sensory landscape assessment: The contribution of acoustic perception to landscape evaluation. J. Acoust. Soc. Am. 2014, 136, 3200–3210. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Jeon, J.Y.; Jo, H.I. Effects of audio-visual interactions on soundscape and landscape perception and their influence on satisfaction with the urban environment. Build. Environ. 2020, 169, 106544. [Google Scholar] [CrossRef] [Scilit]
  35. Relph, E. Place and Placelessness; Pion: London, UK, 1976. [Google Scholar]
  36. Norberg-Schulz, C. Genius Loci: Towards a Phenomenology of Architecture; Rizzoli: New York, NY, USA, 1980. [Google Scholar]
  37. Baudrillard, J. The Consumer Society: Myths and Structures; Sage: London, UK, 1998. [Google Scholar]
  38. Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef] [Scilit]
  39. Chen, T.; Guestrin, C. XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 785–794. [Google Scholar] [CrossRef] [Scilit]
  40. Yan, Y.; Xu, X.; Huang, Y.; Zhang, Q. Micro-scale land-use functional patterns and driving mechanisms in historic districts: A multi-source GIS and interpretable machine learning approach. Front. Environ. Sci. 2026, 14, 1784087. [Google Scholar] [CrossRef] [Scilit]
  41. Huang, Y.; Shi, Z.; Chen, Y.; Zhou, K.; Ying, Z.; Zheng, L.; Yang, S. Urban form evolution and transformation of traditional water towns based on space syntax and GIS: Evidence from ancient Wenzhou city (16th to the 20th century). Front. Earth Sci. 2025, 13, 1520643. [Google Scholar] [CrossRef] [Scilit]
  42. Huang, Y.; Huang, Y.; Chen, Y.; Song, J.; Yang, S.; Huang, L.; Zheng, L.; Gao, Y. The evolution and construction of Shan-shui cities: Evidence from the ancient city of Hangzhou from the sixth to the twenty-first century via geographical information systems and space syntax. Front. Earth Sci. 2025, 13, 1551117. [Google Scholar] [CrossRef] [Scilit]
  43. Zhu, Y.; Huang, Y. HGIS-based analysis of urban morphological evolution in historic Kaifeng. npj Herit. Sci. 2026, 14, 32. [Google Scholar] [CrossRef] [Scilit]
  44. Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30; Guyon, I., von Luxburg, U., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2017; pp. 4765–4774. [Google Scholar]
Figure 1. Location and current conditions of the Ximen Street Historic District, Qujing City: (a) location of Qujing City in Yunnan Province; (b) location and geographical setting of Qilin District within Qujing City; (c) location of the Ximen Street Historic District within Qilin District; and (d) current conditions of the Ximen Street Historic District (Source: Compiled by the authors based on materials provided by the Qujing Planning Bureau).
Figure 1. Location and current conditions of the Ximen Street Historic District, Qujing City: (a) location of Qujing City in Yunnan Province; (b) location and geographical setting of Qilin District within Qujing City; (c) location of the Ximen Street Historic District within Qilin District; and (d) current conditions of the Ximen Street Historic District (Source: Compiled by the authors based on materials provided by the Qujing Planning Bureau).
Buildings 16 02871 g001
Figure 2. Sampling design, streetscape image examples and spatial structure of the Ximen Street Historic Area: (a) distribution of 340 sampling points and workflow of field-collected streetscape image acquisition; (b) representative streetscape images of different spatial types; and (c) spatial classification of streets, lanes, public spaces and green spaces (Source: Drawn and photographed by the authors).
Figure 2. Sampling design, streetscape image examples and spatial structure of the Ximen Street Historic Area: (a) distribution of 340 sampling points and workflow of field-collected streetscape image acquisition; (b) representative streetscape images of different spatial types; and (c) spatial classification of streets, lanes, public spaces and green spaces (Source: Drawn and photographed by the authors).
Buildings 16 02871 g002
Figure 3. Research framework (Source: authors).
Figure 3. Research framework (Source: authors).
Buildings 16 02871 g003
Figure 4. Human validation survey workflow for visual and soundscape homogenization perception (Source: authors).
Figure 4. Human validation survey workflow for visual and soundscape homogenization perception (Source: authors).
Buildings 16 02871 g004
Figure 5. Streetscape semantic composition and bivariate spatial patterns of visual homogenization indicators: (a) distribution of the top 25 semantic segmentation features; (b1) façade patina and material vernacularity; (b2) architectural complexity and interface organic-ness; (b3) signage standardization and chain brand salience; and (b4) generic decoration and commercial frontage openness (Source: authors).
Figure 5. Streetscape semantic composition and bivariate spatial patterns of visual homogenization indicators: (a) distribution of the top 25 semantic segmentation features; (b1) façade patina and material vernacularity; (b2) architectural complexity and interface organic-ness; (b3) signage standardization and chain brand salience; and (b4) generic decoration and commercial frontage openness (Source: authors).
Buildings 16 02871 g005
Figure 6. Spatial distribution of audio perception indicators: (a) local dialect intensity; (b) commercial broadcasting; (c) sound event diversity; (d) technophony ratio; (e) tourist activity impact; and (f) perceived calmness (Source: authors).
Figure 6. Spatial distribution of audio perception indicators: (a) local dialect intensity; (b) commercial broadcasting; (c) sound event diversity; (d) technophony ratio; (e) tourist activity impact; and (f) perceived calmness (Source: authors).
Buildings 16 02871 g006
Figure 7. Correlation structure and distributional characteristics of visual and audio homogenization indicators: (a) correlation matrix of visual and audio homogenization indicators; (b) distribution of visual indicators; (c) distribution of audio indicators (Source: authors).
Figure 7. Correlation structure and distributional characteristics of visual and audio homogenization indicators: (a) correlation matrix of visual and audio homogenization indicators; (b) distribution of visual indicators; (c) distribution of audio indicators (Source: authors).
Buildings 16 02871 g007
Figure 8. Spatial statistical evidence and interpolated surfaces of visual–audio homogenization mismatch: (a) spatial distribution of visual–audio mismatch; (b) joint distribution of visual and audio homogenization; (c) comparison between visual and audio homogenization indices; (d) mean-mismatch plot of visual–audio homogenization; (e) distribution of visual–audio mismatch values; (f1) interpolated surface of visual homogenization; (f2) interpolated surface of audio homogenization; and (f3) interpolated surface of visual–audio mismatch (Source: authors).
Figure 8. Spatial statistical evidence and interpolated surfaces of visual–audio homogenization mismatch: (a) spatial distribution of visual–audio mismatch; (b) joint distribution of visual and audio homogenization; (c) comparison between visual and audio homogenization indices; (d) mean-mismatch plot of visual–audio homogenization; (e) distribution of visual–audio mismatch values; (f1) interpolated surface of visual homogenization; (f2) interpolated surface of audio homogenization; and (f3) interpolated surface of visual–audio mismatch (Source: authors).
Buildings 16 02871 g008
Figure 9. K-means clustering results and cluster-specific distributions of multisensory homogenization features: (a) silhouette scores under different numbers of clusters, with K = 6 selected for typological interpretation; (b) t-SNE projection of the six cluster types; (c) cluster-specific distributions of visual, audio and overall homogenization indicators visualized using parallel coordinates (Source: authors). (FP = Facade Patina; MV = Material Vernacularity; AC = Architectural Complexity; IO = Interface Organic-ness; SS = Signage Standardization; CS = Chain brand Salience; GD = Generic Decoration; CO = Commercial Openness; LDI = Local Dialect Intensity; CB = Commercial Broadcasting; SED = Sound Event Diversity; TR = Technophony Ratio; TAI = Tourist Activity Impact; PC = Perceived Calmness; VH = Visual Homogenization; AH = Audio Homogenization).
Figure 9. K-means clustering results and cluster-specific distributions of multisensory homogenization features: (a) silhouette scores under different numbers of clusters, with K = 6 selected for typological interpretation; (b) t-SNE projection of the six cluster types; (c) cluster-specific distributions of visual, audio and overall homogenization indicators visualized using parallel coordinates (Source: authors). (FP = Facade Patina; MV = Material Vernacularity; AC = Architectural Complexity; IO = Interface Organic-ness; SS = Signage Standardization; CS = Chain brand Salience; GD = Generic Decoration; CO = Commercial Openness; LDI = Local Dialect Intensity; CB = Commercial Broadcasting; SED = Sound Event Diversity; TR = Technophony Ratio; TAI = Tourist Activity Impact; PC = Perceived Calmness; VH = Visual Homogenization; AH = Audio Homogenization).
Buildings 16 02871 g009
Figure 10. Spatial cross-validation and residual diagnosis of the XGBoost model: (a) observed versus predicted values under five-fold spatial cross-validation; (b) spatial folds used for spatial cross-validation; and (c) spatial distribution of prediction residuals.
Figure 10. Spatial cross-validation and residual diagnosis of the XGBoost model: (a) observed versus predicted values under five-fold spatial cross-validation; (b) spatial folds used for spatial cross-validation; and (c) spatial distribution of prediction residuals.
Buildings 16 02871 g010
Figure 11. SHAP global feature importance and sample-level heatmap for multisensory homogenization: (a) global contribution of visual and audio indicators to the model output; (b) distribution of SHAP. (Source: authors).
Figure 11. SHAP global feature importance and sample-level heatmap for multisensory homogenization: (a) global contribution of visual and audio indicators to the model output; (b) distribution of SHAP. (Source: authors).
Buildings 16 02871 g011aBuildings 16 02871 g011b
Figure 12. SHAP-based main and interaction effects of factors influencing multisensory homogenization: (a) feature importance and interaction network of visual and audio indicators; (b) comparison of main and interaction effects across indicators. (Source: authors).
Figure 12. SHAP-based main and interaction effects of factors influencing multisensory homogenization: (a) feature importance and interaction network of visual and audio indicators; (b) comparison of main and interaction effects across indicators. (Source: authors).
Buildings 16 02871 g012
Figure 13. SHAP dependence plots and threshold effects of factors influencing multisensory homogenization. Light blue histograms indicate the distribution of feature values, blue points represent SHAP values of samples, pink curves show the smoothed fitting results, shaded bands indicate the 95% confidence intervals, and labeled points denote zero-crossing thresholds of SHAP values (Source: authors).
Figure 13. SHAP dependence plots and threshold effects of factors influencing multisensory homogenization. Light blue histograms indicate the distribution of feature values, blue points represent SHAP values of samples, pink curves show the smoothed fitting results, shaded bands indicate the 95% confidence intervals, and labeled points denote zero-crossing thresholds of SHAP values (Source: authors).
Buildings 16 02871 g013
Table 1. Comparison of different types of historic districts in Yunnan.
Table 1. Comparison of different types of historic districts in Yunnan.
CaseLocationTypeTourism and Commercialization Level
Ximen Street Historic District,Qilin District,
Qujing City
Residential–incremental renewal typeMedium–low
Dayan Old Town Historic District,Gucheng District,
Lijiang City
High-intensity tourism–commercial typeHigh
Inner City Historic DistrictJianshui County,
Honghe Prefecture
Cultural tourism–residential mixed typeMedium–high
Weishan Prefectural City Historic DistrictWeishan County, Dali CityTraditional residential–moderate tourism typeMedium–low
Huize Ancient City Historic DistrictHuize County, Qujing CityTraditional administrative–commercial mixed typeModerate
Table 2. Independent variable framework and 0–10 scoring criteria for multisensory homogenization in historic districts (Source: Compiled by the authors).
Table 2. Independent variable framework and 0–10 scoring criteria for multisensory homogenization in historic districts (Source: Compiled by the authors).
ModalityDimensionVariableScoring Criteria (0–10)Direction
VisualVisual authenticity retentionFaçade Patina0 = newly renovated; 10 = strong historic patinaNegative
Material Vernacularity0 = modern/industrial materials; 10 = local/traditional materialsNegative
Architectural Complexity0 = plain façade; 10 = rich traditional detailsNegative
Interface Organic-ness0 = regularized interface; 10 = organic, varied interfaceNegative
Commercial landscape standardizationSignage Standardization0 = diverse local signs; 10 = highly standardized signsPositive
Chain Brand Salience0 = no chain brands; 10 = chain brands dominatePositive
Generic Decoration0 = local/traditional decoration; 10 = generic/mass-produced decorationPositive
Commercial Frontage Openness0 = residential or daily life frontage; 10 = continuous open commercial frontagePositive
Visual Homogenization Index0 = low, 10 = high
AuditorySoundscape locality retentionLocal Dialect Intensity0 = no identifiable local dialect; 10 = local dialect clearly presentNegative
Sound Event Diversity0 = few sound event types; 10 = rich and diverse sound eventsNegative
Commercial technological sound intrusionCommercial Broadcasting Intensity0 = absent; 10 = continuous or looped broadcastingPositive
Tourist Activity Impact0 = absent; 10 = tourist-related sounds dominatePositive
Technophony Ratio0 = minimal mechanical sounds; 10 = traffic or equipment sounds dominatePositive
Perceived Calmness0 = chaotic/disturbed; 10 = calm and clearly layeredNegative
Auditory Homogenization Index0 = low, 10 = high
Note: The values shown are raw indicator scores. For variables marked as negative, higher raw scores indicate stronger heritage authenticity or soundscape locality and therefore a lower expected homogenization risk. For variables marked as positive, higher raw scores indicate greater homogenization risk. When calculating the visual and auditory homogenization indices, negative variables were reverse-coded using x r e v = 10 x , so that higher composite index values consistently indicated greater homogenization.
Table 3. Training and test performance and generalization gaps of eight regression models.
Table 3. Training and test performance and generalization gaps of eight regression models.
ModelTrain R2ΔR2Train RMSETrain MAE
MLR0.8540.0710.2580.195
Ridge0.8540.0700.2580.194
SVR0.9960.1110.0450.043
Random Forest0.9790.1270.0970.072
Extra Trees1.0000.1210.0000.000
XGBoost0.9980.1110.0290.022
LightGBM0.9950.1230.0480.030
CatBoost0.9960.0900.0440.034
Table 4. Comparative performance of eight regression models on the independent test set and five-fold cross-validation.
Table 4. Comparative performance of eight regression models on the independent test set and five-fold cross-validation.
ModelTest R2Test RMSETest MAE5-Fold CV R25-Fold CV RMSE5-Fold CV MAE
MLR0.7830.3100.2160.821 ± 0.0620.281 ± 0.0490.207 ± 0.027
Ridge0.7840.3090.2160.821 ± 0.0620.280 ± 0.0490.206 ± 0.027
SVR0.8850.2260.1280.911 ± 0.0260.198 ± 0.0300.134 ± 0.017
Random Forest0.8520.2560.1680.843 ± 0.0270.266 ± 0.0300.194 ± 0.023
Extra Trees0.8790.2310.1450.874 ± 0.0220.238 ± 0.0250.171 ± 0.019
XGBoost0.8860.3450.2280.908 ± 0.0340.201 ± 0.0400.131 ± 0.017
LightGBM0.8720.2380.1350.883 ± 0.0420.227 ± 0.0460.153 ± 0.026
CatBoost0.9060.2040.1000.932 ± 0.0340.170 ± 0.0430.108 ± 0.012
Table 5. Global Moran’s I results for visual and auditory homogenization indices under different distance thresholds (Source: Compiled by the authors).
Table 5. Global Moran’s I results for visual and auditory homogenization indices under different distance thresholds (Source: Compiled by the authors).
TypeDistanceMoran’s IExpected Indexz-Scorep-ValueRole
Visual40 m0.448925−0.00295010.067740p < 0.001Sensitivity check
50 m0.384152−0.00295011.947115p < 0.001Main analytical distance
60 m0.365670−0.00295012.591266p < 0.001Sensitivity check
Auditory40 m0.212447−0.0029500.002014p < 0.001Sensitivity check
50 m0.166666−0.0029500.001050p < 0.001Main analytical distance
60 m0.128422−0.0029500.000857p < 0.001Sensitivity check
Table 6. Comparison of alternative K-means cluster solutions. (Source: Compiled by the authors).
Table 6. Comparison of alternative K-means cluster solutions. (Source: Compiled by the authors).
Number of ClustersSilhouette CoefficientWithin-Cluster Sum of Squares (WCSS)Calinski–Harabasz IndexDavies–Bouldin Index
K = 20.18574005.4763.672.1949
K = 30.12313655.3150.922.1233
K = 40.13463334.0147.901.9874
K = 50.15313054.5846.761.8174
K = 60.15782791.2147.121.7230
K = 70.15592614.2745.551.6855
K = 80.15462475.7143.761.6190
K = 90.15332376.2141.511.6553
K = 100.16012275.7240.031.6536
Table 7. Model performance under random train–test split and five-fold spatial cross-validation.
Table 7. Model performance under random train–test split and five-fold spatial cross-validation.
Validation MethodFoldTraining SamplesTesting SamplesR2RMSEMAE
Random train–test split0.8860.3450.228
Spatial cross-validation1259810.9240.2950.222
2266740.8260.4080.266
3268720.9160.2730.205
4281590.8810.3070.238
5286540.7930.4390.300
Mean ± SD0.868 ± 0.0570.344 ± 0.0740.246 ± 0.038
Table 8. Governance priorities and feasibility considerations for different multisensory homogenization types (Source: Compiled by the authors).
Table 8. Governance priorities and feasibility considerations for different multisensory homogenization types (Source: Compiled by the authors).
TypeGovernance FocusFeasibility Considerations
VisualSoundscapeCoordinated MeasuresCostStakeholder ResistanceLegal
Framework
1Maintain traditional materials, façade details, and street interfaces; control new standardized signsPreserve resident conversations, local dialects, and daily activity sounds; restrict new amplified sound sourcesPrioritize preventive conservation and avoid intensive commercial insertionLow–mediumLow
mainly new businesses
Conservation planning; shop sign rules
2Maintain existing façade and shop sign coherence; avoid excessive Scenographic treatmentRegulate broadcasting time, volume, and loudspeaker orientationPrioritize source control and prevent soundscape disturbance from spreading to adjacent areasLow–mediumMedium
broadcasting-dependent businesses
Urban management; noise control
3Retain traditional materials, architectural details, and street-interface scaleOptimize equipment placement, vehicle stopping areas, and apply vibration reduction or concealed installationReduce technical noise without compromising traditional characterMedium–highMedium
cost and construction disturbance
Equipment rules; fire safety; noise control
4Retain diverse façades, materials, and interfaces; avoid uniform renewalControl the cross-boundary spread of broadcasting, equipment noise, and traffic noiseProvide soundscape buffers at lanes, street corners, and courtyard entrancesMediumMedium
multiple sound-source actors
Public-space management; noise control
5Regulate chain brand signage, generic decoration, and overly open commercial interfaces; reinforce local materialsControl amplified broadcasting, background music, equipment noise, and high-intensity commercial soundsSimultaneously address high-risk visual and auditory elements through priority renewal reviewHighHigh
brand visibility and income concerns
Urban design; shop sign rules; noise control
6Maintain traditional façades, local materials, and street scalePreserve resident conversations, dialect use, and daily activity sounds while reducing broadcasting and equipment noiseDifferentiate daily sounds from disturbance sources and maintain soundscape layeringMediumMedium
different sound perceptions
Community governance; noise control
Note: Type 1–6 correspond respectively to the six clusters: low commercial disturbance, commercial soundscape intensification, technical sound disturbance, soundscape convergence, high commercial homogenization, and mixed soundscape types. The table offers implementation-oriented references, which should be adapted to local planning, funding, management capacity, and stakeholder negotiation.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Qin, Y.; Yang, D.; Huang, Y. When Visual Conservation Meets Auditory Homogenization: Multimodal Evidence from Ximen Street Historic District, Qujing, China. Buildings 2026, 16, 2871. https://doi.org/10.3390/buildings16142871

AMA Style

Qin Y, Yang D, Huang Y. When Visual Conservation Meets Auditory Homogenization: Multimodal Evidence from Ximen Street Historic District, Qujing, China. Buildings. 2026; 16(14):2871. https://doi.org/10.3390/buildings16142871

Chicago/Turabian Style

Qin, Yuxin, Dayu Yang, and Yuhao Huang. 2026. "When Visual Conservation Meets Auditory Homogenization: Multimodal Evidence from Ximen Street Historic District, Qujing, China" Buildings 16, no. 14: 2871. https://doi.org/10.3390/buildings16142871

APA Style

Qin, Y., Yang, D., & Huang, Y. (2026). When Visual Conservation Meets Auditory Homogenization: Multimodal Evidence from Ximen Street Historic District, Qujing, China. Buildings, 16(14), 2871. https://doi.org/10.3390/buildings16142871

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop