Next Article in Journal
When Advice Isn’t Trusted: Privacy, Transparency, and Accountability Risks Driving AI Mistrust and Consumer Resistance in Financial Advisory Services
Previous Article in Journal
How Corporate FinTech Enhances ESG Performance: An Integrated Framework of Resources, Technology, and Governance
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Semantic Segmentation for Walkability Assessment in Southeast Asian Streetscapes

1
Department of Housing and Interior Design, Chungbuk National University, Cheongju 28644, Republic of Korea
2
Lee Kuan Yew Centre for Innovative Cities, Singapore University of Technology and Design, Singapore 487372, Singapore
*
Author to whom correspondence should be addressed.
Sustainability 2026, 18(3), 1355; https://doi.org/10.3390/su18031355
Submission received: 12 December 2025 / Revised: 17 January 2026 / Accepted: 26 January 2026 / Published: 29 January 2026

Abstract

Walkable urban environments are increasingly recognized as essential for sustainable mobility, public health, and social well-being. While macro-scale indicators of walkability are widely used, growing evidence highlights the importance of street-level physical conditions experienced at eye level. Advances in computer vision and street view imagery (SVI) offer new opportunities to quantify such streetscape characteristics, yet the applicability of existing semantic segmentation models in developing urban contexts remains underexplored. This study evaluates the suitability of five state-of-the-art semantic segmentation models for streetscape analysis using crowdsourced SVI from Phnom Penh, Cambodia. Through a comparative analysis, Oneformer was identified as the most suitable semantic segmentation model, uniquely successful in identifying street vendors through surrogate semantic class (base) and street furniture. A rigorous quantitative validation using manually annotated images confirmed the model’s reliability, achieving an mIoU of 65.7% within the complex urban fabric of Phnom Penh. This performance stems from OneFormer’s unified task-conditioned framework, which integrates semantic, instance, and panoptic information within a single query. Such an architecture ensures enhanced boundary stability and semantic coherence by consolidating visual noise into meaningful units, making it particularly robust for processing the irregular street elements typical of Southeast Asian cities. Applying the selected model revealed pronounced spatial variation in streetscape composition across three neighborhoods, reflecting distinct development stages and levels of informality. These findings suggest that carefully selected pretrained models can yield analytically useful representations of streetscape conditions in data-constrained settings, supporting more context-sensitive and inclusive urban analysis in rapidly developing cities.

1. Introduction

Cities increasingly recognize that walkable urban environments are essential for supporting sustainable mobility, public health, and social well-being. Streets, in particular, play a crucial role in facilitating social interaction, accommodating diverse human activities, and connecting residents with the built environment [1]. Beyond serving as mobility corridors, streets function as public spaces where movement, social engagement, commercial activity, and environmental exposure occur simultaneously. As global planning agendas shift from vehicle-centric to pedestrian-friendly urban design [2], understanding the physical characteristics of streetscapes that shape walking behavior has become a critical priority for urban planners and public health researchers.
A large body of research shows that walking is influenced by multiple built environment factors, including destination accessibility, street network configuration, land-use diversity, and the quality of streetscape elements [3]. Among these factors, streetscape features are particularly relevant because they reflect the immediate physical conditions experienced at eye level. Recent studies further demonstrate that eye-level streetscape features extracted from street view imagery (SVI) are consistently associated with walking behavior [4,5,6]. These findings underscore the importance of developing systematic approaches to quantify streetscape characteristics to better understand how street environments function and how they shape pedestrian experience.
Traditional approaches to measuring walkability, including field audits, systematic social observation, and Geographic Information System (GIS)-based spatial indicators, have contributed to valuable insights but face practical and methodological limitations. Field-based assessments are labor-intensive and difficult to scale across large urban areas [7], while GIS-based measures often fail to capture nuanced eye-level visual cues that influence pedestrian perception and comfort [8]. These challenges have led to a growing interest in computational approaches that leverage large-scale SVI for urban analysis.
In recent years, advances in computer vision, particularly semantic segmentation, have enabled the automated extraction of detailed streetscape features from SVI at unprecedented spatial scale. However, most existing applications have focused on regulated urban environments in high-income contexts. Consequently, substantial uncertainty remains regarding the transferability of these methods to rapidly developing cities characterized by informal street activities and heterogeneous built forms.
To address this gap, this study examines the transferability of semantic segmentation to the complex urban environment of Phnom Penh, Cambodia. We focus on identifying the technical challenges and opportunities of using pretrained models to capture both formal and informal street components, characterizing physical streetscape conditions as environmental proxies for the pedestrian environment rather than predicting individual walking behaviors. This approach allows for a more reliable quantification of streetscapes in contexts where official spatial data are often outdated or unavailable.
The contribution of this study is twofold. First, from a methodological perspective, it provides a rigorous suitability assessment of five state-of-the-art semantic segmentation models to identify more robust architecture for capturing fragmented informal street elements. This includes a replicable workflow for extracting pedestrian-level indicators in data-constrained environments. Second, from an empirical perspective, the study offers a comparative analysis of streetscape compositions across three distinct neighborhoods in Phnom Penh, revealing how varying levels of urban informality and planned development shape the pedestrian environment. In doing so, this research contributes to the development of scalable, visually grounded methods for understanding street-level conditions in diverse and rapidly changing urban contexts.
The remainder of the paper is organized as follows. Section 2 reviews the conceptual background on walkability and urban streetscape analysis, together with recent advances in the use of SVI and semantic segmentation for walkability. Section 3 describes the study area, data sources, and analytical methods adopted in the study. Section 4 presents the result of the semantic segmentation analysis, including model suitability assessment and neighborhood-level streetscape characteristics. Section 5 discusses the findings in relation to walkability and the broader implications for streetscape analysis in data-constrained urban contexts. Finally, Section 6 concludes the paper by summarizing key contributions, acknowledging limitations, and outlining directions for future research.

2. Background

2.1. Walkability and Urban Streetscape

Walkability has become a primary concern in urban planning due to its links with sustainable mobility, public health, and social well-being. Walkability generally refers to the extent to which the built environment enables or encourages walking behavior and physical activity [9]. Walking behavior is influenced by a wide range of built environment attributes, often summarized through the widely used 5Ds framework: density, diversity, design, destination accessibility, and distance to transit [10]. These macro-scale neighborhood indicators have long served as a foundation for understanding travel behavior and walkability.
However, growing evidence suggests that macro-scale indicators of the built environment alone are insufficient for characterizing pedestrian experiences and walking behavior. Street-level conditions, reflecting what people see and encounter at eye level, play a critical role in shaping walking decisions. Beyond the simple identification of physical components, contemporary site analysis in urban design and planning emphasizes a micro-scale approach that perceives the contextual interaction between the built environment and user behavior [11]. From this perspective, street-level elements including informal street vendors function as essential micro-scale nodes that shape the qualitative pedestrian experience far more than mere physical metrics might suggest. Previous studies demonstrate that walking is strongly influenced by the quality of streetscape features, such as building frontages, vegetation, street furniture, and visual openness [6,7,12,13]. These elements constitute the immediate micro-scale environment through which pedestrians navigate and interact with urban space.
Increasingly, empirical research highlights the importance of micro-scale streetscape conditions in explaining walking behavior. Differences in street-level environmental features have been shown to account for variation in walking levels even among neighborhoods with similar macro-scale characteristics [4]. Eye-level streetscape features extracted from SVI have also been found to exhibit strong association with walking behavior, underscoring the explanatory value of visually perceived street conditions [5]. Moreover, recent work suggests that neither macro-scale nor micro-scale attributes alone are sufficient to fully represent walkability. By integrating macro-level urban form indicators with micro-level streetscape features derived from SVI using computer vision, composite walkability indices align more closely with pedestrian-rated walking environment satisfaction than indices based solely on macro-scale measures [14].
Together, this body of research indicates that a comprehensive understanding of walkability requires analytical approaches capable of capturing both macro-level urban form and micro-scale streetscape conditions with high visual fidelity. This recognition has motivated the growing adoption of large-scale visual data and computer vision techniques in urban research, enabling more direct and systematic representations of pedestrians’ everyday street-level experiences.

2.2. Semantic Segmentation for Streetscape Analysis

Semantic segmentation has emerged as a key computer vision technique for analyzing the urban environment, as it enables pixel-level classification of streetscape elements directly from visual data. When applied to SVI, segmentation provides an eye-level representation of urban environments that closely align with pedestrians’ visual experience. The increasing use of semantic information extracted from SVI in city planning and urban analytics research has been widely documented [15], reflecting its growing importance as a data source for understanding street-level urban conditions.
Recent advances in deep learning have further strengthened the integration of SVI and computer vision for urban planning and design [16]. In particular, semantic segmentation models have enabled large-scale, automated extraction of diverse streetscape features that are difficult to capture using conventional spatial datasets.
Existing applications can be broadly grouped into several thematic areas. First, segmentation has been widely used to quantify environmental and natural streetscape features, including street greenery, shading, and visual exposure to natural elements. These measures provide objective indicators of vegetation coverage and environmental conditions along streets and have been linked to health and sustainability outcomes [6,17,18].
Second, segmentation outputs have supported assessments of urban form and design quality by quantifying street-level characteristics such as visual enclosure, building frontages, and sky visibility. These features capture the spatial structure and visual coherence of streetscapes that shape pedestrian perception and movement [19,20,21].
Third, SVI-based segmentation has been applied to mapping mobility and active transport infrastructure, including sidewalks, pedestrian crossings, and bike lanes, and roadway space. These applications enable detailed evaluation of pedestrian- and cycling-related conditions at scale, supporting the analysis of active mobility environments [22,23].
Beyond physical characterization, segmentation-based analysis has also been extended to examine perceptual and experiential dimensions of street environments. Previous studies have used visual features extracted from SVI to estimate perceived safety, aesthetic quality, comfort, and street vitality, demonstrating how computer vision can link observable physical form with socially relevant street-level experiences [24,25,26].
Within this growing body of work, a consistent set of street-level components has been identified as particularly relevant to walkability. Semantic segmentation enables these components to be quantified directly from SVI, allowing systematic assessment of features that shape pedestrian experience. Previous studies indicate that continuous and unobstructed sidewalks, street greenery and tree canopy, and human-scale building frontages are positively associated with walking comfort and pedestrian activity, as they contribute to perceived safety, visual interest, and thermal comfort along streets [6,7,12]. Visual enclosure and balanced sky view, often captured through building-street proportions and sky visibility, have also been linked to more comfortable and legible walking environments [20,21].
In contrast, high road surface dominance, excessive vehicular presence, and fragmented pedestrian infrastructure are commonly associated with reduced walkability, reflecting environments that prioritize motorized movement over pedestrian use [22,23]. Importantly, many of these components correspond directly to semantic classes that can be reliably extracted through segmentation-based analysis of SVI. These findings suggest that semantic segmentation provides a suitable analytical bridge between visual street-level data and walkability-relevant indicators, enabling consistent measurement of physical streetscape attributes that have been repeatedly linked to pedestrian-oriented environments.
Despite these methodological advances, the application of SVI-based segmentation remains concentrated in North America, Europe, and high-income Asian cities. Consequently, empirical evidence on how pretrained segmentation models perform in urban contexts with substantially different street conditions remains limited. In many developing urban contexts, systematic walkability assessment is further constrained by limited institutional capacity, financial resources, and the time-intensive nature of conventional data collection methods. Detailed field audits and high-resolution spatial datasets often require levels of funding and administrative support that are difficult to sustain. Under such conditions, scalable approaches based on crowdsourced SVI and pretrained segmentation models hold particular promise, yet their reliability and applicability remain insufficiently examined.
Taken together, these limitations highlight the need for investigation into the reliability of existing semantic segmentation models for assessing streetscape conditions and walkability in developing urban contexts.

3. Materials and Methods

3.1. Study Area

Phnom Penh, the capital city of Cambodia, is characterized by pronounced heterogeneity in its streetscapes and the widespread presence of informal street-based economic activity. The city has an estimated population of approximately 2.04 million, with a density of 6925 persons per square kilometer [27]. Across many parts of the city, street space is intensively used by multiple actors, including pedestrians, motorbikes, and mobile street vendors, resulting in highly variable street-level conditions.
Street vendors constitute a prominent component of Phnom Penh’s informal urban economy. They are often highly mobile as they seek commercial opportunities while avoiding regulatory enforcement, which frequently regards street vending as an obstruction to traffic flow and goods movements [28]. Conflicts between local and migrant vendors further contribute to this mobility, often displacing migrant vendors from main roads to secondary streets and alleyways [28]. These dynamics shape the functional use of street space and contribute to uneven pedestrian conditions across the city. The widespread use of motorbikes and scooters reflects the need for flexible mobility within such constrained and frequently congested street environments.
As shown in Figure 1, the study focused on three neighborhoods—Boeng Keng Kang 1, Tonle Bassac, and Koh Pich—selected to represent contrasting streetscape morphologies, development stages, and levels of informal street activity. Boeng Keng Kang 1 is a dense inner-city area characterized by continuous commercial frontages and intensive pedestrian and street-vending activity. Tonle Bassac presents a heterogeneous urban fabric where automobile-oriented boulevards coexist with older residential streets, resulting in spatially uneven pedestrian environments. Koh Pich, a planned and partially developed artificial island, features wide road corridors and comparatively limited everyday street activity, reflecting its transitional stage of urban development.
The selection of these three neighborhoods was guided by the availability and spatial density of crowdsourced SVI. Compared to other parts of Phnom Penh, these neighborhoods contain relatively high concentrations of Mapillary images, enabling reliable semantic segmentation based on sufficient image coverage. Together, they capture a spectrum of informal street activity and urban form, providing a suitable basis for analyzing streetscape indicators and the visibility of informal economic practices.

3.2. Methods of Analysis

3.2.1. Obtaining and Filtering Streetscape Images

SVI was obtained using the Mapillary’s Application Programming Interface (API) (version 4.0; Mapillary AB, a subsidiary of Meta Platforms, Inc., Malmö, Sweden). Mapillary (https://www.mapillary.com/) is a crowdsourced platform that provides open-access, high-resolution georeferenced SVI for cities worldwide. For the Phnom Penh study area, a total of 24,986 images captured in 2023 were initially collected. Each image is associated with a unique image ID, which was used for retrieval through the Mapillary API. The data collection procedure involved generating multiple spatial bounding boxes across the study area, extracting their geographic coordinates, querying the Mapillary API for all available image IDs within each bounding box, and downloading the corresponding images.
Following data collection, all images were manually screened to ensure analytical suitability to mitigate sampling bias inherent in crowdsourced data, which tends to overrepresent frequently traveled routes. Images were excluded if they exhibited poor visual quality (e.g., blurring or discoloration), a limited field of view due to major obstructions, road segments inaccessible to pedestrians (e.g., expressways or motorways) and indoor environments. In addition, near-duplicate images captured from closely spaced locations or identical viewpoints along the same travel routes were removed. This filtering was intended to prevent repeated captures from disproportionately influencing neighborhood-level averages and biasing streetscape indicators toward frequently traveled segments. Initial screening was conducted by one researcher, and ambiguous cases were subsequently discussed and resolved through author consensus.
Because Mapillary image availability is user-driven and spatially irregular rather than uniformly distributed, sampling density varied across neighborhoods according to differences in street accessibility and image coverage. In particular, Boeng Keng Kang 1 has a relatively small spatial extent but a high density of streets, which likely contributed to its higher image density. After filtering, a final dataset of 8912 images was retained, including 4600 images for Boeng Keng Kang 1, 3002 for Tonle Bassac, and 1310 for Koh Pich. The resulting image counts and sampling densities by neighborhood are summarized in Table 1.

3.2.2. Semantic Segmentation and Model Implementation

Semantic segmentation was employed to extract pixel-level semantic information from SVI images. Recent advances in deep learning have substantially improved the accuracy of semantic scene understanding, enabling reliable extraction of streetscape segmentation elements from large-scale visual data. In this study, five state-of-the-art semantic segmentation models were evaluated. These models are among the highest-performing architectures available, as reflected in their mean Intersection over Union (mIoU) scores on standard benchmark datasets. The mIoU is a standard metric used to evaluate segmentation accuracy by measuring the overlap between predicted and ground-truth labels, with higher values indicating better performance. To ensure consistency in model comparison, all five models were pretrained on the ADE20K dataset. ADE20K is a large scene-centric dataset with pixel-level annotations across 150 semantic classes, including key streetscape elements such as sky, road, grass, vehicles, and pedestrians [29,30].
All segmentation models were deployed using the Hugging Face Transformers library (version 4.28.1; Hugging Face, Inc., New York, NY, USA). Although originally developed for natural language processing (NLP), Hugging Face (https://huggingface.co/) also provides a unified framework for deploying vision-based transformer models, including semantic segmentation architectures. The platform enables standardized model loading, interface, and output handling through a unified API, avoiding the need for complex manual installations from source code repositories. This ensured consistent implementation and reproducibility across all tested models.
The five evaluated models include:
  • OneFormer, a universal transformer-based framework for multi-task image segmentation, trained using task-conditioned joint learning across semantic, instance, and panoptic labels [31];
  • BeiT-L, a vision transformer architecture pretrained through masked image modelling to enhance semantic feature representation, enabling competitive performance in semantic segmentation tasks [32];
  • MaskFormer, a mask classification-based framework that predicts sets of binary masks linked to global class labels, providing a simplified and effective approach to semantic and panoptic segmentation [33];
  • Mask2Former, an extension of MaskFormer that incorporates masked attention to improve localized feature extraction within mask regions, enhancing segmentation accuracy across diverse segmentation tasks [33];
  • SegFormer-B5, a transformer-based segmentation architecture combining hierarchical encoder with a lightweight multilayer perception decoder, achieving efficient and flexible semantic segmentation without reliance on positional encodings [34].

3.3. Model Suitability Assessment for Study Area

To assess the suitability of semantic segmentation models for streetscape analysis in Phnom Penh, all five models were applied to an identical set of SVI images. This approach enabled a direct, image-by-image comparison of model outputs under the same visual conditions. The evaluation focused on two key dimensions relevant to capturing the characteristics of the study area.
First, we examined quantitative performance metrics, including mIoU and the number of semantic classes detected by each model. These metrics provided a baseline measure of the model’s ability to accurately segment visual features across the diverse streetscape. Second, we assessed whether the models could reliably detect informal elements, particularly street vendors, which represent small-scale, non-standard structures critical to understanding functional and social dynamics in Southeast Asian streetscapes. Capturing these elements is essential because they are underrepresented in global segmentation datasets and central to the study’s research objectives.
While both criteria were considered together, special emphasis was placed on the model’s ability to detect street vendors. Models that achieved high overall IoU or detected many classes but failed to capture these informal elements were deemed less suitable. Conversely, models that balanced quantitative performance with accurate identification of street vendors were prioritized. This combined approach of general segmentation quality and task-specific coverage guided the selection of the model most appropriate for large-scale streetscape analysis. Figure 2 illustrates the overall workflow of data preparation, model selection, evaluation, validation, and neighborhood-scale analysis adopted in this study.

4. Results

4.1. Model Selection for AI-Based Streetscape Segmentation

Quantitative evaluation of the five semantic segmentation models applied to the sample set of Mapillary SVI images is summarized in Table 2. OneFormer achieved an overall mIoU of 60.8% while detecting 20 semantic classes, indicating a balance between segmentation breadth and accuracy. Mask2Former achieved the next highest mIoU of 57.7%, although it detected fewer classes (15). By contrast, MaskFormer detected more classes (23) but achieved a lower mIoU of 55.6% and failed to consistently detect informal street vendors. BEiT-L and SegFormer-B5 had both lower mIoU and fewer detected classes, highlighting that a higher number of classes alone does not guarantee practical utility for capturing informal urban components.
Figure 3 illustrates sample segmentation outputs for each model. OneFormer successfully detected street vendors while also capturing the boundaries of individual objects relatively well. By contrast, MaskFormer failed to detect street vendors and often struggled to delineate object boundaries accurately. Mask2Former also showed limitations, failing to detect the diverse range of objects present in the streetscape.
These factors, including its high mIoU, broad class coverage, and detection of street vendors, highlight OneFormer’s suitability for this study. This suitability arises from its technical design. Unlike semantic-only models, OneFormer’s panoptic approach integrates both stuff (background) and thing (discrete objects) queries. This allows the model to handle informal, irregular objects such as street vendors more effectively, producing coherent segmentation masks that capture both larger background elements and smaller non-standard objects [31]. This capability is particularly important in the study area, where informal street vendors do not conform to regular shapes and are embedded within complex urban backgrounds.
Based on this combination of quantitative performance and visual inspection of segmentation outputs, and task-specific detection of informal street vendors, OneFormer was selected for full-scale segmentation of the Mapillary dataset. Its panoptic design ensures that both formal urban structures and non-standard, informal street features are captured accurately, providing a robust foundation for subsequent streetscape analysis.

4.2. Extraction of Streetscape Indicators and Performance Validation

Using the OneFormer model, semantic segmentation was applied to all 8912 Mapillary SVI images, enabling the extraction of streetscape indicators related to pedestrian infrastructure, urban greenery, vehicular presence, and informal economic activities. To ensure the reliability of this process, we first assessed the objectivity of the manual annotations used as a baseline. An inter-rater agreement analysis using a random sample of 210 images yielded an mIoU of 76.9%, indicating high consistency between annotators.
The performance of OneFormer was subsequently validated against this high-quality ground truth. The model demonstrated robust zero-shot performance with an overall mIoU of 65.7% (Table 3). Specifically, the model achieved a recall of 0.69 for base class (used here as a proxy for street vendors), confirming its effectiveness in capturing informal elements that are traditionally difficult to segment. The SVI images were processed at their native high resolution of 2048 × 1152 pixels (16:9 aspect ratio) to preserve visual details. The inference process was highly efficient, averaging 0.35 s per image using an NVIDIA RTX 4000 Ada Generation GPU (20 GB VRAM).
Figure 4 illustrates a representative street segment characterized by a high density of street vendors lining the sidewalks. Compared with the neighborhood averages (Table 4), this segment exhibits a relatively higher presence of street vendors, reflecting a streetscape strongly influenced by informal commercial activity. In this street section, road surfaces accounted for the largest proportion (29.26%), followed by sky (19.78%) and trees (18.85%). Notably, base (10.00%) constituted the fifth-largest class of the segmented area, making it one of the most prominent non-background classes and highlighting the visual significance of informal commercial activities at the street-level.
Specifically for the base class, OneFormer achieved an IoU of 46.5%, a precision of 0.59, and a recall of 0.69 (Table 3). These results indicate that a substantial proportion of regions identified as base by the model correspond to manually annotated street vendors, supporting the use of this class as a proxy for informal commercial activity. While precision suggests some degree of over-identification, the relatively high recall confirms the model’s ability to capture the majority of street vendor instances present in the images.
We hypothesize that this pattern arises from several factors related to the visual characteristics and spatial context of street vendors. Visually, street vendors typically appear as small-scale structures connected to the ground, which may resemble patterns learned by the pre-trained model for the base class. In terms of scale and context, these vendors are relatively small and tend to be distinguished from their surroundings along the ground plane from roads and sidewalks, and vertically from walls, fences, or buildings. Consequently, the model may associate low-lying, consistently patterned structures with the base category. This interpretation is consistent with both visual inspection of the segmentation outputs and the quantitative validation metrics.

4.3. Neighborhood-Level Differences of Streetscape Components

Building on the validated performance of OneFormer across the full Mapillary dataset, this section examines neighborhood-level differences in key streetscape components across the three study areas: Boeng Keng Kang 1 (BKK1), Tonle Bassac (TB), and Koh Pich (KP). Using semantic segmentation outputs, distributions of streetscape components were compared to identify systematic spatial differences in both structural elements and street-level activity.
The analysis focuses on eight streetscape components that were consistently detected and directly relevant to pedestrian environments and informal streets activity: sky, green, building, road, sidewalk, vehicle, person, and base (used as a proxy for street vendor presence). To improve interpretability, functionally similar semantic classes were aggregated using predefined rules. Specifically, the green category combines tree, palm, plant, and grass classes to represent visible urban greenery, while the vehicle category aggregates car, truck, van, bus, minibike, and bicycle classes to capture overall vehicular presence at the pedestrian scale. Other detected classes (e.g., rocks, awning, signboard, etc.) that were not directly consequential to walkability or informal street activity were excluded from the comparative analysis. Although the person class reflects momentary visual presence due to the snapshot nature of SVI, it is retained and interpreted as a supplementary indicator of transient street-level activity or crowding rather than as a direct measure of pedestrian demand.
One-way ANOVA results in Table 4 indicate that all eight streetscape components exhibit statistically significant differences across neighborhoods (F-values ranging from 61.67 to 1287.61, all p < 0.001), confirming that streetscape composition varies systematically across the study areas rather than being driven by isolated street-level observations. Effect size estimates (η2) suggest that structural components—particularly building (0.224) and sky (0.146)—account for the largest share of between-neighborhood variation, highlighting built density and visual openness as the primary dimensions distinguishing these urban environments.
Post hoc comparisons further clarify which neighborhood differences are statistically significant for each streetscape component. For sky, the ordering (KP > TB > BKK1) indicates progressively greater visual openness, with Koh Pich exhibiting the most unobstructed sky views. In contrast, building coverage follows the reverse pattern (BKK1 > KP > TB), reflecting the dense built form of Boeng Keng Kang 1. The green component shows a distinct grouping (TB > BKK1 > KP), indicating that Tonle Bassac has significantly higher visible greenery, while no statistically meaningful difference is observed between Boeng Keng Kang 1 and Koh Pich. For the person class, post hoc results (BKK1, TB > KP) indicate similarly higher levels of momentary pedestrian presence in Boeng Keng Kang 1 and Tonle Bassac compared to Koh Pich.
Notably, the base class—used as a physical proxy for the visual presence of street vendors—exhibits a clear and statistically significant ordering across all three neighborhoods (BKK1 > TB > KP). Although the effect size for base (η2 = 0.011) is relatively small, this reflects the limited proportion of image area occupied by street vendors rather than a lack of meaningful spatial differentiation. The consistent post hoc ordering suggests that informal street vending contributes to neighborhood-level distinctions in streetscape composition, even when its overall visual footprint is modest.
At the neighborhood level, Boeng Keng Kang 1 exhibited the highest proportions of buildings (27.75%), sidewalks (3.01%), base (1.00%), and person (0.65%), reflecting a compact streetscape characterized by dense built form and a fine-grained, grid-like street network. Tonle Bassac displays the highest proportions of greenery (16.20%) and vehicles (18.13%), alongside intermediate levels of base (0.75%) and person (0.60%), indicating a heterogeneous streetscape combining visible landscaping, vehicular presence, and localized street activity. In contrast, Koh Pich shows the highest proportions of sky (27.99%) and roads (20.35%), coupled with the lowest levels of base (0.53%) and persons (0.16%), reflecting a more open and low-density streetscape with comparatively limited momentary street activity.
The distributional patterns visualized in Figure 5 further support these findings. For key streetscape components such as building (F(2, 8910) = 1287.61, p < 0.001), sky (F(2, 8910) = 760.86, p < 0.001), and road (F(2, 8910) = 247.37, p < 0.001), the non-overlapping distributional spreads—characterized by distinct separations in both medians (indicated by the orange lines) and interquartile ranges—underscore the robustness of the observed neighborhood differences. The close alignment between neighborhood-level means (represented by red dots) and their respective medians confirms that these spatial patterns are pervasive across each study area rather than being driven by isolated street segments. Notably, the base class (F(2, 2449) = 13.48, p < 0.001) also follows this consistent distributional trend, reinforcing its validity as a physical proxy for street vendors.

5. Discussion

This study examined the applicability of five state-of-the-art semantic segmentation models for extracting streetscape indicators from crowdsourced SVI in a developing Southeast Asian city characterized by high levels of urban informality. The comparative analysis revealed notable differences in model capability, particularly in detecting non-standard and informal street features such as street vendors. Among the evaluated models, OneFormer produced the most contextually reliable segmentation outputs, effectively balancing overall accuracy with sensitivity to informal streetscape elements.
Importantly, the results demonstrate that model transferability in data-constrained and informality-rich contexts cannot be assessed solely through aggregate accuracy metrics such as mIoU. While several models achieved comparable overall performance, they differed markedly in their ability to detect small-scale, irregular, and visually ambiguous elements. The successful detection of street vendors—operationalized through the base class—demonstrates that sensitivity to informal features is a more critical determinant of model utility in the Global South streetscapes than overall performance alone.
The observed technical advantage of OneFormer warrants further consideration. Its strong performance can be plausibly attributed to its panoptic learning formulation, which jointly models stuff (background) and things (discrete objects). Informal street vendors often occupy an ambiguous visual position between object and background, making them difficult to isolate using purely semantic segmentation approaches. In contrast, OneFormer’s panoptic representation allows such elements to be treated as distinct objects while preserving their contextual relationship with the surrounding street environment. In addition, OneFormer’s multi-task conditioning may contribute to more stable boundary delineation across objects of different sizes. This property is particularly relevant for detecting small, irregular structures that do not conform to standard object shapes, such as street vendors. Although the possibility of a favorable alignment with the Phnom Penh visual domain cannot be ruled out, these architectural characteristics provide a technically grounded explanation for OneFormer’s performance and suggest promising directions for future domain adaptation of fine-tuning.
However, it is important to note that most semantic models are primarily trained on datasets from regulated Global North cities, which feature discrete standardized street layouts and clearly defined object boundaries [35]. In contrast, the Global South’s streetscape is characterized by a high degree of informality, where street vendors and their associated equipment often create visually entangled environments that defy clear object separation [36]. Our findings suggest that the observed performance gaps are not inherent model failures, but rather a reflection of the morphological gap between the urban environment represented in training datasets and the complex, informal realities of Global South cities.
Beyond the specific context of Phnom Penh, this study offers broader methodological insights for informality-aware streetscape measurement. The findings indicate that urban analytics in the Global South benefit from model architectures capable of simultaneously representing structured infrastructure and irregular street-level activity. Moreover, the proposed workflow—combining crowdsourced SVI with post-segmentation filtering—demonstrates a scalable and cost-effective alternative for environments where official GIS datasets are missing or outdated. This approach provides a practical blueprint for consistent street-level benchmarking across cities with similar informal street economies.
At the neighborhood scale, segmentation-derived indicators revealed statistically robust spatial variation in streetscape composition. Differences in built density, openness, and informal activity were statistically captured, underscoring the capacity of SVI-based indicators to reflect meaningful contrasts in urban form. For example, Boeng Keng Kang 1 exhibits a compact street network with small block sizes and continuous mixed-use frontages, corresponding to higher levels of pedestrian presence and informal street activity. Tonle Bassac presents a more heterogeneous urban fabric in which large automobile-oriented arterials coexist with older residential blocks, leading to uneven pedestrian conditions across short distances. Koh Pich, characterized by wide and planned street corridors, shows comparatively limited everyday street activity, reflecting its transitional stage of development within the city.
Importantly, informal street activity should not be interpreted through a simplistic normative lens. While street vendors may contribute to urban vitality and social interaction, they can also constrain pedestrian movement when they encroach upon sidewalks or reduce effective walking space. In this study, pixel-based indicators are interpreted as environmental proxies rather than direct behavioral measures; a higher visual presence of vendors suggests potential spatial occupation but does not, by itself, establish accessibility outcomes. The explicit identification of informal elements through semantic segmentation enables these nuanced and context-dependent relationships to be examined.

6. Conclusions

This study demonstrates the feasibility and value of applying current semantic segmentation models to crowdsourced SVI in a developing Southeast Asian context. While the models were not originally designed or optimized for Phnom Penh’s urban characteristics, our evaluation shows that existing pretrained architectures, when carefully validated, can yield reliable representations of street-level physical conditions. By deploying the most suitable model, this research produced spatially detailed indicators that offer a practical, scalable alternative to conventional datasets which are often outdated or unavailable in rapidly developing cities.
This baseline evaluation provides a necessary starting point for applying computer vision to informal urban environments. By identifying how existing models respond to non-standard street elements, this work establishes an empirical basis for future efforts to refine these tools for the Global South. The ability to quantify features such as informal street activities despite the inherent visual complexity suggests a scalable methodological pathway for providing objective environmental proxies in data-limited environments.
Despite these contributions, several limitations should be acknowledged. The availability and distribution of crowdsourced SVI were uneven across neighborhoods, potentially affecting representativeness. The segmentation models relied on training datasets dominated by Global North imagery, limiting their sensitivity to region-specific streetscapes. Manual filtering of imagery, while manageable for the current dataset, limits scalability for citywide applications. Additionally, the analysis was restricted to static visual attributes and did not capture temporal variations or behavioral dimensions related to street use.
Future research should address these gaps to enhance the robustness and policy relevance of SVI-based assessments. A key priority is the development of region-specific segmentation datasets that include the informal urban features identified in this study. This would enable domain-specific fine-tuning of model architectures to better capture the unique characteristics of Southeast Asian cities. Additionally, automating image pre-screening pipelines and integrating visually derived indicators with mobility, environmental, or behavioral data would provide a more comprehensive understanding of urban environments. Such advancements will support more evidence-informed planning for walkable and sustainable cities in data-constrained contexts.

Author Contributions

Conceptualization, Y.C. and S.C.; methodology, Y.C., D.H.D.X. and S.C.; validation, Y.C. and D.H.D.X.; formal analysis, D.H.D.X.; investigation, Y.C. and D.H.D.X.; data curation, D.H.D.X.; writing—original draft preparation, Y.C., D.H.D.X. and S.C.; writing—review and editing, Y.C.; visualization, Y.C. and D.H.D.X.; supervision, Y.C. and S.C.; funding acquisition, Y.C. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (No. RS-2025-00559786).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available on request from the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
GISGeographic Information System
SVIStreet View Imagery
APIApplication Programming Interface
mIoUMean Intersection Over Union
NLPNatural Language Processing

References

  1. Li, X.; Ratti, C.; Seiferling, I. Mapping urban landscapes along streets using Google Street View. In Proceedings of the International Cartographic Conference, Washington, DC, USA, 2–7 July 2017; Springer International Publishing: Cham, Switzerland, 2017; pp. 341–356. [Google Scholar]
  2. Southworth, M. Designing the walkable city. J. Urban Plan. Dev. 2005, 131, 246–257. [Google Scholar] [CrossRef] [Scilit]
  3. Saelens, B.E.; Handy, S.L. Built environment correlates of walking: A review. Med. Sci. Sports Exerc. 2008, 40, S550–S566. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Sallis, J.F.; Cain, K.L.; Conway, T.L.; Gavand, K.A.; Millstein, R.A.; Geremia, C.M.; King, A.C. Is your neighborhood designed to support physical activity? A brief streetscape audit tool. Prev. Chronic Dis. 2015, 12, E141. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Koo, B.W.; Guhathakurta, S.; Botchwey, N. How are neighborhood and street-level walkability factors associated with walking behaviors? A big data approach using street view images. Environ. Behav. 2022, 54, 211–241. [Google Scholar] [CrossRef] [Scilit]
  6. Huang, X.; Zeng, L.; Liang, H.; Li, D.; Yang, X.; Zhang, B. Comprehensive walkability assessment of urban pedestrian environments using big data and deep learning techniques. Sci. Rep. 2024, 14, 26993. [Google Scholar] [CrossRef] [Scilit]
  7. Ewing, R.; Hajrasouliha, A.; Neckerman, K.M.; Purciel-Hill, M.; Greene, W. Streetscape features related to pedestrian activity. J. Plan. Educ. Res. 2016, 36, 5–15. [Google Scholar] [CrossRef] [Scilit]
  8. Li, X.; Cai, B.; Ratti, C. Using street-level images and deep learning for urban landscape studies. Landsc. Archit. Front. 2018, 6, 20–31. [Google Scholar] [CrossRef] [Scilit]
  9. Westenhoefer, J.; Nouri, E.; Reschke, M.L.; Seebach, F.; Buchcik, J. Walkability and urban built environments—A systematic review of health impact assessments. BMC Public Health 2023, 23, 518. [Google Scholar] [CrossRef] [Scilit]
  10. Ewing, R.; Cervero, R. Travel and the built environment: A meta-analysis. J. Am. Plan. Assoc. 2010, 76, 265–294. [Google Scholar] [CrossRef] [Scilit]
  11. Yahia, M.W.; Abdalla, S.B.; Sukkar, A.; Saleem, A.A.; Maksoud, A.M. Towards better site analysis in architectural and urban design: Adapting experiential learning theory in post-COVID architectural teaching methods. Arch. Des. Res. 2023, 36, 51–65. [Google Scholar] [CrossRef] [Scilit]
  12. Yin, L. Street-level urban design qualities for walkability: Combining 2D and 3D GIS measures. Comput. Environ. Urban Syst. 2017, 64, 288–296. [Google Scholar] [CrossRef] [Scilit]
  13. Angel, A.; Cohen, A.; Nelson, T.; Plaut, P. Evaluating the relationship between walking and street characteristics based on big data and machine learning analysis. Cities 2024, 151, 105111. [Google Scholar] [CrossRef] [Scilit]
  14. Ki, D.; Chen, Z.; Lee, S.; Lieu, S. A novel walkability index using Google Street View and deep learning. Sustain. Cities Soc. 2023, 99, 104896. [Google Scholar] [CrossRef] [Scilit]
  15. Crooks, A.; See, L. Leveraging street-level imagery for urban planning. Environ. Plan. B Urban Anal. City Sci. 2022, 49, 773–776. [Google Scholar] [CrossRef] [Scilit]
  16. Biljecki, F.; Ito, K. Street view imagery in urban analytics and GIS: A review. Landsc. Urban Plan. 2021, 215, 104217. [Google Scholar] [CrossRef] [Scilit]
  17. Li, X.; Zhang, C.; Li, W.; Ricard, R.; Meng, Q.; Zhang, W. Assessing street-level urban greenery using Google Street View and a modified green view index. Urban For. Urban Green. 2015, 14, 675–685. [Google Scholar] [CrossRef] [Scilit]
  18. Lu, Y. Using Google Street View to investigate the association between street greenery and physical activity. Landsc. Urban Plan. 2019, 191, 103435. [Google Scholar] [CrossRef] [Scilit]
  19. Yin, L.; Wang, Z. Measuring visual enclosure for street walkability using machine learning and Google Street View imagery. Appl. Geogr. 2016, 76, 147–153. [Google Scholar] [CrossRef] [Scilit]
  20. Ye, Y.; Zeng, W.; Shen, Q.; Zhang, X.; Lu, Y. The visual quality of streets: A human-centred continuous measurement based on machine learning and street view images. Environ. Plan. B Urban Anal. City Sci. 2019, 46, 1439–1457. [Google Scholar] [CrossRef] [Scilit]
  21. Chen, S.; Biljecki, F. Automatic assessment of public open spaces using street view imagery. Cities 2023, 137, 104329. [Google Scholar] [CrossRef] [Scilit]
  22. Ito, K.; Biljecki, F. Assessing bikeability with street view imagery and computer vision. Transp. Res. Part C Emerg. Technol. 2021, 132, 103371. [Google Scholar] [CrossRef] [Scilit]
  23. Dai, S.; Zhao, W.; Wang, Y.; Huang, X.; Chen, Z.; Lei, J.; Jia, P. Assessing spatiotemporal bikeability using multisource geospatial big data: A case study of Xiamen, China. Int. J. Appl. Earth Obs. Geoinf. 2023, 125, 103539. [Google Scholar]
  24. Naik, N.; Kominers, S.D.; Raskar, R.; Glaeser, E.L.; Hidalgo, C.A. Computer vision uncovers predictors of physical urban change. Proc. Natl. Acad. Sci. USA 2017, 114, 7571–7576. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Li, S.; Ma, S.; Tong, D.; Jia, Z.; Li, P.; Long, Y. Associations between the quality of street space and attributes of the built environment using large volumes of street view images. Environ. Plan. B Urban Anal. City Sci. 2021, 49, 1197–1211. [Google Scholar] [CrossRef] [Scilit]
  26. Nathvani, R.; Cavanaugh, A.; Suel, E.; Bixby, H.; Clark, S.N.; Metzler, A.B.; Nimo, J.; Moses, J.B.; Baah, S.; Arku, R.E.; et al. Measurement of urban vitality with time-lapsed street-view images and object detection. Int. Soc. Photogramm. Remorte Sens. 2025, 221, 251–264. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. National Institute of Statistics. Statistical Yearbook of Cambodia 2021; Ministry of Planning: Phnom Penh, Cambodia, 2021. Available online: https://www.nis.gov.kh/nis/yearbooks/StatisticalYearbookofCambodia2021.pdf (accessed on 1 December 2025).
  28. Eidse, N.; Turner, S.; Oswin, N. Contesting street spaces in a socialist city: Itinerant vending-scapes and the everyday politics of mobility in Hanoi, Vietnam. Ann. Am. Assoc. Geogr. 2016, 106, 340–349. [Google Scholar] [CrossRef] [Scilit]
  29. Zhou, B.; Zhao, H.; Puig, X.; Fidler, S.; Barriuso, A.; Torralba, A. Scene parsing through the ADE20K dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017. [Google Scholar]
  30. Zhou, B.; Zhao, H.; Puig, X.; Xiao, T.; Fidler, S.; Barriuso, A.; Torralba, A. Semantic understanding of scenes through the ADE20K dataset. Int. J. Comput. Vis. 2019, 127, 302–321. [Google Scholar] [CrossRef] [Scilit]
  31. Jain, J.; Li, J.; Chiu, M.T.; Hassani, A.; Orlov, N.; Shi, H. OneFormer: One transformer to rule universal image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 2989–2998. [Google Scholar]
  32. Bao, H.; Dong, L.; Wei, F. BEiT: BERT pre-training of image transformers. arXiv 2021, arXiv:2106.08254. [Google Scholar]
  33. Cheng, B.; Schwing, A.; Kirillov, A. Per-pixel classification is not all you need for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 19–25 June 2021. [Google Scholar]
  34. Xie, E.; Wang, W.; Yu, Z.; Anandkumar, A.; Alvarez, J.M.; Luo, P. SegFormer: Simple and efficient design for semantic segmentation with transformers. Adv. Neural Inf. Process. Syst. 2021, 34, 12077–12090. [Google Scholar]
  35. Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 3213–3223. [Google Scholar]
  36. Ibrahim, M.R.; Haworth, J.; Cheng, T. URBAN-i: From urban scenes to mapping slums, transport modes, and pedestrians in cities using deep learning and computer vision. Environ. Plan. B Urban Anal. City Sci. 2021, 48, 76–93. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Location of the study area within central Phnom Penh, Cambodia.
Figure 1. Location of the study area within central Phnom Penh, Cambodia.
Sustainability 18 01355 g001
Figure 2. Overview of the proposed methodological workflow.
Figure 2. Overview of the proposed methodological workflow.
Sustainability 18 01355 g002
Figure 3. Segmentation Results of Semantic Segmentation Models.
Figure 3. Segmentation Results of Semantic Segmentation Models.
Sustainability 18 01355 g003
Figure 4. Segmentation results of OneFormer with class percentages.
Figure 4. Segmentation results of OneFormer with class percentages.
Sustainability 18 01355 g004
Figure 5. Distribution of key streetscape components across neighborhoods shown as box plots with 95% confidence intervals.
Figure 5. Distribution of key streetscape components across neighborhoods shown as box plots with 95% confidence intervals.
Sustainability 18 01355 g005
Table 1. SVI dataset and sampling density by neighborhood.
Table 1. SVI dataset and sampling density by neighborhood.
NeighborhoodNumber of ImagesArea (km2)Density (Images/km2)
Boeng Keng Kang 146001.054381.0
Tonle Bassac30021.621853.1
Koh Pich13101.231065.0
Total89123.902285.1
Table 2. Performance and informality detection of semantic segmentation models.
Table 2. Performance and informality detection of semantic segmentation models.
ModelmIoU (%)No. of ClassesStreet VendorStreet Furniture
OneFormer60.820YesYes
Mask2Former57.715YesNo
BEiT-L57.06NoNo
MaskFormer55.623NoNo
SegFormer-B551.89NoYes
Table 3. OneFormer performance evaluation based on manual annotations by streetscape classes (N = 210).
Table 3. OneFormer performance evaluation based on manual annotations by streetscape classes (N = 210).
ClassIoU (%)PrecisionRecallF1-Score
Sky92.40.930.990.96
Green85.40.930.920.92
Building77.10.790.970.87
Road74.80.760.980.86
Sidewalk57.60.720.740.73
Vehicle57.10.680.780.73
Person34.70.370.850.52
Base (Street vendor proxy)46.50.590.690.63
Overall (mIoU)65.7--0.78
Table 4. Differences in key streetscape components across neighborhoods (N = 8912).
Table 4. Differences in key streetscape components across neighborhoods (N = 8912).
ClassBoeng Keng Kang 1Tonle BassacKoh PichANOVA (p)η2Post Hoc
MeanSDMeanSDMeanSD
Sky18.92 7.5325.1410.4427.999.12<0.0010.146KP > TB > BKK1
Green13.819.7716.2010.6513.528.26<0.0010.014TB > BKK1, KP
Building27.759.9116.7110.3018.158.89<0.0010.224BKK1 > KP > TB
Road15.437.8614.937.5220.357.95<0.0010.053KP > BKK1 > TB
Sidewalk3.012.052.271.902.501.29<0.0010.031BKK1 > KP > TB
Vehicle15.458.0218.139.3112.546.00<0.0010.048TB > BKK1 > KP
Person0.650.980.600.750.160.22<0.0010.037BKK1, TB > KP
Base1.001.650.751.490.530.71<0.0010.011BKK1 > TB > KP
Note: Post hoc results are based on independent t-tests with Bonferroni correction (p < 0.05).
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Choi, Y.; Xiang, D.H.D.; Chng, S. Semantic Segmentation for Walkability Assessment in Southeast Asian Streetscapes. Sustainability 2026, 18, 1355. https://doi.org/10.3390/su18031355

AMA Style

Choi Y, Xiang DHD, Chng S. Semantic Segmentation for Walkability Assessment in Southeast Asian Streetscapes. Sustainability. 2026; 18(3):1355. https://doi.org/10.3390/su18031355

Chicago/Turabian Style

Choi, Yunkyung, Darren Ho Di Xiang, and Samuel Chng. 2026. "Semantic Segmentation for Walkability Assessment in Southeast Asian Streetscapes" Sustainability 18, no. 3: 1355. https://doi.org/10.3390/su18031355

APA Style

Choi, Y., Xiang, D. H. D., & Chng, S. (2026). Semantic Segmentation for Walkability Assessment in Southeast Asian Streetscapes. Sustainability, 18(3), 1355. https://doi.org/10.3390/su18031355

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop