Urban Morphology-Oriented Streetscape Segmentation via Hierarchical Transformer and Frequency-Aware Feature Learning
Abstract
1. Introduction
1.1. Urban Morphology Background and Research Significance
1.2. Research Gap and Limitations of Existing Methods
1.3. Scope, Motivation, and Research Objective
1.4. Overview and Contributions
1.4.1. Method Overview
1.4.2. Main Contributions
2. Related Work
2.1. Backbone Networks for Semantic Segmentation
2.2. Multi-Scale Feature Fusion and Frequency-Aware Representation
2.3. Urban Street-View Datasets
2.4. Practical Relevance and Target Users
3. Methods
3.1. Experimental Data
3.1.1. Dataset Selection and Characteristics
3.1.2. Urban Morphology-Based Semantic Reconfiguration
3.1.3. Dataset Limitations and Preprocessing Considerations
3.2. Model Architecture
3.3. Fusion Module Architecture
3.4. Wave Fusion Module
3.5. Loss Function
3.6. Evaluation Metrics
3.7. Methodological Novelty Clarification
4. Results
4.1. Experimental Setup
4.2. Comparative Results
4.3. Comparative Experimental Results and Analysis
4.3.1. Backbone Network Comparison Experiments
4.3.2. Comparative Experiments on Feature Fusion Networks
4.3.3. Comparative Experiments on Decoders
4.4. Ablation Studies and Analysis
4.4.1. Ablation Study and Analysis of Loss Functions
4.4.2. Ablation Study and Analysis of Key Components
4.4.3. AAblation Study and Analysis of WaveFusion
4.5. Morphological Interpretability Analysis
5. Discussion
5.1. Overall Morphological Interpretation
5.2. Hierarchical Representation and Urban Spatial Structure
5.3. Multi-Scale Fusion and Spatial Continuity
5.4. Frequency-Aware Refinement and Boundary Morphology
5.5. Comparison with Existing Methods in Urban Context
5.6. Morphological Interpretability and Architectural Relevance
5.6.1. Case-Driven Urban Morphology Analysis Based on Feature Activation
5.6.2. Architectural and Urban Design Usability Analysis
5.7. Limitations and Future Work
5.8. Urban Morphological Implications of the Six-Class Reconfiguration
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| HieraWaveSeg | Hierarchical Wavelet-based Segmentation Network |
| DWT | Discrete Wavelet Transform |
| PAN | Path Aggregation Network |
| FPN | Feature Pyramid Network |
| CSPNet | Cross Stage Partial Network |
| LL | Low–Low (Approximation subband) |
| LH | Low–High (Horizontal detail subband) |
| HL | High–Low (Vertical detail subband) |
| HH | High–High (Diagonal detail subband) |
| BCE | Binary Cross-Entropy |
| LCE | Weighted Cross-Entropy Loss |
| LDice | Weighted Dice Loss |
| OA | Overall Accuracy |
| IoU | Intersection over Union |
| mIoU | mean Intersection over Union |
| TP | True Positive |
| FP | False Positive |
| TN | True Negative |
| FN | False Negative |
| PIDNet | Proportional–Integral–Derivative Network |
| EWMA | Exponentially Weighted Moving Average |
| SAM2 | Segment Anything Model 2 |
References
- Kropf, K. The Handbook of Urban Morphology; John Wiley & Sons: Hoboken, NJ, USA, 2018. [Google Scholar]
- Lynch, K. The Image of the City; MIT Press: Cambridge, MA, USA, 1964. [Google Scholar]
- Huang, J.; Lu, X.X.; Sellers, J.M. A global comparative analysis of urban form: Applying spatial metrics and remote sensing. Landsc. Urban Plan. 2007, 82, 184–197. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.; Du, Y.; Ma, Y.; Liu, D.; Chen, L. Investigating Spatial Variation Characteristics and Influencing Factors of Urban Green View Index Based on Street View Imagery—A Case Study of Luoyang, China. Sustainability 2025, 17, 10208. [Google Scholar] [CrossRef] [Scilit]
- He, N.; Li, G. Urban neighbourhood environment assessment based on street view image processing: A review of research trends. Environ. Chall. 2021, 4, 100090. [Google Scholar] [CrossRef] [Scilit]
- Lu, X.; Li, Q.; Ji, X.; Sun, D.; Meng, Y.; Yu, Y.; Lyu, M. Impact of streetscape built environment characteristics on human perceptions using street view imagery and deep learning: A case study of Changbai Island, Shenyang. Buildings 2025, 15, 1524. [Google Scholar] [CrossRef] [Scilit]
- Dabove, P.; Daud, M.; Olivotto, L. Revolutionizing urban mapping: Deep learning and data fusion strategies for accurate building footprint segmentation. Sci. Rep. 2024, 14, 13510. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Oliveira, V. Urban form and the socioeconomic and environmental dimensions of cities. J. Urban. Int. Res. Placemaking Urban Sustain. 2024, 17, 1–23. [Google Scholar] [CrossRef] [Scilit]
- Ulku, I.; Akagündüz, E. A survey on deep learning-based architectures for semantic segmentation on 2d images. Appl. Artif. Intell. 2022, 36, 2032924. [Google Scholar] [CrossRef] [Scilit]
- Zuo, Z.; Shuai, B.; Wang, G.; Liu, X.; Wang, X.; Wang, B.; Chen, Y. Convolutional recurrent neural networks: Learning spatial dependencies for image representation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Boston, MA, USA, 7–12 June 2015; pp. 18–26. [Google Scholar]
- Pereira, G.A.; Hussain, M. A review of transformer-based models for computer vision tasks: Capturing global context and spatial relationships. arXiv 2024, arXiv:2408.15178. [Google Scholar] [CrossRef] [Scilit]
- Pohlen, T.; Hermans, A.; Mathias, M.; Leibe, B. Full-resolution residual networks for semantic segmentation in street scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 4151–4160. [Google Scholar]
- Chen, L.; Fu, Y.; Gu, L.; Zheng, D.; Dai, J. Spatial frequency modulation for semantic segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 9767–9784. [Google Scholar] [CrossRef] [Scilit]
- Clifton, K.; Ewing, R.; Knaap, G.J.; Song, Y. Quantitative analysis of urban form: A multidisciplinary review. J. Urban. 2008, 1, 17–45. [Google Scholar] [CrossRef] [Scilit]
- Mishra, B.; Dahal, A.; Luintel, N.; Shahi, T.B.; Panthi, S.; Pariyar, S.; Ghimire, B.R. Methods in the spatial deep learning: Current status and future direction. Spat. Inf. Res. 2022, 30, 215–232. [Google Scholar] [CrossRef] [Scilit]
- Karndacharuk, A.; Wilson, D.J.; Dunn, R. A review of the evolution of shared (street) space concepts in urban environments. Transp. Rev. 2014, 34, 190–220. [Google Scholar] [CrossRef] [Scilit]
- Ankareddy, R.; Delhibabu, R. Dense segmentation techniques using deep learning for urban scene parsing: A review. IEEE Access 2025, 13, 34496–34517. [Google Scholar] [CrossRef] [Scilit]
- Yu, B.; Yang, A.; Chen, F.; Wang, N.; Wang, L. SNNFD, spiking neural segmentation network in frequency domain using high spatial resolution images for building extraction. Int. J. Appl. Earth Obs. Geoinf. 2022, 112, 102930. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Bai, X.; Wang, T. Boundary finding based multi-focus image fusion through multi-scale morphological focus-measure. Inf. Fusion 2017, 35, 81–101. [Google Scholar] [CrossRef] [Scilit]
- Karimi, K. The configurational structures of social spaces: Space syntax and urban morphology in the context of analytical, evidence-based design. Land 2023, 12, 2084. [Google Scholar] [CrossRef] [Scilit]
- Tarkhan, N.; Szcześniak, J.T.; Reinhart, C. Façade feature extraction for urban performance assessments: Evaluating algorithm applicability across diverse building morphologies. Sustain. Cities Soc. 2024, 105, 105280. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Zhang, Y.; Zeng, A.; Pan, D.; Zhang, X. Learning hierarchical representations in temporal and frequency domains for time series forecasting. In Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision (PRCV); Springer: Berlin/Heidelberg, Germany, 2023; pp. 91–103. [Google Scholar]
- Chen, J.; Zhao, X.; Wang, H.; Yan, J.; Yang, D.; Xie, K. Portraying heritage corridor dynamics and cultivating conservation strategies based on environment spatial model: An integration of multi-source data and image semantic segmentation. Herit. Sci. 2024, 12, 419. [Google Scholar] [CrossRef] [Scilit]
- Wu, C.; Wang, J.; Wang, M.; Kraak, M.J. Machine learning-based characterisation of urban morphology with the street pattern. Comput. Environ. Urban Syst. 2024, 109, 102078. [Google Scholar] [CrossRef] [Scilit]
- Serra, M.L.L.A. Anatomy of an Emerging Metropolitan Territory-Towards an Integrated Analytical Framework for Metropolitan Morphology. Ph.D. Thesis, Universidade do Porto (Portugal), Porto, Portugal, 2014. [Google Scholar]
- He, H.; Xiong, W.; Zhou, F.; He, Z.; Zhang, T.; Sheng, Z. Topology-Aware Multi-View Street Scene Image Matching for Cross-Daylight Conditions Integrating Geometric Constraints and Semantic Consistency. ISPRS Int. J. Geo-Inf. 2025, 14, 212. [Google Scholar] [CrossRef] [Scilit]
- Zhong, L.; Guo, W.; Zheng, J.; Yan, L.; Xia, J.; Zhang, D.; Li, Q. HPAN: Hierarchical Part-Aware Network for Fine-Grained Segmentation of Street View Imagery. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 7794–7810. [Google Scholar] [CrossRef] [Scilit]
- Fernando, K.R.M.; Tsokos, C.P. Dynamically weighted balanced loss: Class imbalanced learning and confidence calibration of deep neural networks. IEEE Trans. Neural Netw. Learn. Syst. 2021, 33, 2940–2951. [Google Scholar] [CrossRef] [Scilit]
- Dong, G.; Yan, Y.; Shen, C.; Wang, H. Real-time high-performance semantic image segmentation of urban street scenes. IEEE Trans. Intell. Transp. Syst. 2020, 22, 3258–3274. [Google Scholar] [CrossRef] [Scilit]
- Khan, A.; Sohail, A.; Zahoora, U.; Qureshi, A.S. A survey of the recent architectures of deep convolutional neural networks. Artif. Intell. Rev. 2020, 53, 5455–5516. [Google Scholar] [CrossRef] [Scilit]
- Al Mushayt, N.S.; Dal Cin, F.; Barreiros Proença, S. New lens to reveal the street interface. A morphological-visual perception methodological contribution for decoding the public/private edge of arterial streets. Sustainability 2021, 13, 11442. [Google Scholar] [CrossRef] [Scilit]
- Hosseinzadeh, R.; Sadeghzadeh, M. Attention mechanisms in transformers: A general survey. J. AI Data Min. 2025, 13, 359–368. [Google Scholar]
- Ryali, C.; Hu, Y.T.; Bolya, D.; Wei, C.; Fan, H.; Huang, P.Y.; Aggarwal, V.; Chowdhury, A.; Poursaeed, O.; Hoffman, J.; et al. Hiera: A hierarchical vision transformer without the bells-and-whistles. In Proceedings of the International Conference on Machine Learning. PMLR, Honolulu, HI, USA, 23–29 July 2023; pp. 29441–29454. [Google Scholar]
- Zhang, Z.; Yuan, D.; Zhou, Y.; Yang, R. No Trade-Offs: Unified Global, Local, and Multi-Scale Context Modeling for Building Pixel-Wise Segmentation. Remote Sens. 2026, 18, 472. [Google Scholar] [CrossRef] [Scilit]
- Czerkauer-Yamu, C. Strategic Planning for the Development of Sustainable Metropolitan Areas Using a Multi-Scale Decision Support System–The Vienna Case. Ph.D. Thesis, Université de Franche-Comté, Besançon, France, 2012. [Google Scholar]
- Wu, J.; Li, J.; Yang, J.; Mei, S. Wavelet-integrated deep neural networks: A systematic review of applications and synergistic architectures. Neurocomputing 2025, 657, 131648. [Google Scholar] [CrossRef] [Scilit]
- Cinnamon, J.; Jahiu, L. Panoramic street-level imagery in data-driven urban research: A comprehensive global review of applications, techniques, and practical considerations. ISPRS Int. J. Geo-Inf. 2021, 10, 471. [Google Scholar] [CrossRef] [Scilit]
- Varghese, S.; Bayzidi, Y.; Bar, A.; Kapoor, N.; Lahiri, S.; Schneider, J.D.; Schmidt, N.M.; Schlicht, P.; Huger, F.; Fingscheidt, T. Unsupervised temporal consistency metric for video segmentation in highly-automated driving. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA, 14–19 June 2020; pp. 336–337. [Google Scholar]
- Zhong, T.; Ye, C.; Wang, Z.; Tang, G.; Zhang, W.; Ye, Y. City-scale mapping of urban façade color using street-view imagery. Remote Sens. 2021, 13, 1591. [Google Scholar] [CrossRef] [Scilit]
- Xia, Y.; Yabuki, N.; Fukuda, T. Development of a system for assessing the quality of urban street-level greenery using street view images and deep learning. Urban For. Urban Green. 2021, 59, 126995. [Google Scholar] [CrossRef] [Scilit]
- Carneiro, C.; Morello, E.; Voegtle, T.; Golay, F. Digital urban morphometrics: Automatic extraction and assessment of morphological properties of buildings. Trans. GIS 2010, 14, 497–531. [Google Scholar] [CrossRef] [Scilit]
- Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The cityscapes dataset for semantic urban scene understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 3213–3223. [Google Scholar]
- Cordts, M.; Omran, M.; Ramos, S.; Scharwächter, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; Schiele, B. The cityscapes dataset. In Proceedings of the CVPR Workshop on the Future of Datasets in Vision; IEEE: New York, NY, USA, 2015; Volume 2, pp. 1–8. [Google Scholar]
- Shijie, X.; Dong, Z.; Dan, T. Multi-scale feature fusion network for real-time semantic segmentation of urban street scenes: Enhancing detail retention and accuracy: X. Shijie et al. Vis. Comput. 2025, 41, 7799–7815. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Peng, L.; Wu, C.; Zhang, J. Street View Imagery (SVI) in the built environment: A theoretical and systematic review. Buildings 2022, 12, 1167. [Google Scholar] [CrossRef] [Scilit]
- Sharifi, A. Resilient urban forms: A review of literature on streets and street networks. Build. Environ. 2019, 147, 171–187. [Google Scholar] [CrossRef] [Scilit]
- Athanasiadis, T.; Mylonas, P.; Avrithis, Y.; Kollias, S. Semantic image segmentation and object labeling. IEEE Trans. Circuits Syst. Video Technol. 2007, 17, 298–312. [Google Scholar] [CrossRef] [Scilit]
- Arribas-Bel, D.; Fleischmann, M. Spatial Signatures-Understanding (urban) spaces through form and function. Habitat Int. 2022, 128, 102641. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Li, C. Characterizing urban spatial structure through built form typologies: A new framework using clustering ensembles. Land Use Policy 2024, 141, 107166. [Google Scholar] [CrossRef] [Scilit]
- Ma, Y.; Wang, L.; Zhang, J. Quantifying Spatial Openness and Visual Perception in Historic Urban Environments. Buildings 2025, 15, 3295. [Google Scholar] [CrossRef] [Scilit]
- Jiang, F.; Ma, J. Predicting urban vitality at regional scales: A deep learning approach to modelling population density and pedestrian flows. Smart Cities 2025, 8, 58. [Google Scholar] [CrossRef] [Scilit]
- Awad, M.M. A Morphological Model for Extracting Road Networks from High-Resolution Satellite Images. J. Eng. 2013, 2013, 243021. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.; Fu, Y.; Gu, L.; Yan, C.; Harada, T.; Huang, G. Frequency-aware feature fusion for dense image prediction. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 10763–10780. [Google Scholar] [CrossRef] [Scilit]
- Xiao, T.; Liu, Y.; Huang, Y.; Li, M.; Yang, G. Enhancing multiscale representations with transformer for remote sensing image semantic segmentation. IEEE Trans. Geosci. Remote Sens. 2023, 61, 1–16. [Google Scholar] [CrossRef] [Scilit]
- Yoshida, H.; Omae, M. An approach for analysis of urban morphology: Methods to derive morphological properties of city blocks by using an urban landscape model and their interpretations. Comput. Environ. Urban Syst. 2005, 29, 223–247. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Fan, X.; Jiang, J.; Liu, R.; Luo, Z. Learning a deep multi-scale feature ensemble and an edge-attention guidance for image fusion. IEEE Trans. Circuits Syst. Video Technol. 2021, 32, 105–119. [Google Scholar] [CrossRef] [Scilit]
- Kuang, Z.; Zhang, J.; Li, Y.; Fukuda, T. Preserving architectural heritage in urban renewal: A stable diffusion model framework for automated historical facade generation. npj Herit. Sci. 2025, 13, 256. [Google Scholar] [CrossRef] [Scilit]
- Yang, F.; Li, W.; Li, L.; Yang, M.; Zhang, J. DWSF-Net: A Dynamic Wavelet-based Spatial-frequency Fusion Network for Multispectral Object Detection. IEEE Trans. Multimed. 2026; early access.











| Semantic Category | Morphological Function | Quantitative Indicators |
|---|---|---|
| Build | Spatial enclosure system | Enclosure ratio; Façade continuity; Skyline structure |
| Main Road | Spatial accessibility framework | Accessibility; Spatial hierarchy; Network connectivity |
| Road Append | Human-scale interface system | Interface permeability; Pedestrian interaction intensity |
| Natural | Environmental and visual background | Spatial openness; Environmental quality |
| Dynamic | Activity and behavioral layer | Urban vitality; Activity intensity |
| Semantic Category | Corresponding Spatial Function Description | CamVid Subclasses | CamVid Proportion (%) | Cityscapes Subclasses | Cityscapes Proportion (%) | RGB |
|---|---|---|---|---|---|---|
| Build | “Hard Edge” and physical enclosure framework of urban space | Building, Bridge, Tunnel, Archway, Wall, Column | 26.635% | Building, Fence, Wall, Bridge, Tunnel | 14.479% | (255, 0, 0) Red |
| Main Road | Fundamental framework and flow base of horizontal urban space | Road, LaneMkgsDriv, LaneMkgsNonDriv | 29.004% | road | 21.5786% | (0, 0, 255) Blue |
| Road Append | Human scale elements and “Soft Edge” of street interfaces | TrafficLight, TrafficCone, SignSymbol, Misc_Sign | 8.023% | Sidewalk, Pole, Parking, Traffic sign | 5.4542% | (255, 255, 0) Yellow |
| Natural | Ecological regulation and visual background of the built environment | Tree, VegetationMisc, Sky | 27.050% | Vegetation, Sky, Terrain, Ground | 13.1537% | (0, 255, 0) Green |
| Dynamic | Indicators of activity intensity and urban vitality | Pedestrian, Child, Bicyclist, Animal, Car, SUV | 6.278% | Car, Person, Bicycle, Dynamic, Truck | 5.6598% | (255, 0, 255) Magenta |
| Background | Semantically neutral areas and non-structural boundaries | Void | 3.010% | static | 39.6738% | (0, 0, 0) Black |
| Total | Full coverage of urban street-view elements | all | 100.000% | all | 100.000% | - |
| Experimental Setting | Details |
|---|---|
| CPU | Intel(R) Xeon(R) Gold 6346 CPU @ 3.10 GHz |
| GPU | NVIDIA A40 ×2 |
| Operating System | Linux 6.8.0-101-generic 22.04.1-Ubuntu x86-64 |
| Python Version | Python 3.10.12 |
| Pytorch Version | PyTorch 2.8.0+cu128 |
| CUDA Version | CUDA 12.2 |
| Training Setting | Details |
|---|---|
| Learning Rate | 0.00001 |
| Total Epochs | 200 |
| Weight Decay | 0.005 |
| Input Dimensions | 640 × 640 |
| Training Set Split | 60% |
| Validation Set Split | 30% |
| Test Set Split | 10% |
| Models | Backbone | Accuracy | Precision | Recall | F1-Score | IoU | Params |
|---|---|---|---|---|---|---|---|
| Camvid dataset | |||||||
| DeepLabV3Plus | resnet101 | 82.15 | 68.33 | 72.53 | 69.79 | 57.46 | 45.78 M |
| resnext101 | 79.73 | 66.22 | 71.11 | 67.92 | 55.01 | 45.44 M | |
| mit_b3 | 86.83 | 74.66 | 76.79 | 75.40 | 64.24 | 45.23 M | |
| Hiera_tiny | 86.45 | 74.22 | 77.13 | 75.24 | 64.96 | 39.25 M | |
| Segformer | resnet101 | 80.44 | 66.18 | 70.45 | 67.70 | 55.29 | 43.94 M |
| resnext101 | 78.01 | 63.52 | 67.19 | 64.73 | 51.96 | 43.60 M | |
| mit_b3 | 86.06 | 73.32 | 75.47 | 74.09 | 62.74 | 44.60 M | |
| Hiera_tiny | 86.62 | 73.25 | 75.45 | 74.96 | 62.25 | 38.85 M | |
| Unet | resnet101 | 75.59 | 64.65 | 66.38 | 63.00 | 49.96 | 51.62 M |
| resnext101 | 73.40 | 60.55 | 65.48 | 61.48 | 48.53 | 51.28 M | |
| mit_b3 | 82.55 | 69.08 | 75.02 | 70.76 | 58.79 | 47.36 M | |
| Hiera_tiny | 82.21 | 68.56 | 75.45 | 69.82 | 58.62 | 41.75 M | |
| PIDNet | PIDNet-L | 74.75 | 61.23 | 66.05 | 62.49 | 49.18 | 36.99 M |
| HieraWaveSeg | Hiera-transformer_tiny | 91.60 | 81.86 | 83.11 | 81.97 | 73.15 | 44.49 M |
| Cityscapes dataset | |||||||
| DeepLabV3Plus | resnet101 | 86.06 | 82.90 | 82.94 | 82.49 | 71.38 | 45.78 M |
| resnext101 | 85.27 | 82.18 | 82.10 | 81.63 | 70.21 | 45.44 M | |
| mit_b3 | 88.23 | 85.61 | 85.22 | 84.09 | 75.05 | 45.23 M | |
| Hiera_tiny | 88.56 | 86.24 | 85.64 | 85.26 | 75.29 | 39.25 M | |
| Segformer | resnet101 | 85.03 | 81.78 | 81.35 | 80.93 | 69.22 | 43.94 M |
| resnext101 | 84.34 | 81.02 | 80.54 | 80.07 | 68.01 | 43.60 M | |
| mit_b3 | 87.96 | 85.07 | 84.80 | 84.55 | 74.25 | 44.60 M | |
| Hiera_tiny | 87.89 | 85.34 | 84.14 | 84.10 | 74.55 | 38.85 M | |
| Unet | resnet101 | 85.94 | 82.88 | 82.47 | 82.12 | 70.85 | 51.62 M |
| resnext101 | 85.22 | 81.89 | 81.71 | 81.25 | 69.64 | 51.28 M | |
| mit_b3 | 88.21 | 85.61 | 85.04 | 85.04 | 74.99 | 47.36 M | |
| Hiera_tiny | 88.25 | 84.11 | 85.97 | 85.26 | 74.49 | 41.75 M | |
| PIDNet | PIDNet-L | 77.81 | 74.48 | 73.35 | 72.67 | 58.59 | 36.99 M |
| HieraWaveSeg | Hiera-transformer_tiny | 89.84 | 87.51 | 87.04 | 86.96 | 77.79 | 44.49 M |
| Backbone | Accuracy | Precision | Recall | F1-Score | IoU | Params |
|---|---|---|---|---|---|---|
| Efficientnet_b5 | 83.62 | 70.21 | 74.32 | 71.69 | 59.97 | 35.29 M |
| Mvitv2_tiny | 91.55 | 81.38 | 82.48 | 81.87 | 72.87 | 37.98 M |
| Resnet50 | 89.40 | 77.47 | 79.91 | 78.56 | 68.74 | 48.13 M |
| Resnext50 | 89.25 | 77.31 | 79.75 | 78.4 | 68.43 | 47.62 M |
| Swin-Transformer_tiny | 89.40 | 78.02 | 80.15 | 78.95 | 68.98 | 41.94 M |
| Hiera-Transformer_tiny | 91.60 | 81.86 | 83.11 | 81.97 | 73.15 | 44.49 M |
| NeckFusion | Accuracy | Precision | Recall | F1-Score | IoU | Params |
|---|---|---|---|---|---|---|
| ChannelMapper WithPooling | 90.27 | 78.86 | 81.18 | 79.90 | 70.39 | 35.96 M |
| GDFusion | 90.49 | 79.22 | 81.44 | 80.23 | 70.80 | 36.81 M |
| ASFF | 90.83 | 80.09 | 81.57 | 80.76 | 71.40 | 38.28 M |
| PANet | 90.68 | 79.84 | 81.68 | 80.67 | 71.19 | 41.23 M |
| UniFusion | 90.80 | 79.82 | 81.95 | 80.77 | 71.41 | 38.34 M |
| PAN (Ours) | 91.60 | 81.86 | 83.11 | 81.97 | 73.15 | 44.49 M |
| Deoder | Accuracy | Precision | Recall | F1-Score | IoU | Params |
|---|---|---|---|---|---|---|
| Mask2Former Decoder | 87.92 | 75.30 | 81.32 | 77.33 | 67.09 | 50.13 M |
| RFDetrSegment Decoder | 90.33 | 79.15 | 81.05 | 79.99 | 70.35 | 39.94 M |
| SAM2 Decoder | 91.27 | 80.74 | 82.34 | 81.44 | 72.23 | 50.13 M |
| SegFormer Decoder | 90.52 | 79.47 | 81.35 | 80.31 | 70.79 | 39.82 M |
| Unet Decoder | 91.07 | 80.59 | 82.17 | 81.30 | 71.96 | 45.86 M |
| HieraWaveSeg Decoder (Ours) | 91.60 | 81.86 | 83.11 | 81.97 | 73.15 | 44.49 M |
| No. | (CW) | (CW) | : | Accuracy | Precision | Recall | IoU |
|---|---|---|---|---|---|---|---|
| 1 | × | × | 1:0 | 87.21 | 74.33 | 80.51 | 66.82 |
| 2 | ✔ | × | 1:0 | 87.95 | 75.12 | 81.33 | 67.29 |
| 3 | × | × | 0:1 | 89.15 | 78.43 | 80.14 | 68.32 |
| 4 | × | ✔ | 0:1 | 89.84 | 79.56 | 80.45 | 69.11 |
| 5 | × | × | 0.5:0.5 | 90.74 | 80.12 | 82.04 | 71.55 |
| 6 | ✔ | × | 0.5:0.5 | 91.22 | 80.44 | 82.59 | 72.29 |
| 7 | × | ✔ | 0.5:0.5 | 91.31 | 80.62 | 82.78 | 72.43 |
| 8 | ✔ | ✔ | 0.5:0.5 | 91.60 | 81.86 | 83.11 | 73.15 |
| Neck Fusion | MLP | Wave Fusion | Accuracy | Precision | Recall | F1-Score | IoU |
|---|---|---|---|---|---|---|---|
| - | - | - | 86.97 | 73.68 | 77.84 | 75.32 | 64.55 |
| √ | - | - | 90.37 | 79.03 | 81.45 | 80.11 | 70.52 |
| - | √ | - | 88.80 | 76.50 | 79.33 | 77.71 | 67.60 |
| - | - | √ | 90.35 | 78.95 | 81.43 | 80.07 | 70.47 |
| √ | √ | - | 90.74 | 79.66 | 81.72 | 80.57 | 71.16 |
| - | √ | √ | 90.49 | 79.02 | 81.41 | 80.09 | 70.65 |
| √ | - | √ | 90.45 | 79.33 | 81.67 | 80.39 | 70.83 |
| √ | √ | √ | 91.60 | 81.86 | 83.11 | 81.97 | 73.15 |
| WaveFusion | Accuracy | Precision | Recall | F1-Score | IoU | Glops | Inference Time | ||
|---|---|---|---|---|---|---|---|---|---|
| DWT Edge Attention | Concat | Resnet | |||||||
| √ | √ | - | 91.09 | 80.58 | 82.11 | 81.27 | 71.99 | 114.14 | 0.0441 s |
| - | √ | √ | 91.07 | 80.52 | 82.04 | 81.21 | 71.91 | 107.54 | 0.0218 s |
| √ | - | √ | 89.77 | 78.44 | 80.93 | 79.56 | 69.61 | 94.33 | 0.0145 s |
| √ | √ | √ | 91.60 | 81.86 | 83.11 | 81.97 | 73.15 | 114.14 | 0.0144 s |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Guan, X.; Luo, K. Urban Morphology-Oriented Streetscape Segmentation via Hierarchical Transformer and Frequency-Aware Feature Learning. Buildings 2026, 16, 2180. https://doi.org/10.3390/buildings16112180
Guan X, Luo K. Urban Morphology-Oriented Streetscape Segmentation via Hierarchical Transformer and Frequency-Aware Feature Learning. Buildings. 2026; 16(11):2180. https://doi.org/10.3390/buildings16112180
Chicago/Turabian StyleGuan, Xiyue, and Kejun Luo. 2026. "Urban Morphology-Oriented Streetscape Segmentation via Hierarchical Transformer and Frequency-Aware Feature Learning" Buildings 16, no. 11: 2180. https://doi.org/10.3390/buildings16112180
APA StyleGuan, X., & Luo, K. (2026). Urban Morphology-Oriented Streetscape Segmentation via Hierarchical Transformer and Frequency-Aware Feature Learning. Buildings, 16(11), 2180. https://doi.org/10.3390/buildings16112180
