Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (220)

Search Parameters:
Keywords = urban scene segmentation

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
28 pages, 71265 KB  
Article
Sharing Cultural Values Through 3D Point-Cloud-Based Documentation of Transylvanian Heritage
by Alina Elena Voinea, Calin Neamtu and Virgil Pop
Remote Sens. 2026, 18(16), 2841; https://doi.org/10.3390/rs18162841 - 21 Aug 2026
Viewed by 141
Abstract
This paper presents a pilot educational workflow that couples 3D remote sensing with heritage-driven pedagogy by engaging architecture master’s students in the documentation and digital archiving of Transylvanian cultural sites. Using terrestrial and mobile 3D scanning, students documented multiple typologies—wooden churches (Târgușor, Tioltiur), [...] Read more.
This paper presents a pilot educational workflow that couples 3D remote sensing with heritage-driven pedagogy by engaging architecture master’s students in the documentation and digital archiving of Transylvanian cultural sites. Using terrestrial and mobile 3D scanning, students documented multiple typologies—wooden churches (Târgușor, Tioltiur), historical ensembles (Mociu, Coplean), industrial sites (1 Mai–Luduș, Vânătorilor–Luduș), and an urban street segment (Potaissa)—to generate dense point clouds that served as the basis for geometric reconstruction, semantic interpretation, and condition assessment. The study describes how the characteristics of different construction systems (timber, brick, stone, mixed structures) relate to point-cloud quality, survey coverage, and subsequent CAD/BIM drafting, with attention to the qualitative reading of minor deformations in wooden churches and of degradation patterns in masonry and industrial buildings. We also consider how artefacts in the data (noise, occlusions, registration errors) affect scene understanding and the interpretation of derived observations relevant to condition assessment and, prospectively, to monitoring. For the Tioltiur dual-sensor case, the TLS and SLAM datasets were compared through an internal CloudCompare registration check (final RMS 0.1121 on 50,000 points, fixed scale 1.0 and theoretical overlap 100%), surface-density displays (r = 0.005 for the Z+F dataset and for the GeoSLAM dataset), fitted-wall-plane readings (dip values around 89 deg. and 85 deg.) and a longitudinal section documenting roof/vault deformation. Beyond technical performance, the paper examines the self-reported formative impact on students’ digital skills and their understanding of cultural values, arguing that participation in 3D data acquisition, processing, and interpretation positions them as co-creators of a living digital archive. Pre- and post-workshop questionnaires (n = 13 each) are analysed descriptively—counts, percentages and medians with interquartile ranges—because the two instruments are unmatched and carry no shared identifier, so no paired test is applied; post-workshop self-ratings of technical competence, heritage understanding, archival awareness and collaboration were consistently high (medians 4–5), with uneven access to VR the main gap. By connecting point-cloud-based documentation workflows with heritage education, the project outlines a transferable, monitoring-ready baseline model in which 3D remote sensing supports both careful documentation and the transmission of regional identity and cultural meaning in architectural training. As an exploratory pilot with a small, self-reported sample, the study reports descriptive and qualitative findings rather than validated metric or statistical results. Full article
Show Figures

Figure 1

21 pages, 11093 KB  
Article
A Lightweight RGB-LiDAR Feature Recalibration Network for Large-Scale 3D Scene Understanding
by Weifeng Zhai and Zexi Tan
Optics 2026, 7(4), 59; https://doi.org/10.3390/opt7040059 - 13 Aug 2026
Viewed by 173
Abstract
Semantic segmentation of large-scale 3D point clouds is a fundamental task in robotic perception, semantic mapping, and urban scene understanding. Existing methods mainly rely on geometric information, which limits their ability to distinguish semantic categories with similar spatial structures. To address this issue, [...] Read more.
Semantic segmentation of large-scale 3D point clouds is a fundamental task in robotic perception, semantic mapping, and urban scene understanding. Existing methods mainly rely on geometric information, which limits their ability to distinguish semantic categories with similar spatial structures. To address this issue, this paper proposes a lightweight cross-modal feature learning framework that adaptively integrates geometric coordinates and RGB color information. By exploiting the complementary characteristics of spatial structure and visual appearance during feature encoding, the proposed method enhances feature discriminability while maintaining a compact model scale. Experiments on the Semantic3D dataset show that the proposed method achieves an mIoU of 87.1%, outperforming the original RandLA-Net and several representative approaches. Additional runtime and LiDAR-only cross-dataset experiments indicate the potential of the proposed structure for online outdoor point cloud perception. Full article
(This article belongs to the Topic Optical and Laser Scanning: Systems and Applications)
Show Figures

Figure 1

29 pages, 45575 KB  
Article
Fine-Grained Urban Vegetation Segmentation Under Two Imaging Views Based on Scale-Aware Mixture of Experts and Scene-Specific Optimization
by Yuhe Hu, Yujie Li, Nan Chen, Yuzhen Zhang, Yangle Jin, Yiqiu Chen and Jia Wang
Remote Sens. 2026, 18(16), 2701; https://doi.org/10.3390/rs18162701 - 11 Aug 2026
Viewed by 272
Abstract
High-precision urban vegetation mapping is essential for assessing carbon sink capacities, mitigating the urban heat island effect, and supporting sustainable development. Although deep learning and high-resolution remote sensing have advanced automated vegetation monitoring, existing models still face challenges when a common segmentation architecture [...] Read more.
High-precision urban vegetation mapping is essential for assessing carbon sink capacities, mitigating the urban heat island effect, and supporting sustainable development. Although deep learning and high-resolution remote sensing have advanced automated vegetation monitoring, existing models still face challenges when a common segmentation architecture is evaluated under different imaging geometries. In this study, Cityscapes and ISPRS Vaihingen are treated as two independent benchmarks representing perspective street-level imagery and orthographic aerial imagery, rather than as simultaneous cross-view inputs. “Background dominance” caused by perspective distortion and the “gridding artifacts” inherent in orthographic textures severely constrain segmentation accuracy across varying vegetation scales, particularly for small targets. To address these limitations, we propose a Scale-Aware Mixture of Experts (SA-MoE) architecture for fine-grained vegetation segmentation under two distinct imaging views, together with a scene-specific optimization strategy. The core SA-MoE framework consists of two main components. First, the spatial gating network uses a temperature polarization mechanism with τ = 0.5 to adjust the initial logit maps, sharpening expert-weight differences while preserving stable gradient propagation. Second, we use a heterogeneous expert group with five parallel branches: a pixel-level expert, three spatial experts with different dilation rates, and a global average-pooling expert. A dynamic pixel-level weighted fusion mechanism is then applied, decoupling feature extraction from receptive-field allocation. Furthermore, to address the heterogeneity of “hard samples” and “label noise” across the two benchmark settings, we introduce a scene-specific optimization strategy. Our findings show that the Focal-Dice (FD) loss is more suitable for perspective scenes with severe target imbalance and hard-to-classify vegetation targets, whereas the Cross-Entropy (CE) loss is more robust to boundary jitter in orthographic imagery. Comparative experiments on the Cityscapes (perspective view) and ISPRS Vaihingen (orthographic view) datasets reveal that SA-MoE achieves a highly competitive balance between computational efficiency and fine-grained segmentation, particularly in micro-target recall. Notably, the recall for extra-small (XS) scale targets in the aerial dataset improved by 3.21 percentage points compared to the second-best model. For the street-level dataset, our model achieved competitive global performance in terms of Overall Accuracy (OA), Precision, and F1-Score. However, we also observed a performance trade-off, where Transformer-based models maintained an advantage in preserving fine boundary details for these extra-small targets. In the routing analysis, we observed a pattern that we refer to as “receptive field inversion”, in which the model assigns lower weights to large-dilation experts for large canopy regions in orthophotos. We interpret this pattern as a plausible routing hypothesis. Overall, SA-MoE offers an efficient and adaptive solution for urban vegetation mapping under two imaging views. Full article
(This article belongs to the Special Issue Innovations in Remote Sensing Image Analysis)
Show Figures

Figure 1

27 pages, 9653 KB  
Article
Uncertainty-Aware Vision-Based Landing-Site Perception for Autonomous UAV Landing in Urban Environments
by Jingjing Qian, Yang Cheng, Junhong Wu, Bing Liu and Wei Dai
Drones 2026, 10(8), 610; https://doi.org/10.3390/drones10080610 - 7 Aug 2026
Viewed by 218
Abstract
Autonomous UAV landing in urban scenes requires a perception module that can identify candidate landing surfaces, reject structural and dynamic hazards, and express uncertainty before a downstream controller commits to a landing maneuver. This study reformulates UAV landing perception as a unified three-class [...] Read more.
Autonomous UAV landing in urban scenes requires a perception module that can identify candidate landing surfaces, reject structural and dynamic hazards, and express uncertainty before a downstream controller commits to a landing maneuver. This study reformulates UAV landing perception as a unified three-class landing-safety segmentation problem by relabeling UAVid, UDD6, and VDD into candidate landing area, structural obstacle, and critical hazard classes. The Landing3 dataset is constructed with an implementation-consistent relabeling protocol in which critical hazards override other labels and only sufficiently large connected components of source-specific candidate classes are retained as landing candidates. Two real-time segmentation models are evaluated under this unified task: PIDNet, a CNN-based multi-branch model with explicit boundary modeling, and SCTNet, a Transformer-guided model with semantic alignment for long-range context modeling. A post hoc conformal prediction (CP) module then converts softmax outputs into pixel-wise prediction sets, and the final candidate landing region is extracted only from pixels whose prediction set is the singleton candidate-landing class. Experiments compare the two models in terms of best validation mIoU, class-wise IoU, resolution-dependent accuracy–speed trade-offs, power-constrained FPS and complexity, qualitative candidate-area visualization, and CP-derived safe-area quality. A fixed-checkpoint sensitivity analysis further shows that the comparative model ranking remains stable across the tested connected-component threshold settings. SCTNet provides a lighter model and higher throughput under most tested resolutions and power limits, whereas PIDNet preserves higher safe-area recall, safe IoU, and spatial coherence after CP filtering. These results show that reliable UAV landing perception requires joint consideration of cross-dataset task definition, real-time model efficiency, and uncertainty-aware candidate-area extraction, rather than semantic segmentation accuracy alone. Full article
(This article belongs to the Section Innovative Urban Mobility)
Show Figures

Figure 1

45 pages, 52572 KB  
Article
Multi-Sensor Fusion SLAM Based on LiDAR, IMU and GPS for Structured Urban Scenes
by Jiajia Lu, Yue Shen, Xu Wang and Fuyang Ke
J. Imaging 2026, 12(8), 345; https://doi.org/10.3390/jimaging12080345 - 30 Jul 2026
Viewed by 329
Abstract
Aiming at the current SLAM (Simultaneous Localization and Mapping) algorithms in urban scenarios, which have problems such as elevation drift, odometry drift, and the appearance of false loop closures, a tightly coupled SLAM method with LiDAR and inertial guidance is proposed. In the [...] Read more.
Aiming at the current SLAM (Simultaneous Localization and Mapping) algorithms in urban scenarios, which have problems such as elevation drift, odometry drift, and the appearance of false loop closures, a tightly coupled SLAM method with LiDAR and inertial guidance is proposed. In the front-end, a raster-based point cloud feature extraction method is introduced, enabling simultaneous segmentation and extraction of line, surface, and ground features. Utilizing the alignment results of line and surface features as the initial value for ground point alignment, interpolation weights are determined based on roll and pitch angle errors, effectively reducing global elevation errors through frame-by-frame constraints. The back-end employs an error state-based Kalman filter (ESKF) for GPS and IMU data fusion, enhancing the validity of true state estimation. A Scan Context loop closure detection method is designed, augmented by GPS detection as an auxiliary loop closure constraint to mitigate false loop closures. A global factor graph optimization model is also proposed. Experimental results demonstrate that, compared to existing open-source algorithms, the proposed method exhibits improved performance in structured urban scenes, reducing the average RMSE APE by 47.4% compared with LiDAR-only methods and by 22.9% compared with tightly coupled LiDAR-inertial methods. This work highlights the potential of multi-sensor fusion SLAM for achieving high-precision 3D localization and mapping in complex urban environments. Full article
(This article belongs to the Section Computer Vision and Pattern Recognition)
Show Figures

Figure 1

18 pages, 48650 KB  
Article
PMS-Net: A Real-Time Instance Segmentation Framework with Large Receptive Fields for Urban Driving Scenes
by Ling Zhang and Zhuang Xiong
Algorithms 2026, 19(8), 620; https://doi.org/10.3390/a19080620 - 24 Jul 2026
Viewed by 521
Abstract
With the rapid development of edge AI chips and autonomous driving algorithms, autonomous driving perception systems are evolving toward higher accuracy, lower latency, and lightweight deployment. This trend places greater demands on real-time instance segmentation algorithms in terms of multi-scale feature representation, spatial [...] Read more.
With the rapid development of edge AI chips and autonomous driving algorithms, autonomous driving perception systems are evolving toward higher accuracy, lower latency, and lightweight deployment. This trend places greater demands on real-time instance segmentation algorithms in terms of multi-scale feature representation, spatial detail modeling, and edge deployment efficiency. To address these challenges, this paper proposes PMS-Net (Progressive Multi-Scale Network), a network designed for real-time instance segmentation. PMS-Net adopts a progressive multi-scale feature modeling mechanism that progressively enlarges the receptive field while integrating semantic and fine-grained spatial information across different scales. This enables efficient collaboration between local features and global contextual information, thereby enhancing scale-awareness and feature representation while maintaining a lightweight architecture. In addition, efficient feature encoding and dynamic feature reconstruction are incorporated to further improve spatial alignment, boundary recovery, and semantic continuity for complex scene modeling. Experimental results show that PMS-Net achieves 36.7% Mask mAP50 and 175 FPS on the Cityscapes dataset, outperforming the baseline by 2.9%. Deployment experiments on the NVIDIA Jetson Orin NX platform further demonstrate that PMS-Net achieves 35.2% Mask mAP50 and 98 FPS, improving the baseline by 3.7% and 14.0%, respectively. These results validate the effectiveness and practicality of PMS-Net for real-time edge-deployed autonomous driving applications. Full article
Show Figures

Figure 1

23 pages, 2800 KB  
Article
A Topology-Aware Lane Instance Segmentation Network Based on Dynamic Snake Convolution and High-Resolution Feature Pyramid
by Sen Wang, Baohua Guo, Weifan Gu, Xiaoyu Zhang, Anthony Sigama and David Bassir
Electronics 2026, 15(14), 3213; https://doi.org/10.3390/electronics15143213 - 21 Jul 2026
Viewed by 318
Abstract
Accurate lane detection is essential for autonomous driving, but the elongated geometry of lane markings and the information loss caused by perspective projection still make continuous lane perception challenging. Existing Convolutional Neural Network (CNN) methods are limited by fixed receptive fields and repeated [...] Read more.
Accurate lane detection is essential for autonomous driving, but the elongated geometry of lane markings and the information loss caused by perspective projection still make continuous lane perception challenging. Existing Convolutional Neural Network (CNN) methods are limited by fixed receptive fields and repeated downsampling, which can weaken narrow curvilinear structures and distant lane details. To address these issues, this paper proposes an enhanced single-stage lane instance segmentation framework based on YOLOv11. Dynamic Snake Convolution (DSConv) is integrated into the feature extraction network as an adaptive sampling operator for finite-width lane-mask representations, enabling local sampling trajectories to better follow narrow curved lane structures. A high-resolution P2 branch (Stride = 4) is further constructed in the Feature Pyramid Network (FPN) to compensate for fine-grained spatial details lost in deep feature maps. A reproducible point-to-mask training and mask-to-point evaluation pipeline is also established for the sparse TuSimple annotations. Experiments on the TuSimple benchmark, which mainly represents structured daytime highway scenes, show that the proposed method improves mask Precision, Recall, and mAP50 by 8.5, 9.0, and 6.8 percentage points, respectively, compared with YOLOv11-seg, while reaching 103.09 FPS under the reported testing configuration. Qualitative results on a newly constructed CULane subset further indicate improved lane continuity under reflections, illumination changes, night scenes, and complex urban backgrounds, while failure cases reveal remaining sensitivity to lane-like linear structures and weak lane visibility. The results suggest a practical accuracy–efficiency trade-off within the YOLO-style segmentation framework, although robustness in broader real-world scenarios still requires further verification. Full article
Show Figures

Figure 1

33 pages, 10785 KB  
Article
Lightweight Semantic Perception from UAV-Borne Visual Sensors via Conflict-Suppressed Heterogeneous Expert Distillation
by Feng Ouyang, Yongpeng Ding, Miao Qin, Weiting Xie and Chao Zhou
Sensors 2026, 26(14), 4509; https://doi.org/10.3390/s26144509 - 16 Jul 2026
Viewed by 423
Abstract
UAV-borne visual sensors provide high-resolution aerial observations for low-altitude scene understanding, urban monitoring, traffic observation, emergency inspection, and infrastructure assessment. However, semantic perception from UAV visual sensor data remains challenging because aerial images often contain dense small objects, elongated road structures, fragmented boundaries, [...] Read more.
UAV-borne visual sensors provide high-resolution aerial observations for low-altitude scene understanding, urban monitoring, traffic observation, emergency inspection, and infrastructure assessment. However, semantic perception from UAV visual sensor data remains challenging because aerial images often contain dense small objects, elongated road structures, fragmented boundaries, scale variations caused by flight-altitude changes, oblique viewpoints, and strict onboard or edge computational constraints. To address these challenges, this paper proposes MEKD-UAVSeg, a lightweight semantic perception framework based on conflict-suppressed heterogeneous expert distillation. During training, a Transformer-based semantic expert provides global contextual understanding and region-level class consistency, while a Mamba-based spatial expert provides complementary structural guidance for roads, roofs, boundaries, and other continuous aerial structures. Both experts are used only during training, and the final inference model remains a compact CNN-based segmentation network. In addition, UAV-aware density and hard-region priors are designed to emphasize small-object-dense areas, boundary-sensitive regions, rare classes, and uncertain aerial categories. A conflict-suppressed reliability routing strategy is further developed to reduce inconsistent supervision between heterogeneous experts and selectively transfer reliable knowledge to the student model. Experiments on UAVid and UDD6 demonstrate that the proposed framework achieves a favorable accuracy–efficiency trade-off compared with representative CNN-, Transformer-, Mamba-, and hybrid-based UAV segmentation methods, without introducing expert-induced inference complexity. Full article
(This article belongs to the Section Vehicular Sensing)
Show Figures

Figure 1

25 pages, 12058 KB  
Article
DMSF-Net: A Dual-Encoder Multi-Source Feature Fusion Network for Fine-Grained Urban Green Space Segmentation
by Longzhen Jiao, Xiaoyong Zhang, Linlin Lu, Wei Cai, Shusheng Yin, Manbin Yuan, Zhengchao Chen and Qingting Li
Remote Sens. 2026, 18(14), 2308; https://doi.org/10.3390/rs18142308 - 9 Jul 2026
Viewed by 523
Abstract
The mapping and monitoring of urban green space (UGS) are of great significance for ecological assessment and sustainable development in urban settlements. However, for fine-grained classification tasks in high-resolution remote sensing imagery, existing methods suffer from significant challenges due to the high spectral [...] Read more.
The mapping and monitoring of urban green space (UGS) are of great significance for ecological assessment and sustainable development in urban settlements. However, for fine-grained classification tasks in high-resolution remote sensing imagery, existing methods suffer from significant challenges due to the high spectral similarity between low-growing and dense vegetation, as well as the complexity of spatial structures. To overcome these challenges, this paper proposes a Dual-Encoder Multi-Source Feature Fusion Network (DMSF-Net) for fine-grained urban green space segmentation. The proposed method constructs a parallel encoding structure for RGB and auxiliary features (NDVI and LBP), introduces an Adaptive Feature Fusion Module (AFFM) during the encoding phase to achieve dynamic weighted fusion of cross-source features, and designs a Boundary-Aware Up-Sampling Module (BAM) during the decoding phase to strengthen the representation of complex boundary regions through joint modeling of regional semantics and boundary information. Experimental results on a self-constructed UrbanGreen dataset and the publicly available Vaihingen dataset demonstrate the superior performance of DMSF-Net over existing mainstream methods across several evaluation metrics, achieving mIoU values of 82.27% and 74.73%, with improvements of 1.07% and 0.57% over the best baselines, respectively. The model demonstrates particularly strong discrimination capability for the fine-grained category of low vegetation. Ablation experiments further validate the usefulness of each structural module, with AFFM playing a key role in overall performance improvement, while the BAM improves boundary delineation as observed in visual comparisons. Through the synergistic integration of multi-source feature information and structural optimization, DMSF-Net effectively enhances fine-grained UGS segmentation in complex urban scenes, thereby providing an effective approach for high-resolution remote sensing-based urban ecological monitoring. Full article
(This article belongs to the Special Issue Monitoring Urban Environment from Space)
Show Figures

Figure 1

25 pages, 6125 KB  
Article
MCPF-Net: Multi-Stage LiDAR-Image Collaborative Perception Fusion Network for Point Cloud Semantic Segmentation of Urban Scenes
by Huchen Li, Wubiao Huang, Xiangda Lei, Bin Liu, Haibing Liu, Shihan Chen and Fei Deng
Remote Sens. 2026, 18(13), 2218; https://doi.org/10.3390/rs18132218 - 6 Jul 2026
Viewed by 499
Abstract
Multimodal fusion unlocks the potential of point cloud semantic segmentation, thereby driving advancements in surface observation and visual perception tasks. Although light detection and ranging (LiDAR) systems capture precise 3D structural geometry and optical images provide rich semantic and textural information, existing fusion [...] Read more.
Multimodal fusion unlocks the potential of point cloud semantic segmentation, thereby driving advancements in surface observation and visual perception tasks. Although light detection and ranging (LiDAR) systems capture precise 3D structural geometry and optical images provide rich semantic and textural information, existing fusion methods struggle with limited cross-modal perception and insufficient information complementarity. To address these limitations, we propose a multi-stage LiDAR-image collaborative perception fusion network (MCPFNet) for point cloud semantic segmentation of urban scenes. At the middle fusion stage, the network incorporates an elevation-guided geometric-aware fusion module and a semantic-aware cross-attention fusion module to enable bidirectional feature injection between LiDAR and image modalities. In the late fusion stage, a bidirectional adaptive fusion module further refines semantic representations through gated weighting and bidirectional cross-attention mechanisms. Extensive experiments on three multimodal datasets with different resolutions, i.e., ISPRS Vaihingen, N3C-California, and UAVScenes, demonstrate that MCPFNet outperforms existing fusion methods, achieving mIoUs of 74.51%, 95.15%, and 62.76%, respectively. Hence, our multi-stage fusion and bidirectional interaction strategy is more reliable and accurate than existing methods in performing segmentation across diverse and complex urban scenes. Full article
Show Figures

Figure 1

23 pages, 2122 KB  
Article
DSD-Mamba: Dual-Stream Semantic Segmentation of Remote Sensing Imagery via Dense-Sparse Fusion
by Xinyi Feng, Shaochen Jiang, Liejun Wang and Beibei Gao
Sensors 2026, 26(12), 3864; https://doi.org/10.3390/s26123864 - 17 Jun 2026
Viewed by 430
Abstract
High-resolution remote sensing image segmentation is important for urban mapping but remains challenging because of spectral ambiguity, large scale variations, fragmented elongated structures, and background interference. This study aims to improve semantic segmentation in complex aerial scenes by combining local feature extraction, selective [...] Read more.
High-resolution remote sensing image segmentation is important for urban mapping but remains challenging because of spectral ambiguity, large scale variations, fragmented elongated structures, and background interference. This study aims to improve semantic segmentation in complex aerial scenes by combining local feature extraction, selective multi-scale fusion, and global sequence modeling. We propose DSD-Mamba, an asymmetric dual-stream architecture with a ResNet-18 encoder. The Dense-Sparse Pyramid Fusion Module aligns multi-level features and applies dual Top-k selective value aggregation for cross-scale response filtering and background-response suppression. This Top-k operation is used as a feature-selection mechanism and is not intended to reduce the theoretical memory footprint of dense attention. Scale-Aware Strip Attention refines skip connections through horizontal and vertical dependency modeling, and the Dual-Stream Context Decoder combines a Mamba-based global branch with a CNN-based local branch during upsampling. Experiments were conducted on UAVid, ISPRS Vaihingen, and ISPRS Potsdam under a single-model inference protocol without test-time augmentation. DSD-Mamba achieved mIoU scores of 73.4%, 85.2%, and 87.2%, respectively. Ablation experiments on Vaihingen showed that DSPFM, SASA, and DSCD improved performance over the baseline when evaluated in this setting, with the full model reaching the highest mIoU. The method improves segmentation accuracy under the tested protocols, although its higher FLOPs indicate an accuracy-oriented rather than lightweight design. Full article
Show Figures

Figure 1

19 pages, 5482 KB  
Article
MAD-SAR: A Multi-Agent Agentic Engineering Framework for Landslide Detection Using Sentinel-1 SAR Imagery
by Kohei Arai
Information 2026, 17(6), 597; https://doi.org/10.3390/info17060597 - 15 Jun 2026
Viewed by 1066
Abstract
Rapid and accurate detection of landslide-affected areas is critical for disaster response and risk mitigation. Sentinel-1 SAR imagery offers all-weather, day-and-night observation capability, but existing deep learning approaches treat landslide detection as a single-pass segmentation problem, which limits performance in complex terrain where [...] Read more.
Rapid and accurate detection of landslide-affected areas is critical for disaster response and risk mitigation. Sentinel-1 SAR imagery offers all-weather, day-and-night observation capability, but existing deep learning approaches treat landslide detection as a single-pass segmentation problem, which limits performance in complex terrain where backscatter changes are confounded by soil moisture, surface roughness, urban double bounce, shadow, and layover effects. MAD-SAR, a rule-based agentic framework that coordinates anomaly detection, super-resolution, object detection, and semantic segmentation under a planning orchestrator and a physics-aware validation engine is proposed. The orchestrator selects specialist modules, their execution order, and the number of refinement iterations according to a scene complexity score computed from SAR-derived statistics. The physics-aware validation engine cross-checks every candidate detection against backscatter change thresholds, DEM-derived slope constraints, and radar geometry masks before any detection is committed to the output. MAD-SAR is evaluated on three Japanese disaster datasets: Hiroshima 2018, Kumamoto 2016, and Ibaraki 2019. On the held-out Ibaraki test event, the framework achieves an F1-score of 0.863 and IoU of 0.759, outperforming all baselines and reducing false alarms by 45% relative to standalone SegFormer. Ablation results confirm that each module contributes to the final performance. These results suggest that multi-module orchestration with embedded physical validation can meaningfully improve SAR-based landslide mapping, though broader validation across regions, sensor configurations, and failure mechanisms remains necessary. Full article
(This article belongs to the Special Issue AI-Based Image Processing and Computer Vision, 2nd Edition)
Show Figures

Figure 1

20 pages, 2220 KB  
Article
R2KAN-U-Net: A Novel Architecture Integrating Kolmogorov–Arnold Networks with Residual U-Net for Robust Traffic Sign Segmentation
by Taha Ben-Abbou, Houda El Omrani, Khalid El Fazazy, Mohamed Adnane Mahraz, Hamid Tairi and Jamal Riffi
Sensors 2026, 26(12), 3797; https://doi.org/10.3390/s26123797 - 15 Jun 2026
Viewed by 459
Abstract
Traffic sign segmentation is a fundamental component of intelligent transportation systems and autonomous driving, where reliable pixel-level perception is required under challenging real-world conditions such as illumination variations, occlusion, scale diversity, and complex urban backgrounds. In this work, we propose Residual–Recurrent Kolmogorov–Arnold Network [...] Read more.
Traffic sign segmentation is a fundamental component of intelligent transportation systems and autonomous driving, where reliable pixel-level perception is required under challenging real-world conditions such as illumination variations, occlusion, scale diversity, and complex urban backgrounds. In this work, we propose Residual–Recurrent Kolmogorov–Arnold Network U-Net (R2KAN-U-Net), where “R2” denotes the integration of residual convolutional learning and recurrent KAN-based feature refinement. The proposed architecture combines residual U-Net feature extraction, multi-scale KAN fusion, and recurrent KAN refinement to improve pixel-level traffic sign segmentation under challenging road-scene conditions. The proposed framework integrates three complementary components: (1) residual convolutional blocks for stable feature propagation; (2) a multi-scale KAN fusion bottleneck for capturing contextual information at different receptive fields; and (3) recurrent KAN refinement modules for iterative enhancement of discriminative features. Unlike conventional convolutional architectures, the proposed KAN-based formulation replaces linear transformations with learnable univariate functions, enabling adaptive nonlinear feature modeling. We conduct extensive experiments on a custom dataset containing 9300 annotated urban traffic scene images, as well as on the ADE20K and Cityscapes benchmarks. On the custom dataset, the proposed R2KAN-U-Net achieved a Dice coefficient of 0.92 and an IoU score of 0.89, providing a strong accuracy–efficiency trade-off for traffic-sign foreground segmentation. It achieves competitive segmentation accuracy compared with recent CNN-, transformer-, and state-space-based segmentation models while using fewer parameters and lower computational cost. Additional low-light experiments demonstrate improved segmentation stability, with R2KAN-U-Net achieving the highest low-light Dice score of 0.88 and a competitive low-light IoU of 0.79. Furthermore, the proposed architecture maintains competitive computational efficiency with only 24 M parameters, 44.8 G FLOPs, and near-real-time inference at 13 ms per image. The experimental results demonstrate that integrating KAN-based function-space learning with residual and multi-scale feature refinement provides an effective and computationally efficient solution for robust traffic sign segmentation in complex driving environments. Full article
(This article belongs to the Section Sensors and Robotics)
Show Figures

Figure 1

23 pages, 7625 KB  
Article
MultiDecNet: An Ensemble-Based Semantic Segmentation Architecture for Urban Scene Understanding
by Büşra Emek Soylu and Mehmet Serdar Güzel
Information 2026, 17(6), 540; https://doi.org/10.3390/info17060540 - 1 Jun 2026
Viewed by 603
Abstract
Semantic segmentation is a fundamental task in computer vision that aims to assign a categorical label to each pixel in an image, facilitating dense and detailed scene understanding. This pixel-level classification is especially crucial in autonomous driving, where accurate environmental perception is vital [...] Read more.
Semantic segmentation is a fundamental task in computer vision that aims to assign a categorical label to each pixel in an image, facilitating dense and detailed scene understanding. This pixel-level classification is especially crucial in autonomous driving, where accurate environmental perception is vital for dependable object detection and safe decision-making. In this study, we propose MultiDecNet, a novel multi-decoder semantic segmentation framework designed to capture both macroscopic scene layouts and fine-grained spatial boundaries in complex urban environments. Drawing inspiration from classical networks, MultiDecNet incorporates a parallel dual-branch decoding strategy that simultaneously leverages the multi-scale context modeling of the Pyramid Pooling Module (PPM) and the structural refinement capabilities of Atrous Spatial Pyramid Pooling (ASPP). To explore the impact of modern backbone representations, we structurally modernize the feature extraction pipeline by introducing the contemporary ConvNeXt convolutional architecture as an alternative to traditional ResNet101 backbones. We extensively evaluate and compare the baseline configurations alongside our proposed MultiDecNet using both ResNet101 and ConvNeXt-Large backbones on the benchmark Cityscapes dataset. The quantitative assessments demonstrate that the MultiDecNet architecture consistently provides highly competitive performance within the scope of this comparative study, with the MultiDecNet-ConvNeXt variant achieving favorable overall scores among the evaluated methods. Furthermore, a granular, class-wise IoU and training dynamics analysis reveals that while traditional networks retain competitive boundaries for localized minority targets, the modern ConvNeXt backbone ensures faster convergence stability and balanced contextual mastery over large-scale driving layouts. Ultimately, these findings offer critical insights into architectural synergy and backbone selection, presenting a robust, scalable, and well-balanced solution for advanced autonomous navigation systems. Full article
(This article belongs to the Special Issue Computer Vision for Security Applications, 2nd Edition)
Show Figures

Graphical abstract

38 pages, 42009 KB  
Article
Urban Morphology-Oriented Streetscape Segmentation via Hierarchical Transformer and Frequency-Aware Feature Learning
by Xiyue Guan and Kejun Luo
Buildings 2026, 16(11), 2180; https://doi.org/10.3390/buildings16112180 - 29 May 2026
Viewed by 614
Abstract
Semantic segmentation of street-view imagery has become an important computational tool for urban morphological analysis and the evaluation of street spatial quality. However, existing methods still struggle in complex urban environments. Major challenges include large variations in building façade scales, degradation of boundary [...] Read more.
Semantic segmentation of street-view imagery has become an important computational tool for urban morphological analysis and the evaluation of street spatial quality. However, existing methods still struggle in complex urban environments. Major challenges include large variations in building façade scales, degradation of boundary information, and severe class imbalance. These issues limit the ability of current models to capture structurally meaningful urban forms. To address these challenges, this study proposes a high-resolution street-view segmentation framework, termed HieraWaveSeg. The model aims not only to improve pixel-level segmentation accuracy but also to enhance the interpretability of urban morphology through structured representations of street space. Specifically, a Hiera Transformer backbone is employed to capture hierarchical spatial semantics. A Path Aggregation Network is further introduced to strengthen cross-scale feature interaction and improve structural consistency in complex scenes. In addition, a Wave Fusion module based on the Haar wavelet transform is incorporated to preserve fine-grained architectural details by enhancing high-frequency boundary and texture information during decoding. Unlike conventional segmentation approaches that primarily focus on object recognition, this study introduces a morphology-oriented semantic reconfiguration strategy. This strategy reorganizes original categories into functionally meaningful urban units. As a result, the segmentation outputs can be more directly linked to urban morphological indicators, such as façade continuity, spatial enclosure, and interface permeability, thereby improving interpretability in architectural and urban design contexts. To further address class imbalance, a composite loss function combining weighted cross-entropy and Dice loss is adopted, together with a median frequency balancing strategy. Experimental results on the CamVid and Cityscapes datasets demonstrate that the proposed method consistently outperforms several state-of-the-art baselines in both segmentation accuracy and structural preservation. Beyond quantitative improvements, the results indicate that the proposed framework generates more coherent and morphologically meaningful urban representations, supporting further quantitative analysis in urban morphology and architectural studies. Full article
Show Figures

Figure 1

Back to TopTop