Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (2,718)

Search Parameters:
Keywords = image semantic segmentation

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
25 pages, 1589 KB  
Article
SDFR-Net: A Stage-Asymmetric Spectral Diffusion and Frequency–Spatial Refinement Network for Brain Tumor MRI Segmentation
by Jingshi Lei, Hongwei Deng, Xicheng Fu, Yi Lei, Lei Xu and Qiangfei Wang
Symmetry 2026, 18(9), 1417; https://doi.org/10.3390/sym18091417 - 23 Aug 2026
Abstract
Accurate brain tumor segmentation from multi-modal magnetic resonance imaging (MRI) is essential for clinical diagnosis and treatment planning. However, effectively capturing long-range contextual information and fine lesion boundaries under limited computational budgets remains challenging. In this work, we propose SDFR-Net, a lightweight stage-asymmetric [...] Read more.
Accurate brain tumor segmentation from multi-modal magnetic resonance imaging (MRI) is essential for clinical diagnosis and treatment planning. However, effectively capturing long-range contextual information and fine lesion boundaries under limited computational budgets remains challenging. In this work, we propose SDFR-Net, a lightweight stage-asymmetric Spectral Diffusion and Frequency–Spatial Refinement Network for efficient 2.5D brain tumor MRI segmentation. Instead of applying identical processing across all hierarchical stages, SDFR-Net adopts stage-dependent spectral diffusion, stage-selective conditional refinement, and asymmetric cross-stage frequency-grid allocation to accommodate the distinct semantic and frequency characteristics of shallow and deep representations. The network consists of a Spectral Diffusion Encoder for spectral-domain contextual propagation, a Frequency–Spatial Enhancement Module for adaptive refinement of multi-scale skip features, and a lightweight Conditional Refinement Decoder for lesion-aware reconstruction. Experiments on the BraTS 2019 and BraTS 2020 datasets demonstrate that SDFR-Net achieves whole-tumor Dice scores of 0.856 and 0.880, respectively, while requiring only 1.33 M parameters. Ablation comparisons of stage-selective FiLM injection and symmetric versus asymmetric frequency-grid schedules further support the stage-asymmetric design. These results indicate that SDFR-Net provides a favorable accuracy–efficiency trade-off for resource-constrained brain tumor MRI segmentation. Full article
(This article belongs to the Section A: Computer Science)
Show Figures

Figure 1

27 pages, 17265 KB  
Article
How Visual Elements Shape Perceived Spatial Quality in Urban Waterfront Space: An Explainable Machine Learning Approach for Urban Landscape Planning
by Wenhan Li, Yinzhe Li, Gaoming Liang, Congxi Liu, Dezheng Kong and Yan Feng
Sustainability 2026, 18(16), 8610; https://doi.org/10.3390/su18168610 - 21 Aug 2026
Viewed by 109
Abstract
As China’s urbanization shifts toward quality-oriented development, urban regeneration increasingly prioritizes the perceived quality of public spaces to enhance urban vitality and advance sustainable urban living. This study takes Zhengzhou’s Dongfeng Canal, a revitalized urban core waterfront, as a case to develop a [...] Read more.
As China’s urbanization shifts toward quality-oriented development, urban regeneration increasingly prioritizes the perceived quality of public spaces to enhance urban vitality and advance sustainable urban living. This study takes Zhengzhou’s Dongfeng Canal, a revitalized urban core waterfront, as a case to develop a human–machine collaborative analytical framework for exploring nonlinear relationships between visual environmental features and human spatial quality perception. By integrating 779 geolocated panoramic images with volunteers’ subjective rating data, this study adopts deep learning-based semantic segmentation to quantify eight objective visual indicators (e.g., greenness, color diversity, spatial structure). A random forest (RF) model links these indicators to three perceptual dimensions: scenic beauty, safety, and recreational value. Adopting explainable artificial intelligence (SHAP and PDPs), the results indicate that: (1) greenness is positively associated with positive perceptions but exhibits a significant threshold effect; (2) color diversity and waterfront accessibility substantially improve user experience, while excessive uniformity and extreme openness negatively affect perceived spatial quality. These findings challenge the simplistic linear “more-is-better” assumption in urban design and highlight the value of balanced, context-sensitive spatial interventions. This study provides evidence-based, segment-specific strategies for urban waterfront regeneration, advancing people-centered planning that integrates ecological functionality, social inclusivity, and long-term sustainability via Geospatial Artificial Intelligence (GeoAI) and geospatial analytics. Full article
Show Figures

Figure 1

34 pages, 2181 KB  
Article
Weakly Supervised Remote Sensing Segmentation via Decoupled Cross-Modal Distillation and Semantic-Guided Refinement
by Jing Li, Yulin Cao, Xiantao Jiang, Dong Zhao and Dan Zhang
Remote Sens. 2026, 18(16), 2843; https://doi.org/10.3390/rs18162843 - 21 Aug 2026
Viewed by 95
Abstract
Pixel-level annotation of remote sensing imagery is costly, motivating weakly supervised semantic segmentation (WSSS) using only image-level labels. However, class activation maps (CAMs) often highlight only discriminative sub-regions and fail to separate adjacent land-cover regions, particularly in remote sensing scenes characterized by densely [...] Read more.
Pixel-level annotation of remote sensing imagery is costly, motivating weakly supervised semantic segmentation (WSSS) using only image-level labels. However, class activation maps (CAMs) often highlight only discriminative sub-regions and fail to separate adjacent land-cover regions, particularly in remote sensing scenes characterized by densely co-occurring land-cover classes and substantial variations in object scale. To address these limitations, we propose a three-stage framework that integrates complementary priors from Contrastive Language–Image Pre-training (CLIP), Self-Distillation with No Labels version 2 (DINOv2), and the Segment Anything Model (SAM). First, a lightweight CLIP adapter aligns vision–language priors with remote sensing imagery, while sigmoid-based multi-label decoupled distillation replaces class-competitive distillation with independent class-wise supervision, producing more complete CAMs. Second, DINOv2-guided feature clustering decomposes large merged regions before SAM prompt generation, while Spatial–Semantic Constraints are used to construct confidence-guided point-and-box prompts and reject excessively expanded or semantically inconsistent masks, thereby generating reliable pseudo-labels. Finally, a compact segmentation network is initialized with the weights learned in Stage 1 and retrained using the refined pseudo-labels generated in Stage 2, eliminating the need for foundation models during inference. Experiments on the Potsdam, LoveDA, and DeepGlobe datasets show that the proposed method achieves mean intersection over union (mIoU) scores of 53.16%, 52.66%, and 62.98%, respectively, outperforming state-of-the-art WSSS baselines by 6.55, 1.16, and 1.27 percentage points, respectively. These results demonstrate the effectiveness and generalizability of the proposed framework across diverse remote sensing scenarios under image-level supervision. Full article
Show Figures

Figure 1

18 pages, 3993 KB  
Article
Rail Light-Strip Abnormality Analysis from Color Inspection Images Using an Improved SegFormer and Geometric Rules
by Haoran Song, Yuntao Gou, Ning Wang, Le Wang, Junbo Liu, Shengchun Wang, Chengliang Xia, Qiang Han and Zichen Gu
Sensors 2026, 26(16), 5292; https://doi.org/10.3390/s26165292 - 21 Aug 2026
Viewed by 135
Abstract
Rail light-strip morphology reflects the wheel-rail contact condition. Reliable automatic analysis remains difficult. The strip is narrow and has weak boundaries, while specular reflection, rail-head texture and trackside background interfere with color inspection images. This study proposes a segmentation-guided geometric method for rail [...] Read more.
Rail light-strip morphology reflects the wheel-rail contact condition. Reliable automatic analysis remains difficult. The strip is narrow and has weak boundaries, while specular reflection, rail-head texture and trackside background interfere with color inspection images. This study proposes a segmentation-guided geometric method for rail light-strip abnormality analysis. An improved SegFormer jointly segments the background, rail-head and light-strip regions. A boundary detail enhancement module refines weak rail-head and light-strip contours. Focal Loss emphasizes minority and hard boundary pixels. The rail-head mask provides the geometric reference for extracting the light-strip centerline, eccentricity, width sequence and connected-component morphology. The predicted masks are ordered using the corrected mileage record. Every 1000 original-resolution rows then form a consecutive 1 m detection unit. When a geometric rule is triggered, the method reports that unit’s 1 m mileage interval together with its eccentricity, width-change or local-integrity measurement. The model achieves 95.67% mean Intersection over Union (mIoU) on 3520 annotated images. It detects 845 of 876 positive units, with 96.46% recall, 89.23% precision and 92.70% F1-score. The resulting records identify abnormal 1 m mileage intervals and report the corresponding eccentricity, width-change, or local-integrity measurements for targeted manual review. Full article
Show Figures

Figure 1

34 pages, 4855 KB  
Article
PC-PLF: Path-Conditioned Per-Layer LoRA Fusion for Open-Vocabulary ROADWork Segmentation
by Ping Wu, Zhi-Ren Pan, Bo Qiu, Jian-Ping Wu and Shao-Jiang Zheng
Information 2026, 17(8), 805; https://doi.org/10.3390/info17080805 - 20 Aug 2026
Viewed by 179
Abstract
Construction work zones are a difficult case for open-vocabulary semantic segmentation. Their layouts are temporary, safety-relevant objects that are often small and long-tailed, and generic models readily confuse them with background. We address these failures inside an LoRA adapter space rather than retraining [...] Read more.
Construction work zones are a difficult case for open-vocabulary semantic segmentation. Their layouts are temporary, safety-relevant objects that are often small and long-tailed, and generic models readily confuse them with background. We address these failures inside an LoRA adapter space rather than retraining the backbone. Using only ROADWork training data, we audit a CAT-Seg RoadWork LoRA for false-positive- and recall-dominated cases and pair them with anchor images to train a residual adapter. Path-conditioned per-layer LoRA fusion (PC-PLF) then distributes a global correction budget across adapted layers using each layer’s first-order tangent magnitude along the stored factor path. Under group-disjoint out-of-fold evaluation on ROADWork, the complete method raises the mIoU from 61.72 for the RoadWork LoRA baseline to 62.29, with a shared-budget allocation gain of 0.33 mIoU over uniform fusion. Most of the total improvement appears before per-layer allocation. Uniform residual fusion contributes 0.69 points over the baseline, confirming that failure-driven residual training supplies the larger share; PC-PLF contributes a smaller allocation effect when tested on the same trained base-residual pair. The allocation effect is reproducible across four curation rules but near zero under SAN architecture transfer and MUSES second-target-domain evaluation. Three-group and text-weighted controls do not recover the full gain. Improvements concentrate in several long-tail safety classes. Cross-architecture, cross-dataset, and calibration audits define the operating regime rather than universal advantage. Residual curation supplies the larger share of the improvement; layer-wise allocation contributes a smaller, pair-specific gain. Full article
(This article belongs to the Topic Artificial Neural Networks for Visual Learning)
Show Figures

Figure 1

29 pages, 17754 KB  
Article
Structure-Prior-Guided Multi-Stage Cross-Modal Collaborative Network for RGB-D Semantic Segmentation
by Yifan Yu, Zhiwei Zhong, Fan Min and Song Deng
J. Imaging 2026, 12(8), 394; https://doi.org/10.3390/jimaging12080394 - 20 Aug 2026
Viewed by 167
Abstract
Red–green–blue and depth (RGB-D) semantic segmentation combines appearance cues from RGB images with geometric information from depth maps, but sensor noise, missing measurements, and boundary-inconsistent depth responses can introduce conflicting evidence during cross-modal fusion. We propose the Structure-Prior-Guided Network (SPGNet), a dual-branch, multi-stage [...] Read more.
Red–green–blue and depth (RGB-D) semantic segmentation combines appearance cues from RGB images with geometric information from depth maps, but sensor noise, missing measurements, and boundary-inconsistent depth responses can introduce conflicting evidence during cross-modal fusion. We propose the Structure-Prior-Guided Network (SPGNet), a dual-branch, multi-stage framework that follows a correction-before-fusion strategy. At each feature scale, SPGNet estimates a learned structure prior from cross-modal agreement and discrepancy. The Cross-Modal Correction Module (CCM) uses this prior to regulate bidirectional information transfer, suppressing unreliable responses while retaining complementary cues. The Dual-branch Enhancement Fusion Module (DEF) then enhances the corrected RGB and depth features and integrates them through shared-representation-guided interaction, after which a lightweight multi-scale decoder produces the segmentation output. Under a unified training and evaluation protocol, SPGNet achieved three-run mean Intersection over Union (mIoU) scores of 50.845% on NYU Depth V2 and 48.457% on SUN RGB-D. Compared with the best reproduced baseline on each dataset, SPGNet improved mean mIoU by 2.111 and 0.899 percentage points, respectively. These results suggest that separating reliability-oriented correction from multimodal fusion can limit the propagation of unreliable cross-modal responses and improve indoor RGB-D semantic segmentation performance. Full article
(This article belongs to the Section AI in Imaging)
Show Figures

Figure 1

37 pages, 5649 KB  
Article
AB-SAM: A SAM-Based Asymmetric Boundary-Aware Model for the Semantic Segmentation of Small and Medium-Sized Landslides
by Jiting Tang, Zhiwei Liang, Suli Guo, Bin Tong, Jun’an Chen, Guoliang Sun, Jiaxing Liu, Can Wang, Dong Li and Xin Zhou
Geomatics 2026, 6(4), 92; https://doi.org/10.3390/geomatics6040092 - 20 Aug 2026
Viewed by 87
Abstract
Small- and medium-sized landslides frequently occur in clusters and exhibit fragmented morphologies, irregular boundaries, and spectral characteristics similar to surrounding roads, bare soil, and sparsely vegetated surfaces, making their automated extraction from remote sensing imagery challenging. Although the Segment Anything Model (SAM) provides [...] Read more.
Small- and medium-sized landslides frequently occur in clusters and exhibit fragmented morphologies, irregular boundaries, and spectral characteristics similar to surrounding roads, bare soil, and sparsely vegetated surfaces, making their automated extraction from remote sensing imagery challenging. Although the Segment Anything Model (SAM) provides strong general-purpose segmentation capabilities, its direct application to landslide mapping is limited by the geoscience domain gap and its dependence on external prompts. This study proposes the Asymmetric Boundary-aware Segment Anything Model (AB-SAM), a parameter-efficient adaptation of SAM for automated landslide semantic segmentation. AB-SAM integrates three task-specific components. First, the offline Multi-Feature Variation-Guided Prompting (MF-VGP) module generates cached auxiliary bounding boxes from registered pre- and post-event images without accessing ground-truth masks. Second, the Asymmetric Feature Augmentation (AFA) strategy combines geometric perturbation, CutMix, and asymmetric dual-branch supervision, in which a Hint-free branch serves as the primary optimization pathway and a lower-weight box-guided branch provides auxiliary spatial supervision. Third, the Boundary-Aware Morphological Prompting (BAMP) module injects trainable boundary-aware morphological information into the largely frozen SAM image encoder. During validation, testing, and application, only the Hint-free branch is retained, enabling inference using post-event imagery without external point, box, or mask prompts. On the fixed, spatially disjoint Zixing test set, AB-SAM achieved an overall accuracy of 96.171%, a precision of 68.149%, a recall of 60.011%, an F1-score of 63.822%, a landslide-class Intersection over Union of 46.867%, and a mean Intersection over Union of 71.452%. Repeated experiments with three random seeds showed low run-to-run variation. Direct evaluation without retraining on the Hokkaido Iburi-Tobu dataset yielded a mean Intersection over Union of 66.136%, providing evidence of cross-region and cross-event transferability. These results demonstrate that AB-SAM provides a practical parameter-efficient framework for automated, hint-free landslide segmentation, although further evaluation across additional regions, sensors, and landslide-size distributions remains necessary. Full article
Show Figures

Figure 1

25 pages, 28843 KB  
Article
UNet-DFH: A Semantic Segmentation Network Combining Multi-Scale Edge Fusion and Attention-Deformable Modules for Sugarcane Mapping in Heterogeneous Karst Regions
by Yanling Lu, Jinshuang Liu, Jingwen Li, Li Zhang and Jizheng Wan
Remote Sens. 2026, 18(16), 2815; https://doi.org/10.3390/rs18162815 - 20 Aug 2026
Viewed by 188
Abstract
In karst regions, sugarcane mapping faces challenges from fragmented fields, undulating terrain, spectral confusion, and persistent cloud cover, which limit traditional optical remote sensing. To address these issues, we propose a fine-scale extraction framework that integrates Sentinel-2 optical and Sentinel-1 synthetic aperture radar [...] Read more.
In karst regions, sugarcane mapping faces challenges from fragmented fields, undulating terrain, spectral confusion, and persistent cloud cover, which limit traditional optical remote sensing. To address these issues, we propose a fine-scale extraction framework that integrates Sentinel-2 optical and Sentinel-1 synthetic aperture radar (SAR) imagery through image-level fusion, and introduces a UNet-DFH network with a Multi-Scale Edge Fusion (MSEF) module and an Attention-Deformable Fusion Module (ADFM). This study makes three core contributions: (1) we construct a dedicated optical–SAR collaborative sugarcane extraction dataset for typical karst regions, alleviating the scarcity of multimodal labeled samples; (2) we propose the UNet-DFH network, where MSEF enhances boundary preservation and topological detail in shallow decoding stages, while ADFM improves robustness to geometric deformation and local misalignment in deep semantic stages; (3) we demonstrate that the joint mechanism of edge-preserving filtering and deformable adaptation yields a synergistic effect in addressing the precision–recall trade-off. Experiments in a typical karst area of Guangxi, China, demonstrate that optical–SAR fusion achieves an IoU of 80.08% and an OA of 92.09% during the sugar accumulation and maturity stage. During the more challenging tillering stage, UNet-DFH maintains relatively stable performance under optical-only conditions, with an IoU of 72.98%, Recall of 82.78%, and OA of 92.12%. Moreover, optical–SAR fusion improves Recall by 5.5 percentage points over optical-only inputs (from 83.54% to 89.04%), while Precision exhibits a moderate decrease from 91.89% to 88.84%, reflecting the expected trade-off associated with speckle noise. These results confirm the complementary value of multimodal data and the effectiveness of the proposed modules in preserving fragmented plot boundaries and improving segmentation performance in complex karst terrain. The framework offers a promising approach for high-precision crop mapping in the studied karst agricultural landscape. Full article
Show Figures

Figure 1

18 pages, 3692 KB  
Article
Semantic Segmentation by Semantic Proportions
by Halil Ibrahim Aysel, Xiaohao Cai and Adam Prugel-Bennett
Sensors 2026, 26(16), 5262; https://doi.org/10.3390/s26165262 - 19 Aug 2026
Viewed by 207
Abstract
Semantic segmentation is a critical task in computer vision aiming to identify and classify individual pixels in an image, with numerous applications, for example, in autonomous driving and medical image analysis. However, semantic segmentation can be highly challenging, particularly due to the need [...] Read more.
Semantic segmentation is a critical task in computer vision aiming to identify and classify individual pixels in an image, with numerous applications, for example, in autonomous driving and medical image analysis. However, semantic segmentation can be highly challenging, particularly due to the need for large amounts of annotated data. Annotating images is a time-consuming and costly process, often requiring expert knowledge and significant effort; moreover, saving the annotated images could dramatically increase the storage space. In this paper, we propose a novel approach for semantic segmentation, requiring only rough information about the proportions of individual semantic classes, hereafter referred to as semantic proportions (SPs), rather than the necessity of ground-truth segmentation maps. This greatly simplifies the data annotation process and thus will significantly reduce the annotation time, cost and storage space, opening up new possibilities for semantic segmentation tasks where obtaining the full ground-truth segmentation maps may not be feasible or practical. Our proposed method of utilising semantic proportions can (i) further be utilised as a booster in the presence of ground-truth segmentation maps to gain performance without extra data and model complexity, and (ii) also be seen as a parameter-free plug-and-play module, which can be attached to existing deep neural networks designed for semantic segmentation. Extensive experimental results demonstrate the good performance of our method compared to benchmark methods that rely on ground-truth segmentation maps. Utilising semantic proportions suggested in this work offers a promising direction for future semantic segmentation research. Full article
Show Figures

Figure 1

26 pages, 7007 KB  
Article
OVR-GS: Open-Vocabulary 3D Object Removal via Semantic Gaussian Selection and Local Diffusion-Guided Completion
by Yongpeng Ding, Feng Ouyang, Jiawei Fan, Ting Chen and Hongyan Xu
Sensors 2026, 26(16), 5258; https://doi.org/10.3390/s26165258 - 19 Aug 2026
Viewed by 199
Abstract
Camera-reconstructed 3D scenes often require offline visual cleanup before inspection, presentation, or reuse as renderable virtual-scene assets. Representative applications include removing temporary furniture, parked vehicles, equipment, signage, and other distracting or obsolete objects from reconstructed indoor and outdoor environments. Such editing requires not [...] Read more.
Camera-reconstructed 3D scenes often require offline visual cleanup before inspection, presentation, or reuse as renderable virtual-scene assets. Representative applications include removing temporary furniture, parked vehicles, equipment, signage, and other distracting or obsolete objects from reconstructed indoor and outdoor environments. Such editing requires not only accurate target localization across viewpoints but also plausible recovery of the previously occluded background. Existing methods often depend on manually specified masks or category-restricted detectors, while projection-based pipelines independently inpaint multiple views and subsequently refine the 3D representation, potentially introducing cross-view appearance and geometry inconsistencies. We present OVR-GS (Open-Vocabulary Removal in Gaussian Splatting), an instruction-driven object-removal framework for pre-trained 3D Gaussian Splatting (3DGS) scenes. Given a free-form instruction, a language parser generates target-oriented queries and a textual background-completion condition. Grounding DINO and the Segment Anything Model (SAM) produce multi-view candidate masks, which are filtered using Contrastive Language–Image Pre-training (CLIP). The proposed Semantic-Aware Gaussian Selector (SAGS) aggregates rendering-contribution-weighted mask evidence, groups spatially coherent candidates, and identifies the target Gaussian subset through rendered-cluster semantic verification. After removal, new Gaussians are initialized from boundary-adjacent primitives and interior samples and optimized locally using Score Distillation Sampling (SDS), while the original background remains fixed. On IMFine, SPIn-NeRF, and Inpaint360GS, OVR-GS achieves peak signal-to-noise ratio (PSNR) values of 19.78, 17.82, and 24.62 dB and Fréchet inception distance (FID) values of 142.30, 148.60, and 34.80, respectively. The results demonstrate the effectiveness of localized Gaussian optimization for instruction-driven cleanup of reconstructed environments before visual inspection, presentation, or reuse as renderable virtual-scene assets. Full article
(This article belongs to the Section Optical Sensors)
Show Figures

Figure 1

23 pages, 8737 KB  
Article
AgriUFM: Unconditional-Flow-Matching-Based Generative Model for Creating Image–Mask Pairs of Agricultural Pests and Disease
by Haocheng Kong, Lei Liu, Haotian Bai, Xiaoyu Li and Yuefeng Du
Agriculture 2026, 16(16), 1777; https://doi.org/10.3390/agriculture16161777 - 19 Aug 2026
Viewed by 219
Abstract
Pests and diseases are key biological stress factors affecting crop yield and quality. Semantic segmentation enables pixel-level localization and severity characterization, but its performance and generalization are constrained by the high cost of high-quality pixel-level annotations, limited labeled samples, and class imbalance in [...] Read more.
Pests and diseases are key biological stress factors affecting crop yield and quality. Semantic segmentation enables pixel-level localization and severity characterization, but its performance and generalization are constrained by the high cost of high-quality pixel-level annotations, limited labeled samples, and class imbalance in agricultural datasets. We propose AgriUFM, an unconditional flow-matching framework for joint image–mask generation in agricultural pest and disease scenarios. By learning a unified continuous probability flow over the joint distribution, the framework is designed to promote structural co-evolution and spatial consistency between generated RGB images and masks. Across four evaluated datasets, AgriUFM achieved lower FID and rFID than the evaluated GAN- and diffusion-based comparators, whereas IS performance was dataset-dependent. Within the evaluated ablation configurations, uniform time sampling with 25 sampling steps and the midpoint ODE solver yielded the most favourable observed quality–efficiency trade-off. Under the held-out test protocol, the joint UFM strategy achieved higher image–mask correspondence than the M2I and I2M conditional variants. In the evaluated downstream settings, AgriUFM-generated augmentation improved MIoU and PA for U-Net and TransUNet. These results indicate that joint distribution modelling is a promising approach for structurally coherent generative augmentation in the agricultural imaging tasks studied. Full article
Show Figures

Figure 1

32 pages, 7877 KB  
Article
DFSA: Dynamic-Feature Collaborative Optimization and Semantic-Alignment Network for UAV Cross-View Geo-Localization
by Xiaojia Yan, Zhangsong Shi, Shiyan Sun, Huihui Xu, Huimin Zhu, Qingping Hu, Weiming Zhu and Yinglei Li
Drones 2026, 10(8), 632; https://doi.org/10.3390/drones10080632 - 19 Aug 2026
Viewed by 216
Abstract
Cross-view geo-localization (CVGL) is a critical technology used in unmanned aerial vehicles (UAVs) and widely applied in navigation and target localization tasks. However, owing to the extreme perspective disparity between UAV oblique views and satellite vertical views, CVGL still involves significant challenges, including [...] Read more.
Cross-view geo-localization (CVGL) is a critical technology used in unmanned aerial vehicles (UAVs) and widely applied in navigation and target localization tasks. However, owing to the extreme perspective disparity between UAV oblique views and satellite vertical views, CVGL still involves significant challenges, including geometric distortion caused by viewpoint differences, drastic appearance inconsistencies, and the difficulty in bridging semantic gaps between heterogeneous data. To address these issues, we propose a novel CVGL method named dynamic-feature collaborative optimization and semantic-alignment network (DFSA), designed to extract robust feature representations and achieve fine-grained alignment. Specifically, the DFSA employs a residual-based vision transformer as the backbone to capture global context while alleviating the training instability and feature collapse often associated with standard transformers. To bridge the semantic gap between global and local features, we design a feature optimization module comprising a local feature enhancer and a global feature aggregator. This module establishes a closed-loop collaborative system that facilitates top-down semantic guidance and bottom-up detail feedback. Furthermore, we introduce a semantic segmentation and alignment module that adaptively partitions images into semantic regions based on feature response distributions, shifting the matching granularity from the global level to the semantic region level to effectively overcome feature mismatches caused by positional offsets and scale variations. Extensive experiments conducted on the University-1652 and SUES-200 datasets demonstrate the superior image retrieval performance of the proposed DFSA. Specifically, DFSA achieves a Recall@1 of 94.87% and an Average Precision (AP) of 95.32% on the University-1652 dataset and maintains highly competitive Recall@1 performances between 96.83% and 99.25% across various altitudes on the SUES-200 dataset. These results validate the model’s effectiveness in handling extreme viewpoint changes for UAV-based cross-view image retrieval tasks. Full article
Show Figures

Figure 1

25 pages, 11254 KB  
Article
MemGeoSeg: Location-Aware Semantic Segmentation with Spatially Indexed Memory for Repetitive Driving Scenarios
by Huei-Yung Lin and Jou-An Tsai
Smart Cities 2026, 9(8), 134; https://doi.org/10.3390/smartcities9080134 - 19 Aug 2026
Viewed by 195
Abstract
Current visual perception techniques for self-driving vehicles mainly focus on the generalization across diverse scenes, and they often overlook the valuable spatial consistency present in the repetitive driving routes such as public transit lines, delivery and shuttle services. In this paper, we introduce [...] Read more.
Current visual perception techniques for self-driving vehicles mainly focus on the generalization across diverse scenes, and they often overlook the valuable spatial consistency present in the repetitive driving routes such as public transit lines, delivery and shuttle services. In this paper, we introduce MemGeoSeg, which is a novel multi-modal framework that enhances semantic segmentation by exploiting scene repetitions through GPS-guided spatial priors and historical memory. Our approach introduces a hierarchical GPS embedding module, which is a spatially indexed memory bank that accumulates location-specific visual knowledge and a cross-modal fusion mechanism with contrastive learning. To validate the idea of improving visual perception with repetitive driving scenarios, a new dataset, RMTD-AD, is constructed for evaluation. It contains over 13,000 annotated images across various weather and lighting conditions on repeated routes. Extensive experiments conducted on the dataset have demonstrated that MemGeoSeg significantly outperforms the state-of-the-art baseline, achieving an mIoU of 76.5% compared to SegFormer’s 71.8% (a 4.7 percentage-point improvement), with particularly strong gains in challenging scenarios like low-light and adverse weather conditions. The result shows that there are substantial benefits to incorporating geographical contexts and historical memory for location-aware perception in intelligent vehicles. Full article
Show Figures

Figure 1

18 pages, 5607 KB  
Article
Improving Cross-Organ Generalization in Histopathology Segmentation via Evidence-Guided Vision–Language Query Decoding
by Biwen Meng, Jiahao Wang and Jingxin Liu
Electronics 2026, 15(16), 3691; https://doi.org/10.3390/electronics15163691 - 18 Aug 2026
Viewed by 112
Abstract
Domain shift remains a major obstacle to robust histopathology image segmentation, especially when models trained on several source organs are deployed to unseen anatomical sites. This study addresses cross-organ adenocarcinoma segmentation by introducing an evidence-guided vision–language segmentation framework that incorporates pathology-relevant morphological evidence [...] Read more.
Domain shift remains a major obstacle to robust histopathology image segmentation, especially when models trained on several source organs are deployed to unseen anatomical sites. This study addresses cross-organ adenocarcinoma segmentation by introducing an evidence-guided vision–language segmentation framework that incorporates pathology-relevant morphological evidence into dense mask prediction. The proposed method uses a pathology vision–language encoder to extract image and text representations, a Semantic Query Booster to form image-aware segmentation queries, and an evidence-guided query recalibration that integrates positive tumor-supporting evidence and negative misleading evidence. Experiments were conducted on cross-organ adenocarcinoma datasets from the COSAS challenge under a source-only domain generalization setting, with colorectum, stomach, and pancreas as source domains and ampullary, gallbladder, and intestine as unseen target domains. The proposed framework achieved the highest pooled performance on the seen, unseen, and overall evaluation sets among the compared segmentation, domain generalization, and foundation model-based systems. These findings support the use of structured pathology evidence for cross-organ tumor segmentation under source-only training. Full article
(This article belongs to the Special Issue Advances in Real-Time Image Processing)
Show Figures

Figure 1

21 pages, 4909 KB  
Article
Sequential 2D–3D Recognition for Privacy-Sensitive Object Extraction from 3D Point Clouds
by Yusuke Shinwashi, Etsuji Kitagawa, Satoshi Abiko, Kennosuke Takada and Ryo Kato
Big Data Cogn. Comput. 2026, 10(8), 277; https://doi.org/10.3390/bdcc10080277 - 18 Aug 2026
Viewed by 101
Abstract
With the growing use of digital twins and 3D city models, 3D point cloud data have become increasingly important. Such data, however, may contain privacy-sensitive objects, including people and vehicles, which poses challenges for public release and secondary use. This study proposes a [...] Read more.
With the growing use of digital twins and 3D city models, 3D point cloud data have become increasingly important. Such data, however, may contain privacy-sensitive objects, including people and vehicles, which poses challenges for public release and secondary use. This study proposes a sequential 2D–3D recognition framework for extracting privacy-sensitive objects by integrating 2D image recognition and 3D point cloud recognition. The proposed framework first detects candidate regions in images and associates them with the corresponding 3D point cloud through multi-view projection, after which 3D semantic segmentation is applied only to the candidate point cloud. By restricting 3D recognition to candidate regions, the proposed method suppresses background-point contamination while reducing unnecessary 3D processing. We evaluate the method using SfM-derived 3D point clouds containing people and vehicles. The results show that the proposed method achieves higher F-scores than the selected direct 2D-projection and 3D-only baselines, reflecting a better balance between precision and recall. These findings suggest that sequentially combining 2D image recognition with 3D point cloud recognition provides an effective approach for privacy-sensitive object extraction and supports the privacy-preserving publication and secondary use of digital twins and 3D city models. Full article
Show Figures

Figure 1

Back to TopTop