Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (1,705)

Search Parameters:
Keywords = high-level semantics

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
31 pages, 6298 KB  
Article
HDSMNet: Height-Guided Sparse Cross-Modal Fusion for High-Resolution Remote Sensing Semantic Segmentation
by Hanxu Gu, Jian Hu, Li Wang, Jianwen Wang, Yujie Wang, Yapeng Zhou and Nan Wang
Remote Sens. 2026, 18(17), 2992; https://doi.org/10.3390/rs18172992 - 3 Sep 2026
Abstract
High-resolution remote sensing semantic segmentation requires the joint modeling of local details, global semantics, and height-derived geometric structures, and it provides an important basis for urban object mapping, land-cover analysis, and fine-grained spatial understanding. However, in complex urban scenes, fine-grained boundaries, small objects, [...] Read more.
High-resolution remote sensing semantic segmentation requires the joint modeling of local details, global semantics, and height-derived geometric structures, and it provides an important basis for urban object mapping, land-cover analysis, and fine-grained spatial understanding. However, in complex urban scenes, fine-grained boundaries, small objects, inter-class similarity, and spectral confusion can still weaken the stability of pixel-level prediction. To enhance discriminative dense feature representations in high-resolution remote sensing images, we propose HDSMNet, a dual-branch multimodal semantic segmentation network designed for optical–nDSM data. The network separately extracts appearance and semantic features from optical imagery and height–structural features from nDSM, and introduces a Height-Guided Sparse Cross-Modal Fusion (HGSCF) module. Rather than treating nDSM as an additional feature source for generic fusion, HGSCF derives contextual representations, local feature contrasts, and structural-discontinuity cues from encoded nDSM features and uses them to guide sparse anchor-based interaction between optical and height features. This design enhances discriminative dense feature representations through interaction with a compact set of geometry-guided anchors. To complement HGSCF at the output stage, HDSMNet further adapts a Context-Guided Refinement (CGR) path that combines intermediate-response-guided contextual aggregation with dynamic feature modulation. This supplementary path recalibrates decoder features for output refinement. Experiments on the ISPRS Potsdam and Vaihingen datasets show that HDSMNet achieves mIoU values of 86.57% and 84.22%, respectively; ablation results further identify HGSCF as the main contributor to the observed improvement. Full article
Show Figures

Figure 1

27 pages, 7364 KB  
Article
Aerial Scene Classification Using Hierarchical and Multi-Stage Swin Transformer Features
by Elif Kanca Gulsoy, Selen Ayas, Tolgahan Gulsoy, Elif Baykal Kablan, Esra Tunc Gormus and Alin Achim
Remote Sens. 2026, 18(17), 2987; https://doi.org/10.3390/rs18172987 - 3 Sep 2026
Abstract
Aerial image classification remains challenging due to complex spatial structures, high intra-class variability, and the need to capture semantic information across multiple spatial scales. This study investigates the potential of the Swin Transformer architecture to address the limitations of conventional deep learning approaches, [...] Read more.
Aerial image classification remains challenging due to complex spatial structures, high intra-class variability, and the need to capture semantic information across multiple spatial scales. This study investigates the potential of the Swin Transformer architecture to address the limitations of conventional deep learning approaches, particularly their restricted ability to model long-range contextual dependencies in high-resolution aerial images. The analysis focuses on the contribution of hierarchical feature representations extracted from different stages of the Swin Transformer to classification performance. Experiments were conducted on the AID, UCM21, and NWPU-RESISC45 benchmark datasets using multiple model variants and varying training proportions. The results show that combining intermediate and high-level features with the classification head achieves competitive and generally improved classification performance compared to other configurations, emphasizing the importance of multi-scale semantic feature representation. Additional analyses reveal that input image size and data augmentation strategies influence performance, while attention maps are used as an auxiliary visualization tool to support qualitative interpretation by highlighting relevant spatial regions. Overall, the findings indicate that hierarchical attention mechanisms and multi-stage feature integration can provide an effective representation strategy for Transformer-based aerial image classification. Full article
Show Figures

Figure 1

36 pages, 12571 KB  
Article
Boot-Shaped Terrain Screening and Deep Learning Semantic Segmentation for Landslide-Hazard Candidate Extraction from Airborne LiDAR DEM: A Case Study in Zhenxiong County, China
by Bowen Du, Xiangcong Meng, Junchen Ye, Bin Tong and Yueping Yin
Remote Sens. 2026, 18(17), 2983; https://doi.org/10.3390/rs18172983 - 3 Sep 2026
Abstract
Automated screening of geomorphologically defined landslide potential-hazard candidates from high-resolution topographic data remains challenging in mountainous regions where comprehensive field inventories are unavailable. This study proposes a two-stage framework for extracting rule-defined boot-shaped terrain candidates from airborne LiDAR digital elevation model (DEM) data. [...] Read more.
Automated screening of geomorphologically defined landslide potential-hazard candidates from high-resolution topographic data remains challenging in mountainous regions where comprehensive field inventories are unavailable. This study proposes a two-stage framework for extracting rule-defined boot-shaped terrain candidates from airborne LiDAR digital elevation model (DEM) data. First, an expert-informed screening rule formalizes a steep-upper–gentle-lower terrain morphology using representative longitudinal profiles of slope units, and the screened units are converted into rule-derived reference masks. Second, semantic segmentation models are trained to approximate these reference patterns directly from DEM-derived raster inputs. Four architectures—U-Net, U-Net++, DeepLabV3+, and SegFormer-B0—were evaluated using 406 patches of 256 × 256 pixels at 2 m resolution from four LiDAR-covered subregions in Zhenxiong County, China. Under spatially grouped three-fold cross-validation, DeepLabV3+ with DEM + slope-gradient input and Dice + Focal loss achieved a mean pixel-level F1-score of 0.351, mIoU of 0.557, and Patch-F1 of 0.814. Input-feature experiments showed that slope-gradient information was particularly informative, whereas the incremental contribution of aspect was configuration-dependent; DEM + slope was retained as a parsimonious two-channel input. Sensitivity analysis showed that the rule-derived candidate definition changed materially with the screening parameters. Leave-one-subregion-out evaluation yielded a macro-averaged F1 of 0.321, indicating measurable within-county cross-subregion transfer. However, whole-area evaluation under natural candidate prevalence reduced the macro-average F1 to 0.072 at a fixed threshold and 0.095 using validation-derived operating thresholds. These results indicate that the proposed model is best interpreted as a raster-based surrogate for rule-derived geomorphological screening rather than as an independently validated landslide detector. Full article
Show Figures

Figure 1

32 pages, 32908 KB  
Article
Spatial Prompt and Wavelet Mamba-Based Multi-Scale Cross-Domain Feature Fusion Network for Segmentation of Mining-Disturbed Land
by Jianing Song, Jiangyuan Wang, Yanyan Qin, Zhe Liu, Jian Feng and Xianju Li
Remote Sens. 2026, 18(17), 2963; https://doi.org/10.3390/rs18172963 - 2 Sep 2026
Abstract
Segmentation of mining-disturbed land is important for eco-geological environment monitoring. Although existing methods possess strong segmentation capabilities, mining-disturbed land exhibits irregular edges, different spatial sizes, and global texture variability, which lead to difficulties in extracting discriminative features, thereby limiting segmentation accuracy. This study [...] Read more.
Segmentation of mining-disturbed land is important for eco-geological environment monitoring. Although existing methods possess strong segmentation capabilities, mining-disturbed land exhibits irregular edges, different spatial sizes, and global texture variability, which lead to difficulties in extracting discriminative features, thereby limiting segmentation accuracy. This study built an RGB-based binary semantic segmentation dataset, covering typical mining-disturbed land in Fujian Province of China. Then a spatial prompt and wavelet Mamba-based multi-scale cross-domain feature fusion network (SWDF-Net) was proposed. (1) Wavelet Mamba-based dual-frequency collaborative enhancement module: high-low frequency information was decoupled by wavelet transform, and collaboratively enhanced by fusion of multiple local details and Mamba-guided global information. It can highlight the high-frequency edge features and low-frequency texture patterns of mining-disturbed land. (2) Cross-domain feature alignment and fusion module: cross-domain statistical calibration and channel conditional modulation were used to narrow spatial–frequency feature distribution gap, which facilitates cross-domain feature alignment and fusion. (3) Spatial prompt-based multi-scale feature weighted fusion module: pixel-level weight maps were generated by spatial prior prompt derived from a spatial decoder and edge-gated branch, which adaptively fuses former multi-scale dual-domain features. The SWDF-Net achieved the best Intersection over Union of 72.22% for mining-disturbed land and performed competitively on the ISPRS Vaihingen and Potsdam datasets. Full article
(This article belongs to the Special Issue Deep Learning for Remote Sensing Image Segmentation)
Show Figures

Figure 1

29 pages, 19353 KB  
Article
Complex-Valued HRU-Net with Cross-Gated Attention for PolSAR Semantic Segmentation
by Xiaochun Xie, Pin Xin, Lingjuan Yu, Miaomiao Liang, Yuting Guo and Xuan Jiao
Remote Sens. 2026, 18(17), 2947; https://doi.org/10.3390/rs18172947 - 2 Sep 2026
Abstract
In recent years, U-Net-based architectures have been widely applied to polarimetric synthetic aperture radar (PolSAR) semantic segmentation. However, successive downsampling may lead to the loss of fine spatial details, while conventional U-Net-style decoder directly concatenates encoder features with the corresponding decoder features without [...] Read more.
In recent years, U-Net-based architectures have been widely applied to polarimetric synthetic aperture radar (PolSAR) semantic segmentation. However, successive downsampling may lead to the loss of fine spatial details, while conventional U-Net-style decoder directly concatenates encoder features with the corresponding decoder features without explicitly accounting for their semantic discrepancy, potentially introducing redundant or irrelevant information and weakening feature discrimination. To address these limitations, this paper proposes a lightweight complex-valued high-resolution U-Net (CV-HRU-Net) with a complex-valued cross-gated attention (CV-CGA) module for PolSAR semantic segmentation. CV-HRU-Net employs a complex-valued high-resolution network (CV-HRNet) as the encoder to maintain high-resolution representations through parallel multi-resolution streams, while a complex-valued U-Net (CV-U-Net) decoder progressively incorporates multi-resolution high-level semantic features for pixel-wise prediction. To improve encoder–decoder feature interaction, CV-CGA adaptively calibrates the core decoder features using encoder information. Specifically, CV-CGA integrates the Convolutional Block Attention Module, Transformer-style cross-attention with decoder features as queries and encoder features as keys and values, and adaptive gated recalibration to enhance semantic selectivity and boundary representation. Experiments on two airborne and two spaceborne PolSAR datasets demonstrate that the proposed network achieves accurate land-cover segmentation and precise boundary delineation by jointly exploiting polarimetric phase relationships, fine spatial details, and multi-resolution semantic information. Furthermore, CV-CGA substantially improves segmentation accuracy and boundary F1 scores while introducing only marginal model-size overhead. Full article
(This article belongs to the Section Engineering Remote Sensing)
Show Figures

Figure 1

20 pages, 3447 KB  
Article
A Linear Attention Framework with Dual-Axis Multi-Scale Fusion for Fine-Grained Eucalyptus Change Detection
by Guangjin Li, Liyang You, Jirong Ding, Xu Tang, Jianjun Chen, Haoyu Wang and Haotian You
Remote Sens. 2026, 18(17), 2944; https://doi.org/10.3390/rs18172944 - 1 Sep 2026
Viewed by 133
Abstract
The fine-scale monitoring of plantation cover disappearance and appearance is challenging because these changes are often expressed as weak within-class variations in high-resolution images. This study proposes MLLAForestCD, a three-class pixel-level semantic change-detection network for Eucalyptus plantations. The model uses an MLLA encoder [...] Read more.
The fine-scale monitoring of plantation cover disappearance and appearance is challenging because these changes are often expressed as weak within-class variations in high-resolution images. This study proposes MLLAForestCD, a three-class pixel-level semantic change-detection network for Eucalyptus plantations. The model uses an MLLA encoder to model the long-range spatial context with efficient linear attention, while a dual-axis change extractor reorganizes paired bi-temporal features through complementary layouts before contextual interaction. Multi-scale fusion then combines semantic cues with boundary-level details. We further construct the Eucalyptus Change Detection Dataset (ECDD), which contains plantation scenes with weak spectral contrast, fragmented boundaries, and directional canopy textures. Under the retained patch-level training/validation split, MLLAForestCD achieves an F1-score of 96.66% and an mIoU of 93.59%. After separate training and evaluation based on WHU-CD, it achieves an F1-score of 97.29% and an IoU of 90.10%; this result reflects performance under an independent WHU-CD training protocol. Finally, annual change maps from 2020 to 2023 are used to derive the most recently detected plantation-appearance time within the observation window. The resulting product is interpreted as a recent stand-renewal event map and requires independent forestry records before biological stand age can be inferred. Full article
(This article belongs to the Section Forest Remote Sensing)
Show Figures

Figure 1

15 pages, 7594 KB  
Article
Speaker-Mediated Route Knowledge Transfer for Goal-Oriented Vision-and-Language Navigation
by Ju Han, Xiaoyan Li, Boyue Wang, Yongli Hu and Baocai Yin
Electronics 2026, 15(17), 3937; https://doi.org/10.3390/electronics15173937 - 1 Sep 2026
Viewed by 120
Abstract
Goal-oriented vision-and-language navigation (VLN) requires an agent to reach a target from a high-level instruction that does not prescribe a route. Training therefore faces two problems. First, terminal and distance-based feedback leave partial trajectories without route-semantic supervision. Second, a single annotated path introduces [...] Read more.
Goal-oriented vision-and-language navigation (VLN) requires an agent to reach a target from a high-level instruction that does not prescribe a route. Training therefore faces two problems. First, terminal and distance-based feedback leave partial trajectories without route-semantic supervision. Second, a single annotated path introduces single-reference route bias because several routes may satisfy the same goal. We propose Speaker-Mediated Route Knowledge Transfer (SMRKT), which comprises three components. To provide the route knowledge required by both problems, we trained a trajectory-to-language speaker on route-following VLN data and froze it as a training-time evaluator in the cross-granularity route-prior transfer process. To address the lack of route-semantic supervision, Speaker Reconstruction Progress Reward converts consecutive reconstruction-loss changes into intermediate feedback. To address single-reference route bias, Disagreement-Aware Trajectory Reuse interprets the speaker score with terminal outcome, retaining successful low-score paths as alternative demonstrations and failed high-score paths as hard replay cases. The navigator retains the original goal-oriented instruction, and the speaker is absent at inference. In DUET-based validation on two goal-oriented VLN datasets, SMRKT improves Success Rate (SR) by 1.42 percentage points on the unseen validation setting of the REVERIE dataset, by 1.32 percentage points on the Unseen Houses split of the SOON dataset, and by 2.22 percentage points on the Unseen Instructions split of the SOON dataset; SPL and OSR rise in some settings and fall in others. The results support transferred route knowledge as a training signal, with environmental success retained as the final task criterion. Full article
Show Figures

Figure 1

18 pages, 6164 KB  
Article
An Optimized YOLO11n Model with Multi-Stage Attention Mechanisms for Pitaya Disease Detection
by Zhi Qiu, Zhaomin Shi, Yuechao Sun, Deyun Mo, Xingzao Ma, Guangbin Wang and Caimao Su
Horticulturae 2026, 12(9), 1089; https://doi.org/10.3390/horticulturae12091089 - 1 Sep 2026
Viewed by 143
Abstract
At present, the diagnosis of pitaya diseases is chiefly dependent on human expertise, a process that is plagued by issues such as low identification efficiency and high subjectivity. In order to achieve rapid and accurate identification of diseases that are prevalent in pitaya, [...] Read more.
At present, the diagnosis of pitaya diseases is chiefly dependent on human expertise, a process that is plagued by issues such as low identification efficiency and high subjectivity. In order to achieve rapid and accurate identification of diseases that are prevalent in pitaya, this paper focuses on such diseases and proposes a machine-vision detection method based on an improved YOLO11n algorithm. The enhanced architecture incorporates MSCAM for edge-feature enhancement, ACmix to balance local and global perception, and PMHSA to optimize deep semantic flow, collectively addressing the challenges of small-target loss and false positives in pitaya disease detection. The experimental findings demonstrate that the improved model attains a detection accuracy of 95.8%, a recall rate of 92.3%, and an mAP50 of 95.9%. A comparison of the baseline model with the improved model reveals a reduction in the number of parameters representing a decrease of approximately 14.2%. Furthermore, the computational complexity indicates a decline of approximately 3.07%, and a reduction in model size by approximately 4.73%. This approach led to enhanced disease-recognition performance while preserving high levels of accuracy. This study proposes a technically sound solution for the intelligent development of disease-detection models for pitaya. Full article
Show Figures

Figure 1

22 pages, 1769 KB  
Article
Reliable or Just Accurate? A Cross-Dataset Audit of Early-Warning Models Under Course-Level Distribution Shift
by Athanasios Angeioplastis, Markos Konstantakis and Alkiviadis Tsimpiris
Computers 2026, 15(9), 572; https://doi.org/10.3390/computers15090572 - 1 Sep 2026
Viewed by 139
Abstract
Early-warning systems based on Learning Management System (LMS) traces are commonly evaluated using random or unseen-student splits, even when deployment requires predictions in courses that were absent during model development. In this study, we distinguish predictive accuracy within familiar course environments from reliability [...] Read more.
Early-warning systems based on Learning Management System (LMS) traces are commonly evaluated using random or unseen-student splits, even when deployment requires predictions in courses that were absent during model development. In this study, we distinguish predictive accuracy within familiar course environments from reliability under course-level distribution shift. We evaluate an institutional Moodle deployment from the International Hellenic University (IHU; 1284 student-course observations, 924 students, 35 courses) alongside two public behavioral benchmarks, OULAD (N = 22,437) and Riestra (N = 25,260). Six semantically harmonized behavioral features were measured at 10%, 25%, 33%, and 50% observation cutoffs and evaluated under random-row, unseen-student, and unseen-course validation. With Extra Trees, unseen-course balanced accuracy was stable and above chance for OULAD (0.619–0.686) and Riestra (0.584–0.639), whereas IHU performance ranged from 0.470 to 0.582 and was statistically indistinguishable from chance at the earliest cutoff. The early IHU result persisted across Logistic Regression, Random Forest, Extra Trees, and Histogram Gradient Boosting. A grade-scale audit identified substantial inconsistencies between configured and observed course maxima; an observed-maximum sensitivity specification materially changed class prevalence and prevalence-sensitive metrics but did not significantly improve unseen-course balanced accuracy or ROC-AUC. Repeated size-matched controls attenuated the Riestra–IHU difference but left the OULAD advantage largely unchanged. OULAD leave-one-presentation-out and IHU leave-one-mixed-course-out analyses were consistent with grouped five-fold validation. The findings show that high within-sample accuracy is not sufficient evidence of operational reliability. Reported performance is associated with validation design, label definition, course composition, and dataset scale. We recommend unseen-course evaluation, label-definition audits, course-aware uncertainty estimates, and multi-model robustness checks as standard reporting practices for LMS early-warning systems intended for use beyond their original training courses. Full article
Show Figures

Figure 1

24 pages, 7734 KB  
Article
EGDNet: An Event-Guided Frequency-Aware Network with Cross-Modal Attention for Robust Weak Signature UAV Detection
by Ziming Tang, Zhiming Liu, Peilun Sun and Jie Deng
Remote Sens. 2026, 18(17), 2925; https://doi.org/10.3390/rs18172925 - 1 Sep 2026
Viewed by 147
Abstract
With the rapid expansion of the low-altitude economy, detecting micro unmanned aerial vehicles (UAVs) in complex urban environments faces severe challenges. In optical surveillance, micro-UAVs typically appear as weak and small targets lacking distinct spatial semantics, rendering conventional appearance-based feature extraction highly ineffective. [...] Read more.
With the rapid expansion of the low-altitude economy, detecting micro unmanned aerial vehicles (UAVs) in complex urban environments faces severe challenges. In optical surveillance, micro-UAVs typically appear as weak and small targets lacking distinct spatial semantics, rendering conventional appearance-based feature extraction highly ineffective. Furthermore, extreme lighting conditions and complex backgrounds easily submerge these subtle textures into noise. While event cameras can uniquely capture the high-frequency rotational characteristics of UAV rotors thanks to their microsecond temporal resolution, their sparse asynchronous outputs are highly susceptible to environmental clutter. To address this, we propose EGDNet, an event-guided frequency-aware dual-modal detection network optimized for robust UAV sensing. First, we developed a custom co-axial beam-splitting hardware platform to construct a strictly pixel-aligned dual-modal dataset of 43,090 frames, eliminating inter-modal parallax at the physical level. Second, our framework introduces a novel frequency-domain deconstruction pipeline utilizing a Temporal Fast Fourier Transform (T-FFT) to isolate the high-frequency mechanical rotor vibrations from low-frequency background clutter. Additionally, a Cross-Modal Synergistic Enhancement Module (CMSEM) and a Cross-Scale Adaptive Fusion Pyramid Network (CSAFPN) are designed to leverage event-derived high-frequency details for RGB feature restoration and to preserve sub-pixel target signatures during deep downsampling. Experimental results demonstrate that EGDNet achieves a mean average precision of 74.63% (AP@0.5) and an F1 score of 78.57%, significantly outperforming state-of-the-art fusion architectures. This work provides a robust visual perception solution for intelligent low-altitude surveillance. Full article
Show Figures

Figure 1

24 pages, 30390 KB  
Article
Managing Street–Level Visible Greenery in High–Density Historic Districts: Spatial Differentiation and Associated Factors in Guangzhou, China
by Yin Ding, Changdong Ye, Long Zhou, Zhiqiang Yi and Yanjun Ou
Land 2026, 15(9), 1611; https://doi.org/10.3390/land15091611 - 31 Aug 2026
Viewed by 117
Abstract
In high–density historic districts, heritage conservation, compact urban morphology, and intensive land use constrain opportunities for large–scale greening, increasing the importance of street–level visible greenery. Taking Guangzhou’s historic districts as a case, this study quantifies the street–level Green View Index (GVI) from 45,380 [...] Read more.
In high–density historic districts, heritage conservation, compact urban morphology, and intensive land use constrain opportunities for large–scale greening, increasing the importance of street–level visible greenery. Taking Guangzhou’s historic districts as a case, this study quantifies the street–level Green View Index (GVI) from 45,380 Baidu Street View sampling points using Mask2Former semantic segmentation and analyzes 892 valid 150 m × 150 m grid cells. Global and Local Moran’s I were used to characterize spatial dependence and clustering. Spearman correlation assessed exploratory bivariate associations, while ordinary least squares (OLS), residual spatial–autocorrelation diagnostics, the spatial lag model (SLM), and the spatial error model (SEM) were used to evaluate multivariable associations and residual spatial dependence. Mean GVI was 16.74% at the point level and 16.70% at the grid level. Global Moran’s I was 0.354 (z = 19.585, permutation p = 0.001), indicating significant positive spatial autocorrelation. Across the tested grid resolutions, SVI and SEI remained significantly negatively associated with GVI, while NDVI and NDBI retained significant positive and negative associations, respectively. Additional specifications excluding SVI and/or SEI further supported the consistency of the NDVI and NDBI associations. SEM provided the best fit among the compared models, with substantial spatial dependence remaining in the error structure after the measured covariates were controlled. Overall, visible greenery was significantly associated with vegetation background, built–up intensity, and street–interface characteristics. The findings provide a spatial evidence base for context–sensitive land management and micro–renewal aimed at improving the continuity and street–level visibility of existing greenery while respecting historic urban fabric. Full article
Show Figures

Figure 1

33 pages, 30121 KB  
Article
Geometry-Enhanced Point–Voxel Fusion with Active Learning for Label-Efficient Point Cloud Semantic Segmentation
by Cheng Zhang, Fei Meng, Yichang Qiu, Haofei Zhao, Yefei Liu and Jianfeng Huang
Appl. Sci. 2026, 16(17), 8650; https://doi.org/10.3390/app16178650 - 31 Aug 2026
Viewed by 85
Abstract
Point cloud semantic segmentation is fundamental for 3D scene understanding and has been widely used in autonomous driving and infrastructure inspection applications. However, its performance is often limited by insufficient representation of local geometric structures and the high cost of point-wise annotation. To [...] Read more.
Point cloud semantic segmentation is fundamental for 3D scene understanding and has been widely used in autonomous driving and infrastructure inspection applications. However, its performance is often limited by insufficient representation of local geometric structures and the high cost of point-wise annotation. To address these issues, this paper proposes GeoFuse-AL, a label-efficient segmentation framework that integrates a geometry-guided point–voxel network with multi-cue active learning. The base model, GeoFuseNet, builds on a hybrid point–voxel backbone and incorporates a Local Geometry Prototype Attention module to enhance object boundaries, fine-grained structures, and local geometric patterns. An Adaptive Channel Fusion module is further designed to improve feature interaction between point-level details and voxel-level context. To reduce annotation dependence, a Multi-Cue Diversity Active Sampling strategy combines prediction uncertainty, color-gradient variation, geometric curvature, and feature-space clustering to select informative and diverse samples. Experiments on S3DIS and SemanticKITTI demonstrate that the proposed model achieves mIoU scores of 63.3% and 61.7%, respectively, outperforming several representative methods. Under limited annotation settings, the proposed strategy reaches 99.2% of fully supervised performance with only 15% labeled data on S3DIS and 97.6% with only 5% labeled data on SemanticKITTI. These results demonstrate that GeoFuse-AL improves segmentation accuracy while substantially reducing annotation requirements. Full article
Show Figures

Figure 1

25 pages, 5154 KB  
Article
X-Net: Hybrid Transformer–CNN Framework with Multi-Aspect Attention for Precise Colon Polyp Segmentation
by Benyamin Mirab Golkhatmi and Mohammad Hossein Moattar
AI 2026, 7(9), 336; https://doi.org/10.3390/ai7090336 - 30 Aug 2026
Viewed by 386
Abstract
Accurate localization of colon polyps plays a crucial role in the early diagnosis of colon cancer. In this paper, we propose a novel network architecture, X-Net, which integrates the Pyramid Vision Transformer (PVT) and EfficientNet-B0. We selected EfficientNet-B0 due to its optimized design [...] Read more.
Accurate localization of colon polyps plays a crucial role in the early diagnosis of colon cancer. In this paper, we propose a novel network architecture, X-Net, which integrates the Pyramid Vision Transformer (PVT) and EfficientNet-B0. We selected EfficientNet-B0 due to its optimized design and computational efficiency, enabling it to effectively capture fine-grained local features. Meanwhile, we use PVT V2-b1 (Pb1) because of its hierarchical structure and self-attention mechanism, which enhance the model’s ability to preserve multi-scale features. By concatenating these two networks in part of the decoder, the proposed model effectively leverages both local and global features. Additionally, the X-Net architecture introduces three blocks: Shape-Aware Enhancement Block (SAEB), Multi-Scale Feature Block (MSFB), and Multi-Aspect Attention Block (MAAB). The SAEB utilizes differences in initial-layer features extracted from the encoder to filter noise in polyp images. The MSFB extracts and enhances multi-scale features from the backbone, effectively integrating semantic information, which improves polyp segmentation accuracy and enhances robustness against variations in polyp size and shape. Finally, the MAAB employs high-level features from the outputs of the MSFB and SAEB to attend to the foreground, background, and boundaries. To address the class imbalance problem, we propose a hybrid loss function combining Tversky Loss and Binary Cross-Entropy Loss. For evaluation, we trained the proposed network on four polyp datasets: Kvasir-SEG, CVC-ClinicDB, CVC-T, and ETIS. The results demonstrate that the model performs well across different datasets with varying image quality and characteristics. Full article
Show Figures

Figure 1

24 pages, 19703 KB  
Article
STG: Structured Topology of Gridpoints for Occluded Pedestrian Detection
by Tian Qiu, Jifeng Shen and Xin Zuo
Sensors 2026, 26(17), 5496; https://doi.org/10.3390/s26175496 - 30 Aug 2026
Viewed by 181
Abstract
Pedestrian detection in crowds is a challenging problem in computer vision. Existing occlusion-handling methods heavily rely on expensive visible-box annotations to locate visible body parts, posing severe limitations in label acquisition cost and open-world generalization. To break through this limitation, we propose a [...] Read more.
Pedestrian detection in crowds is a challenging problem in computer vision. Existing occlusion-handling methods heavily rely on expensive visible-box annotations to locate visible body parts, posing severe limitations in label acquisition cost and open-world generalization. To break through this limitation, we propose a novel Structured Topology of Gridpoints (STG) framework. Operating strictly under standard full-box annotations without any extra visibility supervision, STG aims to achieve implicit, fine-grained local semantic compensation. Specifically, we formulate a coarse-to-fine reasoning paradigm consisting of three interactive stages. To mitigate the high spatial complexity and eliminate background redundancy, we first introduce a Saliency-Aware Feature Filtering (SAFF) mechanism, which leverages gridpoint heatmaps to filter out low-confidence pedestrian candidates. Second, a query-guided Across-Instance Feature Interaction (AIFI) model is designed to utilize inter-instance spatial relationships to propagate missing context from highly visible individuals to their occluded neighbors. Finally, we devise a prior-guided Inner-Instance Gridpoints Interaction (I2GI) model to achieve fine-grained structured part-level feature completion, which dynamically aggregates vital localized cues from diverse human parts to reconstruct holistic pedestrian representations. Extensive experiments on the CityPersons, CrowdHuman, and WiderPerson datasets demonstrate the effectiveness and efficiency of our proposed method. Specifically, STG achieves a log-average miss rate of 7.41% on Reasonable and 32.05% on Heavy Occlusion subsets of CityPersons, while running at up to 16 FPS, outperforming existing part-based methods under full-box supervision. Full article
(This article belongs to the Special Issue Image Processing and Analysis for Object Detection: 3rd Edition)
Show Figures

Figure 1

27 pages, 46478 KB  
Article
FMRS-YOLO: A Feature-Modulated and Redundancy-Suppressed YOLO for UAV Remote Sensing Object Detection
by Qianxu Ren, Yong He, Yufeng Li and Qingzhou Li
Appl. Sci. 2026, 16(17), 8624; https://doi.org/10.3390/app16178624 - 29 Aug 2026
Viewed by 174
Abstract
Unmanned aerial vehicle (UAV) remote sensing object detection remains challenging because aerial targets are often small, densely distributed, and embedded in complex background clutter, whereas onboard deployment requires compact computation and real-time inference. Existing real-time detectors commonly retain a conventional P3 [...] Read more.
Unmanned aerial vehicle (UAV) remote sensing object detection remains challenging because aerial targets are often small, densely distributed, and embedded in complex background clutter, whereas onboard deployment requires compact computation and real-time inference. Existing real-time detectors commonly retain a conventional P3P5 detection hierarchy, in which deep low-resolution stages may consume computation after fine-grained small-object cues have already been weakened. To address this issue, this paper presents feature-modulated and redundancy-suppressed YOLO (FMRS-YOLO), a UAV-oriented lightweight detector that improves the accuracy–efficiency trade-off through the redundancy-reduction (RR) scale allocation strategy and targeted feature refinement. Instead of extending the terminal hierarchy to the conventional P5 scale, FMRS-YOLO replaces the original stride-2 P5 transition with a stride-1 high-level transformation and omits the corresponding P5 prediction branch, thereby concentrating detection on the P3 and P4 feature levels where small aerial targets retain more informative spatial cues. To enhance feature representation under the compact RR scale allocation framework, we propose Gaussian cross-feature modulation (GCFM) to strengthen spatial–semantic interaction, and design attention-weighted parallel pyramid fusion (AWPPF) to improve clutter-aware multi-scale aggregation while preserving unpooled spatial details. Experiments are conducted separately on VisDrone-2019, a drone-based object detection benchmark used as the primary benchmark, and UCAS-AOD, an aerial object detection benchmark for aircraft and cars used as an independent complementary benchmark. Compared with scale-matched YOLOv11 baselines on VisDrone-2019, FMRS-YOLO improves mAP by 1.9–2.4 percentage points and mAP50 by 2.3–2.9 percentage points, while reducing parameters by 41.0–64.9%, reducing FLOPs by 1.6–20.7%, and increasing FPS by 9.3–10.4%. On UCAS-AOD, FMRS-YOLOn achieves 71.3% mAP, 98.8% mAP50, and 88.3% mAP75, showing favorable localization performance under a compact model scale. Ablation and visualization results further indicate that the proposed architecture improves foreground focus, suppresses background interference, and strengthens small-object localization. These results demonstrate that FMRS-YOLO provides a practical accuracy–efficiency trade-off for UAV remote sensing object detection. Full article
(This article belongs to the Section Aerospace Science and Engineering)
Show Figures

Figure 1

Back to TopTop