Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (868)

Search Parameters:
Keywords = semantic segmentation algorithm

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
23 pages, 3265 KB  
Article
A Weakly Supervised Segmentation Algorithm Based on Local–Global Class Labelling Comparison
by Binyu Guo, Laibao Yu, Yiming Yang and Chunzhi Wang
Appl. Sci. 2026, 16(17), 8496; https://doi.org/10.3390/app16178496 - 26 Aug 2026
Abstract
As a key pixel-level analysis technology, semantic segmentation is widely deployed in autonomous driving and medical imaging. Fully supervised segmentation relies on labour-intensive pixel-wise annotations, so weakly supervised semantic segmentation (WSSS) with only image-level labels has attracted wide attention. Existing Vision Transformer (ViT)-based [...] Read more.
As a key pixel-level analysis technology, semantic segmentation is widely deployed in autonomous driving and medical imaging. Fully supervised segmentation relies on labour-intensive pixel-wise annotations, so weakly supervised semantic segmentation (WSSS) with only image-level labels has attracted wide attention. Existing Vision Transformer (ViT)-based WSSS methods suffer from two critical limitations: ViT’s global self-attention mechanism leads to insensitivity to local small target features and incomplete foreground activation; its class-agnostic attention maps frequently misactivate background regions as foreground objects, introducing heavy noise. To tackle these two issues, this paper proposes a single-stage weakly supervised segmentation algorithm based on local–global class labelling comparison. First, we design a local–global class labelling comparison (LTG) module. By feeding both original images and randomly cropped local patches into ViT, we adopt InfoNCE contrastive loss to align local class tokens with global class tokens, enhancing the feature integrity of local target regions and suppressing background false activation. Second, a class-aware stimulus module (CSM) is embedded into ViT’s multi-head attention branch. It injects category semantic constraints into self-attention to generate class-aware attention maps, guiding the model to focus on real foreground targets and reduce background interference. Finally, we construct a feature fusion class-aware activation map (FFCAM) by fusing ViT global output features and CSM class-aware attention features to generate high-quality pseudo-labels for segmentation training. Extensive experiments are conducted on PASCAL VOC 2012 and MS COCO 2014 datasets. Our method achieves 78.5% mIoU on the validation set of PASCAL VOC 2012 and 50.9% mIoU on MS COCO 2014, showing competitive numerical performance among the compared single-stage ViT-based WSSS approaches. Ablation experiments verify the independent and joint effectiveness of LTG, CSM and FFCAM. The proposed method effectively improves the completeness of target activation regions and suppresses background noise, and it maintains strong generalization for slender, small and texture-sparse objects. In future work, we will further lightweight the ViT backbone to reduce computational overhead for embedded deployment. Full article
Show Figures

Figure 1

36 pages, 82509 KB  
Article
A TLS-Based Framework for the Realization of Digital Twin Basemaps Applied to an Adaptively Reused Heritage Building
by Mohamed H. Salaheldin, Ahmed Shaker and Songnian Li
Appl. Sci. 2026, 16(16), 8306; https://doi.org/10.3390/app16168306 - 20 Aug 2026
Viewed by 208
Abstract
The transition toward urban-scale digital twin and smart city management requires survey-grade 3D basemaps, yet conventional documentation remains time-consuming and prone to inaccuracies. While Terrestrial Laser Scanning (TLS) offers rapid 3D acquisition, capturing complex, GNSS-denied multi-story interiors frequently causes cumulative registration errors and [...] Read more.
The transition toward urban-scale digital twin and smart city management requires survey-grade 3D basemaps, yet conventional documentation remains time-consuming and prone to inaccuracies. While Terrestrial Laser Scanning (TLS) offers rapid 3D acquisition, capturing complex, GNSS-denied multi-story interiors frequently causes cumulative registration errors and isolated indoor–outdoor data silos. To address this, this study proposes a comprehensive typology-agnostic framework for developing high-fidelity digital twin basemaps. Treating the building as a unified spatial network, the methodology systematically mitigates error propagation through strategic linkage planning, rigid shell-first registration, continuous vertical core anchoring (via stairwells), and adaptive multi-source data fusion. Implemented on an adaptively reused heritage building, the developed basemap achieved an absolute georeferencing accuracy of 30.0 mm (RMSE) against an independent total station control network, alongside a mean relative error of 3.11 mm. Comparative analysis against legacy 2D CAD floor plans revealed simplified geometric representations and categorical dimensional deviations of up to 29.8 cm. Demonstrating its practical utility, the point cloud-centric geometric hub avoids forced geometric idealization, successfully supporting direct immersive visualization, architectural slicing (floor plans, sections, elevations), and multi-LOD algorithmic planar segmentation. This spatially constrained acquisition strategy bypasses legacy limitations, delivering a mathematically verified 3D reality capture essential for smart facility management, heritage conservation, and downstream semantic intelligence. Full article
Show Figures

Figure 1

34 pages, 21458 KB  
Article
Adaptive Flight Maneuver Boundary Localization via Spectral Entropy-Weighted Multi-Channel Spectrogram Fusion
by Shansong Song, Wei Han, Bing Wan, Xiangyi Liu, Xichao Su, Chao Li and Yunyang Cao
Entropy 2026, 28(8), 922; https://doi.org/10.3390/e28080922 - 17 Aug 2026
Viewed by 116
Abstract
To address ambiguous maneuver boundaries, background interference, and uneven multi-sensor quality in long-duration flight parameter recordings, this paper proposes an adaptive flight maneuver boundary localization algorithm that integrates spectral entropy-weighted multi-channel spectrogram fusion with attitude-constrained structural correction. Multi-channel Short-Time Fourier Transform (STFT) spectrograms [...] Read more.
To address ambiguous maneuver boundaries, background interference, and uneven multi-sensor quality in long-duration flight parameter recordings, this paper proposes an adaptive flight maneuver boundary localization algorithm that integrates spectral entropy-weighted multi-channel spectrogram fusion with attitude-constrained structural correction. Multi-channel Short-Time Fourier Transform (STFT) spectrograms are first constructed from flight parameter time series. Spectral entropy (SE) is introduced to quantify the uncertainty of each channel’s time–frequency energy distribution and is combined with the maneuver activation ratio (MAR) and the linear contrast ratio (LCR) to form objective credibility weights, thereby suppressing channels dominated by aerodynamic turbulence and high frequency structural vibration. Normal overload soft gating and logarithmic noise floor subtraction are then applied to obtain an enhanced fused spectrogram, from which candidate intervals are extracted by low band energy thresholding. Finally, roll and pitch angle steady-state priors refine the event structure through local boundary refinement, cross-segment expansion/chain merging, and semantic post-processing, recovering continuous maneuvers fragmented by instantaneous energy valleys. On the held-out test sorties (SE_018–SE_020; 61 annotated intervals), the proposed algorithm achieves Precision, Recall, and F1-scores of 0.967. On the full primary corpus of 20 sorties (461 intervals), used for ablation and sensitivity analyses, the corresponding figures are Precision 0.934, Recall 0.959, and F1 0.946, with start and end boundary mean absolute errors of 1.484 s and 1.471 s. Under the same IoU protocol, consistent superiority is observed against learning-based baselines, and an independent external set of 10 sorties yields F1 = 0.938. The results indicate that entropy-constrained multi-sensor time–frequency fusion mainly improves maneuver/background separability, whereas attitude-constrained structural correction restores the integrity of long continuous maneuvers. Full article
(This article belongs to the Section Signal and Data Analysis)
Show Figures

Figure 1

28 pages, 14275 KB  
Article
Energy Optimization-Based Segmentation and Extraction of 3D Sonar Point Clouds for Complex Underwater Structures
by Junchi Dong, Zilong Li, Hao Li, Shaobo Li and Yunlong Wu
Remote Sens. 2026, 18(16), 2759; https://doi.org/10.3390/rs18162759 - 15 Aug 2026
Viewed by 207
Abstract
3D sonar is the primary technical means for the detection and health assessment of underwater structures. However, due to the harsh constraints of underwater imaging conditions, these point clouds experience severe noise interference and inconspicuous features. This complicates structural surface extraction and hinders [...] Read more.
3D sonar is the primary technical means for the detection and health assessment of underwater structures. However, due to the harsh constraints of underwater imaging conditions, these point clouds experience severe noise interference and inconspicuous features. This complicates structural surface extraction and hinders the engineering monitoring process. To address this challenge, we propose an object-based energy-optimization segmentation method for precise extraction. First, the Voxel Cloud Connectivity Segmentation (VCCS) algorithm transforms massive discrete point clouds into semantically coherent supervoxel objects, significantly reducing computational complexity. Next, a global energy optimization framework is constructed, integrating a data term that characterizes structural surfaces and a smoothness term based on spatial neighborhood constraints. Through a graph-cut optimization procedure, this model enables the joint extraction of conspicuous and inconspicuous surfaces. Finally, a region-growing algorithm with multi-attribute constraints effectively filters typical sonar noise. Experimental results demonstrate the method’s effectiveness, achieving an average F1-score of 89.19%. This approach maintains high segmentation accuracy and provides reliable technical support for underwater structure detection and assessment. Full article
(This article belongs to the Special Issue Remote Sensing for Maritime Monitoring)
Show Figures

Figure 1

20 pages, 43201 KB  
Article
Application of Deep Learning Semantic Segmentation Models in Remote Sensing-Based Cropland Non-Grain and Non-Agriculturalization Monitoring: A Comparative Study
by Zhao Deng, Ming Cheng, Junde Xie, Tianyong Wan, Pengzhi Yang, Jianbo Tan, Sixue Xia, Jia Zhang and Xin Wu
Land 2026, 15(8), 1437; https://doi.org/10.3390/land15081437 - 9 Aug 2026
Viewed by 257
Abstract
Cropland non-grain and non-agriculturalization monitoring (CNNM) is of great significance for ensuring national food security. Recently, semantic segmentation has been a promising approach that can fully exploit the advantages of high-resolution remote sensing imagery. It is of great significance to clarify the performance [...] Read more.
Cropland non-grain and non-agriculturalization monitoring (CNNM) is of great significance for ensuring national food security. Recently, semantic segmentation has been a promising approach that can fully exploit the advantages of high-resolution remote sensing imagery. It is of great significance to clarify the performance of semantic segmentation algorithms for CNNM and their influencing factors. This paper contributes to research on deep learning-based semantic segmentation models for CNNM by investigating the robustness and generalization ability of detection models in complex scenarios. To investigate these models, seven mainstream segmentation models were examined with respect to two self-constructed unmanned aerial vehicle (UAV)-based cropland monitoring datasets under various experimental settings. The experimental results reveal several key findings: For the datasets employed in the present study, remote sensing semantic segmentation models demonstrate superior performance relative to generic visual models. Increasing backbone depth does not guarantee significant performance improvements. The recognition challenges posed by task-specific categories such as agricultural facility land further underscore the need for tailored deep learning semantic segmentation algorithms for CNNM. These experimental results provide valuable insights into the practical performance of deep learning semantic segmentation models and offer useful guidance for future semantic segmentation research in the field of CNNM. Full article
Show Figures

Figure 1

41 pages, 2283 KB  
Article
PartSense-IP: Part-Aware Vision–Language Sensor Fusion for Visual–Semantic Consistency Evaluation of IP Prototypes
by Yangfan Feng and Wen Zhao
Sensors 2026, 26(16), 5052; https://doi.org/10.3390/s26165052 - 9 Aug 2026
Viewed by 239
Abstract
Evaluating whether an intellectual property (IP) prototype faithfully preserves the visual identity and semantic intent of its original concept design is an important yet challenging task in product design and creative prototyping. Existing evaluation practices mainly rely on manual inspection or global image-level [...] Read more.
Evaluating whether an intellectual property (IP) prototype faithfully preserves the visual identity and semantic intent of its original concept design is an important yet challenging task in product design and creative prototyping. Existing evaluation practices mainly rely on manual inspection or global image-level similarity comparison, which are subjective, difficult to reproduce, and insufficient for localizing identity-critical deviations. To address this problem, this paper proposes PartSense-IP, a part-aware vision–language sensor fusion framework for visual–semantic consistency evaluation of IP prototypes. The proposed framework takes a 2D concept image, an optional textual design description, and multi-view RGB-D sensor observations of a prototype as inputs. It first constructs a multi-view prototype representation and decomposes both the concept and prototype observations into design-relevant parts. Dense visual features, color and shape descriptors, and vision–language semantic embeddings are then extracted to evaluate part-level consistency. A Part-Aware Visual–Semantic Consistency Fusion (PVCF) algorithm is further developed to integrate shape, color, local visual similarity, semantic alignment, and cross-view stability into a unified IP consistency score. In addition to scalar scoring, PartSense-IP generates localized difference maps, 3D inconsistency visualization, and interpretable design feedback for prototype refinement. Experiments on the proposed IP-ProtoSense evaluation protocol demonstrate that PartSense-IP outperforms representative vision–language, dense-visual, segmentation-based, and 3D multimodal baselines in consistency scoring, inconsistency detection, localization, ablation, and robustness evaluation. Full article
(This article belongs to the Section Optical Sensors)
Show Figures

Figure 1

26 pages, 10571 KB  
Article
A Task-Oriented Segmentation Algorithm Based on Residual Super-Resolution and Edge-Aware Learning for Real-Time Agricultural Vision
by Esther Gascó, Clara I. López-González, Gonzalo Pajares and Eva Besada-Portas
Algorithms 2026, 19(8), 650; https://doi.org/10.3390/a19080650 - 6 Aug 2026
Viewed by 260
Abstract
Accurate vegetation-soil segmentation is a key component of intelligent agricultural vision systems operating under limited computational resources. Existing CNN-based approaches generally improve segmentation accuracy by increasing network complexity or relying on high-resolution imagery, limiting their suitability for real-time embedded applications. This paper proposes [...] Read more.
Accurate vegetation-soil segmentation is a key component of intelligent agricultural vision systems operating under limited computational resources. Existing CNN-based approaches generally improve segmentation accuracy by increasing network complexity or relying on high-resolution imagery, limiting their suitability for real-time embedded applications. This paper proposes a lightweight task-oriented CNN segmentation algorithm that integrates residual structural reconstruction, spectral augmentation, semantic segmentation, and edge-aware refinement within a unified multi-task optimization framework. Unlike conventional super-resolution methods, the residual module is specifically designed to recover segmentation-relevant structural features, including vegetation contours, crop-row geometry, and vegetation-soil transitions, without increasing spatial resolution. The algorithm combines four computational stages-spectral augmentation, pseudo low-resolution generation, residual enhancement, and edge-guided refinement-to improve internal feature representations while preserving computational efficiency. Experiments on the Crop Row Benchmark Dataset demonstrate competitive segmentation performance (BF1 = 0.95, mIoU = 0.92) with real-time inference (54.7 FPS). Comparative experiments, ablation studies, computational complexity analysis, and Grad-CAM-based interpretability analysis demonstrate the effectiveness of the proposed algorithm and the complementary contribution of its computational modules. The proposed formulation provides an efficient and interpretable CNN-based solution for resource-constrained agricultural vision systems. Full article
(This article belongs to the Collection Algorithms for Computer Vision Applications)
Show Figures

Graphical abstract

21 pages, 22855 KB  
Article
Airborne Point Cloud Fusion with Local Plane Constraints for Advanced Semantic Consistency
by Shahoriar Parvaz, Felicia N. Teferle, Abdul Nurunnabi, Roderik Lindenbergh and Luis A. Leiva
Remote Sens. 2026, 18(15), 2598; https://doi.org/10.3390/rs18152598 - 5 Aug 2026
Viewed by 310
Abstract
Point cloud fusion is crucial in geospatial analysis, combining data from multiple sources (e.g, LiDAR and photogrammetry) to provide a more complete and accurate environmental representation. However, integrating airborne hybrid sensors or cross-source point clouds remains challenging due to variations in geometric accuracy, [...] Read more.
Point cloud fusion is crucial in geospatial analysis, combining data from multiple sources (e.g, LiDAR and photogrammetry) to provide a more complete and accurate environmental representation. However, integrating airborne hybrid sensors or cross-source point clouds remains challenging due to variations in geometric accuracy, data precision, gaps, and sensor attributes. Despite recent advancements, these challenges remain and are among the most demanding aspects in geospatial data processing for remote sensing applications. We propose a new point cloud fusion algorithm that leverages local plane constraints to achieve advanced semantic consistency. The proposed method dynamically fits local planes to the target point clouds, enabling robust alignment of source points to these planes. Evaluation on two real-world datasets demonstrates significant gains in accuracy and preservation of geometric details. Our algorithm also improves the accuracy of downstream tasks such as semantic segmentation. In our experiment, the overall accuracy for the Dudelange dataset increases from 48.5% to 80.1%, and that for the Dublin dataset increases from 72.9% to 88.0%. While challenges persist with sparse and noisy datasets, experimental results highlight the effectiveness of the proposed method, offering valuable insights for maximizing the potential of cross-source point cloud data. Full article
Show Figures

Figure 1

18 pages, 48650 KB  
Article
PMS-Net: A Real-Time Instance Segmentation Framework with Large Receptive Fields for Urban Driving Scenes
by Ling Zhang and Zhuang Xiong
Algorithms 2026, 19(8), 620; https://doi.org/10.3390/a19080620 - 24 Jul 2026
Viewed by 526
Abstract
With the rapid development of edge AI chips and autonomous driving algorithms, autonomous driving perception systems are evolving toward higher accuracy, lower latency, and lightweight deployment. This trend places greater demands on real-time instance segmentation algorithms in terms of multi-scale feature representation, spatial [...] Read more.
With the rapid development of edge AI chips and autonomous driving algorithms, autonomous driving perception systems are evolving toward higher accuracy, lower latency, and lightweight deployment. This trend places greater demands on real-time instance segmentation algorithms in terms of multi-scale feature representation, spatial detail modeling, and edge deployment efficiency. To address these challenges, this paper proposes PMS-Net (Progressive Multi-Scale Network), a network designed for real-time instance segmentation. PMS-Net adopts a progressive multi-scale feature modeling mechanism that progressively enlarges the receptive field while integrating semantic and fine-grained spatial information across different scales. This enables efficient collaboration between local features and global contextual information, thereby enhancing scale-awareness and feature representation while maintaining a lightweight architecture. In addition, efficient feature encoding and dynamic feature reconstruction are incorporated to further improve spatial alignment, boundary recovery, and semantic continuity for complex scene modeling. Experimental results show that PMS-Net achieves 36.7% Mask mAP50 and 175 FPS on the Cityscapes dataset, outperforming the baseline by 2.9%. Deployment experiments on the NVIDIA Jetson Orin NX platform further demonstrate that PMS-Net achieves 35.2% Mask mAP50 and 98 FPS, improving the baseline by 3.7% and 14.0%, respectively. These results validate the effectiveness and practicality of PMS-Net for real-time edge-deployed autonomous driving applications. Full article
Show Figures

Figure 1

19 pages, 4439 KB  
Article
An Algorithm for Fine-Grained Content Extraction and Understanding in Short Videos
by Yanqi Wan, Shuya Zhang, Yi Xu, Kunfang Zhang, Heyi Wang and Mingzheng Liu
Data 2026, 11(7), 179; https://doi.org/10.3390/data11070179 - 20 Jul 2026
Viewed by 533
Abstract
We integrated communication theory with advanced computer vision techniques to propose a novel approach for fine-grained content extraction from short videos. Unlike methods focused on summarization or subtitle generation for longer videos, our approach emphasizes extracting detailed content and understanding the intricate narrative [...] Read more.
We integrated communication theory with advanced computer vision techniques to propose a novel approach for fine-grained content extraction from short videos. Unlike methods focused on summarization or subtitle generation for longer videos, our approach emphasizes extracting detailed content and understanding the intricate narrative structure of short videos. By employing scene segmentation, similarity-based filtering algorithms, and support vector machines, the method identifies keyframes that capture precise visual details. Further, it generates semantically accurate textual descriptions using the mPLUG model, enabling an in-depth understanding of video content. Using a dataset of short videos from the cultural and tourism domain, we validated the proposed method. Experimental results demonstrate that our approach achieves high precision in identifying and understanding detailed visual elements, effectively bridging the gap between visual representation and semantic meaning. Additionally, the study explores the influence of different video content types, interference factors, and image description models on fine-grained content extraction, highlighting its potential for improving intelligent analysis of short-video data. Full article
(This article belongs to the Special Issue Vision-Based AI in the Real World: Data, Robustness and Deployment)
Show Figures

Figure 1

42 pages, 50781 KB  
Article
Urban Outdoor Thermal Environment Analysis Based on Semantic Segmentation and Morphology Indicators: A Case Study of Residential Blocks in Wuhan
by Hongying Wang, Lin Cai and Kai Guo
Buildings 2026, 16(14), 2870; https://doi.org/10.3390/buildings16142870 - 19 Jul 2026
Viewed by 184
Abstract
Rapid urbanization has intensified urban heat issues. Previous studies often relied on subjective block selection and rarely integrated vegetation data. This study extracted vegetation from Wuhan’s satellite imagery and combined it with building geometry to generate large-scale 3D block models. Typical blocks were [...] Read more.
Rapid urbanization has intensified urban heat issues. Previous studies often relied on subjective block selection and rarely integrated vegetation data. This study extracted vegetation from Wuhan’s satellite imagery and combined it with building geometry to generate large-scale 3D block models. Typical blocks were identified by clustering, and thermal environments were simulated using ENVI-met to establish regression models. POI and spatial analyses validated the results. The study found that 1. t-SNE outperforms PCA and UMAP in dimensionality reduction. 2. K-means surpasses GMM, DBSCAN, and Spectral in clustering. 3. SVFave, FAall, VDW, VAR, and BBA are critical for block morphology classification and block outdoor thermal assessment. 4. The final ridge regression model based on these indices achieved high R2 values (0.805, 0.507, and 0.855), indicating excellent model performance. 5. The blocks in Cluster 1 (west of the Yangtze River) exhibit higher mean air temperatures. 6. the blocks in Cluster 2 (new areas) have high vegetation coverage, causing larger temperature differences between the inside and outside of blocks. This study provides a comprehensive workflow for urban block morphology classification and thermal assessment. Full article
Show Figures

Figure 1

25 pages, 2996 KB  
Article
ReViTA-Unet: An Enhanced Semantic Segmentation Model for Automated Morphometric Analysis of Macrobrachium rosenbergii
by Dawei Sun, Qi Chen, Guanghui Yu, Xinran Han, Chen Li, Chengquan Zhou and Hongbao Ye
Sensors 2026, 26(14), 4570; https://doi.org/10.3390/s26144570 - 19 Jul 2026
Viewed by 479
Abstract
Accurate morphometric analysis of Macrobrachium rosenbergii is essential for selective breeding, growth monitoring, and precision aquaculture, yet conventional manual measurements are labor-intensive, time-consuming, and prone to operator variability. This study presents ReViTA-UNet, an automated, non-contact morphometric analysis framework based on an enhanced semantic [...] Read more.
Accurate morphometric analysis of Macrobrachium rosenbergii is essential for selective breeding, growth monitoring, and precision aquaculture, yet conventional manual measurements are labor-intensive, time-consuming, and prone to operator variability. This study presents ReViTA-UNet, an automated, non-contact morphometric analysis framework based on an enhanced semantic segmentation network coupled with a geometric topology refinement algorithm to accurately extract multiple morphological traits. The proposed framework integrates complementary feature extraction to improve segmentation of elongated anatomical structures and complex body boundaries. A complete automated measurement system was subsequently developed to convert segmented images into biologically meaningful morphometric parameters. The results demonstrated that ReViTA-UNet achieved a Dice coefficient of 97.7%, a mean Intersection over Union (mIoU) of 96.7%, a precision of 98.3%, and a recall of 98.4%, outperforming eight representative semantic segmentation models. The automated measurement system achieved a mean absolute percentage error of 1.83% for body length, with strong agreement with manual measurements (R2 = 0.987), while maintaining high accuracy for other major morphometric traits. These results indicate that the proposed framework provides an accurate and efficient solution for automated prawn phenotyping under controlled imaging conditions. It establishes a practical foundation for future intelligent aquaculture applications following validation under commercial farming environments. Full article
(This article belongs to the Section Smart Agriculture)
Show Figures

Graphical abstract

20 pages, 5470 KB  
Article
Fine-Scale Spartina alterniflora Mapping Using Advanced Deep Learning and High-Resolution UAV Imagery
by Zegang Chen, Yixing Tang, Yulong Hu, Bin Chu, Zhong Long, Dong Xie, Yunfei Zhang and Zhipan Wang
Remote Sens. 2026, 18(14), 2375; https://doi.org/10.3390/rs18142375 - 16 Jul 2026
Viewed by 429
Abstract
Accurate mapping of Spartina alterniflora (S. alterniflora) is critical for coastal conservation. While deep learning has improved mapping, current algorithms that rely on moderate-resolution satellite imagery struggle to detect minute, early-stage invasive patches amid complex backgrounds with similar textures (e.g., water [...] Read more.
Accurate mapping of Spartina alterniflora (S. alterniflora) is critical for coastal conservation. While deep learning has improved mapping, current algorithms that rely on moderate-resolution satellite imagery struggle to detect minute, early-stage invasive patches amid complex backgrounds with similar textures (e.g., water ripples, native vegetation). To overcome this resolution bottleneck, we construct UAV-SaSeg, the first high-resolution UAV semantic segmentation dataset targeting these complicated scenarios. Furthermore, we propose DINOsegNext, a novel mapping model integrating the DINOv3 foundation model with a Global Filter (GF) multi-frequency convolution. This architecture leverages generalized knowledge and spatial-frequency feature decoupling to filter out confusing background noise effectively. Validated in real-world intertidal zones, DINOsegNext outperforms state-of-the-art models, achieving an IoU of 0.7420 and an F1-Score of 0.8519. Crucially, it maintains high computational efficiency, with an inference latency of merely 8.5 ms per image, making it well-suited for onboard edge computing. This work provides urgently needed high-resolution data and a robust technical path for the precise, efficient monitoring of invasive species. Full article
Show Figures

Figure 1

28 pages, 69523 KB  
Article
FLDO-LKNet: An Efficient Method for Segmenting Stone Cells in Rubber Tree Bark with Robustness to Staining Differences and Morphological and Scale Changes
by Yeling Peng, Yuanyuan Zhang, Hao Zeng, Yiqing Zeng, Meixi Pan, Zhongyang Peng, Yang Liu, Yongling Xia, Peng Wang, Mingfang He and Yaowen Hu
Plants 2026, 15(14), 2178; https://doi.org/10.3390/plants15142178 - 16 Jul 2026
Viewed by 379
Abstract
Stone cells are an important structural component of rubber tree bark and are closely associated with traits such as cracking propensity, bark hardness, stress tolerance and latex production. However, high-precision segmentation methods that are robust to staining variations, morphological diversity, and scale changes [...] Read more.
Stone cells are an important structural component of rubber tree bark and are closely associated with traits such as cracking propensity, bark hardness, stress tolerance and latex production. However, high-precision segmentation methods that are robust to staining variations, morphological diversity, and scale changes are still missing. To address these challenges in stone-cell histological images, such as the inconsistent staining intensity, the diverse shapes and sizes of cells, and the interference of cell debris, we proposed an automated semantic segmentation network called FLDO-LKNet, which formulates stone-cell delineation as a pixel-wise semantic segmentation task. Specifically, we introduce an LDFE module to recalibrate backbone features at the channel level, thereby mitigating the effects of staining differences. A KCFA attention mechanism is designed to better capture complex morphology and scale variation. In addition, we develop an FLDO optimization algorithm with a performance-feedback-based dynamic learning rate adjustment strategy to enhance robustness against training instability caused by debris interference. We further construct a dataset of 1084 stone-cell images collected from CATAS to support model training and evaluation. Experimental results demonstrate that FLDO-LKNet achieves 75.18% mIoU, 98.5% accuracy, and 82.21% sensitivity. Overall, as a dedicated semantic segmentation network, the proposed method enables high-precision pixel-level segmentation of stone cells, which may facilitate subsequent studies of stone-cell development and genetic functions and shows potential for agricultural applications. Full article
(This article belongs to the Special Issue Advances in Artificial Intelligence for Plant Research—2nd Edition)
Show Figures

Figure 1

26 pages, 18794 KB  
Article
DWFSeg: A Dynamic Multiscale Feature Fusion and Dual Attention-Enhanced Network for High-Precision Water Body Segmentation Based on Super-Resolution Remote Sensing Imagery
by Ziwei Li, Bingjie Liang, Jianzhong Guo, Ning Li, Weiran Luo, Baowei Zhang, Jiali Guo, Weizhen Zhang, Yan Zhou, Yuezhen Guo and Yishan Li
Remote Sens. 2026, 18(14), 2271; https://doi.org/10.3390/rs18142271 - 8 Jul 2026
Viewed by 352
Abstract
Remote sensing imagery provides a primary data source for large-scale surface water body monitoring, which is crucial for quantifying climate-related hydrological impacts, supporting flood control, and sustaining integrated water resource management. However, remote sensing images generally face the trade-off between spatial resolution and [...] Read more.
Remote sensing imagery provides a primary data source for large-scale surface water body monitoring, which is crucial for quantifying climate-related hydrological impacts, supporting flood control, and sustaining integrated water resource management. However, remote sensing images generally face the trade-off between spatial resolution and temporal coverage. To address this issue, the Real-ESRGAN super-resolution algorithm is employed to reconstruct temporally continuous, wide-coverage medium-resolution imagery to a 2.5 m resolution, effectively improving its capability to identify sub-pixel river boundaries. Water body segmentation (WBS) is an effective method for fine-detail surface water extraction. Nonetheless, when applied in complex hydrological environments, it still faces several limitations, such as ambiguous delineation of land–water boundaries and the difficulty in capturing multiscale water body characteristics. To address these issues, a Dynamic Weight Fusion SegFormer (DWFSeg) network is constructed, integrating a MixVision Transformer (MVT) encoder with a multiscale decoding architecture. Specifically, a Dynamic Multiscale Feature Fusion (DMFF) mechanism is proposed, which adaptively assigns semantic-guided fusion weights to multiscale feature water bodies. Furthermore, the Dual Attention-Enhanced (DAE) module strengthens discriminative essential features and suppresses background noise in both channel and spatial dimensions. Evaluated on a self-constructed super-resolution imagery dataset (SID) and the public GID, DWFSeg achieves overall accuracies of 98.08% and 96.14%, respectively. It outperforms representative benchmark models across multiple quantitative metrics, while maintaining competitive inference efficiency and favorable segmentation stability. Ablation studies verify the effectiveness and necessity of each proposed component. The presented network provides a reliable technical solution and supports refined water resource evaluation and sustainable watershed management. Full article
Show Figures

Figure 1

Back to TopTop