Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (517)

Search Parameters:
Keywords = superpixels

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
27 pages, 6835 KB  
Article
A Reliability-Aware Cross-Branch Contrastive Graph Convolutional Network for Hyperspectral Image Processing
by Runhao Zhang, Wanzhang Wang, Wei Feng, Fei Yu and Haize Hu
Algorithms 2026, 19(9), 774; https://doi.org/10.3390/a19090774 - 9 Sep 2026
Abstract
Hyperspectral images contain abundant spectral information and provide fine-grained spatial representations. However, their high dimensionality, severe spectral redundancy, subtle inter-class differences, mixed boundary regions, and limited labeled samples pose significant challenges to accurate classification. Convolutional neural networks (CNNs) have limited capability in modeling [...] Read more.
Hyperspectral images contain abundant spectral information and provide fine-grained spatial representations. However, their high dimensionality, severe spectral redundancy, subtle inter-class differences, mixed boundary regions, and limited labeled samples pose significant challenges to accurate classification. Convolutional neural networks (CNNs) have limited capability in modeling non-Euclidean structural relationships, whereas graph convolutional networks (GCNs) are susceptible to the quality of superpixel segmentation and noise propagation over graph structures. To address these issues in hyperspectral image classification, this paper proposes a Reliability-Aware Cross-Branch Contrastive Graph Convolutional Network (RACB-CGCN). The proposed method employs a dual-branch CNN–GCN architecture to extract pixel-level local spectral–spatial features and superpixel-level structural features, respectively. A superpixel reliability estimation and propagation control mechanism is introduced to assess node reliability based on the discrepancy between pixel-level features and superpixel-reconstructed features. This mechanism effectively suppresses the propagation of noisy information caused by impure superpixels and mixed boundary regions. Meanwhile, a cross-branch supervised contrastive learning strategy is developed to enhance semantic consistency between the CNN and GCN branches, thereby improving intra-class compactness and inter-class separability. In addition, a class-adaptive fusion module is designed to dynamically adjust the contributions of the two branches according to the feature characteristics of different land-cover classes. Experimental results demonstrate that the proposed method effectively exploits the complementary information between pixel-level fine-grained features and superpixel-level structural features, leading to improved classification accuracy and robustness in hyperspectral image classification. Full article
Show Figures

Figure 1

32 pages, 89911 KB  
Article
Homogeneous Terrain Unit Extraction by Integrating Superpixel Segmentation and Multiscale Region Merging: A Case Study in the Deeply Incised Valleys of Southeastern Tibet
by Zhongkang Yang, Shishu Zhang, Jianhui Deng, Jingen Ma, Qingchun Li, Jinbing Wei and Siyuan Zhao
Remote Sens. 2026, 18(17), 3028; https://doi.org/10.3390/rs18173028 - 4 Sep 2026
Viewed by 219
Abstract
Mapping mountain surfaces requires spatial units that represent both hillslope-scale structure and local within-slope terrain heterogeneity. Hydrological slope units provide limited representation of within-slope objects, whereas general object-based segmentation is sensitive to fragmentation and scale selection. We developed a homogeneous terrain unit extraction [...] Read more.
Mapping mountain surfaces requires spatial units that represent both hillslope-scale structure and local within-slope terrain heterogeneity. Hydrological slope units provide limited representation of within-slope objects, whereas general object-based segmentation is sensitive to fragmentation and scale selection. We developed a homogeneous terrain unit extraction framework based on superpixel segmentation and multiscale region merging (SSM-HTU), in which initial slope units serve as local statistical references and within-slope terrain objects are generated through slope-unit-conditioned morphometric representation, superpixel initialization, distribution-sensitive region merging, and a nested partition hierarchy. The framework was applied to the 5136 km2 Yuqu River Basin in southeastern Tibet. Of 971 expert-interpreted reference HTUs, 680 were reserved for independent geometric evaluation; 1329 historical landslides were additionally used for supplementary spatial association analysis across mapping-unit schemes. Relative to the eCognition Multiresolution Segmentation (MSS) baseline, SSM-HTU showed a slight decrease in Precision from 0.8432 to 0.8340, while Recall (directional reference-object coverage) increased from 0.7615 to 0.8011, area-weighted IoU from 0.6710 to 0.6945, and Boundary F1 at a 12.5 m tolerance from 0.5980 to 0.6810, indicating greater reference-object coverage, spatial overlap, and boundary correspondence without uniform improvement across all geometric metrics. Across four geomorphological zones, area-weighted IoU ranged from 0.671 to 0.724 and Boundary F1 from 0.651 to 0.709, with non-monotonic regional variation. Mapping-unit schemes also yielded factor-dependent spatially stratified associations, underscoring the importance of spatial support in downstream statistical analysis. SSM-HTU therefore provides an object-based mapping framework for representing local within-slope terrain heterogeneity within a hillslope-scale statistical context in deeply incised valleys. Full article
Show Figures

Figure 1

41 pages, 63568 KB  
Article
Vision-Based Automated Inspection of Box Meals for Food Portion Defects and Foreign Object Detection
by Hong-Dar Lin, Guan-Ming Chen and Chou-Hsien Lin
Sensors 2026, 26(17), 5636; https://doi.org/10.3390/s26175636 - 4 Sep 2026
Viewed by 200
Abstract
Automated inspection of prepared meals is important for improving food safety, quality assurance, and production efficiency. However, vision-based inspection remains challenging because boxed meals contain multiple adjacent food items with irregular shapes, varying portion sizes, and visually similar appearances, while oily surfaces may [...] Read more.
Automated inspection of prepared meals is important for improving food safety, quality assurance, and production efficiency. However, vision-based inspection remains challenging because boxed meals contain multiple adjacent food items with irregular shapes, varying portion sizes, and visually similar appearances, while oily surfaces may introduce specular reflections that degrade image quality. This study presents a vision-based framework for multi-object recognition and quantitative defect analysis in complex meal images, with Chinese-style lunch boxes used as representative test samples. The framework integrates region-of-interest (ROI) extraction, fixed-grid regional feature representation, deep neural network (DNN) classification, flood-filling post-processing, and empirical food-quantity thresholds. The lunch-box ROI is first extracted using the Hough transform, followed by median filtering to suppress reflection noise. The ROI is partitioned into 6 × 6 regions, from which the mean and standard deviation of RGB, HSV, and CIE Lab* color components are extracted and classified using a DNN. Flood filling is subsequently applied to refine the classification results, and category-specific empirical thresholds are used to identify missing food items and insufficient portions. Foreign objects are detected as an additional abnormal category. Under the evaluated experimental conditions, the proposed framework achieved an overall image classification rate (CR) of 96.45%, a defective-image detection rate (1−β) of 97.76%, a normal-image false alarm rate (α) of 5.34%, and a defective-image misclassification rate (γ) of 0.82%. Sensitivity experiments involving illumination variation, two lunch-box configurations with different food compositions, and conveyor-based image acquisition further demonstrated the feasibility and stability of the framework under the evaluated laboratory and prototype conditions. These findings support the feasibility of the proposed lightweight framework for vision-based quality inspection of representative boxed meals, while broader validation across meal types, contaminants, and industrial production environments remains necessary. Full article
(This article belongs to the Special Issue Sensing and Imaging for Defect Detection: 2nd Edition)
Show Figures

Figure 1

26 pages, 8441 KB  
Article
Explainable Superpixel-Guided Graph Vision Transformer for Hyperspectral Image Analysis
by Jieli Chen, Kah Phooi Seng, Chee Shen Lim, Li-Minn Ang and Jeremy Smith
Sensors 2026, 26(17), 5513; https://doi.org/10.3390/s26175513 - 31 Aug 2026
Viewed by 188
Abstract
Hyperspectral imaging provides rich spectral–spatial information for fine-grained material discrimination, but effective and interpretable modeling remains challenging because land-cover regions often have irregular spatial structures and class-specific spectral responses. Conventional methods typically rely on fixed grid patches or local neighborhoods, which may not [...] Read more.
Hyperspectral imaging provides rich spectral–spatial information for fine-grained material discrimination, but effective and interpretable modeling remains challenging because land-cover regions often have irregular spatial structures and class-specific spectral responses. Conventional methods typically rely on fixed grid patches or local neighborhoods, which may not align with natural object boundaries, whereas pure superpixel or graph models may lose fine pixel-level details. This paper proposes an explainable superpixel-guided graph vision transformer (ESG-ViT) for hyperspectral image classification and analysis. The proposed framework contains two complementary branches: a graph superpixel vision transformer (GS-ViT) that represents hyperspectral scenes as adaptive superpixel tokens and injects graph topology into self-attention, and a windowed pixel vision transformer (WP-ViT) that preserves dense local spectral–spatial details through efficient local attention. The two representations are adaptively fused for pixel-wise classification. To support interpretability, the model further derives class-wise superpixel relevance maps and spectral channel importance from gradient responses, revealing both the spatial regions and wavelength channels that contribute to each category. Experiments on multiple benchmark hyperspectral datasets demonstrate that the proposed method improves classification accuracy while producing clearer, more human-aligned explanations. Visualization results show that the model highlights meaningful class-related superpixel regions and assigns distinct spectral-channel importance patterns to different classes. These results indicate that the proposed framework provides an accurate and explainable alternative to conventional patch-based transformers for hyperspectral image analysis. Full article
Show Figures

Figure 1

33 pages, 6003 KB  
Article
Unsupervised Gaussian-Noise-Robust Remote Sensing Change Detection via FRFCM-IRM Change Intensity Modeling and SEEDSAM-Constrained HCRF
by Lei Fan, Jiaxin Song, Yikun Li, Yuxi Hu and Yingang Ren
Remote Sens. 2026, 18(16), 2821; https://doi.org/10.3390/rs18162821 - 20 Aug 2026
Viewed by 240
Abstract
Remote sensing change detection technology is widely used in land-use monitoring, urban planning, and disaster assessment. However, during imaging and transmission, bi-temporal remote sensing images are vulnerable to Gaussian noise, which makes it difficult for change detection algorithms to distinguish truly changed areas [...] Read more.
Remote sensing change detection technology is widely used in land-use monitoring, urban planning, and disaster assessment. However, during imaging and transmission, bi-temporal remote sensing images are vulnerable to Gaussian noise, which makes it difficult for change detection algorithms to distinguish truly changed areas from noise-affected regions. To address this issue, this study proposes an unsupervised Gaussian-noise-robust change detection algorithm, termed FRIH-SEEDSAM. The proposed method first applies the Fast and Robust Fuzzy C-Means (FRFCM) algorithm to perform noise-resistant fuzzy clustering on bi-temporal remote sensing images. To establish reliable correspondences between the clustering results, the Integrated Region Matching (IRM) algorithm is introduced to construct weighted matching relationships while reducing the influence of abnormal memberships. The change intensity of spatially corresponding pixels is then calculated to generate a more stable change intensity map. Subsequently, the change intensity map is input into the Hybrid Conditional Random Field (HCRF) to infer pixel-level change labels, where the object potential function is constructed from the segmentation results of the Superpixels Extracted via Energy-Driven Sampling (SEEDS)-guided Segment Anything Model (SEEDSAM), which uses the centroids of the SEEDS superpixel regions as point prompts for the SAM, thereby enhancing change-label consistency within the same changed object region. The experimental results show that the FRIH-SEEDSAM algorithm maintains stable change detection performance across different datasets and under varying Gaussian noise levels. It outperforms the comparison algorithms in terms of several accuracy evaluation indicators, including Kappa and F1. Furthermore, even when the Gaussian noise variance increases to 0.05, Kappa remains at 0.8 or above on multiple dataset images. Full article
Show Figures

Figure 1

24 pages, 1761 KB  
Article
Superpixel-Level Joint-Sparse and Graph-Regularized Framework for Hyperspectral Image Classification
by Tugcan Dundar
Remote Sens. 2026, 18(16), 2699; https://doi.org/10.3390/rs18162699 - 11 Aug 2026
Viewed by 361
Abstract
Hyperspectral image classification (HSIC) remains challenging because high-dimensional spectral signatures must be interpreted together with spatially coherent land-cover structures, particularly when labeled samples are limited. This paper presents a superpixel-based spectral–spatial HSIC method called SJSGR, which combines joint-sparse representation with graph Laplacian regularization. [...] Read more.
Hyperspectral image classification (HSIC) remains challenging because high-dimensional spectral signatures must be interpreted together with spatially coherent land-cover structures, particularly when labeled samples are limited. This paper presents a superpixel-based spectral–spatial HSIC method called SJSGR, which combines joint-sparse representation with graph Laplacian regularization. The HSI is first partitioned into homogeneous superpixel regions so that neighbouring pixels with similar spectral characteristics can be represented jointly rather than independently. For each superpixel, a similarity-aware weighting matrix is constructed between the training dictionary and the superpixel samples, encouraging the coefficient matrix to select more label-consistent and representative training atoms. To further preserve local manifold structure, graph Laplacian regularization is incorporated into the optimization objective, enforcing smooth and coherent representation coefficients among neighboring pixels within each superpixel. The resulting unified formulation integrates spectral correlation, spatial consistency, and local geometric structure, and is solved by the alternating-direction method of multipliers (ADMM). Classification is then performed by assigning each superpixel to the class with the minimum reconstruction error. Experiments are conducted on three real-world HSI datasets called Indian Pines, Pavia University and Fanglu to compare the proposed framework with several sparse representation and graph-based HSIC methods. Experimental results on these datasets reveal the capability of the proposed method, obtaining overall accuracies of 98.12%, 98.04%, and 98.26% under 10%, 1% and 1% labeled samples, respectively. Besides obtaining nearly 1% higher overall accuracy than the compared methods under these low-training-sample distributions, the SJSGR also provided better classification performance even under much more limited numbers of training samples. The findings suggest that superpixel-guided sparse representation with local manifold regularization is a promising direction for effective spectral–spatial HSIC. Full article
Show Figures

Figure 1

22 pages, 7908 KB  
Article
Disentangling Spectrally Similar Urban Vegetation via Semantic Segmentation-Guided Object Analysis and Multi-Periodic Phenological Features
by Chenglong Zhu, Xi Cheng, Tao Liu, Haoyu Wang, Hao Lei, Haiyu Wang and Zhanfeng Shen
Remote Sens. 2026, 18(15), 2623; https://doi.org/10.3390/rs18152623 - 6 Aug 2026
Viewed by 254
Abstract
Fine-grained classification of urban green spaces (UGSs) is important for urban ecological assessment and management but remains challenging because of spectral similarity among vegetation types and inaccurate object delineation in complex urban environments. This study proposes a pixel-to-object framework that combines semantic segmentation-guided [...] Read more.
Fine-grained classification of urban green spaces (UGSs) is important for urban ecological assessment and management but remains challenging because of spectral similarity among vegetation types and inaccurate object delineation in complex urban environments. This study proposes a pixel-to-object framework that combines semantic segmentation-guided object construction with multi-periodic phenological modeling. A semantic green-space mask derived from 0.27 m very-high-resolution imagery constrains superpixel segmentation to generate spatially coherent, boundary-aware green space object-level patches (GSOPs). Pixel-level temporal representations are then derived from Sentinel-2 normalized difference vegetation index (NDVI) time series using TimesNet, aggregated into GSOP-level phenological features, and combined with spatial attributes to classify urban trees, grasslands, and farmlands. Applied to the built-up area of Chengdu, China, the framework achieved an overall accuracy of 91.6%, with F1-scores of 92.5%, 91.9%, and 87.6% for urban trees, grasslands, and farmlands, respectively. Ablation experiments showed that removing phenological features reduced overall accuracy by 13.1 percentage points and decreased the F1-scores of grasslands and farmlands by 16.0 and 23.0 percentage points, respectively. These results demonstrate that semantically constrained object delineation and phenological information jointly reduce boundary fragmentation and improve the discrimination of spectrally similar urban vegetation types. Full article
Show Figures

Figure 1

17 pages, 7391 KB  
Article
Improved YOLOv8 Weed Segmentation Method Based on Dual-ViT
by Weihan Wu, Kaiwen Huang, Haonan Ji, Tujia Chen and Xueshen Chen
Agriculture 2026, 16(15), 1675; https://doi.org/10.3390/agriculture16151675 - 3 Aug 2026
Viewed by 385
Abstract
To address inaccurate weed segmentation under crop overlap, occlusion, and complex field backgrounds, this study developed a combined method integrating DViT-YOLOv8-seg with confidence-guided SLIC voting. The dataset contained 1872 field images (800 × 600 pixels) of Guangzhou soft-stem lettuce and four common weed [...] Read more.
To address inaccurate weed segmentation under crop overlap, occlusion, and complex field backgrounds, this study developed a combined method integrating DViT-YOLOv8-seg with confidence-guided SLIC voting. The dataset contained 1872 field images (800 × 600 pixels) of Guangzhou soft-stem lettuce and four common weed species: Eleusine indica, Digitaria sanguinalis, Portulaca oleracea, and Amaranthus blitum. All weed species were merged into one weed class, while lettuce, soil, and other field regions were treated as non-weed. Real-ESRGAN and data augmentation enhanced the training samples; Dual-ViT strengthened global–local feature interaction; GSConv reduced redundant computation; BiFPN improved multi-scale fusion; and SLIC refined ambiguous boundaries. After super-resolution preprocessing and three-fold expansion, baseline mPA increased by 10.9 percentage points. The improved network achieved 88.3% mPA at 7.9 GFLOPs, corresponding to +3.6 percentage points and -1.0 GFLOPs relative to the baseline. SLIC voting increased FWIoU to 95.6%, 4.1 percentage points above the network without SLIC. Compared with YOLOv5-seg and Fast-SCNN, mPA improved by 2.0 and 6.5 percentage points, respectively; GFLOPs were 87.7% and 95.5% lower than those of YOLOv5-seg and DeepLabv3+, respectively. The method therefore provides a favorable trade-off between segmentation accuracy and theoretical network computation for complex lettuce field imagery. Full article
(This article belongs to the Section Crop Protection, Diseases, Pests and Weeds)
Show Figures

Figure 1

21 pages, 1153 KB  
Article
A Comparative Analysis of Gradient-Based, Edge-Based, and Segmentation-Based Data Augmentation Methods for Early Diagnosis of Alzheimer’s Disease Using Neuroimaging Modalities and Deep Learning
by Muhammad Dawood, Usman Rasheed, Waqas Ahmad, Ahsan Bin Tufail and Afnan Albahli
Symmetry 2026, 18(8), 1254; https://doi.org/10.3390/sym18081254 - 23 Jul 2026
Viewed by 408
Abstract
Alzheimer’s disease (AD) is a neurodegenerative disorder that causes progressive damage to brain neurons, leading to declines in cognitive and behavioral abilities. This deterioration often results in changes in personality and increasing difficulty in thinking and memory over time. Although there is no [...] Read more.
Alzheimer’s disease (AD) is a neurodegenerative disorder that causes progressive damage to brain neurons, leading to declines in cognitive and behavioral abilities. This deterioration often results in changes in personality and increasing difficulty in thinking and memory over time. Although there is no cure, early detection is crucial as it allows for more effective management and care. Advances in deep learning have significantly improved the accuracy of brain scan analysis for diagnostic purposes. In this study, we utilized the publicly available Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset consisting of subjects diagnosed with AD, Mild Cognitive Impairment (MCI), and Normal Control (NC). Each participant has either Magnetic Resonance Imaging (MRI) or Positron Emission Tomography (PET) neuroimaging data, ensuring representation across heterogeneous modalities. The research focuses on comparing gradient-based, edge-based, and segmentation-based data augmentation techniques for early AD detection using neuroimaging and deep learning approaches, particularly 3D Convolutional Neural Networks (3D CNNs). Various augmentation methods were applied, including directional gradient, azimuth gradient direction, numerical gradient, Sobel horizontal edge filter, superpixel oversegmentation, and Canny edge detection. These techniques are evaluated in both binary and multiclass classification tasks involving MRI and PET scans. The results indicate that optimal performance varied depending on the task and modality. For PET-based classification, directional gradient performed the best for AD vs. NC binary classification, achieving an accuracy of 87.24%, while Canny edge detection was most effective for AD vs. MCI binary classification and AD-MCI-NC multiclass classification tasks, achieving accuracies of 72.77% and 59.04%, respectively. For MCI vs. NC, the best result (accuracy = 64.32%) is achieved by combining azimuth gradient direction with Sobel filtering. In contrast, for the MRI-based AD vs. NC classification task, the highest performance is achieved without applying augmentation (balanced accuracy = 60.90%). This research confirms the efficacy of data augmentation methods in the early diagnosis of AD in clinical settings. Full article
(This article belongs to the Section A: Computer Science)
Show Figures

Figure 1

23 pages, 15531 KB  
Article
Anchor-Level Spectral–Spatial Graph Clustering for Hyperspectral Images
by Chaodie Liu, Jianxiong Luo, Fei Li, Qianyao Qiang and Feiping Nie
Remote Sens. 2026, 18(13), 2172; https://doi.org/10.3390/rs18132172 - 3 Jul 2026
Viewed by 375
Abstract
Hyperspectral image (HSI) clustering aims to partition pixels into distinct clusters by leveraging spectral and spatial features, thereby providing crucial support for the interpretation and information extraction of hyperspectral data. However, due to high spectral variability, complex spatial distribution, and noise interference, HSI [...] Read more.
Hyperspectral image (HSI) clustering aims to partition pixels into distinct clusters by leveraging spectral and spatial features, thereby providing crucial support for the interpretation and information extraction of hyperspectral data. However, due to high spectral variability, complex spatial distribution, and noise interference, HSI clustering still faces considerable challenges. Graph-based clustering represents a prominent learning framework and achieves competitive performance on HSI analysis. However, most existing methods ignore spatial information and suffer from high computational cost, rendering them incapable of effectively dealing with large-scale HSIs. To address the aforementioned challenges, this paper proposes an anchor-level spectral–spatial graph clustering (ASSGC) model for HSIs. The proposed ASSGC employs a band-wise median strategy within each superpixel to generate representative anchors to suppress noise and outlier effects. A novel distance metric is designed to integrate spectral features and spatial positions to effectively identify neighbors and construct a spectral–spatial joint affinity matrix at the anchor-level, thereby reducing computational burden and memory consumption. Subsequently, spectral clustering is applied to obtain anchor labels, which are propagated to the corresponding superpixels to achieve full-image clustering. Experiments on four HSI datasets yield ACC of 64.13% on Indian Pines, 71.33% on Pavia University, 87.86% on Salinas, and 99.23% on Salinas A, demonstrating that the proposed ASSGC outperforms several existing state-of-the-art methods while maintaining low time complexity. Full article
Show Figures

Figure 1

30 pages, 14827 KB  
Article
A Superpixel-Guided Spectral–Spatial Fusion Network for Hyperspectral Scene Classification
by Yan Wang, Xinyao Li, Baisen Liu, Jianxin Chen and Weili Kong
Remote Sens. 2026, 18(13), 2124; https://doi.org/10.3390/rs18132124 - 1 Jul 2026
Viewed by 503
Abstract
In recent years, research on remote sensing scene classification (RSSC) has mainly focused on high-resolution imagery, which provides limited spectral information, whereas hyperspectral imaging (HSI) offers richer cues about material properties and compositional structure. Despite its potential, hyperspectral scene classification (HSI-SC) remains challenging [...] Read more.
In recent years, research on remote sensing scene classification (RSSC) has mainly focused on high-resolution imagery, which provides limited spectral information, whereas hyperspectral imaging (HSI) offers richer cues about material properties and compositional structure. Despite its potential, hyperspectral scene classification (HSI-SC) remains challenging because pixel- or patch-based representations fail to preserve spatial structures and regional boundaries. In addition, labeled hyperspectral samples are often scarce, making it difficult to learn stable class-discriminative representations from high-dimensional spectral observations. To address these issues, this paper proposes a dual-branch fusion framework. Superpixels are used to aggregate high-dimensional spectral signals into compact, boundary-aware tokens. The spectral branch is initialized with pretrained model weights and further adapted via a lightweight adaptation strategy for efficient transfer under limited supervision. In parallel, a pseudo-RGB spatial branch complements structural and textural information. Spectral and spatial features are fused additively to generate a more discriminative scene representation. Experimental results demonstrate that the proposed method outperforms compared hyperspectral scene classification approaches. Full article
Show Figures

Figure 1

16 pages, 52629 KB  
Article
Automatic Segmentation and Recognition of the Microstructure of High-Strength Low-Alloy Steel
by Lu Wang, Ziying Ren, Baoyu Song, Bing Wang, Qiaochuan Chen, Jingjing Wang, Tianpeng Zhou and Yuexing Han
Materials 2026, 19(12), 2554; https://doi.org/10.3390/ma19122554 - 12 Jun 2026
Viewed by 349
Abstract
Metallographic microstructure analysis is essential for understanding the evolution of steel microstructures during heat treatment and mechanical processing. However, accurate analysis of optical micrographs remains difficult because of blurred grain boundaries, grayscale inhomogeneity within grains, and irregular grain morphologies. To address these issues, [...] Read more.
Metallographic microstructure analysis is essential for understanding the evolution of steel microstructures during heat treatment and mechanical processing. However, accurate analysis of optical micrographs remains difficult because of blurred grain boundaries, grayscale inhomogeneity within grains, and irregular grain morphologies. To address these issues, this work proposes an automated metallographic image-processing method based on superpixels, DPSS (dual-phase steel segmentation), with the main contribution focused on microstructure segmentation. First, image contrast and boundary visibility are enhanced by edge detection and sharpening. Then, superpixel segmentation is combined with extracted edge information to improve boundary localization and preserve irregular grain morphology, enabling more complete extraction of grain or particle regions from optical images. The proposed method is validated on optical micrographs of Mn-Si low-alloy steel, and the results show that it provides more accurate and complete segmentation than conventional ImageJ (Version: 1.54f)-based processing. Based on the segmented regions, a lightweight neural network is further used for phase identification. The final classification recognition accuracy can reach 99.91%. This classification result serves to demonstrate that the improved segmentation results can provide more reliable inputs for subsequent microstructure recognition. Overall, the proposed method offers an effective and automated solution for metallographic image segmentation and supports more accurate downstream phase analysis. Full article
(This article belongs to the Section Metals and Alloys)
Show Figures

Figure 1

25 pages, 2289 KB  
Article
Superpixel Random Selection Random Walk Multi-Branch Depthwise Convolutional Neural Network for Hyperspectral Image Classification
by Kai Zhang, Xinwei Jiang and Zhihua Cai
Sensors 2026, 26(11), 3558; https://doi.org/10.3390/s26113558 - 3 Jun 2026
Viewed by 458
Abstract
Convolutional neural networks (CNNs) and training-free CNN variants have been successfully applied to hyperspectral image (HSI) processing and analysis. Training-free CNNs have shown promising feature extraction performance, which could effectively address the issue of typical CNNs being highly parameterized; however, inevitable noise and [...] Read more.
Convolutional neural networks (CNNs) and training-free CNN variants have been successfully applied to hyperspectral image (HSI) processing and analysis. Training-free CNNs have shown promising feature extraction performance, which could effectively address the issue of typical CNNs being highly parameterized; however, inevitable noise and redundancy in the randomly selected training-free convolutional kernels often leads to unsatisfactory performance. To address this issue, we propose Superpixel Random Selection Random Walk Multi-Branch Depthwise Convolutional Neural Network (SRSRWMD-CNN). Specifically, we propose a novel training-free convolutional neural network characterized by inter-layer multi-scale integration and intra-layer grouping. Various superpixels groups are first generated through multi-scale superpixel segmentation algorithms, then the predetermined number of superpixels are randomly sampled from these groups to serve as training-free convolution kernels. This mechanism enables adaptive computation of HSI feature maps without costly model training in the feature extraction stage, allowing the network to effectively capture a multi-scale spectral–spatial feature representation. Additionally, we propose a multi-branch depthwise convolution strategy that mitigates feature learning errors while significantly enhancing feature representation capabilities. A random walk strategy is employed to expand the receptive field and enhance the robustness of the training-free convolution kernels. Finally, the multi-scale spectral–spatial features are concatenated with the multiple convolutional stages to fuse salient shallow and deep features for accurate HSI classification. Extensive experiments demonstrate that the proposed method achieves superior performance compared to state-of-the-art algorithms. Full article
(This article belongs to the Special Issue High-Frequency Spectroscopy and Imaging: Techniques and Applications)
Show Figures

Figure 1

29 pages, 22126 KB  
Article
Mask-Guided Feature Routing and Adaptive Context Modeling for Wide-FoV UAV Object Detection in IoT Remote Sensing
by Lingfan Wu, Yachun Feng, Hong Zhang and Yawei Li
Remote Sens. 2026, 18(11), 1753; https://doi.org/10.3390/rs18111753 - 30 May 2026
Cited by 1 | Viewed by 562
Abstract
Object detection in wide-field-of-view (wide-FoV) unmanned aerial vehicle (UAV) imagery for Internet of Things (IoT) remote sensing applications requires accurate recognition of tiny objects under severe background redundancy and extreme scale variation. As the field of view expands, conventional dense detectors tend to [...] Read more.
Object detection in wide-field-of-view (wide-FoV) unmanned aerial vehicle (UAV) imagery for Internet of Things (IoT) remote sensing applications requires accurate recognition of tiny objects under severe background redundancy and extreme scale variation. As the field of view expands, conventional dense detectors tend to waste substantial computation on non-informative regions, while feature downsampling and static receptive fields often cause the dilution of foreground information and scale confusion. To address these issues, we propose MFRC-Det, a unified framework built upon two complementary principles: mask-guided feature routing and adaptive context modeling. Specifically, a Superpixel-Masking Generator (SP-Masker) is introduced to estimate an image-space soft foreground prior by comparing Simple Linear Iterative Clustering (SLIC) superpixel histograms with a peripheral background reference, propagating the resulting scores on a superpixel adjacency graph, and projecting the refined region-level scores back to a pixel-level routing mask. Guided by these priors, a Greedy-Cutter (G-Cutter) converts dense feature maps into compact, foreground-focused patches without repeated backbone evaluation on cropped image regions, thereby reducing redundant background computation while preserving local structural coherence. On top of the retained regions, an Adaptive Receptive-field Selection Network (ARSNet) aggregates multi-scale contextual responses from several learnable receptive-field candidate branches. ARSNet predicts spatial selection weights conditioned on the input features, allowing each location to emphasize a suitable receptive-field response for object representation. Experimental results on VisDrone-DET and UAVDT demonstrate that MFRC-Det achieves competitive detection accuracy with favorable computational efficiency. Specifically, MFRC-Det obtains 36.1% AP, 60.4% AP50, and 38.5 FPS on VisDrone-DET and 21.3% AP, 36.8% AP50, and 37.4 FPS on UAVDT. These results validate the effectiveness of mask-guided feature routing and adaptive context modeling for wide-FoV UAV object detection and suggest their potential value for computation-efficient aerial perception in IoT remote sensing applications. Full article
Show Figures

Figure 1

31 pages, 17767 KB  
Article
Integration of Superpixel Segmentation, Convolutional Neural Networks and Vision Transformers for Automatic Benthic Habitats Classification
by Hassan Mohamed and Kazuo Nadaoka
Remote Sens. 2026, 18(11), 1711; https://doi.org/10.3390/rs18111711 - 26 May 2026
Viewed by 513
Abstract
Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) have achieved significant success in various computer vision applications, including the classification of high-resolution imagery. However, a notable limitation of these deep learning approaches is their tendency to inadequately preserve the precise edges and shapes [...] Read more.
Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) have achieved significant success in various computer vision applications, including the classification of high-resolution imagery. However, a notable limitation of these deep learning approaches is their tendency to inadequately preserve the precise edges and shapes of target objects. In contrast, Object-Based Image Analysis (OBIA) offers a methodology that emphasizes the preservation of object boundaries by segmenting images into meaningful objects. Combining CNNs and ViTs with OBIA leverages the feature extraction capabilities of these deep learning algorithms and the boundary-preserving advantages of OBIA, leading to enhanced classification accuracy and improved delineation of object boundaries in high-resolution images. Still, the main challenge for combining these methods lies in effectively aligning the irregularly shaped image objects produced by OBIA with the regular image patches required by CNNs and ViT architectures. In this study, we propose a novel approach that integrates superpixel segmentation with CNNs and ViTs for the automatic classification of benthic habitats using high-resolution orthomosaic images. Initially, the Simple Linear Iterative Clustering (SLIC) algorithm was applied to segment the high-resolution orthomosaic images into superpixels. Subsequently, the central points of the resulting superpixels were utilized to generate square image patches. These patches performed as inputs for ConvNeXt-Base and EfficientNet-B0 pre-trained CNNs to extract fine-grained features and Dinov2 ViTs to extract high-level features. Then, a Support Vector Machine (SVM) classifier was trained using these attributes to classify benthic habitats. Eventually, the classification label derived from the SVM defined the class of each superpixel segment. This method achieved an average overall accuracy of 0.96 in classifying benthic habitats. Overall, we demonstrate that combining CNNs, ViTs, and superpixel segmentation is an effective approach to benthic habitats classification, providing accurate high-resolution maps of heterogeneous reef environments. Full article
(This article belongs to the Section Ocean Remote Sensing)
Show Figures

Figure 1

Back to TopTop