Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (978)

Search Parameters:
Keywords = blur dataset

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
37 pages, 3015 KB  
Article
Deepfake Detection via Frequency-Aware Vision Transformer and Bidirectional Cross-Attention Fusion with Post-Processing Robustness
by Wasin Alkishri, Shahid Kamal and Jabar Yousif
Information 2026, 17(9), 819; https://doi.org/10.3390/info17090819 - 26 Aug 2026
Abstract
Today, the use of increasingly ubiquitous synthetic media, or ‘deepfakes’, has become a risk to online trust, information integrity and individual security and is being created by artificial intelligence (AI). The current approaches are mainly based on either spatial features of CNNs or [...] Read more.
Today, the use of increasingly ubiquitous synthetic media, or ‘deepfakes’, has become a risk to online trust, information integrity and individual security and is being created by artificial intelligence (AI). The current approaches are mainly based on either spatial features of CNNs or high-level semantic representations of Vision Transformer; both have major drawbacks in effectively leveraging multi-domain forensic cues. This paper presents FAViT (Frequency-Aware Vision Transformer), a hybrid architecture capable of jointly utilizing spatial- and frequency-domain forensic information by the means of a bidirectional cross-attention fusion scheme. We use an 11-channel forensic tensor in each face image (including per-channel Fast Fourier Transform (FFT) magnitude maps, Discrete Wavelet Transform (DWT) sub-bands, channel noise residual maps, Sobel gradient magnitude and channels of Error Level Analysis (ELA)). A Frequency Branch CNN processes this multi-domain tensor and the original RGB image is encoded with a pretrained ViT-B/16 spatial branch. The two streams are combined through the bidirectional cross-attention which allows the model to localize both spatial and spectral manipulation artifacts. We also present an adversarial cleaning simulation pipeline which partitions the training process with five post-processing attack methods, namely GFPGAN neural face restoration, learned autoencoder cleaning, etc., to increase resistance to real-world forensic defenses. Tests of FaceForensics++ C23 (7926 images, consisting of four manipulation types) show that FAViT attains F1-score of 86.22, AUC-ROC of 94.26 and accuracy of 85.55 on the held-out test set. The strength analysis of 21 attack conditions shows that the max degradation in AUC is 30.3, with specific strengths in GFPGAN restoration (AUC = 98.51). Robustness is evaluated based on 21 post-processing attack cases that include JPEG compression, Gaussian blurring, down-sampling, and GFDGAN neural-based restoration; it should be noted that robustness against gradient-based adaptive attacks requires additional attention. Testing on the CIFAKE and Celeb-DF v2 datasets reveals some limitations of domain generalization. Full article
(This article belongs to the Special Issue Artificial Intelligence for Signal, Image and Video Processing)
Show Figures

Graphical abstract

16 pages, 9671 KB  
Article
A Lightweight Semantic Segmentation for Terrestrial Oil Spill Detection
by Keyong Shao and Honglian Cao
Appl. Sci. 2026, 16(17), 8458; https://doi.org/10.3390/app16178458 - 25 Aug 2026
Abstract
Accurate and timely monitoring of terrestrial oil spills is vital for ecological conservation and safe oilfield operations. To address the challenges of segmenting terrestrial oil spills in UAV remote sensing imagery, including blurred boundaries, irregular shapes, and complex background interference, we propose Fluid-SegFormer, [...] Read more.
Accurate and timely monitoring of terrestrial oil spills is vital for ecological conservation and safe oilfield operations. To address the challenges of segmenting terrestrial oil spills in UAV remote sensing imagery, including blurred boundaries, irregular shapes, and complex background interference, we propose Fluid-SegFormer, a fluid-aware semantic segmentation model based on the lightweight SegFormer architecture. Fluid-SegFormer employs a Mix Transformer (MiT-B0) encoder to extract hierarchical multi-scale features and integrates a hierarchical fluid-aware optimization framework. Specifically, the Local Noise Gating (LNG) module suppresses background noise, the Horizontal–Vertical Perception Attention (HVPA) module enhances the structural representation of irregular oil spill regions, and the Fluid Soft Boundary Refinement Decoder (FSBRD) recovers fine boundary details. Experiments on a newly constructed high-resolution UAV terrestrial oil spill dataset demonstrate that Fluid-SegFormer achieves an mIoU of 87.84%, an IoU of 77.56%, and a Precision of 91.42%, effectively balancing computational efficiency and segmentation accuracy. These results demonstrate the potential of Fluid-SegFormer for practical deployment in UAV-based oil spill monitoring on edge devices. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

24 pages, 40484 KB  
Article
BC-GECO2: A Coarse and Fine Aggregate Segmentation and Counting Method for Hydraulic Concrete with Dense Depth Feature Fusion and Edge Enhancement
by Jiandong Wu, Baijing Wu, Jianwei Deng, Long Ma, Shuhong Liu and Shufan Zhang
Infrastructures 2026, 11(9), 297; https://doi.org/10.3390/infrastructures11090297 - 25 Aug 2026
Abstract
To reduce aggregate gradation counting errors caused by over-segmentation and under-segmentation of stacked and clustered aggregates with mixed types and diverse spatial distributions in hydraulic concrete, this study proposes BC-GECO2, a coarse and fine aggregate segmentation and counting method. Firstly, a BAHiera feature [...] Read more.
To reduce aggregate gradation counting errors caused by over-segmentation and under-segmentation of stacked and clustered aggregates with mixed types and diverse spatial distributions in hydraulic concrete, this study proposes BC-GECO2, a coarse and fine aggregate segmentation and counting method. Firstly, a BAHiera feature extraction network is designed to extract multi-scale deep features through edge-aware attention. In addition, a DFG-Edge module is developed to enhance the boundary features of densely distributed aggregates by integrating wavelet transform with a gated fusion mechanism, thereby alleviating the loss of small aggregate features during downsampling. Secondly, a CSFM-GFFCA module is constructed, in which a dual-branch structure is employed to adaptively fuse adjacent-scale features, strengthen the edge responses of densely distributed small aggregates, and enhance cross-layer feature interaction. Finally, a joint optimization function combining Focal loss and counting loss is established to guide the model toward hard-to-classify pixels, especially boundary pixels, thereby improving segmentation integrity and counting accuracy. Experiments conducted on an aggregate dataset collected from practical construction sites show that, compared with the baseline GECO2 model, the proposed method improves the average segmentation IoU, Dice, and BIoU by 2.92%, 5.04%, and 2.83%, respectively, while reducing the average counting MAE and RMSE by 6.92 and 15.65, respectively. Moreover, BC-GECO2 exhibits superior robustness and generalization capability under different stacking densities and blurred-boundary scenarios, providing technical support for the intelligent development of rapid concrete gradation detection. Full article
(This article belongs to the Section Infrastructures Materials and Constructions)
Show Figures

Figure 1

28 pages, 8603 KB  
Article
Event-Guided Image Reconstruction for Nighttime Dynamic Scenes
by Qingjiao Meng, Ji Li and Yan Jin
J. Imaging 2026, 12(9), 399; https://doi.org/10.3390/jimaging12090399 - 23 Aug 2026
Viewed by 66
Abstract
Image reconstruction in nighttime dynamic scenes is challenged by low illumination, long exposure, rapid camera or object motion, and sensor noise. Conventional RGB cameras, therefore, struggle to recover both sufficient brightness and clear structural details in nighttime dynamic scenes. To address this problem, [...] Read more.
Image reconstruction in nighttime dynamic scenes is challenged by low illumination, long exposure, rapid camera or object motion, and sensor noise. Conventional RGB cameras, therefore, struggle to recover both sufficient brightness and clear structural details in nighttime dynamic scenes. To address this problem, we propose an event-guided image reconstruction method for nighttime dynamic visual perception. The method constructs a multi-channel event voxel representation by jointly encoding event count, event intensity, timestamp distribution, and blurred-frame intensity priors. A parameter-efficient local–global reconstruction network is then designed to restore fine-grained textures and model holistic structures. In addition, edge-alignment and blur-alignment constraints are introduced to improve geometric consistency and imaging plausibility. Experiments on the HQF and REDS datasets show that the proposed method outperforms existing methods in the MSE, PSNR, and SSIM. Compared with DeblurSR, it reduces the MSE by 14.81% on HQF and 10.00% on REDS, while improving the PSNR by 1.603 dB and 1.053 dB, respectively. Qualitative results further show sharper edges, lower structural errors, and better edge consistency. Low illumination, dynamic blur, rapid brightness variation, and event noise are also common degradation factors in nighttime UAV imaging, making the investigated problem technically relevant to that setting. However, because neither REDS nor HQF was acquired during an actual UAV flight, the reported results establish benchmark-level reconstruction performance rather than UAV-specific operational effectiveness. Full article
Show Figures

Figure 1

23 pages, 10390 KB  
Article
SSDM-Net: A Spatial–Spectral Distillation Mamba Network for Hyperspectral Image Super-Resolution
by Anjie Chen, Shunli Liu, Qiao Luo, Zhengyong Feng and Weichao Yang
Electronics 2026, 15(17), 3768; https://doi.org/10.3390/electronics15173768 - 22 Aug 2026
Viewed by 109
Abstract
Hyperspectral image super-resolution (HSI SR) focuses on enhancing the spatial resolution of HSIs while preserving their inherent spectral information. Existing single-image HSI SR methods still suffer from blurred spatial edges and spectral distortion. Although numerous spatial–spectral enhancement networks can enhance spatial–spectral feature extraction, [...] Read more.
Hyperspectral image super-resolution (HSI SR) focuses on enhancing the spatial resolution of HSIs while preserving their inherent spectral information. Existing single-image HSI SR methods still suffer from blurred spatial edges and spectral distortion. Although numerous spatial–spectral enhancement networks can enhance spatial–spectral feature extraction, they often lead to a cumbersome network architecture. To address these issues, we propose a Spatial–Spectral Distillation Mamba Network, called SSDM-Net, for HSI SR, which contains a main reconstruction branch and two training-only auxiliary branches for spatial and spectral knowledge distillation. Specifically, the spatial and spectral auxiliary branches, which are utilized exclusively during training, provide edge-aware guidance and capture spectral correlations, respectively. During training, the spatial–spectral knowledge is transferred to the main branch. During inference, the auxiliary branches are removed, improving reconstruction quality without extra computational burden. In the main branch, a Mamba-based spatial–spectral global enhancement module processes spatial and latent inter-channel sequences using selective scanning whose cost is linear in the processed sequence lengths when the feature dimensions are fixed. In addition, a dynamic loss weighting strategy is developed to balance reconstruction, distillation, and auxiliary losses during optimization. Comprehensive experiments conducted on the CAVE and Houston datasets with three scale factors demonstrate that SSDM-Net produces more accurate reconstruction results than existing representative HSI SR methods. Cross-dataset experiments on the Harvard dataset further suggest that the method can maintain competitive reconstruction performance under the evaluated cross-dataset settings. Full article
(This article belongs to the Topic Computational Intelligence in Remote Sensing: 3rd Edition)
Show Figures

Figure 1

22 pages, 3833 KB  
Article
Structural Consistency-Aware LiDAR Super-Resolution Method
by Jun Zeng, Chunqiu Xia, Hongwei Zhang, Hongchao Gao, Ziyang Wang and Mingjun Li
Photonics 2026, 13(9), 802; https://doi.org/10.3390/photonics13090802 - 22 Aug 2026
Viewed by 124
Abstract
Existing LiDAR super-resolution methods primarily aim to increase point cloud density or improve coordinate reconstruction accuracy. However, they tend to introduce blurred edges and distorted planar surfaces during reconstruction, making it difficult to preserve the consistency of local scene geometry. To address this [...] Read more.
Existing LiDAR super-resolution methods primarily aim to increase point cloud density or improve coordinate reconstruction accuracy. However, they tend to introduce blurred edges and distorted planar surfaces during reconstruction, making it difficult to preserve the consistency of local scene geometry. To address this issue, this paper proposes a structural consistency-aware LiDAR super-resolution method that aims to preserve the local geometric relationships of the reconstructed point cloud with respect to the ground-truth point cloud in edge and planar regions. Specifically, complementary observations from adjacent frames are first fused using multi-scale dilated convolutions. An anisotropic Swin Transformer and a Coordinate-Aware Structure Enhancement (CASE) module are then employed to accommodate the horizontally dense and vertically sparse sampling pattern of LiDAR, strengthen long-range geometric modeling, and reduce the loss of critical structural information. During training, a local curvature-based structural consistency loss is designed to separately constrain edge sharpness and planar smoothness. During inference, prediction uncertainty and point cloud height are combined to adaptively remove low-confidence points, further improving the geometric reliability of the reconstructed point cloud. Experiments on the KITTI dataset show that the proposed method achieves an MAE of 0.4916 and an IoU of 0.4633, outperforming the representative comparison methods on both metrics. When the reconstructed point clouds are applied to A-LOAM, the average RTE and RRE values are reduced by 34.6% and 31.2%, respectively. In addition, experiments on the self-collected CSU-SLAM dataset provide preliminary evidence of the applicability of the proposed method to indoor and outdoor scenes under a different LiDAR configuration. Full article
(This article belongs to the Special Issue Computational Imaging)
Show Figures

Figure 1

60 pages, 2148 KB  
Review
Atmospheric Turbulence Mitigation in the Deep Learning Era: A Critical Review from CNNs and GANs to Transformers, Diffusion, Mamba, and Physics-Informed Models
by Nurul Jannah, Teddy Surya Gunawan, Mira Kartiwi, Nadirah Abdul Rahim and Ali Sophian
Big Data Cogn. Comput. 2026, 10(8), 282; https://doi.org/10.3390/bdcc10080282 - 21 Aug 2026
Viewed by 112
Abstract
Anyone who has watched a distant scene shimmer above hot pavement has seen atmospheric turbulence destroy image detail. In long-range imaging, turbulence produces spatially varying blur, geometric warping, scintillation, and temporal instability. Recovering the underlying scene is therefore an ill-posed inverse problem, and [...] Read more.
Anyone who has watched a distant scene shimmer above hot pavement has seen atmospheric turbulence destroy image detail. In long-range imaging, turbulence produces spatially varying blur, geometric warping, scintillation, and temporal instability. Recovering the underlying scene is therefore an ill-posed inverse problem, and learned priors must compensate for distortions that simplified optical models capture only partially. We use turbulence mitigation as the umbrella term for all countermeasures and turbulence restoration for its computational core, the estimation of a clean image from degraded observations. From a cognitive-computing perspective, mitigation is not merely image enhancement. It is an uncertainty-constrained visual inference problem in which an intelligent system must reconstruct, interpret, and act on observations relayed through a stochastic physical channel. This critical review examines how the deep learning era has reshaped turbulence mitigation, with physics as the foundation for understanding degradation and designing inductive biases. It traces the architectural progression from convolutional and adversarial networks to Transformers, denoising diffusion models, Mamba and other state-space architectures, and physics-informed frameworks. For each family, we ask a common question: How does it treat the aleatoric uncertainty intrinsic to a random optical channel and the epistemic uncertainty introduced by scarce and simulator-dominated training data? The review also analyzes datasets, simulation strategies, loss functions, and evaluation metrics, and it separates the small body of shared-protocol benchmark evidence from the far larger body of self-reported results that cannot be compared across studies. Persistent obstacles include the synthetic-to-real domain gap, the scarcity of paired real turbulence data, the mismatch between fidelity metrics and downstream task performance, and the computational cost that limits operational deployment. We close with an AI-centered agenda in which uncertainty quantification stands alongside domain adaptation as a first-order priority. Full article
(This article belongs to the Special Issue Machine Learning and Image Processing: Applications and Challenges)
Show Figures

Figure 1

31 pages, 22406 KB  
Article
HiFi-Det: Collaborative Multi-Scale Frequency-Domain Feature Optimization for Crown-of-Thorns Starfish Detection in Complex Underwater Environments
by Sirong Qian, Yuewen Huang, Meng Wang, Houlei Jia, Xiaoyong Mei and Fudan Zheng
J. Mar. Sci. Eng. 2026, 14(16), 1523; https://doi.org/10.3390/jmse14161523 - 17 Aug 2026
Viewed by 189
Abstract
Outbreaks of the Crown-of-Thorns Starfish (COTS, Acanthaster spp.) are a leading biological driver of coral cover loss, making timely and accurate population monitoring essential for reef management. Conventional diver-based surveys are labor-intensive and prone to missed detections, motivating automated detection from underwater imagery. [...] Read more.
Outbreaks of the Crown-of-Thorns Starfish (COTS, Acanthaster spp.) are a leading biological driver of coral cover loss, making timely and accurate population monitoring essential for reef management. Conventional diver-based surveys are labor-intensive and prone to missed detections, motivating automated detection from underwater imagery. However, COTS detection in complex underwater scenes still faces three major challenges. First, COTS individuals are often very small and carry limited discriminative information, making them inherently difficult to detect. Second, low underwater contrast and complex coral textures blur target boundaries and cause targets to be easily confused with the background. Third, ecological monitoring values recall more highly than precision—missing a COTS individual is far more costly than a false alarm—yet the recall of existing detectors remains insufficient. To address these challenges, we propose HiFi-Det (High-resolution Frequency-integration Detector), a collaborative multi-scale frequency-domain feature optimization method built on YOLO11. HiFi-Det integrates three complementary enhancements: a high-resolution detection branch that strengthens feature representation for small targets; wavelet transform convolution (WTConv) modules in the backbone and neck that apply band-separated processing in the wavelet domain to improve discrimination of COTS targets from low-contrast, textured coral backgrounds; and a WIoUv3 bounding box regression loss that dynamically focuses on ordinary-quality samples to improve recall while maintaining precision. On the public Great Barrier Reef dataset, HiFi-Det attains 81.02% F2 and 87.54% mAP@50, surpassing the YOLO11 baseline by 3.00% and 2.57%, respectively, while keeping the parameter count essentially unchanged relative to the YOLO11s baseline (within 3%), so that the accuracy gains are obtained without inflating model size. Ablation studies confirm the synergy of the three components: the high-resolution branch preserves spatial details, WTConv suppresses background textures, and WIoUv3 further curbs false positives while sustaining high recall. Applying the same recipe to a larger YOLO11m backbone yields HiFi-Det-m, which likewise improves over that backbone in both F2 and recall, indicating that the approach is a transferable recipe rather than a single fixed architecture. These results show that task-specific architectural and training designs can effectively adapt generic detectors to the demands of underwater ecological monitoring. Full article
(This article belongs to the Section Marine Biology)
Show Figures

Figure 1

40 pages, 89575 KB  
Article
BFMambaNet: Boundary-Frequency-Guided Global Semantic Mamba Network for Fine-Grained Camellia oleifera Leaf Disease Segmentation
by Xuanhao Li, Fulin Su, Yongming Yan, Shaofeng Peng, Lin Li, Fangying Wan and Ruifeng Liu
Plants 2026, 15(16), 2493; https://doi.org/10.3390/plants15162493 - 17 Aug 2026
Viewed by 171
Abstract
Camellia oleifera leaf disease segmentation under natural field conditions is important for precision plant protection but remains challenging because lesions often show small target areas, blurred boundaries, uneven illumination, complex backgrounds, and coexisting symptoms. To address these problems, this paper proposes BFMambaNet, a [...] Read more.
Camellia oleifera leaf disease segmentation under natural field conditions is important for precision plant protection but remains challenging because lesions often show small target areas, blurred boundaries, uneven illumination, complex backgrounds, and coexisting symptoms. To address these problems, this paper proposes BFMambaNet, a Boundary-Frequency-guided Global Semantic Mamba Network for fine-grained disease segmentation. The model adopts an encoder-decoder framework and introduces a Global Semantic Mamba-based spatial selective feature modeling block to capture long-range lesion context and reduce semantic confusion. A gated wavelet spatial enhancement block is further designed to strengthen high-frequency boundary details while suppressing noisy responses. During training, boundary-frequency auxiliary supervision guides contour localization and pathological texture recovery without additional manual boundary labels. A reinforcement-learning-guided adaptive loss controller adjusts class-wise reweighting factors and loss-component weights according to the training state, improving optimization stability. A pixel-level dataset containing 1400 images and seven disease categories was constructed for evaluation. Experimental results show that BFMambaNet achieves 92.39% Precision, 91.43% Recall, 91.26% Dice, and 85.46% mIoU, outperforming representative CNN-based, Transformer-based, and Mamba-based models. Evaluations on environmental subsets confirm superior robustness, outperforming VMamba by 3.70% mIoU under uneven illumination, 3.55% mIoU under complex backgrounds, and 5.10% mIoU under coexisting symptoms. Cross-dataset validation on Apple leaf diseases further proves its generalization with 3.39% mIoU and 3.84% Dice improvements over U-Mamba, while maintaining a competitive inference speed of 30 FPS. Qualitative results also show clearer boundaries, fewer missed small lesions, and more stable predictions in complex field scenarios. Full article
(This article belongs to the Special Issue Advances in Artificial Intelligence for Plant Research—2nd Edition)
Show Figures

Figure 1

23 pages, 5177 KB  
Article
Tea Shoot Category Detection Based on UAV Remote Sensing
by Zhaoxia Liu, Meng Tan, Baijuan Wang, Xiaoxue Guo, Jing Zhao and Shihao Zhang
Agronomy 2026, 16(16), 1576; https://doi.org/10.3390/agronomy16161576 - 17 Aug 2026
Viewed by 191
Abstract
This study proposes an object detection network named HR-YOLOv10-S for UAV remote sensing-based detection of three tea shoot categories. The network aims to solve key problems in UAV images under complex tea plantation environments, including large target scale changes, excessive background information, and [...] Read more.
This study proposes an object detection network named HR-YOLOv10-S for UAV remote sensing-based detection of three tea shoot categories. The network aims to solve key problems in UAV images under complex tea plantation environments, including large target scale changes, excessive background information, and motion blur. Based on the YOLOv10 framework, the proposed network incorporates Shape Weights and Scale Adjustment Factors into the bounding-box regression loss to jointly account for target shape and scale variations, thereby improving the geometric consistency and localization accuracy between predicted and ground-truth bounding boxes. This mechanism improves the matching ability between predicted bounding boxes and real targets. The Rectangular Self Calibrated Module is introduced to improve the network ability to model complex spatial structure information, so it can capture target edge features accurately and improve localization. The Histogram Transformer is added to make full use of the global statistical distribution features of images and reduce interference from complex background noise. Experimental results on the test dataset showed that HR-YOLOv10-S achieved a Precision of 91.11%, a Recall of 88.91%, an mAP@0.5 of 93.33%, and an F1-score of 89.99%. Compared with the baseline YOLOv10 model, these metrics increased by 8.68, 4.03, 5.18, and 6.36 percentage points, respectively. Under the same experimental settings, HR-YOLOv10-S also achieved higher values for these four metrics than SSD, CornerNet, and RT-DETR. These findings suggest that the proposed model can improve the detection performance of three tea shoot categories under conditions involving target-scale variation, background interference, and partial occlusion. Therefore, HR-YOLOv10-S provides a potentially useful approach for UAV remote sensing-based detection of three tea shoot categories and smart tea plantation management, although further validation on independent datasets is needed to assess its broader generalizability. Full article
Show Figures

Figure 1

24 pages, 13299 KB  
Article
BCNet: Boundary-Constrained Remote Sensing Change Detection Network Based on Vision Foundation Models
by Shenbo Liu, Dongxue Zhao, Huang He and Lijun Tang
Remote Sens. 2026, 18(16), 2760; https://doi.org/10.3390/rs18162760 - 15 Aug 2026
Viewed by 213
Abstract
Limited by the diversity and complexity of real-world scenes, existing remote sensing change detection methods often suffer from insufficient fine-grained semantic understanding and blurred boundaries of change targets. To address these issues, this paper proposes a boundary-constrained remote sensing change detection network based [...] Read more.
Limited by the diversity and complexity of real-world scenes, existing remote sensing change detection methods often suffer from insufficient fine-grained semantic understanding and blurred boundaries of change targets. To address these issues, this paper proposes a boundary-constrained remote sensing change detection network based on vision foundation models (BCNet). BCNet employs a differential modeling approach and multi-branch guidance mechanism to design a differential detail enhancement module, amplifying fine-grained semantic information. Through cross-layer feature alignment, stepwise fusion, and edge-sensitive modeling, it constructs a multi-scale edge enhancement module that enhances perception of minute variations and edge details, fully leveraging the universal semantic representation capabilities of the vision foundation model. In addition, an edge feature constraint mechanism is introduced that applies dual guidance and supervision during the feature fusion and output stages. This mechanism achieves refined delineation of change region boundaries and significantly mitigates the issue of boundary blurring. Experimental results on four mainstream datasets, namely LEVIR-CD, WHU-CD, NJDS and MSRS-CD, demonstrate that BCNet outperforms 13 state-of-the-art methods in terms of key metrics including F1 and IoU. Against the best VFM-based baseline, BCNet obtains F1 score gains of 0.21%, 0.71%, 6.33% and 0.63% on the above four datasets. Specifically, the proposed method exhibits superior detection accuracy and edge detail preservation capabilities in complex regions. Full article
(This article belongs to the Section Remote Sensing Image Processing)
Show Figures

Figure 1

18 pages, 2452 KB  
Article
DoubleTransU-Net: Enhancing Teeth Segmentation in Panoramic Dental X-Ray Images
by Manal Touahri and Aissam Berrahou
Algorithms 2026, 19(8), 685; https://doi.org/10.3390/a19080685 - 15 Aug 2026
Viewed by 179
Abstract
Accurate teeth segmentation in panoramic dental radiographs remains a challenging task due to high image noise, low contrast, the similarity in intensity between teeth and surrounding tissues, and blurred tooth boundaries. To address these challenges, we propose DoubleTransU-Net, a dual-stage hybrid CNN–Transformer architecture [...] Read more.
Accurate teeth segmentation in panoramic dental radiographs remains a challenging task due to high image noise, low contrast, the similarity in intensity between teeth and surrounding tissues, and blurred tooth boundaries. To address these challenges, we propose DoubleTransU-Net, a dual-stage hybrid CNN–Transformer architecture that combines progressive segmentation refinement with global contextual feature learning. The first stage generates an initial tooth segmentation, while the second stage progressively refines ambiguous tooth regions to improve boundary delineation and segmentation accuracy. In addition, Atrous Spatial Pyramid Pooling (ASPP) modules capture multi-scale contextual information, whereas squeeze-and-excitation (SE) blocks enhance discriminative feature representations through channel-wise feature recalibration. The proposed model was evaluated on two public panoramic dental X-ray datasets, UFBA-UESC (1500 images) and Tufts (1000 images), and compared against several state-of-the-art segmentation models, including U-Net, DoubleU-Net, Attention U-Net, TransUNet, and DeepLabv3+. On the UFBA-UESC dataset, DoubleTransU-Net achieved an Accuracy of 95.54%, a Dice coefficient of 93.72%, an Intersection over Union (IoU) of 88.18%, a Precision of 93.57%, and a Recall of 94.24%. On the Tufts dataset, it achieved an Accuracy of 91.83%, a Dice coefficient of 92.93%, an IoU of 86.80%, a Precision of 92.27%, and a Recall of 93.97%. These results demonstrate that DoubleTransU-Net consistently outperforms existing state-of-the-art segmentation methods while exhibiting strong robustness and generalization across different panoramic dental datasets, highlighting its effectiveness for tooth semantic segmentation in panoramic dental X-ray images. Full article
Show Figures

Figure 1

25 pages, 4272 KB  
Article
WPSeg-Net: A Boundary-Guided Dual-Branch Network for Molten Pool Segmentation
by Xin Heng, Yongjing Wang, Wanqiang Zhang and Jin Huang
Appl. Sci. 2026, 16(16), 8110; https://doi.org/10.3390/app16168110 - 14 Aug 2026
Viewed by 151
Abstract
The geometric morphology and dynamic evolution of the molten pool are key indicators of heat input, metal transfer, and solidification behavior during welding, making them critical for quality monitoring and process control. To address the challenges of molten pool image segmentation under complex [...] Read more.
The geometric morphology and dynamic evolution of the molten pool are key indicators of heat input, metal transfer, and solidification behavior during welding, making them critical for quality monitoring and process control. To address the challenges of molten pool image segmentation under complex conditions, including strong arc interference, blurred boundaries, small foreground regions, and significant shape variations, a MIG welding-based visual acquisition system is established to construct a multi-condition dataset. Based on this dataset, an efficient boundary-guided dual-branch network, WPSeg-Net, is proposed. The network employs ResNet-34 as the encoder, integrates a lightweight Transformer module for global context modeling, and adopts BiFPN for multi-scale feature fusion. A region branch and a boundary branch with a mutual guidance mechanism are further designed to improve segmentation accuracy and boundary refinement. Experimental results demonstrate that WPSeg-Net achieves an IoU of 92.83% and a Dice score of 96.28%, improving by 0.98% and 0.58% over DeepLabV3+, respectively. Meanwhile, the proposed method reduces the parameter size from 22.43 MB to 21.78 MB and decreases FLOPs from 62.25 G to 31.90 G, achieving a better balance between segmentation performance and computational efficiency. Full article
Show Figures

Figure 1

30 pages, 1840 KB  
Article
Weak Ridge-Flow Prior-Guided Fingerprint Reconstruction Under Severe Degradation
by Haiyong Xie, Lin Wang, Yonghao Dai and Yunqian Cheng
Computers 2026, 15(8), 527; https://doi.org/10.3390/computers15080527 - 14 Aug 2026
Viewed by 187
Abstract
Fingerprint enhancement plays an important role in recovering identity-related ridge structures from degraded fingerprints. However, existing methods primarily focus on local texture restoration and may struggle to preserve ridge continuity and structural consistency under severe degradation conditions, including ridge fragmentation, diffusion blur, and [...] Read more.
Fingerprint enhancement plays an important role in recovering identity-related ridge structures from degraded fingerprints. However, existing methods primarily focus on local texture restoration and may struggle to preserve ridge continuity and structural consistency under severe degradation conditions, including ridge fragmentation, diffusion blur, and partial information loss. In this paper, we observe that degraded fingerprints may retain incomplete ridge-flow information that can provide useful structural guidance for fingerprint reconstruction. Based on this observation, we propose a conditional generative adversarial network guided by a weak ridge-flow prior (WRP-cGAN) for degraded fingerprint enhancement. The proposed method treats the estimated ridge-flow information as a weak structural prior rather than an exact structural constraint and introduces prior-conditioned feature modulation to adaptively incorporate structural cues during reconstruction. The framework is jointly optimized using adversarial, image-space reconstruction, ridge-flow orientation-consistency, and gradient-consistency losses to improve ridge continuity, structural coherence, and local detail preservation. On the NIST SD301-derived test set, the proposed method increases the median NFIQ2 score from 9 to 44, improves the minutiae-restoration F1-score from 0.2507 to 0.5426, and increases the SourceAFIS Rank-1 identification rate from 39% to 86%. An additional qualitative evaluation on FVC2004 DB1 provides preliminary evidence of cross-dataset transferability without fine-tuning. These results suggest that weak ridge-flow priors provide useful structural guidance for degraded fingerprint reconstruction and improve recognition-oriented fingerprint quality under the degradation conditions considered in this study. Full article
(This article belongs to the Section AI-Driven Innovations)
Show Figures

Figure 1

25 pages, 23777 KB  
Article
Medical Textile Stain Detection Based on Chemically Enhanced Visualization and Deep Semantic Segmentation
by Wenjie Min, Junfeng He, Zhenping Wan, Jinde Chen, Zhixiang Zou and Yuandong Mo
J. Imaging 2026, 12(8), 380; https://doi.org/10.3390/jimaging12080380 - 13 Aug 2026
Viewed by 210
Abstract
Pre-wash sorting of medical textiles is essential for hospital infection control, yet accurate stain detection remains challenging because visually apparent stains often have blurred boundaries, whereas dried urine stains lack distinguishable optical features. This study proposes a medical textile stain detection method integrating [...] Read more.
Pre-wash sorting of medical textiles is essential for hospital infection control, yet accurate stain detection remains challenging because visually apparent stains often have blurred boundaries, whereas dried urine stains lack distinguishable optical features. This study proposes a medical textile stain detection method integrating chemically enhanced visualization with deep semantic segmentation. Dimethylaminocinnamaldehyde (DMACA) was used to convert latent urine stains into chemically developed stains with orange–red visual features. Based on the spatial color difference ΔE in the L*a*b* color space, 0.0183 mol/L was selected as the most suitable DMACA concentration among those tested. A dataset of 1974 images was constructed, including blood stains, chemically developed urine stains, medication stains, and uncontaminated textiles. A cascaded preprocessing strategy was applied to enhance stain boundaries and suppress textile texture noise, after which an Enhanced semantic segmentation model incorporating residual feature extraction, multiscale feature fusion, and transfer learning was used for pixel-level recognition. The IoU values for blood stains, chemically developed urine stains, and medication stains were 88.11%, 82.67%, and 89.62%, respectively. The average time required for image preprocessing and network inference was 15.39 ms per image. An input-level ablation comparison showed that DMACA-based color development increased the urine-stain IoU from 3.07% to 86.23%, demonstrating its substantial contribution to latent urine-stain detection. These results support the feasibility of integrating front-end chemical feature enhancement with back-end semantic segmentation for multiclass medical textile stain recognition under the current experimental conditions. Full article
(This article belongs to the Section Image and Video Processing)
Show Figures

Figure 1

Back to TopTop