Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (30)

Search Parameters:
Keywords = super-resolution semantic segmentation

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
26 pages, 18794 KB  
Article
DWFSeg: A Dynamic Multiscale Feature Fusion and Dual Attention-Enhanced Network for High-Precision Water Body Segmentation Based on Super-Resolution Remote Sensing Imagery
by Ziwei Li, Bingjie Liang, Jianzhong Guo, Ning Li, Weiran Luo, Baowei Zhang, Jiali Guo, Weizhen Zhang, Yan Zhou, Yuezhen Guo and Yishan Li
Remote Sens. 2026, 18(14), 2271; https://doi.org/10.3390/rs18142271 - 8 Jul 2026
Viewed by 232
Abstract
Remote sensing imagery provides a primary data source for large-scale surface water body monitoring, which is crucial for quantifying climate-related hydrological impacts, supporting flood control, and sustaining integrated water resource management. However, remote sensing images generally face the trade-off between spatial resolution and [...] Read more.
Remote sensing imagery provides a primary data source for large-scale surface water body monitoring, which is crucial for quantifying climate-related hydrological impacts, supporting flood control, and sustaining integrated water resource management. However, remote sensing images generally face the trade-off between spatial resolution and temporal coverage. To address this issue, the Real-ESRGAN super-resolution algorithm is employed to reconstruct temporally continuous, wide-coverage medium-resolution imagery to a 2.5 m resolution, effectively improving its capability to identify sub-pixel river boundaries. Water body segmentation (WBS) is an effective method for fine-detail surface water extraction. Nonetheless, when applied in complex hydrological environments, it still faces several limitations, such as ambiguous delineation of land–water boundaries and the difficulty in capturing multiscale water body characteristics. To address these issues, a Dynamic Weight Fusion SegFormer (DWFSeg) network is constructed, integrating a MixVision Transformer (MVT) encoder with a multiscale decoding architecture. Specifically, a Dynamic Multiscale Feature Fusion (DMFF) mechanism is proposed, which adaptively assigns semantic-guided fusion weights to multiscale feature water bodies. Furthermore, the Dual Attention-Enhanced (DAE) module strengthens discriminative essential features and suppresses background noise in both channel and spatial dimensions. Evaluated on a self-constructed super-resolution imagery dataset (SID) and the public GID, DWFSeg achieves overall accuracies of 98.08% and 96.14%, respectively. It outperforms representative benchmark models across multiple quantitative metrics, while maintaining competitive inference efficiency and favorable segmentation stability. Ablation studies verify the effectiveness and necessity of each proposed component. The presented network provides a reliable technical solution and supports refined water resource evaluation and sustainable watershed management. Full article
Show Figures

Figure 1

19 pages, 5482 KB  
Article
MAD-SAR: A Multi-Agent Agentic Engineering Framework for Landslide Detection Using Sentinel-1 SAR Imagery
by Kohei Arai
Information 2026, 17(6), 597; https://doi.org/10.3390/info17060597 - 15 Jun 2026
Viewed by 472
Abstract
Rapid and accurate detection of landslide-affected areas is critical for disaster response and risk mitigation. Sentinel-1 SAR imagery offers all-weather, day-and-night observation capability, but existing deep learning approaches treat landslide detection as a single-pass segmentation problem, which limits performance in complex terrain where [...] Read more.
Rapid and accurate detection of landslide-affected areas is critical for disaster response and risk mitigation. Sentinel-1 SAR imagery offers all-weather, day-and-night observation capability, but existing deep learning approaches treat landslide detection as a single-pass segmentation problem, which limits performance in complex terrain where backscatter changes are confounded by soil moisture, surface roughness, urban double bounce, shadow, and layover effects. MAD-SAR, a rule-based agentic framework that coordinates anomaly detection, super-resolution, object detection, and semantic segmentation under a planning orchestrator and a physics-aware validation engine is proposed. The orchestrator selects specialist modules, their execution order, and the number of refinement iterations according to a scene complexity score computed from SAR-derived statistics. The physics-aware validation engine cross-checks every candidate detection against backscatter change thresholds, DEM-derived slope constraints, and radar geometry masks before any detection is committed to the output. MAD-SAR is evaluated on three Japanese disaster datasets: Hiroshima 2018, Kumamoto 2016, and Ibaraki 2019. On the held-out Ibaraki test event, the framework achieves an F1-score of 0.863 and IoU of 0.759, outperforming all baselines and reducing false alarms by 45% relative to standalone SegFormer. Ablation results confirm that each module contributes to the final performance. These results suggest that multi-module orchestration with embedded physical validation can meaningfully improve SAR-based landslide mapping, though broader validation across regions, sensor configurations, and failure mechanisms remains necessary. Full article
(This article belongs to the Special Issue AI-Based Image Processing and Computer Vision, 2nd Edition)
Show Figures

Figure 1

19 pages, 2788 KB  
Article
Universal Image Segmentation with Arbitrary Granularity for Efficient Pest Monitoring
by L. Minh Dang, Sufyan Danish, Muhammad Fayaz, Asma Khan, Gul E. Arzu, Lilia Tightiz, Hyoung-Kyu Song and Hyeonjoon Moon
Horticulturae 2025, 11(12), 1462; https://doi.org/10.3390/horticulturae11121462 - 3 Dec 2025
Viewed by 880
Abstract
Accurate and timely pest monitoring is essential for sustainable agriculture and effective crop protection. While recent deep learning-based pest recognition systems have significantly improved accuracy, they are typically trained for fixed label sets and narrowly defined tasks. In this paper, we present RefPestSeg, [...] Read more.
Accurate and timely pest monitoring is essential for sustainable agriculture and effective crop protection. While recent deep learning-based pest recognition systems have significantly improved accuracy, they are typically trained for fixed label sets and narrowly defined tasks. In this paper, we present RefPestSeg, a universal, language-promptable segmentation model specifically designed for pest monitoring. RefPestSeg can segment targets at any semantic level, such as species, genus, life stage, or damage type, conditioned on flexible natural language instructions. The model adopts a symmetric architecture with self-attention and cross-attention mechanisms to tightly align visual features with language embeddings in a unified feature space. To further enhance performance in challenging field conditions, we integrate an optimized super-resolution module to improve image quality and employ diverse data augmentation strategies to enrich the training distribution. A lightweight postprocessing step refines segmentation masks by suppressing highly overlapping regions and removing noise blobs introduced by cluttered backgrounds. Extensive experiments on a challenging pest dataset show that RefPestSeg achieves an Intersection over Union (IoU) of 69.08 while maintaining robustness in real-world scenarios. By enabling language-guided pest segmentation, RefPestSeg advances toward more intelligent, adaptable monitoring systems that can respond to real-time agricultural demands without costly model retraining. Full article
Show Figures

Figure 1

23 pages, 11997 KB  
Article
Deep Learning-Driven Automatic Segmentation of Weeds and Crops in UAV Imagery
by Jianghan Tao, Qian Qiao, Jian Song, Shan Sun, Yijia Chen, Qingyang Wu, Yongying Liu, Feng Xue, Hao Wu and Fan Zhao
Sensors 2025, 25(21), 6576; https://doi.org/10.3390/s25216576 - 25 Oct 2025
Cited by 22 | Viewed by 2352
Abstract
Accurate segmentation of crops and weeds is essential for enhancing crop yield, optimizing herbicide usage, and mitigating environmental impacts. Traditional weed management practices, such as manual weeding or broad-spectrum herbicide application, are labor-intensive, environmentally harmful, and economically inefficient. In response, this study introduces [...] Read more.
Accurate segmentation of crops and weeds is essential for enhancing crop yield, optimizing herbicide usage, and mitigating environmental impacts. Traditional weed management practices, such as manual weeding or broad-spectrum herbicide application, are labor-intensive, environmentally harmful, and economically inefficient. In response, this study introduces a novel precision agriculture framework integrating Unmanned Aerial Vehicle (UAV)-based remote sensing with advanced deep learning techniques, combining Super-Resolution Reconstruction (SRR) and semantic segmentation. This study is the first to integrate UAV-based SRR and semantic segmentation for tobacco fields, systematically evaluate recent Transformer and Mamba-based models alongside traditional CNNs, and release an annotated dataset that not only ensures reproducibility but also provides a resource for the research community to develop and benchmark future models. Initially, SRR enhanced the resolution of low-quality UAV imagery, significantly improving detailed feature extraction. Subsequently, to identify the optimal segmentation model for the proposed framework, semantic segmentation models incorporating CNN, Transformer, and Mamba architectures were used to differentiate crops from weeds. Among evaluated SRR methods, RCAN achieved the optimal reconstruction performance, reaching a Peak Signal-to-Noise Ratio (PSNR) of 24.98 dB and a Structural Similarity Index (SSIM) of 69.48%. In semantic segmentation, the ensemble model integrating Transformer (DPT with DINOv2) and Mamba-based architectures achieved the highest mean Intersection over Union (mIoU) of 90.75%, demonstrating superior robustness across diverse field conditions. Additionally, comprehensive experiments quantified the impact of magnification factors, Gaussian blur, and Gaussian noise, identifying an optimal magnification factor of 4×, proving that the method was robust to common environmental disturbances at optimal parameters. Overall, this research established an efficient, precise framework for crop cultivation management, offering valuable insights for precision agriculture and sustainable farming practices. Full article
(This article belongs to the Special Issue Smart Sensing and Control for Autonomous Intelligent Unmanned Systems)
Show Figures

Figure 1

23 pages, 1108 KB  
Article
HADQ-Net: A Power-Efficient and Hardware-Adaptive Deep Convolutional Neural Network Translator Based on Quantization-Aware Training for Hardware Accelerators
by Can Uğur Oflamaz and Müştak Erhan Yalçın
Electronics 2025, 14(18), 3686; https://doi.org/10.3390/electronics14183686 - 18 Sep 2025
Cited by 3 | Viewed by 1937
Abstract
With the increasing demand for implementing deep-learning models on devices on resource-constrained devices, the development of power-efficient neural networks has become imperative. This paper introduces HADQ-Net, a novel framework for optimizing deep convolutional neural networks (CNNs) through Quantization-Aware Training (QAT). By compressing 32-bit [...] Read more.
With the increasing demand for implementing deep-learning models on devices on resource-constrained devices, the development of power-efficient neural networks has become imperative. This paper introduces HADQ-Net, a novel framework for optimizing deep convolutional neural networks (CNNs) through Quantization-Aware Training (QAT). By compressing 32-bit floating-point (FP32) precision weights and activation values to lower bit-widths, HADQ-Net significantly reduces memory footprint and computational complexity while maintaining high accuracy. We propose adaptive quantization limits based on the statistical properties of each layer or channel, coupled with normalization techniques, to enhance quantization efficiency and accuracy. The framework includes algorithms for QAT, quantized convolution, and quantized inference, enabling efficient deployment of deep CNN models on edge devices. Extensive experiments across tasks such as super-resolution, classification, object detection, and semantic segmentation demonstrate the trade-offs between accuracy, model size, and computational efficiency under various quantization levels. Our results highlight the superiority of QAT over post-training quantization methods and underscore the impact of quantization types on model performance. HADQ-Net achieves significant reductions in memory footprint, computational complexity, and energy consumption, making it ideal for resource-constrained environments without sacrificing performance. Full article
Show Figures

Figure 1

19 pages, 2806 KB  
Article
SP-IGAN: An Improved GAN Framework for Effective Utilization of Semantic Priors in Real-World Image Super-Resolution
by Meng Wang, Zhengnan Li, Haipeng Liu, Zhaoyu Chen and Kewei Cai
Entropy 2025, 27(4), 414; https://doi.org/10.3390/e27040414 - 11 Apr 2025
Cited by 6 | Viewed by 1638
Abstract
Single-image super-resolution (SISR) based on GANs has achieved significant progress. However, these methods still face challenges when reconstructing locally consistent textures due to a lack of semantic understanding of image categories. This highlights the necessity of focusing on contextual information comprehension and the [...] Read more.
Single-image super-resolution (SISR) based on GANs has achieved significant progress. However, these methods still face challenges when reconstructing locally consistent textures due to a lack of semantic understanding of image categories. This highlights the necessity of focusing on contextual information comprehension and the acquisition of high-frequency details in model design. To address this issue, we propose the Semantic Prior-Improved GAN (SP-IGAN) framework, which incorporates additional contextual semantic information into the Real-ESRGAN model. The framework consists of two branches. The main branch introduces a Graph Convolutional Channel Attention (GCCA) module to transform channel dependencies into adjacency relationships between feature vertices, thereby enhancing pixel associations. The auxiliary branch strengthens the correlation between semantic category information and regional textures in the Residual-in-Residual Dense Block (RRDB) module. The auxiliary branch employs a pretrained segmentation model to accurately extract regional semantic information from the input low-resolution image. This information is injected into the RRDB module through Spatial Feature Transform (SFT) layers, generating more accurate and semantically consistent texture details. Additionally, a wavelet loss is incorporated into the loss function to capture high-frequency details that are often overlooked. The experimental results demonstrate that the proposed SP-IGAN outperforms state-of-the-art (SOTA) super-resolution models across multiple public datasets. For the X4 super-resolution task, SP-IGAN achieves a 0.55 dB improvement in Peak Signal-to-Noise Ratio (PSNR) and a 0.0363 increase in Structural Similarity Index (SSIM) compared to the baseline model Real-ESRGAN. Full article
Show Figures

Figure 1

15 pages, 5686 KB  
Article
Integrating Super-Resolution with Deep Learning for Enhanced Periodontal Bone Loss Segmentation in Panoramic Radiographs
by Vungsovanreach Kong, Eun Young Lee, Kyung Ah Kim and Ho Sun Shon
Bioengineering 2024, 11(11), 1130; https://doi.org/10.3390/bioengineering11111130 - 8 Nov 2024
Cited by 7 | Viewed by 2516
Abstract
Periodontal disease is a widespread global health concern that necessitates an accurate diagnosis for effective treatment. Traditional diagnostic methods based on panoramic radiographs are often limited by subjective evaluation and low-resolution imaging, leading to suboptimal precision. This study presents an approach that integrates [...] Read more.
Periodontal disease is a widespread global health concern that necessitates an accurate diagnosis for effective treatment. Traditional diagnostic methods based on panoramic radiographs are often limited by subjective evaluation and low-resolution imaging, leading to suboptimal precision. This study presents an approach that integrates Super-Resolution Generative Adversarial Networks (SRGANs) with deep learning-based segmentation models to enhance the segmentation of periodontal bone loss (PBL) areas on panoramic radiographs. By transforming low-resolution images into high-resolution versions, the proposed method reveals critical anatomical details that are essential for precise diagnostics. The effectiveness of this approach was validated using datasets from the Chungbuk National University Hospital and the Kaggle data portal, demonstrating significant improvements in both image resolution and segmentation accuracy. The SRGAN model, evaluated using the Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index Measure (SSIM) metrics, achieved a PSNR of 30.10 dB and an SSIM of 0.878, indicating high fidelity in image reconstruction. When applied to semantic segmentation using a U-Net architecture, the enhanced images resulted in a dice similarity coefficient (DSC) of 0.91 and an intersection over union (IoU) of 84.9%, compared with 0.72 DSC and 65.4% IoU for native low-resolution images. These results underscore the potential of SRGAN-enhanced imaging to improve PBL area segmentation and suggest broader applications in medical imaging, where enhanced image clarity is crucial for diagnostic accuracy. This study also highlights the importance of further research to expand the dataset diversity and incorporate clinical validation to fully realize the benefits of super-resolution techniques in medical diagnostics. Full article
Show Figures

Figure 1

18 pages, 12417 KB  
Article
An Object-Aware Network Embedding Deep Superpixel for Semantic Segmentation of Remote Sensing Images
by Ziran Ye, Yue Lin, Baiyu Dong, Xiangfeng Tan, Mengdi Dai and Dedong Kong
Remote Sens. 2024, 16(20), 3805; https://doi.org/10.3390/rs16203805 - 13 Oct 2024
Cited by 7 | Viewed by 3399
Abstract
Semantic segmentation forms the foundation for understanding very high resolution (VHR) remote sensing images, with extensive demand and practical application value. The convolutional neural networks (CNNs), known for their prowess in hierarchical feature representation, have dominated the field of semantic image segmentation. Recently, [...] Read more.
Semantic segmentation forms the foundation for understanding very high resolution (VHR) remote sensing images, with extensive demand and practical application value. The convolutional neural networks (CNNs), known for their prowess in hierarchical feature representation, have dominated the field of semantic image segmentation. Recently, hierarchical vision transformers such as Swin have also shown excellent performance for semantic segmentation tasks. However, the hierarchical structure enlarges the receptive field to accumulate features and inevitably leads to the blurring of object boundaries. We introduce a novel object-aware network, Embedding deep SuperPixel, for VHR image semantic segmentation called ESPNet, which integrates advanced ConvNeXt and the learnable superpixel algorithm. Specifically, the developed task-oriented superpixel generation module can refine the results of the semantic segmentation branch by preserving object boundaries. This study reveals the capability of utilizing deep convolutional neural networks to accomplish both superpixel generation and semantic segmentation of VHR images within an integrated end-to-end framework. The proposed method achieved mIoU scores of 84.32, 90.13, and 55.73 on the Vaihingen, Potsdam, and LoveDA datasets, respectively. These results indicate that our model surpasses the current advanced methods, thus demonstrating the effectiveness of the proposed scheme. Full article
(This article belongs to the Special Issue Advances in Deep Learning Approaches in Remote Sensing)
Show Figures

Figure 1

19 pages, 6395 KB  
Article
Dmg2Former-AR: Vision Transformers with Adaptive Rescaling for High-Resolution Structural Visual Inspection
by Kareem Eltouny, Seyedomid Sajedi and Xiao Liang
Sensors 2024, 24(18), 6007; https://doi.org/10.3390/s24186007 - 17 Sep 2024
Cited by 6 | Viewed by 3317
Abstract
Developments in drones and imaging hardware technology have opened up countless possibilities for enhancing structural condition assessments and visual inspections. However, processing the inspection images requires considerable work hours, leading to delays in the assessment process. This study presents a semantic segmentation architecture [...] Read more.
Developments in drones and imaging hardware technology have opened up countless possibilities for enhancing structural condition assessments and visual inspections. However, processing the inspection images requires considerable work hours, leading to delays in the assessment process. This study presents a semantic segmentation architecture that integrates vision transformers with Laplacian pyramid scaling networks, enabling rapid and accurate pixel-level damage detection. Unlike conventional methods that often lose critical details through resampling or cropping high-resolution images, our approach preserves essential inspection-related information such as microcracks and edges using non-uniform image rescaling networks. This innovation allows for detailed damage identification of high-resolution images while significantly reducing the computational demands. Our main contributions in this study are: (1) proposing two rescaling networks that together allow for processing high-resolution images while significantly reducing the computational demands; and (2) proposing Dmg2Former, a low-resolution segmentation network with a Swin Transformer backbone that leverages the saved computational resources to produce detailed visual inspection masks. We validate our method through a series of experiments on publicly available visual inspection datasets, addressing various tasks such as crack detection and material identification. Finally, we examine the computational efficiency of the adaptive rescalers in terms of multiply–accumulate operations and GPU-memory requirements. Full article
(This article belongs to the Special Issue Feature Papers in Fault Diagnosis & Sensors 2024)
Show Figures

Figure 1

14 pages, 1877 KB  
Article
Multi-Resolution Learning and Semantic Edge Enhancement for Super-Resolution Semantic Segmentation of Urban Scene Images
by Ruijun Shu and Shengjie Zhao
Sensors 2024, 24(14), 4522; https://doi.org/10.3390/s24144522 - 12 Jul 2024
Cited by 4 | Viewed by 2758
Abstract
Super-resolution semantic segmentation (SRSS) is a technique that aims to obtain high-resolution semantic segmentation results based on resolution-reduced input images. SRSS can significantly reduce computational cost and enable efficient, high-resolution semantic segmentation on mobile devices with limited resources. Some of the existing methods [...] Read more.
Super-resolution semantic segmentation (SRSS) is a technique that aims to obtain high-resolution semantic segmentation results based on resolution-reduced input images. SRSS can significantly reduce computational cost and enable efficient, high-resolution semantic segmentation on mobile devices with limited resources. Some of the existing methods require modifications of the original semantic segmentation network structure or add additional and complicated processing modules, which limits the flexibility of actual deployment. Furthermore, the lack of detailed information in the low-resolution input image renders existing methods susceptible to misdetection at the semantic edges. To address the above problems, we propose a simple but effective framework called multi-resolution learning and semantic edge enhancement-based super-resolution semantic segmentation (MS-SRSS) which can be applied to any existing encoder-decoder based semantic segmentation network. Specifically, a multi-resolution learning mechanism (MRL) is proposed that enables the feature encoder of the semantic segmentation network to improve its feature extraction ability. Furthermore, we introduce a semantic edge enhancement loss (SEE) to alleviate the false detection at the semantic edges. We conduct extensive experiments on the three challenging benchmarks, Cityscapes, Pascal Context, and Pascal VOC 2012, to verify the effectiveness of our proposed MS-SRSS method. The experimental results show that, compared with the existing methods, our method can obtain the new state-of-the-art semantic segmentation performance. Full article
(This article belongs to the Special Issue Advances in Automated Driving: Sensing and Control)
Show Figures

Figure 1

18 pages, 5061 KB  
Article
Generating 10-Meter Resolution Land Use and Land Cover Products Using Historical Landsat Archive Based on Super Resolution Guided Semantic Segmentation Network
by Dawei Wen, Shihao Zhu, Yuan Tian, Xuehua Guan and Yang Lu
Remote Sens. 2024, 16(12), 2248; https://doi.org/10.3390/rs16122248 - 20 Jun 2024
Cited by 3 | Viewed by 4608
Abstract
Generating high-resolution land cover maps using relatively lower-resolution remote sensing images is of great importance for subtle analysis. However, the domain gap between real lower-resolution and synthetic images has not been permanently resolved. Furthermore, super-resolution information is not fully exploited in semantic segmentation [...] Read more.
Generating high-resolution land cover maps using relatively lower-resolution remote sensing images is of great importance for subtle analysis. However, the domain gap between real lower-resolution and synthetic images has not been permanently resolved. Furthermore, super-resolution information is not fully exploited in semantic segmentation models. By solving the aforementioned issues, a deeply fused super resolution guided semantic segmentation network using 30 m Landsat images is proposed. A large-scale dataset comprising 10 m Sentinel-2, 30 m Landsat-8 images, and 10 m European Space Agency (ESA) Land Cover Product is introduced, facilitating model training and evaluation across diverse real-world scenarios. The proposed Deeply Fused Super Resolution Guided Semantic Segmentation Network (DFSRSSN) combines a Super Resolution Module (SRResNet) and a Semantic Segmentation Module (CRFFNet). SRResNet enhances spatial resolution, while CRFFNet leverages super-resolution information for finer-grained land cover classification. Experimental results demonstrate the superior performance of the proposed method in five different testing datasets, achieving 68.17–83.29% and 39.55–75.92% for overall accuracy and kappa, respectively. When compared to ResUnet with up-sampling block, increases of 2.16–34.27% and 8.32–43.97% were observed for overall accuracy and kappa, respectively. Moreover, we proposed a relative drop rate of accuracy metrics to evaluate the transferability. The model exhibits improved spatial transferability, demonstrating its effectiveness in generating accurate land cover maps for different cities. Multi-temporal analysis reveals the potential of the proposed method for studying land cover and land use changes over time. In addition, a comparison of the state-of-the-art full semantic segmentation models indicates that spatial details are fully exploited and presented in semantic segmentation results by the proposed method. Full article
(This article belongs to the Special Issue AI-Driven Mapping Using Remote Sensing Data)
Show Figures

Figure 1

31 pages, 19170 KB  
Article
Semantic-Aware Fusion Network Based on Super-Resolution
by Lingfeng Xu and Qiang Zou
Sensors 2024, 24(11), 3665; https://doi.org/10.3390/s24113665 - 5 Jun 2024
Cited by 6 | Viewed by 3303
Abstract
The aim of infrared and visible image fusion is to generate a fused image that not only contains salient targets and rich texture details, but also facilitates high-level vision tasks. However, due to the hardware limitations of digital cameras and other devices, there [...] Read more.
The aim of infrared and visible image fusion is to generate a fused image that not only contains salient targets and rich texture details, but also facilitates high-level vision tasks. However, due to the hardware limitations of digital cameras and other devices, there are more low-resolution images in the existing datasets, and low-resolution images are often accompanied by the problem of losing details and structural information. At the same time, existing fusion algorithms focus too much on the visual quality of the fused images, while ignoring the requirements of high-level vision tasks. To address the above challenges, in this paper, we skillfully unite the super-resolution network, fusion network and segmentation network, and propose a super-resolution-based semantic-aware fusion network. First, we design a super-resolution network based on a multi-branch hybrid attention module (MHAM), which aims to enhance the quality and details of the source image, enabling the fusion network to integrate the features of the source image more accurately. Then, a comprehensive information extraction module (STDC) is designed in the fusion network to enhance the network’s ability to extract finer-grained complementary information from the source image. Finally, the fusion network and segmentation network are jointly trained to utilize semantic loss to guide the semantic information back to the fusion network, which effectively improves the performance of the fused images on high-level vision tasks. Extensive experiments show that our method is more effective than other state-of-the-art image fusion methods. In particular, our fused images not only have excellent visual perception effects, but also help to improve the performance of high-level vision tasks. Full article
(This article belongs to the Special Issue Multi-Modal Image Processing Methods, Systems, and Applications)
Show Figures

Figure 1

20 pages, 1863 KB  
Article
Denoising Diffusion Probabilistic Model with Adversarial Learning for Remote Sensing Super-Resolution
by Jialu Sui, Qianqian Wu and Man-On Pun
Remote Sens. 2024, 16(7), 1219; https://doi.org/10.3390/rs16071219 - 30 Mar 2024
Cited by 19 | Viewed by 5203
Abstract
Single Image Super-Resolution (SISR) for image enhancement enables the generation of high spatial resolution in Remote Sensing (RS) images without incurring additional costs. This approach offers a practical solution to obtain high-resolution RS images, addressing challenges posed by the expense of acquisition equipment [...] Read more.
Single Image Super-Resolution (SISR) for image enhancement enables the generation of high spatial resolution in Remote Sensing (RS) images without incurring additional costs. This approach offers a practical solution to obtain high-resolution RS images, addressing challenges posed by the expense of acquisition equipment and unpredictable weather conditions. To address the over-smoothing of the previous SISR models, the diffusion model has been incorporated into RS SISR to generate Super-Resolution (SR) images with enhanced textural details. In this paper, we propose a Diffusion model with Adversarial Learning Strategy (DiffALS) to refine the generative capability of the diffusion model. DiffALS integrates an additional Noise Discriminator (ND) into the training process, employing an adversarial learning strategy on the data distribution learning. This ND guides noise prediction by considering the general correspondence between the noisy image in each step, thereby enhancing the diversity of generated data and the detailed texture prediction of the diffusion model. Furthermore, considering that the diffusion model may exhibit suboptimal performance on traditional pixel-level metrics such as Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity (SSIM), we showcase the effectiveness of DiffALS through downstream semantic segmentation applications. Extensive experiments demonstrate that the proposed model achieves remarkable accuracy and notable visual enhancements. Compared to other state-of-the-art methods, our model establishes an improvement of 189 for Fréchet Inception Distance (FID) and 0.002 for Learned Perceptual Image Patch Similarity (LPIPS) in a SR dataset, namely Alsat, and achieves improvements of 0.4%, 0.3%, and 0.2% for F1 score, MIoU, and Accuracy, respectively, in a segmentation dataset, namely Vaihingen. Full article
Show Figures

Figure 1

16 pages, 7437 KB  
Article
Using Super-Resolution for Enhancing Visual Perception and Segmentation Performance in Veterinary Cytology
by Jakub Caputa, Maciej Wielgosz, Daria Łukasik, Paweł Russek, Jakub Grzeszczyk, Michał Karwatowski, Szymon Mazurek, Rafał Frączek, Anna Śmiech, Ernest Jamro, Sebastian Koryciak, Agnieszka Dąbrowska-Boruch, Marcin Pietroń and Kazimierz Wiatr
Life 2024, 14(3), 321; https://doi.org/10.3390/life14030321 - 28 Feb 2024
Cited by 1 | Viewed by 2235
Abstract
The primary objective of this research was to enhance the quality of semantic segmentation in cytology images by incorporating super-resolution (SR) architectures. An additional contribution was the development of a novel dataset aimed at improving imaging quality in the presence of inaccurate focus. [...] Read more.
The primary objective of this research was to enhance the quality of semantic segmentation in cytology images by incorporating super-resolution (SR) architectures. An additional contribution was the development of a novel dataset aimed at improving imaging quality in the presence of inaccurate focus. Our experimental results demonstrate that the integration of SR techniques into the segmentation pipeline can lead to a significant improvement of up to 25% in the mean average precision (mAP) metric. These findings suggest that leveraging SR architectures holds great promise for advancing the state-of-the-art in cytology image analysis. Full article
(This article belongs to the Section Animal Science)
Show Figures

Figure 1

23 pages, 15534 KB  
Article
Super-Resolution Semantic Segmentation of Droplet Deposition Image for Low-Cost Spraying Measurement
by Jian Liu, Shihui Yu, Xuemei Liu, Guohang Lu, Zhenbo Xin and Jin Yuan
Agriculture 2024, 14(1), 106; https://doi.org/10.3390/agriculture14010106 - 8 Jan 2024
Cited by 11 | Viewed by 3691
Abstract
In-field in situ droplet deposition digitization is beneficial for obtaining feedback on spraying performance and precise spray control, the cost-effectiveness of the measurement system is crucial to its scalable application. However, the limitations of camera performance in low-cost imaging systems, coupled with dense [...] Read more.
In-field in situ droplet deposition digitization is beneficial for obtaining feedback on spraying performance and precise spray control, the cost-effectiveness of the measurement system is crucial to its scalable application. However, the limitations of camera performance in low-cost imaging systems, coupled with dense spray droplets and a complex imaging environment, result in blurred and low-resolution images of the deposited droplets, which creates challenges in obtaining accurate measurements. This paper proposes a Droplet Super-Resolution Semantic Segmentation (DSRSS) model and a Multi-Adhesion Concave Segmentation (MACS) algorithm to address the accurate segmentation problem in low-quality droplet deposition images, and achieve a precise and efficient multi-parameter measurement of droplet deposition. Firstly, a droplet deposition image dataset (DDID) is constructed by capturing high-definition droplet images and using image reconstruction methods. Then, a lightweight DSRSS model combined with anti-blurring and super-resolution semantic segmentation is proposed to achieve semantic segmentation of deposited droplets and super-resolution reconstruction of segmentation masks. The weighted IoU (WIoU) loss function is used to improve the segmented independence of droplets, and a comprehensive evaluation criterion containing six sub-items is used for parameter optimization. Finally, the MACS algorithm continues to segment the remained adhesive droplets processed by the DSRSS model and corrects the bias of the individual droplet regions by regression. The experiments show that when the two weight parameters α and β in WIoU are 0.775 and 0.225, respectively, the droplet segmentation independence rate of DSRSS on the DDID reaches 0.998, and the IoU reaches 0.973. The MACS algorithm reduces the droplet adhesion rate in images with a coverage rate of more than 30% by 15.7%, and the correction function reduces the coverage error of model segmentation by 3.54%. The parameters of the DSRSS model are less than 1 M, making it possible to run it on embedded platforms. The proposed approach improves the accuracy of spray measurement using low-quality droplet deposition image and will help to scale-up of fast spray measurements in the field. Full article
Show Figures

Figure 1

Back to TopTop