Next Article in Journal
A Spatial Reconstruction Method of Ionospheric foF2 Based on High Accuracy Surface Modeling Theory
Next Article in Special Issue
Semantic Labeling of High-Resolution Images Combining a Self-Cascaded Multimodal Fully Convolution Neural Network with Fully Conditional Random Field
Previous Article in Journal
UVIO: Adaptive Kalman Filtering UWB-Aided Visual-Inertial SLAM System for Complex Indoor Environments
Previous Article in Special Issue
Stepwise Attention-Guided Multiscale Fusion Network for Lightweight and High-Accurate SAR Ship Detection
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

GLFuse: A Global and Local Four-Branch Feature Extraction Network for Infrared and Visible Image Fusion

1
School of Computer Science and Technology, Guangdong University of Technology, Guangzhou 510006, China
2
School of Integrated Circuits, Guangdong University of Technology, Guangzhou 510006, China
3
School of Automation, Guangdong University of Technology, Guangzhou 510006, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2024, 16(17), 3246; https://doi.org/10.3390/rs16173246
Submission received: 16 July 2024 / Revised: 17 August 2024 / Accepted: 30 August 2024 / Published: 1 September 2024

Abstract

Infrared and visible image fusion integrates complementary information from different modalities into a single image, providing sufficient imaging information for scene interpretation and downstream target recognition tasks. However, existing fusion methods often focus only on highlighting salient targets or preserving scene details, failing to effectively combine entire features from different modalities during the fusion process, resulting in underutilized features and poor overall fusion effects. To address these challenges, a global and local four-branch feature extraction image fusion network (GLFuse) is proposed. On one hand, the Super Token Transformer (STT) block, which is capable of rapidly sampling and predicting super tokens, is utilized to capture global features in the scene. On the other hand, a Detail Extraction Block (DEB) is developed to extract local features in the scene. Additionally, two feature fusion modules, namely the Attention-based Feature Selection Fusion Module (ASFM) and the Dual Attention Fusion Module (DAFM), are designed to facilitate selective fusion of features from different modalities. Of more importance, the various perceptual information of feature maps learned from different modality images at the different layers of a network is investigated to design a perceptual loss function to better restore scene detail information and highlight salient targets by treating the perceptual information separately. Extensive experiments confirm that GLFuse exhibits excellent performance in both subjective and objective evaluations. It deserves note that GLFuse effectively improves downstream target detection performance on a unified benchmark.
Keywords: image fusion; infrared and visible image fusion; global and local feature extraction; attention mechanism; deep learning image fusion; infrared and visible image fusion; global and local feature extraction; attention mechanism; deep learning

Share and Cite

MDPI and ACS Style

Zhao, G.; Hu, Z.; Feng, S.; Wang, Z.; Wu, H. GLFuse: A Global and Local Four-Branch Feature Extraction Network for Infrared and Visible Image Fusion. Remote Sens. 2024, 16, 3246. https://doi.org/10.3390/rs16173246

AMA Style

Zhao G, Hu Z, Feng S, Wang Z, Wu H. GLFuse: A Global and Local Four-Branch Feature Extraction Network for Infrared and Visible Image Fusion. Remote Sensing. 2024; 16(17):3246. https://doi.org/10.3390/rs16173246

Chicago/Turabian Style

Zhao, Genping, Zhuyong Hu, Silu Feng, Zhuowei Wang, and Heng Wu. 2024. "GLFuse: A Global and Local Four-Branch Feature Extraction Network for Infrared and Visible Image Fusion" Remote Sensing 16, no. 17: 3246. https://doi.org/10.3390/rs16173246

APA Style

Zhao, G., Hu, Z., Feng, S., Wang, Z., & Wu, H. (2024). GLFuse: A Global and Local Four-Branch Feature Extraction Network for Infrared and Visible Image Fusion. Remote Sensing, 16(17), 3246. https://doi.org/10.3390/rs16173246

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop