remotesensing-logo

Journal Browser

Journal Browser

Image Fusion and Object Detection Using Multi-Modal Remote Sensing Data

A Special Issue of Remote Sensing (ISSN 2072-4292) belonging to the section "Remote Sensing Image Processing".

Deadline for manuscript submissions: closed (30 April 2026) | Viewed by 14297

Editors


E-Mail Website
Guest Editor
College of Information and Electrical Engineering, China Agricultural University, Beijing, China
Interests: thermal infrared remote sensing; deep learning

E-Mail Website
Guest Editor
College of Information, Central University of Finance and Economics, Beijing, China
Interests: Image processing; Deep learning

Special Issue Information

Dear Colleagues,

Multi-modal remote sensing data provide abundant information about ground objects from multiple perspectives. The image fusion technique aims to extract complementary information from source images of different modalities and combine the information to generate images/data with abundant information, which is believed to be conducive to subsequent tasks, such as object detection. Image fusion and object detection have thus attracted extensive attention from the remote sensing community. Despite remarkable successes that have been achieved in recent decades, challenges remain regarding the fast development of multi-modal data collection and artificial intelligence techniques. Considering both data properties and model ability, improving the performance of image fusion and object detection can be possible.

This Special Issue aims to study all techniques designed for image fusion and object detection from multi-modal remote sensing data, ranging from multispectral, hyperspectral, panchromatic, thermal, and SAR data. Topics from conventional machine learning techniques and model-driven techniques to the most recent deep learning techniques may be covered. Applications and reviews about the aforementioned techniques are also welcome for submission to this Special Issue, as they can give helpful guidance for future-related studies. Articles may address, but are not limited, to the following topics:

  • Remote sensing image fusion;
  • Remote sensing image classification;
  • Remote sensing image instance segmentation;
  • Object detection and tracking;
  • Benchmark dataset creation;
  • Deep neural network optimization;
  • Multi-modal data analysis and fusion;
  • Applications of image fusion;
  • Applications of object detection;
  • Review of images fusion;
  • Review of object detection;
  • Review of image classification.

Dr. Bin Yang
Dr. Xin Ye
Dr. Jing Li
Guest Editors

Manuscript Submission Information

Manuscripts should be submitted online at www.mdpi.com by registering and logging in to this website. Once you are registered, click here to go to the submission form. Manuscripts can be submitted until the deadline. All submissions that pass pre-check are peer-reviewed. Accepted papers will be published continuously in the journal (as soon as accepted) and will be listed together on the special issue website. Research articles, review articles as well as short communications are invited. For planned papers, a title and short abstract (about 250 words) can be sent to the Editorial Office for assessment.

Submitted manuscripts should not have been published previously, nor be under consideration for publication elsewhere (except conference proceedings papers). All manuscripts are thoroughly refereed through a single-anonymized peer-review process. A guide for authors and other relevant information for submission of manuscripts is available on the Instructions for Authors page. Remote Sensing is an international peer-reviewed open access semimonthly journal published by MDPI.

Please visit the Instructions for Authors page before submitting a manuscript. The Article Processing Charge (APC) for publication in this open access journal is 2700 CHF (Swiss Francs). Submitted papers should be well formatted and use good English. Authors may use MDPI's English editing service prior to publication or during author revisions.

Keywords

  • image fusion
  • object detection
  • multi-modal data
  • deep learning
  • benchmark dataset
  • remote sensing data analysis

Benefits of Publishing in a Special Issue

  • Ease of navigation: Grouping papers by topic helps scholars navigate broad scope journals more efficiently.
  • Greater discoverability: Special Issues support the reach and impact of scientific research. Articles in Special Issues are more discoverable and cited more frequently.
  • Expansion of research network: Special Issues facilitate connections among authors, fostering scientific collaborations.
  • External promotion: Articles in Special Issues are often promoted through the journal's social media, increasing their visibility.
  • Reprint: MDPI Books provides the opportunity to republish successful Special Issues in book format, both online and in print.

Further information on MDPI's Special Issue policies can be found here.

Published Papers (7 papers)

Order results
Result details
Select all
Export citation of selected articles as:

Research

26 pages, 3445 KB  
Article
Significance-Preserving Progressive Network for Infrared and Visible Image Fusion
by Jingsui Li, Xiaorun Li, Shu Xiang and Shuhan Chen
Remote Sens. 2026, 18(14), 2328; https://doi.org/10.3390/rs18142328 - 12 Jul 2026
Viewed by 363
Abstract
Fusing infrared and visible images can effectively compensate for the inherent limitations of each modality in different scenes, resulting in fused images that contain richer information. However, existing methods often struggle to balance global dependency modeling with local detail preservation and to effectively [...] Read more.
Fusing infrared and visible images can effectively compensate for the inherent limitations of each modality in different scenes, resulting in fused images that contain richer information. However, existing methods often struggle to balance global dependency modeling with local detail preservation and to effectively coordinate heterogeneous local and global features during fusion. To address these issues, this paper proposes a Significance-Preserving Progressive Fusion Network (SiPFusion). First, a progressive feature extraction framework was designed, which hierarchically extracts multi-scale local features using CNNs and then models long-range dependencies across scales via a Transformer-based global module. To adaptively integrate local-global complementary features, a significance-preserving fusion module was designed to obtain significance attention maps with a spatial selection mechanism, enabling dynamic fusion of multi-source features. Furthermore, we propose a significance similarity loss function that leverages intermediate feature guidance to enhance structural consistency and preserve salient-region information in the fused image. Extensive experiments on the MSRS, RoadScene, and TNO datasets demonstrate that SiPFusion achieves competitive visual quality and strong overall quantitative performance against 15 state-of-the-art fusion methods, obtaining leading results on most evaluated metrics. Full article
Show Figures

Figure 1

19 pages, 2430 KB  
Article
LMFusion: Breaking the Computational Barrier for Multimodal Classification in Remote Sensing
by Shenbo Zhou, Sibo He, Daixun Li, Weiying Xie and Yunsong Li
Remote Sens. 2026, 18(12), 1972; https://doi.org/10.3390/rs18121972 - 13 Jun 2026
Viewed by 360
Abstract
Multi-modal land cover classification plays an important role in remote sensing applications such as urban monitoring and environmental analysis. By integrating complementary information from hyperspectral imagery (HSI) and LiDAR data, multimodal learning can significantly improve classification performance. However, existing Transformer-based fusion methods often [...] Read more.
Multi-modal land cover classification plays an important role in remote sensing applications such as urban monitoring and environmental analysis. By integrating complementary information from hyperspectral imagery (HSI) and LiDAR data, multimodal learning can significantly improve classification performance. However, existing Transformer-based fusion methods often suffer from high computational complexity and inefficient cross-modal interaction modeling, which limits their applicability in resource-constrained scenarios. To address these challenges, we propose LMFusion, an efficient framework for multimodal feature learning. Specifically, LMFusion enables efficient bidirectional feature interaction through a linear-complexity cross-attention mechanism and enhances long-range spatial-spectral representation learning with Mamba-based state space modeling, thereby achieving effective multimodal dependency modeling with linear computational complexity. In addition, a selective quantization-aware optimization strategy is introduced to support multiple bit-width settings (down to 1-bit), yielding a more compact and efficient model while improving representation robustness under low-bit constraints. Extensive experiments on the Houston2013, MUUFL, and Augsburg datasets demonstrate the effectiveness of LMFusion. It achieves overall accuracies of 95.84%, 94.95%, and 99.05%, respectively, consistently outperforming representative multimodal classification methods and showing strong potential for accurate and efficient multimodal remote sensing classification. Full article
Show Figures

Figure 1

23 pages, 53610 KB  
Article
Multispectral Sparse Cross-Attention Guided Mamba Network for Small Object Detection in Remote Sensing
by Wen Xiang, Yamin Li, Liu Duan, Qifeng Wu, Jiaqi Ruan, Yucheng Wan and Sihan Wu
Remote Sens. 2026, 18(3), 381; https://doi.org/10.3390/rs18030381 - 23 Jan 2026
Cited by 2 | Viewed by 1674
Abstract
Remote sensing small object detection remains a challenging task due to limited feature representation and interference from complex backgrounds. Existing methods that rely exclusively on either visible or infrared modalities often fail to achieve both accuracy and robustness in detection. Effectively integrating cross-modal [...] Read more.
Remote sensing small object detection remains a challenging task due to limited feature representation and interference from complex backgrounds. Existing methods that rely exclusively on either visible or infrared modalities often fail to achieve both accuracy and robustness in detection. Effectively integrating cross-modal information to enhance detection performance remains a critical challenge. To address this issue, we propose a novel Multispectral Sparse Cross-Attention Guided Mamba Network (MSCGMN) for small object detection in remote sensing. The proposed MSCGMN architecture comprises three key components: Multispectral Sparse Cross-Attention Guidance Module (MSCAG), Dynamic Grouped Mamba Block (DGMB), and Gated Enhanced Attention Module (GEAM). Specifically, the MSCAG module selectively fuses RGB and infrared (IR) features using sparse cross-modal attention, effectively capturing complementary information across modalities while suppressing redundancy. The DGMB introduces a dynamic grouping strategy to improve the computational efficiency of Mamba, enabling effective global context modeling. In remote sensing images, small objects occupy limited areas, making it difficult to capture their critical features. We design the GEAM module to enhance both global and local feature representations for small object detection. Experiments on the VEDAI and DroneVehicle datasets show that MSCGMN achieves mAP50 scores of 83.9% and 84.4%, outperforming existing state-of-the-art methods and demonstrating strong competitiveness in small object detection tasks. Full article
Show Figures

Graphical abstract

25 pages, 25629 KB  
Article
DSEPGAN: A Dual-Stream Enhanced Pyramid Based on Generative Adversarial Network for Spatiotemporal Image Fusion
by Dandan Zhou, Lina Xu, Ke Wu, Huize Liu and Mengting Jiang
Remote Sens. 2025, 17(24), 4050; https://doi.org/10.3390/rs17244050 - 17 Dec 2025
Cited by 2 | Viewed by 721
Abstract
Many deep learning-based spatiotemporal fusion (STF) methods have been proven to achieve high accuracy and robustness. Due to the variable shapes and sizes of objects in remote sensing images, pyramid networks are generally introduced to extract multi-scale features. However, the down-sampling operation in [...] Read more.
Many deep learning-based spatiotemporal fusion (STF) methods have been proven to achieve high accuracy and robustness. Due to the variable shapes and sizes of objects in remote sensing images, pyramid networks are generally introduced to extract multi-scale features. However, the down-sampling operation in the pyramid structure may lead to the loss of image detail information, affecting the model’s ability to reconstruct fine-grained targets. To address this issue, we propose a novel Dual-Stream Enhanced Pyramid based on Generative Adversarial Network (DSEPGAN) for the spatiotemporal fusion of remote sensing images. The network adopts a dual-stream architecture to separately process coarse and fine images, tailoring feature extraction to their respective characteristics: coarse images provide temporal dynamics, while fine images contain rich spatial details. A reversible feature transformation is embedded in the pyramid feature extraction stage to preserve high-frequency information, and a fusion module employing large-kernel and depthwise separable convolutions captures long-range dependencies across inputs. To further enhance realism and detail fidelity, adversarial training encourages the network to generate sharper and more visually convincing fusion results. The proposed DSEPGAN is compared with widely used and state-of-the-art STF models in three publicly available datasets. The results illustrate that DSEPGAN achieves superior performance across various evaluation metrics, highlighting its notable advantages for predicting seasonal variations in highly heterogeneous regions and abrupt changes in land use. Full article
Show Figures

Figure 1

20 pages, 8646 KB  
Article
Fine-Grained Multispectral Fusion for Oriented Object Detection in Remote Sensing
by Xin Lan, Shaolin Zhang, Yuhao Bai and Xiaolin Qin
Remote Sens. 2025, 17(22), 3769; https://doi.org/10.3390/rs17223769 - 20 Nov 2025
Cited by 2 | Viewed by 2256
Abstract
Infrared–visible-oriented object detection aims to combine the strengths of both infrared and visible images, overcoming the limitations of a single imaging modality to achieve more robust detection with oriented bounding boxes under diverse environmental conditions. However, current methods often suffer from two issues: [...] Read more.
Infrared–visible-oriented object detection aims to combine the strengths of both infrared and visible images, overcoming the limitations of a single imaging modality to achieve more robust detection with oriented bounding boxes under diverse environmental conditions. However, current methods often suffer from two issues: (1) modality misalignment caused by hardware and annotation errors, leading to inaccurate feature fusion that degrades downstream task performance; and (2) insufficient directional priors in square convolutional kernels, impeding robust object detection with diverse directions, especially in densely packed scenes. To tackle these challenges, in this paper, we propose a novel method, Fine-Grained Multispectral Fusion (FGMF), for oriented object detection in the paired aerial images. Specifically, we design a dual-enhancement and fusion module (DEFM) to obtain the calibrated and complementary features through weighted addition and subtraction-based attention mechanisms. Furthermore, we propose an orientation aggregation module (OAM) that employs large rotated strip convolutions to capture directional context and long-range dependencies. Extensive experiments on the DroneVehicle and VEDAI datasets demonstrate the effectiveness of our proposed method, yielding impressive results with accuracies of 80.2% and 66.3%, respectively. These results highlight the effectiveness of FGMF in oriented object detection within complex remote sensing scenarios. Full article
Show Figures

Figure 1

26 pages, 62665 KB  
Article
FAMHE-Net: Multi-Scale Feature Augmentation and Mixture of Heterogeneous Experts for Oriented Object Detection
by Yixin Chen, Weilai Jiang and Yaonan Wang
Remote Sens. 2025, 17(2), 205; https://doi.org/10.3390/rs17020205 - 8 Jan 2025
Cited by 8 | Viewed by 3336
Abstract
Object detection in remote sensing images is essential for applications like unmanned aerial vehicle (UAV)-assisted agricultural surveys and aerial traffic analysis, facing unique challenges such as low resolution, complex backgrounds, and the variability of object scales. Current detectors struggle with integrating spatial and [...] Read more.
Object detection in remote sensing images is essential for applications like unmanned aerial vehicle (UAV)-assisted agricultural surveys and aerial traffic analysis, facing unique challenges such as low resolution, complex backgrounds, and the variability of object scales. Current detectors struggle with integrating spatial and semantic information effectively across scales and often omit necessary refinement modules to focus on salient features. Furthermore, a detector head that lacks a meticulous design may face limitations in fully understanding and accurately predicting based on the enriched feature representations. These deficiencies can lead to insufficient feature representation and reduced detection accuracy. To address these challenges, this paper introduces a novel deep-learning framework, FAMHE-Net, for enhancing object detection in remote sensing images. Our framework features a consolidated multi-scale feature enhancement module (CMFEM) with integrated Path Aggregation Feature Pyramid Network (PAFPN), utilizing our efficient atrous channel attention (EACA) within CMFEM for enhanced contextual and semantic information refinement. Additionally, we introduce a sparsely gated mixture of heterogeneous expert heads (MOHEH) to adaptively aggregate detector head outputs. Compared to the baseline model, FAMEH-Net demonstrates significant improvements, achieving a 0.90% increase in mean Average Precision (mAP) of the DOTA dataset and a 1.30% increase in mAP12 of HRSC2016 datasets. These results highlight the effectiveness of FAMEH-Net in object detection within complex remote sensing images. Full article
Show Figures

Figure 1

21 pages, 57724 KB  
Article
MDSCNN: Remote Sensing Image Spatial–Spectral Fusion Method via Multi-Scale Dual-Stream Convolutional Neural Network
by Wenqing Wang, Fei Jia, Yifei Yang, Kunpeng Mu and Han Liu
Remote Sens. 2024, 16(19), 3583; https://doi.org/10.3390/rs16193583 - 26 Sep 2024
Cited by 7 | Viewed by 3624
Abstract
Pansharpening refers to enhancing the spatial resolution of multispectral images through panchromatic images while preserving their spectral features. However, existing traditional methods or deep learning methods always have certain distortions in the spatial or spectral dimensions. This paper proposes a remote sensing spatial–spectral [...] Read more.
Pansharpening refers to enhancing the spatial resolution of multispectral images through panchromatic images while preserving their spectral features. However, existing traditional methods or deep learning methods always have certain distortions in the spatial or spectral dimensions. This paper proposes a remote sensing spatial–spectral fusion method based on a multi-scale dual-stream convolutional neural network, which includes feature extraction, feature fusion, and image reconstruction modules for each scale. In terms of feature fusion, we propose a multi cascade module to better fuse image features. We also design a new loss function aim at enhancing the high degree of consistency between fused images and reference images in terms of spatial details and spectral information. To validate its effectiveness, we conduct thorough experimental analyses on two widely used remote sensing datasets: GeoEye-1 and Ikonos. Compared with the nine leading pansharpening techniques, the proposed method demonstrates superior performance in multiple key evaluation metrics. Full article
Show Figures

Figure 1

Back to TopTop