ISDG-Net: Efficient RGB–Infrared Object Detection for Remote Sensing Imagery
Highlights
- We have proposed the ISDG-Net, a lightweight RGB–infrared detection framework, which integrates modules such as IBC-Conv, DySparse, Detect-SASD, and GIS to achieve efficient cross-modal fusion.
- This framework demonstrated extremely high accuracy on the VEDAI, M3FD and LLVIP datasets, while using only 42 million parameters. An extremely low cost of computing resources, with 1.13 billion floating-point operations.
- This tailored, integrated architecture successfully resolves the differences between various modes and reduces the excessive suppression of densely arranged small objects in complex and dimly lit environments.
- By achieving high-precision detection with extremely low computational overhead, this model makes it possible to deploy remote sensing object detection on resource-constrained edge platforms.
Abstract
1. Introduction
1.1. Background and Significance
- (1)
- Inherent limitations of single-modal image object detection: The visible light modality depends on illumination conditions and is prone to feature blurring or even target loss in scenes such as night-time or rainy weather; although the infrared modality is unaffected by illumination, it cannot capture key discriminative information such as color and texture, resulting in increased difficulty in distinguishing similar targets and a significant drop in recognition accuracy.
- (2)
- Multi-modal detection accuracy needs to be improved: Existing fusion methods mostly adopt fixed weights or simple concatenation to process dual-modal features, making them inflexible to the fluctuating distributions of modal information in different scenes; some models ignore the global-local information synergy during the course of generating and blending cross-domain representations, thereby deteriorating the recognition accuracy for diminutive targets and under complex backgrounds.
1.2. Main Contributions
- (1)
- We proposed ISDG-Net, a two-stream target recognition architecture tailored for complex, low-illumination remote sensing environments. Overcoming single-sensor limitations, it significantly enhances multi-scale object localization. Evaluations on VEDAI, M3FD, and LLVIP datasets demonstrate that ISDG-Net surpasses state-of-the-art models, achieving 55.1% mAP@0.5 on VEDAI and 93.7% on LLVIP, proving its exceptional accuracy, rapid convergence, and robust cross-domain generalization.
- (2)
- We design an inverted bottleneck module based on channel separation (IBC-Conv) to construct an efficient dual-modal feature extraction backbone, addressing edge deployment challenges. By utilizing inverted residuals and depthwise separable convolutions, it independently extracts visible textures and infrared characteristics, preventing mutual interference. This lossless semantic transmission significantly reduces parameter redundancy, improving cross-modal fusion and detection accuracy in complex scenes.
- (3)
- We integrate and adapt DySparse, a dynamic sparse Transformer module designed to alleviate the computational burden of global modeling in high-resolution imagery. By incorporating Bi-Level Routing Attention, DySparse employs a query-aware strategy to sparsely connect only the most semantically relevant regions. This approach reduces complexity while effectively suppressing background noise, thereby enhancing global context perception for small objects.
- (4)
- We develop a detection head integrating adaptive spatial feature fusion (Detect-SASD) to overcome modality heterogeneity and scale deviations. Utilizing a learnable adaptive weight map, it dynamically adjusts pixel-level dual-modal fusion ratios for accurate multi-scale alignment. This mechanism significantly enhances feature robustness under low illumination and complex backgrounds, effectively solving missed detections of multi-scale objects.
- (5)
- We adapt the CIoU and Soft-NMS mechanisms to construct a geometry-aware greedy IoU selector (GIS, distinct from Geographic Information Systems) to address the false suppression problem of traditional NMS in densely arranged scenes. By introducing multi-dimensional CIoU metrics (overlap area, center distance, aspect ratio), GIS reconstructs the suppression strategy without increasing training costs. This effectively distinguishes highly overlapping adjacent objects, significantly improving recall rates in dense vehicle scenes.
2. Related Works
2.1. Traditional Object Detection Algorithms
2.2. Visible-Infrared Object Detection
3. Method
3.1. Overall Architecture
3.2. IBC-Conv: Mid-Level Dual-Modal Feature Extraction and Fusion Backbone
3.3. Detect SASD (Learned Adaptive Spatial Detection)
3.4. Dynamic Sparse Transformer (DySparse)
3.5. GreedyIoU Selector (GIS)
4. Experiments
4.1. Datasets and Evaluation Metrics
4.1.1. Dual-Modal Image Object Detection Dataset
- (1)
- VEDAI
- (2)
- M3FD
- (3)
- LLVIP
4.1.2. Evaluation Metrics
4.2. Implementation Details
4.3. Ablation Experiments
4.4. Comparative Experiments
4.5. Generalization Experiments
4.6. Evaluation Experiment
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Liu, J.; Zhang, J.; Ni, Y.; Chi, W.; Qi, Z. Small-object detection in remote sensing images with super-resolution perception. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 15721–15734. [Google Scholar] [CrossRef]
- Wu, T.; Dong, Y. YOLO-SE: Improved YOLOv8 for remote sensing object detection and recognition. Appl. Sci. 2023, 13, 12977. [Google Scholar] [CrossRef]
- Sakaridis, C.; Dai, D.; Van Gool, L. ACDC: The adverse conditions dataset with correspondences for semantic driving scene understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 10765–10775. [Google Scholar]
- Hu, L.; Qin, M.; Zhang, F.; Du, Z.; Liu, R. RSCNN: A CNN-based method to enhance low-light remote-sensing images. Remote Sens. 2021, 13, 62. [Google Scholar] [CrossRef]
- Shen, R.; Zhang, X.; Xiang, Y. AFFNet: Attention mechanism network based on fusion feature for image cloud removal. Int. J. Pattern Recognit. Artif. Intell. 2022, 36, 2254014. [Google Scholar] [CrossRef]
- Ai, J.; Tian, R.; Luo, Q.; Jin, J.; Tang, B. Multi-scale rotation-invariant Haar-like feature integrated CNN-based ship detection algorithm of multiple-target environment in SAR imagery. IEEE Trans. Geosci. Remote Sens. 2019, 57, 10070–10087. [Google Scholar] [CrossRef]
- Ai, J.; Mao, Y.; Luo, Q.; Jia, L.; Xing, M. SAR target classification using the multikernel-size feature fusion-based convolutional neural network. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5214313. [Google Scholar] [CrossRef]
- Xue, W.; Ai, J.; Zhu, Y.; Sun, X.; Zhang, Y.; Gao, G. LMCNet: Light-weight modality compensation network for salient ship detection under missing modality conditions. IEEE Trans. Aerosp. Electron. Syst. 2026, 62, 6547–6560. [Google Scholar] [CrossRef]
- Huang, Q.; Sun, H.; Wang, Y.; Yuan, Y.; Guo, X.; Gao, Q. Ship detection based on YOLO algorithm for visible images. IET Image Process. 2023, 18, 481–492. [Google Scholar] [CrossRef]
- Chen, Z.; Xiang, W.; Lin, Z.; Yang, K.; Liu, Y.; Shi, Z. Alignment-assisted frequency fusion network for RGB-infrared vehicle detection. Neurocomputing 2025, 647, 130505. [Google Scholar] [CrossRef]
- Zhao, G.; Zhu, J.; Jiang, Q.; Feng, S.; Wang, Z. Edge feature enhanced transformer network for RGB and infrared image fusion based object detection. Infrared Phys. Technol. 2025, 147, 105824. [Google Scholar] [CrossRef]
- Hao, T.; Yang, J.; Zhang, S.; Wu, S. EEF: Energy score-guided feature enhancement fusion method for RGB and thermal infrared images object detection. Signal Process. 2026, 239, 110231. [Google Scholar] [CrossRef]
- Meng, F.; Hong, A.; Tang, H.; Tong, G. FQDNet: A fusion-enhanced quad-head network for RGB-infrared object detection. Remote Sens. 2025, 17, 1095. [Google Scholar] [CrossRef]
- Yuan, M.; Shi, X.; Wang, N.; Wang, Y.; Wei, X. Improving RGB-infrared object detection with cascade alignment-guided transformer. Inf. Fusion 2024, 105, 102246. [Google Scholar] [CrossRef]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the 29th Annual Conference on Neural Information Processing Systems (NIPS 2015), Montreal, QC, Canada, 7–12 December 2015; pp. 91–99. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16×16 words: Transformers for image recognition at scale. In Proceedings of the 9th International Conference on Learning Representations (ICLR 2021), Vienna, Austria, 3–7 May 2021. [Google Scholar]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In Proceedings of the 16th European Conference on Computer Vision (ECCV 2020), Glasgow, UK, 23–28 August 2020; pp. 213–229. [Google Scholar]
- Shehzadi, T.; Hashmi, K.A.; Liwicki, M.; Stricker, D.; Afzal, M.Z. Object detection with transformers: A review. Sensors 2025, 25, 6025. [Google Scholar] [CrossRef]
- Oquab, M.; Darcet, T.; Moutakanni, T.; Vo, H.; Szafraniec, M.; Khalidov, V.; Fernandez, P.; Haziza, D.; Massa, F.; El-Nouby, A.; et al. DINOv2: Learning robust visual features without supervision. arXiv 2024, arXiv:2304.07193. [Google Scholar] [CrossRef]
- Hassija, V.; Palanisamy, B.; Chatterjee, A.; Mandal, A.; Chakraborty, D.; Kumar, D.; Pandey, A. Transformers for vision: A survey on innovative methods for computer vision. IEEE Access 2025, 13, 95496–95523. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
- Shen, L.; Lang, B.; Song, Z. CA-YOLO: Model optimization for remote sensing image object detection. IEEE Access 2023, 11, 64769–64781. [Google Scholar] [CrossRef]
- Wang, C.Y.; Bochkovskiy, A.; Liao, H.Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 7464–7475. [Google Scholar]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-time end-to-end object detection. In Proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS 2024), Vancouver, BC, Canada, 10–15 December 2024. [Google Scholar]
- Jocher, G.; Chaurasia, A.; Qiu, J. YOLO11 by Ultralytics. 2024. Available online: https://github.com/ultralytics/ultralytics (accessed on 10 February 2026).
- Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-centric real-time object detectors. arXiv 2025, arXiv:2502.12524. [Google Scholar]
- Lei, M.; Li, S.; Wu, Y.; Hu, H.; Zhou, Y.; Zheng, X.; Ding, G.; Du, S.; Wu, Z.; Gao, Y. YOLOv13: Real-time object detection with hypergraph-enhanced adaptive visual perception. arXiv 2025, arXiv:2506.17733. [Google Scholar]
- Xie, S.; Zhou, M.; Wang, C.; Huang, S. CSPPartial-YOLO: A lightweight YOLO-based method for typical objects detection in remote sensing images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 388–399. [Google Scholar] [CrossRef]
- Tang, Q.; Su, C.; Tian, Y.; Zhao, S.; Yang, K.; Hao, W.; Feng, X.; Xie, M. YOLO-SS: Optimizing YOLO for enhanced small object detection in remote sensing imagery. J. Supercomput. 2025, 81, 303. [Google Scholar] [CrossRef]
- Fan, K.; Li, Q.; Li, Q.; Zhong, G.; Chu, Y.; Le, Z. YOLO-Remote: An object detection algorithm for remote sensing targets. IEEE Access 2024, 12, 155654–155665. [Google Scholar] [CrossRef]
- Wei, J.; Su, S.; Zhao, Z.; Tong, X.; Hu, L.; Gao, W. Infrared pedestrian detection using improved UNet and YOLO through sharing visible light domain information. Measurement 2023, 221, 113442. [Google Scholar] [CrossRef]
- Wen, M.; Li, C.; Xue, Y.; Xu, M.; Xi, Z.; Qiu, W. YOFIR: High precise infrared object detection algorithm based on YOLO and FasterNet. Infrared Phys. Technol. 2025, 144, 105627. [Google Scholar] [CrossRef]
- Wang, L.; Zhang, X.; Song, Z.; Bi, J.; Zhang, G.; Wei, H.; Tang, L.; Yang, L.; Li, J.; Jia, C.; et al. Multi-modal 3D object detection in autonomous driving: A survey and taxonomy. IEEE Trans. Intell. Veh. 2023, 8, 3781–3798. [Google Scholar] [CrossRef]
- Hwang, S.; Park, J.; Kim, N.; Choi, Y.; Kweon, I.S. Multispectral pedestrian detection: Benchmark dataset and baseline. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 1037–1045. [Google Scholar]
- Wang, C.; Yang, J.; Sun, D.; Gao, Q.; Liu, Q.; Wang, T.; Hu, A.; Wang, L. Air-to-ground target detection and tracking based on dual-stream fusion of unmanned aerial vehicle. J. Field Robot. 2025, 42, 3582–3599. [Google Scholar] [CrossRef]
- Bao, C.; Cao, J.; Hao, Q.; Cheng, Y.; Ning, Y.; Zhao, T. Dual-YOLO architecture from infrared and visible images for object detection. Sensors 2023, 23, 2934. [Google Scholar] [CrossRef] [PubMed]
- Zhang, H.; Fromont, É.; Lefèvre, S.; Avignon, B. Guided attentive feature fusion for multispectral pedestrian detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 5–9 January 2021; pp. 72–80. [Google Scholar]
- Tian, D.; Yan, X.; Zhou, D.; Wang, C.; Zhang, W. IV-YOLO: A lightweight dual-branch object detection network. Sensors 2024, 24, 6181. [Google Scholar] [CrossRef]
- Sun, X.; Zhu, Y.; Huang, H. Specificity-Guided Cross-Modal Feature Reconstruction for RGB-Infrared Object Detection. IEEE Trans. Intell. Transp. Syst. 2024, 25, 950–961. [Google Scholar] [CrossRef]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.C. MobileNetV2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 4510–4520. [Google Scholar]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 7132–7141. [Google Scholar]
- Liu, S.; Huang, D.; Wang, Y. Learning spatial fusion for single-shot object detection. arXiv 2019, arXiv:1911.09516. [Google Scholar] [CrossRef]
- Li, X.; Wang, W.; Wu, L.; Chen, S.; Hu, X.; Li, J.; Tang, J.; Yang, J. Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 9479–9488. [Google Scholar]
- Zhu, L.; Wang, X.; Ke, Z.; Zhang, W.; Lau, R.W.H. BiFormer: Vision Transformer with Bi-Level Routing Attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 10323–10333. [Google Scholar]
- Zhang, T.; Gao, G.; Zhang, X. Glance-Focus-Gaze: A novel Eagle-Eye Vision-Inspired Panorama-Population-Individual progressive screening paradigm to capture ships in SAR images. ISPRS J. Photogramm. Remote Sens. 2026, 235, 241–260. [Google Scholar] [CrossRef]
- Bodla, N.; Singh, B.; Chellappa, R.; Davis, L.S. Soft-NMS-Improving object detection with one line of code. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 5561–5569. [Google Scholar]
- Zheng, Z.; Wang, P.; Liu, W.; Li, J.; Ye, R.; Ren, D. Distance-IoU loss: Faster and better learning for bounding box regression. In Proceedings of the 34th AAAI Conference on Artificial Intelligence, New York, NY, USA, 7–12 February 2020; pp. 12993–13000. [Google Scholar]
- Razakarivony, S.; Jurie, F. Vehicle detection in aerial imagery: A small target detection benchmark. J. Vis. Commun. Image Represent. 2016, 34, 187–203. [Google Scholar] [CrossRef]
- Liu, J.; Fan, X.; Huang, Z.; Wu, G.; Liu, R.; Zhong, W.; Luo, Z. Target-aware dual adversarial learning and a multi-scenario multi modality benchmark to fuse infrared and visible for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 5802–5811. [Google Scholar]
- Jia, X.; Zhu, C.; Li, M.; Tang, W.; Zhou, W. LLVIP: A visible-infrared paired dataset for low-light vision. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 3496–3504. [Google Scholar]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland, 6–12 September 2014; pp. 740–755. [Google Scholar]
- Yu, Z. RT-DETR-iRMB: A lightweight real-time small object detection method. In Proceedings of the 2024 IEEE 6th Advanced Information Management, Communicates, Electronic and Automation Control Conference (IMCEC), Chongqing, China, 24–26 May 2024. [Google Scholar] [CrossRef]
- Gong, Z.; Xiao, G.; Shi, Z.; Chen, R.; Yu, J. MSGA-Net: Progressive feature matching via multi-layer sparse graph attention. IEEE Trans. Circuits Syst. Video Technol. 2024, 34, 5765–5775. [Google Scholar] [CrossRef]
- Tang, L.; Zhang, H.; Xu, H.; Ma, J. Rethinking the necessity of image fusion in high-level vision tasks: A practical infrared and visible image fusion network based on progressive semantic injection and scene fidelity. Inf. Fusion 2024, 99, 101870. [Google Scholar] [CrossRef]










| Environment | Parameters |
|---|---|
| CPU | Intel Core Ultra 9-285K |
| GPU | NVIDIA GeForce RTX 5090 |
| Memory | 64 GB |
| Language | Python 3.12 |
| Framework | PyTorch 2.8.0 |
| CUDA Version | 12.8 |
| Epochs | Batch Size | Workers | Optimizer | Amp | Initial Learning Rate | LR Schedule | Input Resolution |
|---|---|---|---|---|---|---|---|
| 200 | 24 | 20 | SGD | False | 0.01 | Linear Decay | 512 × 512 |
| Method | IBC-Conv | SASD | DySparse | Param (M) | FLOPs (G) | mAP@0.5 (%) | mAP@[0.5:0.95] (%) |
|---|---|---|---|---|---|---|---|
| Baseline | × | × | × | 3.6 | 9.9 | 50.8 | 31.1 |
| 1 | √ | × | × | 2.4 (↓1.2) | 8.5 (↓1.4) | 52.3 (↑1.5) | 31.3 (↑0.2) |
| 2 | × | √ | × | 5.2 (↑1.6) | 13.1 (↑3.2) | 50.0 (↓0.8) | 31.0 (↑0.1) |
| 3 | × | × | √ | 3.6 | 9.2 (↓0.7) | 48.9 (↓1.9) | 29.5 (↓1.6) |
| 4 | × | √ | √ | 5.3 (↑1.7) | 12.6 (↑2.7) | 50.4 (↓0.4) | 29.7 (↓1.4) |
| 5 | √ | × | √ | 2.5 (↓1.1) | 7.8 (↓2.1) | 48.0 (↓2.8) | 29.2 (↓1.9) |
| 6 | √ | √ | × | 4.1 (↑0.5) | 11.7 (↑1.8) | 51.9 (↑1.1) | 31.8 (↑0.7) |
| 7 | √ | √ | √ | 4.2 (↑0.6) | 11.3 (↑1.4) | 55.1 (↑4.3) | 33.8 (↑2.7) |
| Method | Param (M) | FLOPs (G) | Precision (%) | Recall (%) | mAP@0.5 (%) | mAP@[0.5:0.95] (%) |
|---|---|---|---|---|---|---|
| YOLOv10-mid-fusion [24] | 3.7 | 11.3 | 56.1 | 52.0 | 52.5 | 32.3 |
| YOLOv11-mid-fusion [25] | 3.7 | 9.6 | 57.7 | 52.7 | 51.0 | 31.2 |
| YOLOv12-mid-fusion [26] | 4.0 | 9.1 | 57.4 | 50.6 | 50.6 | 30.6 |
| YOLOv13-mid-fusion [27] | 3.5 | 9.9 | 63.2 | 50.5 | 50.8 | 31.1 |
| iRMB-fusion [52] | 4.1 | 11.8 | 52.9 | 52.1 | 49.3 | 28.7 |
| MSGA-fusion [53] | 4.7 | 14.4 | 57.8 | 53.5 | 50.8 | 30.9 |
| ISDG-Net (Ours) | 4.2 | 11.3 | 63.7 | 52.9 | 55.1 | 33.8 |
| Method | Param (M) | FLOPs (G) | Precision (%) | Recall (%) | mAP@0.5 (%) | mAP@[0.5:0.95] (%) |
|---|---|---|---|---|---|---|
| YOLOv10-mid-fusion [24] | 3.7 | 11.3 | 83.9 | 67.3 | 73.4 | 49.0 |
| YOLOv11-mid-fusion [25] | 3.7 | 9.6 | 83.1 | 64.2 | 71.1 | 46.6 |
| YOLOv12-mid-fusion [26] | 4.0 | 9.1 | 84.9 | 61.4 | 70.0 | 46.3 |
| YOLOv13-mid-fusion [27] | 3.5 | 9.9 | 84.5 | 65.9 | 72.7 | 48.6 |
| ISDG-Net (Ours) | 4.2 | 11.3 | 88.8 | 69.7 | 77.1 | 55.8 |
| Method | Param (M) | FLOPs (G) | Precision (%) | Recall (%) | mAP@0.5 (%) | mAP@[0.5:0.95] (%) |
|---|---|---|---|---|---|---|
| MSGA-fusion [53] | 4.7 | 13.7 | 85.4 | 65.3 | 72.06 | 48.50 |
| iRMB-fusion [52] | 4.1 | 11.8 | 83.3 | 66.7 | 73.31 | 48.67 |
| PSFM-fusion [54] | 4.7 | 14.4 | 86.2 | 64.4 | 72.71 | 48.5 |
| Ours | 4.2 | 11.3 | 88.8 | 69.7 | 77.10 | 55.8 |
| Method | Param (M) | FLOPs (G) | Precision (%) | Recall (%) | mAP@0.5 (%) | mAP@[0.5:0.95] (%) |
|---|---|---|---|---|---|---|
| MSGA-fusion [53] | 4.7 | 13.7 | 91.3 | 89.4 | 93.1 | 62.2 |
| iRMB-fusion [52] | 4.1 | 11.8 | 91.5 | 88.7 | 92.2 | 61.3 |
| PSFM-fusion [54] | 4.7 | 14.4 | 93.4 | 87.4 | 93.4 | 63.5 |
| Ours | 4.2 | 11.3 | 91.5 | 89.7 | 93.7 | 63.6 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Gao, Y.; Cheng, X.; Li, Y.; Xu, D.; Sun, D.; Hu, Y. ISDG-Net: Efficient RGB–Infrared Object Detection for Remote Sensing Imagery. Remote Sens. 2026, 18, 1570. https://doi.org/10.3390/rs18101570
Gao Y, Cheng X, Li Y, Xu D, Sun D, Hu Y. ISDG-Net: Efficient RGB–Infrared Object Detection for Remote Sensing Imagery. Remote Sensing. 2026; 18(10):1570. https://doi.org/10.3390/rs18101570
Chicago/Turabian StyleGao, Yaoyue, Xinru Cheng, Yimeng Li, Dawei Xu, Desheng Sun, and Yaoyi Hu. 2026. "ISDG-Net: Efficient RGB–Infrared Object Detection for Remote Sensing Imagery" Remote Sensing 18, no. 10: 1570. https://doi.org/10.3390/rs18101570
APA StyleGao, Y., Cheng, X., Li, Y., Xu, D., Sun, D., & Hu, Y. (2026). ISDG-Net: Efficient RGB–Infrared Object Detection for Remote Sensing Imagery. Remote Sensing, 18(10), 1570. https://doi.org/10.3390/rs18101570
