Mamba-Based Infrared and Visible Images Fusion Method
Highlights
- A Mamba-based fusion framework is proposed, integrating Multi-Path Mamba (MPMamba) and Dual-path Mamba Attention Fusion (DMAF) modules, significantly enhancing information retention and structural consistency in fused images.
- Experiments on the MSRS dataset demonstrate that the proposed method outperforms state-of-the-art approaches in key metrics such as Information Entropy, Spatial Frequency, and Mutual Information, and exhibit strong generalization in object detection tasks.
- This study provides a novel paradigm for multimodal image fusion that combines global modeling capability with linear computational complexity, facilitating applications in autonomous driving and nighttime surveillance.
- Through dual-path feature decoupling and dynamic attention calibration, the method achieves a better balance between preserving thermal targets from infrared images and texture details from visible images, offering higher-quality inputs for downstream vision tasks.
Abstract
1. Introduction
- A Multi-path Mamba module is designed to perform multi-scale characterization of texture details in visible images and thermal radiation features in infrared images, providing rich feature representations for cross-modal fusion.
- A Dual-path Mamba Attention Fusion (DMAF) module is proposed, which leverages the global modeling capability and dynamic feature selection mechanism of the Mamba model to separately process common and distinct features from visible and infrared images, and achieves adaptive fusion through the Convolutional Block Attention Module (CBAM).
- Experiments conducted on the MSRS dataset validate the superiority of the proposed method in both objective metrics and subjective quality.
2. Materials
2.1. Deep Learning-Based Fusion
2.2. State Space Models
3. Methods
3.1. Background
| Algorithm 1 |
| 1: |
| 2: |
| 3: |
| 4: |
| 5: |
| 6: |
| 7: |
| 8: |
| 9: |
| 10: |
| 11: |
| 12: |
| 13: |
| 14: |
| 15: |
| 16: |
| 17: |
| 18: |
| 19: |
| 20: |
| 21: |
| 22: |
| 23: |
| 24: |
| 25: |
| 26: |
| 27: |
| 28: |
| 29: |
3.2. Overall Architecture
3.3. MPMamba Module
3.4. Dual-Path Mamba Attention Fusion Module
- (1)
- Common Feature Path
- (2)
- Differential Feature Path
- (1)
- Channel Attention Module
- (2)
- Spatial Attention Module
4. Results
4.1. Training
4.2. Dataset
4.3. Experimental Setup
4.4. Evaluation Metric
- (a)
- Information Richness
- (b)
- Structural Fidelity
- (c)
- Detail and Clarity
- (d)
- Contrast and Saliency
4.5. Comparison with Existing Methods
4.6. Generalization Experiment
4.7. Efficiency Comparison
4.8. Ablation Study
- (1)
- The complementary design of multi-path convolutions and the Mamba block in MPMamba captures both local details and global context, overcoming the limitations of a single scale or a single operator’s receptive field;
- (2)
- The dual-path Mamba attention fusion module explicitly separates and integrates shared and complementary cross-modal features, followed by adaptive enhancement through channel-spatial attention, thereby achieving precise and robust information fusion. Consequently, our method exhibits superior comprehensive performance in both quantitative metrics and visual perception.
4.9. Object Detection Performance Evaluation
5. Discussion
5.1. Interpretation of Results and Methodological Innovation
5.2. Implications and Broader Context
5.3. Limitations and Future Research Directions
- Dynamic and Adaptive Fusion Mechanisms: The current fusion strategy employs a fixed architecture. Future work could explore more dynamic mechanisms, such as content-adaptive weighting or gating networks, to decide “what” and “how” to fuse at different regions or feature levels. This aligns with the growing focus on developing adaptive fusion strategies that can handle varying scene complexities and modality-specific information strengths.
- Advanced Feature Alignment Techniques: While the DMAF module provides a foundational decoupling-and-alignment strategy, the challenge of robust feature matching under significant modality gaps persists. Future improvements could investigate more sophisticated alignment techniques inspired by recent progress in multimodal and multi-view learning, such as employing gradient-aware alignment losses or exploring non-local correlation modules, to further enhance the precision of cross-modal feature integration.
- Extension to More Modalities and Tasks: The Mamba-based architecture is promising for fusing more than two modalities. Exploring its capacity for unified multimodal (e.g., visible, infrared, depth) representation learning and for other downstream perception tasks like segmentation is a logical next step.
- End-to-End Task-Specific Optimization: While we evaluated on detection, the fusion network was not trained end-to-end with the task loss. Joint optimization could further tailor the fused features to maximize performance for specific applications, creating a tighter synergy between fusion and perception.
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Appendix A. Evaluation Metric Formulas
- (1)
- Entropy (EN):
- (2)
- Standard Deviation (SD)
- (3)
- Spatial Frequency (SF)
- (4)
- Mutual Information (MI)
- (5)
- Visual Information Fidelity (VIF):
- (6)
- Quality Assessment of Blended Features (Qabf):
- (7)
- Structural Similarity Index Measure (SSIM)
- (8)
- Sum of Correlation Differences (SCD)
References
- Ma, J.; Tang, L.; Xu, M.; Zhang, H.; Xiao, G. STDFusionNet: An Infrared and Visible Image Fusion Network Based on Salient Target Detection. IEEE Trans. Instrum. Meas. 2021, 70, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Cao, Y.; Guan, D.; Huang, W.; Yang, J.; Cao, Y.; Qiao, Y. Pedestrian detection with unsupervised multispectral feature learning using deep neural networks. Inf. Fusion 2019, 46, 206–217. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Demiris, Y. Visible and Infrared Image Fusion Using Deep Learning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 10535–10554. [Google Scholar] [CrossRef]
- Toet, A. Image fusion by a ratio of low-pass pyramid. Pattern Recognit. Lett. 1989, 9, 245–253. [Google Scholar] [CrossRef] [Scilit]
- Toet, A.; Van Ruyven, L.J.; Valeton, J.M. Merging Thermal And Visual Images by A Contrast Pyramid. Opt. Eng. 1989, 28, 789–792. [Google Scholar] [CrossRef] [Scilit]
- Toet, A. A morphological pyramidal image decomposition. Pattern Recognit. Lett. 1989, 9, 255–261. [Google Scholar] [CrossRef] [Scilit]
- Ibrahim, R.; Alirezaie, J.; Babyn, P. Pixel level jointed sparse representation with RPCA image fusion algorithm. In Proceedings of the 2015 38th International Conference on Telecommunications and Signal Processing (TSP), Prague, Czech Republic, 9–11 July 2015; IEEE: New York, NY, USA, 2015; pp. 592–595. [Google Scholar] [CrossRef] [Scilit]
- Aishwarya, N.; Bennila, C. Thangammal An image fusion framework using novel dictionary based sparse representation. Multimed. Tools Appl. 2017, 76, 21869–21888. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Liu, Y.; Sun, P.; Yan, H.; Zhao, X.; Zhang, L. IFCNN: A general image fusion framework based on convolutional neural network. Inf. Fusion 2020, 54, 99–118. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Xu, H.; Xiao, Y.; Guo, X.; Ma, J. Rethinking the Image Fusion: A Fast Unified Image Fusion Network based on Proportional Maintenance of Gradient and Intensity. AAAI 2020, 34, 12797–12804. [Google Scholar] [CrossRef] [Scilit]
- Xu, H.; Ma, J.; Jiang, J.; Guo, X.; Ling, H. U2Fusion: A Unified Unsupervised Image Fusion Network. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 502–518. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. arXiv 2023, arXiv:1706.03762. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, H.; Zhang, J.; Cai, Y.; Chen, H.; Hu, X.; Gan, Z.; Wang, Y.; Wang, C.; Wu, Y.; Xie, L. MobileMamba: Lightweight Multi-Receptive Visual Mamba Network. arXiv 2024, arXiv:2411.15941. [Google Scholar] [CrossRef] [Scilit]
- Tay, Y.; Dehghani, M.; Bahri, D.; Metzler, D. Efficient Transformers: A Survey. ACM Comput. Surv. 2023, 55, 1–28. [Google Scholar] [CrossRef] [Scilit]
- Gu, A.; Dao, T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv 2024, arXiv:2312.00752. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; Wu, X.-J. DenseFuse: A Fusion Approach to Infrared and Visible Images. IEEE Trans. Image Process. 2019, 28, 2614–2623. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; Wu, X.-J.; Durrani, T. NestFuse: An Infrared and Visible Image Fusion Architecture Based on Nest Connection and Spatial/Channel Attention Models. IEEE Trans. Instrum. Meas. 2020, 69, 9645–9656. [Google Scholar] [CrossRef] [Scilit]
- Ma, J.; Yu, W.; Liang, P.; Li, C.; Jiang, J. FusionGAN: A generative adversarial network for infrared and visible image fusion. Inf. Fusion 2019, 48, 11–26. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Huo, H.; Li, C.; Wang, R.; Feng, Q. AttentionFGAN: Infrared and Visible Image Fusion Using Attention-Based Generative Adversarial Networks. IEEE Trans. Multimed. 2021, 23, 1383–1396. [Google Scholar] [CrossRef] [Scilit]
- Yue, J.; Fang, L.; Xia, S.; Deng, Y.; Ma, J. Dif-Fusion: Toward High Color Fidelity in Infrared and Visible Image Fusion With Diffusion Models. IEEE Trans. Image Process. 2023, 32, 5705–5720. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Zhu, J.; Li, C.; Chen, X.; Yang, B. CGTF: Convolution-Guided Transformer for Infrared and Visible Image Fusion. IEEE Trans. Instrum. Meas. 2022, 71, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Ma, J.; Tang, L.; Fan, F.; Huang, J.; Mei, X.; Ma, Y. SwinFusion: Cross-domain Long-range Learning for General Image Fusion via Swin Transformer. IEEE/CAA J. Autom. Sin. 2022, 9, 1200–1217. [Google Scholar] [CrossRef] [Scilit]
- Xu, H.; Liang, P.; Yu, W.; Jiang, J.; Ma, J. Learning a Generative Model for Fusing Infrared and Visible Images via Conditional Generative Adversarial Network with Dual Discriminators. In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, Macao, China, 10–16 August 2019; International Joint Conferences on Artificial Intelligence Organization: Sydney, Australia, 2019; pp. 3954–3960. [Google Scholar] [CrossRef] [Scilit]
- Jie, Y.; Xu, Y.; Li, X.; Zhou, F.; Lv, J.; Li, H. FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution. Inf. Fusion 2025, 121, 103146. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Zhao, Z.; Yang, Q.; Nie, R.; Cao, J.; Pu, Y. AFDFusion: An adaptive frequency decoupling fusion network for multi-modality image. Expert Syst. Appl. 2025, 263, 125694. [Google Scholar] [CrossRef] [Scilit]
- Guan, F.; Zhao, N.; Wang, H.; Fang, Z.; Zhang, J.; Yu, Y.; Jiang, L.; Huang, H. Dual-branch transformer framework with gradient-aware weighting feature alignment for robust cross-view geo-localization. Inf. Fusion 2026, 127, 103808. [Google Scholar] [CrossRef] [Scilit]
- Han, P.; Chen, C. An efficient cross-view image fusion method based on selected state space and hashing for promoting urban perception. Inf. Fusion 2025, 115, 102737. [Google Scholar] [CrossRef] [Scilit]
- Gu, A.; Goel, K.; Ré, C. Efficiently Modeling Long Sequences with Structured State Spaces. arXiv 2022, arXiv:2111.00396. [Google Scholar] [CrossRef] [Scilit]
- Smith, J.T.H.; Warrington, A.; Linderman, S.W. Simplified State Space Layers for Sequence Modeling. arXiv 2023, arXiv:2208.04933. [Google Scholar] [CrossRef] [Scilit]
- Fu, D.Y.; Dao, T.; Saab, K.K.; AThomas, W.; Rudra, A.; Ré, C. Hungry Hungry Hippos: Towards Language Modeling with State Space Models. arXiv 2023, arXiv:2212.14052. [Google Scholar] [CrossRef] [Scilit]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model. arXiv 2024, arXiv:2401.09417. [Google Scholar] [CrossRef] [Scilit]
- Qiao, Y.; Yu, Z.; Guo, L.; Chen, S.; Zhao, Z.; Sun, M.; Wu, Q.; Liu, J. VL-Mamba: Exploring State Space Models for Multimodal Learning. arXiv 2024, arXiv:2403.13600. [Google Scholar] [CrossRef] [Scilit]
- Yu, W.; Wang, X. MambaOut: Do We Really Need Mamba for Vision? arXiv 2024, arXiv:2405.07992. [Google Scholar] [CrossRef] [Scilit]
- Lee, Y.; Kim, J.; Willette, J.; Hwang, S.J. MPViT: Multi-Path Vision Transformer for Dense Prediction. arXiv 2021, arXiv:2112.11010. [Google Scholar] [CrossRef] [Scilit]


















| Layer | Size | Stride | Channel (Input) | Channel (Output) | Activation |
|---|---|---|---|---|---|
| C1 | 3 | 1 | 1 | 16 | Leaky RELU |
| C2 | 3 | 1 | 16 | 32 | Leaky RELU |
| C3 | 3 | 1 | 32 | 16 | Leaky RELU |
| C4 | 3 | 1 | 16 | 1 | Leaky RELU |
| Layer | Size | Stride | Groups | Channel (Input) | Channel (Output) | Activation |
|---|---|---|---|---|---|---|
| DWConv (3 × 3 Conv) | 3 | 1 | 32 | 32 | 32 | Hardswish |
| PWConv (3 × 3 Conv) | 1 | 1 | 1 | 32 | 32 | |
| DWConv (5 × 5 Conv) | 3 | 1 | 32 | 32 | 32 | Hardswish |
| DWConv (5 × 5 Conv) | 3 | 1 | 32 | 32 | 32 | |
| PWConv (5 × 5 Conv) | 1 | 1 | 1 | 32 | 32 | |
| DWConv (7 × 7 Conv) | 3 | 1 | 32 | 32 | 32 | Hardswish |
| DWConv (7 × 7 Conv) | 3 | 1 | 32 | 32 | 32 | |
| DWConv (7 × 7 Conv) | 3 | 1 | 32 | 32 | 32 | |
| PWConv (7 × 7 Conv) | 1 | 1 | 1 | 32 | 32 | |
| 1 × 1 Conv (ConvBlock) | 1 | 1 | 32 | 16 | Hardswish | |
| DWConv (ConvBlock) | 3 | 1 | 16 | 16 | 16 | |
| 1 × 1 Conv (ConvBlock) | 1 | 1 | 16 | 32 | ||
| 1 × 1 Conv | 1 | 1 | 128 | 32 | Hardswish |
| Layer | Size | Stride | Channel (Input) | Channel (Output) | Activation |
|---|---|---|---|---|---|
| conv1 | 3 | 1 | 32 | 16 | LeakyRELU |
| conv2 | 3 | 1 | 16 | 32 | LeakyRELU |
| conv3 | 3 | 1 | 32 | 16 | LeakyRELU |
| conv4 | 3 | 1 | 16 | 32 | LeakyRELU |
| 1 × 1 conv | 1 | 1 | 64 | 32 | LeakyRELU |
| Layer | Input Dimension | Output Dimension |
|---|---|---|
| Input | [B,64,H,W] | |
| Channel Attention Mechanism | ||
| MaxPool | [B,64,H,W] | [B,64,1,1] |
| AvgPool | [B,64,H,W] | [B,64,1,1] |
| Conv2d (64 → 4) | [B,64,1,1] | [B,4,1,1] |
| ReLU | [B,64,1,1] | [B,4,1,1] |
| Conv2d (4 → 64) | [B,4,1,1] | [B,64,1,1] |
| addition & sigmoid | [B,64,1,1] × 2 | [B,64,1,1] |
| channel weighting | [B,64,H,W] × [B,64,1,1] | [B,64,H,W] |
| Spatial Attention Mechanism | ||
| Channel MaxPool | [B,64,H,W] | [B,1,H,W] |
| Channel AvgPool | [B,64,H,W] | [B,1,H,W] |
| concatenation | [B,1,H,W] × 2 | [B,2,H,W] |
| 7 × 7 conv | [B,2,H,W] | [B,1,H,W] |
| sigmoid | [B,1,H,W] | [B,1,H,W] |
| space weighting | [B,64,H,W] × [B,64,1,1] | [B,64,H,W] |
| Output | [B,64,H,W] |
| Parameter | Configuration |
|---|---|
| CPU model | Intel(R) Core(TM) i9-13900K |
| GPU model | NVIDIA GeForce RTX 4080 |
| Operating system | Windows 10 |
| Deep learning frame | Pytorch1.10.0 |
| GPU accelerator | CUDA12.7 |
| Integrated development environment | Pycharm |
| Scripting language | Python3.8 |
| Neural network accelerator | CUDNN8.2.0 |
| Parameter | Configuration |
|---|---|
| Neural network optimizer | Adam |
| Learning rate | 0.001 |
| Stage 1 Training epochs | 10 |
| Stage 2 Training epochs | 50 |
| Momentum β1 | 0.9 |
| Momentum β2 | 0.999 |
| Batch size | 4 |
| λ | 100 |
| λ1 | 2 |
| λ2 | 2 |
| λ3 | 5 |
| Datasets: MSRS Fusion Dataset | ||||||||
|---|---|---|---|---|---|---|---|---|
| EN | SD | SF | MI | SCD | VIF | Qabf | SSIM | |
| IFCNN | 6.18 | 26.61 | 4.73 | 1.73 | 1.33 | 0.73 | 0.54 | 0.48 |
| U2Fusion | 5.22 | 22.8 | 3.57 | 1.37 | 1.15 | 0.51 | 0.38 | 0.38 |
| RCGAN | 5.90 | 23.8 | 3.49 | 1.60 | 1.16 | 0.55 | 0.40 | 0.37 |
| MFEIF | 5.91 | 33.58 | 3.55 | 2.06 | 1.50 | 0.72 | 0.54 | 0.34 |
| SwinFusion | 6.62 | 42.99 | 5.01 | 3.30 | 1.69 | 0.93 | 0.66 | 0.51 |
| Ours | 6.65 | 38.98 | 5.69 | 3.35 | 1.50 | 0.89 | 0.68 | 0.48 |
| Datasets: TNO & RoadScene Dataset | ||||||||
|---|---|---|---|---|---|---|---|---|
| EN | SD | SF | MI | SCD | VIF | Qabf | SSIM | |
| IFCNN | 6.5946 | 31.402 | 7.2917 | 1.6544 | 1.7135 | 0.5852 | 0.483 | 0.4258 |
| U2Fusion | 6.5208 | 27.1557 | 6.4177 | 1.2312 | 1.7328 | 0.5468 | 0.4221 | 0.5017 |
| RCGAN | 6.4169 | 26.4547 | 4.9966 | 1.485 | 1.5423 | 0.5404 | 0.4177 | 0.4385 |
| MFEIF | 6.5388 | 30.5476 | 4.0907 | 1.8147 | 1.7716 | 0.6194 | 0.4073 | 0.5051 |
| SwinFusion | 6.6812 | 37.8414 | 6.3565 | 2.2236 | 1.7496 | 0.6589 | 0.4963 | 0.5064 |
| Ours | 7.1021 | 40.4312 | 7.7482 | 2.0288 | 1.7142 | 0.5716 | 0.4599 | 0.5093 |
| Method | IFCNN | U2Fusion | SwinFusion | MFEIF | RCGAN | Ours |
|---|---|---|---|---|---|---|
| Size (M) | 0.0836 | 0.6592 | 0.9737 | 0.1581 | 0.1137 | 0.1261 |
| FLOPs (G) | 39.93 | 405.17 | 292.53 | 25.36 | 45.60 | 50.88 |
| Time (s) | 0.201 | 0.361 | 1.467 | 1.423 | 0.238 | 0.392 |
| EN | SD | SF | MI | SCD | VIF | Qabf | SSIM | |
|---|---|---|---|---|---|---|---|---|
| Feature Extraction | ||||||||
| I. Use Mamba block instead of MPMamba block | 6.54 | 38.65 | 5.59 | 3.28 | 1.47 | 0.87 | 0.57 | 0.41 |
| Feature Fusion | ||||||||
| II. Additive strategy | 6.37 | 35.89 | 5.18 | 2.87 | 1.35 | 0.69 | 0.61 | 0.42 |
| III. Maximum value strategy | 6.26 | 37.48 | 5.45 | 3.16 | 1.27 | 0.76 | 0.65 | 0.45 |
| Ours | 6.65 | 38.98 | 5.69 | 3.35 | 1.50 | 0.89 | 0.68 | 0.48 |
| AP@0.5 | AP@0.9 | |||||
|---|---|---|---|---|---|---|
| Person | Car | All | Person | Car | All | |
| Infrared | 0.955 | 0.661 | 0.808 | 0.89 | 0.538 | 0.714 |
| Visible | 0.679 | 0.943 | 0.811 | 0.617 | 0.902 | 0.76 |
| IFCNN | 0.949 | 0.919 | 0.934 | 0.887 | 0.873 | 0.88 |
| U2Fusion | 0.959 | 0.931 | 0.945 | 0.876 | 0.898 | 0.887 |
| SwinFusion | 0.934 | 0.927 | 0.931 | 0.867 | 0.889 | 0.878 |
| MFEIF | 0.934 | 0.926 | 0.93 | 0.861 | 0.881 | 0.871 |
| RCGAN | 0.935 | 0.918 | 0.926 | 0.856 | 0.875 | 0.866 |
| Ours | 0.96 | 0.928 | 0.946 | 0.894 | 0.900 | 0.889 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
He, J.; Cheng, J.; Liu, T.; Cheng, B.; Pan, X.; Cai, Y. Mamba-Based Infrared and Visible Images Fusion Method. Remote Sens. 2026, 18, 636. https://doi.org/10.3390/rs18040636
He J, Cheng J, Liu T, Cheng B, Pan X, Cai Y. Mamba-Based Infrared and Visible Images Fusion Method. Remote Sensing. 2026; 18(4):636. https://doi.org/10.3390/rs18040636
Chicago/Turabian StyleHe, Jinsong, Jianghua Cheng, Tong Liu, Bang Cheng, Xiaoyi Pan, and Yahui Cai. 2026. "Mamba-Based Infrared and Visible Images Fusion Method" Remote Sensing 18, no. 4: 636. https://doi.org/10.3390/rs18040636
APA StyleHe, J., Cheng, J., Liu, T., Cheng, B., Pan, X., & Cai, Y. (2026). Mamba-Based Infrared and Visible Images Fusion Method. Remote Sensing, 18(4), 636. https://doi.org/10.3390/rs18040636

