MRMAFusion: A Multi-Scale Restormer and Multi-Dimensional Attention Network for Infrared and Visible Image Fusion
Abstract
1. Introduction
- We design an encoder architecture that vertically integrates improved Restormer blocks with convolutional blocks, jointly preserving local details and global dependencies during multi-scale downsampling. It preserves both local and long-range dependency information, reducing detail loss caused by downsampling across different scales.
- We adopt nested connections in both horizontal encoder and decoder, which fully utilize the features output by intermediate convolutional blocks. Additionally, we integrate the SimAm attention within the encoder’s convolutional blocks to train 3D attention weights. It is combined with the fusion strategy of 1D spatial attention and 2D channel attention to form multidimensional attention, which improves network fusion performance.
- We conducted extensive experiments on the TNO, RoadScene and NIR public datasets. Our proposed method outperforms 10 representative deep learning image fusion methods on both subjective and objective measures.
2. Related Work
2.1. Deep Learning-Based Image Fusion
2.2. Transformer-Based Fusion Methods
2.3. Multi-Dimensional Attention Mechanisms
3. Method
3.1. Network Architecture
| Algorithm 1 ESB Module (Encoder SimAM Block) |
Require: Input feature |
Ensure: Output feature
|
| Algorithm 2 RTB Module (Restormer Transformer Block) |
Require: Input feature |
Ensure: Output feature
|
3.2. Fusion Strategy
3.3. Loss Function
4. Experiments
4.1. Experimental Details
4.2. Value Selection and Ablation Experiments
4.3. Compared Methods and Evaluation Metrics
4.4. Computational Complexity Analysis
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Zhang, X.; Demiris, Y. Visible and infrared image fusion using deep learning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 10535–10554. [Google Scholar] [CrossRef]
- Zhang, Y.; Liu, Y.; Sun, P.; Yan, H.; Zhao, X.; Zhang, L. IFCNN: A general image fusion framework based on convolutional neural network. Inf. Fusion 2020, 54, 99–118. [Google Scholar] [CrossRef]
- Kang, X.; Yin, H.; Duan, P. Global–local feature fusion network for visible–infrared vehicle detection. IEEE Geosci. Remote Sens. Lett. 2024, 21, 1–5. [Google Scholar] [CrossRef]
- Chen, J.; Ding, J.; Ma, J. HitFusion: Infrared and visible image fusion for high-level vision tasks using transformer. IEEE Trans. Multimed. 2024, 26, 10145–10159. [Google Scholar] [CrossRef]
- Zhang, X.; Ye, P.; Leung, H.; Gong, K.; Xiao, G. Object fusion tracking based on visible and infrared images: A comprehensive review. Inf. Fusion 2020, 63, 166–187. [Google Scholar] [CrossRef]
- Chen, J.; Li, X.; Luo, L.; Mei, X.; Ma, J. Infrared and visible image fusion based on target-enhanced multiscale transform decomposition. Inf. Sci. 2020, 508, 64–78. [Google Scholar] [CrossRef]
- Luo, X.; Zhang, Z.; Zhang, B.; Wu, X. Image fusion with contextual statistical similarity and nonsubsampled shearlet transform. IEEE Sens. J. 2016, 17, 1760–1771. [Google Scholar] [CrossRef]
- Liu, Y.; Yang, X.; Zhang, R.; Albertini, M.K.; Celik, T.; Jeon, G. Entropy-based image fusion with joint sparse representation and rolling guidance filter. Entropy 2020, 22, 118. [Google Scholar] [CrossRef]
- Zhang, Q.; Li, G.; Cao, Y.; Han, J. Multi-focus image fusion based on non-negative sparse representation and patch-level consistency rectification. Pattern Recognit. 2020, 104, 107325. [Google Scholar] [CrossRef]
- Ma, J.; Zhou, Z.; Wang, B.; Zong, H. Infrared and visible image fusion based on visual saliency map and weighted least square optimization. Infrared Phys. Technol. 2017, 82, 8–17. [Google Scholar] [CrossRef]
- Li, H.; Wu, X.J. DenseFuse: A fusion approach to infrared and visible images. IEEE Trans. Image Process. 2018, 28, 2614–2623. [Google Scholar] [CrossRef] [PubMed]
- Li, H.; Wu, X.J.; Durrani, T. NestFuse: An infrared and visible image fusion architecture based on nest connection and spatial/channel attention models. IEEE Trans. Instrum. Meas. 2020, 69, 9645–9656. [Google Scholar] [CrossRef]
- Li, H.; Wu, X.J.; Kittler, J. RFN-Nest: An end-to-end residual fusion network for infrared and visible images. Inf. Fusion 2021, 73, 72–86. [Google Scholar] [CrossRef]
- Zhao, Z.; Xu, S.; Zhang, C.; Liu, J.; Zhang, J.; Li, P. DIDFuse: Deep image decomposition for infrared and visible image fusion. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, Yokohama, Japan, 11–17 July 2020; pp. 970–976. [Google Scholar]
- Wang, Z.; Wang, J.; Wu, Y.; Xu, J.; Zhang, X. UNFusion: A unified multi-scale densely connected network for infrared and visible image fusion. IEEE Trans. Circuits Syst. Video Technol. 2021, 32, 3360–3374. [Google Scholar] [CrossRef]
- Zhang, H.; Xu, H.; Xiao, Y.; Guo, X.; Ma, J. Rethinking the image fusion: A fast unified image fusion network based on proportional maintenance of gradient and intensity. Proc. Aaai Conf. Artif. Intell. 2020, 34, 12797–12804. [Google Scholar] [CrossRef]
- Xu, H.; Ma, J.; Le, Z.; Jiang, J.; Guo, X. FusionDN: A unified densely connected network for image fusion. Proc. Aaai Conf. Artif. Intell. 2020, 34, 12484–12491. [Google Scholar] [CrossRef]
- Xu, H.; Ma, J.; Jiang, J.; Guo, X.; Ling, H. U2Fusion: A unified unsupervised image fusion network. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 44, 502–518. [Google Scholar] [CrossRef]
- Ma, J.; Yu, W.; Liang, P.; Li, C.; Jiang, J. FusionGAN: A generative adversarial network for infrared and visible image fusion. Inf. Fusion 2019, 48, 11–26. [Google Scholar] [CrossRef]
- Ma, J.; Xu, H.; Jiang, J.; Mei, X.; Zhang, X.P. DDcGAN: A dual-discriminator conditional generative adversarial network for multi-resolution image fusion. IEEE Trans. Image Process. 2020, 29, 4980–4995. [Google Scholar] [CrossRef]
- Ma, J.; Tang, L.; Fan, F.; Huang, J.; Mei, X.; Ma, Y. SwinFusion: Cross-domain long-range learning for general image fusion via swin transformer. IEEE/CAA J. Autom. Sin. 2022, 9, 1200–1217. [Google Scholar] [CrossRef]
- Wang, Z.; Chen, Y.; Shao, W.; Li, H.; Zhang, L. SwinFuse: A residual swin transformer fusion network for infrared and visible images. IEEE Trans. Instrum. Meas. 2022, 71, 1–12. [Google Scholar] [CrossRef]
- Tang, W.; He, F.; Liu, Y.; Duan, Y.; Si, T. DATFuse: Infrared and visible image fusion via dual attention transformer. IEEE Trans. Circuits Syst. Video Technol. 2023, 33, 3159–3172. [Google Scholar] [CrossRef]
- Luo, X.; Wang, J.; Zhang, Z.; Wu, X.J. A full-scale hierarchical encoder-decoder network with cascading edge-prior for infrared and visible image fusion. Pattern Recognit. 2024, 148, 110192. [Google Scholar] [CrossRef]
- Zhou, Z.; Siddiquee, M.M.R.; Tajbakhsh, N.; Liang, J. UNet++: Redesigning skip connections to exploit multiscale features in image segmentation. IEEE Trans. Med. Imaging 2019, 39, 1856–1867. [Google Scholar] [CrossRef]
- Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.H. Restormer: Efficient transformer for high-resolution image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 5728–5739. [Google Scholar]
- Yang, L.; Zhang, R.Y.; Li, L.; Xie, X. SimAM: A simple, parameter-free attention module for convolutional neural networks. In Proceedings of the International Conference on Machine Learning, PMLR, Virtual, 18–24 July 2021; pp. 11863–11874. [Google Scholar]
- Liu, Y.; Chen, X.; Ward, R.K.; Wang, Z.J. Image fusion with convolutional sparse representation. IEEE Signal Process. Lett. 2016, 23, 1882–1886. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems 30, Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 10012–10022. [Google Scholar]
- Su, W.; Huang, Y.; Li, Q.; Zuo, F.; Liu, L. Infrared and visible image fusion based on adversarial feature extraction and stable image reconstruction. IEEE Trans. Instrum. Meas. 2022, 71, 104853. [Google Scholar] [CrossRef]
- Liu, J.; Sun, G.; Zheng, B.; Dong, L. TCIGFusion: A two-stage correlated feature interactive guided network for infrared and visible image fusion. Opt. Lasers Eng. 2025, 195, 109265. [Google Scholar] [CrossRef]
- Liu, J.; Sun, G.; Zheng, B.; Dong, L. TSDGFusion: A text and semantic dual-guided model for infrared and visible image fusion. Displays 2025, 91, 103266. [Google Scholar] [CrossRef]
- Li, H.; Wu, X.J. CrossFuse: A novel cross attention mechanism based infrared and visible image fusion approach. Inf. Fusion 2024, 103, 102147. [Google Scholar] [CrossRef]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
- Li, J.; Huo, H.; Li, C.; Wang, R.; Feng, Q. AttentionFGAN: Infrared and visible image fusion using attention-based generative adversarial networks. IEEE Trans. Multimed. 2020, 23, 1383–1396. [Google Scholar] [CrossRef]
- Tang, X.; Zhao, J.; Cui, G.; Tian, H.; Shi, Z.; Hou, C. MFAGAN: A multiscale feature-attention generative adversarial network for infrared and visible image fusion. Infrared Phys. Technol. 2023, 133, 104796. [Google Scholar] [CrossRef]
- Liu, X.; Wang, R.; Huo, H.; Yang, X.; Li, J. An attention-guided and wavelet-constrained generative adversarial network for infrared and visible image fusion. Infrared Phys. Technol. 2023, 129, 104570. [Google Scholar] [CrossRef]
- Wang, Y.; Pu, J.; Miao, D.; Zhang, L.; Zhang, L.; Du, X. SCGRFuse: An infrared and visible image fusion network based on spatial/channel attention mechanism and gradient aggregation residual dense blocks. Eng. Appl. Artif. Intell. 2024, 132, 107898. [Google Scholar] [CrossRef]
- Lou, M.; Zhang, S.; Zhou, H.Y.; Yang, S.; Wu, C.; Yu, Y. TransXNet: Learning both global and local dynamics with a dual dynamic token mixer for visual recognition. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 11534–11547. [Google Scholar] [CrossRef]
- Liu, Z.; Blasch, E.; Xue, Z.; Zhao, J.; Laganiere, R.; Wu, W. Objective assessment of multiresolution image fusion algorithms for context enhancement in night vision: A comparative study. IEEE Trans. Pattern Anal. Mach. Intell. 2011, 34, 94–109. [Google Scholar] [CrossRef]
- Roberts, J.W.; Van Aardt, J.A.; Ahmed, F.B. Assessment of image fusion procedures using entropy, image quality, and multispectral classification. J. Appl. Remote Sens. 2008, 2, 023522. [Google Scholar]
- Eskicioglu, A.M.; Fisher, P.S. Image quality measures and their performance. IEEE Trans. Commun. 1995, 43, 2959–2965. [Google Scholar] [CrossRef]
- Rao, Y.J. In-fibre Bragg grating sensors. Meas. Sci. Technol. 1997, 8, 355. [Google Scholar] [CrossRef]
- Piella, G.; Heijmans, H. A new quality metric for image fusion. In Proceedings of the 2003 International Conference on Image Processing, Barcelona, Spain, 14–17 September 2003; Volume 3, pp. 173–176. [Google Scholar]
- Han, Y.; Cai, Y.; Cao, Y.; Xu, X. A new image fusion performance metric based on visual information fidelity. Inf. Fusion 2013, 14, 127–135. [Google Scholar] [CrossRef]
- Tang, L.; Yuan, J.; Zhang, H.; Jiang, X.; Ma, J. PIAFusion: A progressive infrared and visible image fusion network based on illumination aware. Inf. Fusion 2022, 83, 79–92. [Google Scholar] [CrossRef]
- Zhang, H.; Ma, J. SDNet: A versatile squeeze-and-decomposition network for real-time image fusion. Int. J. Comput. Vis. 2021, 129, 2761–2785. [Google Scholar] [CrossRef]











| Layer | Input Channel | Output Channel | Resolution | Activation | |
|---|---|---|---|---|---|
| Encoder | Conv1 | 1 | 16 | ReLU | |
| Conv2 | 16 | 32 | ReLU | ||
| Conv3 | 32 | 48 | ReLU | ||
| Conv4 | 48 | 64 | ReLU | ||
| RTB20 | 16 | 16 | – | ||
| RTB30 | 32 | 32 | – | ||
| RTB40 | 48 | 48 | – | ||
| ESB20 | 48 | 64 | – | ||
| ESB30 | 80 | 96 | – | ||
| ESB31 | 208 | 256 | – | ||
| ESB40 | 112 | 128 | – | ||
| ESB41 | 288 | 304 | – | ||
| ESB42 | 752 | 1024 | – | ||
| Decoder | DCB10 | 80 | 16 | – | |
| DCB11 | 96 | 16 | – | ||
| DCB12 | 112 | 16 | – | ||
| DCB13 | 16 | 1 | – | ||
| DCB20 | 320 | 64 | – | ||
| DCB21 | 384 | 64 | – | ||
| DCB30 | 1280 | 256 | – |
| Dataset | EN | MI | SF | SD | VIF | Qabf | |
|---|---|---|---|---|---|---|---|
| TNO | 1 | 7.01034 | 3.68497 | 9.52341 | 41.48464 | 0.87745 | 0.53287 |
| 5 | 7.01878 | 3.73317 | 10.16536 | 41.62811 | 0.89373 | 0.53723 | |
| 10 | 7.00937 | 3.65345 | 10.08234 | 41.39382 | 0.87380 | 0.53041 | |
| 100 | 7.02507 | 3.64551 | 9.89553 | 41.92570 | 0.87154 | 0.52993 | |
| 1000 | 7.02071 | 3.71342 | 9.67823 | 41.76024 | 0.88941 | 0.53368 | |
| NIR | 1 | 7.35249 | 4.71012 | 17.24790 | 49.15533 | 0.94346 | 0.60329 |
| 5 | 7.36145 | 4.75487 | 18.55717 | 49.30941 | 0.95292 | 0.60347 | |
| 10 | 7.35683 | 4.66922 | 18.57130 | 49.11026 | 0.93332 | 0.59808 | |
| 100 | 7.35312 | 4.68580 | 17.89285 | 49.12542 | 0.93445 | 0.59828 | |
| 1000 | 7.35465 | 4.73748 | 16.82413 | 49.09030 | 0.94849 | 0.60192 | |
| RoadScene | 1 | 7.36990 | 3.78614 | 12.78249 | 49.15461 | 0.71346 | 0.50884 |
| 5 | 7.41757 | 3.86197 | 13.99177 | 50.71819 | 0.72641 | 0.50922 | |
| 10 | 7.37889 | 3.76332 | 12.78923 | 49.32286 | 0.71280 | 0.50637 | |
| 100 | 7.39970 | 3.79374 | 13.87034 | 50.25727 | 0.71312 | 0.50404 | |
| 1000 | 7.39404 | 3.82886 | 13.25628 | 49.99016 | 0.71900 | 0.50654 |
| Dataset | Configuration | EN | MI | SF | SD | VIF | Qabf |
|---|---|---|---|---|---|---|---|
| TNO | Plain Conv (NO-RTB) | 7.02100 | 3.70090 | 9.45826 | 41.80793 | 0.88441 | 0.53715 |
| 1-RTB | 7.02439 | 3.69958 | 9.78312 | 41.96736 | 0.88294 | 0.53321 | |
| 2-RTB | 7.02002 | 3.70835 | 10.22783 | 41.75639 | 0.88329 | 0.53317 | |
| Ours (3-RTB) | 7.01878 | 3.73317 | 10.16536 | 41.62811 | 0.89373 | 0.53723 | |
| NIR | Plain Conv (NO-RTB) | 7.35783 | 4.74665 | 17.12581 | 49.22418 | 0.95328 | 0.60671 |
| 1-RTB | 7.35881 | 4.76207 | 17.89427 | 49.30651 | 0.95231 | 0.60452 | |
| 2-RTB | 7.35669 | 4.72750 | 18.34712 | 49.16089 | 0.95109 | 0.60420 | |
| Ours (3-RTB) | 7.36145 | 4.75487 | 18.55717 | 49.30941 | 0.95292 | 0.60347 | |
| RoadScene | Plain Conv (NO-RTB) | 7.38180 | 3.81768 | 12.45813 | 49.62981 | 0.72311 | 0.51772 |
| 1-RTB | 7.40297 | 3.83066 | 13.23624 | 50.36707 | 0.72375 | 0.51005 | |
| 2-RTB | 7.37539 | 3.80123 | 13.78318 | 49.35612 | 0.71946 | 0.51164 | |
| Ours (3-RTB) | 7.41757 | 3.86197 | 13.99177 | 50.71819 | 0.72641 | 0.50922 |
| Dataset | Metrics | EN | MI | SF | SD | VIF | Qabf |
|---|---|---|---|---|---|---|---|
| TNO | NO-simAM | 7.00268 | 3.60090 | 8.98532 | 41.20689 | 0.85804 | 0.52504 |
| CA-Att(1D) | 7.00477 | 3.65112 | 9.34719 | 41.17492 | 0.87462 | 0.53072 | |
| SP-Att(2D) | 7.01964 | 3.70683 | 9.78295 | 41.80817 | 0.88800 | 0.53379 | |
| SC-Att | 7.00651 | 3.68526 | 10.05813 | 41.26345 | 0.88292 | 0.53387 | |
| SimAM (3D, Ours) | 7.01878 | 3.73317 | 10.16536 | 41.62811 | 0.89373 | 0.53723 | |
| NIR | NO-simAM | 7.35590 | 4.61427 | 16.23814 | 49.05813 | 0.92143 | 0.59314 |
| CA-Att(1D) | 7.35818 | 4.68576 | 17.12583 | 49.11274 | 0.93844 | 0.59818 | |
| SP-Att(2D) | 7.35627 | 4.75444 | 18.39205 | 49.12062 | 0.95043 | 0.60229 | |
| SC-Att | 7.34832 | 4.68820 | 17.89417 | 47.84280 | 0.94231 | 0.60728 | |
| SimAM (3D, Ours) | 7.36145 | 4.75487 | 18.55717 | 49.30941 | 0.95292 | 0.60347 | |
| RoadScene | NO-simAM | 7.35847 | 3.68774 | 11.56924 | 48.50586 | 0.69606 | 0.50042 |
| CA-Att(1D) | 7.37939 | 3.76182 | 12.34853 | 49.29348 | 0.71203 | 0.50494 | |
| SP-Att(2D) | 7.41484 | 3.83969 | 13.99528 | 50.03931 | 0.72135 | 0.50848 | |
| SC-Att | 7.39128 | 3.80355 | 13.62984 | 49.77502 | 0.71542 | 0.50424 | |
| SimAM (3D, Ours) | 7.41757 | 3.86197 | 13.99177 | 50.71819 | 0.72641 | 0.50922 |
| Dataset | Dwconv Scale | EN | MI | SF | SD | VIF | Qabf |
|---|---|---|---|---|---|---|---|
| TNO | Single () | 7.01263 | 3.68957 | 9.34218 | 41.35286 | 0.87847 | 0.53105 |
| Single () | 7.01728 | 3.70193 | 9.67532 | 41.51981 | 0.88172 | 0.53264 | |
| Dual ( + ) | 7.01571 | 3.71829 | 9.98015 | 41.57893 | 0.88931 | 0.53648 | |
| Multi-scale (Ours) | 7.01878 | 3.73317 | 10.16536 | 41.62811 | 0.89373 | 0.53723 | |
| NIR | Single () | 7.35083 | 4.70297 | 17.23918 | 49.01573 | 0.94256 | 0.59631 |
| Single () | 7.35476 | 4.72183 | 17.78429 | 49.12685 | 0.94609 | 0.60142 | |
| Dual ( + ) | 7.35861 | 4.74305 | 18.21073 | 49.24316 | 0.95028 | 0.60076 | |
| Multi-scale (Ours) | 7.36145 | 4.75487 | 18.55717 | 49.30941 | 0.95292 | 0.60347 | |
| RoadScene | Single () | 7.39582 | 3.80217 | 12.56342 | 50.02178 | 0.71428 | 0.50305 |
| Single () | 7.40156 | 3.82189 | 13.12067 | 50.25314 | 0.71847 | 0.50519 | |
| Dual ( + ) | 7.41281 | 3.84532 | 13.67085 | 50.53278 | 0.72351 | 0.50983 | |
| Multi-scale (Ours) | 7.41757 | 3.86197 | 13.99177 | 50.71819 | 0.72641 | 0.50922 |
| Dataset | Weighting Strategy | EN | MI | SF | SD | VIF | Qabf |
|---|---|---|---|---|---|---|---|
| TNO | Simple Average | 7.00534 | 3.67182 | 9.23847 | 41.54913 | 0.87645 | 0.52934 |
| Learned Weights | 7.01839 | 3.70823 | 9.87152 | 41.39772 | 0.88756 | 0.53418 | |
| Self-adaptive (Ours) | 7.01878 | 3.73317 | 10.16536 | 41.62811 | 0.89373 | 0.53723 | |
| NIR | Simple Average | 7.34523 | 4.68934 | 16.78519 | 48.89417 | 0.93845 | 0.59283 |
| Learned Weights | 7.36734 | 4.73821 | 17.89831 | 49.18934 | 0.94923 | 0.60156 | |
| Self-adaptive (Ours) | 7.36145 | 4.75487 | 18.55717 | 49.30941 | 0.95292 | 0.60347 | |
| RoadScene | Simple Average | 7.39185 | 3.79841 | 12.34829 | 50.01493 | 0.71382 | 0.50187 |
| Learned Weights | 7.41023 | 3.84912 | 13.56412 | 50.72581 | 0.72193 | 0.50712 | |
| Self-adaptive (Ours) | 7.41757 | 3.86197 | 13.99177 | 50.71819 | 0.72641 | 0.50922 |
| Dataset | SimAM Placement | EN | MI | SF | SD | VIF | Qabf |
|---|---|---|---|---|---|---|---|
| TNO | Decoder Only | 7.00156 | 3.65294 | 9.12934 | 41.18453 | 0.86521 | 0.52677 |
| Encoder + Decoder | 7.01245 | 3.68956 | 9.78530 | 41.75234 | 0.87892 | 0.53156 | |
| Encoder Only (Ours) | 7.01878 | 3.73317 | 10.16536 | 41.62811 | 0.89373 | 0.53723 | |
| NIR | Decoder Only | 7.34823 | 4.67334 | 16.89047 | 48.89023 | 0.93156 | 0.59479 |
| Encoder + Decoder | 7.35776 | 4.71823 | 17.78923 | 49.12793 | 0.94934 | 0.59878 | |
| Encoder Only (Ours) | 7.36145 | 4.75487 | 18.55717 | 49.30941 | 0.95292 | 0.60347 | |
| RoadScene | Decoder Only | 7.38317 | 3.76543 | 12.58494 | 49.56274 | 0.70662 | 0.49612 |
| Encoder + Decoder | 7.40183 | 3.81934 | 13.40986 | 50.37527 | 0.71523 | 0.50343 | |
| Encoder Only (Ours) | 7.41757 | 3.86197 | 13.99177 | 50.71819 | 0.72641 | 0.50922 |
| Method | Params (M) | FLOPs (G) | Latency (ms) | FPS |
|---|---|---|---|---|
| DenseFuse | 0.0742 | 5.7661 | 1.2722 | 786.02 |
| NestFuse | 2.733 | 76.211 | 14.938 | 66.95 |
| RFN-Nest | 7.524 | 111.106 | 11.169 | 89.53 |
| FusionGAN | 1.327 | 62.519 | 4.049 | 246.94 |
| PIAFusion | 0.491 | 38.658 | 2.940 | 340.11 |
| IFCNN | 0.074 | 8.544 | 0.659 | 1518.14 |
| PMGI | 0.042 | 2.7661 | 1.773 | 563.99 |
| SDNet | 0.067 | 1.8728 | 0.851 | 1174.68 |
| U2Fusion | 0.659 | 43.170 | 2.957 | 338.20 |
| UNFusion | 31.122 | 81.789 | 11.044 | 90.55 |
| Ours | 32.513 | 83.140 | 10.652 | 93.87 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Dong, L.; Sun, G.; Zhang, H.; Luo, W. MRMAFusion: A Multi-Scale Restormer and Multi-Dimensional Attention Network for Infrared and Visible Image Fusion. Appl. Sci. 2026, 16, 946. https://doi.org/10.3390/app16020946
Dong L, Sun G, Zhang H, Luo W. MRMAFusion: A Multi-Scale Restormer and Multi-Dimensional Attention Network for Infrared and Visible Image Fusion. Applied Sciences. 2026; 16(2):946. https://doi.org/10.3390/app16020946
Chicago/Turabian StyleDong, Liang, Guiling Sun, Haicheng Zhang, and Wenxuan Luo. 2026. "MRMAFusion: A Multi-Scale Restormer and Multi-Dimensional Attention Network for Infrared and Visible Image Fusion" Applied Sciences 16, no. 2: 946. https://doi.org/10.3390/app16020946
APA StyleDong, L., Sun, G., Zhang, H., & Luo, W. (2026). MRMAFusion: A Multi-Scale Restormer and Multi-Dimensional Attention Network for Infrared and Visible Image Fusion. Applied Sciences, 16(2), 946. https://doi.org/10.3390/app16020946

