In this section, the effectiveness of the proposed lightweight Mamba-INN dual-branch fusion network is evaluated on both infrared-visible image fusion and medical image fusion tasks. The experiments are designed from four perspectives: quantitative comparison, qualitative visual evaluation, lightweight performance analysis, and ablation study. Through these evaluations, we examine not only the fusion quality of the proposed method but also its computational efficiency and deployment potential. To ensure fairness, other comparison models, such as CDDFuse, were retrained using the same training set, input resolution, number of training iterations, and hardware environment as our proposed model. All complexity and speed metrics were recalculated for the same input size.
4.2. Comparative Models
To comprehensively evaluate the fusion performance of the proposed method, this paper selects several classical and state-of-the-art methods for comparison, including DenseFuse [
2], DIDFuse [
3], U2Fusion [
4], SDNet [
9], RFNet [
6], TarDAL [
5], DeFusion [
37], ReCoNet [
21], CoCoNet [
38], and CDDFuse [
1].
4.2.1. Quantitative Comparison
In the task of infrared-visible light image fusion, this paper quantitatively evaluates the proposed method on the MSRS, TNO, and RoadScene datasets and compares it with methods such as DIDFuse, U2Fusion, SDNet, TarDAL, DeFusion, ReCoNet, and CDDFuse.
As shown in
Table 1, the proposed method achieves competitive performance across all three infrared-visible datasets. On the MSRS dataset, our method obtains strong results in terms of MI, SF, VIF, and Qabf, suggesting that it can effectively integrate salient infrared targets with texture details from visible images. Compared with CDDFuse, the proposed method does not rank first on every single metric, but it maintains comparable fusion quality with a much smaller model size and lower computational cost. This demonstrates that the proposed architecture achieves a favorable balance between fusion performance and model complexity.
On the TNO dataset, our method also demonstrates strong generalization capabilities. This dataset contains a large number of nighttime, low-light, and complex background scenes, which place high demands on target enhancement and detail preservation. Experimental results show that our method maintains stable performance on metrics such as MI, SF, VIF, and Qabf, indicating that the Mamba-INN dual-branch architecture effectively balances global structural modeling and local detail preservation.
To further illustrate the overall performance on the TNO dataset, a radar chart of the main evaluation metrics is presented in
Figure 7. Each axis on the radar chart represents a normalized evaluation metric; the larger the area enclosed by the lines, the more balanced the method’s overall performance across multiple metrics. The proposed method covers a relatively large area across multiple metrics and performs particularly well in MI, SF, VIF, and Qabf. This indicates that the method achieves a balanced fusion result in terms of information preservation, detail representation, and edge structure maintenance. Although CDDFuse still has advantages in certain metrics such as EN and SD, the proposed method achieves comparable or better performance on several key perceptual and structural metrics with much lower computational complexity.
On the RoadScene dataset, our method continues to show stable quantitative performance. RoadScene contains complex traffic scenes with roads, vehicles, pedestrians, buildings, and varying illumination conditions. These characteristics require the fusion model to preserve background structures while enhancing salient infrared targets. The results show that the proposed method maintains a good level of information content and image clarity, while avoiding obvious structural distortion. Compared with traditional or lightweight methods, the proposed method achieves stronger performance in information retention and detail representation.
A comprehensive analysis of results across three IR-visible light datasets reveals that our method does not achieve the best performance on every individual metric, but rather strikes a more balanced outcome between fusion performance and model complexity. Compared to CDDFuse, our method significantly reduces the number of parameters and FLOPs by simplifying the Mamba global branch, streamlining the INN detail branch, and compressing the decoding structure, while maintaining competitive fusion quality. Therefore, our method is more suitable for IR-Vis fusion scenarios with resource constraints or high real-time requirements.
To further evaluate cross-modal adaptability, experiments were conducted on MRI-CT, MRI-PET, and MRI-SPECT fusion tasks. Medical image fusion requires preserving anatomical structures and tissue boundaries while integrating functional information. MRI provides soft-tissue details, CT highlights bone structures, and PET/SPECT reflects metabolic or perfusion responses. Thus, effective fusion should balance structural detail preservation with functional information integration.
To distinguish cross-task generalization from task-specific adaptation, two experimental settings were adopted. “Ours” denotes the model trained on the infrared-visible fusion task and directly tested on medical image fusion datasets, which evaluates the transferability of the proposed architecture. “Ours*” denotes the model retrained on medical image datasets, which evaluates its adaptability to medical modality distributions. Similarly, CDDFuse* represents the retrained version of CDDFuse on medical image fusion datasets. This setting provides a clearer comparison between inherent architectural generalization and task-specific optimization.
As shown in
Table 2, the proposed method achieves stable quantitative results across MRI-CT, MRI-PET, and MRI-SPECT fusion tasks. For MRI-CT fusion, our method effectively combines soft-tissue information from MRI with high-density structural information from CT, producing fused images with clear contours and comprehensive anatomical representation. For MRI-PET and MRI-SPECT fusion, the model introduces significant functional responses from PET or SPECT while preserving the structural details of MRI. These results suggest that the proposed Mamba-INN dual-branch architecture can balance structural preservation and functional information integration.
The comparison between Ours and Ours* further shows that retraining on medical datasets can improve the model’s adaptation to medical modality distributions. Nevertheless, even without medical-specific retraining, the directly transferred model still achieves competitive performance, which indicates that the proposed dual-branch modeling strategy has a certain degree of cross-modal transferability. The Mamba-inspired branch contributes to large-scale structural dependency modeling, while the lightweight INN branch helps retain local edges and fine details. Their complementary roles allow the model to adapt to different types of multimodal image fusion tasks.
To further validate the generalization ability of the proposed model on unseen datasets, we conducted a zero-shot evaluation on the official LLVIP [
39] test set. No additional training or fine-tuning was performed on the model during the evaluation. The results are shown in
Table 3. As can be seen from the table, although our model was not trained on the LLVIP dataset, its fusion performance still demonstrates good stability. For the EN metric, Ours achieved a score of 7.35, only slightly lower than CDDFuse’s 7.44; while it is slightly lower than CDDFuse on the SD metric, it achieves 0.68 and 0.91 on structural and edge preservation metrics such as Qabf and SSIM, respectively, demonstrating an advantage in preserving detail and structural information. This indicates that the proposed Mamba-INN dual-branch architecture can maintain relatively stable fusion quality when faced with unseen infrared-visible light datasets, demonstrating a certain degree of cross-dataset generalization capability.
To further evaluate the naturalness of the fused images, we employed the NIQE metric for testing.
Table 4 lists the NIQE test results for different methods across four datasets, comparing CDDFuse with our proposed method (Ours). The experimental results show that our method achieves NIQE values comparable to those of CDDFuse on most datasets, indicating that it delivers stable and competitive performance in preserving image naturalness.
In summary, the proposed method demonstrates good stability and adaptability in medical image fusion. Although it does not obtain the best result on every individual metric, it maintains competitive fusion quality with a significantly smaller number of parameters and lower computational cost than CDDFuse and CDDFuse*. These results indicate that the proposed lightweight architecture is not only effective for infrared-visible fusion but also has potential for medical multimodal fusion scenarios where computational efficiency and inference speed are important.
4.2.2. Qualitative Comparison
To provide a more intuitive evaluation of visual fusion quality, qualitative comparisons were conducted for both infrared-visible and medical image fusion tasks. Compared with quantitative metrics, visual comparison can better reflect target saliency, texture preservation, structural clarity, contrast balance, and artifact suppression in fused images.
For the infrared-visible fusion task on the TNO dataset, the visual results are shown in
Figure 8. The columns in the figure correspond to the source image and the results of different fusion methods, respectively, and are used to compare the saliency of infrared targets and the ability to preserve visible-light textures. Some images in the comparison were obtained from CoCoNet [
38]. It can be observed that different methods show different fusion tendencies. Some methods enhance infrared targets effectively but weaken visible background textures and edge structures. Other methods preserve part of the visible details but fail to sufficiently highlight salient infrared targets, resulting in weak contrast between targets and background. In contrast, the proposed method enhances infrared targets while retaining texture and structural information from visible images. In low-light or complex-background regions, road surfaces, building boundaries, and object contours are clearly preserved, and no obvious brightness imbalance or over-smoothing can be observed.
A closer inspection of the locally enlarged regions further confirms the advantage of the proposed method. Traditional fusion methods tend to produce blurred edges, discontinuous textures, or insufficient local contrast in detail-rich regions. CDDFuse achieves strong visual quality, but its network structure is relatively complex. The proposed method obtains comparable visual results with a much more compact architecture. This benefit mainly comes from the cooperative design of the lightweight Mamba-inspired branch and the INN detail branch. The Mamba-inspired branch helps maintain large-scale structural consistency, while the INN branch reduces the loss of high-frequency textures and edges during fusion.
To further analyze how information from the two source modalities is inherited in the fused image, residual maps between the fused image and the infrared and visible images are presented in
Figure 9. The top row shows the infrared image, the visible light image, and the fused image; the bottom row shows the residual maps between the fused image and the infrared image, and between the fused image and the visible light image, respectively. The residual map between the fused image and the infrared image mainly highlights background structures and texture regions, indicating that visible-light details are effectively introduced into the fused result. Meanwhile, the residual map between the fused image and the visible image is mainly concentrated around foreground target contours and their surrounding regions, suggesting that infrared thermal targets are successfully enhanced. These residual distributions provide additional evidence that the proposed method can achieve complementary fusion between infrared saliency and visible texture information.
For medical image fusion, the visual comparisons are shown in
Figure 10. MRI images primarily show anatomical structures and tissue boundaries, while PET images primarily provide information on functional metabolism. Different methods present obvious differences in anatomical structure preservation and functional information integration. Some methods enhance high-response regions from functional modalities but tend to weaken MRI structural details or blur tissue boundaries. Other methods preserve MRI structures but fail to adequately represent PET or SPECT functional responses. In contrast, the proposed method preserves anatomical structures and tissue boundaries from MRI while incorporating complementary information from CT, PET, or SPECT. The resulting fused images exhibit clearer structural contours and more complete functional representation.
For MRI-CT fusion, the proposed method simultaneously retains MRI soft-tissue structures and CT high-density bone boundaries, improving the structural clarity and stability of the fused images. For MRI-PET and MRI-SPECT fusion, the proposed method highlights functional response regions while preserving MRI anatomical details, avoiding excessive smoothing of functional information or masking of structural information. These visual results indicate that the proposed method is suitable not only for target-texture fusion in infrared-visible scenarios but also for structure-function fusion in medical multimodal imaging.
Overall, the qualitative results show that the proposed method achieves stable and balanced visual performance. Compared with traditional or lightweight fusion methods, it better preserves edges, textures, and structural information. Compared with more complex models such as CDDFuse, it achieves comparable visual quality with significantly fewer parameters and lower computational complexity. This further confirms the effectiveness of the proposed lightweight Mamba-INN dual-branch architecture.
4.3. Lightweight Performance Comparison
To evaluate the lightweight advantage of the proposed model, the number of parameters and FLOPs were compared with those of mainstream image fusion methods. The results are reported in
Table 5. The parameter and FLOP values of some comparison methods were obtained from CoCoNet [
38].
As shown in
Table 5, the proposed model contains only 0.24 M parameters and requires 24.04 GFLOPs. Both values are substantially lower than those of U2Fusion, DenseFuse, SwinFusion, FMamba-S, FMamba-L, CDDFuse, and other comparison methods. Compared with CDDFuse, a representative state-of-the-art dual-branch fusion model, our method reduces the parameter count by approximately 79.8% and the computational complexity by approximately 79.5%. This means that the proposed model achieves an overall complexity reduction of nearly 80% while maintaining competitive fusion quality.
To further analyze the relationship between model complexity and fusion performance,
Figure 11 provides a comprehensive comparison of the number of parameters, VIF, and MI metrics across different methods. The horizontal axis represents the number of model parameters; the further to the left, the lighter the model. The vertical axis represents VIF; higher values indicate better visual information fidelity. The size of the bubbles represents MI; larger bubbles indicate that more mutual information has been preserved. Our method achieves high VIF values and large MI bubbles even with a low number of parameters, indicating that it maintains good information retention and fusion quality while significantly reducing model complexity.
It is worth noting that, compared with CDDFuse, the proposed method achieves a comparable or higher VIF value with far fewer parameters while maintaining a high MI level. This demonstrates that the proposed lightweight dual-branch framework can effectively reduce model scale without significantly compromising fusion quality. The performance-complexity comparison confirms that the proposed method has strong practical potential for deployment on computationally constrained platforms.
The lightweight advantage mainly comes from three aspects. First, the simplified Mamba-inspired branch reduces the computational burden of global modeling. Second, the streamlined INN branch preserves local details with limited parameter overhead. Third, the compact encoder–decoder structure, module reuse strategy, and low-dimensional channel design jointly reduce redundant computation in the feature extraction, fusion, and reconstruction stages. These designs allow the model to achieve both low complexity and high fusion performance, making it suitable for real-time image fusion on embedded devices, mobile terminals, and other resource-limited platforms.
4.4. Computational Efficiency and Scalability
To further assess the computational efficiency of the proposed model, this section analyzes its parameter count, FLOPs, inference latency, and frame rate under different input resolutions. The comparison in
Table 5 has already shown that the proposed method has significantly fewer parameters and lower computational complexity than most mainstream fusion methods. In particular, compared with CDDFuse, the proposed method reduces both parameters and FLOPs by about 80%, which verifies the effectiveness of the simplified Mamba branch, lightweight INN branch, and module reuse strategy.
In addition to fixed-resolution complexity comparison, cross-resolution efficiency tests were conducted to evaluate the scalability of the proposed model. Four input sizes were used: 128 × 128, 256 × 256, 512 × 512, and 1024 × 1024. The parameter count, FLOPs, inference latency, and FPS were recorded for each resolution. All tests were conducted on an NVIDIA GeForce RTX 3060 Laptop GPU with a batch size of 1. To obtain stable latency measurements, each input size was tested after 30 warm-up runs, followed by 100 formal inference runs. The final results were averaged. FLOPs were measured using the THOP tool, with two single-channel images as model input.
As shown in
Table 6, the parameter count remains constant at 0.24 M as the input resolution increases. This demonstrates that the model size is independent of spatial resolution and confirms the structural compactness of the proposed network. FLOPs increase steadily with input size. When the input resolution increases from 128 × 128 to 256 × 256, the number of pixels increases by four times, and the FLOPs rise from 6.01 G to 24.04 G. When the resolution further increases to 512 × 512 and 1024 × 1024, the FLOPs reach 96.16 G and 384.63 G, respectively. This trend is consistent with the increase in spatial resolution, indicating that the computational growth of the proposed model is stable and predictable.
In terms of inference speed, the proposed method achieves an average latency of 23.00 ms at 128 × 128 resolution, corresponding to 43.48 FPS. This indicates that the model can meet high real-time requirements under small input sizes. At the commonly used 256 × 256 resolution, the latency is 94.74 ms, and the frame rate reaches 10.56 FPS, which still reflects reasonable inference efficiency on a laptop-class GPU. When the input size increases to 512 × 512 and 1024 × 1024, the latency rises to 381.35 ms and 1540.04 ms, while the FPS decreases to 2.62 and 0.65, respectively. Although high-resolution inputs still introduce considerable computational pressure, the model does not show abnormal complexity growth.
To further address deployment requirements in resource-constrained scenarios, this paper builds upon existing resolution scalability experiments by adding CPU-only constrained inference tests. Unlike the multi-resolution inference experiments in
Table 6, which were conducted using an NVIDIA GeForce RTX 3060 Laptop GPU, this experiment is performed entirely in a CPU environment without GPU acceleration to more closely simulate inference conditions under computational resource constraints. During testing, the input resolution was fixed at 256 × 256, the batch size was set to 1, and the number of CPU threads was set to 1, 2, 4, and 8, respectively, to analyze the model’s inference latency, frame rate, model storage size, and additional memory usage under different computational resource configurations. It should be noted that this experiment is not equivalent to deployment testing on real embedded hardware platforms; its purpose is to provide a reproducible analysis of CPU-constrained inference, thereby further validating the deployment potential of the proposed lightweight model.
Table 7 presents the results of the CPU-only constrained inference tests. As shown, the storage size of the proposed model is only 0.9925 MB, indicating that the model has a small storage footprint and is suitable for model loading and deployment in resource-constrained scenarios. Under single-threaded CPU conditions, the model’s average inference latency was 2605.30 ms, with a frame rate of 0.3838 FPS; when the number of CPU threads was increased to 2, the inference latency decreased to 1575.28 ms, and the frame rate improved to 0.6348 FPS; when the number of threads was further increased to 4, the inference latency decreased to 862.23 ms, and the frame rate increased to 1.1598 FPS; under 8-thread conditions, the model achieved the best inference efficiency under the experimental setup, with the average inference latency further reduced to 497.23 ms and the frame rate increased to 2.0111 FPS. Compared to the single-threaded configuration, inference latency under the 8-thread condition was reduced by approximately 80.9%, indicating that the model can effectively leverage multi-threaded CPU computing resources.
Furthermore, in terms of additional memory usage, the memory overhead under different thread configurations remained within a manageable range, at 345.30 MB, 247.92 MB, 266.86 MB, and 372.38 MB, respectively. Although the inference speed under CPU-only conditions remains lower than that on GPU platforms, these experimental results demonstrate that the method proposed in this paper possesses stable CPU inference capabilities and good multithreading scalability while maintaining a compact model size. Therefore, combined with the analysis in
Table 5 and
Table 6 regarding the number of parameters, FLOPs, and inference efficiency at different resolutions, this further demonstrates that the lightweight Mamba-INN dual-branch network proposed in this paper has certain application potential in resource-constrained deployment scenarios.
Overall, the proposed method exhibits good real-time potential for small- and medium-resolution inputs and maintains stable computational scalability at higher resolutions. It should be noted that inference latency depends on multiple factors, including hardware platform, deep learning framework, operator optimization, and implementation details. Therefore, the latency results reported here are mainly used to analyze the computational trend of the model under different input scales rather than to provide hardware-independent speed conclusions. Nevertheless, the results confirm that the proposed Mamba-INN dual-branch framework provides a compact and efficient model basis for real-time or near-real-time multimodal image fusion.
Furthermore, a Pareto front analysis was conducted on the MRI-PET dataset to evaluate the relationship between model complexity and fusion performance, as shown in
Figure 12. The horizontal axis represents the number of model parameters and is plotted on a logarithmic scale; the vertical axis represents the overall performance score, which is calculated based on multiple normalized fusion metrics. The green dashed line indicates the Pareto front; methods located on this front achieve an optimal trade-off between the number of parameters and performance. Our method lies on the Pareto front and achieves a high overall score with a small number of parameters, demonstrating its advantages in terms of model efficiency and performance stability for medical image fusion tasks.
Since the magnitude and distribution ranges of the aforementioned six metrics vary greatly (e.g., SF is typically greater than 20, while Qbaf is less than 1), simply summing them lacks scientific justification. Therefore, we propose a method for calculating a normalized Comprehensive Performance Score. First, we use Min-Max Normalization to map the
i-th metric
of the
j-th method to the interval [0, 1]:
where
represents the raw score of the
i-th method on the
j-th evaluation metric,
represents the normalized score, and
represents the number of evaluation metrics.
represents the overall evaluation score for method
i. For metrics where higher values are better, positive normalization is applied.
Next, we calculate the average of the normalized metrics and map them linearly to a standard 100-point scale ranging from 60 to 100 to enhance the clarity of the visualization. The final composite
for the
j-th method is calculated as follows (where
N = 6 is the total number of selected metrics):
As illustrated in
Figure 12, the horizontal axis represents the number of parameters on a logarithmic scale, while the vertical axis denotes the comprehensive performance score. The proposed method lies on the Pareto frontier and forms a favorable boundary in terms of both efficiency and performance. Specifically, it achieves the highest composite score of 99.5 while using only 0.24 M parameters, making it the lightest model among the compared methods. In contrast, CDDFuse obtains the second-highest score of 95.2 but requires 1.19 M parameters, nearly five times that of the proposed model. This result further confirms that the proposed decoupled Mamba-INN architecture can achieve strong fusion performance with a substantially reduced model size.
To provide a clearer numerical comparison corresponding to the Pareto analysis,
Table 8 reports the parameter counts and comprehensive performance scores of different methods on the MRI-PET dataset. The comprehensive score is calculated based on six normalized evaluation metrics, including SF, MI, SCD, VIF, Qabf, and SSIM. This table allows a direct comparison between fusion performance and model complexity, thereby further illustrating the efficiency advantage of the proposed method.
4.5. Ablation Studies
To validate the effectiveness of each key module, we conducted ablation experiments on the infrared-visible light image fusion task, with the results shown in
Table 8. In this table, “heavy baseline” refers to the baseline model without any lightweight design; “w/o Mamba,” “w/o INN,” and “w/o module reuse” denote the removal of the Mamba branch, the INN branch, and the module reuse strategy, respectively; and “Full module” refers to the complete model proposed in this paper.
As shown in
Table 9, the heavy baseline has the highest model complexity, with 1.19 M parameters and 116.85 G FLOPs, but its fusion performance is not optimal. In contrast, the Full module contains only 0.24 M parameters and 24.04 G FLOPs, representing reductions of approximately 79.83% and 79.43% compared to the heavy baseline, respectively. It achieves the best results on Qabf while maintaining competitive performance on MI and VIF, indicating that our method can maintain stable fusion quality while significantly reducing complexity.
Specifically, the w/o Mamba variant achieves the highest MI value, but its Qabf is lower than that of the Full module, indicating that the model’s ability to model global structure and maintain edge consistency declines without the Mamba branch. The w/o INN variant achieves the highest SF, but its Qabf remains lower than that of the Full module, suggesting that higher spatial frequency does not necessarily correspond to better edge information propagation, and that the INN branch plays a positive role in preserving details and textures. For the w/o module reuse variant, both the number of parameters and FLOPs are higher than those of the Full module, and the VIF drops to 0.76, indicating that the module reuse strategy can effectively reduce redundant computations while maintaining consistency in the feature extraction and fusion processes.
In summary, the Mamba branch, INN branch, and module reuse strategy play crucial roles in global structure modeling, local detail preservation, and model lightweighting, respectively. The full model achieved optimal Qabf and stable overall performance with the lowest number of parameters and computational cost, validating the effectiveness of the proposed lightweight Mamba-INN dual-branch architecture.