In this section, we first describe the experimental settings and provide comprehensive details of our training procedure. Subsequently, we systematically evaluate the proposed method’s fusion performance both qualitatively and quantitatively on multiple public datasets and typical degradation scenarios. Finally, we conduct ablation studies to validate the effectiveness of our approach.
4.2. Fusion Performance Evaluation
This section comprehensively evaluates the fusion performance of our method across multiple test datasets through systematic qualitative and quantitative analyses.
Qualitative Comparisons:
Figure 4 shows qualitative comparisons with several state-of-the-art fusion methods on multiple datasets. The red boxes mark representative local regions. In these regions, our method, enabled by prompt-driven condition-aware modeling, exhibits three key advantages. First, it preserves fine-grained details, as shown by the clearer ground textures in the first group. Second, it enhances structural delineation, producing more discernible contours of the fog-obscured person in the fourth group. Third, it improves chromatic fidelity, yielding more natural and coherent color reproduction in the second group. Methods such as CDDFuse, BDLFusion, CMTFusion, and LRRNet exhibit reduced detail preservation in certain regions, resulting in smoother appearances, as evidenced by the highlighted regions. IGNet presents variations in color distribution in some cases, which are distinctly reflected in the sky region of the second group. Although Text-IF and LDFusion incorporate textual guidance, their results still show limited clarity in fine-scale structures, particularly at edges and around small objects. MUFusion shows slight reductions in contrast and structural detail in multiple cases, particularly in the wheel region of the third group, which may reduce its distinguishability. TarDAL and YDTR demonstrate different smoothing characteristics. TarDAL tends to preserve higher intensity responses, as observed in the brighter human regions in the first group. In contrast, YDTR produces smoother results with reduced local structural distinctness. DAFusion enhances brightness in low-light regions; however, slight color deviations are perceptible in the first and third groups, and the wheel boundary in the third group appears less well-defined.
Quantitative Comparisons:
Table 1 presents quantitative results on four benchmark datasets. Although the optimal methods vary across individual metrics and datasets, our method achieves the best overall ranking (RoR) on all datasets. In particular, notable advantages are observed in the MI and VIF metrics. The MI metric indicates that SF
2M improves the preservation of informative content, while the VIF metric suggests better alignment with human visual perception in structural representation and detail reconstruction. For structural preservation, our method achieves the best SCD metric on the MSRS dataset and near-optimal results on the remaining datasets, reflecting consistent structural integrity under diverse imaging conditions. The highest AG metric is achieved on the LLVIP dataset, indicating that ADPM effectively adapts to low-illumination conditions and enhances edge and detail representation. On the TNO dataset, our method achieves the highest SF metric, suggesting effective recovery of high-frequency details. Although it does not achieve the highest EN metric, our method attains near-leading performance, indicating competitive information richness. In comparison, LDFusion attains high EN and SF metrics with textual guidance, indicating strong high-frequency responses and information representation; however, its performance in the SCD, MI, and VIF metrics suggests a limited capacity to balance structural preservation and noise suppression. DAFusion attains relatively high EN and AG metrics, reflecting effectiveness in content restoration and detail enhancement under certain conditions, but remains inferior in overall performance. Collectively, the consistent first-place RoR ranking across the evaluated datasets suggests the effectiveness of our method, indicating its capability to jointly preserve informative content, structural integrity, and high-frequency details.
4.3. Fusion Performance Under Degraded Conditions
We evaluated our method with several image restoration and fusion methods. To address various image degradations, we applied representative restoration models to preprocess the source images before fusion: NeuralBR [
65] for low-light enhancement, DTR [
66] for raindrop removal, AirNet [
67] for denoising and contrast adjustment, and IAT [
68] for exposure correction. All methods were evaluated using publicly available pretrained models.
Qualitative Comparisons:
Figure 5 presents qualitative comparisons between our method and several image restoration and fusion approaches under different conditions. The red boxes indicate representative local regions for comparison. Within these regions, our method exhibits significant advantages in preserving salient targets and representing fine-grained details. In contrast, DAFusion and Text-IF demonstrate limited structural information retention under challenging scenarios. In particular, under degradations such as rain and noise, fine-scale structural consistency is compromised, as evidenced in the fourth group where the overhead wires exhibit reduced continuity. In addition, residual noise persists in the outputs of DAFusion, as observed in the fifth group, indicating insufficient noise suppression. Meanwhile, under low-contrast infrared conditions, Text-IF shows reduced discriminability between targets and the background in the third group, reflecting inadequate target saliency representation. Other methods (e.g., TarDAL and CDDFuse) rely on restoration-based preprocessing and tend to preserve already visible content in complex scenarios, while their capability to represent structures in low-visibility regions remains limited.
Quantitative Comparisons:
Figure 6 presents quantitative comparisons between our method and several image restoration and fusion methods under various conditions. Our method maintains consistently competitive performance across all evaluation metrics. The advantages in the MI and VIF metrics indicate that our method achieves more effective information retention and better alignment with human visual perception, which can be attributed to the prompt-driven condition-aware modeling. Meanwhile, the NIQE and PIQE results remain competitive, demonstrating its capability to generate high-quality fused images. In comparison, although DAFusion and TarDAL achieve competitive NIQE scores in specific scenarios, their relatively weaker performance in other metrics suggests limited overall balance and consistency. Similarly, while LDFusion performs well in PIQE, its inferior MI and VIF indicate a weaker ability to preserve informative and structurally relevant content. Overall, our method achieves leading performance in MI and VIF while maintaining competitive results in NIQE and PIQE, resulting in a more balanced trade-off across different evaluation criteria under diverse imaging conditions.
4.4. Results of Infrared–Visible Object Detection
To evaluate infrared and visible image fusion for multimodal object detection, we conducted experiments on the LLVIP dataset, which contains low-light street scenes. The dataset was split into 3000 image pairs for training, 300 for validation, and 163 for testing. Each fusion method was paired with a YOLOv5 detection model, with all models trained under identical settings. We evaluated performance using Precision (P), Recall (R), mAP@0.5, and mAP@0.5:0.95. As shown in
Figure 7, our method achieves superior detection accuracy under low-light conditions, demonstrating its effectiveness in challenging visual environments. As reported in
Table 2, the quantitative results on the LLVIP dataset further substantiate this observation, with our method achieving the top ranking in both mAP@0.5 and mAP@0.5:0.95, thereby indicating consistently stable detection performance across a range of IoU thresholds. This advantage stems from the frequency-domain fusion strategy, which preserves structural consistency and fine details, enabling more discriminative and reliable features for detection. In terms of the P metric, our method ranks second, indicating accurate localization capability. However, the relatively lower R score suggests reduced target recall, indicating that some targets may be missed in certain scenarios.
4.5. Ablation Study
Importance of each component: To assess the contribution of each component, we perform ablation experiments on the LLVIP dataset, as summarized in
Table 3. We begin with a baseline model based on parameter-matched
convolutional fusion, where all evaluation metrics remain at relatively low levels. The performance improves steadily as SF
2M, ADPM, and CSAC are progressively incorporated. The inclusion of SF
2M yields consistent improvements across multiple metrics, owing to its ability to effectively integrate structural and detail information in the frequency domain. The incorporation of ADPM and CSAC leads to further performance gains, underscoring their efficacy in enriching the representation of degradation-related information. The full model, which integrates all three modules, achieves the best overall performance. These results show that each component contributes positively and that their combination leads to consistent performance gains.
Impact of ADPM and CSAC: To further analyze the roles of ADPM and CSAC, we design additional comparative experiments. In addition to removing each module individually, we replace ADPM with static prompts for comparison, as shown in
Table 4 and
Table 5. The results indicate that, although static prompts achieve marginal gains on certain metrics compared to ADPM alone, the joint integration of ADPM and CSAC leads to more consistent overall performance across evaluation criteria, highlighting their effectiveness in jointly modeling and mitigating adverse imaging conditions.
To further assess the degradation discrimination capability of ADPM and CSAC, we employ t-SNE visualization to analyze feature distributions. As illustrated in
Figure 8, removing ADPM leads to significant overlap of feature embeddings across different image conditions. When ADPM is replaced with static prompts, the boundaries between different categories remain ambiguous, resulting in limited discriminability of the feature distributions. Furthermore, removing CSAC causes the feature boundaries of degradation types such as Noise and Rain to become more blurred, making them difficult to distinguish effectively. In contrast, the full model produces more compact intra-class distributions while significantly enlarging the inter-class separation under the same conditions. These results substantiate the effectiveness of ADPM and CSAC in jointly capturing adverse imaging variations and enhancing feature discriminability.
Effectiveness of FFU and PFU: To evaluate the effectiveness of the FFU and PFU modules, we conduct ablation experiments on the MSRS dataset, with results summarized in
Table 6. In these experiments, PFU is replaced with a parameter-matched
convolution. The results show that removing either FFU or PFU results in performance deterioration across all metrics, with varying degrees of decline. Notably, FFU contributes more substantial gains, highlighting its role in enhancing informative frequency components while suppressing noise-dominated regions. The inclusion of PFU further promotes cross-modal information integration, improving structural consistency in the fused results. Combining both modules yields additional improvements, demonstrating their complementary roles in enhancing overall fusion performance.
Hyperparameter analysis: We conduct an ablation study on the weighting hyperparameter
in the overall training objective. Specifically,
is varied from 1.3 to 1.7 with a step size of 0.1. As shown in
Table 7, the model performance varies with different
values. Performance improves as
increases from 1.3 to 1.5, and then decreases when
is further increased. The best results are obtained at
across most evaluation metrics. This setting provides a balanced contribution between the fusion loss and the CSAC loss
.
We further analyze the impact of the threshold hyperparameter
on model performance. As shown in
Table 8, increasing
from 0.7 to 0.9 leads to consistent improvements across all metrics, with the optimal overall performance attained at a value of 0.9. When
is further increased to 1.0, most metrics exhibit noticeable performance deterioration, although VIF shows a marginal gain. These results indicate that a moderate
facilitates the selection of informative prompts, whereas a larger
may introduce redundancy or less relevant signals. Accordingly,
is adopted in all experiments.
Analysis of Fusion Losses: We conduct an ablation study on four fusion loss components, including the intensity loss (
), structural similarity loss (
), color consistency loss (
), and gradient loss (
). These losses constrain the fusion process from different aspects. As shown in
Table 9, the removal of any loss term results in varying degrees of performance deterioration. In contrast, integrating all four loss terms yields superior performance across most metrics, indicating that their complementary constraints enhance feature optimization and overall fusion quality.
Evaluation of runtime and computational complexity: To evaluate the computational efficiency of our method, we select a representative text-guided fusion approach, Text-IF, as a baseline for comparison.
Table 10 reports the number of learnable parameters, FLOPs, and runtime for different methods across multiple datasets. Our method achieves comparable runtime while maintaining relatively low computational cost, indicating a balanced trade-off between performance and efficiency.