Abstract
RGBT segmentation is a challenging task in the area of computer vision. Current advanced networks for RGBT segmentation focus on extracting deeper discriminative features from RGB and thermal images to provide richer semantic information for the fusion features to the decoder. However, excessively mining deeper semantic features only makes the model redundant. Simultaneously, lacking shallow spatial features leads to difficulties in guaranteeing accurate localization of targets. We believe that the features provided by images can be categorized into three types: edge, patch, and semantics. Only by synchronously taking into account the extraction of all three types of features can models achieve accurate classification on the basis of precise localization. Therefore, we propose Depth Adaption SegNet for RGB-T Segmentation (DASNet). According to the characteristics of the three types of features, we extract semantics, patch, and edge features from the deep, middle, and shallow stages respectively. We specifically design the cross-attention semantics module, patch activation module, and edge enhancement module to perform feature extraction. In addition, in order to efficiently fuse features from different categories, we design a deep-emphasis fusion module to fuse the output features of the modules. Compared to advanced methods, qualitative and quantitative experiments show that DASNet exhibits state-of-the-art performance on the CNN-based RGBT segmentation task.
1. Introduction
RGBT segmentation is a challenging task of computer vision [1,2,3]. The objective is to assign a class label to each pixel in the RGB image using both the RGB and corresponding thermal infrared (TIR) images In conditions of adverse nocturnal weather and low-light environments, RGB images exhibit limitations within the realm of applicability. In light of this, the task incorporates TIR images to provide additional discriminative features for segmentation tasks under conditions such as solar glare and low-light environments. Common RGBT models are typically constructed in a bottom-up structure, wherein input features are passed through predefined modules, with the output of each module serving as the input to the next. Throughout this process, progressively deeper feature representations of the input data are extracted, culminating in the application of upsampling operations to generate the final mask.
While this architecture enables the model to learn more abstract and higher-level features from the raw data, it overlooks the distinctions among features at different levels. We take inspiration from [4] that features at different levels exhibit varying degrees of sensitivity to different types of information. For instance, shallow-level features are most compatible with edge information and can extract rich boundary details. Similarly, intermediate-level features are effective at capturing patch information, while deep-level features exhibit higher sensitivity to semantic information. Simultaneously, a module is necessary to merge different categories of features.
Upon this foundation, we propose DASNet for RGBT segmentation. We design three distinct feature extraction modules, denoted as cross-attention semantics module (CSM), patch activation module (PAM), and edge enhancement module (EEM), tailored for processing feature maps at different depths, capturing boundary, patch, and semantic information, respectively. Additionally, we design a deep-emphasis fusion module (DEF) to amalgamate the output features from these various modules. Ultimately, the final mask is generated through a decoder.
The main characteristics and contributions of the proposed method are summarized as follows:
- We propose DASNet to accomplish the RGBT segmentation task. This network employs a novel hierarchical divergence structure, utilizing differentiation modules to handle features at various depths, maximizing the exploitation of feature potentials, thereby providing an enhanced set of discriminative features for segmentation.
- We design three distinct differentiation modules: The CSM focuses on semantic feature extraction, the PAM concentrates on capturing patch features within features, and the EEM is specifically for extracting boundary information from input images. Furthermore, these modules incorporate fusion operations for RGB and TIR modalities to comprehensively capture multi-modal characteristics.
- We design a DEF module to integrate the outputs of the differentiation modules. This fusion module not only accomplishes the fusion of diverse output features but also introduces multi-scale insights into the model, thereby enhancing model diversity.
2. Related Work
In this section, we review the literature about RGB semantic segmentation and RGB-T semantic segmentation.
2.1. RGB Semantic Segmentation
Recently, RGB semantic segmentation based on convolutional neural networks has made amazing progress by many researchers. Long et al. proposed the Fully Convolutional Network (FCN) [5]. FCN is the first end-to-end CNN-based semantic segmentation method, which integrates the semantic features from the deep, coarse network layers with the surface-level details gleaned from the shallow layers. Badrinarayanan et al. proposed SegNet [6]. The encoder of SegNet employs the computed pooling indices from the corresponding encoder’s max-pooling steps to perform nonlinear upsampling. Ronneberger et al. proposed U-Net [7], which comprises a contracting pathway that captures contextual information and a symmetric expansive pathway that facilitates precise localization.
Inspired by the above works, RGB semantic segmentation has made great progress. Furthermore, researchers endeavor to employ more efficient means for feature extraction. Chen et al. introduced DeepLab v1 [8], where they first proposed dilated convolutions. The authors also employed a conditional random field as a post-processing step. Chen et al. introduced Atrous Spatial Pyramid Pooling in DeepLab v2 [9], employing varying dilation rates in different branches to obtain multi-scale image representations. Chen et al. proposed DeepLab v3 [10]. To introduce diverse dilation rates, the authors employed a multigrid approach within residual blocks. They also incorporated image-level features into the Atrous Spatial Pyramid Pooling module.
In addition, researchers have made significant strides in improving the architecture of neural networks. Zhao et al. proposed PSPNet [11]. The model modifies the underlying ResNet [12] architecture by introducing dilated convolutions, which incorporates auxiliary losses in the middle layers of ResNet to optimize holistic learning. Lin et al. proposed RefineNet [13], which utilizes multi-resolution inputs by merging extracted features. The network effectively captures background information from a larger image region by chained residual pooling.
2.2. RGB-T Semantic Segmentation
With the popularization of sensors and imaging technology, many researchers have introduced TIR images to complete the challenging segmentation task in RGB semantic segmentation. Ha et al. proposed MFNet [14], which was the first network designed for RGBT segmentation. The authors established a novel RGB-T dataset for that task. Sun et al. proposed RTFNet [15], which employs a novel encoder–decoder architecture. Sun et al. proposed FuseSeg [16], which utilizes a Monte Carlo dropout technique to construct a Bayesian uncertainty algorithm, enabling the analysis of uncertainty in semantic segmentation results.
In addition, the cross-attention mechanism has been widely used in RGB-T semantic segmentation. Zhang et al. proposed ABMDRNet [17]. The model initially employs a method based on image-to-image bidirectional transformation to bridge the modality gap in multimodal data. It adaptively selects distinct multimodal features for RGB-T semantic segmentation. Zhou et al. proposed GMNet [18]. The authors, taking into account the characteristics of low-level and high-level features, designed two distinct fusion modules to enhance pertinent details. Deng et al. proposed FEANet [19]. The model explores and enhances multi-level features from both channel and spatial perspectives, directing increased attention towards high-resolution features obtained from the fusion images.
3. Approach
In this section, we present the proposed DASNet in detail. Firstly, we introduce the network overview of DASNet. Secondly, we elaborate the proposed CSM, PAM, EEM, and DEF, respectively. Thirdly, we present the loss function.
3.1. Overview of the Proposed Method
As depicted in Figure 1, the proposed DASNet is based on the feature fusion paradigm, including a feature extractor, three specific modules, and a decoder. For the feature extractor, we adopt two parallel ResNet-152 backbones, named RGB branch and TIR branch, to extract multi-modality features from RGB and TIR images. We denote the five convolution blocks in the RGB branch and TIR branch as and respectively. The corresponding output RGB and TIR features are denoted as . The input size is denoted as , so is , is , and . Here, the RGB and TIR branches share parameters, which not only keeps the extracted features in the same feature space but also reduces the number of parameters. Furthermore, we employ two identical convolutional layers (i.e., shared parameters) to project the same level of RGB and TIR features, to , respectively, with fewer channels (i.e., ).
Figure 1.
DASNet’s core design principle is that features at different depths carry qualitatively different information. Multiple sets of RGB and TIR features from different stages are passed to the corresponding modules (EEM, PAM, and CSM). Then, the output features of each module are passed to the decoder. The corresponding transmission relationship in the figure uses the same color, shape, and dotted line type. Before passing the output features of each module to the decoder, we additionally set corresponding heads for supervision. Heads are convolutional layers. Among them, the number of channels for the features of the CSM head outputs is the total number of the dataset categories, while for the PAM and EEM, it is 2.
We divide the extracted basic cross-modal features into three levels, that is, as the high level, as the middle level, and as the low level. We arrange the CSM, PAM, and EEM for these three levels of features to achieve semantics extraction, patch activation, and edge enhancement respectively.
Concretely, the CSM enables the network to enhance the semantic information of all objects under semantic supervision. The PAM enables the network to highlight the position regions of all objects of different categories. The EEM enables the network to extract edge information of objects with the help of edge supervision. These three specific modules are the core components of our DASNet. With the informative features generated by the above modules from the five-level features, we use a deep-emphasis fusion module (DEF) in the decoder for the fusion of different informative features to achieve the accurate segmentation result ; the decoder is composed of three DEF modules, an element-wise addition operation, a convolutional layer, and an upsampling operation.
3.2. Cross-Attention Semantics Module
As we all know, high-level features contain rich semantic information. Cross-modality RGB and TIR features are complementary, which is more conducive to semantics extraction. Inspired by object segmentation works exploring the cross-attention modeling of target objects in two different modalities, consecutive video frames and features of successive levels, we propose a cross-attention semantics module to model the pixel-level correlation of multi-modality high-level semantic features to highlight the targeted semantics. We illustrate the CSM in Figure 2. The input features of the CSM are and . The whole process of CSM can be divided into correlation fore fusion, cross-attention modeling, and multi-modality fusion.
Figure 2.
The PAM and CSM are complementary rather than interchangeable: the PAM adapts its attention type to the depth of the features it receives (spatial attention for shallow stages, channel attention for deeper stages), while the CSM must resolve cross-modal misalignment before feature interaction between modalities. The (right) part is the PAM, while the other is the CSM. In the CSM, the weight matrix obtained through the correlation fusion is spatial-wise multiplied by the features before the multi-modality fusion, in order to highlight the interested position. The convolution layers in the figure are omitted.
3.2.1. Correlation Fore Fusion
Due to significant disparities between two modalities, direct fusion yields suboptimal results, which is confirmed by the corresponding ablation study listed in Table 1. Therefore, we designed the correlation fore fusion (where ‘Fore’ indicates ‘Before’) to narrow the inter-modality gap in the feature space. We reshape into . Then, we use matrix multiplication to get the correlation matrix between two different modalities, which can be formulated as follows:
where × is the matrix multiplication, is the matrix transpose operation and . We reshape into . Then, we use convolution to adjust the channel of and get the as the spatial attention of the multi-modality.
Table 1.
Ablation experiment about the module details. The top result is highlighted in red color.
3.2.2. Cross-Attention Modeling
Owing to environmental factors such as illumination, certain categories prove challenging to discern within RGB features. Element-wise multiplication serves to accentuate regions of interest from the TIR features within the RGB feature space. Therefore, we first perform an element-wise multiplication operation on and , generating . Then, we model the pixel-level correlation of and , respectively, to collaboratively identify objects in cross-modal features, which can be formulated as follows:
where is the co-attention operation, and are the multi-modality correlation features. The co-attention operation computes the correlation matrix of two input features through matrix multiplication. Then, it transfers the valuable semantics of the correlation matrix to the output feature.
3.2.3. Multi-Modality Fusion
is the spatial attention matrix of multi-modality features. and have a strong representation of the targeted semantics. We design the multi-modality fusion to fuse them effectively. We adopt a fusion scheme that mixes multiple multiplications and additions to combine multi-modality features, which can be formulated as follows:
where ⊗ is the spatial-wise multiplication, ⊕/∗ are the element-wise summation/multiplication, © is the concatenation, and is the output feature of the CSM. We omit the convolutional layer.
Through the above multi-modality fusion operation, we can obtain informative features of targeted semantics. Furthermore, as shown in Figure 3, we attach a semantic head (SemHead) after the CSM to achieve more accurate targeted semantics with the semantic supervision.
Figure 3.
EEM exploits the asymmetry between modalities rather than simply combining them. When the features of the two modalities are element-wise subtracted, the subtraction relationship is exchanged to obtain two output features. Specifically, of the two outputs from the element-wise subtraction, one is to use RGB features to subtract TIR features, while the other uses TIR to subtract RGB features. The convolution layers in the figure are ignored.
3.3. Patch Activation Module
Middle-level multi-modality features (i.e., ) facilitate the localization of all targets within the image, offering sufficient patch information. To achieve this, we devised the PAM to activate patch regions of objects within middle-level features at different scales. Concretely, our PAM is based on the attention mechanism, which enables us to activate specific regions and establish stronger connections between features.
We illustrate the PAM in Figure 2, whose inputs are and . We first use element-wise summation and multiplication to get to take advantage of the two types of complementary feature combinations.
Then, distinct attention mechanisms are employed for features at various stages. For the 2nd stage, given its proximity to the pixel level, we employ bidirectional spatial attention. To enhance the sharpness of edges within the patch, we further integrate the obtained fused features with edge features from the 1st stage.
To sum up, we briefly formulate the above process as follows:
where ⊗ is the spatial-wise multiplication, ⊕ is the element-wise summation, is the spatial attention operation, is after the convolution and downsampling layers, and is the output feature of the PAM from the 2nd stage.
For the 3rd stage, in order to extract deeper discriminative features, we employ bidirectional channel attention. We incorporate with output from the 3rd stage for the multi-scale fusion. To sum up, we briefly formulate the above process as follows:
where ⊙ is the channel-wise multiplication, is the channel attention operation, is after the convolution and downsampling layers, and is the output feature of the PAM from the 3rd stage.
3.4. Edge Enhancement Module
Low-level multi-modality features, being the closest to the pixel level, encapsulate rich spatial edge features. Incorporating edge features into the model offers two advantages. First, it introduces a third discriminative feature into the entire model, enhancing its hierarchical structure and robustness. Second, it provides additional edge features to the PAM, improving the results of the patch analysis.
We illustrate the EEM in Figure 3, whose inputs are . We employ element-wise subtraction to obtain modality-specific regions of interest. For instance, subtracting from , and subsequently applying the ReLU activation function, yields , representing regions of interest for RGB but disinterest for TIR. Similarly, we obtain . To enable a single modality to capture regions of interest from the other modality, we perform element-wise multiplication between and , as well as between and , resulting in . Similar to the PAM, we use element-wise summation and multiplication between and to get . Then, we add and , transferring to a parallel structure with dilated convolutions to extract multi-scale detail information. Finally, we aggregate multi-scale features by concatenation, generating the output feature of the EEM, . We briefly summarize the above process as follows:
where represents the multi-head dilated convolutions.
3.5. Deep-Emphasis Fusion Module
To efficiently fuse the output features from different modules, as illustrated in Figure 4, we designed a deep-emphasis fusion module. We first adopted a hybrid scheme including summation, multiplication, and concatenation to combine multi-scale features. Then, a residual structure was designed to amplify the weight of deep features. Finally, upsampling was employed to adjust the size of the output features. We briefly summarize the above process as follows:
where is the output feature of the module at the ith stage, is the output feature of DEF from the th stage, . We omit the convolutional and upsampling layers.
Figure 4.
DEF’s key design choice is to enhance rather than merely concatenate deep features. is the output feature of the module at the ith stage. is the output feature of DEF from the th stage. These two features are fused through DEF to obtain , which is the output feature of DEF from the ith stage.
4. Experiments and Results
All experiments used an input resolution of 640 × 480 for both training and testing. A single random seed of 42 was fixed for all runs, and all reported results were obtained from one training run per configuration. Implementation was in Python 3.10.16 with PyTorch 2.1.0 and torchvision 0.16.0, on an NVIDIA GeForce RTX 3090 card. The MFNet dataset is a publicly available RGB-T semantic segmentation benchmark for autonomous driving. It comprises 1569 spatially aligned RGB–thermal image pairs collected in urban environments under both daytime and nighttime conditions, officially partitioned into 820 daytime and 749 nighttime pairs. Eight obstacle classes commonly encountered in driving scenarios, including car, person, bike, curve, car stop, guardrail, color cone, and bump, are annotated. The dataset is divided into training, validation, and test sets, with the training set containing 50% of the daytime and nighttime images, and the validation and test sets each containing 25% of the daytime and nighttime images.
Following the empirical studies in [20,21], all quantitative results reported in this paper were obtained from a single training run with a fixed random seed. We do not report means, standard deviations, or significance tests, and therefore we do not claim that the observed performance differences are statistically significant. Known sources of run-to-run variability in our setting include random weight initialization of the non-pretrained layers, batch ordering, and the stochastic augmentation pipeline. Our use of the official MFNet split removes data-partitioning variability, and ImageNet-pretrained backbone weights reduce sensitivity to initialization, but these measures do not substitute for repeated runs. Related experimental results on MFNet benchmark are shown in Table 2. We therefore present our results as a single-run comparison under a fixed, fully reported protocol, and we explicitly refrain from any claim of statistical superiority over the compared methods.
During the training phase, we utilized RGB and TIR images at their original resolutions as inputs to the network, employing random flipping and cropping for data augmentation. The batch size during model training was set to four, with an initial learning rate of . The parameters of the feature extractor were initialized with a pretrained ResNet-152 model. The parameters of other convolutional layers were initialized with the KAIM method [22]. We set the number of training epochs to 500. During the testing phase, we directly input the RGB-T image pairs at their original resolutions into the trained DASNet without any post-processing, obtaining segmentation results.
Table 2.
Quantitative comparison on the test set of the MFNet dataset. ‘-’ means that the authors do not provide the corresponding information. The top two results in each column are highlighted in red and blue.
4.1. Comparisons with SOTA Methods
We compared DASNet with advanced CNN-based models on RGB/RGB-D/RGB-T tasks. For a fair comparison, building upon the foundations laid by [17,33], we further adapted certain RGB semantic segmentation methods to accommodate RGB-T image pairs. We employed default parameter settings for the RGB-T, RGB-D, and modified RGB models, subsequently retraining them on the same dataset. On the MFNet dataset, we compared our approach with state-of-the-art CNN-based models, including RGB-T (i.e., SpiderMesh-152 [48] and CACFNet [47]), RGB-D (i.e., ACNet [31]) and RGB (i.e., HRNet [26]). The quantitative experimental results of our method and all other approaches on the MFNet dataset are presented in Table 2, encompassing performance metrics for eight classes and overall performance. Overall, our approach achieves the best performance on the crucial mIoU metric, demonstrating robust adaptability to diverse scenes. Specifically, on the mIoU metric, our approach outperforms the second-best method (SpiderMesh-152 [48]) by 0.2%. Across all categories, our approach demonstrates excellent performance in the Car Stop, Car, and Color Cone classes. Particularly, the Car Stop category secures the first position, surpassing the second-best method (GCGLNet [40]) by 1.1%. The Car and Color Cone categories achieve the second position. We observe that DASNet achieves a notably low IoU on the Guardrail class (5.2%). We attribute this to the extreme class imbalance (Guardrail is one of the rarest classes in MFNet) and its thin, elongated structure, which is difficult for standard convolutions and pooling operations to capture. While our EEM enhances edge features, it is not specifically designed for extremely thin objects. This highlights a limitation of our current design and suggests that specialized strategies—such as class-reweighted loss functions, dedicated thin-structure modules, or synthetic data augmentation—are needed to improve performance on such underrepresented classes.
Moreover, the quantitative experimental results of our model and other methods on the daytime and nighttime test sets of the MFNet dataset are presented in Table 3. Our approach demonstrated superiority in both of these scenarios, particularly on the mIoU metric. This suggests that the strategy employed by our segmentation model is effective. The edge–patch–semantic feature extraction pathway progressively enables our method to extract discriminative features, achieving RGB-T segmentation. The above quantitative analysis clearly indicates the effectiveness of our DASNet on the MFNet dataset.
Table 3.
Comparative results (%) in daytime and at nighttime. The top two results in each column are highlighted in red and blue.
We acknowledge that the total FLOPs of DASNet are higher than lightweight backbones due to the employment of the dual ResNet-152 feature extractor. However, benefiting from the shared-weight strategy, the parameter count is only 89.7M, which is even lower than CACFNet (198.6M). As shown in Table 4, we compare DASNet with GCGLNet and CACFNet from the perspectives of params, FLOPs, FPS, and mIoU. While DASNet’s FLOPs are substantially higher than GCGLNet, it attains the highest mIoU of the three methods with less than half the parameters of CACFNet. DASNet achieves an inference speed of 24 FPS. This indicates that our method maintains feasible efficiency for real-world autonomous driving deployment.
Table 4.
Complexity comparisons of various models. FPS indicates frames per second.
In addition, we conducted a partial visualization experiment. Figure 5 presents the segmentation results of our model. The first row corresponds to the original RGB images. The second row depicts the original TIR images. The third row illustrates the segmentation results obtained by our model. The fourth row represents the ground truth. The first to fourth columns correspond to daytime scenes, while the fifth to eighth columns represent nighttime scenes. These segmented images visually showcase the effectiveness of our model on the MFNet dataset, demonstrating robust segmentation performance in both daytime and nighttime scenarios.
Figure 5.
Results of qualitative experiments.
Additionally, we visualize the outputs of the EEM and PAM in Figure 6 and Figure 7. Figure 6 corresponds to daytime scenes, while Figure 7 represents nighttime scenes. The first and second rows depict the original RGB and TIR images, respectively. The third row visualizes the output features from the EEM. The fourth row visualizes the output features from the PAM. From the output feature maps of the EEM, we observe that our model accurately extracts and sharpens the edge features of the objects in both daytime and nighttime scenes. In the feature maps from the PAM, we find that the regions of interest in both scenarios are highlighted, with particularly accurate delineation in nighttime scenes. This validates the effectiveness of the EEM and PAM.
Figure 6.
Experiment results on daytime scenes with corresponding feature maps. We used the output mask of the EEM head as the feature map of the EEM for visualization. The EEM highlights the edge areas of the target effectively. Considering that visualization pays more attention to spatial features, we chose the output mask of the head of the 2nd-stage PAM as the feature map of the PAM for visualization. Also, for easy observation, we set the visualization threshold to 0.5. The PAM effectively pays more attention to the area around the target position.
Figure 7.
Experiment results on nighttime scenes with corresponding feature maps. The settings are the same as for the daytime scenes. Compared with the daytime scenes, the performance of the PAM is better at nighttime. The position of interest is more precise.
4.2. Ablation Experiments
We conducted a comprehensive ablation study to demonstrate the effectiveness of individual modules and key structures within the modules on the MFNet dataset. Specifically, we firstly employed ablation experiments to showcase the separate and combined contributions, as well as the effectiveness, of the four modules (EEM, PAM, CSM, and DEF). Subsequently, we established the effectiveness of individual components within each of the four modules. For all ablation experiments, we trained the model with the same training parameters as described in Section 4. We employed mIoU as the performance evaluation metric.
4.2.1. The Individual and Combined Effectiveness of the Modules
We introduced three modules, the CSM, PAM, and EEM, to extract semantic, patch, and edge features, respectively. Initially, we proposed four different variants to assess the segmentation performance when each of the three modules was used separately. These variants were: (1) baseline, (2) baseline + EEM, (3) baseline + PAM, and (4) baseline + CSM.
We employed element-wise addition and convolution to replace the modules that were not equipped in each variant. The experimental results are presented in Table 5. The “Baseline” achieved an mIoU of only 55.4%, which is 2.7% lower than our complete DASNet, indicating that these three modules indeed contributed to improving segmentation accuracy. With the assistance of the EEM, PAM or CSM, the segmentation performance of Variants 2, 3, and 4 improved compared to that of the “Baseline,” with increases of 0.5%, 1.0%, and 0.8%, respectively.
Table 5.
Ablation experiments. The top result is highlighted in red color. ✓ denotes the adoption of the corresponding module.
Next, we introduced three variants to assess the combined contributions of the three modules: (5) Baseline + PAM + CSM, (6) Baseline + EEM + CSM, and (7) Baseline + EEM + PAM. The segmentation metrics for the three variants demonstrated further improvements when two modules were combined compared to using a single module. Subsequently, we simultaneously incorporated all three modules in the model. The segmentation metrics for that variant showed improved performance compared to Variants 5, 6, and 7, with increases of 1.6%, 1.1%, and 0.5%, respectively. These experimental results confirm the collective contributions of the three modules when used in conjunction.
We then introduced the DEF module for multi-scale fusion, integrating features of two different types (patch and semantic). To demonstrate the effectiveness of the fusion module, we designed Variant 8. That variant was based on the simultaneous use of the three modules, replacing the DEF module in the decoder with concatenation and convolution. The segmentation metrics for Variant 8 decreased by 1.3%. This result demonstrates the effectiveness of the DEF module. The above experimental results collectively confirm the effectiveness of all the modules designed in the DASNet.
4.2.2. Effectiveness of Each Component of EEM
To validate the effectiveness of individual components within the EEM, we introduced two variants: (1) w/o dilated convolution and (2) using element-wise addition instead of the multi-modality fusion. The experimental results are illustrated in Table 1. Firstly, we observe a decrease of 1.7% in segmentation metrics when we use regular convolution instead of dilated convolution. This is primarily attributed to the smaller receptive field of regular convolution compared to dilated convolution. The smaller receptive field causes the EEM to overlook some boundaries of target objects.
Secondly, compared to the multi-modality fusion, the segmentation metrics of the model with element-wise addition in the EEM decrease by 0.4%. This is mainly because the multi-modality fusion incorporates element-wise addition, element-wise multiplication, and matrix concatenation. These three methods have corresponding advantages in feature fusion, which are equally important.
4.2.3. Effectiveness of Each Component of PAM
To validate the effectiveness of individual components within the PAM, we used three variants: (1) w/o channel attention, using spatial attention instead, (2) w/o spatial attention, using channel attention instead, and (3) swapping the application positions of the two attention mechanisms. The experimental results are illustrated in Table 1. Firstly, we observe a 1.5% decrease in mIoU when only using the spatial attention mechanism compared to the segmentation metrics of the original network. When only using the channel attention mechanism, the mIoU decreases by 0.5%. Due to the difference in the extent of the decrease, we conclude that the spatial attention is more crucial than the channel attention in the PAM. This is because the features processed by the PAM are closer to the pixel level, which is more suitable for spatial attention.
Secondly, when we swap the positions of the two attention mechanisms in the network, the segmentation metrics decrease by 0.3%. We attribute this to the characteristics of the two attention mechanisms. Spatial attention is suitable for operating on feature maps close to the pixel level, effectively highlighting regions of interest. Channel attention is more suitable for operating on feature maps farther from the pixel level, extracting deeper information. The usage positions of the two attention mechanisms differ.
4.2.4. Effectiveness of Each Component of CSM
To validate the effectiveness of individual components within the CSM, we used three variants: (1) w/o co-attn, (2) w/o correlation fusion, and (3) using element-wise addition instead of the multi-modality fusion. The experimental results are illustrated in Table 1. Firstly, using element-wise addition results in a decrease of 1.3% in segmentation metrics compared to co-attn. We believe that the co-attn mechanism calculates the similarity between different positions of two features. The co-attn mechanism, compared to element-wise addition, allows for a tighter integration of the two features, optimizing segmentation performance.
Secondly, we observe a decrease of 0.9% in segmentation metrics when the correlation fusion module is excluded. This is because the correlation fusion module enables the module to explore deep positional features, passing these features to the subsequent fused features. The absence of these positional features results in a CSM lacking localization capabilities, leading to inaccurate semantic positioning. Then, as mentioned in the ablation experiments for the EEM, the multi-modality fusion incorporates element-wise addition, element-wise multiplication, and matrix concatenation. The effectiveness of the multi-modality fusion surpasses the use of any single method.
4.2.5. Effectiveness of Each Component of DEF
To validate the effectiveness of the residual structure within the DEF module, we introduced a single variant: w/o res-structure. We found that when removing the residual structure in the DEF module, the segmentation metrics of the network decreased by 0.7%. This is because the residual structure is designed to strengthen the weight of deep features in the final output fusion features. The absence of the residual structure results in the dominance of shallow edge and patch features, leading to a decrease in semantic accuracy.
5. Conclusions
In this paper, we proposed a novel Depth Adaption SegNet for RGB-T Segmentation. In order to fully utilize the characteristics of features from different stages, our network employed three different modules to extract corresponding discriminative information from features of varying levels of depth. Our network extracted semantics, patch, edge and discriminative information from features on deep-to-shallow stages. The network also utilized an effective approach to integrate these features. Specifically, a cross-attention semantics module elicited semantic discriminative features for target classification from deep features. We employed attention mechanisms to construct a patch activation module extracting patch-based discriminative characteristics from middle-level features. In addition, we designed an edge enhancement module leveraging dilated convolutions to extract boundary information from shallow features. Moreover, to effectively integrate output features of various categories, we proposed a deep-emphasis fusion module in the decoder to fuse them. Notably, DASNet maintains a balance between spatial localization and depth-wise semantics, which has not been effectively addressed by other existing RGB-T segmentation methods. Qualitative and quantitative experiments showed that our Depth Adaption SegNet exhibited impressive segmentation performance.
We acknowledge several limitations in the current study. First, our quantitative evaluation was limited to the MFNet dataset. While MFNet provides rich day/night variability and is the most widely used benchmark for RGB-T urban scene parsing, we recognize that evaluating on an additional dataset such as PST900 would further strengthen our generalization claims. Second, while our computational analysis showed that DASNet achieved practical inference speed, further optimization for embedded deployment remains an important direction. Third, while our method achieved impressive mIoU, its performance on certain challenging categories, notably Guardrail, remained suboptimal, which we attribute to limited training samples and the structural complexity of such objects. Finally, our results were based on a single training run and were not accompanied by standard deviations or significance testing. We therefore do not claim statistical significance for the reported margins over competing methods. Conducting multiple runs with different random seeds and reporting mean performance with standard deviation are important directions for future work.
Author Contributions
Conceptualization, S.Z., H.Z. and Y.Z.; methodology, S.Z. and C.Z.; software, S.Z. and C.Z.; validation, S.Z., H.Z. and B.L.; formal analysis, C.Z., S.Z. and B.L.; investigation, S.Z., B.L. and Y.Z.; resources, H.Z., B.L. and Y.Z.; data curation, S.Z. and C.Z.; writing—original draft preparation, S.Z., H.Z. and C.Z.; writing—review and editing, B.L. and Y.Z.; visualization, C.Z., S.Z. and B.L.; supervision, B.L. and Y.Z.; project administration, B.L. and S.Z.; funding acquisition, S.Z. All authors have read and agreed to the published version of the manuscript.
Funding
This research was funded by the National Natural Science Foundation of China (62602669), the Fundamental Research Funds for the Central Universities (2025QN1162).
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The raw data supporting the conclusions of this article will be made available by the authors on request.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Hong, Y.; Tarimo, S.A.; Woo, J. Pancreas Segmentation Using a Two-Stage Pipeline of Faster R-CNN and TransUNet. Appl. Sci. 2026, 16, 5764. [Google Scholar] [CrossRef] [Scilit]
- Majanga, V.; Mnkandla, E.; Luo, Y.; Oladele, D. QEEF: A Quantitative Explainability Evaluation Framework for CNN and Vision Transformer-Based Segmentation Models in Dental Images. Appl. Sci. 2026, 16, 7133. [Google Scholar] [CrossRef] [Scilit]
- Yue, T.; Huang, H.; Wang, Q.; Song, B.; Chen, Y. A Multimodal Deep Learning Framework for Accurate Wildfire Segmentation Using RGB and Thermal Imagery. Appl. Sci. 2025, 15, 10268. [Google Scholar] [CrossRef] [Scilit]
- Li, G.; Wang, Y.; Liu, Z.; Zhang, X.; Zeng, D. RGB-T semantic segmentation with location, activation, and sharpening. IEEE Trans. Circuits Syst. Video Technol. 2022, 33, 1223–1235. [Google Scholar] [CrossRef] [Scilit]
- Long, J.; Shelhamer, E.; Darrell, T. Fully convolutional networks for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 3431–3440. [Google Scholar]
- Badrinarayanan, V.; Kendall, A.; Cipolla, R. Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 2481–2495. [Google Scholar] [CrossRef] [Scilit]
- Ronneberger, O.; Fischer, P.; Brox, T. U-net: Convolutional networks for biomedical image segmentation. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, 5–9 October 2015; Proceedings, Part III 18; Springer: Berlin/Heidelberg, Germany, 2015; pp. 234–241. [Google Scholar]
- Chen, L.C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. Semantic image segmentation with deep convolutional nets and fully connected crfs. arXiv 2014, arXiv:1412.7062. [Google Scholar]
- Chen, L.C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 834–848. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.C.; Papandreou, G.; Schroff, F.; Adam, H. Rethinking atrous convolution for semantic image segmentation. arXiv 2017, arXiv:1706.05587. [Google Scholar]
- Zhao, H.; Shi, J.; Qi, X.; Wang, X.; Jia, J. Pyramid scene parsing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2881–2890. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Lin, G.; Milan, A.; Shen, C.; Reid, I. Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1925–1934. [Google Scholar]
- Ha, Q.; Watanabe, K.; Karasawa, T.; Ushiku, Y.; Harada, T. MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes. In Proceedings of the 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2017; pp. 5108–5115. [Google Scholar]
- Sun, Y.; Zuo, W.; Liu, M. RTFNet: RGB-thermal fusion network for semantic segmentation of urban scenes. IEEE Robot. Autom. Lett. 2019, 4, 2576–2583. [Google Scholar] [CrossRef] [Scilit]
- Sun, Y.; Zuo, W.; Yun, P.; Wang, H.; Liu, M. FuseSeg: Semantic segmentation of urban scenes based on RGB and thermal data fusion. IEEE Trans. Autom. Sci. Eng. 2020, 18, 1000–1011. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Zhao, S.; Luo, Y.; Zhang, D.; Huang, N.; Han, J. ABMDRNet: Adaptive-weighted bi-directional modality difference reduction network for RGB-T semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 2633–2642. [Google Scholar]
- Zhou, W.; Liu, J.; Lei, J.; Yu, L.; Hwang, J.N. GMNet: Graded-feature multilabel-learning network for RGB-thermal urban scene semantic segmentation. IEEE Trans. Image Process. 2021, 30, 7790–7802. [Google Scholar] [CrossRef] [Scilit]
- Deng, F.; Feng, H.; Liang, M.; Wang, H.; Yang, Y.; Gao, Y.; Chen, J.; Hu, J.; Guo, X.; Lam, T.L. FEANet: Feature-enhanced attention network for RGB-thermal real-time semantic segmentation. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2021; pp. 4467–4473. [Google Scholar]
- Wang, S. Class-aware temporal and contextual contrastive framework for semi-supervised automated fault detection and diagnosis in air handling units. Energy Build. 2026, 358, 117233. [Google Scholar] [CrossRef] [Scilit]
- Wang, S. Automated fault detection and diagnosis of ahus via tabular-based methods using operational data from a large office building. J. Comput. Civ. Eng. 2026, 40, 04026018. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proceedings of the IEEE International Conference on Computer Vision, Boston, MA, USA, 7–12 June 2015; pp. 1026–1034. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Nashville, TN, USA, 20–25 June 2021; pp. 10012–10022. [Google Scholar]
- Yu, C.; Wang, J.; Peng, C.; Gao, C.; Yu, G.; Sang, N. Bisenet: Bilateral segmentation network for real-time semantic segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 325–341. [Google Scholar]
- Fu, J.; Liu, J.; Tian, H.; Li, Y.; Bao, Y.; Fang, Z.; Lu, H. Dual attention network for scene segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 3146–3154. [Google Scholar]
- Sun, K.; Xiao, B.; Liu, D.; Wang, J. Deep high-resolution representation learning for human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 5693–5703. [Google Scholar]
- Chen, X.; Lin, K.Y.; Wang, J.; Wu, W.; Qian, C.; Li, H.; Zeng, G. Bi-directional cross-modality feature propagation with separation-and-aggregation gate for RGB-D semantic segmentation. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 561–577. [Google Scholar]
- Wang, W.; Neumann, U. Depth-aware cnn for rgb-d segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 135–150. [Google Scholar]
- Hazirbas, C.; Ma, L.; Domokos, C.; Cremers, D. Fusenet: Incorporating depth into semantic segmentation via fusion-based cnn architecture. In Proceedings of the Computer Vision–ACCV 2016: 13th Asian Conference on Computer Vision, Taipei, Taiwan, 20–24 November 2016; Revised Selected Papers, Part I 13; Springer: Berlin/Heidelberg, Germany, 2017; pp. 213–228. [Google Scholar]
- Karen, S.; Andrew, Z. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv 2014, arXiv:1409.1556. [Google Scholar]
- Hu, X.; Yang, K.; Fei, L.; Wang, K. Acnet: Attention based network to exploit complementary features for rgbd semantic segmentation. In Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP); IEEE: Piscataway, NJ, USA, 2019; pp. 1440–1444. [Google Scholar]
- Huang, G.; Liu, Z.; Van Der Maaten, L.; Weinberger, K.Q. Densely connected convolutional networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 4700–4708. [Google Scholar]
- Zhou, W.; Dong, S.; Xu, C.; Qian, Y. Edge-aware guidance fusion network for RGB–thermal scene parsing. Proc. AAAI Conf. Artif. Intell. 2022, 36, 3571–3579. [Google Scholar] [CrossRef] [Scilit]
- Zhou, W.; Dong, S.; Lei, J.; Yu, L. MTANet: Multitask-aware network with hierarchical multimodal fusion for RGB-T urban scene understanding. IEEE Trans. Intell. Veh. 2022, 8, 48–58. [Google Scholar] [CrossRef] [Scilit]
- Zhou, W.; Lin, X.; Lei, J.; Yu, L.; Hwang, J.N. MFFENet: Multiscale feature fusion and enhancement network for RGB–Thermal urban road scene parsing. IEEE Trans. Multimed. 2021, 24, 2526–2538. [Google Scholar] [CrossRef] [Scilit]
- Fu, Y.; Chen, Q.; Zhao, H. CGFNet: Cross-guided fusion network for RGB-thermal semantic segmentation. Vis. Comput. 2022, 38, 3243–3252. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Zhou, W.; Cui, Y.; Yu, L.; Luo, T. GCNet: Grid-like context-aware network for RGB-thermal semantic segmentation. Neurocomputing 2022, 506, 60–67. [Google Scholar] [CrossRef] [Scilit]
- Dong, S.; Zhou, W.; Qian, X.; Yu, L. GEBNet: Graph-enhancement branch network for RGB-T scene parsing. IEEE Signal Process. Lett. 2022, 29, 2273–2277. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Mao, H.; Wu, C.Y.; Feichtenhofer, C.; Darrell, T.; Xie, S. A convnet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 11976–11986. [Google Scholar]
- Gong, T.; Zhou, W.; Qian, X.; Lei, J.; Yu, L. Global contextually guided lightweight network for RGB-thermal urban scene understanding. Eng. Appl. Artif. Intell. 2023, 117, 105510. [Google Scholar] [CrossRef] [Scilit]
- Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adam, H. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv 2017, arXiv:1704.04861. [Google Scholar]
- Cai, Y.; Zhou, W.; Zhang, L.; Yu, L.; Luo, T. DHFNet: Dual-decoding hierarchical fusion network for RGB-thermal semantic segmentation. Vis. Comput. 2023, 40, 169–179. [Google Scholar] [CrossRef] [Scilit]
- Zhao, S.; Liu, Y.; Jiao, Q.; Zhang, Q.; Han, J. Mitigating modality discrepancies for RGB-T semantic segmentation. IEEE Trans. Neural Netw. Learn. Syst. 2023, 35, 9380–9394. [Google Scholar] [CrossRef] [Scilit]
- Zhao, S.; Zhang, Q. A Feature Divide-and-Conquer Network for RGB-T Semantic Segmentation. IEEE Trans. Circuits Syst. Video Technol. 2022, 33, 2892–2905. [Google Scholar] [CrossRef] [Scilit]
- Frigo, O.; Martin-Gaffe, L.; Wacongne, C. DooDLeNet: Double DeepLab enhanced feature fusion for thermal-color semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–20 June 2022; pp. 3021–3029. [Google Scholar]
- Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 801–818. [Google Scholar]
- Zhou, W.; Dong, S.; Fang, M.; Yu, L. CACFNet: Cross-modal attention cascaded fusion network for RGB-T urban scene parsing. Proc. IEEE Trans. Intell. Veh. 2023, 9, 1919–1929. [Google Scholar] [CrossRef] [Scilit]
- Fan, S.; Wang, Z.; Wang, Y.; Liu, J. SpiderMesh: Spatial-aware Demand-guided Recursive Meshing for RGB-T Semantic Segmentation. arXiv 2023, arXiv:2303.08692. [Google Scholar]
- Pohlen, T.; Hermans, A.; Mathias, M.; Leibe, B. Full-resolution residual networks for semantic segmentation in street scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 4151–4160. [Google Scholar]
- Yu, C.; Wang, J.; Peng, C.; Gao, C.; Yu, G.; Sang, N. Learning a discriminative feature network for semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 1857–1866. [Google Scholar]
- Sakaridis, C.; Dai, D.; Gool, L.V. Guided curriculum model adaptation and uncertainty-aware evaluation for semantic nighttime image segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Korea, 27 October–2 November 2019; pp. 7374–7383. [Google Scholar]
- Huang, Z.; Wang, X.; Huang, L.; Huang, C.; Wei, Y.; Liu, W. Ccnet: Criss-cross attention for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Korea, 27 October–2 November 2019; pp. 603–612. [Google Scholar]
- He, J.; Deng, Z.; Zhou, L.; Wang, Y.; Qiao, Y. Adaptive pyramid context network for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 7519–7528. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.






