Next Article in Journal
Maize Detection and Row Extraction Using Maize–YOLO and IPM–Clustering Method for Autonomous Agricultural Navigation
Previous Article in Journal
FedMIR: Multimodal Federated Learning with Missing Modality Imputation and Distribution-Aware Routing
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Scene-Aware Degradation Universal Re-Identification Framework for Adverse Weather

1
School of Computer and Artificial Intelligence, Wuhan University of Technology, Wuhan 430070, China
2
School of Computer Science, Hubei University of Technology, Wuhan 430068, China
*
Author to whom correspondence should be addressed.
Sensors 2026, 26(10), 2951; https://doi.org/10.3390/s26102951
Submission received: 2 February 2026 / Revised: 11 April 2026 / Accepted: 28 April 2026 / Published: 8 May 2026
(This article belongs to the Section Sensing and Imaging)

Highlights

What are the main findings?
  • The proposed ScA-UniReID framework, built upon CLIP, effectively addresses the challenge of ReID under coupled adverse weather (e.g., rain and fog) by dynamically disentangling identity semantics from degradation artifacts using dual textual prompts and an adaptive control module.
What are the implications of the main findings?
  • This work provides a novel cross-modal paradigm that moves beyond conventional image-enhancement or robust-feature approaches, offering a new solution for ReID in complex, real-world environments where multiple degradations co-occur.

Abstract

Vision-based Re-identification (ReID) is crucial for intelligent surveillance yet remains vulnerable to adverse-weather degradations such as rain and fog, which simultaneously corrupt visual clarity and identity-specific cues. Existing image-enhancement and robust-feature paradigms struggle when multiple degradations co-occur, while recent CLIP-based ReID models have scarcely examined cross-modal alignment under weather distortions. To bridge this gap, we propose ScA-UniReID, a Scene-Aware Degradation Universal ReID framework built upon CLIP’s dual-encoder architecture. ScA-UniReID introduces dual textual prompts—target-oriented for identity features and degradation-oriented for weather noise—and an adaptive control module that dynamically re-weights them to disentangle identity semantics from degradation artifacts. Extensive experiments on pedestrian and maritime ReID benchmarks under diverse adverse-weather protocols show that ScA-UniReID outperforms state-of-the-art methods and generalizes robustly to unseen conditions, validating its efficacy and universality.

1. Introduction

Target re-identification (ReID) aims to retrieve and recognize specific individuals across large-scale image databases captured by different cameras and over extended periods [1]. As a cornerstone of intelligent perception, it underpins a wide range of applications—from smart transportation, public security [2], and maritime surveillance [3] to smart retail and urban management. Yet, the performance of any vision-based recognition system is fundamentally limited by image clarity and content integrity [4]. In real-world complex environments, images are frequently degraded by noise, motion blur, and adverse weather conditions such as rain, fog, or low-light scenarios, which significantly hinder the practical effectiveness of ReID systems.
To counteract the adverse effects of degraded images, existing research has generally pursued two complementary technical avenues [5,6]. The first operates at the input level: image-enhancement pipelines that denoise, dehaze, or super-resolve imagery before it reaches the recognition network [7,8]. The second avenue focuses on the model level: designing more robust feature extractors [9,10] that learn identity representations resilient to various image corruptions. Despite notable successes, empirical studies reveal that both strategies suffer limited generalization when multiple degradation factors co-occur—such as heavy rain at night—because they struggle to simultaneously restore visual clarity and preserve identity-specific cues.
In recent years, the rise of Vision-Language Pre-training (VLP) has provided new solutions for re-identification research. Cross-modal models such as CLIP (Contrastive Language-Image Pre-training) [11] employ large-scale image-text contrastive learning to jointly optimize visual and text encoders, mapping multimodal data into a unified semantic embedding space. This enables the alignment of semantically relevant samples and the separation of irrelevant ones, thereby modeling high-level semantic correspondences across modalities, and achieves strong semantic alignment capability and excellent zero-shot transfer performance [12]. Within the ReID community, researchers have begun to exploit learnable textual prompts to steer visual encoders toward identity-relevant regions, thereby enhancing fine-grained discriminative power [13,14]. However, most CLIP-based ReID studies have concentrated on standard, well-controlled scenarios. They have yet to systematically investigate how rain, fog, or other degradation factors distort cross-modal alignment, leaving models vulnerable in complex, real-world environments [15,16].
To address the limitations mentioned above, this paper proposes a Scene-Aware Degradation Universal ReID framework (SCA-UniReID). Built upon the dual-encoder architecture of CLIP [17], the proposed method introduces a scene-aware degradation modeling mechanism to explicitly characterize environmental factors such as rain, fog, and low-light conditions. Specifically, we develop a Scene-Aware Degradation CLIP (SCA-CLIP), which incorporates a scene perceiver to learn degradation-aware representations and guide the visual encoder to focus on identity-relevant features under adverse conditions.
Furthermore, we design a dual-textual semantic guidance mechanism consisting of a target-oriented prompt and a scene-aware prompt. The former enhances identity-discriminative information, while the latter adaptively captures scene-specific degradation characteristics. An adaptive control module dynamically balances the contributions of the two prompts, enabling effective disentanglement of identity semantics from degradation noise (i.e., decomposing the latent representation into interpretable and minimally correlated subspaces). This design allows the model to achieve robust cross-modal alignment while accounting for environmental interference, thereby significantly improving ReID performance under complex, real-world conditions.
  • We systematically analyze the challenges posed by rain-and-fog coupled degradations to existing ReID systems and reveal the limitations of both image-enhancement and robust-feature paradigms under extreme weather.
  • We propose SCA-UniReID, a scene-aware universal ReID framework that integrates dual textual prompts—target-oriented and degradation-oriented—into a CLIP-style dual-encoder architecture, enabling fine-grained disentanglement of identity semantics from weather noise while preserving discriminability.
  • We conduct comprehensive experiments on ship and pedestrian benchmarks under multiple adverse-weather protocols; results demonstrate that SCA-UniReID surpasses state-of-the-art methods and maintains strong generalization across unseen conditions.

2. Related Works

2.1. Image Preprocessing Methods in Re-Identification

Complex environments can cause image degradation, making it difficult for ReID models to extract stable identity features. Image enhancement methods aim to improve image quality through preprocessing, enabling better visibility in complex environments, thus improving feature extraction performance. Jiao et al. [18] combined super-resolution convolutional networks with ReID networks to enhance the re-identification performance of low-resolution images. To further improve the scale adaptability of super-resolution methods, Wang et al. [19] adopted a cascaded SRGAN structure, progressively reconstructing missing details to improve the super-resolution techniques’ ability to adapt to different scales. Zhang et al. [20] proposed a frequency-domain modeling framework based on wavelet decomposition, which enhances detail recovery capability through multi-scale feature decomposition and makes the feature distribution conform to the atmospheric scattering model. In addition, Liu et al. [21] proposed a prior model based on inverse haze density correction, which models the transmittance via pixel-level gamma correction, improving the generalization performance of the method in various degraded scenarios. Mao et al. [22] proposed the FFSR module and designed a dual-branch module to extract resolution-invariant features, further optimizing the detailed representation of target re-identification. Huang et al. [23] proposed a pedestrian re-identification framework that addresses illumination changes. By utilizing the Retinex theory for illumination decomposition, they designed a bottom-up attention network aimed at eliminating interference in low-light environments. In real-world surveillance scenarios, adverse weather conditions such as rain and haze can significantly degrade image quality, thereby affecting the extraction of discriminative features for person re-identification (ReID). To address this issue, existing studies have explored joint deraining and dehazing methods toenhance the visibility of degraded images. For example, Ragini et al. [24] proposed a Single-Stage V-Shaped Network (S2VSNet) for end-to-end image deraining and dehazing. Similarly, Xie et al. [25] combined the dark channel prior (DCP) with deep learning models, further improving image restoration performance through atmospheric light modeling. Although these methods primarily focus on low-level visual enhancement, they effectively reduce environmental interference and provide more reliable inputs for subsequent high-level tasks. Despite the progress of these methods under degraded conditions, they still have limitations. Since image enhancement is merely a preprocessing step for the ReID task, it is somewhat disconnected from the target recognition task. The enhanced image may have domain shifts compared to the real high-quality image, leading to feature extraction and matching biases, which ultimately affect recognition performance.

2.2. Discriminative Feature Learning Methods in Re-Identification

Different from strategies centered on image enhancement, such technical routes target the ReID task for direct optimization, allowing feature extractors to steadily capture identity-related representations within complicated scenes, and bypassing the risks of information missing and domain migration that may be triggered by image enhancement operations [26]. At present, relevant studies are mainly split into two branches: the first is degradation-robust learning via feature decoupling, and the second is end-to-end modeling driven by multi-task synergistic optimization [27].
The former branch (degradation-robust learning via feature decoupling) is designed to strengthen the environmental adaptability of ReID models in complex scenarios. Huang et al. [28] put forward a degradation-agnostic learning architecture that extracts identity features while suppressing degradation interference in an unsupervised manner. Zeng et al. developed a dual-branch network structure to offset the disturbance of illumination fluctuations on identity characterization for all-weather application scenes. Moreover, to quantitatively evaluate the anti-degradation performance of conventional ReID models, Chen et al. [29] built a specialized benchmark dataset for degradation robustness testing in object re-identification, and further established a universal ReID baseline with strong generalization capability on this basis. Kanwal et al. [30] pioneered a feature integration framework that combines dark channel prior knowledge with transfer learning, merging global semantic features and local prior clues to elevate the robustness of object re-identification systems.
Joint optimization-driven end-to-end schemes strive to boost image clarity and recognition precision synchronously via shared feature representation and multi-task collaborative learning [31]. For pedestrian re-identification tasks, Zheng et al. [32] tackled the cross-resolution matching challenge by proposing a bilateral resolution-aware identity modeling approach. Lu et al. presented an Illumination Distillation Framework (IDF) that integrates illumination enhancement and knowledge distillation mechanisms, enhancing the model’s adaptability to low-light environments and lifting nighttime ReID accuracy effectively.
In the vehicle re-identification domain, Chen et al. [33] designed a Semi-supervised Joint Defogging Learning (SJDL) framework for foggy-scene ReID tasks. This framework realizes end-to-end cooperative training of defogging and re-identification, and shares degradation-free features across tasks to eliminate domain shift problems existing in traditional two-step pipelines. Besides, the research team further proposed the RVSL training paradigm [34], which fulfills vehicle ReID without relying on manual annotations and clean reference images. For ship re-identification, most existing methods are only applicable to calm sea conditions, whereas the marine environment features strong randomness and is frequently affected by harsh factors like high humidity, heavy fog and low visibility. Yasod et al. [35] developed a thermal imaging-based ship monitoring solution, which is distinctive for conducting fine-grained local feature matching via ship contour information instead of employing standard re-identification pipelines.

2.3. Vision-Language Learning

In recent years, the Vision-Language Pretraining (VLP) paradigm, particularly the introduction of CLIP (Contrastive Language-Image Pretraining), has significantly advanced multimodal representation learning between images and language. CLIP establishes a deep connection between visual content and natural language through contrastive learning on large-scale image-text pairs, demonstrating strong transferability across various domains. In this “pretraining and fine-tuning” paradigm, the quality of the pretrained model plays a crucial role in the optimization difficulty and performance of downstream tasks. Fine-tuning methods based on prompts or adapters have been widely applied in the vision domain. CLIP-Adapter [13] adds a lightweight adapter module on top of CLIP’s image and text encoders to improve task-specific fine-tuning effectiveness. Li et al. [14] first proposed the CLIP-ReID model, using learnable prompts to guide the visual encoder to extract more semantically rich image features. Yu et al. [12] further introduced the CLIP-driven Semantic Discovery Network, which partially addresses the semantic consistency issue between modalities. Based on the above research, we aim to further utilize CLIP’s vision-text representation capability to extract degradation characteristics of complex scenes, thereby advancing the development of the object re-identification field.

3. Preliminaries

In this section, we first review CLIP, followed by the introduction of the proposed Scene-Aware CLIP and Scene-Aware Prompts.

3.1. CLIP Framework

CLIP is a large-scale pretrained vision-language model composed of two encoders: the image encoder I(·) and the text encoder T(·). The image encoder I(·) is mainly based on two backbone networks, ResNet-50 and ViT-B/16, while the text encoder T(·) is implemented using Transformer blocks. In classification tasks, CLIP adopts a simple yet effective method to construct text prompts, usually using templates like “A photo of a [CLS],” where [CLS] is replaced with specific class names. On this basis, CoOp [36] introduces learnable text contexts, represented as “[v]1[v]2...[v]m[CLS].” This method dynamically generates meta-tokens based on different input images and, by combining learnable context vectors, enhances the model’s transferability and performance. By leveraging the capabilities of CLIP and CoOp, deeper semantic understanding can be achieved, allowing for more robust and accurate modeling in re-identification tasks under complex scenarios. Given an image, CLIP calculates the similarity between the image embedding and the text prompt embedding. The entire process involves projecting both the image and text into a shared high-dimensional space for quantitative evaluation of their relationship. The specific objective function is as follows:
s V i , T i = V i ⋅ T i = g v ( i m g i ) ⋅ g T ( t e x t i )
L i 2 t i = − l o g e x p s V i , T i ∑ k = 1 N e x p s V i , T k ,   L t 2 i i = − l o g e x p s V i , T i ∑ k = 1 N e x p s V k , T i
In this context, g V ⋅ and g T ⋅ are linear layers that project the embeddings into the cross-modal embedding space. L i 2 t i represents the image-to-text contrastive loss, while L t 2 i ( i ) represents the text-to-image contrastive loss.

3.2. Scene-Aware-Degradation Module

The Scene-Aware Degradation (SCA) module is a copy of the CLIP image encoder that introduces a small number of zero-initialized connection modules to dynamically control the image encoder, optimizing its feature extraction ability in degraded environments. The core function of SCA is to adjust the output of the image encoder, guiding the model to focus on the target features and suppressing the influence of noise from weather-related degradations (such as rain or fog) on recognition performance.
This study designs corresponding control modules for two different backbone networks: ResNet-50 and ViT-16. In the ResNet-50 structure, the Scene-Aware module uses the output of each layer as a hidden control signal. After passing through the zero-initialized connection modules, the control signal is added to the target encoder. The initialization module consists of convolutional layers, normalization layers, and ReLU functions, with all parameters initially set to zero to ensure that the early stages of training do not interfere with the extraction of target features. As training progresses, the control signal gradually adjusts the behavior of the encoder, focusing more on the target area and ignoring the interference from degradation factors.
In the ViT-16 structure, the control signal comes from the output of the Transformer blocks. These outputs are then combined with relevant layers of the target encoder. By adding the control signal, the predictions are adjusted. The Transformer blocks are connected by a simple fully connected neural network, with these connections also using a zero-initialization strategy to allow for gradual weight adjustments during training, adapting to different environmental conditions.

3.3. Scene-Aware Prompts

In the CLIP model, traditional fixed prompts struggle to accurately characterize complex degradation patterns such as rain, fog, and low-light conditions, and fail to adapt to the regional variability of degradation features. To address this, we propose the ScA-CLIP scene-aware prompt learning mechanism: as shown in Figure 1b, we take degradation types such as clean, rain, and fog as learnable prompts to construct dynamic scene prompt templates; as illustrated in Figure 1a, we use comparative learning to align degraded scene prompts with degraded image features, and clean scene prompts with clean image features within the CLIP embedding space. This not only strengthens the model’s ability to perceive degraded regions but also enables the model to adapt to various adverse weather scenarios without additional annotations, effectively resolving the issue of insufficient adaptability caused by dataset limitations, while reducing the interference of high-level semantic information on the scene-aware module.
The design of the scene-aware prompt is as follows:
p r o m p t s c = A   p h o t o   o f   a   v e s s e l   i n   a   X 1 … X n s c e n e
In this context, X i ( i = 1,2 … , n ) represents randomly initialized learnable labels that learn different degradation factors of the image. This format of the prompt is capable of learning various scene features and adapting to multiple coupled degradation factors, such as rain, fog, low light, etc. As a result, the model takes into full consideration the impact of environmental factors on the target features when performing cross-modal feature alignment.

4. Implementation Method

In this section, we introduce the core idea of our scene semantic-aware feature decoupling network, which aims to separate target vessel semantics from complex, degraded scene interference to enhance visual perception robustness under adverse weather. The network consists of two core stages:
First, semantic prompt construction and feature guidance, as illustrated in Figure 2, where we build learnable scene-aware prompts (e.g., clean, rain, fog) to achieve contrastive alignment between text and degraded/clean image features in the CLIP embedding space, guiding the model to focus precisely on degraded regions;
Second, target-scene encoding and feature decoupling, as illustrated in the figure, the image encoder is split into a target encoder and a scene-aware degradation module, paired with two controller designs based on ResNet-50 and ViT-16 to dynamically regulate feature flow, effectively decoupling target features from scene degradation features while mitigating interference from high-level semantic noise.

4.1. Overview of the Scene Semantic-Aware Feature Decoupling Network Framework

Building upon the previously proposed scene-aware CLIP, this section further develops the scene semantic-aware feature decoupling method, ScA-UniReID, aimed at enhancing the model’s adaptability to different target categories and rain/fog environments. The framework integrates cross-modal semantic guidance with a target-scene decoupling strategy to optimize feature modeling capabilities. This enables the model to precisely extract identity features even under varying degradation conditions, effectively suppressing degradation interference and improving the model’s generalization in rainy or foggy environments.
Specifically, ScA-UniReID uses a two-stage training strategy. In the first stage, the text encoder is trained to optimize the text-image alignment ability through adjustable prompts, allowing the model to adapt to different degradation environments and improving its semantic understanding of identity features. This provides more accurate guidance for subsequent target feature learning. In the second stage, the target encoder and the scene-aware module are jointly trained. The scene-aware module serves as an auxiliary branch and uses a control mechanism to guide the target encoder’s focus on identity information, thereby enhancing the model’s feature extraction stability under various degradation conditions. The training process is as shown in Algorithm 1.
Algorithm 1 Training Process of Scene Semantic-Aware Feature Decoupling Network Framework
Stage 1: Load Pre-trained Models  I i d ⋅ , I s c ⋅ , T ⋅ , Initialize Learnable Parameters
Output: Learnable Text Prompts
1.   f o r   e p o c h = 1 … M   d o
2.    The target encoder I i d ⋅ and the scene-aware module I s c ⋅ extract image features respectively
3.    Define dual text learnable prompts as shown in Equation (3)
4.    The text encoder T ⋅ encodes the dual text as shown in Equation (4)
5.    Update the prompts via backpropagation as shown in Equation (7)
6.  End for
Stage 2: Load Pre-trained Models  I i d ⋅ , I s c ⋅ , T ⋅ , Dual Text Prompts  p r o m p t i d   a n d   p r o m p t s c
Output: Trained Target Encoder I i d ⋅ and Scene−Aware Module I s c ⋅
7.   f o r   e p o c h = 1 … N   d o
8.    The text encoder T(⋅) extracts text features from the dual text prompts
9.    The target encoder I i d ⋅ and scene-aware module I s c ⋅ extract image features respectively
10.     Update the parameters of I i d ⋅   a n d   I s c ⋅ via backpropagation (9)
11.     Fix the parameters of I s c ⋅ and update   I i d ⋅   via backpropagation (10)
12.   End for

4.2. Semantic Prompt Construction and Feature Guidance

In complex scenarios, target identity features are highly coupled with mixed rain and fog noise, making traditional fixed text prompts unable to model both target semantics and degradation patterns precisely. To address this, the first stage of ScA-UniReID (whose overall framework is shown in Figure 3) adopts a dual-text prompting mechanism and a dedicated scene-aware degradation pipeline:
As illustrated in Figure 3, we construct target-scene dual-text prompts by inserting learnable degradation tokens (e.g., rain, fog) into a CLIP-compatible prompt template. This design explicitly decouples semantic modeling: the target prompt guides the frozen target encoder to extract clean vessel identity features, while the scene-aware prompt drives the Scene-Aware Degradation Module—a parallel branch encoder that extracts low-level degradation features (e.g., raindrop textures, fog scattering) and aligns them with degradation semantics in the CLIP embedding space. To further regulate feature flow and avoid semantic interference, we introduce a lightweight controller module that dynamically gates the propagation of target and scene features in the backbone, ensuring effective decoupling of identity and degradation signals.
Via the contrastive loss L i 2 t + L t 2 i , we align the target encoder features with the clean text prompt and the scene module features with the degradation text prompt in the CLIP embedding space. Experimental results validate the effectiveness of these components: on the VesselReID_Adverse and Market_Adverse datasets, ScA-UniReID outperforms the baseline CLIP-ReID and other state-of-the-art methods, achieving 63.2% mAP and 75.9% Rank-1 accuracy on VesselReID_Adverse, with consistent improvements across both ResNet-50 and ViT-16 backbones. This confirms that the collaborative design of dual prompting, the scene-aware degradation module, and the controller enables the model to dynamically adapt to diverse degradation conditions while effectively separating target identity from noise, thus significantly enhancing recognition performance.
In the first stage, to effectively utilize the CLIP text encoder, this paper designs target-scene dual-text prompts, which are used to model target identity information and rain-fog degradation information, in order to enhance text-image alignment ability. The target text prompt and scene-aware prompt are defined as follows:
p r o m p t i d = A   p h o t o   o f   X 1 … X n   v e s s e l   i n   c l e a n   s c e n e .
p r o m p t s c = A   p h o t o   o f   a   v e s s e l   i n   a   X 1 … X n   s c e n e .
Using the pre-trained identity encoder I i d ⋅ ,   s c e n e   e n c o d e r   I s c ⋅ and text encoder T ⋅ , we extract the target identity feature V i d and V s c and bilingual semantic text features. During the training phase, by freezing the parameters of I i d ⋅ and I s c ⋅ , we focus on optimizing the text tokens to learn contextual representations, thereby obtaining a unique textual representation for each identity (ID) and its corresponding scene. The formulation is as follows:
T i d = T ( p r o m p t i d )
T s c = T p r o m p t s c
Finally, based on the principle of Equation (1), the image-text contrastive objective function is defined as follows:
L i 2 t i = − l o g e x p s V i d i , T i d i ∑ k = 1 N e x p s V i d i , T i d i − l o g e x p s V s c i , T s c i ∑ k = 1 N e x p s V s c i , T s c i
In this principle, L i 2 t denotes the image-to-text contrastive loss, L t 2 i denotes the text-to-image contrastive loss, and S ⋅ ⋅ represents the similarity function. N is the batch size. Since multiple images in a single batch may belong to the same identity, this implies that there may be multiple positive samples. Therefore, the computation of the text-to-image contrastive objective function is as follows:
L t 2 i y i = − 1 P y i ∑ p ∈ P y i l o g e x p s V i d p , T i d y i ∑ k = 1 N e x p s V i d k , T i d y i 1 P y i ∑ p ∈ P y i l o g e x p s V s c p , T s c y i ∑ k = 1 N e x p s V s c k , T s c y i
Therefore, the final objective function for the first-stage training is as follows:
L stage 1 = L i 2 t + L t 2 i

4.3. Target-Scene Encoding and Feature Disentanglement

In the second stage, we use an identity encoder and a scene encoder for contrastive learning to effectively disentangle target identity features from rain/fog noise (whose overall framework is shown in Figure 4). The scene encoder is initialized by copying the identity encoder, and then zero-initialized layers are added to the intermediate layers to enable conditional interactive learning with the identity encoder. The scene encoder continuously learns rain/fog degradation features from images, generating an implicit representation vector of the noise, which is then fed into the identity encoder as a conditional signal. Contrastive learning is performed between the two encoders to guide the identity encoder to reduce its focus on noisy regions and enhance its attention to the target identity feature regions.
In our study, contrastive learning objective function is employed to ensure effective alignment of image features and text descriptions in the embedding space. The objective function is defined as follows:
L c o n V , T = − 1 N ∑ i = 1 N q i l o g e x p s V i , T i ∑ j = 1 N e x p s V j , T j
N represents the number of paired embeddings in a training batch, and q i denotes label smoothing. The optimization goal of this function is to maximize the cosine similarity between correctly paired text-image embeddings, while increasing the distance from incorrectly paired samples, enabling degraded images to find the most matching textual descriptions in the embedding space.
To jointly optimize the target identity features and noise features, this section further defines a combined objective function. By using the scene encoder to conditionally control the identity encoder, the latter treats noisy regions as negative samples and target feature regions as positive samples, thereby enhancing its focus on the desired target regions. The objective function is formulated as follows:
L control = L con V i d , T i d + L con V s c , T s c
V i d and V s c denote the target identity features and noise features extracted from the original image by the identity encoder I i d ⋅ and the scene encoder I s c ⋅ respectively. Meanwhile T i d and T s c are the textual representation vectors obtained from the target text and scene-aware prompt words using the text encoder T ⋅ trained in the first stage.
Furthermore, to enhance the adaptation of the target encoder to the target re-identification task, the cross-entropy loss L i d and the triplet loss L t r i are further employed to optimize the identity encoder:
L i d = ∑ i = 1 N − q i l o g p i
L t r i = m a x d p − d q + α , 0
q i denotes the true label of the i-th sample, and p i is the predicted probability of the true label. d p and d q represent the feature distances of the positive and negative sample pairs, respectively, while α is the margin parameter of the triplet loss L t r i . The final objective function for the second stage is as follows:
L stage 2 = λ 1 L control + λ 2 L i d + L t r i

5. Experiment

To evaluate the performance of the proposed ScA-UniReID method in target re-identification tasks, we conduct experiments on both ship and pedestrian datasets and compare our approach with state-of-the-art methods in the field of person and object re-identification. Furthermore, to deeply analyze the contribution of each module to the overall model performance, ablation studies are carried out in this section to investigate the impact of different modules and parameters on the model’s effectiveness. Finally, visualization analyses are provided to demonstrate the recognition performance of different methods under degraded conditions, making the experimental conclusions more intuitive and interpretable.

5.1. Experimental Settings

In the training phase, we adopt modified versions of ResNet-50 and ViT-16 pretrained on CLIP as the backbone networks for feature extraction. The Adam optimizer is used during training, along with data augmentation techniques such as random horizontal flipping, cropping, and erasing. A global attention pooling layer reduces the feature dimension from 2048 to 1024; correspondingly, the text feature dimension is scaled from 512 to 1024 for alignment. For the ViT-16 backbone, the batch size is set to 32, image size to 384 × 256, and the feature dimension is reduced from 768 to 512, while the text feature dimension remains at 512. Additionally, on the VesselReID_Adverse dataset, the batch size is set to 64 and the image size to 384 × 192. Since pedestrian images are generally smaller, following the AGW [37] protocol, the batch size is set to 64 and the image size to 256 × 128 on the Market_Adverse dataset.
In the first training stage, two textual prompt tokens are trained for 60 epochs on each dataset, with an initial learning rate of 3.5 × 10−4, which is adjusted using a cosine annealing scheduler. In the second training stage, the identity encoder and scene encoder use the ResNet-50 backbone and are trained for 120 epochs on the dataset, with an initial learning rate of 3.5 × 10−4. The learning rate is reduced to one-tenth of its current value at the 40th and 70th epochs. When using the ViT-16 backbone, the model is trained for 60 epochs with an initial learning rate of 5 × 10−6, and the learning rate is similarly reduced at the 30th and 50th epochs.
Moreover, to better adapt the network to the target re-identification task when using the ResNet-50 backbone, the two encoders are trained synchronously for the first 60 epochs. Then, in the subsequent 60 epochs, the scene encoder is frozen, and the identity encoder is optimized using Equations (10) and (11).

5.2. Comprehensive Experimental Comparison and Analysis

To comprehensively evaluate the performance of the proposed ScA-UniReID method in the task of person re-identification under rainy and foggy conditions, we conducted extensive experiments on the VesselReID_Adverse and Market_Adverse datasets, comparing it with the baseline model CLIP-ReID [14] and several state-of-the-art re-identification methods, including AGW, TransReID [38], and HRCN [39]. The experimental results indicate that on the VesselReID_Adverse dataset, ScA-UniReID with ResNet-50 as the backbone network achieved a mean Average Precision (mAP) of 63.2% and a Rank-1 accuracy of 75.9%, significantly outperforming the baseline CLIP-ReID, which recorded an mAP of 58.1% and a Rank-1 of 70.2%. This demonstrates the superior feature extraction capability of the proposed method in rainy and foggy environments. However, when using ViT-16 as the backbone, although ScA-UniReID still performed well, its mAP was 61.5%, slightly lower than the ResNet-50 version, suggesting that local feature modeling is critical for re-identification tasks in such scenarios.
Furthermore, as illustrated in Table 1, on the Market_Adverse dataset, ScA-UniReID with ResNet-50 achieved an mAP of 80.8% and a Rank-1 accuracy of 92.0%, showing clear improvements over the baseline CLIP-ReID, which obtained an mAP of 79.7% and a Rank-1 of 91.4%. Interestingly, in terms of Rank-1 accuracy, ScA-UniReID was slightly outperformed by ISM [40], possibly due to the more pronounced dynamic pedestrian features in the Market_Adverse dataset, where identity discrimination relies heavily on local detailed information, and ISM may have an advantage in such cases. However, by optimizing cross-modal alignment through scene-aware text prompt learning, ScA-UniReID surpassed ISM in mAP, highlighting its superior ability to handle dynamic degradation factors and effectively separate target features from noise.
In summary, the experimental results demonstrate that, whether using ResNet-50 or ViT-16 as the backbone, ScA-UniReID exhibits significant advantages in person re-identification tasks under complex environments, particularly in the stability of feature extraction.

5.3. Ablation Studies

5.3.1. Ablation Study on Scene Encoder Architecture

To systematically validate the effectiveness of the proposed scene encoder architecture, this study conducts ablation experiments comparing three representative structural design schemes: (1) A baseline without control, where the encoder structure is identical to that of the image encoder. The encoder is fine-tuned using only pre-trained weights, and it does not exert any control over the identity encoder. (2) An architecture with a scene encoder but without zero initialization, to investigate the impact of parameter initialization on model performance. (3) The complete architecture proposed in this study, which incorporates the scene encoder along with the zero initialization strategy, aiming to maximize the control over the identity encoder.
This experiment is conducted on both ResNet-50 and ViT-16 backbone networks for cross-architecture validation.
As shown in Table 2, when ResNet-50 is used as the backbone, the CLIP+finetune approach achieves an mAP of 60.6%, while ScA-UniReID w/o zero initialization achieves only 52.5%. This indicates that without zero initialization, the scene encoder may cause instability during early training, thus degrading model performance. In contrast, the proposed ScA-UniReID with zero initialization improves the mAP to 63.2%, representing an improvement of 2.6% over CLIP+finetune and 10.7% over the w/o zero version. This validates the critical role of the zero initialization strategy in enhancing model performance.
When ViT-16 is used as the backbone, CLIP+finetune achieves an mAP of 58.1%, while ScA-UniReID w/o zero achieves 57.3%, again indicating that the absence of zero initialization may negatively affect model stability. The proposed ScA-UniReID based on ViT-16 improves the mAP by 1.7% over CLIP+finetune and by 2.5% over the w/o zero version. Although the improvement is relatively smaller compared to ResNet-50, it still confirms the effectiveness of the proposed method.
In summary, the proposed ScA-UniReID framework achieves the best performance across different backbone networks, demonstrating that the scene encoder can effectively guide the identity encoder to enhance its focus on target regions and improve feature extraction capabilities. This leads to improved accuracy in re-identification tasks under complex environmental conditions. Moreover, the experimental results further validate the necessity of the zero initialization strategy, which helps stabilize the early stages of training and fully unleashes the potential of the scene encoder.

5.3.2. Ablation Study on the Objective Function L c o n t r o l

To investigate the role of the proposed objective function L c o n t r o l in model training, we conduct an ablation experiment based on the ViT-16 backbone. The performance of two models—one trained with L c o n t r o l (labeled as “w/L”) and one without it (labeled as “ w / o   L ”)— is compared. As shown in Figure 5, the blue curve (“ w / L ”) outperforms the orange curve (“ w / o   L ”) on both the mAP and Rank-1 metrics, indicating that L c o n t r o l effectively enhances the model’s recognition capability under rainy and foggy conditions. Specifically, this objective function enables the scene encoder to dynamically perceive degradation features in mixed-degraded images through contrastive learning. It then uses this information as a conditional signal to regulate the attention distribution of the identity encoder, guiding it to focus on identity-related regions. During training, this facilitates the disentanglement of target identity features from noise features, accurately separating the identity information of the target and significantly improving the model’s discriminative ability for degraded images.

5.3.3. Ablation Study on Parameters λ 1 and λ 2

To investigate the impact of different values of parameters λ 1 and λ 2 on the overall objective function L s t a g e 2 in the second training stage, this study conducts comparative experiments based on the ViT-16 backbone, setting different weight combinations for λ 1 and λ 2 .
As shown in Table 3, although varying the parameter weights leads to some fluctuations in performance metrics such as mAP and results at different k-values, the overall model performance remains robust. This indicates that the proposed framework exhibits strong generalization with respect to parameter selection. It can achieve satisfactory target re-identification performance without requiring precise parameter tuning, further demonstrating the framework’s good adaptability and stability.

5.4. Visualization Results

5.4.1. Retrieval Performance Comparison

To intuitively demonstrate the retrieval capability of ScA-UniReID under rainy and foggy conditions, this study conducts visualization analysis on the VesselReID_Adverse test set. Specifically, query images along with their top-10 matching results are selected and displayed. The retrieval results are shown in Figure 6.
ScA-UniReID exhibits strong target recognition ability under complex weather conditions such as rain and fog. In contrast, the baseline model shows clear limitations under the same conditions. The visualization results indicate that, in the presence of rain, haze, and wave interference, the baseline model struggles to effectively separate target features from noise, leading to loss of detailed vessel information and degradation in retrieval performance. For example, in the first case, the baseline model is significantly affected by rain interference and fails to capture key structural features, resulting in incorrect matches with targets of similar color but different identity.
In comparison, ScA-UniReID successfully identifies the correct target under the same challenging conditions, validating its strong adaptability to adverse weather environments. This advantage is primarily attributed to the introduction of the scene encoder, which actively learns noise characteristics from degraded regions and optimizes the attention distribution of the identity encoder. As a result, the model remains focused on target-related features even under complex disturbances such as rain and fog.

5.4.2. Feature Visualization

To further verify the feature learning capability of ScA-UniReID under degraded conditions, Figure 6 illustrates a comparison between the target feature attention regions of ScA-UniReID and the baseline method CLIP-ReID in rainy and foggy scenarios.
From the feature visualization results, it can be observed that CLIP-ReID primarily focuses on local high-salience areas of the vessel, such as the bow, stern, and mast, which are structurally prominent parts. as illustrated in Figure 7. However, this localized focus strategy has clear limitations under degraded conditions. When parts of the target are affected by rain or fog noise, the model struggles to form stable global features, leading to incomplete target representation and thus affecting retrieval accuracy.
In contrast, guided by the scene encoder, ScA-UniReID is capable of capturing the overall structure and key identity features of the target more comprehensively. Heatmaps show that ScA-UniReID not only extracts local high-salience features but also evenly distributes attention across the entire outline, hull shape, and distinctive details, ensuring the completeness of the target features. Additionally, under degraded conditions, the scene encoder effectively reduces the model’s focus on degradation features, allowing the identity encoder to concentrate more on identity-related features.
Overall, relying on the scene-aware mechanism, ScA-UniReID enhances the global perception ability of target features, maintaining stable feature representations in rainy and foggy environments, effectively reducing environmental interference. This further validates its advantage in improving recognition robustness.

5.4.3. Robustness Analysis Under Adverse Weather Degradations

In complex environments such as rain and haze, target re-identification faces significant challenges, primarily due to visual information degradation and the superposition of multiple interfering factors. As illustrated in Figure 8b, rainfall introduces randomly distributed rain streaks that occlude key regions of the target, leading to the loss of fine-grained details (first row). Meanwhile, the rain–haze effect reduces image contrast, blurring object contours and textures and thereby weakening discriminative capability (second row). In addition, overcast low-light conditions result in insufficient illumination and increased noise, further affecting the stability of feature extraction (third row). More importantly, these degradation factors often interact and couple with each other, substantially increasing the difficulty of recognition in real-world scenarios and posing higher demands on the robustness and generalization ability of ReID models.
To mitigate these issues, As illustrated in Figure 8a. existing studies often introduce image restoration techniques such as dehazing and deraining as a preprocessing step to improve visual quality. These methods help suppress degradation artifacts, enhance image clarity, and recover structural details, thereby providing more reliable inputs for subsequent feature extraction. However, such approaches mainly focus on low-level visual enhancement and may not fully preserve identity-related discriminative features, which are crucial for robust re-identification under complex degradation conditions.

5.4.4. t-SNE Visualization

To further demonstrate the effectiveness of ScA-UniReID, this section employs the t-SNE visualization method [18], as shown in Figure 9. This analysis compares the feature distribution in the latent space between the baseline model and the proposed method in the second training stage, using 20 randomly selected classes from the dataset.
In Figure 9a, the feature distribution of CLIP-ReID appears scattered, with sample points from the same class spreading over a large area, indicating poor clustering within the same category. Different colors are used to distinguish individual samples for visualization purposes only and do not correspond to specific semantic categories. This suggests that the baseline model struggles to learn stable identity representations under degraded conditions, leading to reduced feature discriminability.
In contrast, as shown in Figure 9b, the features generated by ScA-UniReID exhibit much better intra-class compactness and inter-class separability. The feature points within the same category are more tightly clustered, while different categories are clearly distinguished from each other. Although colors are randomly assigned for visualization, the clustering structure clearly demonstrates improved feature discrimination. This indicates that the proposed method can more accurately extract identity features in rainy and foggy scenarios and effectively differentiate between different classes.

6. Conclusions

This paper presents a novel scene-aware person re-identification approach to confront the fundamental challenge of image degradation arising from the coupled effects of rain, fog, and other adverse factors in real-world scenarios. In practice, severe weather dramatically reduces image contrast and blurs structural details, leading to drastic shifts in the feature space and a sharp decline in model generalization. To systematically alleviate these issues, we intervene at two complementary levels—feature modeling and semantic guidance. First, a dedicated scene encoder together with scene-aware prompt tokens is introduced to inject environmental priors into the encoding stage, enabling the model to dynamically perceive and compensate for rain-and-fog-induced degradation. Second, a contrastive-learning objective between the identity encoder and the scene encoder is employed to explicitly constrain the distributions of target features and environmental noise in the feature space, achieving effective disentanglement. Extensive experiments demonstrate that the proposed ScA-UniReID not only yields significant accuracy improvements under diverse adverse-weather settings but also exhibits strong generalization to previously unseen conditions.

Author Contributions

Conceptualization, S.W. and C.W.; methodology, S.W. and C.W.; software, S.W. and M.Y.; validation, S.W.; formal analysis, S.W.; investigation, M.Y.; resources, S.W.; data curation, S.W.; writing—original draft preparation, S.W. and Y.W.; writing—review and editing, Y.W. and S.W.; visualization, S.W.; project administration, S.W.; funding acquisition, C.W. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the National Natural Science Foundation of China under Grant (62472149).

Data Availability Statement

The original contributions presented in the study are included in the article, further inquiries can be directed to the corresponding author.

Acknowledgments

The authors are thankful to the providers for all the datasets used in this study. We are also thankful to the anonymous reviewers and editors for their comments, which helped to improve this paper.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ReIDRe-identification
ScAScene-Aware Degradation
CLIPContrastive Language-Image Pretraining
VLPVision-Language Pretraining

References

  1. Han, X.; Sheng, H.; Bai, C. Exploring training data-free video generation from a single image via a stable diffusion model. J. Vis. Commun. Image Represent. 2025, 111, 104504. [Google Scholar] [CrossRef] [Scilit]
  2. Oladimeji, D.; Gupta, K.; Kose, N.A.; Gundogan, K.; Ge, L.; Liang, F. Smart transportation: An overview of technologies and applications. Sensors 2023, 23, 3880. [Google Scholar] [CrossRef] [Scilit]
  3. Milledge, T.J.; Pradhan, B.; Shukla, N. Integrating systems methodologies for australian undersea surveillance: A systematic literature review. Syst. Eng. 2025, 28, 471–497. [Google Scholar] [CrossRef] [Scilit]
  4. Zhong, X.; Han, X.; Jia, X.; Huang, W.; Liu, W.; Su, S.; Yu, X.; Ye, M. ICLR: Instance credibility-based label refinement for label noisy person re-identification. Pattern Recognit. 2024, 148, 110168. [Google Scholar] [CrossRef] [Scilit]
  5. Liu, H.; Tian, Y.; Wang, Y.; Pang, L.; Huang, T. Deep relative distance learning: Tell the difference between similar vehicles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 2167–2175. [Google Scholar]
  6. Qian, X.; Fu, Y.; Jiang, Y.-G.; Xiang, T.; Xue, X. Multi-scale deep learning architectures for person re-identification. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 5399–5408. [Google Scholar]
  7. Zhang, R.; Xu, L.; Yang, S.; Wang, L. MambaReID: Exploiting vision mamba for multi-modal object re-identification. Sensors 2024, 24, 4639. [Google Scholar] [CrossRef] [Scilit]
  8. Wang, Y.; Zhang, P.; Liu, X.; Tu, Z.; Lu, H. Unity is strength: Unifying convolutional and transformer features for better person re-identification. IEEE Trans. Intell. Transp. Syst. 2025, 26, 3713–3723. [Google Scholar] [CrossRef] [Scilit]
  9. Zhang, L.; Wang, L.; Wu, Y.; Chen, M.; Zheng, D.; Cai, Y. ISCDFuse: Interval sampling correlation driven visual state space models for multimodal image fusion. Neurocomputing 2025, 640, 130329. [Google Scholar] [CrossRef] [Scilit]
  10. Zhang, Q.; Zhang, M.; Liu, J.; He, X.; Song, R.; Zhang, W. Unsupervised maritime vessel re-identification with multi-level contrastive learning. IEEE Trans. Intell. Transp. Syst. 2023, 24, 5406–5418. [Google Scholar] [CrossRef] [Scilit]
  11. Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning, Virtual Event, 18–24 July 2021; pp. 8748–8763. [Google Scholar]
  12. Yu, X.; Dong, N.; Zhu, L.; Peng, H.; Tao, D. CLIP-driven semantic discovery network for visible-infrared person re-identification. IEEE Trans. Multimed. 2025, 27, 4137–4150. [Google Scholar] [CrossRef] [Scilit]
  13. Gao, P.; Geng, S.; Zhang, R.; Ma, T.; Fang, R.; Zhang, Y.; Li, H.; Qiao, Y. CLIP-adapter: Better vision-language models with feature adapters. Int. J. Comput. Vis. 2024, 132, 581–595. [Google Scholar] [CrossRef] [Scilit]
  14. Li, S.; Sun, L.; Li, Q. CLIP-ReID: Exploiting vision-language model for image re-identification without concrete text labels. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; Volume 37, pp. 1405–1413. [Google Scholar]
  15. Zeng, Z.; Wang, Z.; Wang, Z.; Zheng, Y.; Chuang, Y.Y.; Satoh, S.I. Illumination-adaptive person re-identification. IEEE Trans. Multimed. 2020, 22, 3064–3074. [Google Scholar] [CrossRef] [Scilit]
  16. Lu, A.; Zhang, Z.; Huang, Y.; Zhang, Y.; Li, C.; Tang, J.; Wang, L. Illumination distillation framework for nighttime person re-identification and a new benchmark. IEEE Trans. Multimed. 2023, 26, 406–419. [Google Scholar] [CrossRef] [Scilit]
  17. Liu, W.; Zhong, X.; Zhou, Z.; Jiang, K.; Wang, Z.; Lin, C.-W. Dual-recommendation disentanglement network for view fuzz in action recognition. IEEE Trans. Image Process. 2023, 32, 2719–2733. [Google Scholar] [CrossRef] [Scilit]
  18. Jiao, J.; Zheng, W.-S.; Wu, A.; Zhu, X.; Gong, S. Deep low-resolution person re-identification. In Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018; Volume 32. [Google Scholar]
  19. Wang, Z.; Ye, M.; Yang, F.; Bai, X.; Satoh, S. Cascaded SR-GAN for scale-adaptive low resolution person re-identification. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence {IJCAI-18}, Stockholm, Sweden, 13–19 July 2018; pp. 3891–3897. [Google Scholar]
  20. Zhang, S.; Zhang, X.; Shen, L.; Wan, S.; Ren, W. Wavelet-based physically guided normalization network for real-time traffic dehazing. Pattern Recognit. 2025, 172, 112451. [Google Scholar] [CrossRef] [Scilit]
  21. Liu, Y.; Li, T.; Tan, C.; Ren, W.; Ancuti, C.; Lin, W. Ihdcp: Single image dehazing using inverted haze density correction prior. IEEE Trans. Image Process. 2026, 35, 1448–1461. [Google Scholar] [CrossRef] [Scilit]
  22. Mao, S.; Zhang, S.; Yang, M. Resolution-invariant person re-identification. arXiv 2019, arXiv:1906.09748. [Google Scholar] [CrossRef] [Scilit]
  23. Huang, Y.; Zha, Z.J.; Fu, X.; Zhang, W. Illumination-invariant person re-identification. In Proceedings of the 27th ACM International Conference on Multimedia; Association for Computing Machinery: New York, NY, USA, 2019; pp. 365–373. [Google Scholar]
  24. Ragini, T.; Prakash, K.; Cheruku, R.S. S2VSNet: Single stage V-shaped network for image deraining & dehazing. Digit. Signal Process. 2025, 156, 104786. [Google Scholar]
  25. Shrivastava, P.; Gupta, R.; Moghe, A.A. Joint Deraining and Dehazing Using a CNN with Dark Channel Prior and Atmospheric Light Hybrid Model for Robust Image Restoration. Int. J. Intell. Eng. Syst. 2025, 18, 241–256. [Google Scholar] [CrossRef] [Scilit]
  26. Singh, R.; Sharma, A. STAD-ConvBi-LSTM: Spatio-temporal attention-based deep convolutional Bi-LSTM framework for abnormal activity recognition. J. Vis. Commun. Image Represent. 2025, 110, 104465. [Google Scholar] [CrossRef] [Scilit]
  27. Zhong, X.; Wang, M.; Liu, W.; Yuan, J.; Huang, W. SCPNet: Self-constrained parallelism network for keypoint-based lightweight object detection. J. Vis. Commun. Image Represent. 2023, 90, 103719. [Google Scholar] [CrossRef] [Scilit]
  28. Huang, Y.; Zha, Z.-J.; Fu, X.; Hong, R.; Li, L. Real-world person re-identification via degradation invariance learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 14084–14094. [Google Scholar]
  29. Chen, M.; Wang, Z.; Zheng, F. Benchmarks for corruption invariant person re-identification. In Proceedings of the 35th Conference on Neural Information Processing Systems Datasets and Benchmarks Track, Virtual Event, 6–14 December 2021. [Google Scholar]
  30. Kanwal, S.; Shah, J.H.; Khan, M.A.; Nisa, M.; Kadry, S.; Sharif, M.; Yasmin, M.; Maheswari, M. Person re-identification using adversarial haze attack and defense: A deep learning framework. Comput. Electr. Eng. 2021, 96, 107542. [Google Scholar] [CrossRef] [Scilit]
  31. Xu, S.; Zhou, C.; Xiao, J.; Tao, W.; Dai, T. A dual-branch infrared and visible image fusion network using progressive image-wise feature transfer. J. Vis. Commun. Image Represent. 2024, 102, 104190. [Google Scholar] [CrossRef] [Scilit]
  32. Zheng, W.-S.; Hong, J.; Jiao, J.; Wu, A.; Zhu, X.; Gong, S.; Qin, J.; Lai, J. Joint bilateral-resolution identity modeling for cross-resolution person re-identification. Int. J. Comput. Vis. 2022, 130, 136–156. [Google Scholar] [CrossRef] [Scilit]
  33. Chen, W.-T.; Chen, I.-H.; Yeh, C.-Y.; Yang, H.-H.; Ding, J.-J.; Kuo, S.-Y. SJDL-vehicle: Semi-supervised joint defogging learning for foggy vehicle re-identification. Proc. AAAI Conf. Artif. Intell. 2022, 36, 347–355. [Google Scholar] [CrossRef] [Scilit]
  34. Chen, W.-T.; Chen, I.-H.; Yeh, C.-Y.; Yang, H.-H.; Chang, H.-E.; Ding, J.-J.; Kuo, S.-Y. RVSL: Robust vehicle similarity learning in real hazy scenes based on semi-supervised learning. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; Springer: Cham, Switzerland, 2022; pp. 427–443. [Google Scholar]
  35. Ginige, Y.; Gunasekara, R.; Hewavitharana, D.; Ariyarathne, M.; Rodrigo, R.; Jayasekara, P. Vessel re-identification and activity detection in thermal domain for maritime surveillance. arXiv 2024, arXiv:2406.08294. [Google Scholar] [CrossRef] [Scilit]
  36. Zhou, K.; Yang, J.; Loy, C.C.; Liu, Z. Learning to prompt for Vision-Language Models. Int. J. Comput. Vis. 2022, 130, 2337–2348. [Google Scholar] [CrossRef] [Scilit]
  37. Ye, M.; Shen, J.; Lin, G.; Xiang, T.; Shao, L.; Hoi, S.C.H. Deep learning for person re-identification: A survey and outlook. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 44, 2872–2893. [Google Scholar] [CrossRef] [Scilit]
  38. He, S.; Luo, H.; Wang, P.; Wang, F.; Li, H.; Jiang, W. Transreid: Transformer-based object re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 15013–15022. [Google Scholar]
  39. Zhao, J.; Zhao, Y.; Li, J.; Yan, K.; Tian, Y. Heterogeneous relational complement for vehicle re-identification. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 205–214. [Google Scholar]
  40. Pang, J.; Zhang, D.; Li, H.; Liu, W.; Yu, Z. Hazy re-ID: An interference suppression model for domain adaptation person re-identification under inclement weather condition. arXiv 2021, arXiv:2104.11004. [Google Scholar]
Figure 1. ScA-CLIP: (a) Scene-aware degradation learning in SAD-CLIP; (b) Scene-aware prompt generation.
Figure 1. ScA-CLIP: (a) Scene-aware degradation learning in SAD-CLIP; (b) Scene-aware prompt generation.
Sensors 26 02951 g001
Figure 2. Overview of ScA-UniReID. Our framework comprises: (a) Construction of Adaptive Scene-Aware Perceptron and Controller on ResNet-50 and ViT-16; (b) text encoder based on Scene-Aware Prompts.
Figure 2. Overview of ScA-UniReID. Our framework comprises: (a) Construction of Adaptive Scene-Aware Perceptron and Controller on ResNet-50 and ViT-16; (b) text encoder based on Scene-Aware Prompts.
Sensors 26 02951 g002
Figure 3. Framework of the First Stage Training of the Text Encoder for ScA-UniReID: Dual Prompting and Scene-Aware Feature Alignment.
Figure 3. Framework of the First Stage Training of the Text Encoder for ScA-UniReID: Dual Prompting and Scene-Aware Feature Alignment.
Sensors 26 02951 g003
Figure 4. Framework Diagram of Target-Scene Encoding and Feature Disentanglement in the Second Stage.
Figure 4. Framework Diagram of Target-Scene Encoding and Feature Disentanglement in the Second Stage.
Sensors 26 02951 g004
Figure 5. Comparison Chart of Ablation Study on Objective Function L c o n t r o l .
Figure 5. Comparison Chart of Ablation Study on Objective Function L c o n t r o l .
Sensors 26 02951 g005
Figure 6. Comparison of Retrieval Results Between the Proposed ScA-UniReID Method and the Baseline Model CLIP-ReID.
Figure 6. Comparison of Retrieval Results Between the Proposed ScA-UniReID Method and the Baseline Model CLIP-ReID.
Sensors 26 02951 g006
Figure 7. Feature Attention Comparison Between ScA-UniReID and CLIP-ReID Under Rainy and Foggy Conditions.
Figure 7. Feature Attention Comparison Between ScA-UniReID and CLIP-ReID Under Rainy and Foggy Conditions.
Sensors 26 02951 g007
Figure 8. Qualitative Comparison of Degraded and Restored Images under Adverse Weather Conditions.
Figure 8. Qualitative Comparison of Degraded and Restored Images under Adverse Weather Conditions.
Sensors 26 02951 g008
Figure 9. (a) t-SNE Visualization of the Proposed ScA-UniReID Method; (b) the Baseline Model CLIP-ReID.
Figure 9. (a) t-SNE Visualization of the Proposed ScA-UniReID Method; (b) the Baseline Model CLIP-ReID.
Sensors 26 02951 g009
Table 1. Results of State-of-the-Art Methods on the VesselReID_Adverse and Market_Adverse Datasets (bold indicates the best performance).
Table 1. Results of State-of-the-Art Methods on the VesselReID_Adverse and Market_Adverse Datasets (bold indicates the best performance).
TecnologyConference/
Journal’Year
VesselReID_AdverseMarket_Adverse
mAPRank-1mAPRank-1
AWGTPAMI’2151.264.875.989.9
HRCNICCV’2153.966.372.188.0
CILNeurIPS’2154.468.568.785.4
ISMICME’2154.868.379.292.2
TransReIDICCV’2154.868.078.790.4
RotTransACMMM’2253.666.176.789.2
SJDLAAAI’2256.469.863.483.0
PHACVPR’2352.565.576.288.7
CLIP-ReID (ViT-16)AAAI’2357.670.978.889.9
CLIP-ReIDAAAI’2361.273.379.791.4
DenoiseRepNeurIPS’2453.268.079.889.9
DCCCICASSP’2447.363.4--
DHCCNTCSVT’2436.856.0--
CCLIJCNN’2550.566.4--
ScA-UniReID (ViT-16)-59.871.881.291.6
ScA-UniReID-63.275.980.892.0
Table 2. Comparison Results of Different Scene Encoder Architectures on the VesselReID_Adverse Dataset.
Table 2. Comparison Results of Different Scene Encoder Architectures on the VesselReID_Adverse Dataset.
BackboneTecnologymAPk = 1k = 5k = 10
CLIP+fineune60.673.389.993.6
ResNet-50ScA-UniReID w/o zero52.569.887.892.2
ScA-UniReID(ours)63.275.991.094.6
CLIP+finetune58.170.888.592.8
ViT-16ScA-UniReID w/o zero57.369.987.192.5
ScA-UniReID(ours)59.871.888.793.2
Table 3. Comparison Results of Different Weight Settings on the VesselReID_Adverse Dataset.
Table 3. Comparison Results of Different Weight Settings on the VesselReID_Adverse Dataset.
λ 1 λ 2 mAPk = 1k = 5k = 10
1159.871.888.793.2
1259.872.288.793.2
1359.472.388.392.9
2259.772.588.692.7
3159.373.388.492.9
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wei, S.; Wang, Y.; Yang, M.; Wang, C. A Scene-Aware Degradation Universal Re-Identification Framework for Adverse Weather. Sensors 2026, 26, 2951. https://doi.org/10.3390/s26102951

AMA Style

Wei S, Wang Y, Yang M, Wang C. A Scene-Aware Degradation Universal Re-Identification Framework for Adverse Weather. Sensors. 2026; 26(10):2951. https://doi.org/10.3390/s26102951

Chicago/Turabian Style

Wei, Siwei, Yuxin Wang, Mingxuan Yang, and Chunzhi Wang. 2026. "A Scene-Aware Degradation Universal Re-Identification Framework for Adverse Weather" Sensors 26, no. 10: 2951. https://doi.org/10.3390/s26102951

APA Style

Wei, S., Wang, Y., Yang, M., & Wang, C. (2026). A Scene-Aware Degradation Universal Re-Identification Framework for Adverse Weather. Sensors, 26(10), 2951. https://doi.org/10.3390/s26102951

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop