2.3.1. Masked Prediction Feature Alignment
Typically, self-supervised learning based on a masked model will introduce additional masked information into the embedding computation during the pre-training stage. However, the training data used in the fine-tuning stage lack a masking strategy, and the optimization objectives differ between the two stages, leading to inconsistent learning goals at different stages. Thus, we employ a dual-branch vanilla Transformer encoder module to separately learn masked feature representations
and visible feature representations
.
where
and
denote masked embeddings and unmasked embeddings (visible embeddings),
stands for positional encoding, and
refers to the momentum encoder.
The core insight of representation learning lies in the fact that latent spaces capturing intrinsic data distributions possess far lower dimensionality than original input spaces. Latent embeddings can highlight meaningful temporal patterns embedded in time-series inputs while suppressing irrelevant measurement noise. To realize this property, leveraging the consistency of deep feature space representations to construct pretext tasks for self-supervised learning is a significant approach. Thus, we predict new feature representation of masked region based on the feature representation from visible parts learned by the encoder, while ensuring that the predicted masked features remain consistent with the original masked feature representations learned by the momentum encoder.
Specifically, a cross-attention Transformer module takes unmasked (visible) features and mask position information as inputs to reconstruct masked representations. We reinitialize a masked feature vector based on the masked position as the prediction target, and we then use it as the query of the cross-attention transformer. In addition, we construct the value and key from the feature representation
of the visible input part. Then, we can obtain the masked prediction feature representation
through the cross-attention transformer module. Through the above operations, we obtained masked feature representations from two different views, thus enabling feature representation alignment using contrastive learning. The core principle of contrastive learning is to narrow the embedding distance between positive sample pairs and enlarge the gap between negative pairs. Differently, we only align positive sample pairs constructed by the masked prediction feature representation and its original masked feature representation learned by momentum encoder without considering negative sample pairs. Since the masked feature representations from the momentum encoder and predictor are both continuous latent vectors, we use mean squared error (MSE) loss to form the feature prediction alignment optimization objective
shown in Equation (
4).
where
denotes the stacked cross-attention and feed-forward layers.
is the predicted masked feature from the predictor, and
is the original masked feature extracted by the momentum encoder.
We implement a gradient-stop strategy on the momentum encoder, whose weights are updated via a moving average (MA) rule throughout pre-training. The detailed update rules for the momentum encoder and encoder are formulated as follows:
where
is a momentum coefficient for the momentum-based moving average. With a large value of
, the momentum encoder slowly approximates the encoder. For the proposed method, we find
performs effectively. Further, ∇ is the gradient and
denotes the learning rate for stochastic optimization. As shown in Equation (
6), the direction of updating
completely differs from that of updating
. Finally,
converges to equilibrium by the slow-moving average. At each training iteration, only the encoder and the predictor receive gradient updates derived from alignment loss.
During fine-tuning, only the vanilla Transformer encoder is retained to extract feature representations from complete PDW sub-sequences, whose outputs are fed into the classification head for downstream emitter recognition. The momentum encoder is discarded at this stage, ensuring masked representations are decoupled from unmasked information for stable inference.
2.3.2. Masked Discrete Feature Alignment
Conventional masked self-supervised learning (SSL) paradigms solely rely on reconstruction loss as their core optimization objective. Under this training target, the encoder is forced to prioritize pixel-level or token-level recovery of original input PDW sequences, and the learned feature representations tend to overfit trivial noise and measurement interference unique to the training dataset rather than extracting discriminative semantic patterns related to radar emitter categories. When transferred to downstream radar recognition tasks, such noise-biased embeddings often converge to suboptimal solutions and degrade the model’s classification accuracy, especially on real measured PDW data contaminated by pulse loss and complex electromagnetic clutter.
To address this inherent limitation of single reconstruction loss, we draw inspiration from vector quantization (VQ) theory proposed in [
38] and designed an end-to-end tokenizer module. This module bridges continuous latent feature space and discrete semantic space, converting dense feature representations extracted from masked PDW sub-sequences into sparse, interpretable discrete codewords without introducing extra offline clustering steps. The core of the tokenizer is a codebook embedding matrix
, which stores
K latent prototype vectors with dimension
d identical to the embedding size of PDW sub-sequences.
During the pre-training forward propagation, two groups of masked feature vectors are simultaneously fed into the tokenizer: the original masked embedding
generated from the preprocessing module (serving as the discrete target), and the predicted masked representation
output by the cross-attention predictor branch (serving as the prediction). It should be emphasized that the tokenizer operates on the encoder-input-level
, not on the output of the momentum encoder; the momentum encoder is only used in the MPA branch for continuous feature alignment. For each input sub-sequence vector, we calculate its similarity to every prototype vector in the codebook via cosine similarity, a metric insensitive to vector magnitude and suitable for measuring the matching degree of temporal feature distribution of multi-parameter PDW data. Each masked embedding
and predicted masked representation
are then, respectively, mapped to the codeword with the highest similarity score, and the corresponding serial number of this optimal codeword are recorded as the token index
and
.
can be expressed as the maximum cosine similarity between
and the codebook vectors as shown in Equations (
7) and (
8). Similarly,
is the maximum cosine similarity between
and the codebook vectors.
where
denotes the cosine-similarity logit between predictor output embedding
and each codebook prototype vector
, and
K is the total codebook size.
Softmax with temperature
converts logits
into predicted categorical probability distribution
as shown in Equation (
9). Thus, the probability of predicted prototype index
can be expressed as
, and the probability of target distribution from the original masked input prototype index
can be expressed as
.
Each discrete codeword corresponds to a typical temporal pattern of PDW sub-sequences formed by the coupling of PRI, RF, and PW multi-dimensional pulse parameters. Therefore, the matched token indices naturally serve as intrinsic self-supervised supervision signals that reflect the inherent structural characteristics of masked PDW segments. On this basis, we construct the masked discrete representation alignment optimization objective
, as formulated in Equation (
10). This loss adopts cross-entropy to constrain the token indices assigned to the predicted representation
and the original masked embedding
to be consistent. By minimizing the cross-entropy loss between two sets of token labels, the model is guided to learn discrete embeddings with strong category discrimination, which complements the continuous feature alignment loss of the prediction branch and further improves the robustness of PDW feature representation under noisy real measurement environments. Following the self-distillation paradigm, a stop-gradient is applied to the target-side token distribution
, and only the prediction-side distribution
receives gradients. This prevents representational collapse by ensuring the target remains a stable objective that the predictor must learn to match.
During the forward pass, the index of the nearest embedding is assigned as the current embedding vector’s discrete codeword. During the backward pass, the stop-gradient operation takes effect, and the non-differentiable argmax mapping is resolved via the straight-through estimator (STE) [
39]. The gradient bypasses the codebook vectors and is propagated directly to the masked feature vectors prior to discrete assignment, meaning the non-differentiability issue is resolved. In addition, the codebook is updated exclusively via an exponential moving average (EMA) algorithm without gradient backpropagation, as formulated in the Equations (11a–c). At training step
t, each prototype
is updated by weighting the
closest input vectors
at the current step and its historical value from step
.
counts the total number of patches assigned to prototype
i,
accumulates weighted feature sums, and
is the EMA discount factor.