Next Article in Journal
A Modified Generalized Orthogonal Matching Pursuit Imaging Algorithm for High-Resolution Spaceborne iFMCW-SAR
Previous Article in Journal
Evaluating the Vertical Accuracy of Global DEMs Using ICESat-2 and Its Cascading Impact on HAND-Based Flood Modeling in a Low-Gradient Coastal Plain
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

MSPaDet: A Multi-Scale Phase-Aware Denoising Method for Target Detection in SAR Images

College of Computer and Mathematics, Central South University of Forestry and Technology, Changsha 410004, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(10), 1513; https://doi.org/10.3390/rs18101513
Submission received: 17 March 2026 / Revised: 7 May 2026 / Accepted: 8 May 2026 / Published: 11 May 2026
(This article belongs to the Section AI Remote Sensing)

Highlights

What are the main findings?
  • Explicit phase-coherent multi-scale frequency decomposition provides a principled mechanism for separating weak SAR targets from clutter, as it suppresses low-coherence speckle responses while preserving structurally meaningful directional cues.
  • Joint modeling of phase and energy is shown to yield a more discriminative and structurally stable representation than single-cue enhancement, enabling more faithful characterization of weak, irregular, and anisotropically scattered targets under complex background interference.
What are the implication of the main findings?
  • The results suggest that SAR target detection can benefit from moving beyond purely spatial enhancement or coarse magnitude-based frequency selection toward explicit phase-aware, direction-sensitive, and multi-scale representation learning.
  • By improving the structural fidelity and stability of target representations under speckle, clutter, and anisotropic scattering, the proposed strategy offers a practically meaningful route toward more reliable SAR perception in Earth observation, maritime surveillance, and safety-critical monitoring applications.

Abstract

Synthetic Aperture Radar (SAR) target detection remains challenging due to coherent speckle corruption, weak-scattering targets with degraded structural cues, and cross-scale inconsistencies under anisotropic scattering. To tackle these challenges, this paper presents MSPaDet, a novel multi-scale phase-aware denoising detection framework that advances SAR target detection by deeply integrating phase coherence with multi-scale representation learning. The proposed method introduces explicit dual-tree complex wavelet transform decomposition to generate direction-selective complex sub-bands, enabling fine-grained sub-band modulation. Within the framework, an SCFRDeno module suppresses speckle-dominant responses while preserving high-frequency structures via phase-coherence-guided reweighting, and a PaSCA block further refines features through input-adaptive spatial focusing and region reweighting. Extensive experiments on public SAR detection benchmarks—including MSAR, SAR-Aircraft-1.0, and SARDet-100K—demonstrate that our approach consistently outperforms state-of-the-art methods in detection accuracy, robustness, and cross-scenario generalization, with moderate computational cost, showing promising potential for practical deployment in Earth observation and safety monitoring systems.

1. Introduction

Synthetic Aperture Radar (SAR) is an active microwave imaging sensor that enables all-weather observation by penetrating clouds, rain, and fog [1]. Reliable detection and localization of ships, vehicles, and critical infrastructure in complex land–sea backgrounds are crucial for Earth observation applications and safety-critical decision-making [2]. However, as shown in Figure 1, real-world SAR targets are often small, weak, and irregular, and are heavily affected by speckle and clutter. Coherent speckle degrades texture cues and target–background contrast [3,4], while small and weak targets provide limited cues that are easily submerged by noise [5,6]. Moreover, anisotropic scattering induces cross-scale response instability, impairing multi-scale fusion and generalization.
Image resolution directly affects target visibility and target–background separability in SAR target detection [7]. In practical spaceborne SAR systems, high-resolution imaging is commonly achieved through spotlight or sliding spotlight modes. Spotlight SAR provides finer azimuth resolution for detailed observation, whereas sliding spotlight SAR offers a practical trade-off between high resolution and wider azimuth coverage. Although these acquisition modes improve structural detail representation, they also make speckle, clutter, and weak-target detection more challenging, which further motivates the need for robust feature modeling in SAR target detection [8,9].
Before the rise of deep learning, SAR detection largely relied on clutter modeling and threshold-based schemes such as CFAR, often assisted by saliency analysis, clustering, or polarization information to improve detectability [1,2,10,11]. Recent deep learning methods have gradually addressed the problems of heavy speckle, small targets, irregular targets, and scale instability by embedding denoising and detail-preserving modules into detectors or integrating conventional frequency attention (as illustrated in Figure 2), strengthening multi-scale fusion and training strategies, and introducing consistency constraints and adaptive mechanisms to mitigate performance fluctuations caused by resolution variations [12,13,14,15,16,17,18]. SAR ship detection is also an important branch of SAR target detection, especially in maritime surveillance and large-scene ocean observation. Recent studies, such as LS-SSDD-v1.0 for small-ship detection and scale-aware dimension-wise attention networks for SAR small-ship instance segmentation, further highlight that small object scale, weak contours, and complex background interference remain key challenges in SAR target detection [19,20]. Despite these advances, common limitations remain: spatial enhancement still struggles to denoise without erasing weak structural cues; scale- and orientation-related information is under-exploited for irregular, scale-sensitive targets; and fixed model parameters adapt poorly to changing noise and clutter, limiting robustness and generalization.
We argue that overcoming these bottlenecks calls for explicitly introducing scale- and orientation-aware decomposition in the frequency domain. Such representations tend to concentrate structure while dispersing noise, and phase cues are less sensitive to magnitude fluctuations, providing complementary directional information. Moreover, because SAR targets are local and scale-dependent, coarse global frequency selection can suppress weak targets, motivating fine-grained sub-band modulation with input-adaptive refinement.
Based on these insights, we propose MSPaDet, a frequency-domain multi-scale phase-aware detection framework. MSPaDet introduces feature-level DTCWT to decompose features into multi-scale, multi-orientation complex sub-bands and modulates them using phase and energy cues for structure enhancement and speckle suppression. Within MSPaDet, SCFRDeno performs phase-consistency-guided sub-band reweighting and reconstruction to suppress coherent noise while preserving high-frequency details. The embedded PaSCA block then refines the fused features via phase-direction cues, spatial focusing, and region selection, improving stability under clutter and anisotropic scattering.
Our main contributions are as follows:
  • We propose MSPaDet, a phase-aware detection framework using DTCWT for explicit directional sub-band modulation to enhance weak targets in speckle noise.
  • We introduce SCFRDeno–PaSCA, which leverages phase consistency, local energy, and phase-direction cues for sub-band denoising and input-adaptive refinement.
  • Experiments on MSAR, SAR-Aircraft-1.0, and SARDet-100K show superior accuracy and robustness with modest overhead, confirming practical utility.
The remainder of this paper is organized as follows. Section 2 reviews related work. Section 3 presents the MSPaDet framework and its core modules, SCFRDeno and PaSCA. Section 4 describes the experimental setup and results, followed by ablation studies. Section 6 concludes this paper.

2. Related Work

2.1. Object Detection in SAR Imagery

Prior studies on SAR target detection can be broadly summarized into three directions: speckle suppression, detector design, and generalization across heterogeneous imaging conditions.
Speckle suppression. Coherent speckle degrades texture and edge cues in SAR imagery [21,22]. Existing speckle suppression methods can be broadly divided into imaging-level despeckling approaches and network-level feature enhancement methods [23,24,25]. Imaging-level methods, such as traditional filtering and model-based despeckling, are usually applied as preprocessing steps before detection. Although these methods can reduce speckle noise, overly aggressive despeckling may smooth weak target boundaries and erase local structural details that are important for detection. Recent works therefore tend to incorporate speckle suppression into detection networks, including soft-threshold quantization with contextual fusion [26], denoising-aware detection frameworks, and joint denoising–detection learning strategies such as DS-YOLO [27]. Compared with image-level preprocessing, network-level enhancement is generally more task-aware, but many existing methods still mainly rely on magnitude or spatial-domain modulation and rarely exploit phase consistency and directional scattering cues explicitly.
Detector design. From the perspective of framework design, existing SAR target detectors can be roughly categorized into two-stage and one-stage methods. Two-stage detectors first generate candidate proposals and then refine them for classification and localization, which can improve localization accuracy but usually introduce higher computational cost. One-stage detectors directly predict object categories and bounding boxes, and are widely adopted in SAR settings due to their favorable speed–accuracy trade-off [14,28,29,30]. Lightweight and anchor-free detectors further improve efficiency and flexibility for dense or small SAR targets [31,32]. Transformer-based and end-to-end detectors have also been introduced to enhance global context modeling and multi-scale representation for arbitrarily oriented targets [33,34], e.g., LD-Det [35]. Many methods also strengthen feature fusion and recalibration to improve small and irregular target detection [36]. However, most detector designs still focus on spatial-domain feature aggregation or architectural refinement, while explicit frequency-domain phase coherence and direction-aware scattering representations remain insufficiently explored.
Generalization and domain adaptation. Domain shifts caused by sensor and imaging variations have motivated adaptation-oriented studies, including domain-adaptive Transformer detectors for unlabeled multi-source SAR data [37] and unsupervised adaptation with cross-domain feature interaction and data balancing [38]. SARDet-100K further highlights the challenge of cross-scale generalization at the benchmark level [27]. Frequency-aware modeling has also been explored, such as polar-Fourier spatial–frequency fusion [39] and in-detector schemes that suppress speckle-dominant responses while retaining discriminative cues. Overall, most existing methods still focus on spatial-domain learning or coarse frequency selection, with limited explicit modeling of phase coherence and directional scattering cues.

2.2. Adaptive Enhancement and Frequency-Domain Modulation

Spatially adaptive enhancement and region selection. Input-adaptive spatial enhancement has progressed beyond conventional re-weighting toward designs that emphasize receptive-field diversity and region-wise heterogeneity, as exemplified by RepLKNet [40], InternImage [41], and RFAConv [42].
Frequency-domain enhancement and selective modulation. Several works incorporate frequency information for selective enhancement, including frequency-guided strategies [43], frequency-domain attention for distillation [44], and frequency-enhanced pixel-attention frameworks [45]. However, most frequency-domain enhancement methods focus on magnitude modulation with global or static selection, and rarely model directional structure or phase consistency explicitly [46].

3. Method

3.1. Overview

The macro design philosophy of MSPaDet is to enhance SAR targets in a feature-level, frequency-domain, and phase-aware manner. Instead of relying on image-level despeckling, which may smooth weak target boundaries and high-frequency structural details, MSPaDet embeds speckle suppression directly into the detection pipeline. The design is guided by three principles: multi-scale and multi-directional decomposition for weak structure representation, phase-coherence-guided reweighting for speckle suppression, and adaptive refinement across spatial, channel, and regional dimensions for target-related response enhancement. Accordingly, SCFRDeno performs phase-aware frequency-domain reconstruction, while PaSCA further refines the reconstructed features in a task-adaptive manner.
The overall MSPaDet framework is shown in Figure 3. Built upon standard object detectors, MSPaDet introduces modular SCFRDeno modules for feature-level speckle suppression and structural enhancement. SCFRDeno performs explicit multi-scale wavelet decomposition and phase-aware sub-band reweighting, which suppresses coherent speckle while preserving high-frequency structural cues. The resulting representations are more consistent across scales and orientations, alleviating cross-scale inconsistency under anisotropic scattering and improving robustness to scale variation. With low parameter overhead and a plug-and-play design, SCFRDeno can be inserted into backbone stages or FPN fusion layers; in our implementation, it is placed at a key multi-scale fusion node to maximize its impact on feature quality and detection accuracy.

3.2. SCFRDeno

MSPaDet integrates the modular SCFRDeno module for feature-level speckle suppression and structural enhancement. Notably, PaSCA is an internal sub-module of SCFRDeno, which further performs input-adaptive refinement on the frequency-enhanced features. In this subsection, we introduce the frequency-domain formulation of SCFRDeno and define the intermediate feature F freq that is subsequently refined by PaSCA. The detailed design of PaSCA is deferred to the dedicated section.
Algorithm 1 SCFRDeno
  1:
Input: Image tensor I
  2:
Output: Detection results B
  3:
function Init
  4:
      Initialize backbone f bb (ResNet-50, outputs { C 2 , C 3 , C 4 , C 5 } )
  5:
      Initialize FPN base (out_channels =   256 , start_level =   1 , num_outs =   5 )
  6:
      Initialize DTCWT (levels =   2 ) and PerBandNorm
  7:
      Initialize PATM with kernels ( 3 × 11 ) and ( 5 × 9 )
  8:
      Initialize PaSCA blocks (Phase, Channel, Spatial, Group-Selection)
  9:
      Initialize detection head f head
10:
end function
11:
function Forward(I)
12:
       X Preprocess ( I )
13:
       ( C 2 , C 3 , C 4 , C 5 ) f bb ( X )
14:
      Build FPN features { P 0 , P 1 , P 2 , P 3 , P 4 }
15:
      for  t = 0 to 1 do
16:
             S t P t
17:
             ( L t , { H t ( d ) } ) DTCWT ( S t )
18:
             ( L t , { H t ( d ) } ) PerBandNorm ( L t , { H t ( d ) } )
19:
            if Phase consistency enabled then
20:
                   ( L t , { H t ( d ) } ) PhaseReweight ( L t , { H t ( d ) } )
21:
            end if
22:
             F t dtcwt FuseBands ( L t , { H t ( d ) } )
23:
             F t patm PATM ( S t )
24:
             F ˜ t Fuse ( F t dtcwt , F t patm )
25:
             P t PaSCA ( F ˜ t )
26:
      end for
27:
       B f head ( { P 0 , P 1 , P 2 , P 3 , P 4 } )
28:
      return  B
29:
end function
DTCWT decomposition. SCFRDeno adopts Dual-Tree Complex Wavelet Transform (DTCWT) to obtain a multi-scale, multi-orientation complex representation. In the one-dimensional case, the low-pass and high-pass decompositions at the c-th layer are
a c ( t ) [ k ] = ( x h 0 ( t ) ) 2 j ,
d c ( t ) [ k ] = ( x h 1 ( t ) ) 2 j ,
t { a ,   b } .
The complex wavelet coefficient is formed by the real and imaginary outputs:
w c = d c a + i d c b .
In the two-dimensional case, six directional sub-bands are obtained via separable row–column filtering, producing complex coefficients
W ( c , θ ) = D ( c , θ ) a + i D ( c , θ ) b , θ { ± 15 ° ,   ± 45 ° ,   ± 75 ° } .
At each scale, DTCWT yields six oriented high-pass sub-bands along with a low-pass residual. The amplitude and phase are computed as
| W ( c , θ ) | = D ( c , θ ) a 2 + D ( c , θ ) b 2 ,
W ( c , θ ) = atan 2 D ( c , θ ) b , D ( c , θ ) a .
Phase consistency and sub-band reweighting. We measure local phase coherence on each directional sub-band using
ρ c , θ ( p ) = 1 | N ( p ) | q N ( p ) exp i W c , θ ( q ) W c , θ ( p ) , ρ c , θ ( p ) [ 0 , 1 ] .
where N ( p ) is a spatial neighborhood. In practical applications, N ( p ) is defined as a k × k local window on each directional sub-band ( k = 3 in this paper). We then apply phase-aware re-weighting to suppress low-coherence sub-band components:
W ˜ ( c , θ ) ( p ) = ρ c , θ ( p ) · W ( c , θ ) ( p ) .
Normalization and feature aggregation. To prevent high-energy amplitude components from dominating subsequent feature learning while preserving phase information for structural orientation characterization, normalization is applied only to the amplitude:
W ˜ ( c , θ ) ( p ) = | W ˜ ( c , θ ) ( p ) | μ ( c , θ ) σ ( c , θ ) + ε e i W ˜ ( c , θ ) ( p ) .
The normalized coefficients preserve phase information for structural characterization. We concatenate the real and imaginary parts of all scales and orientations along the channel dimension to obtain a multi-scale, multi-directional complex representation. Here, “phase-aware” refers to the analytic sub-band phase of the DTCWT coefficients, which encodes local orientation and phase coherence, rather than the electromagnetic phase of raw complex SAR measurements.
The concatenated features are then compressed and aligned using 1 × 1 convolutions to produce F freq . In MSPaDet, F freq is refined by PaSCA and passed to the detection head for classification and localization.

3.3. PaSCA

Figure 4 illustrates the proposed PaSCA module, which refines the frequency-enhanced feature (e.g., F freq from SCFRDeno) via complementary phase-, spatial-, and semantic-aware recalibration. PaSCA is composed of three components: (1) phase-aware component fusion, (2) cascaded channel–spatial adaptation, and (3) group-wise spatial gating.
Phase-aware component fusion. To better accommodate anisotropic scattering and speckle interference, we model intermediate features with an amplitude–phase representation and perform direction-aware fusion. Given input features
x R B × C × H × W ,
mapped to the complex plane, each component is represented as a complex wave
z ˜ c = | z c | e i θ c , c = 1 , 2 , , n ,
where | z c | R B × C × H × W represents the characteristic amplitude, while θ c R B × 1 × H × W denotes phase information used to describe structural orientation. The phase is generated through a lightweight per-channel mapping
θ c = Θ ( x c , W θ ) .
where x c R B × C × H × W , the phase estimation module Θ ( · ) outputs θ C R B × C × H × W , and its learnable parameters satisfy W θ R 1 × C × k h × k w . According to Euler’s formula, the complex waveform can be decomposed into real and imaginary parts
z ˜ c = | z c | cos θ c + i | z c | sin θ c .
To capture directional variability, we employ two directional branches
z h = W h t ( | z | cos θ h ) + W h i ( | z | sin θ h ) , z v = W v t ( | z | cos θ v ) + W v i ( | z | sin θ v ) .
where | z | R B × C × H × W and θ h , θ v R B × C × H × W . The learnable projection weights are defined as W h t , W h i , W v t , W v i R C out × C in × 1 × 1 , and the modulated outputs satisfy z h , z v R B × C out × H × W . When two waveforms are superimposed, the resulting amplitude is influenced by their phase difference:
| z r | = | z i | 2 + | z c | 2 + 2 | z i | | z c | cos ( θ c θ i ) .
This motivates phase-modulated aggregation:
o c = k W c k t z k cos θ k + W c k i z k sin θ k .
In implementation, we concatenate the two directional outputs and obtain a compact direction-aware feature
f dir = Conv ( Concat [ z h , z v ] ) .
We further introduce a channel-semantic branch f c and compute branch weights via global pooling and a two-layer MLP:
α i = Softmax ( MLP ( GAP ( f i ) ) ) , i { h ,   v ,   c } .
The final fused feature is obtained by adaptive aggregation:
F = i { h , v , c } α i f i .
Cascaded channel–spatial adaptation. We apply a lightweight two-stage correction on F to enhance salient targets and suppress background interference. First, channel statistics are computed by global average pooling and global max pooling:
F avg c = GAP ( F ) , F max c = GMP ( F ) .
They are fed into a shared two-layer MLP:
g ( x ) = W 1 δ ( W 0 x ) ,
yielding the channel attention map
M c ( F ^ ) = σ g ( F avg c ) + g ( F max c ) R B × C × 1 × 1 .
The channel-enhanced feature is
F ^ = F ^ ( 1 + α M c ( F ^ ) ) .
Next, we compute channel-wise average and max maps
F avg s = Avg c ( F ^ ) , F max s = Max c ( F ^ ) .
and obtain the spatial attention map
M s ( F ^ ) = σ F k h × k w [ F avg s ; F max s ] .
Finally, spatial correction is applied as
Y = F ^ ( 1 + β M s ( F ^ ) ) .
Group-wise spatial gating. To further accommodate heterogeneous clutter and spatially varying noise sensitivity, we introduce group-wise spatial gating for finer-grained region selection. Given an input feature map
x R B × C × H × W ,
we first construct a compact spatial descriptor
s = Avg c ( x ) + Max c ( x ) R B × H × W ,
y = vec ( s ) R B × ( H W ) × 1 .
We generate multi-granularity candidates with G 1 = { 2 ,   4 ,   8 ,   16 } and compute grouping features
u g = ReLU ( Conv 1 d g ( y ) ) , g G 1 .
A micro-selective operator performs soft selection over candidates
α g = exp ( ϕ ( u g ) ) g G 1 exp ( ϕ ( u g ) ) , g G 1 .
t = g G 1 α g u g .
We then construct another candidate set G 1 = { 2 ,   4 ,   8 ,   16 } and generate gate candidates conditioned on t:
v g = σ ( Conv 1 d g ( t ) ) , g G 2 .
The final gate vector is obtained by
a = Sel ( t , { v g g G 2 } ) R B × ( H W ) × 1 .
Finally, we reshape and apply the spatial gate:
A = reshape ( a ) R B × 1 × H × W ,
x = A x .
The group configuration G = { 2 ,   4 ,   8 ,   16 } is adopted to provide multi-granularity spatial gating. Smaller groups capture coarse and stable spatial responses, while larger groups enable finer regional selection for weak or small targets. This setting balances global stability and local adaptivity with a limited candidate space.

4. Experiments

4.1. Setup

4.1.1. Dataset

We evaluate MSPaDet on three public SAR detection benchmarks: SAR-Aircraft-1.0 (16,463 aircraft from seven classes), MSAR/HISEA-1 (60,396 targets from four classes), and SARDet-100K (16,598 images with 245,653 instances over six classes).

4.1.2. Implementation Details

All experiments were conducted on the autoDL platform, utilizing a single NVIDIA RTX 3090 GPU. Optimization was performed within the MMRotate framework, employing the DAdaptAdam optimizer with an automatically adjusted learning rate.
For MSAR, images were resized to 256 × 256 and randomly flipped ( p = 0.5 ); training ran for 36 epochs with a batch size of 64, learning rate 1.0, and weight decay 0.05.
For SAR-Aircraft-1.0, raw images were cropped into 512 × 512 patches (200-pixel overlap) and trained for 12 epochs with a batch size of 32.
For SARDet-100K, images were resized to 512 × 512 , flipped with p = 0.5 , and trained for 12 epochs with a batch size of 12. The detailed information of each dataset is presented in Table 1. It should be noted that Pre. denotes the pretraining dataset; IN denotes the ImageNet-1K pretrained weights.

4.2. Test Results

4.2.1. Results on the MSAR Dataset

Table 2 compares MSPaDet with 20 state-of-the-art detectors on MSAR. MSPaDet achieves 69.53% mAP under AP’07 and 70.25% under AP’12, obtaining the best performance among all methods. Compared with DenoDet, MSPaDet improves mAP by about 1% while using fewer parameters and FLOPs, indicating a favorable accuracy–efficiency trade-off.

4.2.2. Results on the SAR-Aircraft-1.0 Dataset

As reported in Table 3, MSPaDet achieves 68.65% mAP and 69.78% AP12 on SAR-Aircraft-1.0, outperforming all 17 compared methods. Relative to S4Det, MSPaDet improves mAP by about 4% and AP12 by 4.6%, suggesting improved robustness to speckle and complex background structures.

4.2.3. Results on the SARDet-100K Dataset

Table 4 shows results on SARDet-100K. MSPaDet surpasses DiffDet4SAR by 0.9 mAP and maintains consistently stronger performance across diverse scenes. For small-target detection, MSPaDet further outperforms the strongest competing model, MSDFEN, by 0.3 percentage points in mAP s , providing direct evidence of its advantage on small objects. It also exhibits smaller accuracy drops across target types and imaging settings, reflecting better robustness.

4.2.4. Cross-Dataset Transfer Evaluation

To further assess the transferability of MSPaDet beyond independent in-domain evaluation, we conduct a cross-dataset aircraft-level transfer experiment. Specifically, the model is trained on SAR-Aircraft-1.0 and directly evaluated on the aircraft subset of SARDet-100K without target-domain fine-tuning. Since SAR-Aircraft-1.0 contains seven aircraft subclasses, these subclasses are merged into a single aircraft category for aircraft-level evaluation.
As shown in Table 5, MSPaDet achieves an AP of 0.4101 and a Recall of 0.6100 under the cross-dataset setting. For contextual reference, the in-domain mAP on SAR-Aircraft-1.0 and SARDet-100K is 0.6865 and 0.5833, respectively. These results suggest that the proposed method maintains non-trivial transfer capability under heterogeneous imaging conditions, even without adaptation to the target domain.
We note that this experiment is intended to provide supplementary evidence of cross-dataset transferability rather than a strict leave-one-sensor-out evaluation, since reliable image-level sensor labels are not consistently available for all samples in the target dataset.

4.2.5. Comparison with Traditional Despeckling-Based Detection Pipelines

To further verify the necessity of the proposed phase-aware feature-level denoising strategy, we compare MSPaDet with traditional image-level despeckling methods embedded into the same detection pipeline. Specifically, Lee filtering, non-local means (NLM), and SAR-BM3D are applied to both the training and testing images as preprocessing modules, while the annotation files remain unchanged. The same baseline detector is then retrained and evaluated on the corresponding despeckled datasets.
As shown in Figure 5, traditional despeckling methods improve the baseline detector to some extent, but their gains remain limited. MSPaDet achieves the highest mAP, outperforming Lee, NLM, and SAR-BM3D by 5.38, 5.10, and 4.57 percentage points, respectively. These results indicate that image-level despeckling can suppress speckle interference, but may also weaken weak target boundaries and local structural details. In contrast, the proposed feature-level phase-aware denoising better balances speckle suppression and structure preservation for SAR target detection.

4.3. Ablation Study

4.3.1. Ablation on Frequency-Domain Modeling

We investigate whether jointly modeling energy and phase in the frequency domain is more effective than using either cue alone. As reported in Table 6, we ablate SCFRDeno by retaining only the energy branch, only the phase branch, or the full energy–phase design. Using only energy improves mAP on MSAR by 1.12%, and using only phase yields a 1.36% gain, while the full dual-branch collaboration achieves a 3.00% improvement over the baseline without additional computational overhead. These results suggest clear complementarity between energy and phase cues for capturing structural details and directional patterns under speckle contamination. In addition, FP64 provides only a marginal 0.26% mAP increase but introduces substantial overhead; we therefore use FP32 in all experiments to balance accuracy and efficiency.

4.3.2. Ablation Study of Dual-Branch Modeling

Table 7 compares the spatial-only baseline with a variant that introduces DTCWT-based frequency-domain decomposition. Incorporating DTCWT yields consistent improvements across metrics, including a 1.34% mAP gain on MSAR, with the largest benefit on small targets. These results support the advantage of frequency-domain modeling for speckle-robust and direction-sensitive feature representation. Notably, SCFRDeno is a SAR-oriented phase-aware enhancement built on DTCWT sub-bands, rather than a plain stack of wavelet filters, enabling adaptive suppression of speckle-dominant components while preserving structural details.

4.3.3. Sensitivity Analysis of Key Hyperparameters

To justify the default phase-coherence window size, we conduct a sensitivity analysis by evaluating k = { 1 ,   3 ,   5 ,   7 } on MSAR, SAR-Aircraft-1.0, and SARDet-100K while keeping all other settings unchanged. As shown in Figure 6, the performance on all three datasets peaks at k =   3 , indicating that this window size provides the best overall trade-off between local stability and preservation of weak structural details. When k =   1 , the neighborhood is too small to support stable local phase-coherence estimation and is therefore more sensitive to speckle fluctuations. In contrast, larger windows such as k =   5 and k =   7 tend to mix target and background responses, which weakens the local structural details of weak targets. Therefore, k =   3 is adopted as the default setting.
We further analyze the effect of DTCWT decomposition depth by comparing L = { 1 ,   2 ,   3 ,   4 } on the three datasets. As shown in Figure 7, L =   2 yields the best overall performance across all three datasets. A shallow decomposition such as L =   1 is insufficient to capture multi-scale structural responses, whereas deeper decompositions such as L =   3 and L =   4 introduce additional redundancy and may dilute weak local target details. Therefore, a two-level DTCWT decomposition is used as the default setting.

4.3.4. Complexity Analysis

To provide a clearer analysis of the computational overhead introduced by the proposed modules, we compare the parameter counts and FLOPs of four variants: the baseline detector, baseline + PaSCA, baseline + SCFRDeno, and the full MSPaDet. As shown in Figure 8, the baseline detector contains 28.78 M parameters and 55.07 G FLOPs. After adding PaSCA, the model reaches 33.59 M parameters and 59.82 G FLOPs, corresponding to increases of 16.7% and 8.6%, respectively. After adding SCFRDeno, the model reaches 44.06 M parameters and 61.85 G FLOPs, corresponding to increases of 53.1% and 12.3%, respectively. The full MSPaDet contains 49.50 M parameters and 64.60 G FLOPs, corresponding to overall increases of 72.0% and 17.3% over the baseline detector. These results show that PaSCA is relatively compact, while SCFRDeno introduces a larger parameter overhead but keeps the FLOP increase moderate. Overall, MSPaDet achieves improved detection performance with moderate computational overhead.

4.4. Visualization of Detection Results

To further support the above findings, we visualize feature responses with Grad-CAM. On MSAR, the baseline FCOS exhibits pronounced high-frequency speckle artifacts in FPN-level features, whereas MSPaDet produces substantially cleaner and more target-focused activations. On SAR-Aircraft-1.0, MSPaDet better separates targets from cluttered backgrounds and suppresses spurious responses around highly reflective structures. Overall, the visual evidence is consistent with the quantitative gains, indicating that SCFRDeno enhances structural cues while reducing speckle-dominant interference in the feature space.

4.4.1. Small-Target Analysis and Feature Visualization

To further explain the small-target improvement, we visualize representative tiny targets smaller than 5 × 5 pixels from MSAR, SAR-Aircraft-1.0, and SARDet-100K. As shown in Figure 9, MSPaDet produces more concentrated feature responses than MSDFEN around tiny targets and better suppresses background speckle and clutter, indicating improved target–background separability. This benefit mainly comes from the multi-scale and multi-directional representation of DTCWT, the suppression of random speckle by phase-coherence modeling, and the subsequent feature refinement that enhances weak target responses and boundary cues.

4.4.2. False Positive Cases

As shown in Figure 10, MSPaDet reduces false alarms compared with the FCOS baseline. In MSAR, FCOS is prone to triggering on cluttered high-energy background regions (e.g., ports and coastlines), whereas MSPaDet suppresses these responses after frequency-domain reconstruction (Figure 10a). On SAR-Aircraft-1.0, MSPaDet better separates structural high-frequency cues from random noise, leading to fewer spurious detections (Figure 10b).

4.4.3. Miss Detection Cases

Figure 11 shows that MSPaDet alleviates missed detections for small or low-reflectivity targets. On MSAR, SCFRDeno enhances weak yet structurally consistent signals via sub-band reweighting, producing more concentrated activations on micro targets and improving recall. On SAR-Aircraft-1.0, the embedded PaSCA refinement further strengthens salient responses in cluttered scenes, helping maintain high recall.

4.4.4. Misclassification

As illustrated in Figure 12, MSPaDet yields more class-discriminative Grad-CAM responses in complex scenes. The combination of frequency-domain enhancement and spatial refinement produces more target-focused activations and suppresses background-induced ambiguity, which helps reduce misclassification.

5. Discussion

This work addresses the coexistence of weak target responses and speckle in high-frequency SAR features. MSPaDet distinguishes them by exploiting directional phase coherence, preserving stable scattering cues while suppressing unstable high-frequency interference. Experiments on MSAR, SAR-Aircraft-1.0, and SARDet-100K show consistent improvements, especially on SAR-Aircraft-1.0, where aircraft targets often appear as sparse and incomplete scattering structures. Ablation studies confirm that energy and phase provide complementary cues: energy highlights strong scatterers, while phase coherence captures stable directional structures.
Compared with Lee filtering, NLM, and SAR-BM3D, MSPaDet better matches the detection objective by performing denoising in feature space rather than smoothing the input image. It also complements existing SAR detectors by adding an explicit scattering-coherence constraint.
However, feature denoising alone cannot fully address cross-dataset domain shifts. In addition, the phase information comes from DTCWT feature coefficients rather than raw complex SAR data, and SCFRDeno increases model parameters. Future work will explore raw complex phase information, phase-aware adaptation, and lightweight sub-band selection.

6. Conclusions

We present MSPaDet for SAR target detection, which leverages complex wavelet sub-bands and phase-coherence modulation for speckle-robust feature enhancement. MSPaDet achieves consistent state-of-the-art performance on MSAR, SAR-Aircraft-1.0, and SARDet-100K, with a 1.0% mAP improvement in MSAR and a 0.57% gain over SARDet-100K.

Author Contributions

N.C.: Conducted the experiments, processed and analyzed the data, implemented the software, and drafted the manuscript. X.X.: Conceived and oversaw the study; provided expert guidance throughout the research. Y.L.: Wrote and revised the manuscript. All authors have read and agreed to the published version of the manuscript.

Funding

Science Research Projects of Hunan Provincial Education Department (Grant No. 25B0310).

Data Availability Statement

The data analyzed in this study are publicly available from the MSAR dataset (https://radars.ac.cn/web/data/getData?dataType=MSAR (accessed on 20 September 2025), the SAR-Aircraft-1.0 dataset (https://radars.ac.cn/web/data/getData?newsColumnId=f896637b-af23-4209-8bcc-9320fceaba19, accessed on 20 September 2025), and the SARDet-100K dataset (https://github.com/zcablii/SARDet_100K, accessed on 20 September 2025). The code used in this study is available from the corresponding author upon reasonable request.

Acknowledgments

Acknowledgments:This work was supported in part by the Natural Science Foundation of Hunan Province under Grant 2025JJ50397, and by the Science Research Projects of Hunan Provincial Education Department under Grant No. 25B0310. We gratefully acknowledge the financial support from both funding agencies. We also extend our sincere thanks to the following individuals for their specific contributions:Xuyu Xiang: Conceived and oversaw the study; provided expert guidance throughout the research.Yuanjing Luo: Wrote and revised the manuscript.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Tirandaz, Z.; Akbarizadeh, G.; Kaabi, H. PolSAR Image Segmentation Based on Feature Extraction and Data Compression Using Weighted Neighborhood Filter Bank and Hidden Markov Random Field–Expectation Maximization. Measurement 2020, 153, 107432. [Google Scholar] [CrossRef]
  2. Quan, S.; Zhang, T.; Xing, S.; Wang, X.; Yu, Q. Maritime Ship Detection with Concise Polarimetric Characterization Pattern. Int. J. Appl. Earth Obs. Geoinf. 2024, 131, 103954. [Google Scholar] [CrossRef]
  3. Chen, S.; Cui, X.; Wang, X.; Xiao, S. Speckle-Free SAR Image Ship Detection. IEEE Trans. Image Process. 2021, 30, 5969–5983. [Google Scholar] [CrossRef]
  4. Yasir, M.; Liu, S.; Xu, M.; Wan, J.; Sheng, H.; Nazir, S.; Zhang, X.; Isiacik Colak, A.T. YOLOv8-BYTE: Ship Tracking Algorithm Using Short-Time Sequence SAR Images for Disaster Response Leveraging GeoAI. Int. J. Appl. Earth Obs. Geoinf. 2024, 128, 103771. [Google Scholar] [CrossRef]
  5. Karwowska, K.; Slesinski, J.; Wierzbicki, D. Effectiveness of YOLO Variants for Small Object Detection in SAR Images Using a New Dataset. Sci. Rep. 2025, 15, 45405. [Google Scholar] [CrossRef]
  6. Liu, W.; Qin, J. Focus and Learn: Boosting Deep Multi-View Clustering via Hard Instance Awareness. Inf. Fusion 2026, 127, 103724. [Google Scholar] [CrossRef]
  7. Wu, B.; Liu, C.; Chen, J. A Review of Spaceborne High-Resolution Spotlight/Sliding Spotlight Mode SAR Imaging. Remote Sens. 2025, 17, 38. [Google Scholar] [CrossRef]
  8. Wang, Y.; Li, J.; Yang, J.; Sun, B. A Novel Spaceborne Sliding Spotlight Range Sweep Synthetic Aperture Radar: System and Imaging. Remote Sens. 2017, 9, 783. [Google Scholar] [CrossRef]
  9. Zhang, Z.; Yu, W.; Zheng, M.; Zhao, L.; Zhou, Z.X. Phase Mismatch Calibration for Dual-Channel Sliding Spotlight SAR-GMTI. Remote Sens. 2022, 14, 617. [Google Scholar] [CrossRef]
  10. Liu, T.; Yang, Z.; Marino, A.; Gao, G.; Yang, J. Robust CFAR Detector Based on Truncated Statistics for Polarimetric Synthetic Aperture Radar. IEEE Trans. Geosci. Remote Sens. 2020, 58, 6731–6747. [Google Scholar] [CrossRef]
  11. Zalpour, M.; Akbarizadeh, G.; AlaeiSheini, N. A New Approach for Oil Tank Detection Using Deep Learning Features with Control False Alarm Rate in High-Resolution Satellite Imagery. Int. J. Remote Sens. 2020, 41, 2239–2262. [Google Scholar] [CrossRef]
  12. Tang, G.; Zhao, H.; Claramunt, C.; Zhu, W.; Wang, S.; Wang, Y.; Ding, Y. PPA-Net: Pyramid Pooling Attention Network for Multi-Scale Ship Detection in SAR Images. IEEE Trans. Pattern Anal. Mach. Intelling 2023, 15, 2855. [Google Scholar] [CrossRef]
  13. Mao, Q.; Li, Y.; Zhu, Y. A Hierarchical Feature Fusion and Attention Network for Automatic Ship Detection from SAR Images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 13981–13994. [Google Scholar] [CrossRef]
  14. Wang, H.; Shi, J.; Karimian, H.; Liu, F.; Wang, F. YOLOSAR-Lite: A Lightweight Framework for Real-Time Ship Detection in SAR Imagery. Int. J. Digit. Earth 2024, 17, 2405525. [Google Scholar] [CrossRef]
  15. Li, Y.; Li, X.; Li, W.; Hou, Q.; Liu, L.; Cheng, M.; Yang, J. SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 10–15 December 2024. [Google Scholar]
  16. Zhang, X.; Yang, X.; Li, Y.; Yang, J.; Cheng, M.; Li, X. RSAR: Restricted State Angle Resolver and Rotated SAR Benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 10–17 June 2025. [Google Scholar]
  17. Du, Y.; Chen, Y.; Huang, L.; Yang, Y.; Ghamisi, P.; Du, Q. SUMMIT: A SAR Foundation Model with Multiple Auxiliary Tasks Enhanced Intrinsic Characteristics. Int. J. Appl. Earth Obs. Geoinf. 2025, 141, 104624. [Google Scholar] [CrossRef]
  18. Ma, W.; Yang, X.; Zhu, H.; Wang, X.; Yi, X.; Wu, Y.; Hou, B.; Jiao, L. NRENet: Neighborhood Removal-and-Emphasis Network for Ship Detection in SAR Images. Int. J. Appl. Earth Obs. Geoinf. 2024, 131, 103927. [Google Scholar] [CrossRef]
  19. Zhang, T.; Zhang, X.; Ke, X.; Zhan, X.; Shi, J.; Wei, S.; Pan, D.; Li, J.; Su, H.; Zhou, Y.; et al. LS-SSDD-v1.0: A Deep Learning Dataset Dedicated to Small Ship Detection from Large-Scale Sentinel-1 SAR Images. Remote Sens. 2020, 12, 2997. [Google Scholar] [CrossRef]
  20. Ke, X.; Zhang, T.; Shao, Z. Scale-aware dimension-wise attention network for small ship instance segmentation in synthetic aperture radar images. J. Appl. Remote Sens. 2023, 17, 046504. [Google Scholar] [CrossRef]
  21. Lee, J.S. Speckle Analysis and Smoothing of Synthetic Aperture Radar Images. Comput. Graph. Image Process. 1981, 17, 24–32. [Google Scholar] [CrossRef]
  22. Parrilli, S.; Poderico, M.; Angelino, C.V.; Verdoliva, L. A Nonlocal SAR Image Denoising Algorithm Based on LLMMSE Wavelet Shrinkage. IEEE Trans. Geosci. Remote Sens. 2012, 50, 606–616. [Google Scholar] [CrossRef]
  23. Argenti, F.; Lapini, A.; Bianchi, T.; Alparone, L. A Tutorial on Speckle Reduction in Synthetic Aperture Radar Images. IEEE Geosci. Remote Sens. Mag. 2013, 1, 6–35. [Google Scholar] [CrossRef]
  24. Li, J.; Yu, Z.; Yu, L.; Cheng, P.; Chen, J.; Chi, C. A Comprehensive Survey on SAR ATR in Deep-Learning Era. Remote Sens. 2023, 15, 1454. [Google Scholar] [CrossRef]
  25. Zhou, Z.; Cui, Z.; Tang, K.; Tian, Y.; Pi, Y.; Cao, Z. Gaussian Meta-Feature Balanced Aggregation for Few-Shot Synthetic Aperture Radar Target Detection. ISPRS J. Photogramm. Remote Sens. 2024, 208, 89–106. [Google Scholar] [CrossRef]
  26. Ying, L.; Miao, D.; Zhang, Z. A Robust One-Stage Detector for SAR Ship Detection with Sequential Three-Way Decisions and Multi-Granularity. Inf. Sci. 2024, 667, 120436. [Google Scholar] [CrossRef]
  27. Shen, Y.F.; Gao, Q. DS-YOLO: A SAR Ship Detection Model for Dense Small Targets. Radioengineering 2025, 34, 407–421. [Google Scholar] [CrossRef]
  28. Xu, X.; Zhang, X.; Zhang, T. Lite-YOLOv5: A Lightweight Deep Learning Detector for On-Board Ship Detection in Large-Scene Sentinel-1 SAR Images. Remote Sens. 2022, 14, 1018. [Google Scholar] [CrossRef]
  29. Yasir, M.; Liu, S.; Pirasteh, S.; Xu, M.; Sheng, H.; Wan, J.; de Figueiredo, F.A.P.; Aguilar, F.J.; Li, J. YOLOShipTracker: Tracking Ships in SAR Images Using Lightweight YOLOv8. Int. J. Appl. Earth Obs. Geoinf. 2024, 134, 104137. [Google Scholar] [CrossRef]
  30. Dong, J.; Feng, J.; Tang, X. OptiSAR-Net: A Cross-Domain Ship Detection Method for Multisource Remote Sensing Data. IEEE Trans. Geosci. Remote Sens. 2024, 62, 4709311. [Google Scholar] [CrossRef]
  31. Xu, W.; Guo, Z.; Huang, P.; Tan, W.; Gao, Z. Towards Efficient SAR Ship Detection: Multi-Level Feature Fusion and Lightweight Network Design. Remote Sens. 2025, 17, 2588. [Google Scholar] [CrossRef]
  32. He, S. Multiscale Task-Decoupled Oriented SAR Ship Detection Based on a Size-Aware Balanced Strategy. Remote Sens. 2025, 17, 2257. [Google Scholar] [CrossRef]
  33. Wu, B.; Liu, C.; Chen, Z.; Zhang, S.; Chen, J. A Novel Star Feature Decoupling Network for Multiscale SAR Image Ship Detection. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2026. [Google Scholar] [CrossRef]
  34. Sun, Z.; Leng, X.; Zhang, X.; Zhou, Z.; Xiong, B.; Ji, K.; Kuang, G. Arbitrary-Direction SAR Ship Detection Method for Multiscale Imbalance. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5208921. [Google Scholar] [CrossRef]
  35. Chen, B.; Xue, F.; Song, H. Lightweight Transformer Detector for SAR Ship Detection. IEEE Trans. Pattern Anal. Mach. Intelling 2024, 16, 237. [Google Scholar] [CrossRef]
  36. Zhang, Y.; Chi, J.; Yang, G.; Chen, C.; Yu, T. SAR-NanoShipNet: A Scale-Adaptive Network for Robust Small Ship Detection in SAR Imagery. ISPRS J. Photogramm. Remote Sens. 2026, 232, 262–279. [Google Scholar] [CrossRef]
  37. Zhao, S.; Luo, Y.; Zhang, T.; Guo, W.; Zhang, Z. A Domain Specific Knowledge Extraction Transformer Method for Multisource Satellite-Borne SAR Images Ship Detection. ISPRS J. Photogramm. Remote Sens. 2023, 198, 16–29. [Google Scholar] [CrossRef]
  38. Yang, Y.; Chen, J.; Sun, L.; Zhou, Z.; Huang, Z.; Wu, B. Unsupervised Domain-Adaptive SAR Ship Detection Based on Cross-Domain Feature Interaction and Data Contribution Balance. Remote Sens. 2024, 16, 420. [Google Scholar] [CrossRef]
  39. Li, D.; Liang, Q.; Liu, H.; Liu, Q.; Liu, H.; Liao, G. A Novel Multidimensional Domain Deep Learning Network for SAR Ship Detection. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5203213. [Google Scholar] [CrossRef]
  40. Ding, X.; Zhang, X.; Han, J.; Ding, G. Scaling up Your Kernels to 31 × 31: Revisiting Large Kernel Design in CNNs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; IEEE: New York, NY, USA, 2022; pp. 11963–11975. [Google Scholar] [CrossRef]
  41. Wang, W.; Dai, J.; Chen, Z.; Huang, Z.; Li, Z.; Zhu, X.; Hu, X.; Lu, T.; Lu, L.; Li, H.; et al. InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; IEEE: New York, NY, USA, 2022; pp. 14408–14419. [Google Scholar]
  42. Zhang, X.; Liu, C.; Yang, D.; Song, T.; Ye, Y.; Li, K.; Song, Y. RFAConv: Innovating Spatial Attention and Standard Convolutional Operation. arXiv 2024, arXiv:2304.03198. [Google Scholar] [CrossRef]
  43. Cui, Y.; Tao, Y.; Bing, Z.; Ren, W.; Gao, X.; Cao, X.; Huang, K.; Knoll, A. Selective Frequency Network for Image Restoration. In Proceedings of the Eleventh International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
  44. Pham, C.; Nguyen, V.; Le, T.; Phung, D.; Carneiro, G.; Do, T. Frequency Attention for Knowledge Distillation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 3–8 January 2024; IEEE: New York, NY, USA, 2024; pp. 2277–2286. [Google Scholar] [CrossRef]
  45. Chen, J.; Duanmu, C.; Long, H. Large Kernel Frequency-Enhanced Network for Efficient Single Image Super-Resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Seattle, WA, USA, 17–18 June 2024; IEEE: New York, NY, USA, 2024; pp. 6317–6326. [Google Scholar] [CrossRef]
  46. Ni, K.; Wang, P.; Zheng, Z.; Zhong, Y. Complex-valued mix transformer for SAR ship detection. ISPRS J. Photogramm. Remote Sens. 2026, 231, 1–16. [Google Scholar] [CrossRef]
  47. Tian, Z.; Wu, Q.; Xu, J.; Zhong, Y. FCOS: Fully Convolutional One-Stage Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; IEEE: New York, NY, USA, 2019; pp. 9627–9636. [Google Scholar] [CrossRef]
  48. Li, X.; Wang, W.; Wu, L.; Chen, S.; Hu, X.; Li, J.; Tang, J.; Yang, J. Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection. In Proceedings of the 34 International Conference on Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 6–12 December 2020. [Google Scholar]
  49. Yang, Z.; Liu, S.; Hu, H.; Wang, L.; Lin, S. RepPoints: Point Set Representation for Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October –2 November 2019; IEEE: New York, NY, USA, 2019; pp. 9657–9666. [Google Scholar] [CrossRef]
  50. Zhou, X.; Wang, D.; Krähenbühl, P. Objects as Points. arXiv 2019, arXiv:1904.07850. [Google Scholar] [CrossRef]
  51. Wang, W.; Chen, K.; Kim, P.; Yoon, K.; Liu, Z.; Joo, K.; Liu, Y.-H. Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction Without Convolution. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 568–578. [Google Scholar] [CrossRef]
  52. Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; IEEE: New York, NY, USA, 2017; pp. 2980–2988. [Google Scholar]
  53. Feng, C.; Zhong, Y.; Gao, Y.; Scott, M.R.; Huang, W. TOOD: Task-Aligned One-Stage Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 3490–3499. [Google Scholar] [CrossRef]
  54. Chen, Z.; Yang, C.; Li, Q.; Zhao, F.; Zha, Z.; Wu, F. Dicentang: Your Dense Object Detector. In Proceedings of the 29th ACM International Conference on Multimedia (ACM MM), Virtual Event, China, 20–24 October 2021; Association for Computing Machinery: New York, NY, USA, 2021; pp. 4939–4948. [Google Scholar]
  55. Chen, Q.; Wang, Y.; Yang, T.; Zhang, X.; Cheng, J.; Sun, J. You Only Look One-Level Feature. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; IEEE: New York, NY, USA, 2021; pp. 13039–13048. [Google Scholar] [CrossRef]
  56. Ge, Z.; Liu, S.; Wang, F.; Li, Z.; Sun, J. YOLOX: Exceeding YOLO Series in 2021. arXiv 2021, arXiv:2107.08430. [Google Scholar] [CrossRef]
  57. Zhang, M.; Zhu, Y.; Li, L.; Guo, J.; Liu, Z.; Li, Y. S4Det: Breadth and Accurate Sine Single-Stage Ship Detection for Remote Sense SAR Imagery. Remote Sens. 2025, 17, 900. [Google Scholar] [CrossRef]
  58. Liu, Z.; Mao, H.; Wu, C.; Feichtenhofer, C.; Darrell, T.; Xie, S. A ConvNet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; IEEE: New York, NY, USA, 2022; pp. 11976–11986. [Google Scholar]
  59. Woo, S.; Debnath, S.; Hu, R.; Chen, X.; Liu, Z.; Kweon, I.S.; Xie, S. ConvNeXt V2: Co-Designing and Scaling ConvNets with Masked Autoencoders. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; IEEE: New York, NY, USA, 2023; pp. 16133–16142. [Google Scholar]
  60. Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-End Object Detection with Transformers. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; Springer: Cham, Switzerland, 2020; pp. 213–229. [Google Scholar] [CrossRef]
  61. Zhang, M.; Li, Y.; Guo, J.; Li, Y.; Gao, X. BurgsVO: Burgs-Associated Vertex Offset Encoding Scheme for Detecting Rotated Ships in SAR Images. Remote Sens. 2025, 17, 388. [Google Scholar] [CrossRef]
  62. Liu, S.; Li, F.; Zhang, H.; Yang, X.; Qi, X.; Su, H.; Zhu, J.; Zhang, L. DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR. In Proceedings of the International Conference on Learning Representations, Online, 25–29 April 2022. [Google Scholar]
  63. Meng, D.; Chen, X.; Fan, Z.; Zeng, G.; Li, H.; Yuan, Y.; Sun, L.; Wang, J. Conditional DETR for Fast Training Convergence. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; IEEE: New York, NY, USA, 2021; pp. 3651–3660. [Google Scholar] [CrossRef]
  64. Zhou, J.; Xiao, C.; Peng, B.; Liu, Z.; Liu, L.; Liu, Y.; Li, X. DiffDet4SAR: Diffusion-Based Aircraft Target Detection Network for SAR Images. IEEE Geosci. Remote Sens. Lett. 2024, 21, 4007905. [Google Scholar] [CrossRef]
  65. Wang, T.; Zeng, Z.; Zhou, S.; Xu, Q. A Multi-Scale Discrete Feature Enhancement Network with Augmented Reversible Transformation for SAR Automatic Target Recognition. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 5135–5156. [Google Scholar] [CrossRef]
  66. Dai, Y.; Zou, M.; Li, Y.; Li, X.; Ni, K.; Yang, J. DenoDet: Attention as Deformable Multi-Subspace Feature Denoising for Target Detection in SAR Images. arXiv 2024, arXiv:2406.02833. [Google Scholar] [CrossRef]
  67. Zhang, M.; Yang, Z.; Guo, J.; Li, Y. PDE-Guided Diverse Feature Learning for SAR Rotated Ship Detection. Remote Sens. 2025, 17, 2998. [Google Scholar] [CrossRef]
  68. Zhang, S.; Chi, C.; Yao, Y.; Lei, Z.; Li, S.Z. Bridging the Gap Between Anchor-Based and Anchor-Free Detection via Adaptive Training Sample Selection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13-19 June 2020; IEEE: New York, NY, USA, 2020; pp. 9759–9768. [Google Scholar] [CrossRef]
Figure 1. Representative images for remote SAR target detection.
Figure 1. Representative images for remote SAR target detection.
Remotesensing 18 01513 g001
Figure 2. Existing frequency attention module.
Figure 2. Existing frequency attention module.
Remotesensing 18 01513 g002
Figure 3. Simplified flowchart of the MSPaDet framework.
Figure 3. Simplified flowchart of the MSPaDet framework.
Remotesensing 18 01513 g003
Figure 4. Diagram of the SCFRDeno and PaSCA modules.
Figure 4. Diagram of the SCFRDeno and PaSCA modules.
Remotesensing 18 01513 g004
Figure 5. Comparison with traditional despeckling methods embedded in the same detection pipeline on SAR-Aircraft-1.0.
Figure 5. Comparison with traditional despeckling methods embedded in the same detection pipeline on SAR-Aircraft-1.0.
Remotesensing 18 01513 g005
Figure 6. Effect of phase-coherence window size k on detection performance.
Figure 6. Effect of phase-coherence window size k on detection performance.
Remotesensing 18 01513 g006
Figure 7. Effect of DTCWT decomposition depth on detection performance.
Figure 7. Effect of DTCWT decomposition depth on detection performance.
Remotesensing 18 01513 g007
Figure 8. Parameter and FLOPs comparison of MSPaDet variants relative to the baseline detector.
Figure 8. Parameter and FLOPs comparison of MSPaDet variants relative to the baseline detector.
Remotesensing 18 01513 g008
Figure 9. Qualitative comparison and feature-map visualization for tiny targets within 5 × 5 pixels across MSAR, SAR-Aircraft-1.0, and SARDet-100K.
Figure 9. Qualitative comparison and feature-map visualization for tiny targets within 5 × 5 pixels across MSAR, SAR-Aircraft-1.0, and SARDet-100K.
Remotesensing 18 01513 g009
Figure 10. Comparison of feature-map denoising produced by different detection methods.
Figure 10. Comparison of feature-map denoising produced by different detection methods.
Remotesensing 18 01513 g010
Figure 11. Misdetection examples for SAR targets; each subfigure shows one case.
Figure 11. Misdetection examples for SAR targets; each subfigure shows one case.
Remotesensing 18 01513 g011
Figure 12. Misclassification examples for SAR targets; each subfigure shows one case.
Figure 12. Misclassification examples for SAR targets; each subfigure shows one case.
Remotesensing 18 01513 g012
Table 1. Satellite and sensor details of the used datasets.
Table 1. Satellite and sensor details of the used datasets.
DatasetResolutionBandPolarizationSatellites/Sensors
MSAR≤1 mCHH/HV/VH/VVHISEA-1
SAR-Aircraft-1.01 mCSingleGF-3
SARDet-100K0.1–25 mC/Ka/Ku/XHH/HV/VH/VV/SingleAirborne SAR; GF-3;
HISEA-1; RadarSat-2;
S-1; TerraSAR-X;
PALSAR-2
Table 2. Comparison with state-of-the-art methods on the MSAR dataset.
Table 2. Comparison with state-of-the-art methods on the MSAR dataset.
MethodPre.#P↓FLOPs↓AP’07AP’12
mAPShipAirBrgOilmAPShipAirBrgOil
One-stage
FCOS [47]IN32.12 M12.89 G66.2287.8742.4272.6261.9868.5588.5643.3375.3466.97
GFL [48]IN32.27 M13.08 G65.9487.9140.8672.7562.2467.6188.9639.8475.6765.98
RepPoints [49]IN36.82 M12.12 G48.0176.687.7949.8357.5146.9280.761.8849.9457.84
CenterNet [50]IN32.12 M12.88 G65.8888.5739.7573.6161.6267.9589.8039.9475.7566.31
PVT-T [51]IN21.39 M10.10 G28.4257.588.8410.2337.0424.9558.340.225.4735.75
RetinaNet [52]IN36.39 M13.14 G32.4162.731.4222.1745.5432.0263.671.6619.7845.27
TOOD [53]IN32.03 M12.62 G65.2887.7441.4869.7062.1967.0288.5840.9472.6165.94
DDOD [54]IN32.27 M29.02 G59.5085.2961.8670.1220.7461.0187.4565.3973.4917.73
YOLOF [55]IN42.41 M6.58 G45.0968.097.7957.5146.9543.2970.421.0158.7744.95
YOLOX [56]IN8.94 M2.13 G67.1088.5962.1960.1657.4469.1289.3165.7663.2358.19
S4 Det [57]IN32.72 M12.09 G64.0186.6437.5069.7362.1965.2687.9135.4471.8265.86
Two-stage
ConvNeXt [58]IN45.06 M26.39 G55.2280.0462.2770.777.7956.1083.9965.7975.820.10
ConvNeXtV2 [59]IN105.00 M40.71 G55.0880.0662.2970.187.7956.0084.1665.5573.560.70
LSKNet [60]IN30.98 M22.52 G56.8380.0162.3177.217.7957.4485.5665.7577.421.04
BurgsVO [61]IN63.10 M45.62 G66.9788.6640.6970.9867.5368.5389.0741.9776.8666.22
End2End
DAB-DETR [62]IN43.70 M10.49 G42.5878.952.1750.9038.3142.8782.900.2851.0637.26
Conditional DETR [63]IN43.35 M9.79 G40.9677.338.5052.8325.1738.7181.561.3053.0420.34
DiffDet4 SAR [64]IN53.23 M12.89 G68.3787.2146.1476.0161.2668.5288.1244.2777.3066.51
MSDFEN [65]IN40.23 M13.41 G68.3887.4345.2377.0363.8168.7388.8445.3778.5066.96
DenoDet [66]IN34.23 M12.89 G68.6088.1046.5777.6262.1169.9189.4444.7778.3067.11
MSPaDet (Ours)IN27.25 M7.75 G69.5389.3548.676.9263.2570.2590.0549.377.3264.45
Table 3. Comparison with state-of-the-art methods on the SAR-Aircraft-1.0 dataset.
Table 3. Comparison with state-of-the-art methods on the SAR-Aircraft-1.0 dataset.
MethodPre.FLOPsAverage Precision (AP’07)Average Precision (AP’12)
mAPA220A320A330ARJ21B737B787OthermAPA220A320A330ARJ21B737B787Other
One-stage
FCOS [47]IN51.58 G57.1356.7787.6958.1557.6034.5646.8758.2958.2659.1389.6659.1157.5233.8647.5561.00
GFL [48]IN32.27 G61.4054.4081.8487.3857.0340.6656.3252.1462.9455.6985.1189.2560.6040.4856.6152.88
RepPoints [49]IN48.50 G61.6358.7784.2580.7855.7836.0054.2361.6062.5960.2287.5381.4157.3834.5154.3062.79
TOOD [53]IN50.53 G57.1648.4476.7086.3655.0129.8648.3155.4557.4848.3178.4188.4955.0128.4148.0055.71
DDOD [54]IN45.60 G57.1647.5281.8490.6254.4529.8139.0456.8957.6046.9485.9590.8155.9029.3336.9057.37
YOLOF [55]IN26.33 G60.7554.7485.2583.3155.5030.8054.1461.4962.2155.4789.3585.8157.8728.8554.9263.21
YOLOX [56]IN8.53 G58.1558.4384.1192.0446.5118.6951.1356.1759.5559.4185.7792.3850.0117.3855.1556.76
S4 Det [57]IN48.39 G64.6759.4489.4189.7559.7941.9150.9657.4365.1562.5090.0790.6459.2639.1355.2658.93
Two-stage
PRDet [67]IN51.00 G65.3752.3790.8292.0163.1542.7758.8957.5966.3553.2691.0892.2366.1643.7159.6058.38
ConvNeXt [58]IN63.85 G61.9157.9486.7592.0363.8533.0843.5656.1461.4858.0188.2192.1662.1130.6743.5655.61
ConvNeXt V2 [59]IN0.12 T62.5455.6388.2791.6654.2330.7856.8860.3263.2257.2189.3891.8054.4328.7457.6963.26
LSKNet [60]IN53.73 G62.0853.6593.5691.3353.3531.8451.3459.4862.7654.3293.7191.7155.3030.0651.9762.29
End2End
DAB-DETR [62]IN28.94 G48.1254.3282.8913.3456.7629.0248.8451.6548.6655.3085.6013.4758.1127.7048.5451.92
Conditional DETR [63]IN28.09 G56.7548.2483.8873.3356.2531.9455.9547.6957.5248.7787.8074.9756.5730.4256.4947.64
DiffDet4 SAR [64]IN50.49 G67.4663.2291.1490.8756.0845.3260.2465.3569.0666.0391.9392.3457.3047.3561.2467.21
MSDFEN [65]IN73.27 G67.2264.3288.7690.6555.2144.8260.8465.9769.3466.0392.9392.3457.3046.3562.2468.21
DenoDet [66]IN48.53 G63.1059.3285.0488.3254.0840.8250.5463.5564.0661.0387.9389.3455.3039.3550.2465.21
MSPaDet (Ours)IN68.38 G68.6565.7391.4992.0456.3445.4062.3367.2269.7868.0192.5492.3457.2946.5163.5868.19
Table 4. Comparison with state-of-the-art methods on the SARDet-100K dataset.
Table 4. Comparison with state-of-the-art methods on the SARDet-100K dataset.
MethodPre.FLOPs#Params↓mAPAP@50AP@75APsAPmAPlShipAircraftCarTankBridgeHarbor
One-stage
FCOS [47]IN51.57 G32.13 M52.5285.8254.9347.0166.1357.8259.7955.4460.7541.7834.1763.44
GFL [48]IN52.36 G32.27 M54.7184.8658.5749.1466.9960.1563.6257.3361.9944.5036.1164.74
RepPoints [49]IN48.49 G36.82 M51.3686.1353.6946.3662.9653.4860.5555.2060.8340.3934.8256.41
ATSS [68]IN51.57 G32.13 M54.6587.3057.9549.5967.6458.6761.2355.6461.4745.9036.9267.18
PVT-T [51]IN42.19 G21.43 M45.8077.2548.7037.7159.2353.0553.0052.6158.7329.9022.2158.81
TOOD [53]IN50.52 G30.03 M54.3586.5858.1149.9066.4258.3061.9855.3162.2345.6636.3464.94
DDOD [54]IN45.58 G32.21 M53.7286.3456.9349.0364.4057.7262.0955.7862.1843.6836.0462.57
YOLOF [55]IN26.32 G42.46 M42.5374.6542.8833.4355.8953.2752.3252.3452.4122.5623.4452.12
YOLOX [56]IN8.53 G8.94 M33.7866.4731.0128.1942.7628.6545.7846.5353.1325.9612.8418.65
S4 Det [57]IN48.38 G32.72 M55.7187.0259.0250.0768.0960.6964.8458.5464.6744.7836.8164.98
Two-stage
ConvNeXt [58]IN63.84 G45.07 M52.8585.2256.9845.3764.2558.3160.2557.0561.8337.8236.5163.65
ConvNeXtV2 [59]IN0.12 T0.11 G53.6185.7158.6047.3364.3759.2761.1855.5362.9339.3538.8663.79
BurgsVO [61]IN63.97 G45.20 M53.9885.3557.9847.6366.3860.4461.3858.3162.8338.8237.9064.65
End2End
DAB-DETR [62]IN28.94 G43.70 M43.0177.8442.8034.5256.0452.3252.8650.0249.1723.7628.1754.77
DiffDet4 SAR [64]IN69.82 G54.74 M55.7985.5159.4650.3367.4060.3064.8157.0663.7945.5636.2266.94
MSDFEN [65]IN82.73 G75.78 M56.1286.2159.8750.5168.3861.2265.2158.3664.1646.2337.7365.07
DenoDet [66]IN52.69 G65.78 M55.3285.5159.2550.3367.4060.3064.6157.0663.3645.4936.0966.87
MSPaDet (Ours)IN64.60 G49.50 M56.6986.3261.1750.8068.3062.5065.9458.3364.3245.1337.8168.50
Table 5. In-domain and cross-dataset aircraft-level evaluation results.
Table 5. In-domain and cross-dataset aircraft-level evaluation results.
CategoryMetricScore
In-domain reference
SAR-Aircraft-1.0mAP0.6865
SARDet-100KmAP0.5833
Cross-dataset transfer
SAR-Aircraft-1.0 → SARDet-100K (aircraft subset)AP0.4101
SAR-Aircraft-1.0 → SARDet-100K (aircraft subset)Recall0.6100
Table 6. Ablation on frequency-domain decomposition.
Table 6. Ablation on frequency-domain decomposition.
DTCWT/Norm.Phase
Reweight
FPMSARSAR-Aircraft-1.0SARDet-100K
FPSmAP(07)mAP(12)FPSmAP(07)mAP(12)FPSmAP
××32230.465.5266.85167.661.6362.76101.353.21
×32230.166.8867.09167.463.0163.40102.454.91
×32225.066.6468.26167.359.7260.49101.453.81
×64220.366.7168.57152.360.5861.05101.354.32
32215.068.5270.31166.467.6468.61101.955.68
Table 7. Ablation on dual-branch modeling.
Table 7. Ablation on dual-branch modeling.
DTCWT/Norm.PaSCAMSARSAR-Aircraft-1.0SARDet-100K
mAP(07)mAP(12)mAP(07)mAP(12)mAP
××66.5267.8561.6362.7651.66
×67.8668.8164.2064.8454.30
69.5370.2568.6569.7856.69
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Chen, N.; Xiang, X.; Luo, Y. MSPaDet: A Multi-Scale Phase-Aware Denoising Method for Target Detection in SAR Images. Remote Sens. 2026, 18, 1513. https://doi.org/10.3390/rs18101513

AMA Style

Chen N, Xiang X, Luo Y. MSPaDet: A Multi-Scale Phase-Aware Denoising Method for Target Detection in SAR Images. Remote Sensing. 2026; 18(10):1513. https://doi.org/10.3390/rs18101513

Chicago/Turabian Style

Chen, Naxiong, Xuyu Xiang, and Yuanjing Luo. 2026. "MSPaDet: A Multi-Scale Phase-Aware Denoising Method for Target Detection in SAR Images" Remote Sensing 18, no. 10: 1513. https://doi.org/10.3390/rs18101513

APA Style

Chen, N., Xiang, X., & Luo, Y. (2026). MSPaDet: A Multi-Scale Phase-Aware Denoising Method for Target Detection in SAR Images. Remote Sensing, 18(10), 1513. https://doi.org/10.3390/rs18101513

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop