2.1. Study Domain and Dataset Construction
This study focuses on a typical deep-sea area in the Northwest Pacific Ocean. Within this specific physical marine region, the spatial evolution of seawater sound speed with depth forms a typical deep-sea sound channel structure, with the channel axis depth typically distributed at approximately 1000 m. To comprehensively characterize the acoustic propagation characteristics of this area and establish a robust recognition model, systematic definitions of the key physical variables in the experiments are provided in this paper.
To simulate the acoustic radiation characteristics of a typical deep-sea submerged operation platform, the source depth is strictly positioned at 1000 m, which precisely aligns with the deep sound channel axis. Crucially, the source frequency is set to 50 Hz to capture the low attenuation characteristics of low-frequency sound waves during long-range propagation. It is worth noting that while conventional ray-geometric methods may face physical validity challenges at low frequencies (below 100 Hz) due to diffraction effects [
8,
27], the use of the Gaussian beam acoustic propagation code (Bellhop) at 50 Hz in this specific deep-sea study remains theoretically justified and methodologically appropriate. First, in our deep-sea environment, with a maximum vertical depth of 5000 m, the water depth-to-acoustic wavelength ratio remains sufficiently large (
), preserving the macro-scale geometric trajectory validity of the ray paths. Second, unlike rigid geometric rays, Bellhop’s geometric hat beams possess a finite beam width determined by the cross-section of ray tubes, thereby smoothing out unphysical singularities at caustics and extending the method’s low-frequency-applicability threshold. Most importantly, the downstream objective of our deep learning framework is image-based semantic segmentation and macro-pattern classification (identifying the general existence and spatial span of CZs), rather than requiring point-to-point coherent phase and absolute amplitude matching. Thus, Bellhop strikes an optimal operational balance, providing the necessary macroscopic physical realism while maintaining the computational efficiency required to generate a massive training dataset.
In terms of spatial scale, the range of acoustic propagation calculations in this paper is limited to a vertical profile with a horizontal distance of 0–100 km and a depth range of 0–5000 m. This spatial envelope is sufficient to fully cover the evolution process of at least one complete period of the CZ structure, thereby ensuring that the model can capture the periodic physical laws of acoustic field morphology.
The construction of the dataset in this study follows the principles of “reliable physical sources and multi-source data coupling”. To ensure the physical authenticity of the simulation calculations, the WOA18 global gridded reanalysis dataset released by the National Centers for Environmental Information (NCEI) was selected as the original input for the oceanic temperature and salinity fields [
28,
29]. This dataset integrates global in situ observational data collected up to 2018; specifically, the temperature and salinity data of the Northwest Pacific Ocean (120° E–160° E, 0° N–50° N), with a spatial resolution of 0.25° and a temporal resolution of quarterly, were adopted in this study. In terms of depth scale, WOA18 employs an unequal-interval layered structure (102 layers within the depth range of 0–5500 m), which can finely characterize the variations in sound-speed gradient near the thermocline and sound channel axis. By screening deep-sea points with seawater depth no less than 5000 m within this geographic range, a total of 13,747 sets of high-quality temperature and salinity data were obtained. Based on these data, 5000 sound TL images with high environmental diversity were generated using the Bellhop ray tracing model.
In the process of characterizing deep-sea sound field features, TL images can intuitively represent the propagation laws of the sound field and the loss regions of sound rays. Therefore, this study employs the Bellhop model to calculate the acoustic TL. As an efficient sound propagation model in underwater acoustics based on ray tracing and geometric beam approximation, the Bellhop model is widely used to compute TL and multipath structures in marine environments. When determining ray coordinates, Bellhop obtains the amplitude and sound pressure by solving the ray equations. For a general cylindrically symmetric system, the ray equations are expressed as in Equation (
1):
where
and
represent the ray coordinates in a cylindrically symmetric system,
s denotes the arc length along the ray, and
is the sound-speed distribution. This equation is solved via numerical integration to obtain the curvature variation of the ray trajectory, thereby reflecting the influence of the sound-speed gradient on the propagation path.
The calculation of TL is based on the superposition of the sound-pressure field. Bellhop calculates the sound pressure
via the superposition of geometric beams and converts it into TL, as defined in Equation (
2):
where
is the sound-pressure intensity at a typical reference distance. For incoherent TL, the model superimposes the energy of each sound ray and ignores the phase information, and the generated images are more adaptable to complex multipath scenarios. In addition, the output TL images can be processed into ray trajectory plots and TL contour maps after point cloud or digitalization processing, providing multidimensional visualization support for deep-sea sound field feature analysis.
The environment file of the Bellhop model defines the physical boundary conditions and medium parameters of sound propagation. The sound-source frequency is set to 50 Hz and the water medium is defined as a single layer, where nmedia equals 1. The SSP adopts the SVW option for piecewise linear interpolation covering from the sea surface to a depth of 5000 m. The minimum sound speed occurs at the mid-deep layer around 1200 m, forming the deep-sea sound channel axis. The seabed is modeled as a sandy sediment layer with a density of
and a P wave velocity of
. While acoustic propagation under sub-critical depths is inherently sensitive to these sediment variations, which can contribute to apparent CZs via bottom-bounce paths [
8], the resulting morphological uncertainties are mathematically accommodated by the hybrid data augmentation strategy in the subsequent visual feature space. The sound source is positioned at a depth of 1000 m while the receiver grid spans the full water depth of 0 to 5000 m over a horizontal range of 0 to 100 km. The coherent TL calculation mode, denoted as CG (i.e., coherent transmission loss with geometric hat beams), is adopted, with emission angles ranging from
to
. The ray-tracing step size is set to 50 m and the computational domain boundaries are defined by a depth box of 5500 m and a range box of 100 km to guarantee accurate calculation of ray curvature, which is particularly critical for the periodic repetition of CZs in long-range deep-sea propagation scenarios.
It must be explicitly clarified that the Munk profile is utilized in this study solely as an analytic canonical benchmark to illustrate the theoretical formation of convergence zones, rather than as a “ground truth” representing any specific real-world ocean environment.
Figure 1a displays the sound TL image for the ideal Munk profile, which was generated using the aforementioned parameters, where the horizontal axis represents the propagation distance from 0 to 100 km, the vertical axis represents water depth from 0 to 5000 m, and the color bar indicates the sound energy loss intensity. The light-colored regions with TL less than 80 decibels in
Figure 1a correspond to acoustic CZs, which appear as periodic high-energy stripes in the 20 km to 60 km range. The blue regions with TL greater than 90 decibels represent acoustic SZs, which are formed by the combined effects of seabed reflection loss and sound ray refraction.
Figure 1b shows the TL image generated using the WOA18 oceanographic climatological data. In this specific image, the CZs are mainly concentrated in a small area near a horizontal distance of 40 km and a depth of 1000 m. This positional difference primarily arises from the depth deviation of the sound channel axis between the WOA18-derived SSP and the ideal Munk model. Real oceanographic environments, such as those derived from WOA18 profiles or in situ measurements, deviate substantially from the canonical analytic Munk form, leading to highly complex and variable sound field morphologies. Furthermore, this shift is exacerbated by discrepancies between the in situ and theoretical values of the seabed sediment absorption coefficient.
To achieve accurate evaluation of the detection efficiency of oceanic sound fields, this study categorized the sound field images into four types of labels with clear acoustic physical significance, according to the geometric morphology and intensity characteristics of energy distribution in the TL images, as visually illustrated in
Figure 2. The classification criteria not only referred to the attenuation degree of acoustic energy but also incorporated the gain requirements for target detection in sonar engineering [
30,
31]. The detailed quantitative definitions and classification criteria for each type are presented in
Table 1. To eliminate subjective ambiguity in overlapping intervals, the classification is strictly governed by the area proportion and the physical spatial continuity of the acoustic energy (i.e., contiguous refractive caustics versus fragmented multipath bounces). Because the dataset was uniformly annotated by the authors as the definitive ground truth based on these morphological guidelines, rather than by independent blind raters, multi-rater reliability statistics (e.g., Cohen’s
) are inapplicable. Furthermore, while continuous regression models provide high-precision predictions for specific acoustic parameters [
2], our discrete 4-class scheme is explicitly optimized for naval tactical decision making and acts as a robust statistical regularizer to prevent overfitting under data-scarce constraints.
To evaluate the framework’s cross-domain generalization capability under realistically data-scarce conditions, an additional 500 simulated TL images were generated using WOA18 oceanographic data from the South China Sea region (105° E–125° E, 0° N–25° N). These images, derived from a geographically distinct maritime environment, serve as a controlled proxy for the practical scenario in which extensive in situ acoustic measurements from a new operational area are unavailable—a pervasive challenge in underwater acoustics that motivates the entire transfer learning design. The South China Sea dataset is deliberately restricted to 500 samples to emulate the few-shot learning regime. To mitigate the risk of model overfitting under this limited sample size, data augmentation, including mirror flipping, random rotation, and Gaussian noise injection, was applied before the fine-tuning phase. This “source domain–target domain” construction scheme—where the abundant Northwest Pacific data (5000 images) constitute the source and the scarce South China Sea data (500 images) constitute the target—provides a principled and reproducible experimental foundation for systematically evaluating cross-domain transfer learning under controlled yet realistic data constraints.
The complete dataset is partitioned as follows. For the Northwest Pacific dataset (5000 images), a stratified 70%/15%/15% split is adopted, yielding 3500 training, 750 validation, and 750 test samples. This split supports the supervised training and evaluation of both the Stage 1 semantic segmentation model and the Stage 2 dual-channel classification model. The Northwest Pacific training set further serves as the source domain for pretraining the Stage 3 transfer learning model. For the South China Sea dataset (500 images), a stratified 40%/20%/40% split is applied, producing 200 fine-tuning, 100 validation, and 200 test samples, dedicated exclusively to the target-domain adaptation and evaluation of the Stage 3 few-shot transfer learning experiments. All splits are performed via stratified random sampling to preserve the natural class distribution across the four acoustic categories within each subset. The Northwest Pacific and South China Sea datasets are derived from geographically disjoint WOA18 profiles with no spatial overlap, ensuring complete physical separation between the source and target domains and eliminating any risk of data leakage. All performance metrics reported in this paper are evaluated exclusively on the respective held-out test sets, which remain unseen during the entire training and hyperparameter-tuning process.
2.2. Structure-Sensitive Acoustic Field Semantic Segmentation via Deformable Convolution
The physics-prior-augmented cascaded recognition framework proposed in this paper aims to address the challenges of physical mismatch and poor generalization in sound field recognition under complex marine environments. Following the logical architecture of geometric prior extraction, multi-source feature fusion, and cross-domain knowledge transfer, the framework achieves the transformation from raw sound propagation images to high-confidence classification results through three cascaded phases. The primary objective of the first phase is to accurately extract the geometric contours of CZs from complex TL images. To mitigate the nonlinear distortion induced by acoustic refraction, this study proposes a structure-aware improved ResNet-34 network. By integrating deformable convolution and a global statistical feature module, the model enables fine-grained pixel-level segmentation of the acoustic boundaries of sound fields. The architecture of the improved sound field feature region segmentation model is illustrated in
Figure 3.
The ResNet-34 network incorporates a multi-layer residual structure, and this multi-level convolutional architecture can effectively capture both feature details and global information in conventional images [
32]. However, deep residual networks are primarily designed for recognizing natural images and lack targeted optimization for the details and global information of simulated images [
33]. To address this issue, two improvement schemes—shallow feature preservation and adaptive sampling—are integrated into the backbone network.
In natural-image processing, high-level semantic features of deep networks usually receive more attention. Nevertheless, in the sound field segmentation task, the images contain fewer features compared to natural images. Therefore, rather than relying on pyramidal skip connections that would reintroduce low-level texture noise into the physically smooth acoustic field, multi-scale acoustic context is captured at the bottleneck: after the final residual stage, a three-level adaptive pooling module extracts features at 4 × 4, 2 × 2, and 1 × 1 spatial granularities, which are concatenated with the original Stage 4 feature map to form a multi-scale representation at a single resolution level [
34]. This design preserves global acoustic energy distribution information while avoiding the parameter overhead of a symmetrical decoder, which is crucial for maintaining generalization under the data-scarce conditions of marine acoustic applications [
35].
To ensure the sensitivity of the network model to variations in sound ray density, deformable convolution is introduced into the residual blocks of Stage 3. By dynamically adjusting the sampling positions of the convolution kernel through learnable offsets in deformable convolution, the network’s recognition rate for regions with significant changes in sound ray density can be improved, along with the segmentation accuracy of proportion variations [
36].
Deep-sea acoustic field images are often characterized by high local sound intensity and wide global distribution. Single-scale pooling operations tend to lose global density distribution information, resulting in blurred boundaries between sound SZs and CZs [
37]. To this end, this paper designs a multi-scale global statistical feature module at the output of Stage 4. Through three-level pooling and feature rescaling mechanisms, the discrimination ability of the network model for sound SZs and CZs in the images is improved [
38].
Multi-scale pooling: given the Stage 4 output
, three parallel pooling branches are applied at different spatial scales, as shown in Equation (
3):
Each pooled branch is projected to 128 channels via a
convolution, as defined in Equation (
4):
Bilinear upsampling restores each projected feature map to the Stage 4 spatial resolution
, as expressed in Equation (
5):
The original Stage 4 features and the three upsampled branches are concatenated along the channel dimension to form a spatially enriched 896-channel feature map, as computed in Equation (
6):
The global statistical feature module constructed by three-level pooling and spatial concatenation enables not only the global perception of the overall density baseline but also the fine identification of local high-density regions, effectively improving the ability of the network model to detect the boundaries between sound SZs and CZs.
To transform the globally enriched feature representation into dense pixel-level predictions at the original input resolution, the proposed framework follows the FCN-32s paradigm: a purely convolutional prediction head operating at the bottleneck resolution, followed by a single bilinear upsampling that directly restores the original input dimensions. Specifically, the 896-channel feature map produced by the global statistics module (at 1/32 of the input resolution) is passed through a lightweight segmentation head consisting of three convolutional layers: a 3 × 3 convolution, reducing the channels from 896 to 256 (with BatchNorm and ReLU); a second 3 × 3 convolution, providing a further reduction to 128 channels (with BatchNorm and ReLU); and a final 1 × 1 convolution projecting to 2 channels, representing the CZ and SZ class logits. All three layers operate at the same 1/32 resolution, keeping the head computationally lightweight (approximately 0.5 M parameters). The resulting 2-channel logit map is then restored to the original input resolution via a single bilinear interpolation, yielding the final pixel-level segmentation mask. The complete spatial resolution and feature dimension flow through the architecture is illustrated in
Figure 4. This design incorporates no skip connections from earlier stages: the multi-scale acoustic context has already been aggregated within the global statistics module at the bottleneck, and a direct single-step upsampling acts as an implicit regularizer that suppresses overfitting under data-scarce conditions.
The selection of a modified ResNet-34 backbone paired with this lightweight FCN upsampling head—as opposed to over-parameterized canonical segmentation baselines like U-Net or DeepLabv3+—is justified by the unique physical characteristics of ocean acoustic fields. Unlike natural images, which contain discrete, highly heterogeneous objects, sound TL images represent continuous energy fields with periodic, relatively uniform geometric strips. Heavy decoders utilizing extensive Atrous Spatial Pyramid Pooling (ASPP) or symmetrical heavy skip connections introduce massive parameter redundancy. In our data-scarce marine environment scenario, such over-parameterization severely elevates the risk of catastrophic overfitting. Our lightweight FCN-32s single-upsampling design actively suppresses overfitting and guarantees high computational efficiency for edge-sonar deployment while focusing the network’s capacity on structural density deformations.
Considering the particularity of the sound field feature region segmentation task, a joint loss function, , is designed based on the fusion of global feature proportion and local details. Composed of the global area matching loss, denoted as , and the density gradient loss, denoted as , this loss function is used to mitigate the increase in overall misjudgment rate caused by local misclassification.
The global area matching loss, denoted as
, directly constrains the difference in global pixel proportion between the predicted region and the soft boundary label, forcing the network to learn the physical property of sound field energy conservation, as shown in Equation (
7):
Specifically, this global proportion matching loss serves a critical regularization role. As demonstrated in weakly supervised segmentation tasks where rigid boundaries are absent [
39], enforcing a global size constraint effectively prevents the convolutional network from artificially inflating regional contrast scores through the over-segmentation of low-pressure pixels. This mathematically ensures that the model learns physically meaningful convergence zone representations rather than simply grouping dark pixels.
The density gradient loss,
, targets boundary regions with drastic variations in acoustic energy gradients, and introduces a gradient magnitude weight
to make the network focus on hard-to-classify edge samples, which are defined in Equations (
8) and (
9):
where
N is the total number of pixels;
and
represent the predicted probability and soft boundary label of the
i-th pixel, respectively;
is a small constant to prevent division by zero;
is a tunable parameter set to 0.6 in the experiments; and the gradient
is calculated using the Sobel operator. By combining the global area matching loss and the density gradient loss, the final joint loss function is calculated using Equation (
10):
This design ensures the accuracy of the overall proportion and strictly mathematically bounds the segmentation area while significantly improving the boundary consistency of transition regions. Comparative experiments demonstrate that compared with a single loss function, the proposed joint loss function can effectively reduce the local misjudgment and over-segmentation risks caused in the segmentation process.
2.3. Coupling of Physical Priors and Visual Textures via Dual-Channel VGG-Based Architecture
To construct the classification branch, selecting an appropriate feature extraction backbone is critical, especially under data-scarce marine environments. The selection of the standard VGG16 architecture, rather than a modern heavy model (such as Swin Transformer or ConvNeXt [
40]), is rigorously grounded in the specific challenges of acoustic fields. First, modern Vision Transformers (ViTs) lack the inherent inductive biases of CNNs, rendering them highly prone to severe overfitting when trained on extremely small datasets (500 samples) without massive domain-specific pretraining [
41]. Second, transferring weights from foundation models (e.g., DINOv2) pretrained on natural optical images often triggers negative transfer when applied to continuous, abstract acoustic transmission loss (TL) fields due to the fundamental domain gap between natural optical images and continuous acoustic fields [
42]. Standard CNNs like VGG16 operate as stable, structurally homogeneous feature extractors [
43], effectively capturing macro-level acoustic energy fringes with minimal parameter redundancy.
However, the original VGG16 network is typically designed for single-channel input when processing grayscale images, with its convolutional kernel parameters optimized based on a single data modality [
44]. In the preceding phase, by improving the ResNet-34 network, regional segmentation results for raw images were obtained in the task of sound SZ and CZ segmentation. The output files include YOLO-format files for feature regions of sound TL images, as well as the area proportions of sound SZs and CZs. Therefore, when classifying preprocessed images, a dual-channel input design that combines visual features and geometric prior features should be adopted [
45]. This dual-channel design achieves synergistic optimization of purely data-driven texture learning and prior-augmented learning by processing visual features and hand-crafted spatial features in parallel [
46]. Specifically, the raw grayscale TL images are fed into the first channel, which preserves the spatial details of the sound field energy distribution; the extracted geometric prior parameters (including the area proportions and regional coordinates of sound SZs and CZs) are input into the second channel [
47]. These regional coordinates are strictly encoded as bounding box corners. This compact rectangular representation effectively encapsulates the essential macro-scale physical spans (horizontal range and vertical depth) while smoothing out diffuse, noisy acoustic boundaries, thereby acting as a crucial geometric regularizer to prevent small-sample overfitting. The architecture of this improved dual-channel classification model is illustrated in
Figure 5.
To achieve efficient fusion of such bimodal data, modifications to the front-end structure and feature interaction mechanism of the VGG16 network are required. First, a dual-branch parallel processing architecture is constructed at the input layer. For the visual feature branch, the standard convolutional modules of VGG16 can be retained to perform layer-by-layer spatial feature extraction on raw grayscale images, capturing local details such as sound field energy gradients and CZ edge morphology. For the geometric prior feature branch, spatial encoding is performed on the CZ area proportion parameters and the sound-ray-density gradient histogram. The area proportion scalar is converted into a full-resolution heatmap, where pixel values are positively correlated with the positions of CZs; meanwhile, the sound-ray-density gradient histogram is mapped to a feature map through fully connected layers, with bilinear interpolation used to achieve spatial dimension matching. These two components are superimposed to form a geometric prior feature map, which is then input into a lightweight convolutional module to extract global statistical features.
To integrate the vector characteristics of visual and geometric prior features in the dual-channel input, in the feature fusion stage, the model adopts a feature alignment strategy to unify the number of channels between the output of the 4th convolutional block in the visual branch and the output of the 3rd layer in the geometric feature branch. A convolution is used to increase the number of channels of the geometric prior features to 512, ensuring consistency in spatial resolution with the visual branch.
Finally, to address the overfitting issue caused by the excessive number of parameters in the fully connected layers of the traditional VGG16, the model replaces the fully connected layers with global average pooling (GAP): the three fully connected layers at the end of the original network are removed, and GAP is directly applied after feature fusion to reduce the feature map to a 512-dimensional vector. This operation effectively reduces the computational load of model parameters while preserving the global statistical information of spatial features.
The improved model embeds the geometric prior knowledge of the sound field into the deep learning framework through a dual-channel input architecture, solving the problem of disconnection between visual features and geometric parameters in traditional methods. The visual branch focuses on the microscopic representation of the local sound field energy distribution, while the geometric feature branch quantizes the morphological laws of CZs from a global perspective. The two branches achieve complementary enhancement through multi-scale concatenation and feature compression, improving the nonlinear expression ability of the network and achieving comprehensive improvements in both classification efficiency and accuracy.
Considering the specificity of the dual-channel input structure, the training procedure of the improved VGG16 network needs to take into account both the heterogeneous characteristics of multimodal data and the convergence stability of the model. In the data preprocessing stage, the heterogeneity of the dual-channel input requires differential processing of visual and geometric prior data. The input of the visual branch is the original sound TL grayscale image. Linear normalization is used to map the original sound intensity values to the interval
, and bicubic interpolation is adopted to uniformly scale the image to a resolution of
, so as to fully retain the spatial details of the sound field energy distribution [
48]. The CZ area proportion parameters and sound-ray-density gradient histograms input to the geometric feature branch need to be converted into spatially encoded features: first, the area proportion parameters are used to generate a full-resolution heatmap, where the pixel values at the positions corresponding to the CZs are set to the proportion scalar values, and the pixel values of non-CZs are set to zero; subsequently, the 10-dimensional gradient histogram is expanded into a
feature map through bilinear interpolation, which is then pixel-wise superimposed with the heatmap to form a geometric prior feature map; finally, normalization is performed to constrain the geometric prior feature map to the interval
, so as to enhance the training stability of the model. The alignment processing of the dual-channel data ensures the effectiveness of subsequent feature fusion.
The initialization strategy of network parameters directly affects the convergence speed and generalization ability of the model. The visual branch inherits the pretrained VGG16 weights on the ImageNet dataset [
49]; the lightweight convolutional layers of the geometric feature branch are initialized with a normal distribution to adapt to the nonlinear characteristics of the ReLU activation function; the weights of the feature encoding layer are initialized as an identity matrix to ensure the lossless transmission of geometric prior features. The
convolutional layer of the fusion module is initialized with a Xavier uniform distribution to balance the weight allocation of dual-channel features [
50]. Through the above-mentioned differential initialization strategy, the model can establish initial associations between multimodal features at the initial stage of training, laying a foundation for the convergence of subsequent training.
The design of the loss function aims to address the problems of class imbalance and consistency with geometric prior logic. A weighted cross-entropy loss function is adopted, which assigns higher weights to infrequent classes according to the class distribution in the training set, so as to alleviate prediction bias caused by sample imbalance. Meanwhile, a geometric constraint regularization term is introduced to force the model predictions to be logically consistent with geometric prior features. Specifically, the geometric constraint regularization term utilizes an indicator function that applies a linear penalty to the predicted probability of physically contradictory classes (e.g., predicting a low-proportion class when the input CZ area proportion exceeds 70%). Because the area proportion acts purely as a constant external condition during backpropagation, the gradient flows exclusively through the network’s continuous probability outputs, making the penalty strictly end-to-end differentiable.
The physical penalty term
is formally defined as Equation (
11):
where
is the indicator function (evaluating to 1 when the condition holds and 0 otherwise),
denotes the CZ area proportion provided by the Stage 1 segmentation output,
is the set of physically contradictory classes for the high-proportion case, and
is the Softmax-predicted probability assigned to class
c. Because
enters the penalty solely as a constant gating condition during backpropagation, gradients flow exclusively through the continuous Softmax outputs
, ensuring strict end-to-end differentiability.
It should be noted that Equation (
11) encodes only the high-proportion constraint direction (
must be strong convergence). The symmetric low-proportion case (
shadow zone) is deliberately not penalized, because in the 20–50% intermediate range—where the exploitable and weak convergence types are discriminated by spatial continuity rather than area proportion alone—a rigid low-proportion penalty would risk incorrectly suppressing valid weak convergence predictions. The penalty is therefore designed as a conservative, asymmetric safeguard: it intervenes only when the physical evidence is unambiguous (very high CZ proportion), and defers to the data-driven classifier in all other regimes.
The mathematical expression of the joint loss function is given as Equation (
12):
where
and
are hyperparameters, and the optimal ratio is determined through grid search. As a fundamental physical safeguard rather than a tunable heuristic, this penalty qualitatively contributes to the framework by actively eliminating physically absurd misclassifications caused by severe localized noise.
A selective parameter freezing and joint fine-tuning strategy was adopted during training to balance the learning rates of the bimodal features and prevent the large initial gradients of the classification layer from disrupting the pretrained visual weights. The optimization process involved freezing the visual branch to concentrate on training the geometric feature encoding network and the classification layer, allowing them to establish initial discriminative capabilities. Subsequently, the visual network was unfrozen to perform end-to-end joint fine-tuning with a minimal learning rate, ensuring the perfect alignment of the feature extractors and the final decision boundary. A dynamic learning-rate-decay strategy was applied throughout the training process, with the learning rate decayed to half of its original value every ten epochs to achieve stable convergence. The training monitoring mechanism further ensured model stability by incorporating an early-stopping mechanism based on the real-time monitoring of the validation loss and accuracy, which automatically terminates the training process upon sustained performance stagnation. Furthermore, L2 norm clipping was performed on the gradients of the geometric feature branch to avoid gradient explosion caused by differences in feature dimensions.
2.4. Feature Enhancement Based on Convolutional Autoencoder
To address the practical challenge of acoustic data scarcity in unfamiliar maritime regions, where it is typically not feasible to acquire extensive in situ measurements, this study employs the geographically distinct South China Sea simulated dataset (500 images) as a controlled proxy for a data-scarce target domain, while the distribution discrepancy between the Northwest Pacific source domain and the South China Sea target domain (known as domain shift) motivates the adoption of a few-shot classification paradigm based on deep transfer learning in the third stage of the cascaded framework. This section proposes a few-shot sound field classification method based on pretrained transfer learning. Prior to the experiments, the adopted pretrained transfer learning strategy is determined and a hybrid sound field image enhancement method based on Mixup and convolutional autoencoder is proposed. This method effectively suppresses the overfitting phenomenon in the few-shot training process and improves the generalization ability of the sample dataset. During the transfer learning training process, the improved VGG16 model is pretrained on deep-sea sound field data of the Northwest Pacific Ocean, driven by temperature and salinity data from multiple sea areas. A parameter freezing strategy is used to retain the universal feature extraction capability, realizing domain adaptive transfer of multi-scale sound field features and completing the few-shot sound field classification task.
Mixup is an image augmentation technique based on linear interpolation that can enhance images of different categories [
51,
52]. Its core principle is to linearly mix training samples and their corresponding labels to generate new sound field data samples, thereby improving the classification accuracy of the model while enhancing its generalization ability, as calculated using Equation (
13):
where
and
, along with
and
, are two randomly selected sound field image sample pairs;
x denotes the TL matrix,
y represents the corresponding sound field-type label, and
and
represent the new sample after mixing augmentation [
53,
54]. The mixing coefficient
follows a Beta distribution. To ensure that the mixed sound ray distribution conforms to the physical laws of sound propagation in sound SZs and CZs, the sampling range of the mixing coefficient
must be constrained to avoid excessive discrepancies between the augmented images and the actual TL images, thus preventing the periodic characteristics of sound propagation in CZs from being disrupted [
55]. The probability density function plot of the Beta distribution is illustrated in
Figure 6, where the value is affected by the interaction of
and
. A larger
leads to an increase in the training error of the Mixup augmentation method, which is reflected in the results as an increase in the mixing ratio of different images and an enhancement in generalization ability. The data augmentation results under an unfixed mixing coefficient are presented in
Figure 7, where one typical SZ sound field feature and one typical CZ sound field feature are randomly selected for the Mixup operation, and nine augmented images with mixing effects are output.
After mixing and augmenting two sound field region images using the Mixup method, the weights of the two images change with the mixing coefficient. The mixed and augmented images can serve as an expanded dataset to improve the training effect of the source-domain model. The convolutional autoencoder (CAE) is a classic deep learning architecture with prominent advantages in image feature learning and data augmentation. This method adopts CAE to enhance sound field images, where the encoder extracts sound ray propagation features from TL images, performs controllable transformations in the feature space, and finally reconstructs new samples conforming to sound field geometric and structural characteristics via the decoder [
56,
57]. Compared with traditional methods, CAE-based feature enhancement can automatically learn key patterns such as sound-ray-density distribution and CZ texture through convolutional operations and realize complex nonlinear interpolation in the feature space, which better meets the requirements of real-world datasets [
58].
The image enhancement process consists of three steps:
In the image reconstruction stage, the transformed feature vector is fed into the decoder and the reconstructed sound field image is output, which preserves the sound ray distribution characteristics of the original image. The structure of this method is illustrated in
Figure 8. The Mixup method is combined with the convolutional autoencoder-based image enhancement approach to achieve in-depth expansion of the details of image samples.
2.5. Ocean Sound Field Feature Classification Method Based on Transfer Learning
The flowchart of the few-shot sound field feature classification method based on pretrained transfer learning proposed in this section is shown in
Figure 9. The algorithm consists of two main parts: the pretraining of the source-domain model and the training of the target-domain model after transfer learning. In the source-domain learning step, a large number of TL images are first obtained and the relationships between different types CZ ratios and CZ characteristics are constructed. Then, image-level preprocessing is performed, which is similar to the region segmentation in the improved ResNet-34 network mentioned earlier [
59]. After grayscale processing and average calculation of grayscale regions, the proportions of sound SZs and CZs are extracted and fed into the improved VGG16 network of the source domain as geometric prior features for model pretraining. After outputting the recognition results, verification and feedback are conducted to further confirm the feasibility of the source-domain model in recognition and classification.
In the target-domain training step, similar data preprocessing is first implemented on the few-shot target-domain dataset [
60]. Then, transfer learning is used to import the source-domain model, the low-layer parameters of the source-domain model are frozen, and the high-layer parameters are fine-tuned to obtain a network model suitable for the target domain. Finally, a well-trained model for few-shot TL images is obtained through training and validation [
42,
61].
The weighted cross-entropy loss function is adopted to address the class imbalance between the target-domain dataset and the source domain, which is formulated as Equation (
16):
where
is the target class label,
is the predicted probability, and
is the class weight, which can be adjusted appropriately according to the sample proportion of each class in the target domain. In practice, the class weights are set to 0.8 (CZ-S), 1.2 (CZ-E), 1.5 (CZ-W), and 0.8 (SZ), with higher weights assigned to the minority and easily confused classes to mitigate class imbalance. These values are configurable in the released source code. Meanwhile, to reduce the overfitting risk of the model, the regularization loss of the model is calculated using Equation (
17):
where
is the regularization coefficient and
denotes the model parameters.
To quickly adapt to the class space of the target domain during the classification layer training step, the AdamW optimizer is selected, while backbone parameters are frozen to enhance the generalization ability of the model. The initial learning rate is set to
and the batch size is set to 16. A cosine annealing learning rate schedule is adopted to dynamically adjust the learning rate and accelerate the network convergence, as expressed in Equation (
18):
where the maximum learning rate is
, the minimum learning rate is
, and the cycle length is
.
The hierarchical training process is executed in three sequential phases. Initially, for class space adaptation, all convolutional layers are frozen and only the fully connected layers are trained. Operating with a learning rate of for 5 epochs, this phase quickly adapts the network to the class space of the target domain and establishes an initial classification boundary. Subsequently, to achieve feature fine-tuning and fine-grained adjustment, the final six convolutional layers are unfrozen, allowing the fully connected layers and convolutional classification layers to be jointly trained. Utilizing a learning rate of coupled with a cosine annealing schedule, the model is trained for 20 epochs. This stage realizes the fine-grained feature recognition of target-domain data and aligns the feature tuning with the source-domain model. Finally, for global representation alignment, the global statistical module is jointly fine-tuned. The learning rate is maintained at for an additional 15 epochs to further enhance the feature classification capability, specifically for the South China Sea sound TL images, thereby completing the comprehensive cross-domain adaptation protocol.