Next Article in Journal
Numerical Modelling of the Barotropic Tides in the Gulf of Aden
Previous Article in Journal
Simulation Methodologies for Fatigue Damage Estimation in Subsea Power Cable Conductors: A Comparative Study
Previous Article in Special Issue
A Data-Driven Matching Error Compensation Framework for Underwater Gravity Aided Inertial Navigation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

UCA-PPO: USV Path Planning for Moving Underwater Target Tracking with Passive Acoustic Observations and Position Priors

1
National Innovation Institute of Defense Technology, PLA Academy of Military Science, Beijing 100071, China
2
Ocean College, Zhejiang University, Zhoushan 316021, China
3
School of Computer Science, Shanghai Jiao Tong University, Shanghai 200240, China
4
School of Information and Communication Engineering, University of Electronic Science and Technology of China, Chengdu 611731, China
*
Author to whom correspondence should be addressed.
J. Mar. Sci. Eng. 2026, 14(18), 1695; https://doi.org/10.3390/jmse14181695
Submission received: 27 August 2026 / Revised: 3 September 2026 / Accepted: 8 September 2026 / Published: 12 September 2026

Abstract

Autonomous navigation of a USV during moving-underwater-target tracking becomes difficult when passive-acoustic contact is intermittent and periodic position priors are stale or inaccurate. We formulate the task as online path planning and develop UCA-PPO, a quality-conditioned cross-attention extension of PPO. Its actor separately encodes USV context, acoustic cues, and position priors. Prior age and stated error, contact status, bearing confidence, and signal excess form a dynamic observation-quality token that conditions cross-attention fusion, after which a Gaussian policy head produces speed and heading-increment commands. The token is a deterministic engineering descriptor rather than a calibrated Bayesian, posterior, or predictive uncertainty estimate. PPO training uses an asymmetric critic, whereas deployment does not require target truth. The two-dimensional simulator combines GEBCO 2025 bathymetry, World Ocean Atlas 2023 hydrography, and range-dependent acoustic-model transmission loss. The evaluation spans three prior-quality regimes, five training seeds, 70 independently trained policies, and 70,000 held-out episodes, with PPO, CA-PPO, targeted ablations, parameter-matched controls, and three non-learning controllers as references. Under severe prior degradation, UCA-PPO improves return, 5 km dwell, and modeled detection probability over PPO by 9.2%, 10.2%, and 13.6%, respectively, while reducing conditional first-contact time by 26.3%. Its mean advantage over CA-PPO is small, although its across-run dispersion is lower. A 135,000-episode frozen-policy factorial test shows that UCA-PPO leads only when 5 km position error is combined with 30- or 60-decision updates. In a 15,000-episode zero-shot Munk-profile transfer, UCA-PPO retains the highest acquisition success (0.944) but not the highest return or shortest distance. Five 300,000-episode retraining runs without the sole acoustic reward retain acquisition success of 0.966, but reduce 5 km dwell from 0.600 to 0.447 and more than double cumulative turn. Across these controlled acoustic settings, the results support quality-conditioned multisource fusion mainly as a means of improving acquisition robustness when external navigation information is substantially degraded, not general superiority over CA-PPO or field readiness.

1. Introduction

Unmanned surface vehicles (USVs) support environmental observation, hydrographic survey, marine-life monitoring, and submerged-instrument operations [1,2,3]. Their autonomy depends not only on low-level motion control, but also on navigation decisions made from incomplete and time-varying information. One such mission is to keep a USV near a moving underwater target long enough to collect useful passive-acoustic observations. The target continues to move while range remains unmeasured, and the next maneuver may have to be selected without a fresh bearing.
Useful guidance arrives through two incomplete channels. A coarse target-position prior provides global direction only at intervals, and its value fades as the message ages or its position error grows. Passive acoustics are tied more closely to the current geometry, but they supply intermittent contact, a noisy bearing, signal excess, and bearing confidence rather than range. This makes the task an active navigation problem rather than ordinary waypoint pursuit: spatially varying transmission loss couples the selected route to the conditions under which contact may be recovered later.
UCA-PPO does not create another sensor measurement. Instead, the prior and acoustic observation retain separate representations, while prior age, stated error, acoustic contact, signal excess, and bearing confidence are reorganized as an observation-quality token. That token conditions cross-attention before the actor produces speed and heading-increment commands. Here, “uncertainty” is a deterministic description of data quality rather than a calibrated posterior; target truth remains unavailable to the deployed actor.
This formulation places intermittent acoustics, periodic priors, and route-dependent observability in one partially observable navigation problem. The contribution lies in the actor representation rather than the PPO objective itself: deployable observation-quality variables condition the fusion of heterogeneous guidance sources before continuous motion commands are produced. We examine that choice in a RAM environment where acoustic availability varies with route geometry. Standard PPO, ordinary cross-attention, targeted ablations, parameter-matched networks, and three transparent engineering controllers separate the effect of information organization from network width and non-learning guidance. Variation across five independently trained seeds is retained rather than hidden by episode averaging. Figure 1 summarizes the task and sensor interpretation. The resulting evidence is deliberately narrow; it does not establish calibrated uncertainty, uniform superiority over CA-PPO, broad transfer across propagation fields, or field performance.
The primary validated scenario is a cooperative or instrumented acoustically emitting AUV/tag whose periodic external position message is available through an abstract mission-system interface. Marine-animal tracking, drifting objects, and generic non-cooperative targets can differ in signature, behavior, prior availability, and operational constraints and are not validated by the present experiment.
The proposed actor is a bounded mission-level tracking-guidance module for a controlled operating area, not a complete surface-navigation stack. A deployable USV requires a supervisory safety layer that fuses chart, AIS, radar, and visual information; enforces geofences and grounding margins; evaluates collision risk; and overrides tracking commands when COLREGs-compliant avoidance is required [4,5]. Collision-case-derived encounters provide a useful validation layer [6]. Navigation-GPT-type decision support may supervise mission intent, explanations, and safety constraints, while UCA-PPO remains subordinate tracking guidance [7]; this hierarchy is conceptual and untested here.

2. Related Work

2.1. Intelligent USV Navigation and Moving-Target Following

USV path-planning studies cover path following, collision avoidance, station keeping, and coordinated motion. Many path-following formulations provide a geometric cross-track error or a waypoint sequence directly to the controller [2,8]. Underactuated and uncertain vessel models have also been addressed with continuous-control reinforcement learning [9]. These problems are demanding at the control level, but their information structure differs from the present task because the desired path or target state is usually continuously available. Here, the policy must form the path online from periodically updated position information and acoustic observations that may be absent at a decision.
Target following also brings a prediction problem. Wang et al. [10] combined predictive target information with soft actor–critic for end-to-end USV control under nonlinear dynamics and disturbances. That input helped the policy anticipate motion, but it carried substantially more information than the finite-frequency, error-contaminated prior and intermittently missing bearing used here. In the present setting, relative motion is only part of the decision: the chosen trajectory also shapes the quality of future measurements through underwater propagation.
Field systems reveal demands that a controlled simulation leaves aside. The low-cost autonomous surface vehicle reported by Gogendeau et al. [3], for example, combined acoustic tracking for ecological surveys with bathymetric and photogrammetric functions. Its operation depended on sensor installation, platform self-noise, communication, and environmental validation. Those system-level issues sit outside the present path-planning and multisource-fusion study.

2.2. Bearings-Only Estimation and Partial Observation

Bearing measurements constrain direction but not range. Observability improves through relative maneuver and repeated measurements, whereas poor geometry can leave substantial range uncertainty [11]. Extended or unscented filters can be used when a local Gaussian approximation is adequate; particle filters are more flexible when the posterior is non-Gaussian or multimodal [12,13]. A practical estimator-based path-planning system must additionally specify a target process model, detection and missed-contact likelihoods, a resampling rule, and a planning or control module that converts the estimate into USV commands.
Because such modular systems remain an important engineering reference, the evaluation code includes a particle-filter pursuit controller that fuses the declared position prior with detected acoustic bearings before steering toward the posterior mean. Direct prior pursuit and a short-horizon prior MPC are included as simpler controls. These baselines do not use simulator target truth; they expose the benefit, if any, of learned observation-to-action mapping relative to transparent estimation–guidance pipelines.
A POMDP gives the appropriate decision-theoretic description because actions depend on observations rather than the latent state [14,15]. Range is unobserved, contact is stochastic, and the external prior is held between updates. Although a recurrent policy could retain history, the implemented observation already contains a held bearing with age-decayed confidence and the elapsed time since the last scheduled prior update. The three main actors therefore use a single frame, keeping the comparison focused on source fusion and avoiding recurrent-state initialization as another experimental factor. This choice is specific to the present sensor interface; it is not a claim against recurrence elsewhere. A recurrent PPO or attention-over-history baseline remains necessary to test whether longer contact history adds information beyond this engineered short-term state.

2.3. Attention and Deterministic Observation-Quality Features

Classical multisensor fusion records uncertainty, source quality, and information timing [16,17]. Neural attention can also preserve source identity while changing the combination with context [18], but the two ideas are not equivalent. Attention weights are internal activations rather than calibrated probabilities; preferring a token does not necessarily quantify whether its source is trustworthy.
UCA-PPO represents observation quality using quantities already available to the actor: time since the prior update, stated prior-error scale, current contact state, decayed bearing confidence, and normalized signal excess. It maps a deterministic descriptor derived from these variables to the observation-quality token. “Uncertainty” therefore refers to an encoded engineering descriptor, not to Bayesian parameter uncertainty, posterior covariance, or predictive calibration [19]. This distinction is retained throughout the analysis.
The main actors also differ in capacity: PPO has 70,660 parameters, CA-PPO has 189,957, and UCA-PPO has 207,685. The severe-regime extension addresses this confounding by widening PPO and CA-PPO to 207,840 and 208,007 parameters, respectively, while keeping the training-only critic fixed. Equal-capacity constant-, prior-, and acoustic-token controls further separate the dynamic uncertainty descriptor from topology.

2.4. Propagation-Aware Acoustic Simulation

Passive-sonar performance depends on source level, transmission loss, ambient and platform noise, and the detection criterion [20,21,22]. Range-only attenuation is inexpensive but suppresses the effects of seabed geometry, sound-speed refraction, and azimuth. Range-dependent parabolic-equation solvers retain these effects and are widely used for numerical propagation studies [23,24].
Directly solving the parabolic equation inside billions of reinforcement-learning transitions is impractical. The alternative adopted here is an offline RAM database followed by memory-mapped lookup during training. This preserves spatial variation at rollout speed, at the cost of grid discretization and a finite local support. The database is constructed from a GEBCO bathymetry tile and a World Ocean Atlas sound-speed profile. Its source–receiver depth convention and its 20 km radial support are described explicitly in Section 3.3, because both affect the physical interpretation of the learned policy.

3. Materials and Methods

3.1. POMDP and Deployment Information Flow

The process is represented by ( S , O , A , P , r , γ ) . The latent state s t S contains the true USV and AUV states, target-route variables, acoustic state, and held prior. The deployed actor receives o t O and selects a t [ 1 , 1 ] 2 . The environment then advances the platform and target, updates the prior and acoustic contact, and returns reward r t . Figure 2 distinguishes the deployment path from the privileged state-value critic used only during training.
The action maps to a speed command v t u [ 0 , 5 ] ms−1 and a requested heading increment in [ 10 ° , 10 ° ] per 60 s decision. The applied increment is slew-limited to 2° between decisions. With the applied increment Δ ψ t ,
ψ t + 1 u = wrap ( ψ t u + Δ ψ t ) ,
p t + 1 u = p t u + Δ t 1000 v t u cos ψ t + 1 u sin ψ t + 1 u .
Positions are clipped to the 100 km square. The environment therefore evaluates online path planning at the kinematic level rather than vessel maneuvering with a 3-DOF dynamic model.
At reset, the USV position is sampled uniformly from the interior square [ 12 , 88 ] × [ 12 , 88 ] km, its heading is uniform on [ π , π ) , and its speed and applied heading increment are zero. One of the four domain sides is then selected with equal probability as the AUV entry side. The AUV starts 1 km inside that side, with its lateral coordinate sampled uniformly from [ 15 , 85 ] km. Its exit point lies 5 km beyond the opposite boundary. The exit lateral coordinate equals the entry coordinate plus a zero-mean Gaussian perturbation with standard deviation 8 km, clipped to [ 15 , 85 ] km. The initial AUV heading points toward this exit with an additional zero-mean 0.05 rad perturbation.
The AUV subsequently travels at 2 ms−1 toward the sampled exit point. The entry–exit randomization prevents the path-planning policy from encountering a single repeated route. After the USV has applied its current action, the desired AUV motion vector combines a unit vector toward the exit with a repulsive artificial-potential-field term evaluated at the updated USV position,
f t o = p exit p t o p exit p t o 2 + w r I ( D t < R r ) 1 D t 1 R r p t o p t u D t 3 ,
where R r = 15 km and w r = 25 . A zero-mean heading perturbation with standard deviation 0.05 rad is added before the AUV turn is clipped to 10°. The target is therefore reactive: once two policies select different actions, they need not experience identical target paths even when evaluated with the same initial seed. Comparisons use the same seed set rather than assuming pointwise paired trajectories.
A controller-independent target-motion experiment is still required to separate planner performance from policy-dependent target reactions.
With the clipped AUV heading increment written as δ t o , its state is advanced by
ψ t + 1 o = wrap ( ψ t o + δ t o ) ,
p t + 1 o = p t o + Δ t 1000 v o cos ψ t + 1 o sin ψ t + 1 o , v o = 2 m s 1 .
The transition order is fixed. The normalized action is mapped to the physical USV commands, the USV state is advanced, the reactive AUV is advanced using the updated USV position, and the environment then updates the scheduled position prior, acoustic variables, reward, and next observation. At reset ( t = 0 ), one noisy prior message and one acoustic contact trial are generated before the first action. This reset observation initializes the feed-forward actor; the trajectory averages reported in the evaluation are formed from the post-action decisions t = 1 , , T .

3.2. Observation and External Position Prior

For one USV, the deployed observation is
o t = [ o t c , o t a , o t p ] R 16 .
Table 1 lists all entries. The six context variables describe the USV only. Absolute coordinates are divided by the 100 km area scale; heading is represented by sine and cosine; speed and applied turn are divided by their limits.
The acoustic block is
o t a = [ d t , sin β ^ t , cos β ^ t , c t β , S E ¯ t ] .
Here d t is the contact sampled at the current decision and β ^ t is the most recently detected bearing relative to the current USV heading. A contact stores the noisy bearing and initializes its confidence to max ( P d ( t ) , 0.05 ) . On a missed contact, the stored bearing and its initial confidence are retained, while an age counter increases up to 50 decisions. The confidence exposed to the actor is c t β = c t f exp [ min ( t t f , 50 ) / 50 ] ; thus the implementation reaches a nonzero floor after 50 missed decisions rather than deleting the bearing abruptly. The signal-excess entry is S E ¯ t = tanh ( S E t / 10 ) on contact and zero on a miss. This explicit held-bearing state is one reason the compared actors can operate without a recurrent hidden state.
The external prior is updated every I p decisions:
p ^ t k p = p t k o + ϵ t k , ϵ t k N ( 0 , σ p 2 I ) .
The sampled value is held between scheduled updates; no communication latency is modeled. Its five entries are the update flag, prior position relative to the USV and divided by 100 km, elapsed time since the scheduled update divided by I p and clipped at one, and σ p / 100 . The actor therefore knows when an update occurred and its nominal error scale, but not the realized prior error.
The prior is an interface-level mission-system message rather than a direct measurement made by the USV. In a cooperative deployment, it could originate from an AUV navigation solution reported through an acoustic modem or surface relay; in a non-cooperative deployment, it could be the output of a topside multi-sensor tracker or an earlier localization stage. The simulator abstracts those upstream systems by specifying an update interval and an isotropic stated one-standard-deviation error σ p . It does not model the estimator, packet loss, latency, bias, or correlation between consecutive fixes. Directional or systematic bias, anisotropic error, burst loss, delayed delivery, and temporally correlated fixes are therefore untested. Consequently, the results concern how the planner uses a declared prior, not the accuracy of any particular localization technology.

3.3. Propagation-Aware Acoustic Channel

Bathymetry is taken from the stored 112–113° E, 13–14° N GEBCO 2025 tile [25,26] and resampled to the local 100 km domain. Temperature and salinity profiles are taken from World Ocean Atlas 2023 products [27,28]; sound speed is computed with the Mackenzie relation [29]. Figure 3 shows the stored bathymetry, sound-speed profile, and one representative transmission-loss slice.
Offline RAM calculations are stored at each 1 km point of the 101 × 101 source grid and at 10° azimuth increments. Each source has a local 400 × 400 array with 0.1 km cells. The 200 m depth file occupies approximately 6.08 GB in single precision and is accessed through memory mapping, enabling 1024-environment GPU rollouts without solving the propagation model during policy optimization. The environmental fields first define the RAM database. During rollout, the vehicle positions select a local TL value, which is combined with source level, speed-dependent noise, and the detection threshold to obtain signal excess. The resulting contact, bearing, and confidence entries are passed to the actor together with the non-acoustic observation components.
RAM calculations use 100 Hz, a modeled source depth of 150 m, and a receiver depth of 200 m. At runtime, the USV horizontal position indexes the source grid and the 200 m AUV position indexes the local map. The operational interpretation uses reciprocity: a 200 m AUV or tag is the source, and a receiver lowered to 150 m is horizontally co-located with the USV. The simulation does not represent a tether, receiver-depth excursions, surface-craft coupling, or installation noise.
The first two database indices select the nearest 1 km source point, and the last two select the nearest 0.1 km local cell. Lookup is nearest-neighbor in the stored implementation. Values beyond the valid local support, below the seabed, or otherwise invalid use a conservative 120 dB fallback. The local map therefore has a sharp boundary near 20 km, and the lookup contains source-grid and local-cell discretization. The policy is exposed to those numerical features during both training and evaluation and may learn implementation-specific discontinuities rather than only physical propagation structure.
The total noise model is
N L ( v ) = 10 log 10 10 N L 0 / 10 + 10 N L f / 10 v 4 ,
where N L 0 = 65 dB and N L f = 50 dB at 1 ms−1. Signal excess and modeled detection probability are
S E t = S L T L t N L ( v t ) D T ,
P d ( t ) = 0 , S E t 0 , 1 exp ( 0.4 S E t ) , S E t > 0 ,
with S L = 130 dB and D T = 0 dB. A binary contact is sampled from Bernoulli ( P d ( t ) ) . Conditional on contact, the bearing error is Gaussian with standard deviation varying linearly from 0.35 rad at weak modeled detection probability to 0.03 rad at strong modeled detection probability. A missed contact produces no new bearing. The actor never observes latent P d ( t ) , T L t , or true range; it receives only the sampled contact and the acoustic entries listed in Table 1.
These acoustic entries should be interpreted as outputs of an upstream passive processing chain: a detector supplies the contact decision and signal excess, while a directional hydrophone or compact array supplies a bearing estimate and associated confidence. Raw pressure time series, array calibration, beamforming, detection-threshold adaptation, and false-alarm control are outside the present model. The Gaussian bearing-error abstraction also omits non-Gaussian or multimodal errors, array-dependent bias, clutter, false contacts, and competing emitters. This separation is intentional: the study evaluates the downstream observation-to-maneuver policy given detector-level products, and does not claim a new acoustic detector or bearing estimator.

3.4. Online Path-Planning Objective, Reward, and Termination

From a path-planning perspective, the policy should bring the USV closer to the moving underwater target, favor positions and speed states with high modeled detection probability, and avoid persistently large turns. The reward therefore contains a target-proximity term, a one-step closing-progress term, a modeled detection-probability term, and a large-turn penalty. The sequence of speed and heading-increment commands defines the planned path, and the learning objective maximizes the time-discounted cumulative value of this single reward.
The path-planning objective is to learn a stochastic policy π θ ( a t o t ) over normalized actions a t [ 1 , 1 ] 2 . The two components are deterministically mapped to the physical command [ v t u , Δ ψ t ] described above. The policy maximizes the expected discounted return
θ = arg max θ J ( θ ) , J ( θ ) = E τ π θ t = 0 T 1 γ t r ˜ t ,
where γ = 0.99 and r ˜ t = κ r t with the numerical optimization scale κ = 0.01 . This constant scale changes gradient magnitudes but not the policy ordering induced by the reward. The raw per-step reward is
r t = 400 q t D + q t close + 400 P d ( t ) 10 Δ ψ t 10 ° 2 ,
where
q t D = 1 1 + D t / ( 5 km ) 2 ,
q t close = clip D t 1 D t D max step , 1 , 1 ,
and D max step = ( v max u + v o ) Δ t / 1000 = 0.42 km is the largest kinematically possible one-step relative displacement. The four terms have distinct roles. The proximity score q t D favors remaining near the target; q t close gives a dense positive or negative signal for approaching or receding; the P d ( t ) term favors locations and speeds with better modeled acoustic availability and retains sensitivity when P d is small; and the quadratic term discourages large applied heading increments. Repeated acoustic reward can therefore favor path segments that support continued acoustic contact, but contact continuity is not itself a separate reward term.
Reward uses true simulator range and latent modeled detection probability, although neither is provided to the actor. It is consequently a privileged training signal rather than a deployable measurement. Because the latent P d reward is derived from the same RAM lookup used by the simulator, it can transmit environment-specific acoustic assumptions into the learned policy even without leaking target truth at deployment. The no-acoustic-reward retraining in Section 3.6.4 directly tests this dependence. The formulation does not contain a localization-error, coverage, success-rate, or propulsion-energy term; in particular, the turn penalty is a maneuver-magnitude surrogate and should not be interpreted as a vessel-energy model. Return and mean modeled P d also share part of the same simulator model and are not independent evidence. An episode ends when the AUV leaves the operating area or after 1000 decisions; leaving the area is the usual reason that the recorded episode length is below the nominal horizon. The key simulation settings and acoustic parameters are summarized in Table 2.

3.5. UCA-PPO Path Planner

UCA-PPO is a learning-based navigation policy that converts the current USV state, a held position prior, and intermittent passive-acoustic observations into continuous speed and heading commands. It is the only proposed model in this work; PPO, CA-PPO, the ablations, the parameter-matched networks, and the engineering controllers serve as experimental controls.

3.5.1. From PPO to UCA-PPO

All three main learned policies use the same PPO updates, reward, action limits, and Gaussian action parameterization. Their distinction is confined to the actor’s observation-to-feature mapping. The reference PPO actor passes the complete 16-dimensional observation through a 16 256 256 2 multilayer perceptron with tanh activations. Its two outputs are the action means, and a separate trainable two-component log standard deviation controls exploration.
CA-PPO retains a 16 256 base branch but adds a source-aware residual update. The six context variables generate a 64-dimensional query; the five acoustic variables and five prior variables are each encoded as two 64-dimensional tokens, giving four key–value tokens in total. The attended feature is processed by a 130 256 256 fusion MLP and added to the base feature through a learned residual scale. Quality-related variables remain present in the original 16-dimensional observation, but CA-PPO does not encode them as an explicit reliability descriptor.
UCA-PPO leaves this CA-PPO path intact and appends one quality token to the four source tokens. Cross-attention therefore operates on five key–value tokens, while the critic, PPO loss, reward, and action head remain unchanged. This progression matters for interpretation: UCA-PPO modifies how the actor organizes available observations; it is not a new policy-gradient objective.

3.5.2. Common Actor–Critic Optimization

UCA-PPO uses a diagonal Gaussian actor and a training-only state-value critic. A sampled raw Gaussian action is passed through tanh to produce the normalized action, after which the environment applies the speed and heading-increment mappings above; evaluation instead applies tanh to the deterministic action mean. The critic is a two-layer, 256-unit tanh MLP. For the single-USV experiments, its 23-dimensional input contains five USV-state entries, six true AUV-state entries, the two-dimensional exit point, the previous speed and heading-increment commands, normalized episode progress, individual and joint modeled detection probabilities, and the five prior-state entries. This asymmetric critic improves value estimation during training, but the target truth, exit coordinates, and latent detection probabilities are never supplied to the deployed actor.
PPO maximizes the clipped surrogate objective [30]:
L clip ( θ ) = E t min ϱ t ( θ ) A ^ t , clip ( ϱ t ( θ ) , 1 ϵ , 1 + ϵ ) A ^ t ,
where ϱ t is the new-to-old likelihood ratio, A ^ t is the generalized advantage estimate, and ϵ = 0.1 . The implementation uses λ = 0.95 , two PPO epochs, gradient-norm clipping at 0.5, and actor and critic learning rates of 10 4 and 3 × 10 4 , respectively. The same reward and optimization settings are used for all learned controls.

3.5.3. UCA-PPO Actor

The 16-dimensional actor observation is divided into USV context o t c R 6 , acoustic information o t a R 5 , and position-prior information o t p R 5 . A base path embeds the complete observation as z t b = tanh ( W b o t + b b ) R 256 . Separate encoders map the acoustic and prior blocks into two 64-dimensional tokens each, while the context block produces the attention query.
The actor also summarizes current source quality as
ρ p ( t ) = exp ( a t p / τ ) exp ( σ t p / σ 0 ) ,
ρ a ( t ) = d t sigmoid ( S E ¯ t ) c t ,
where a t p and σ t p are normalized prior age and stated error, d t is the contact flag, S E ¯ t is normalized signal excess, c t is bearing confidence, τ = 1 , and σ 0 = 0.05 . The four quantities
u t = [ 1 ρ p ( t ) , 1 ρ a ( t ) , a t p , σ t p ]
are encoded as one 64-dimensional observation-quality token. This token is a deterministic description of observation quality, not a calibrated posterior uncertainty estimate. The two reliability scores are bounded engineering features used only as actor inputs; they are neither fitted to reliability labels nor interpreted as event probabilities. The scales τ and σ 0 are expert-selected rather than learned or calibrated parameters.
The four source tokens and the observation-quality token form Z t . Cross-attention is then computed as
q t = W q e t c ,
K t = W K Z t , V t = W V Z t ,
α t = softmax q t K t / 64 ,
h t = α t V t .
The attended feature is mapped to g t R 256 and added to the base embedding as z t = z t b + η g t , where η is trainable and initialized to 0.1. A Gaussian action head maps z t to two action means and a trainable log standard deviation. Figure 4 summarizes this data flow. At deployment, every input to the actor comes from the vehicle state, the received position prior, or the passive-acoustic processor.
The complete training and deployment procedure is summarized in Algorithm 1.
Algorithm 1 Training and deployment of UCA-PPO. The actor receives only deployable observations; simulator truth is restricted to the training critic and reward.
Require: Environment E ; actor π θ ; critic V ϕ ; episode budget N; rollout size T
Ensure: Trained actor parameters θ
  1:
Initialize θ , ϕ , optimizers, and completed-episode counter n 0
  2:
while  n < N   do
  3:
      Initialize rollout buffer B
  4:
      while  | B | < T  do
  5:
            Observe deployable o t = ( o t c , o t a , o t p ) and training-only critic state s t
  6:
            Encode prior and acoustic tokens; compute ρ p , ρ a , and observation-quality token u t
  7:
            Attend to the five tokens, add the residual base feature, and obtain π θ ( · o t )
  8:
            Sample and apply a t ; store transition, log probability, and V ϕ ( s t ) in B
  9:
      end while
10:
      Compute GAE advantages A ^ t and bootstrapped returns R ^ t
11:
      Update θ with the clipped PPO objective and update ϕ by value regression
12:
      Increase n by completed episodes and save periodic checkpoints
13:
end while
14:
Deployment: discard V ϕ and use the deterministic actor mean
15:
return  θ

3.5.4. Experimental Controls

The learned comparison set is designed to isolate what changes when the quality token is introduced. PPO provides unstructured fusion, while CA-PPO adds source-aware attention without an explicit quality token. Three equal-capacity ablations replace the full descriptor with a constant token, prior-only inputs, or acoustic-only inputs. PPO-M and CA-PPO-M then match the UCA-PPO actor parameter count within 0.2%, providing a separate check on network width.
The critic remains unchanged in every run and is excluded from the actor counts in Table 3. PPO-M differs from UCA-PPO by 155 actor parameters (0.075%), and CA-PPO-M differs by 322 parameters (0.155%).

3.5.5. Non-Learning Engineering Baselines

Three transparent controllers provide non-learning references. Greedy-Prior follows the held prior position. Prior-MPC searches 15 speed–turn combinations over an eight-decision horizon and applies the first action that best approaches the held prior. PF-Pursuit fuses prior messages and detected bearings with a 768-particle constant-velocity filter, then pursues the weighted position mean. All three use the same action limits as the learned policies, and none receives target truth, latent detection probability, or transmission loss. They are reproducible reference controllers, not exhaustive representatives of optimized belief-space planning, adaptive MPC, or estimator–planner architectures. Table 4 summarizes their information boundary; full controller settings are supplied with the reproducibility materials.

3.6. Experimental Design

3.6.1. Prior Regimes and Training Matrix

The three regimes in Table 5 jointly vary update interval and Gaussian position noise. Because the two factors change together, the experiment compares three combined operating regimes; it cannot separately identify the effects of update frequency and position error. Ten, 30, and 60 decisions correspond to 10, 30, and 60 min between scheduled updates.

3.6.2. Training and Implementation

Crossing three learned configurations with three prior regimes and training seeds 0–4 produces 45 independently trained instances. The severe regime also includes three observation-quality-token ablations and two parameter-matched controls, again with five seeds. Existing severe-regime PPO, CA-PPO, and UCA-PPO runs are reused in this extension. In total, the design contains 70 unique training runs distributed over 14 configuration–regime cells.
Each run continues until at least 300,000 episodes have finished. A rollout collects 65,536 transitions from 1024 GPU-vectorized environments, so an individual environment advances 64 decisions between updates without being reset at the rollout boundary. Episodes can span several rollouts. Completed trajectories retain their terminal flag; unfinished ones are bootstrapped from the critic and continue at the next environment step.
PPO still uses every collected transition through GAE, irrespective of whether an episode ends within the update. Episode-return plots have a different constraint: only episodes that terminate in a logging interval contribute a value. The initial interval before any vector environment finishes is consequently left blank rather than shown as zero.
Minibatches contain 16,384 transitions, and each rollout is reused for two PPO epochs. Rewards are scaled by 0.01 during optimization. Checkpoints and training summaries are written every 50 updates. Evaluation uses the final checkpoint.pt produced after the common 300,000-episode budget. A run without this file is classified as incomplete and is not included in the reported matrix, even though the evaluation utility retains fallback checkpoint names for recovery and diagnostic use. No checkpoint is selected from training return or from the evaluation episodes. This fixed-budget rule avoids an architecture-dependent model-selection advantage and permits interrupted runs to resume without changing their output path or naming convention.
Training is executed on a workstation with a 24-core AMD Ryzen Threadripper PRO 7965WX processor 256 GB host memory, and one CUDA device identified as an NVIDIA GeForce RTX 4090 with 48 GB device memory. The 6.08 GB TL array is placed on the GPU for each active process. The launcher uses three independent training processes by default. A separate concurrency benchmark can test one to twelve processes with a sufficiently long warm-up and measurement interval; changing process concurrency does not change the seed list, optimizer, rollout size, or model outputs. PyTorch (version 2.8.0, CUDA 12.8) single-precision tensors are used for the environment, actor–critic, and TL lookup.
The launcher checks the existing result tree before scheduling work and skips a run when its expected final artifacts are already complete. The recorded logs contain the update number, completed-episode count, collection and optimization time, actor and critic losses, approximate KL divergence, clip fraction, and action statistics. These logs are used to audit completion; model comparisons use separate evaluation runs.
At deployment, one actor forward pass is required every 60 s. The critic, PPO update, simulator truth, and 6.08 GB transmission-loss array are absent from this path. UCA-PPO is feed-forward, has no recurrent state, and contains 207,685 actor parameters (0.792 MiB in single precision). A decision rate of 1 / 60 = 0.0167 Hz appears modest, although neither parameter count nor batched training throughput gives batch-one latency on the target processor.
Synchronized batch-one latency was measured for the frozen seed-0 UCA-PPO actor on a separate workstation equipped with an Intel Core Ultra 7 265KF CPU and an NVIDIA GeForce RTX 5070 GPU. Across three repetitions, each using 1000 warm-up passes and 10,000 timed forward passes with gradients disabled, the CPU median latency was 0.183–0.239 ms with a 95th-percentile latency of 0.334–0.408 ms; the synchronized GPU median was 0.612–0.683 ms with a 95th percentile of 1.062–1.307 ms. These measurements cover network inference only, excluding observation preparation and upstream acoustic processing. Although the measured model-level latencies are far below the 60 s decision interval, they do not establish end-to-end hard-real-time performance or deployment readiness, which still require hardware-in-the-loop validation.

3.6.3. Evaluation Protocol

Final checkpoints use the deterministic actor mean. Each is run for 200 episodes at environment seeds 100–104, giving 1000 episodes per trained instance. Initial routes, prior errors, sampled contacts, bearing errors, and target-heading perturbations remain stochastic. Across the 70 trained instances, this protocol yields 350 checkpoint–evaluation-seed records and 70,000 episodes.
Metrics are averaged over the five evaluation seeds within a training seed before the five training seeds are compared. Summary means, standard deviations, open training-seed points, and paired contrasts therefore refer to independent training outcomes rather than to individual evaluation episodes. Figure 5 and Figure 6 additionally show the 25 nested evaluation blocks in each configuration–regime cell as small semi-transparent points. Each block summarizes 200 episodes and is included to expose within-checkpoint Monte Carlo variation, not as an additional independent sample. An audit found the expected 350 JSON results, metadata records, and episode-level CSV files; every record contained 200 episodes, with no duplicate, corrupt, or shortened files.
The three non-learning controllers require no training checkpoint. For configuration parity, the baseline evaluator reads the environment configuration from the completed PPO seed-0 checkpoint for each prior regime, then discards all actor, critic, and optimizer state. Each controller is evaluated for 200 episodes at each of the same environment seeds 100–104, giving 1000 episodes per controller and regime. The prior interval, prior-error scale, RAM database, reward, termination rule, and stochastic sensor model are therefore identical to the learned-policy evaluation.
Because a deterministic baseline has no training initialization, its five environment-seed blocks quantify Monte Carlo sampling variation rather than training variability. Baseline means may be compared with learned-policy means, but their across-block standard deviations must not be interpreted as interchangeable with the across-training-seed standard deviations in Table 6. The supplied batch evaluator records both the five block values and the pooled equal-block mean. No baseline result is used to select or tune a controller parameter.
The primary descriptive endpoints are mean USV–AUV distance, mean modeled P d , and post-acquisition contact continuity. Episode return is retained because it is the optimization target, but it is interpreted alongside physical and contact-based quantities. Because modeled P d also appears in the training reward, return and mean modeled P d are partly coupled and are not treated as two independent confirmations of effectiveness. Secondary metrics are final distance, fraction of decisions within 5 km, acquisition timing, contact coverage, velocity alignment, mean USV speed, and applied turn activity.
Let T denote the number of executed post-reset decisions in an episode, D t the USV–AUV distance after decision t, and P t the corresponding modeled detection probability. The time-averaged physical endpoints and final distance are
D ¯ = 1 T t = 1 T D t , D final = D T ,
R 5 = 1 T t = 1 T I ( D t 5 km ) , P ¯ d = 1 T t = 1 T P t .
Episode return is the undiscounted sum of the raw environment rewards over these T decisions; the factor 0.01 is used only during policy optimization. Thus R 5 measures close-range dwell, whereas P ¯ d describes modeled acoustic availability rather than sampled-contact frequency.
Let b t be the sampled contact event. If the first contact occurs at t f , continuity is
C = 0 , no sampled contact , 1 T t s + 1 t = t s T b t , otherwise ,
where t s = max ( 1 , t f ) . This convention keeps continuity on the same post-action time base as the other trajectory averages if a reset-time contact is recorded at t f = 0 . For N evaluation episodes, contact coverage is
A = 1 N i = 1 N I ( t f , i is observed ) .
The penalized first-acquisition metric equals t f , i when contact occurs and 1000 otherwise. It is useful as a single operational penalty but mixes two effects: whether contact is obtained and how long acquisition takes. The analysis therefore also reports the conditional mean first-contact step over episodes with observed contact, the fraction acquiring by decision 200, and continuity conditional on acquisition. Coverage and conditional acquisition time are kept in separate columns so that success probability is not mistaken for speed among successful episodes.
Early episode termination creates a censoring issue. When an episode ends because the AUV leaves the area before any sampled contact, its observed episode length is used as a right-censoring time. A descriptive Kaplan–Meier estimate is calculated for each trained policy,
S ^ ( t ) = t j t 1 d j n j ,
where d j is the number of first contacts at decision t j and n j is the number of episodes still at risk. The censoring mechanism is not guaranteed to be independent because leaving the area depends on the trajectory; the estimate is consequently used only as a descriptive acquisition diagnostic rather than an inferential survival model.
Velocity alignment is the mean cosine between USV and AUV velocity directions. Turn activity is reported as the mean absolute applied heading increment and its mean step-to-step change. These quantities help distinguish contact performance from aggressive maneuvering, but they are not a complete propulsion-energy model.
With five training seeds, raw seed values and mean ± standard deviation are reported for every comparison. Severe-regime attribution uses seed-paired contrasts and bootstrap 95% intervals over training seeds. These intervals describe sensitivity across the five observed initializations; they are not a substitute for a larger independent sample. An exact two-sided sign-flip test with five pairs has a minimum attainable p value of 0.0625. The analysis therefore emphasizes magnitude, direction, and operational relevance rather than a conventional p < 0.05 declaration. The 70,000 evaluation episodes reduce Monte Carlo variation within fixed checkpoints but do not increase the number of independent training outcomes beyond five.
For each paired contrast, lower-is-better endpoints are sign-reversed so that a positive difference always favors UCA-PPO. The reported percentile interval is obtained from 10,000 resamples of the five paired training-seed differences using a fixed analysis seed. Paired Cohen’s d z is the mean paired difference divided by its sample standard deviation. The exact sign-flip value enumerates all 2 5 assignments under a symmetric zero-difference null. Metrics and comparisons were prespecified by the evaluation script, but no multiplicity correction is applied; the resulting p values are therefore compatibility diagnostics rather than confirmatory tests.

3.6.4. Robustness and Operational Evaluation Protocols

The robustness evaluation comprises four analyses specified before examining their final outcomes. All four analyses are reported only after their prespecified completeness audits. The factorial evaluation contains 135,000 episodes, the Munk transfer 15,000 episodes, the common reward-ablation evaluation 10,000 episodes, and the trajectory-metric derivation 45,000 existing evaluation episodes.
Factorial Prior-Degradation Evaluation
To separate temporal staleness from spatial error, the severe-regime PPO, CA-PPO, and UCA-PPO checkpoints are frozen and evaluated on the full factorial grid
I p { 10 , 30 , 60 } × σ p { 0 , 2 , 5 } km .
The five independently trained policies of each architecture are tested at each cell using evaluation seeds 100–104 and 200 episodes per seed, giving 5000 episodes per architecture–cell while retaining the training seed as the statistical unit. Initial-condition and sensor random seeds are common across architectures and cells. This frozen-policy design asks how a policy trained under severe degradation responds when age and error are varied independently; it does not claim that the same result would be obtained after separately retraining all 27 architecture–cell combinations. Cell means, seed-level standard deviations, the marginal changes associated with I p and σ p , and their interaction are reported without treating episodes as independent replicates.
Sound-Speed-Profile Transfer Evaluation
An additional RAM database is generated over exactly the same 101×101 source grid and GEBCO bathymetry, with unchanged frequency, source and receiver depths, azimuths, range resolution, and TL post-processing. The only environmental intervention is the replacement of the WOA-derived sound-speed profile with the canonical Munk profile [31]:
c ( z ) = c 0 1 + ϵ η + e η 1 , η = 2 ( z z a ) B ,
with c 0 = 1500 m/s, ϵ = 0.00737 , z a = 1300 m, and B = 1300 m. The frozen severe-regime checkpoints are then tested under I p = 60 and σ p = 5 km with the same five evaluation seeds and 200 episodes per checkpoint–seed block. This is a zero-shot acoustic-domain-shift test: no policy is selected or fine-tuned using the Munk-profile results.
Privileged Acoustic-Reward Ablation
Five new UCA-PPO policies were trained in the severe regime for the same 300,000-episode budget and seeds 0–4 after disabling the acoustic-reward switch. In the active single-USV reward of Equation (13), this removes the sole acoustic term, 400 P d ( t ) . There is no separate sampled-contact bonus in the implemented reward: sampled contact remains an actor observation and an evaluation endpoint, not a training reward. The distance/progress and action-smoothness terms, actor observations, asymmetric critic, optimizer, rollout size, and stopping rule remain unchanged. The original-reward and no-acoustic-reward policies were evaluated under a common evaluator and common seeds. Physical distance, close-range dwell, sampled-contact acquisition, continuity, and maneuver metrics are treated as the main comparison; common-reward return and modeled P d are reported only as secondary diagnostics because they contain the simulator quantity deliberately removed during training.
Operational Trajectory Metrics
For every stored episode, path length is computed from the applied post-clipping motion,
L = t = 1 T x t USV x t 1 USV 2 ,
and cumulative maneuvering effort is summarized by Θ = t = 1 T | Δ ψ t | . Contact-acquisition success is the fraction of episodes with at least one sampled contact; the penalized first-acquisition step assigns 1000 to a miss. These quantities can be recovered exactly from the existing trajectory summaries because the applied speed, turn, decision interval, and episode length were logged. Propulsion energy, grounding risk, collision risk, and communication demand are not inferred: the present kinematic simulator has no hydrodynamic power, obstacle, grounding, or network model from which such metrics could be calculated honestly.

4. Results

4.1. Data Completeness and Statistical Unit

All 70 prespecified training instances reached the final checkpoint. The audit covered 350 checkpoint–evaluation-seed records and 70,000 episodes without finding a duplicate key, missing numeric field, corrupt file, or shortened record. An independently trained policy remains the statistical unit: the five evaluation seeds are averaged within each training seed, and variation is then reported across the five training seeds.

4.2. Performance of UCA-PPO Across Prior Regimes

Table 6 and Figure 5 do not support one ordering across prior regimes. With an accurate prior, PPO has the highest return, shortest mean distance, longest 5 km dwell, strongest contact metrics, and least turning. UCA-PPO is less variable than CA-PPO there, but it does not overtake PPO. The three models draw closer under mild degradation. CA-PPO retains a slight average edge, with a return of 405.8 ± 12.7 × 10 3 , mean distance of 9.37 ± 0.19 km, and continuity of 0.532 ± 0.033 ; UCA-PPO is only marginally earlier to first contact.
Severe degradation is where the effect of the quality descriptor becomes visible. Relative to PPO, UCA-PPO raises return from 309.2 to 337.7 × 10 3 (9.2%), 5 km dwell from 0.529 to 0.583 (10.2%), modeled P d from 0.276 to 0.314 (13.6%), and contact coverage from 0.830 to 0.983. Conditional first contact occurs 74.8 decisions, or 26.3%, earlier. As detailed in Table 7, the direction is not confined to a single initialization: return, dwell, modeled P d , first contact, and coverage favor UCA-PPO in all five paired seeds, while mean distance does so in four. This earlier acquisition comes with slightly lower continuity (0.405 versus 0.412) and more turning ( 1.38 ° versus 1.01 ° per decision).
Mean performance in the severe regime is similar for UCA-PPO and CA-PPO, but their run-to-run spread differs. Return and distance SDs are 10.5 × 10 3 and 0.24 km for UCA-PPO, against 45.5 × 10 3 and 1.28 km for CA-PPO. Only two or three of the five paired seeds favor UCA-PPO, depending on the metric. The practical distinction in these runs is a lower incidence of weak outcomes, not a consistent mean advantage over CA-PPO.

4.3. Engineering-Baseline Comparison

The engineering comparison contains 45 completed evaluations: three controllers, three prior regimes, and five environment-seed blocks of 200 episodes. Table 8 summarizes the resulting 9000 episodes. Its SDs describe Monte Carlo variation for a fixed controller, unlike the across-training-seed SDs in Table 6; the two quantify different sources of variation.
With an accurate prior, Prior-MPC leads the engineering baselines in return, mean distance, and 5 km dwell, whereas PF-Pursuit produces much stronger modeled detection probability and continuity. PPO improves return over Prior-MPC by 25.0% and has better acoustic-contact metrics, but Prior-MPC remains closer to the target geometrically. Even in this easiest regime, no single controller leads every endpoint.
PF-Pursuit becomes the strongest engineering reference once the prior is degraded. In the mild regime, CA-PPO improves its return by 20.3%, reduces mean distance by 4.8%, and raises modeled P d and continuity by 66.1% and 68.9%. The severe comparison instead favors UCA-PPO: return is 57.6% higher, mean distance 9.5% lower, 5 km dwell 71.9% higher, and conditional first contact 24.3% earlier. Contact coverage changes from 0.858 to 0.983. These percentages compare fixed-controller Monte Carlo means with means over learned-policy training seeds; they are descriptive and do not constitute a common-sample significance test.

4.4. Ablation Results

Attribution is less tidy than the main PPO comparison. In Table 9 and Figure 6, the constant-token control even has a slightly shorter mean distance (10.85 versus 10.96 km). UCA-PPO has higher mean return, dwell, modeled P d , continuity, and coverage, and reaches first contact 11.8 decisions earlier. Four of five paired seeds share the acquisition advantage, with a bootstrap interval of 2.9–24.5 decisions. Other endpoints differ by less and sometimes reverse across seeds. The data therefore associate the dynamic descriptor with faster acquisition and the absence of a pronounced weak run, not with improvement of every endpoint.
Relative to prior-only, the full descriptor improves return by 7.8%, dwell by 8.0%, modeled P d by 13.4%, and continuity by 10.0%, with the same direction in all five seeds. Acoustic-only is competitive in several runs but includes one pronounced failure, leaving a return SD of 75.5 × 10 3 and a distance SD of 2.21 km. The paired pathways seem to reduce dependence on either source; they do not guarantee the best mean after every initialization.
Widening the comparison networks does not close the gap under the shared optimizer and training budget. UCA-PPO exceeds PPO-M and CA-PPO-M by about 35% in return, 43–45% in dwell, and 63–82% in modeled P d , with the core contrasts being aligned across all five seeds. Parameter count alone is not an adequate explanation for these implementations. Since the wider controls were not separately retuned, the result says little about every possible wide PPO or CA-PPO configuration.

4.5. Robustness and Operational Evaluation Results

4.5.1. Factorial Prior-Degradation Results

All 675 prespecified checkpoint–evaluation-seed blocks completed, comprising 135,000 episodes. The audit found 135 architecture–cell–training-seed aggregates and 27 architecture–cell aggregates, with no duplicate episode keys, incomplete 200-episode blocks, non-finite reported metrics, or incorrect episode counts. Table 10 shows that neither information age nor position error produces a uniform model ordering. At σ p = 0 , PPO has the highest return at all three update intervals. With σ p = 5 km, however, the ordering progressively shifts toward UCA-PPO: its return is essentially tied with PPO at I p = 10 (388.7 versus 388.1 × 10 3 ), exceeds PPO and CA-PPO at I p = 30 (370.0 versus 359.4 and 362.0 × 10 3 ), and remains highest at I p = 60 (342.7 versus 310.7 and 331.9 × 10 3 ). At the most severe joint cell, UCA-PPO improves return by 10.3% over PPO and 3.2% over CA-PPO while retaining acquisition success of 0.983. This descriptive reversal indicates an interaction between stale and spatially inaccurate prior information for the frozen policies; it does not establish that UCA-PPO is uniformly more resistant to either factor in isolation.

4.5.2. Munk-Profile Transfer Results

The Munk RAM generation produced 360,300 valid source–azimuth solutions; the remaining 6936 combinations were exactly the boundary sources whose rays immediately left the 100 km terrain, with no unexpected missing solutions. All 10,201 source tiles were merged into the 200 m database; 65.15% of its 1.63216 × 10 9 entries are finite water-column TL values, every source tile has valid coverage, and the finite range is 42.27–128.13 dB. The subsequent 75 checkpoint–evaluation-seed blocks comprise 15,000 episodes and pass the same duplicate-key, block-size, episode-count, and numeric-field audit as Experiment 1.
Table 11 shows a substantial field shift rather than preservation of the original ranking. All three returns and modeled contact probabilities decrease relative to the WOA-derived database. Under Munk, CA-PPO has the highest return ( 161.8 ± 32.6 × 10 3 ) and modeled P d ( 0.027 ± 0.015 ), while PPO has the shortest mean distance ( 14.11 ± 0.28 km). UCA-PPO instead has the highest acquisition success ( 0.944 ± 0.029 ) and earliest conditional first contact ( 226.1 ± 18.8 decisions), but its return is below CA-PPO, and its mean distance is the largest of the three. Thus, the quality-conditioned policy retains an acquisition advantage under this single zero-shot SSP intervention, not a general performance advantage across acoustic fields or endpoints.

4.5.3. Privileged-Reward Ablation Results

All five no-acoustic-reward runs reached at least 300,000 episodes. Their completion audit verified the severe-regime configuration, seeds 0–4, the UCA actor, disabled acoustic reward, and zero acoustic-reward values in every nonempty training log. The common 5070 evaluation then completed all 50 checkpoint–evaluation-seed blocks and 10,000 episodes without a duplicate key, shortened block, non-finite metric, or incorrect aggregate count. Table 12 shows that removing the privileged acoustic reward does not eliminate acquisition, which remains 0.966 ± 0.064 , but it weakens retention and maneuver efficiency. Relative to the original-reward policies under the same evaluator, 5 km dwell decreases from 0.600 ± 0.032 to 0.447 ± 0.154 , mean distance increases from 10.66 ± 0.41 to 11.54 ± 2.54 km, path length increases from 132.68 ± 3.72 to 165.10 ± 8.65 km, and cumulative turn more than doubles from 1157.9 ± 123.4 ° to 2514.9 ± 705.0 ° . Conditional acquisition is 16.4 decisions later. The much larger across-seed SDs, including one weak run, show that the latent- P d reward materially affects the learned severe-regime behavior; the retained acquisition rate also shows that the observation architecture can still acquire contact without that reward.

4.5.4. Operational Trajectory-Metric Results

Unlike Experiments 1–3, the quantities in Table 13 are available from the completed 45,000 learned-policy evaluation episodes. The audit verified all 225 fixed 200-episode blocks, 45 training-seed aggregates, nine condition aggregates, unique episode keys, and finite derived values. Path length and cumulative absolute turn are trajectory summaries, A is sampled-contact acquisition success, and T f ( 1000 ) is the penalized first-acquisition step. PPO generally travels a shorter path and turns less, so the improved severe-regime acquisition of UCA-PPO is accompanied by greater maneuvering effort rather than being cost-free. Because the kinematic model contains no propulsion map, these results must not be relabeled as energy consumption.

4.6. Representative Severe-Regime Trajectories

Figure 7 shows one common, non-best-case rollout index under severe degradation. The training seed, evaluation seed, and episode index were fixed at 0, 102, and 93, respectively. This episode minimizes the summed robust distance from each model’s overall median on return, mean distance, 5 km dwell, contact continuity, and first-contact time; it was not selected for the largest UCA-PPO advantage. The target reacts to the policy action, so sharing the random seed aligns the initial randomization but does not force the three target paths to remain identical.
In this rollout, PPO first acquires contact at decision 369 and spends 62.4% of the episode within 5 km. CA-PPO and UCA-PPO make contact at decisions 205 and 196, raising the close-range fractions to 76.8% and 77.3%. UCA-PPO has the highest return ( 413.0 × 10 3 ) and modeled P d (0.358), whereas PPO retains the highest continuity (0.597, compared with 0.411 and 0.446). The visual difference resembles the aggregate result: earlier acquisition and longer close-range dwell are evident, but not a simultaneous lead on every endpoint.

5. Discussion

5.1. Implications for Intelligent USV Navigation

An exact prior refreshed every 10 decisions already gives PPO a strong pursuit cue. It leads the physical and contact metrics while turning least; adding tokens offers no visible mean benefit and may increase optimization sensitivity. When updates are 30 decisions apart and carry 2 km error, keeping the sources separate begins to help, and CA-PPO has the best average balance. Neither source is persistently unreliable in that regime, leaving relatively little for the explicit quality token to correct. The ranking is therefore regime-dependent rather than evidence of universal UCA-PPO superiority.
At 60-decision updates and 5 km error, the prior may point toward a stale location for a substantial part of an episode. Acoustic evidence then becomes more consequential. UCA-PPO acquires contact earlier and in more episodes than PPO, and subsequently spends more time near the target and in regions of higher modeled detectability. Its continuity is nevertheless 1.7% lower, and its turns are larger. Acquisition, retention, and maneuver effort are not interchangeable; the result is an acquisition–retention–maneuver trade-off, not simultaneous dominance of every mission endpoint.
The modest mean separation from CA-PPO is also unsurprising. CA-PPO already preserves the acoustic and prior identities and sees the quality-related variables in the common 16-dimensional observation. UCA-PPO introduces no new measurement; it presents a deterministic recombination of those variables as a fifth token. This is an inductive bias, not an information gain. After a fresh bearing arrives, both attention actors operate on the same informative source tokens and may settle into similar behavior despite different acquisition histories.
Nor does the shared reward encourage uncertainty calibration or low seed variance. It rewards proximity, approach, modeled detectability, and restrained turning; interpretability of the observation-quality token is not an optimization target. For every reported UCA-PPO–CA-PPO endpoint in the severe regime, the five-seed bootstrap interval crosses zero and paired wins are divided. CA-PPO’s larger SDs arise from one or more weak runs. The evidence points to fewer observed failures and more repeatable training, but not to a reliable increase in mean performance over CA-PPO. Moreover, return and mean modeled P d partly reuse the same simulator term and cannot be counted as independent confirmations. Resolving a smaller mean effect would require more seeds, direct interventions on token inputs, and separate tuning of CA-PPO.

5.2. Evidence from Controls and Baselines

The ablations leave a fairly narrow interpretation of the observation-quality token. Compared with a trainable constant, its dynamic form is most clearly associated with acquisition time and seed stability; most other endpoint differences are small. Prior-only trails the joint descriptor consistently, while acoustic-only contains a failed seed, suggesting that the two pathways share risk rather than simply raise the mean. PPO-M and CA-PPO-M offer a different check: under the present optimizer and training budget, width alone does not reproduce UCA-PPO. They do not establish that every wide baseline would behave similarly.
None of these comparisons establishes calibrated uncertainty or gives the attention weights a causal meaning. Prior age, error scale, contact state, signal excess, and confidence are already visible to every base actor; UCA-PPO reorganizes them but adds no sensor information. A stronger mechanism claim would need attention logging, interventions on the token inputs, and separately tuned capacity controls.
The non-learning controllers probe a different issue: how end-to-end learning compares with transparent prior pursuit, receding-horizon pursuit, and filter-then-pursue guidance. Results depend on what the prior provides and which endpoint is emphasized. Prior-MPC, for instance, is geometrically closer than PPO when the prior is accurate, yet its acoustic-contact quality is much poorer. A short-horizon point-pursuit objective is not the same as propagation-aware tracking. PF-Pursuit becomes the stronger engineering reference as the prior degrades, but still trails the learned attention policies in return and contact quality. Its constant-velocity process and bearing-only update can leave range unresolved, and the controller pursues a posterior mean rather than future observability. That account is compatible with its lower coverage and later conditional acquisition, but it is not evidence against all estimator–planner designs. The documented baseline settings were fixed rather than exhaustively tuned; the results do not establish superiority over optimized belief-space MPC and multi-hypothesis tracking, which remain relevant for future comparisons.
Inferential precision is set by five training seeds, not by the 70,000 evaluation episodes. Repetition reduces Monte Carlo noise within a checkpoint without creating new training outcomes. Five paired seeds also limit the smallest attainable two-sided exact sign-flip probability to 0.0625. We therefore retain the raw seed points, across-seed SDs, paired directions, effect sizes, and bootstrap intervals, and avoid a conventional significance claim. The severe-regime pattern is visible in these runs, but tail risk remains poorly resolved. Accordingly, the observed dispersion difference is encouraging but sample-limited rather than a definitive stability estimate.
Update interval and position error change together across the three regimes, with a separate checkpoint trained for each cell. Cross-regime differences cannot isolate either degradation factor, nor do they measure zero-shot adaptation of one policy. They describe performance when the expected quality of the prior is known during training.
The 3 × 3 factorial evaluation addresses the first part of this limitation for frozen severe-regime policies by changing the update interval and error independently. PPO leads all zero-error cells, whereas UCA-PPO leads both alternatives when the position error is 5 km and the update interval is 30 or 60 decisions. This ordering shift supports an interaction interpretation rather than uniform resistance to one isolated factor. The interpretation remains bounded: the experiment measures zero-shot response of fixed policies, not the outcome of separately retraining every architecture in every factorial cell. The path-length and cumulative-turn analysis also shows that acquisition gains can require more maneuvering, so return and contact quality should not be read as cost-free operational improvements.
The completed privileged-reward ablation provides a direct simulator-dependence check. Removing the sole 400 P d ( t ) training term preserves high acquisition success (0.966), but reduces 5 km dwell by 25.6%, lengthens the traveled path by 24.4%, and increases cumulative turn by 117.2% under the common evaluator. Its much larger seed dispersion and one weak run show that the acoustic reward materially regularizes severe-regime learning in this simulator. Conversely, the retained acquisition rate means that contact acquisition is not solely produced by that privileged reward. The result therefore supports partial reward dependence rather than either reward independence or complete collapse.

5.3. Limitations and Future Validation

Using a common RAM database for the main comparison keeps that comparison controlled, while the frozen-policy Munk experiment probes one bounded sound-speed-profile shift on the same terrain. The value of the observation-quality token nevertheless remains tied to these synthetic settings: bathymetry, seasonal sound-speed structure, frequency, source and receiver depth, source level, and platform noise reshape T L , S E , contact probability, and bearing error. In turn, ρ a may reach its floor more or less often, changing how long the policy relies on the prior. Normalized signal excess and the decayed confidence of the held bearing help with scale and short-term contact memory, but they do not make the actor invariant to a new propagation field. Persistent contact in a ducted environment might leave little for UCA-PPO to gain; extended shadow zones might force heavier reliance on the prior or move every learned policy outside its training distribution.
The Munk result demonstrates that ranking changes with the endpoint: UCA-PPO retains the highest acquisition success and earliest conditional acquisition, while CA-PPO leads return and modeled P d , and PPO leads mean distance. This mixed outcome resolves neither broad cross-environment behavior nor retrained performance. Retraining each configuration for every database would ask whether the ordering reappears when the environment is known in advance. The two questions should remain separate. Seasonal sound-speed structure, bathymetric sector, operating frequency, and source–receiver depths can be varied while mission geometry and evaluation seeds remain common; changes in ranking, paired effects, acquisition failures, and modeled contact rates would then show how far the current result extends beyond one propagation field.
The Munk-profile experiment in Section 3.6.4 implements one bounded version of the frozen-checkpoint question while holding the terrain and all other RAM settings fixed. UCA-PPO retains the highest acquisition success (0.944) but not the highest return or shortest distance. This endpoint-dependent result remains a single synthetic SSP intervention and does not establish transfer across seasons, bathymetric regions, frequencies, platform noise conditions, or measured environments.
The RAM database retains range- and azimuth-dependent structure that a fixed sensing radius would miss, but it is still a numerical approximation. Nearest-neighbor lookup introduces grid discontinuities, and fallback transmission loss creates an artificial boundary near the modeled range limit. A policy may consequently exploit numerical lookup features as well as physical propagation structure. With only one bathymetric region, two controlled sound-speed profiles, frequency, and source–receiver depth pair, generalization across acoustic environments remains untested.
The simulator omits false alarms, clutter, competing emitters, source-level variation, receiver motion, installation noise, currents, waves, actuator lag, and vessel hydrodynamics. Mean modeled P d is a simulator-derived channel-quality measure, not a field detection statistic. The two-dimensional kinematics and 60 s decision interval support an algorithmic path-planning comparison but not claims about propulsion, endurance, seaworthiness, or deployment readiness. Surface traffic, charted obstacles, grounding margins, COLREGs arbitration, communication protocols, and hardware-in-the-loop validation are also outside the present tracking-guidance model.
Independent propagation calculations would provide a direct check on the transmission-loss lookup. Repeating the evaluation across several bathymetric and sound-speed conditions is the next environmental test; measured platform noise and disturbances would become relevant before hardware-in-the-loop or model-scale trials. Until those steps are completed, the evidence remains an algorithm-level simulation result.

6. Conclusions

The autonomous-navigation problem examined here couples intermittent passive-acoustic cues with a position prior that may be stale or inaccurate. UCA-PPO retains the identity of those sources and conditions their cross-attention fusion on a dynamic quality token before issuing continuous USV motion commands. Across 70 independently trained policies and 70,000 evaluation episodes, its clearest benefit emerged under severe prior degradation. Relative to PPO, return, 5 km dwell, and modeled P d increased by 9.2%, 10.2%, and 13.6%, while conditional acquisition occurred 26.3% earlier. Against PF-Pursuit, return was 57.6% higher, mean distance 9.5% lower, and conditional acquisition 24.3% earlier.
That pattern does not extend uniformly across prior regimes. PPO remained strongest with an accurate prior, and CA-PPO had the best mean performance under mild degradation. The ablations link the dynamic token mainly to acquisition robustness; they neither establish calibrated uncertainty nor show improvement in every endpoint. The common two-dimensional terrain and single added Munk-profile transfer also keep the evidence at the algorithmic level. Whether the same information structure remains useful beyond this simulation will depend on broader cross-environment and higher-fidelity evaluation. Although 70,000 evaluation episodes reduce within-checkpoint Monte Carlo noise, they do not increase the five independent training outcomes; the observed stability pattern is therefore sample-limited. The present result is an algorithm-level simulation study, not evidence of field readiness.
The analysis additionally reports path length, cumulative turn, acquisition success, and penalized acquisition time for the completed evaluations. Under severe degradation, UCA-PPO’s higher acquisition success than PPO is accompanied by a larger cumulative turn, reinforcing the interpretation of a trade-off rather than uniform dominance. The completed factorial experiment further shows that UCA-PPO does not lead when position error is absent, but leads at the two longest update intervals when position error is 5 km; its observed advantage is therefore conditional on the joint degradation rather than a universal response to either factor. In the Munk-profile transfer, UCA-PPO retains the highest acquisition success and earliest conditional acquisition but not the highest return or shortest distance, so the acoustic-profile test likewise supports an endpoint-specific interpretation rather than broad transfer dominance. The completed no-acoustic-reward retraining retains acquisition success of 0.966 but reduces close-range dwell, increases path length and cumulative turn, and substantially increases across-seed dispersion. Thus, the architecture does not require the privileged term merely to acquire contact, but its severe-regime behavior and repeatability materially depend on that simulator-derived training signal.

Author Contributions

J.H. and F.L. conceived the study and developed the methodology; J.H. designed the UCA-PPO algorithm, conducted the simulations, and wrote the original draft; R.Z. and L.C. contributed to software implementation, formal analysis, and acoustic channel modeling; X.T., J.C., and C.D. assisted in data curation, validation, and visualization; F.L. supervised the research and managed the project; J.H. and F.L. revised and edited the manuscript. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

Source code, experiment configurations, seed manifests, processed result tables, and analysis utilities are publicly available at https://github.com/acoustic-hj/UCA-PPO-USV-Tracking (accessed on 7 September 2026). The repository also documents the provenance and construction procedure of the acoustic database and provides access instructions for larger artifacts that are not stored directly in Git. These materials support reproduction of the reported analyses and verification of the experiment matrix.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AUVAutonomous underwater vehicle
CA-PPOCross-attention proximal policy optimization
POMDPPartially observable Markov decision process
PPOProximal policy optimization
RAMRange-dependent acoustic model
TLTransmission loss
UCA-PPOUncertainty cross-attention proximal policy optimization
USVUnmanned surface vehicle

References

  1. Liu, Z.; Zhang, Y.; Yu, X.; Yuan, C. Unmanned surface vehicles: An overview of developments and challenges. Annu. Rev. Control 2016, 41, 71–93. [Google Scholar] [CrossRef] [Scilit]
  2. Qiao, Y.; Yin, J.; Wang, W.; Duarte, F.; Yang, J.; Ratti, C. Survey of deep learning for autonomous surface vehicles in marine environments. IEEE Trans. Intell. Transp. Syst. 2023, 24, 3678–3701. [Google Scholar] [CrossRef] [Scilit]
  3. Gogendeau, P.; Bonhommeau, S.; Fourati, H.; Julien, M.; Contini, M.; Chevrier, T.; Nieblas, A.E.; Bernard, S. An autonomous surface vehicle for acoustic tracking, bathymetric and photogrammetric surveys. Ocean Eng. 2025, 331, 121201. [Google Scholar] [CrossRef] [Scilit]
  4. Namgung, H. Local route planning for collision avoidance of maritime autonomous surface ships in compliance with COLREGs rules. Sustainability 2021, 14, 198. [Google Scholar] [CrossRef] [Scilit]
  5. Namgung, H.; Kim, J.S. Collision risk inference system for maritime autonomous surface ships using COLREGs rules compliant collision avoidance. IEEE Access 2021, 9, 7823–7835. [Google Scholar] [CrossRef] [Scilit]
  6. Lee, J.Y.; Namgung, H.; Kim, J.S. Development of collision-case-based testing scenarios for validating autonomous ship collision-avoidance algorithms. Brodogr. Int. J. Nav. Archit. Ocean Eng. Res. Dev. 2026, 77, 1–47. [Google Scholar] [CrossRef] [Scilit]
  7. Namgung, H.; Kim, J.S.; Jang, D.U. Path planning and collision avoidance technologies for maritime autonomous surface ships: A review of COLREGs compliance, algorithmic trends and the navigation-GPT framework. J. Navig. 2026, 78, 430–450. [Google Scholar] [CrossRef] [Scilit]
  8. Woo, J.; Yu, C.; Kim, N. Deep reinforcement learning-based controller for path following of an unmanned surface vehicle. Ocean Eng. 2019, 183, 155–166. [Google Scholar] [CrossRef] [Scilit]
  9. Qu, X.; Jiang, Y.; Zhang, R.; Long, F. A deep reinforcement learning-based path-following control scheme for an uncertain under-actuated autonomous marine vehicle. J. Mar. Sci. Eng. 2023, 11, 1762. [Google Scholar] [CrossRef] [Scilit]
  10. Wang, Z.; Hu, Q.; Wang, C.; Liu, Y.; Xie, W. Target tracking control for Unmanned Surface Vehicles: An end-to-end deep reinforcement learning approach. Ocean Eng. 2025, 317, 120059. [Google Scholar] [CrossRef] [Scilit]
  11. Aidala, V.J. Kalman filter behavior in bearings-only tracking applications. IEEE Trans. Aerosp. Electron. Syst. 1979, 15, 29–39. [Google Scholar] [CrossRef] [Scilit]
  12. Gordon, N.J.; Salmond, D.J.; Smith, A.F. Novel approach to nonlinear/non-Gaussian Bayesian state estimation. In Proceedings of the IEE Proceedings F (Radar and Signal Processing); IET: Stevenage, UK, 1993; Volume 140, pp. 107–113. [Google Scholar]
  13. Arulampalam, M.S.; Maskell, S.; Gordon, N.; Clapp, T. A tutorial on particle filters for online nonlinear/non-Gaussian Bayesian tracking. IEEE Trans. Signal Process. 2002, 50, 174–188. [Google Scholar] [CrossRef] [Scilit]
  14. Kaelbling, L.P.; Littman, M.L.; Cassandra, A.R. Planning and acting in partially observable stochastic domains. Artif. Intell. 1998, 101, 99–134. [Google Scholar] [CrossRef] [Scilit]
  15. Kurniawati, H. Partially observable markov decision processes and robotics. Annu. Rev. Control Robot. Auton. Syst. 2022, 5, 253–277. [Google Scholar] [CrossRef] [Scilit]
  16. Hall, D.L.; Llinas, J. An introduction to multisensor data fusion. Proc. IEEE 1997, 85, 6–23. [Google Scholar] [CrossRef] [Scilit]
  17. Durrant-Whyte, H.; Henderson, T.C. Multisensor data fusion. In Springer Handbook of Robotics; Springer: Berlin/Heidelberg, Germany, 2016; pp. 867–896. [Google Scholar]
  18. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
  19. Kendall, A.; Gal, Y. What uncertainties do we need in bayesian deep learning for computer vision? Adv. Neural Inf. Process. Syst. 2017, 30, 5574–5584. [Google Scholar]
  20. Urick, R.J. Principles of Underwater Sound-2; McGraw-Hill Book: New York, NY, USA, 1975. [Google Scholar]
  21. Wenz, G.M. Acoustic ambient noise in the ocean: Spectra and sources. J. Acoust. Soc. Am. 1962, 34, 1936–1956. [Google Scholar] [CrossRef] [Scilit]
  22. Ross, D. Mechanics of Underwater Noise; Elsevier: New York, NY, USA, 2013. [Google Scholar]
  23. Jensen, F.B.; Kuperman, W.A.; Porter, M.B.; Schmidt, H.; Tolstoy, A. Computational Ocean Acoustics; Springer: New York, NY, USA, 2011; Volume 2011. [Google Scholar]
  24. Collins, M.D. A split-step Padé solution for the parabolic equation method. J. Acoust. Soc. Am. 1993, 93, 1736–1742. [Google Scholar] [CrossRef] [Scilit]
  25. Weatherall, P.; Bellabad, F.; Bogonko, M.; Boza, X.; Caceres Ferreras, S.; Cornish, N.; Dael, J.; Dorschel, B.; Drennon, H.; Ferrini, V.; et al. The GEBCO_2025 Grid—A Continuous Terrain Model for Oceans and Land at 15 Arc-Second Intervals; NERC EDS British Oceanographic Data Centre NOC: Liverpool, UK, 2025. [Google Scholar]
  26. Tozer, B.; Sandwell, D.T.; Smith, W.H.; Olson, C.; Beale, J.R.; Wessel, P. Global bathymetry and topography at 15 arc sec: SRTM15+. Earth Space Sci. 2019, 6, 1847–1864. [Google Scholar] [CrossRef] [Scilit]
  27. Locarnini, R.A.; Mishonov, A.V.; Baranova, O.K.; Reagan, J.R.; Boyer, T.P.; Seidov, D.; Wang, Z.; Garcia, H.E.; Bouchard, C.; Cross, S.L.; et al. World Ocean Atlas 2023, Volume 1: Temperature; National Oceanic and Atmospheric Administration: Washington, DC, USA, 2024. [Google Scholar]
  28. Reagan, J.R.; Seidov, D.; Wang, Z.; Dukhovskoy, D.; Boyer, T.P.; Locarnini, R.A.; Baranova, O.K.; Mishonov, A.V.; Garcia, H.E.; Bouchard, C.; et al. World Ocean Atlas 2023, Volume 2: Salinity; National Oceanic and Atmospheric Administration: Washington, DC, USA, 2024. [Google Scholar]
  29. Mackenzie, K.V. Nine-term equation for sound speed in the oceans. J. Acoust. Soc. Am. 1981, 70, 807–812. [Google Scholar] [CrossRef] [Scilit]
  30. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
  31. Munk, W.H. Sound channel in an exponentially stratified ocean, with application to SOFAR. J. Acoust. Soc. Am. 1974, 55, 220–226. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Task geometry and sensor-depth interpretation (schematic, not to scale). (a) The deployed policy combines an intermittent noisy position prior with a passive bearing cue while selecting a USV route that changes the subsequent acoustic geometry. The prior marker is deliberately offset from the moving AUV/tag, and target truth remains inside the simulator. The dashed circle denotes finite modeled acoustic support rather than a measured footprint. (b) By acoustic reciprocity, the RAM lookup is interpreted as a 200 m AUV/tag source observed by an idealized hydrophone lowered to 150 m and horizontally co-located with the surface craft. The drawn paths are illustrative; tether dynamics, receiver-depth excursions, and explicit ray tracing are not modeled.
Figure 1. Task geometry and sensor-depth interpretation (schematic, not to scale). (a) The deployed policy combines an intermittent noisy position prior with a passive bearing cue while selecting a USV route that changes the subsequent acoustic geometry. The prior marker is deliberately offset from the moving AUV/tag, and target truth remains inside the simulator. The dashed circle denotes finite modeled acoustic support rather than a measured footprint. (b) By acoustic reciprocity, the RAM lookup is interpreted as a 200 m AUV/tag source observed by an idealized hydrophone lowered to 150 m and horizontally co-located with the surface craft. The drawn paths are illustrative; tether dynamics, receiver-depth excursions, and explicit ray tracing are not modeled.
Jmse 14 01695 g001
Figure 2. Closed-loop information flow (schematic). The 16-dimensional observation contains one frame: six USV context features, five acoustic features, and five prior features. Solid arrows form the deployed feed-forward loop. Simulator truth enters the asymmetric state-value critic and reward only in the lower training lane; the dashed parameter-update path is absent at deployment. No recurrent hidden state is used. The actor-side deployment path and training-only privileged path are explicitly separated so that target truth cannot be mistaken for an actor input.
Figure 2. Closed-loop information flow (schematic). The 16-dimensional observation contains one frame: six USV context features, five acoustic features, and five prior features. Solid arrows form the deployed feed-forward loop. Simulator truth enters the asymmetric state-value critic and reward only in the lower training lane; the dashed parameter-update path is absent at deployment. No recurrent hidden state is used. The actor-side deployment path and training-only privileged path are explicitly separated so that target truth cannot be mistaken for an actor input.
Jmse 14 01695 g002
Figure 3. Acoustic environment used by the simulator: (a) bathymetry, (b) sound-speed profile, and (c) the stored 200 m transmission-loss (TL) layer. The star in panel (c) marks source-grid index ( 50 , 50 ) , and the dashed circle denotes the approximately 20 km lookup support. The database contains a 101 × 101 source grid at 1 km spacing, 36 azimuths, and local 400 × 400 maps at 0.1 km resolution.
Figure 3. Acoustic environment used by the simulator: (a) bathymetry, (b) sound-speed profile, and (c) the stored 200 m transmission-loss (TL) layer. The star in panel (c) marks source-grid index ( 50 , 50 ) , and the dashed circle denotes the approximately 20 km lookup support. The database contains a 101 × 101 source grid at 1 km spacing, 36 azimuths, and local 400 × 400 maps at 0.1 km resolution.
Jmse 14 01695 g003
Figure 4. Network structure of the evaluated UCA-PPO actor (schematic). The 16-dimensional observation comprises 6 USV-context, 5 acoustic, and 5 prior features. The acoustic and prior encoders each produce two 64-dimensional source tokens; a four-component quality descriptor derived from the received acoustic and prior features produces one observation-quality token. A 64-dimensional context query attends to the five source tokens. The resulting feature passes through a 130 256 256 fusion MLP and is added, with learned scale α , to the 16 256 base branch. The Gaussian head maps the fused 256-dimensional feature to the two action means, while a separate trainable two-component log standard deviation specifies exploration. Tiled groups schematically encode changes in layer width; numerical labels give the exact dimensions, and individual cells are visual symbols rather than counted neurons. All actor inputs shown above the training-only boundary are available at deployment; target truth and latent P d are excluded.
Figure 4. Network structure of the evaluated UCA-PPO actor (schematic). The 16-dimensional observation comprises 6 USV-context, 5 acoustic, and 5 prior features. The acoustic and prior encoders each produce two 64-dimensional source tokens; a four-component quality descriptor derived from the received acoustic and prior features produces one observation-quality token. A 64-dimensional context query attends to the five source tokens. The resulting feature passes through a 130 256 256 fusion MLP and is added, with learned scale α , to the 16 256 base branch. The Gaussian head maps the fused 256-dimensional feature to the two action means, while a separate trainable two-component log standard deviation specifies exploration. Tiled groups schematically encode changes in layer width; numerical labels give the exact dimensions, and individual cells are visual symbols rather than counted neurons. All actor inputs shown above the training-only boundary are available at deployment; target truth and latent P d are excluded.
Jmse 14 01695 g004
Figure 5. Main-policy performance across prior regimes: (a) episode return, (b) mean USV–AUV distance, (c) fraction of time within 5 km, and (d) first-contact time conditional on acquisition. Small semi-transparent points show all 25 evaluation-block results in each model–regime combination (five training seeds by five evaluation seeds, 200 episodes per block). Open points show the five training-seed means after averaging the evaluation blocks; the larger model-shaped marker and bar denote the mean ± one standard deviation across those five independent training outcomes. Evaluation blocks are nested repeats, not additional independent training samples. No line is drawn between regimes. Lower is better for distance and first-contact time.
Figure 5. Main-policy performance across prior regimes: (a) episode return, (b) mean USV–AUV distance, (c) fraction of time within 5 km, and (d) first-contact time conditional on acquisition. Small semi-transparent points show all 25 evaluation-block results in each model–regime combination (five training seeds by five evaluation seeds, 200 episodes per block). Open points show the five training-seed means after averaging the evaluation blocks; the larger model-shaped marker and bar denote the mean ± one standard deviation across those five independent training outcomes. Evaluation blocks are nested repeats, not additional independent training samples. No line is drawn between regimes. Lower is better for distance and first-contact time.
Jmse 14 01695 g005
Figure 6. Severe-regime attribution: (a) episode return, (b) mean USV–AUV distance, (c) fraction of time within 5 km, and (d) first-contact time conditional on acquisition. Small semi-transparent points show the 25 evaluation-block results per method (200 episodes per block), while open circles show the five training-seed means. Violin contours are descriptive kernel-density estimates of those five means, and diamonds with bars denote their mean ± SD. Constant, Prior, and Acoustic are equal-capacity observation-quality-token ablations. PPO-M and CA-M match UCA-PPO’s actor parameter count. The five evaluation repeats within a checkpoint are not treated as independent inferential units; inference rests on the five training outcomes and paired analyses rather than on violin width.
Figure 6. Severe-regime attribution: (a) episode return, (b) mean USV–AUV distance, (c) fraction of time within 5 km, and (d) first-contact time conditional on acquisition. Small semi-transparent points show the 25 evaluation-block results per method (200 episodes per block), while open circles show the five training-seed means. Violin contours are descriptive kernel-density estimates of those five means, and diamonds with bars denote their mean ± SD. Constant, Prior, and Acoustic are equal-capacity observation-quality-token ablations. PPO-M and CA-M match UCA-PPO’s actor parameter count. The five evaluation repeats within a checkpoint are not treated as independent inferential units; inference rests on the five training outcomes and paired analyses rather than on violin width.
Jmse 14 01695 g006
Figure 7. Representative trajectories under severe prior degradation ( T p = 60 decisions and σ p = 5 km): (a) PPO, (b) CA-PPO, and (c) UCA-PPO. All panels use training seed 0, evaluation seed 102, and episode 93. Black and blue curves denote the underwater-object and USV paths; green stars are prior updates, and red crosses mark passive-acoustic contacts. Panel headers report episode return R in thousands, mean modeled detection probability P d , and first-contact decision T f . Because the target reacts to the USV, the shared seed does not imply pointwise paired target trajectories.
Figure 7. Representative trajectories under severe prior degradation ( T p = 60 decisions and σ p = 5 km): (a) PPO, (b) CA-PPO, and (c) UCA-PPO. All panels use training seed 0, evaluation seed 102, and episode 93. Black and blue curves denote the underwater-object and USV paths; green stars are prior updates, and red crosses mark passive-acoustic contacts. Panel headers report episode return R in thousands, mean modeled detection probability P d , and first-contact decision T f . Because the target reacts to the USV, the shared seed does not imply pointwise paired target trajectories.
Jmse 14 01695 g007
Table 1. Single-frame actor observation. Target truth, true range, and latent modeled detection probability are excluded.
Table 1. Single-frame actor observation. Target truth, true range, and latent modeled detection probability are excluded.
BlockDim.EntriesScaling or Interpretation
USV context o t c 6 x u , y u , cos ψ u , sin ψ u , v u , Δ ψ u Position/100 km; speed and turn/limits
Acoustic o t a 5 d , sin β ^ , cos β ^ , c β , S E ¯ Contact; held bearing; confidence; tanh ( S E / 10 )
External prior o t p 5 m p , Δ x p , Δ y p , a p , σ p Update; relative position; age; stated error
Table 2. Main simulation and acoustic parameters.
Table 2. Main simulation and acoustic parameters.
ParameterValue
Area/decision interval 100 × 100 km/60 s
Episode horizon1000 decisions
USV/AUV speed5/2 ms−1
USV initialization ( x , y ) U ( [ 12 , 88 ] 2 ) km; ψ U [ π , π ) ; v = 0
AUV entry/exit1 km inside/5 km beyond opposite side
AUV lateral entry/exit jitter U [ 15 , 85 ] km/ N ( 0 , 8 2 ) km, clipped
Heading increment/slew limit10°/2° per decision
AUV avoidance range/weight15 km/25
Acoustic frequency100 Hz
Modeled depth pair150/200 m
RAM maximum range20 km
RAM azimuth/range/depth steps10°/100 m/50 m
S L , N L 0 , N L f , D T 130, 65, 50, 0 dB
Invalid TL/contact memory120 dB/50 decisions
Bearing-noise limits0.03–0.35 rad
Table 3. Learned configurations used to test UCA-PPO. Ablations and matched controls are evaluated only under severe prior degradation.
Table 3. Learned configurations used to test UCA-PPO. Ablations and matched controls are evaluated only under severe prior degradation.
ConfigurationWidthParametersRole
PPO25670,660Unstructured main baseline
CA-PPO256189,957Source-token attention baseline
UCA-PPO (uca)256207,685Proposed full observation-quality token
Constant (uca_zero)256207,685Quality variation removed
Prior-only256207,685Prior quality retained
Acoustic-only256207,685Acoustic quality retained
PPO-M446207,840Parameter-matched PPO control
CA-PPO-M275208,007Parameter-matched CA control
Table 4. Non-learning engineering baselines. All controllers use the same environment action limits as the learned policies and receive no simulator target truth.
Table 4. Non-learning engineering baselines. All controllers use the same environment action limits as the learned policies and receive no simulator target truth.
ControllerPosition PriorAcoustic ObservationDecision Rule
Greedy-PriorHeld positionNoneDistance-scaled direct pursuit
Prior-MPCHeld positionNoneEight-step discrete receding-horizon search
PF-PursuitPosition and stated errorContact, bearing, confidenceParticle-filter mean followed by pursuit
Table 5. Separately trained prior regimes.
Table 5. Separately trained prior regimes.
RegimeCodeInterval σ p
Accuratetp10_sig010 decisions0 km
Mildtp30_sig230 decisions2 km
Severetp60_sig560 decisions5 km
Table 6. Main-policy evaluation. Values are mean ± SD across five training seeds after averaging 1000 episodes within each seed. Return is reported in 10 3 ; D is mean distance, R 5 is the fraction of time within 5 km, C is contact continuity, A is the fraction of episodes with contact, and T f A is first-contact time conditional on acquisition.
Table 6. Main-policy evaluation. Values are mean ± SD across five training seeds after averaging 1000 episodes within each seed. Return is reported in 10 3 ; D is mean distance, R 5 is the fraction of time within 5 km, C is contact continuity, A is the fraction of episodes with contact, and T f A is first-contact time conditional on acquisition.
RegimeModelReturnD (km) R 5 P d ¯ CA T f A
AccuratePPO 421.5 ± 48.5 9.43 ± 0.55 0.789 ± 0.053 0.453 ± 0.083 0.571 ± 0.106 0.979 ± 0.044 177.5 ± 5.0
CA-PPO 348.2 ± 133.9 12.29 ± 6.35 0.642 ± 0.282 0.310 ± 0.155 0.407 ± 0.196 0.942 ± 0.108 219.8 ± 36.3
UCA-PPO 406.1 ± 33.9 10.06 ± 1.07 0.748 ± 0.053 0.423 ± 0.060 0.549 ± 0.070 0.997 ± 0.004 199.9 ± 27.0
MildPPO 399.6 ± 22.6 9.50 ± 0.31 0.743 ± 0.035 0.390 ± 0.044 0.515 ± 0.056 0.990 ± 0.010 211.8 ± 11.0
CA-PPO 405.8 ± 12.7 9.37 ± 0.19 0.744 ± 0.027 0.409 ± 0.026 0.532 ± 0.033 0.999 ± 0.001 199.0 ± 17.9
UCA-PPO 397.4 ± 19.8 9.49 ± 0.33 0.735 ± 0.036 0.399 ± 0.037 0.516 ± 0.047 0.999 ± 0.001 196.4 ± 16.9
SeverePPO 309.2 ± 8.7 11.52 ± 0.37 0.529 ± 0.018 0.276 ± 0.013 0.412 ± 0.015 0.830 ± 0.021 284.2 ± 12.8
CA-PPO 331.6 ± 45.5 11.39 ± 1.28 0.570 ± 0.097 0.314 ± 0.061 0.412 ± 0.077 0.929 ± 0.079 219.8 ± 32.3
UCA-PPO 337.7 ± 10.5 10.96 ± 0.24 0.583 ± 0.023 0.314 ± 0.016 0.405 ± 0.014 0.983 ± 0.016 209.5 ± 13.7
Table 7. Seed-paired severe-regime contrasts between UCA-PPO and PPO. Differences are oriented so that positive values favor UCA-PPO. CI is the 10,000-resample paired bootstrap interval; d z is paired Cohen’s effect size; wins count favorable training seeds. Exact p values are discrete and are not multiplicity-adjusted.
Table 7. Seed-paired severe-regime contrasts between UCA-PPO and PPO. Differences are oriented so that positive values favor UCA-PPO. CI is the 10,000-resample paired bootstrap interval; d z is paired Cohen’s effect size; wins count favorable training seeds. Exact p values are discrete and are not multiplicity-adjusted.
EndpointFavorable Difference95% CI d z WinsExact p
Return 28.47 × 10 3 [ 14.89 , 40.09 ] × 10 3 1.795/50.0625
Mean distance decrease (km)0.556 [ 0.153 , 0.958 ] 1.104/50.1250
5 km dwell0.0539 [ 0.0246 , 0.0806 ] 1.505/50.0625
Modeled P d 0.0376 [ 0.0185 , 0.0566 ] 1.505/50.0625
Contact continuity 0.0070 [ 0.0249 , 0.0097 ] 0.32 2/50.5625
Contact coverage0.1526 [ 0.1374 , 0.1728 ] 6.675/50.0625
Conditional first-contact decrease74.8 [ 56.5 , 90.6 ] 3.495/50.0625
Mean-turn decrease (deg) 0.363 [ 0.479 , 0.258 ] 2.59 0/50.0625
Table 8. Engineering-baseline evaluation, reported as mean ± SD over five environment-seed blocks of 200 episodes. Return is in 10 3 . T f is the penalized first-contact time, with 1000 assigned to an episode without contact.
Table 8. Engineering-baseline evaluation, reported as mean ± SD over five environment-seed blocks of 200 episodes. Return is in 10 3 . T f is the penalized first-contact time, with 1000 assigned to an episode without contact.
RegimeControllerReturnD (km) R 5 P d ¯ CA T f T f A
AccurateGreedy-Prior 219.8 ± 2.9 10.51 ± 0.23 0.782 ± 0.004 0.052 ± 0.004 0.064 ± 0.005 0.819 ± 0.033 322.1 ± 24.6 172.2 ± 5.1
Prior-MPC 337.3 ± 1.8 8.15 ± 0.43 0.800 ± 0.008 0.0426 ± 0.0003 0.0576 ± 0.0004 1.000 ± 0.000 229.3 ± 4.8 229.3 ± 4.8
PF-Pursuit 290.2 ± 1.6 10.20 ± 0.18 0.792 ± 0.005 0.177 ± 0.003 0.225 ± 0.003 1.000 ± 0.000 186.9 ± 4.7 186.9 ± 4.7
MildGreedy-Prior 170.3 ± 1.3 12.51 ± 0.30 0.061 ± 0.005 0.043 ± 0.001 0.056 ± 0.002 0.999 ± 0.002 210.8 ± 7.0 210.0 ± 6.4
Prior-MPC 252.6 ± 3.5 11.08 ± 0.38 0.517 ± 0.011 0.083 ± 0.001 0.110 ± 0.001 0.999 ± 0.002 219.7 ± 7.5 218.9 ± 7.7
PF-Pursuit 337.2 ± 3.2 9.84 ± 0.25 0.630 ± 0.014 0.246 ± 0.004 0.315 ± 0.007 0.999 ± 0.002 207.5 ± 1.9 206.7 ± 2.0
SevereGreedy-Prior 118.1 ± 1.5 15.05 ± 0.23 0.050 ± 0.004 0.016 ± 0.001 0.026 ± 0.005 0.879 ± 0.023 368.4 ± 15.8 281.3 ± 17.2
Prior-MPC 143.1 ± 1.7 14.95 ± 0.19 0.145 ± 0.005 0.028 ± 0.001 0.046 ± 0.002 0.982 ± 0.016 312.1 ± 19.4 299.6 ± 10.4
PF-Pursuit 214.3 ± 10.5 12.11 ± 0.37 0.339 ± 0.032 0.073 ± 0.007 0.100 ± 0.008 0.858 ± 0.016 379.2 ± 26.3 276.7 ± 17.5
Table 9. Severe-regime ablation and parameter-matched evaluation. Reporting follows Table 6.
Table 9. Severe-regime ablation and parameter-matched evaluation. Reporting follows Table 6.
ConfigurationReturnD (km) R 5 P d ¯ CA T f A
PPO 309.2 ± 8.7 11.52 ± 0.37 0.529 ± 0.018 0.276 ± 0.013 0.412 ± 0.015 0.830 ± 0.021 284.2 ± 12.8
CA-PPO 331.6 ± 45.5 11.39 ± 1.28 0.570 ± 0.097 0.314 ± 0.061 0.412 ± 0.077 0.929 ± 0.079 219.8 ± 32.3
UCA-PPO 337.7 ± 10.5 10.96 ± 0.24 0.583 ± 0.023 0.314 ± 0.016 0.405 ± 0.014 0.983 ± 0.016 209.5 ± 13.7
Constant token 331.3 ± 29.6 10.85 ± 0.43 0.574 ± 0.053 0.306 ± 0.061 0.401 ± 0.069 0.954 ± 0.063 221.3 ± 27.2
Prior-only 313.2 ± 14.0 11.29 ± 0.28 0.540 ± 0.024 0.277 ± 0.024 0.368 ± 0.024 0.965 ± 0.027 232.3 ± 23.0
Acoustic-only 299.3 ± 75.5 11.85 ± 2.21 0.501 ± 0.159 0.258 ± 0.103 0.340 ± 0.130 0.943 ± 0.058 229.6 ± 30.1
PPO-M 248.8 ± 58.7 12.96 ± 1.51 0.409 ± 0.128 0.172 ± 0.072 0.227 ± 0.093 0.931 ± 0.034 227.5 ± 14.3
CA-PPO-M 249.3 ± 76.9 14.50 ± 4.74 0.403 ± 0.153 0.192 ± 0.085 0.275 ± 0.099 0.801 ± 0.165 274.5 ± 48.3
Table 10. Frozen-checkpoint 3 × 3 factorial evaluation over 135,000 episodes. Entries are mean ± SD across five independent training seeds after averaging over five evaluation seeds within each training seed. Return is in 10 3 , and A denotes contact-acquisition success.
Table 10. Frozen-checkpoint 3 × 3 factorial evaluation over 135,000 episodes. Entries are mean ± SD across five independent training seeds after averaging over five evaluation seeds within each training seed. Return is in 10 3 , and A denotes contact-acquisition success.
I p σ p (km)PPO ReturnCA-PPO ReturnUCA-PPO ReturnUCA-PPO D ¯ (km)UCA-PPO P ¯ d UCA-PPO A
100 387.7 ± 13.2 387.6 ± 32.1 377.2 ± 19.9 10.02 ± 0.43 0.375 ± 0.031 0.994 ± 0.013
102 392.6 ± 9.7 390.2 ± 32.1 390.7 ± 17.0 9.80 ± 0.39 0.396 ± 0.029 0.996 ± 0.008
105 388.1 ± 8.2 382.1 ± 33.0 388.7 ± 10.6 9.75 ± 0.26 0.388 ± 0.020 1.000 ± 0.001
300 390.1 ± 5.4 385.2 ± 31.5 372.0 ± 17.9 10.00 ± 0.39 0.368 ± 0.029 0.992 ± 0.019
302 391.3 ± 5.9 384.6 ± 31.5 385.1 ± 13.4 9.81 ± 0.34 0.388 ± 0.025 0.997 ± 0.004
305 359.4 ± 6.6 362.0 ± 39.0 370.0 ± 6.6 10.04 ± 0.21 0.360 ± 0.013 0.996 ± 0.004
600 371.3 ± 7.1 366.5 ± 35.1 355.3 ± 21.3 10.23 ± 0.46 0.344 ± 0.032 0.995 ± 0.011
602 367.8 ± 9.6 365.5 ± 38.9 370.7 ± 15.0 10.08 ± 0.40 0.368 ± 0.023 0.996 ± 0.007
605 310.7 ± 7.6 331.9 ± 48.2 342.7 ± 13.5 10.66 ± 0.41 0.322 ± 0.019 0.983 ± 0.018
Table 11. Zero-shot transfer of frozen severe-regime checkpoints to a Munk-profile RAM database over the same terrain. Each acoustic field contains 15,000 episodes; values are the mean ± SD across five training seeds after averaging five evaluation seeds per seed, and return is in 10 3 .
Table 11. Zero-shot transfer of frozen severe-regime checkpoints to a Munk-profile RAM database over the same terrain. Each acoustic field contains 15,000 episodes; values are the mean ± SD across five training seeds after averaging five evaluation seeds per seed, and return is in 10 3 .
Acoustic FieldPolicyReturn D ¯ (km) P ¯ d A T f A
WOA-derivedPPO 309.2 ± 8.7 11.52 ± 0.37 0.276 ± 0.013 0.830 ± 0.021 284.2 ± 12.8
WOA-derivedCA-PPO 331.6 ± 45.5 11.39 ± 1.28 0.314 ± 0.061 0.929 ± 0.079 219.8 ± 32.3
WOA-derivedUCA-PPO 337.7 ± 10.5 10.96 ± 0.24 0.314 ± 0.016 0.983 ± 0.016 209.5 ± 13.7
Munk profilePPO 139.8 ± 3.7 14.11 ± 0.28 0.007 ± 0.002 0.475 ± 0.088 403.5 ± 26.5
Munk profileCA-PPO 161.8 ± 32.6 14.57 ± 1.41 0.027 ± 0.015 0.868 ± 0.112 237.7 ± 44.6
Munk profileUCA-PPO 147.5 ± 4.7 14.68 ± 0.82 0.017 ± 0.003 0.944 ± 0.029 226.1 ± 18.8
Table 12. Severe-regime UCA-PPO retraining with and without the privileged acoustic reward. Both rows are evaluated under the same common evaluator. Values are the mean ± SD across five training seeds after averaging five evaluation seeds per training seed; return is in 10 3 , L is path length, and Θ is cumulative absolute turn.
Table 12. Severe-regime UCA-PPO retraining with and without the privileged acoustic reward. Both rows are evaluated under the same common evaluator. Values are the mean ± SD across five training seeds after averaging five evaluation seeds per training seed; return is in 10 3 , L is path length, and Θ is cumulative absolute turn.
Training RewardReturn D ¯ (km)L (km) R 5 A T f A Θ (deg)
Original 342.7 ± 13.5 10.66 ± 0.41 132.68 ± 3.72 0.600 ± 0.032 0.983 ± 0.018 205.8 ± 18.9 1157.9 ± 123.4
No acoustic reward 241.2 ± 53.7 11.54 ± 2.54 165.10 ± 8.65 0.447 ± 0.154 0.966 ± 0.064 222.3 ± 34.3 2514.9 ± 705.0
Table 13. Additional metrics recovered from completed learned-policy evaluations (mean ± SD across five training seeds).
Table 13. Additional metrics recovered from completed learned-policy evaluations (mean ± SD across five training seeds).
RegimePolicyPath Length L (km)Cumulative Turn Θ (deg)Acquisition Success A T f ( 1000 )
AccuratePPO 124.25 ± 5.12 791.6 ± 86.0 0.979 ± 0.044 194.5 ± 34.7
AccurateCA-PPO 132.41 ± 25.26 2291.7 ± 1318.6 0.942 ± 0.108 262.1 ± 111.8
AccurateUCA-PPO 129.79 ± 7.26 1186.6 ± 519.7 0.997 ± 0.004 202.3 ± 29.3
MildPPO 126.82 ± 2.94 917.4 ± 90.3 0.990 ± 0.010 219.3 ± 16.9
MildCA-PPO 128.72 ± 4.96 1223.5 ± 105.5 0.999 ± 0.001 199.6 ± 18.4
MildUCA-PPO 129.09 ± 2.67 1305.6 ± 257.8 0.999 ± 0.001 197.2 ± 17.6
SeverePPO 132.66 ± 1.34 853.6 ± 21.4 0.830 ± 0.021 405.8 ± 16.9
SevereCA-PPO 134.96 ± 3.50 1202.7 ± 121.5 0.929 ± 0.079 274.5 ± 78.0
SevereUCA-PPO 134.30 ± 4.30 1163.1 ± 128.7 0.983 ± 0.016 222.9 ± 25.4
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Hu, J.; Zhao, R.; Chen, L.; Tu, X.; Chen, J.; Dong, C.; Liu, F. UCA-PPO: USV Path Planning for Moving Underwater Target Tracking with Passive Acoustic Observations and Position Priors. J. Mar. Sci. Eng. 2026, 14, 1695. https://doi.org/10.3390/jmse14181695

AMA Style

Hu J, Zhao R, Chen L, Tu X, Chen J, Dong C, Liu F. UCA-PPO: USV Path Planning for Moving Underwater Target Tracking with Passive Acoustic Observations and Position Priors. Journal of Marine Science and Engineering. 2026; 14(18):1695. https://doi.org/10.3390/jmse14181695

Chicago/Turabian Style

Hu, Jian, Rongyao Zhao, Lu Chen, Xiaoyu Tu, Jiaru Chen, Chengfeng Dong, and Feng Liu. 2026. "UCA-PPO: USV Path Planning for Moving Underwater Target Tracking with Passive Acoustic Observations and Position Priors" Journal of Marine Science and Engineering 14, no. 18: 1695. https://doi.org/10.3390/jmse14181695

APA Style

Hu, J., Zhao, R., Chen, L., Tu, X., Chen, J., Dong, C., & Liu, F. (2026). UCA-PPO: USV Path Planning for Moving Underwater Target Tracking with Passive Acoustic Observations and Position Priors. Journal of Marine Science and Engineering, 14(18), 1695. https://doi.org/10.3390/jmse14181695

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop