Next Article in Journal
Research on Safety Assurance Strategies for Offshore Transfer Operations Based on Floating Hose State Prediction
Next Article in Special Issue
Adaptive Fixed-Time Anti-Saturation Nonsingular Terminal Sliding Mode Control for Multi-AUV Formation Tracking Under Actuator Saturation and Lumped Disturbances
Previous Article in Journal
Practitioner-Informed AI Decision Support for Maritime Accident-Type Risk in Korean Waters
Previous Article in Special Issue
Research on Cooperative Pursuit Strategy of Multiple UUVs Based on Deep Reinforcement Learning in Complex Dynamic Environments
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Robust Adaptive Propagated Interval Observer for Actuator Fault Diagnosis in Underactuated AUVs

1
College of Intelligent Systems Science and Engineering, Harbin Engineering University, Harbin 150001, China
2
Key Laboratory of Intelligent Technology and Application of Marine Equipment, Ministry of Education, Harbin 150001, China
3
Computer and Network Engineering Department, College of Computing, Umm Al-Qura University, Mecca 24231, Saudi Arabia
4
School of Engineering, Nanfang College Guangzhou, Guangzhou 510970, China
*
Author to whom correspondence should be addressed.
J. Mar. Sci. Eng. 2026, 14(15), 1445; https://doi.org/10.3390/jmse14151445
Submission received: 29 June 2026 / Revised: 31 July 2026 / Accepted: 4 August 2026 / Published: 6 August 2026
(This article belongs to the Special Issue Design and Application of Underwater Vehicles—2nd Edition)

Abstract

This paper presents an interval-observer-based actuator fault detection and isolation (FDI) method for underactuated autonomous underwater vehicles (AUVs) under bounded hydrodynamic uncertainty and time-varying ocean currents. A locally frozen linear time-invariant (LTI) representation enables deterministic set-membership analysis, and the robust adaptive propagated interval observer (RAPIO) propagates admissible center–radius state bounds within a Lyapunov framework. Adaptivity is introduced through a reinforcement learning (RL)-augmented uncertainty-bound modulation mechanism, where an offline-trained agent scales a nonnegative channel-wise slack term without modifying the scheduled observer-gain rule or the nominal center predictor. Under the stated observer and disturbance-envelope conditions, positivity, stability, and diagnostic-channel inclusion hold for any bounded learning signal. Actuator loss-of-effectiveness (LoE) faults are represented through the actuator-effectiveness channel and detected through interval-consistency violations, enabling axis-wise isolation of surge, yaw-rate, and pitch-rate actuator faults. The same schedule-blind decision layer is additionally evaluated with structurally distinct additive-bias and stuck/jam actuator models. All stuck/jam events are detected, and bias-magnitude sweeps identify channel-wise 100%-detection boundaries with zero false alarms. A structured 72-case scenario sweep shows reliable detection, strong false-alarm rejection, and acceptable detection delays compared with benchmark observers.

1. Introduction

AUVs have become essential platforms for deep-sea surveys, offshore inspections, and environmental monitoring, where human access is difficult or impossible [1]. However, the performance of these platforms can degrade under strong ocean disturbances and hydrodynamic parameter uncertainties because of the nonlinear added-mass and damping coefficients [2]. Furthermore, in underactuated configurations, limited redundancy and nonlinear coupling amplify these vulnerabilities, and actuator faults such as loss of effectiveness, bias, or jamming can lead to mission failure [3]. Therefore, autonomous, long-duration AUV missions require an integrated fault-tolerant control architecture with a reliable fault detection and isolation (FDI) module [4,5].
In the FDI literature, model-based approaches remain the preferred choice for safety-critical marine applications due to their analytical transparency and formal guarantees [6]. Among these, observer-based techniques are particularly attractive due to their minimal computational demand and structural simplicity [7]. Nevertheless, classical observers such as extended Kalman filters (EKF) and fixed-threshold residual generators are often sensitive to modeling mismatch and external disturbances, which may result in unreliable decisions [8,9].
Set-membership and interval-based observers address these limitations by treating violations of admissible state bounds as deterministic fault indicators [10,11]. Unlike stochastic filters, interval observers propagate bounded uncertainty directly and detect faults through interval-consistency violations [10,12]. For nonlinear AUV dynamics, locally frozen LTI and LPV-inspired representations enable tractable observer synthesis while retaining operating-point dependence [13,14,15]. Similar linearization strategies have also been used for underwater, aerial, and ground-vehicle FDI [16]. Related set-membership developments include ellipsoidal reachability for nonlinear robust FDI [17], unknown-input interval observers [18], and distributed interval/unknown-input observers [19]. However, these methods still rely on fixed or structurally prescribed uncertainty descriptions, limiting their response to time-varying hydrodynamic conditions. Because worst-case uncertainty bounds are selected offline, this fixed-bound assumption—common in interval-observer FDI—can reduce detection performance when operating conditions change [20]. Zonotopic set-membership FDI under event-triggered communication shows the value of bounded-set estimation in multi-agent systems [21], but it still assumes static or slowly varying bound structures. Zhu et al. [22] reduce conservatism using a fixed-gain LPV interval observer bank with offline H / H LMI optimization and an open-loop boundary update indexed by a slowly varying scheduling parameter. This improves mild operating-point variation, but the update does not use runtime residual feedback and can therefore mismatch randomized hydrodynamic disturbances outside the scheduling grid. Convex/polytopic LPV observers similarly offer certificate-oriented synthesis, but their prescribed scheduling structure and direct fault-reconstruction objective do not address residual-adaptive bounded-set consistency under time-varying AUV hydrodynamic and current uncertainty [23].
Stochastic robust filters provide another route. The robust EKF of Song and He [24] models environmental perturbations as multiplicative process noise and evaluates a normalized innovation against a calibrated scalar threshold, but one design-time threshold cannot both suppress Coriolis-induced cross-channel residuals in healthy operation and retain sensitivity to small actuator faults. Online sign-based adaptive interval observers [25] replace conservative constants with sign-based adaptive parameters driven by measurement consistency. Their interval width adapts online, but the sign-based law can lag rapidly varying Coriolis-coupled disturbances, increasing detection delay and false alarms under strong hydrodynamic uncertainty.
To provide greater adaptability, data-driven and reinforcement-learning (RL) approaches have been explored [26,27]. However, when learning is embedded directly into observer gains or residual thresholds, interval inclusion and stability may be compromised. Any resulting guarantee may hold only after the policy has converged and while operation remains within its training distribution, a condition that cannot be assured during deployment in unknown marine environments [28].
The observer design problem has also been studied extensively for fractional-order systems. Observer-based model-reference control for linear fractional-order systems has been developed using Caputo derivatives, Lyapunov analysis, and LMI tools [29,30]. Nonlinear extensions include Hadamard fractional-order one-sided Lipschitz observer/controller design [31], unknown-input observer synthesis for generalized proportional fractional-order systems [32], and practical observers for state and sensor-fault reconstruction [33]. These works address fractional-order dynamics beyond the present integer-order AUV model. Even so, their rigorous observer architectures support the broader principle used here: the learning or adaptation mechanism should be separated from the core observer dynamics to preserve formal guarantees under uncertainty.
These studies leave four linked limitations. Fixed-bound interval observers [16,20] cannot react to real-time residual feedback outside the offline design envelope. LPV/ H / H adaptive methods [22] use open-loop scheduling whose certificates do not cover residual-feedback or learned modulation signals. Stochastic robust filters [24] and sign-based adaptive observers [25] rely on scalar thresholds or lagging adaptation laws. As a result, it is difficult to reject channel-specific false alarms while retaining sensitivity to small faults. RL-assisted methods [26,28] close the adaptation loop, but often place the policy inside the observer core. This makes stability and inclusion conditional on policy convergence.
Accordingly, AUV FDI requires residual-feedback adaptation. Such adaptation should preserve formal inclusion and stability independently of learning convergence. It should also provide channel-specific alarm modulation without changing the observer gains. The novelty of RAPIO is the separation between the certified observer and the learned adaptation signal. Unlike fixed or adaptive interval/set-membership observers whose bounds are prescribed offline, RAPIO updates the admissible tube using runtime residual information. The update is also not tied to an offline scheduling grid, as it would be under LPV/ H / H boundary scheduling. Crucially, it does not alter the scheduled observer-gain rule, nominal center predictor, or residual maps directly. The observer gain-scheduling rule remains fixed, so analytical stability and interval inclusion remain independent of policy convergence under the stated observer and disturbance-envelope assumptions, including out-of-distribution bounded policy outputs.
The main contributions of this work are as follows:
  • An RL-augmented RAPIO architecture is developed in which the frozen SAC policy modulates only bounded tube slack; the observer gains, center predictor, and detector thresholds remain fixed.
  • A locally frozen LTI interval observer is formulated for the underactuated AUV, with Metzler/Hurwitz certification and componentwise inclusion on the diagnostic channels I d = { u , r , q } .
  • An interval-consistency rule combines leaky bound-excess accumulation, confirmation/recovery gates, and center freezing for surge, yaw, and pitch isolation. The same schedule-blind detector identifies additive bias and stuck/jam faults without metadata.
  • A structured 72-case scenario sweep benchmarks RAPIO against an adaptive interval observer and a P2P threshold FDI method, showing 100 % detection on all channels with zero event-level false alarms.
The remainder of this paper is organized as follows. Section 2 gives the formulation of the actuator FDI problem and introduces the system model. Section 3 presents the interval-observer-based FDI approach. It includes observer design, stability properties, and axis-wise fault isolation. Section 4 gives the details of the simulation setup and Section 5 reports the results and comparative analyses. Section 6 discusses the principal findings and trade-offs. Finally, Section 7 concludes the paper.

2. Problem Formulation

To formulate robust interval-based fault diagnosis under bounded uncertainty, we first state the assumptions governing the AUV dynamics and operating environment.
Assumption 1
([34,35]). The AUV dynamics satisfy the nonlinear fault-affected model given in (1).
x ˙ ( t ) = f x ( t ) , u ( t ) + d ( t ) + F f a ( t ) ,
where x ( t ) R 12 is the system state, u ( t ) R 3 is the augmented actuator command, f ( · ) is locally Lipschitz, d ( t ) denotes bounded disturbances encompassing ocean current effects and hydrodynamic uncertainty, and f a ( t ) represents unknown additive actuator fault effects entering the state dynamics through a known distribution matrix F R 12 × 3 . In this work, the augmented command is u = [ δ s , δ r , n ] , where the TNMPC optimizes the stern-plane and rudder deflections [ δ s , δ r ] R 2 and a separate RPM loop generates the propeller speed n. Throughout the manuscript, n x : = 12 denotes the state dimension. The unbold symbol n in the actuator command is reserved for propeller rotational speed (RPM).
Assumption 2.
Let ρ ( t ) R n ρ denote the vector of scheduling variables (a measurable subset of the state and/or input, e.g., ρ ( t ) = ρ ( x ( t ) , u ( t ) ) ) used to parameterize the locally frozen linear model matrices A ( ρ ( t ) ) and B ( ρ ( t ) ) . The scheduling variables satisfy
ρ ( t ) ρ ( t k ) ε ρ , t [ t k , t k + 1 ) ,
for some ε ρ > 0 .
Moreover, the system operates inside compact admissible sets X c R 12 and U c R 3 such that x ( t ) X c and u ( t ) U c for all admissible missions. The admissible sampled command profile is bounded and rate-limited on each diagnostic interval. Saturations, command holds, and actuator-mode changes are treated as sampled inputs. Their residual effect on the frozen model is included in the calibrated disturbance budget. Let z ( t ) = [ x ( t ) , u ( t ) ] . The diagnostic sampling instants satisfy t k + 1 = t k + T s , where T s > 0 is the diagnostic update period. Over Z c = X c × U c , the nonlinear dynamics f ( x , u ) are twice continuously differentiable and there exists a constant L Δ > 0 such that
Δ ( t ) L Δ z ( t ) z ( t k ) 2 , t [ t k , t k + 1 ) ,
where Δ ( t ) denotes the higher-order Taylor remainder. The compact-set condition is required only in the fault-free case; during a fault the true state may exit the interval, which is precisely the detection mechanism.
Remark 1.
Assumption 2 defines a local validity envelope for the frozen model, not a global linear approximation. On the compact set Z c , bounded and rate-limited admissible trajectories admit finite rate bounds v ¯ x , v ¯ u , v ¯ ρ , defined by x ˙ ( t ) v ¯ x , u ˙ ( t ) v ¯ u , and ρ ˙ ( t ) v ¯ ρ on Z c , giving ρ ( t ) ρ ( t k ) v ¯ ρ T s and z ( t ) z ( t k ) 2 ( v ¯ x 2 + v ¯ u 2 ) T s 2 . Substituting into (3) and letting r ¯ i = d ˜ ¯ i d ¯ i phys > 0 denote the available componentwise residual margin, a sufficient sampling condition for keeping the Taylor remainder inside the disturbance envelope is
T s min ε ρ v ¯ ρ , min i r ¯ i L Δ ( v ¯ x 2 + v ¯ u 2 ) .
Here, d ¯ i phys is the componentwise bound allocated to the physical disturbance and frozen-point offset before the Taylor-remainder budget is added, and d ˜ ¯ i is the corresponding aggregate bound defined in (21). If (4) is violated in a new operating regime, inclusion must be restored by reducing T s , refreezing the Jacobian more frequently, or enlarging d ˜ ¯ .
Assumption 3.
The disturbance d ( t ) satisfies the componentwise bounds in (5):
d ̲ d ( t ) d ¯ , t 0 ,
where d ̲ , d ¯ R 12 are known constant vectors corresponding to the full velocity–kinematic state dimension. In practice, the velocity-channel bounds are derived from ocean-current models. In the simulations of Section 4, currents up to 0.3  m/s enter through M 1 τ d . The kinematic states are not directly forced in the lumped disturbance representation of (12); their uncertainty is propagated through the velocity and attitude dynamics.

2.1. Nonlinear AUV Dynamics with Actuator Faults

The AUV motion is described in a body-fixed frame { b } and an inertial North–East–Down (NED) frame { n } , as shown in Figure 1. Body-fixed velocities u, v, and w denote surge, sway, and heave, respectively, and p, q, and r denote roll, pitch, and yaw rates. The inertial coordinates x, y, and z denote North, East, and Down position, while ϕ , θ , and ψ are the roll, pitch, and yaw Euler angles, respectively. The vehicle is actuated in surge by propulsion force τ u , in yaw by control moment τ r , and in vertical motion by stern-plane force τ w , which affects both heave and pitch. Ocean current effects are modeled as an unknown bounded disturbance d ( t ) consistent with Assumption 3.
The body-fixed velocity and inertial position–orientation vectors are defined in (6):
ν = u v w p q r , η = x y z ϕ θ ψ .
The nonlinear 6-DOF AUV dynamics are
M ( ν ) ν ˙ + C ( ν ) ν + D ( ν ) ν + g ( η ) = τ + τ d , η ˙ = J ( η ) ν ,
The kinematic transformation in (7) maps the body-fixed translational and angular velocities into NED position and Euler-angle rates. For the adopted 3–2–1 yaw–pitch–roll convention, the transformation is given by (8):
J ( η ) = R b n ( ϕ , θ , ψ ) 0 3 × 3 0 3 × 3 T ( ϕ , θ ) ,
where the blocks in (9) use c χ = cos χ and s χ = sin χ :
R b n = c ψ c θ c ψ s θ s ϕ s ψ c ϕ c ψ s θ c ϕ + s ψ s ϕ s ψ c θ s ψ s θ s ϕ + c ψ c ϕ s ψ s θ c ϕ c ψ s ϕ s θ c θ s ϕ c θ c ϕ , T = 1 s ϕ tan θ c ϕ tan θ 0 c ϕ s ϕ 0 s ϕ / c θ c ϕ / c θ ,
where, R b n R 3 × 3 rotates a vector from the body frame to NED, and T R 3 × 3 maps the body angular-rate vector [ p , q , r ] to [ ϕ ˙ , θ ˙ , ψ ˙ ] . The Euler representation is used over the regular operating set | θ | < π / 2 , for which c θ 0 and T is nonsingular. In (7), M ( ν ) is the inertia matrix including added-mass effects, C ( ν ) denotes Coriolis and centripetal terms, D ( ν ) represents nonlinear hydrodynamic damping, g ( η ) collects the hydrostatic restoring forces and moments, τ R 6 is the generalized control input, and τ d R 6 represents external disturbances.
The underactuated actuator map is nonlinear in the augmented command and is written as (10):
τ = τ nom ( x , u ) + τ f ( x , u , α ) ,
where u = [ δ s , δ r , n ] collects stern-plane deflection, rudder deflection, and propeller RPM, and α = [ α u , α s , α r ] denotes the actuator effectiveness factors for surge propulsion, stern planes, and rudder. The additive term τ f is therefore an equivalent local force/moment difference induced by the multiplicative LoE model used by the plant, not an independently injected state offset.
Expanding (7) yields the componentwise velocity dynamics in (11):
ν ˙ i = 1 m i F i ( ν ) + τ i + τ i f + d i ,
where m i is the effective inertia (including added mass), F i ( ν ) aggregates the nonlinear hydrodynamics, τ i and τ i f are the nominal and fault-induced inputs, and d i represents the bounded disturbance in channel i. The index i { 1 , , 6 } follows the generalized force ordering (surge, sway, heave, roll, pitch, yaw).
For interval-based diagnosis, the physical disturbance is represented as (12),
d ( t ) = M 1 ( ν ( t ) ) τ d ( t ) 0 6 × 1 R 12 ,
where the lower zero block reflects that the kinematic states are not directly driven by external forces. Ocean current effects appearing in η ˙ = J ( η ) ν enter through the relative-velocity terms embedded in C ( ν ) and D ( ν ) .
Combining the above yields the compact nonlinear state-space model
x ˙ ( t ) = f ( x ( t ) , u ( t ) ) + d ( t ) + F phys ( x , u ) f a ( t ) ,
with state ordering x = [ u , v , w , p , q , r , x , y , z , ϕ , θ , ψ ] . Now, at the physical actuator level, let f a = [ Δ T , Δ Z s , Δ Y r ] collect the fault-induced changes in propeller thrust, stern-plane vertical force, and rudder lateral force. For the actuator geometry used in the nonlinear plant,
τ f = B f eq f a , B f eq = 1 0 0 0 0 1 0 1 0 0 0 0 0 L s 0 0 0 L r * ,
where stern-plane LoE gives Δ Z s = ( α s 1 ) Z u u d s u 2 δ s and Δ M s = ( α s 1 ) Z u u d s u 2 δ s L s , rudder LoE gives Δ Y r = ( α r 1 ) Y u u d r u 2 δ r and Δ N r = ( α r 1 ) N u u d r u 2 δ r L r with L r * = ( N u u d r / Y u u d r ) L r in the adopted coefficient set. The propeller RPM LoE gives Δ T = ( α u 2 1 ) K T n 2 because the plant uses T = K T ( α u n ) 2 . Here K T is the propeller thrust coefficient; Z u u d s , Y u u d r , and N u u d r are the stern-plane vertical-force, rudder lateral-force, and rudder yaw-moment coefficients; and L s and L r are the associated moment arms. The effective ratio L r * makes the equivalent force map reproduce the yaw moment of the implemented hydrodynamic model.
The AUV structure and actuator topology are adapted from the configuration reported in [36], with one propeller drive and mechanically linked stern-plane and rudder pairs, each commanded as one actuator. Consequently, the physically realized thrust and surface deflections satisfy
T act = K T ( α u n ) 2 , δ s , L act = δ s , R act = α s δ s , δ r , U act = δ r , D act = α r δ r , 0 α u , α s , α r 1 ,
where L / R and U / D denote the linked pitch- and yaw-fin pairs. These paired surfaces give three commanded drives associated with the { u , q , r } diagnostic channels. Roll remains part of the coupled 6-DOF motion, but it has no independent actuator in this topology.
The physical actuator-to-state-derivative distribution is defined in (15):
F phys ( x , u ) = M 1 ( x ) B f eq 0 6 × 3 ,
where the scalar entries of f a are evaluated from the effectiveness factors and commanded actuators above. The lower zero block excludes direct excitation of the kinematic states. Because M 1 contains inertial coupling, the upper part of a physical column need not have only one nonzero entry. For compactness, the locally frozen value of F phys ( x , u ) is denoted by F in the subsequent observer analysis.
To represent actuator command-path faults at the plant interface, let a cmd = [ n , δ s , δ r ] and a act denote the commanded and physically realized propeller speed, stern-plane angle, and rudder angle. The nonlinear plant applies LoE, additive-bias, and stuck/jam realizations through
a j act ( k ) = α j a j cmd ( k ) , LoE , 0 α j 1 , sat A j a j cmd ( k ) + b j , additive calibration bias , sat A j a j , stuck / jammed actuator ,
where b j is an additive offset, a j is a fixed jam value, and sat A j enforces the actuator limits. The RAPIO predictor continues to use a cmd and the nominal actuator map, so the resulting deviations are diagnosed from persistent interval inconsistency rather than supplied fault labels.
For diagnosis, the resulting dominant measured-state signatures are associated with the surge, yaw-rate, and pitch-rate channels. Equivalently, the diagnostic support can be written as
G d = e 1 e 6 e 5 , f sig ( k ) = q u ( k ) q r ( k ) q q ( k ) ,
where e i denotes the i-th canonical basis vector in R 12 . This notation is used only to identify the monitored diagnostic channels: surge speed u, yaw rate r, and pitch rate q. The scalars q u ( k ) , q r ( k ) , and q q ( k ) denote the fault-induced diagnostic signatures in those three monitored channels at sample k; they are not additional physical inputs. The underlying simulation applies the selected LoE, bias, or stuck/jam realization at the actuator interface of the nonlinear 6-DOF plant. The resulting perturbation then propagates through the same inertial, hydrodynamic, rotational, and kinematic couplings.

2.2. Locally Frozen Linear Model for Diagnosis

To enable interval-observer-based FDI for the nonlinear dynamics (13), a locally frozen LPV representation treated as piecewise LTI over each sampling interval is adopted. This representation is obtained via first-order Jacobian linearization and is used exclusively within the observer. The nominal center is propagated with the nonlinear diagnostic integrator under α = 1 and the recorded actuator commands, while the locally frozen Jacobian is used to schedule the RAPIO gain, certify the closed-loop cooperative structure, and propagate the interval half-width.
A first-order Taylor expansion about the operating point x ( t k ) , u ( t k ) gives (18):
f x ( t ) , u ( t ) = f x ( t k ) , u ( t k ) + A ( ρ ( t k ) ) x ( t ) x ( t k ) + B ( ρ ( t k ) ) u ( t ) u ( t k ) + Δ ( t ) ,
where A ( ρ ( t k ) ) and B ( ρ ( t k ) ) are the Jacobian matrices of f ( · ) with respect to x and u , respectively, and Δ ( t ) denotes the higher-order linearization residual. For implementation, their dimensions and definitions are made explicit in (19):
A k : = f x ( x ( t k ) , u ( t k ) ) R n x × n x , B k : = f u ( x ( t k ) , u ( t k ) ) R n x × 3 ,
with n x = 12 , x R n x , u R 3 , and Δ ( t ) R n x . Thus, every term in the Taylor expansion is an n x -vector. The notation A k = A ( ρ ( t k ) ) and B k = B ( ρ ( t k ) ) is used below. Under Assumption 2, the Jacobian matrices are treated as constant on [ t k , t k + 1 ) , yielding the locally frozen representation. Because the actuator command is stored and replayed over each diagnostic step, the width update uses the state Jacobian A ( ρ ( t k ) ) and treats command variation and integration mismatch as part of the calibrated disturbance budget.
Absorbing Δ ( t ) into the disturbance term gives the locally frozen LTI diagnostic model
x ˙ ( t ) = A ( ρ ( t k ) ) x ( t ) + B ( ρ ( t k ) ) u ( t ) + d ˜ ( t ) + F f a ( t ) ,
where the aggregated disturbance d ˜ ( t ) = d ( t ) + Δ ( t ) + ξ k ( t ) is unknown but bounded, with ξ k ( t ) = f ( x ( t k ) , u ( t k ) ) A ( ρ ( t k ) ) x ( t k ) B ( ρ ( t k ) ) u ( t k ) denoting the constant linearization offset at the frozen operating point. The remaining quantities in (20) satisfy d ˜ , ξ k R n x , F R n x × 3 , and f a R 3 . Consequently, A k x , B k u , and F f a all belong to R n x ; no scalar–matrix interpretation is required. Under Assumptions 1 and 2, d ˜ ( t ) is bounded componentwise: there exist known vectors d ˜ ̲ and d ˜ ¯ such that
d ˜ ̲ d ˜ ( t ) d ˜ ¯ , t [ t k , t k + 1 ) ,
where d ˜ ( t ) : = d ( t ) + Δ ( t ) + ξ k ( t ) includes the physical disturbance, the higher-order Taylor remainder, and the linearization constant ξ k ( t ) = f ( x ( t k ) , u ( t k ) ) A ( ρ ( t k ) ) x ( t k ) B ( ρ ( t k ) ) u ( t k ) ; the bound (21) is required to hold for this full aggregate. Each component of d ˜ ¯ accounts for the physical disturbance bound from Assumption 3, the linearization residual budget governed by (4), and the frozen-point offset | ξ k ( t ) | .
The frozen model is valid only on Z c = X c × U c ; nonlinear mismatch outside the frozen approximation is absorbed into d ˜ ( t ) . In the reported benchmark, interval violations are evaluated only on the diagnostic axes { u , r , q } . The model uses bounded-error full-state estimates rather than direct noise-free state measurements:
y ( t k ) x ^ ( t k ) = x ( t k ) + n y ( t k ) , | n y , i ( t k ) | e ¯ i ,
where y is the full-state diagnostic output, x ^ is the bounded-error reconstructed state, and n y collects sensor noise, fusion error, and FTESO reconstruction error. The leaky-accumulator floor is set by (23):
δ i floor = e ¯ i + ε buf ,
where ε buf > 0 prevents accumulator growth from bounded estimation error alone.

3. RAPIO-Based Actuator Fault Diagnosis

The proposed diagnosis architecture operates alongside the TNMPC/FTESO control loop, as illustrated in Figure 2. TNMPC generates a single actuator-command vector u ( t ) , whose branches supply the plant, FTESO, and nominal observer; they do not represent separate controller outputs. The FTESO supplies the bounded-error reconstructed measurement defined in (22) without altering the RAPIO state-propagation model.
Using the locally frozen model of Section 2.2, RAPIO propagates an admissible tube for fault-free behavior. Persistent interval violations in I d = { u , r , q } provide evidence of surge, rudder/yaw, and stern-plane/pitch faults, respectively. The learned scalar β ( t ) adapts only the admissible uncertainty bound, while the observer matrices, scheduled gain, and decision structure remain model-based.

3.1. Interval Observer for Locally Frozen Dynamics

Recalling the nonlinear fault-affected AUV dynamics (13), the system is described by
x ˙ ( t ) = f x ( t ) , u ( t ) + d ( t ) + F f a ( t ) ,
where f ( · ) is locally Lipschitz in x , d ( t ) denotes bounded disturbances, and f a ( t ) represents unknown actuator fault inputs.
Let t k denote a sampling instant and define the operating point x ( t k ) , u ( t k ) . A first-order Taylor expansion of f ( x , u ) about this point gives
f ( x ( t ) , u ( t ) ) = f x ( t k ) , u ( t k ) + A k x ( t ) x ( t k ) + B k u ( t ) u ( t k ) + Δ ( t ) ,
where
A k = f x ( x ( t k ) , u ( t k ) ) , B k = f u ( x ( t k ) , u ( t k ) ) ,
and Δ ( t ) denotes the higher-order remainder. Since f is twice continuously differentiable over Z c = X c × U c , Taylor’s theorem ensures that there exists a scalar L Δ > 0 such that the vector remainder Δ ( t ) R n x satisfies
Δ ( t ) L Δ z ( t ) z ( t k ) 2 , t [ t k , t k + 1 ) .
where z ( t ) = [ x ( t ) , u ( t ) ] . Since T s satisfies (4), the trajectory remains inside Z c and there exists Δ ¯ > 0 such that
Δ ( t ) Δ ¯ , t [ t k , t k + 1 ) ,
with Δ ¯ = L Δ sup t [ t k , t k + 1 ) z ( t ) z ( t k ) 2 .
The Jacobian A k is re-evaluated at every sampling instant t k , so the observer tracks operating-point changes at the sampling rate. Criterion (4) keeps the within-sample linearization residual inside the admissible disturbance envelope. High-speed maneuvers or abrupt input changes can still be handled by faster refreezing or by enlarging the disturbance envelope, with the corresponding increase in interval width. Substituting (25) into (24) yields the locally frozen model
x ˙ ( t ) = A k x ( t ) + B k u ( t ) + d ˜ ( t ) + F f a ( t ) ,
where the aggregated disturbance is defined in (29):
d ˜ ( t ) : = d ( t ) + Δ ( t ) + ξ k ( t ) ,
with ξ k ( t ) = f ( x ( t k ) , u ( t k ) ) A k x ( t k ) B k u ( t k ) .
Under (27), the disturbance envelope ϕ d ( t ) R 0 n x is defined as a componentwise radius that upper-bounds the effective aggregate disturbance entering the interval-error dynamics, including physical disturbance, Taylor remainder, frozen-point offset, and bounded measurement/reconstruction error. It is selected before propagation so that
ϕ d ( t ) | d ( t ) | + Δ ¯ 1 n x + | ξ k ( t ) | + | L k | n ¯ y , t ,
where 1 n x is the n x -dimensional all-ones vector, so the scalar Taylor bound Δ ¯ is applied componentwise. Here ϕ d ( t ) is the componentwise disturbance envelope, Δ ¯ bounds the Taylor remainder, ξ k ( t ) is the frozen linearization offset, L k is the observer gain introduced in (43), and n ¯ y bounds measurement/reconstruction error with componentwise entries n ¯ y , i = e ¯ i from (22). Define the effective disturbance entering the observer error dynamics as
d ˜ o ( t ) : = d ˜ ( t ) L k n y ( t ) .
where d ˜ o ( t ) is the disturbance seen by the observer error dynamics and n y ( t ) is the bounded measurement error. The triangle inequality then gives the componentwise bound in (32):
| d ˜ o ( t ) | ϕ d ( t ) , t .
The (33):
ϕ d ( t ) = η | d ^ o ( t ) | + ε 1 n x ,
where η = Δ max is a fixed worst-case proportional uncertainty scaling factor and ε > 0 is a fixed additive robustness margin. The vector d ^ o ( t ) R n x is the FTESO-derived disturbance estimate lifted to the state-dynamics ordering; its first six entries are the generalized dynamic disturbance estimates and its kinematic entries are zero.
Remark 2.
The envelope parameters are fixed offline as η = Δ max = 0.5 and ε = 0.5 . Thus, η is not selected from the realized perturbation level of a scenario case. Online variation comes only through the FTESO estimate | d ^ o ( t ) | : larger estimated disturbances widen the tube, while calm conditions keep it tight. Channel-specific linearization-error allowances are included in the width update using 5 % of the estimated surge speed, 10 % of the estimated yaw rate plus a 0.08  rad/s floor, and 5 % of the estimated heave velocity. Larger η and ε improve inclusion robustness but increase the minimum detectable fault magnitude in (60). These values are verified to satisfy (30) over the maneuvering envelope Z c by offline calibration: the additive floor ε = 0.5 absorbs the worst-case sum | Δ | + | ξ k | + | L k | n ¯ y computed from nominal and perturbed simulation cases, while the proportional term η | d ^ o | tracks the time-varying physical disturbance. Any operating regime with larger Δ (e.g., high-speed maneuvers exceeding the calibration envelope) would require a larger ε or a reduced T s per (4). Table 1 summarizes the qualitative effect of varying both parameters from the nominal design.
The admissible state set in (34) is represented by the two boundary observers x ( t ) and x + ( t ) :
X ( t ) = x R n x | x ( t ) x x + ( t ) .
Let
A o , k : = A k L k C
denote the frozen observer error matrix. The observer gain L k is selected by a deterministic scheduling rule so that A o , k is Metzler and Hurwitz at each sampled operating point.
The two boundary observers are
x ˙ + ( t ) = A o , k x + ( t ) + B k u ( t ) + L k y ( t ) + ϕ d ( t ) + Γ ( t ) ,
x ˙ ( t ) = A o , k x ( t ) + B k u ( t ) + L k y ( t ) ϕ d ( t ) Γ ( t ) ,
where C = I n x = I 12 , y is the reconstructed full-state output defined in (22), and Γ ( t ) R 0 n x is a nonnegative width-slack input, parameterized explicitly in (55). The full-state output C = I 12 is an idealization of direct state access; Proposition 2 shows the inclusion results extend to any bounded-error state reconstructor, covering practical FTESO-based AUV implementations. Equivalently, the observer can be written in center and half-width variables
m ( t ) = 1 2 x + ( t ) + x ( t ) , δ ( t ) = 1 2 x + ( t ) x ( t ) ,
which gives the same plotted tube [ m ( t ) δ ( t ) , m ( t ) + δ ( t ) ] . The center estimate follows a prediction-correction scheme with a separate small scalar gain γ ( 0 , 1 ) :
m k + 1 p = f nom m k c , u k ,
m k + 1 c = m k + 1 p + γ H k + 1 y k + 1 m k + 1 p ,
where f nom ( · ) denotes the nominal plant predictor with α = 1 and the recorded actuator commands. The diagonal matrix H k + 1 = diag ( h 1 , k + 1 , , h n x , k + 1 ) is the detector-gated correction mask. For non-diagnostic states h i , k + 1 = 1 . For diagnostic axes i I d , h i , k + 1 = 0 while the corresponding channel is in the confirmed-alarm freeze state, and h i , k + 1 = 1 otherwise. Detection uses m k + 1 p (before correction); the gated correction m k + 1 c prepares the center for the subsequent prediction step. Thus, when a channel is alarmed, only its measurement-correction term is suspended. The center continues to evolve through the nominal predictor (38).
Remark 3.
The center gain γ and the width gain L k serve distinct roles and are chosen independently. L k is a large Metzler-scheduled gain that ensures A o , k = A k L k C is Metzler and Hurwitz (providing the formal width-bounding certificate). γ is a small scalar ( γ i T s ) chosen so that the center tracks the nominal trajectory without being pulled toward fault-affected measurements. The gate H k + 1 further prevents a confirmed faulty diagnostic channel from dragging its own center toward the fault response. Using L k for the center would cause the center estimate to follow fault-induced measurement excursions, eliminating the fault residual and defeating detection.
The interval certificate is associated with the two boundary observers (35) and (36). In the discrete implementation, the center follows (38) and (39) rather than a continuous Luenberger form; the center-width decomposition m = 1 2 ( x + + x ) , δ = 1 2 ( x + x ) defined in (37) remains valid at each step. A sufficient initialization condition for guaranteed enclosure is x ( 0 ) x ( 0 ) x + ( 0 ) ; in the simulations m 0 c = x ( 0 ) and δ ( 0 ) 0 make this condition trivially satisfied. The formal componentwise enclosure in (40)
x ( t ) x ( t ) x + ( t ) , t 0 ,
is established directly from positivity of Metzler systems in Section 3.2.
Subtracting (36) from (35) gives the boundary separation w ( t ) = x + ( t ) x ( t ) = 2 δ ( t ) :
w ˙ ( t ) = A o , k w ( t ) + 2 ϕ d ( t ) + 2 Γ ( t ) .
where w ( t ) is the interval width, A o , k is the observer error matrix, ϕ d ( t ) is the disturbance envelope, and Γ ( t ) is the additional slack input. Dividing by two and writing δ ( t ) = 1 2 w ( t ) gives the half-width dynamics
δ ˙ ( t ) = A o , k δ ( t ) + ϕ d ( t ) + Γ ( t ) ,
which is the form integrated by the Van Loan augmented matrix exponential in Remark 4. Actuator faults f a ( t ) do not enter (41). Disturbances are absorbed by the boundary separation, whereas persistent actuator faults induce exits from the interval [ x ( t ) , x + ( t ) ] , enabling detection.

3.2. Stability and Interval Inclusion Guarantees

This subsection establishes boundedness, stability, and guaranteed interval inclusion for the tube-based observer of Section 3.1. All results are derived for the fault-free case f a ( t ) 0 .
Assumption 4
([10,15]). For each frozen operating point ρ ( t k ) , the observer gain L k R n x × n x is constructed by the deterministic scheduling rule
[ L k ] i i = i , [ L k ] i j = min 0 , [ A k ] i j μ c , i j ,
where i > 0 are fixed diagonal seeds and μ c > 0 is a cooperative margin, both selected as fixed observer-design parameters. The symbol μ c is used to distinguish this observer-design margin from the vehicle mass m. For every off-diagonal pair ( i , j ) , the closed-loop matrix A o , k : = A k L k C then satisfies the Metzler condition in (44):
[ A o , k ] i j = [ A k ] i j min 0 , [ A k ] i j μ c = max [ A k ] i j , μ c μ c > 0 ,
where [ A o , k ] i j is an off-diagonal entry of the closed-loop observer matrix and μ c is the cooperative margin. The inequality holds regardless of the sign of [ A k ] i j , including all Coriolis and hydrodynamic coupling terms. A o , k is therefore Metzler by construction at every operating point. This gain design is always feasible: for any frozen A k , choosing i to exceed the largest row-sum of positive off-diagonal entries of A k plus a margin is sufficient to make A o , k both Metzler and Hurwitz simultaneously. The diagonal seeds i are selected offline to provide contraction on the tested operating envelope and are verified by the certificate below. At each update, the triple certificate
min i j [ A o , k ] i j 0 , max i Re ( λ i ( A o , k ) ) λ < 0 , min i , j [ e A o , k T s ] i j 0
is imposed on the sampled operating envelope, where λ > 0 is the certified uniform decay margin, λ i ( A o , k ) denotes the i-th eigenvalue, Re ( · ) its real part, and T s the diagnostic update period. The off-diagonal entries of L k are mathematical observer-gain entries used to enforce cooperativity; they are not restricted to be sensor weights. The certificate is verified numerically over the same diagnostic update interval used in the Van Loan width propagation and logged at run time.
Remark 4.
The AUV Jacobian A k carries Coriolis and hydrodynamic coupling terms whose signs change with surge speed, yaw rate and depth. It reflects the underactuated, coupled nature of the vehicle. The scheduling rule (43) handles this by construction. Every off-diagonal entry of A o , k is brought to at least μ c > 0 regardless of the sign of [ A k ] i j . The half-width ODE (42) is then discretized exactly over [ t k , t k + 1 ) via the Van Loan augmented matrix exponential, yielding the update Φ k δ k + ψ k with Φ k , ψ k 0 . The implementation applies the Müller form | Φ k | δ k + | ψ k | as a floating-point safeguard against Jacobian rounding near the cooperative margin and the triple certificate (45) is logged at every update step.
Assumption 5
([10,12]). This condition is a consequence of Assumptions 1–3 and the envelope construction (30); it is stated separately for clarity in the stability proof. The effective disturbance entering the observer error dynamics satisfies (46):
| d ˜ o ( t ) | ϕ d ( t ) ,
where ϕ d ( t ) R n x is componentwise nonnegative, uniformly bounded, and piecewise continuous.
Assumption 6
([10,12,37]). As a design choice that preserves the nonnegative bounded-input structure of the width dynamics, the slack term is parameterized as Γ ( t ) = 1 β ( t ) d slack , d slack 0 , β ( t ) [ 0 , 1 ] , giving the interval width dynamics
w ˙ ( t ) = A o , k w ( t ) + 2 ϕ d ( t ) + 2 1 β ( t ) d slack .
where w ( t ) is the interval width and d slack is the nonnegative maximum slack vector. Because A o , k is Metzler and Hurwitz, these dynamics form a positive and stable comparison system. The learning signal enters only through the nonnegative bounded input 1 β ( t ) d slack .
Lemma 1.
Given w ( 0 ) 0 , the width w ( t ) defined by (47) remains componentwise nonnegative and uniformly ultimately bounded for all t 0 .
Proof. 
Since A o , k is Metzler, e A o , k t is componentwise nonnegative for all t 0 . With b ( t ) = 2 ϕ d ( t ) + 2 ( 1 β ( t ) ) d slack 0 , the variation-of-constants formula gives (48):
w ( t ) = e A o , k t w ( 0 ) + 0 t e A o , k ( t τ ) b ( τ ) d τ .
This proves nonnegativity on [ t k , t k + 1 ) . Since A o , k updates at each t k , the argument applies piecewise. At each t k , continuity of x + and x gives w ( t k ) 0 as the initial condition for the next frozen interval, therefore proving propagating nonnegativity by induction. Boundedness follows from the uniform certificate in (45). Since max i Re ( λ i ( A o , k ) ) λ < 0 for all sampled frozen matrices and b ( t ) is uniformly bounded, the positive switched comparison system admits a common finite envelope. Equivalently, for each frozen interval the componentwise steady bound A o , k 1 b ¯ is finite, and the uniform decay margin prevents growth under the scheduled updates. Here b ¯ : = sup t 0 b ( t ) denotes any finite componentwise upper bound on the nonnegative input b ( t ) = 2 ϕ d ( t ) + 2 Γ ( t ) .    □
Proposition 1.
Let D i > 0 denote the componentwise one-step prediction-error bound calibrated from nominal operation on the diagnostic axes I d :
e ˜ m , i ( k ) : = x i ( k ) m i p ( k ) D i , i I d ,
where e ˜ m , i is the center prediction error before correction. Let Φ k = e A o , k T s and let ψ k denote the Van Loan input generated using d rate , i = D i / T s . If the calibrated input satisfies
| e ˜ m ( k + 1 ) | Φ k | e ˜ m ( k ) | + ψ k ,
then δ i ( 0 ) D i implies δ i ( k ) | e ˜ m , i ( k ) | for all k 0 , and hence, x i ( k ) x i ( k ) x i + ( k ) on I d .
Proof. 
Metzlerness gives Φ k 0 and, since d rate 0 , also ψ k 0 . Combining the width recursion with (50) yields
δ ( k + 1 ) | e ˜ m ( k + 1 ) | Φ k δ ( k ) | e ˜ m ( k ) | 0 .
Because δ ( 0 ) D | e ˜ m ( 0 ) | , induction gives δ ( k ) | e ˜ m ( k ) | for every k, which is equivalent to the stated interval framing.    □
Theorem 1.
Suppose Assumptions 4–6 hold and x ( 0 ) x ( 0 ) x + ( 0 ) . Then, in the absence of actuator faults,
x ( t ) x ( t ) x + ( t ) , t 0 .
where x ( t ) and x + ( t ) are the lower and upper interval bounds and x ( t ) is the true plant state.
Proof. 
Define the upper and lower enclosure errors e + ( t ) = x + ( t ) x ( t ) and e ( t ) = x ( t ) x ( t ) . Subtracting the plant (28) from the upper observer (35) and using y = x + n y together with definition (31) gives L k y + d ˜ ( t ) = L k x + d ˜ o ( t ) , yielding
e ˙ + ( t ) = A o , k e + ( t ) + ϕ d ( t ) + Γ ( t ) d ˜ o ( t ) ,
e ˙ ( t ) = A o , k e ( t ) + ϕ d ( t ) + Γ ( t ) + d ˜ o ( t ) .
The enclosure-error dynamics in (53) and (54) are positive comparison systems. By Assumption 5, ϕ d ( t ) ± d ˜ o ( t ) 0 , and Γ ( t ) 0 . Since A o , k is Metzler, the comparison systems for e + and e are positive. The initialization x ( 0 ) x ( 0 ) x + ( 0 ) gives e + ( 0 ) 0 and e ( 0 ) 0 , hence e + ( t ) , e ( t ) 0 for all t 0 . Therefore, x ( t ) x ( t ) x + ( t ) componentwise, which proves (52).    □

3.3. RL-Augmented Uncertainty Bound Modulation

RL-augmented uncertainty bound modulation is introduced as a bounded exogenous input acting exclusively on the interval width dynamics, with the objective of enhancing disturbance accommodation while preserving the guarantees of Section 3.2. As illustrated in Figure 3, the learning module is structurally decoupled from the center estimation dynamics and the plant model.
The learning signal enters the width dynamics exclusively through a bounded confidence-dependent slack: high confidence tightens the disturbance-rate slack and low confidence widens it:
Γ ( t ) = 1 β ( t ) d slack , β ( t ) [ 0 , 1 ] , d slack 0 ,
where Γ ( t ) is the learning-modulated slack input, β ( t ) is the bounded RL confidence signal, and d slack is the maximum channel-wise slack. The corresponding learning-modulated width dynamics are given by Equation (47), where w ( t ) is the interval width and A o , k , ϕ d ( t ) , β ( t ) , and d slack retain the definitions given above.
By Lemma 1, w ( t ) remains nonnegative and uniformly bounded. Learning therefore enlarges the admissible tube only through the bounded input 2 ( 1 β ( t ) ) d slack and does not modify A o , k or its Metzler/Hurwitz certificate.
The center prediction-error bound (49) is independent of β ( t ) because β modulates only the disturbance-rate slack fed into the Van Loan width update. Consequently, the framing guarantee of Proposition 1 holds for all admissible β ( t ) .
Let e m ( t ) : = x ( t ) m ( t ) denote the interval-center error and let the superscript ss denote its persistent steady value. For a single actuator fault f a ( t ) = f ¯ j e j , decompose the steady-state center error as
e m ss = A o , k 1 d ˜ ss + F f ¯ j e j .
where e m ss is the steady-state center error, d ˜ ss is the steady-state aggregate disturbance, f ¯ j is the persistent fault magnitude in actuator channel j, and e j is the corresponding canonical basis vector in R 3 . Define the scalar fault-to-channel gain
g i j : = [ A o , k 1 F ] i j ,
where g i j is the scalar gain from actuator fault j to state channel i. Thus, | [ A o , k 1 F f ¯ j e j ] i | = g i j | f ¯ j | . Setting κ i = [ A o , k 1 ] i , : 1 and η i = κ i sup t d ˜ ( t ) , the reverse triangle inequality gives
| e m , i ss | g i j | f ¯ j | η i ,
where η i is a disturbance-induced bias bound and is distinct from the pose vector η and the scalar envelope factor used elsewhere. In (60), τ i base > 0 is the fixed accumulator threshold and δ i floor > 0 is the residual noise floor; both are defined operationally in Section 3.4. Here g i j is a nonnegative scalar (the absolute value of a single matrix entry), so its use as the lower-bound coefficient in (58) is justified directly by the reverse triangle inequality without invoking any row-norm argument. Define the certified worst-case steady half-width by (59):
δ ¯ i : = 1 2 sup k A o , k 1 2 ϕ ¯ d + 2 d slack i ,
where δ ¯ i is the certified worst-case steady half-width for channel i, ϕ ¯ d is a componentwise upper bound on ϕ d ( t ) over the operating envelope, and the supremum is evaluated over the certified frozen matrices. This bound uses the coupled positive system directly rather than a scalar approximation. A sufficient offline detectability condition is
g i j | f ¯ j | > δ ¯ i + η i + 1 σ i σ i τ i base + δ i floor ,
where σ i is the channel-specific leak of the accumulator in Section 3.4. The factor ( 1 σ i ) / σ i converts the fixed accumulated gate τ i base into the required persistent residual excess, consistent with Proposition 3 and Remark 7.
Remark 5.
Condition (60) is a conservative offline separation check, not an online decision rule. The conservatism comes from using worst-case tube width, disturbance masking and frozen-Jacobian fault-to-channel gains. These quantities are used only during calibration. Online detection uses the propagated interval tube, directed cross-channel margins, and the fixed leaky accumulator of Section 3.4.
Theorem 2.
Suppose Assumptions 4–6 hold and x ( 0 ) x ( 0 ) x + ( 0 ) . For any bounded measurable β ( t ) [ 0 , 1 ] , the RL-augmented interval observer preserves the componentwise inclusion condition (61) for every certified component i:
x i ( t ) x i ( t ) x i + ( t ) , t 0 ,
where x i ( t ) and x i + ( t ) are the lower and upper bounds of component i, and x i ( t ) is the corresponding true state component, in the absence of actuator faults.
Proof. 
β ( t ) enters only the width subsystem (47) via the slack ( 1 β ) d slack . It does not appear in the center prediction step (38) or in the detector mask H k + 1 of the correction step (39), and does not modify A o , k , L k , or the scheduled-gain rule (43). The Metzler/Hurwitz certificate therefore holds for all β ( t ) [ 0 , 1 ] , and the inclusion result of Theorem 1 applies without modification.    □
Proposition 2.
Suppose any stabilizing feedback law u ( t ) = π ( x ( t ) ) keeps x ( t ) X c for all t 0 , and any state reconstructor supplies x ^ ( t k ) with the componentwise error bound in (62):
| x i ( t k ) x ^ i ( t k ) | e ¯ i , t k , i .
Then Theorems 1 and 2 hold with the noise floor updated to δ i floor = e ¯ i + ε buf , independently of the specific form of π ( · ) or the state reconstructor, provided the innovation-noise term is covered by | L k | n ¯ y in (30). In addition, zero healthy-operation accumulator growth requires the boundary-excess buffer ε buf , i > 0 to dominate the reconstruction error margin.
Proof. 
The center prediction error at step k is e ˜ m ( k ) = x ( k ) m p ( k ) , where m p ( k ) follows the prediction step (38). The nominal predictor uses the recorded command u k 1 , which is the same command applied to the plant (28); the control input therefore cancels in the prediction error regardless of the form of π ( · ) . The scheduled gain L k and frozen Jacobian A k do not depend on the controller structure. Measurement or reconstruction noise n y enters only through the correction step (39) with gain γ H k + 1 ; since 0 h i , k + 1 1 , its contribution is bounded by γ n ¯ y and is absorbed into δ i floor . Condition (30) covers the remaining disturbance terms entering the boundary observers, so Assumption 5 remains satisfied.
Since x ( t ) X c is guaranteed by assumption, the Taylor remainder bound (26) and the slowly varying scheduling condition (2) hold, so Theorem 1 applies without modification to yield x i ( t ) x i ( t ) x i + ( t ) for all i.
When x ^ ( t k ) is used in the residual computation, the apparent boundary-excess residual satisfies ϱ i ( k ) e ¯ i in the fault-free case. If ε buf , i > 0 , then
ϱ i ( k ) e ¯ i < e ¯ i + ε buf , i = δ i floor ,
and the accumulator does not grow in the fault-free case. Without this buffer, the bounded reconstruction error is still included in the residual bound, but zero accumulator growth is not guaranteed. The width dynamics (47) and the RL signal β k are unaffected by the controller or state reconstructor choice, completing the proof.    □
The learned signal β k does not change the observer gain, frozen Jacobian, nominal center predictor, detector gate, or alarm thresholds. It only modulates the sampled disturbance-rate slack used in the width propagation,
d i ( k ) = D i + ( 1 β k ) D i slack T s , i I d ,
where d i ( k ) is the sampled disturbance rate used in the width update, D i is the calibrated one-step disturbance bound, and D i slack is the maximum one-step RL-modulated slack. Its continuous-time counterpart in (55) is d i slack = D i slack / T s . The quantities β k , T s , and I d denote the SAC confidence signal, the diagnostic sampling period, and the diagnostic axis set, respectively. Thus, large β k tightens the tube, while small β k gives a more conservative enclosure. Table 2 summarizes this sign convention and its diagnostic effect.

3.4. Interval-Consistency-Based Fault Detection and Isolation

Building on the interval inclusion guarantee of Theorem 1, the proposed FDI mechanism detects actuator faults through systematic violations of interval consistency. To distinguish it from the yaw-rate state r, the scalar boundary-excess residual is denoted ϱ i ( k ) . In the absence of actuator faults, the inclusion condition (63) holds:
x ( t ) x ( t ) x + ( t ) , t 0 .
At discrete instants t k , the componentwise residual is defined as
ϱ i ( k ) = max { 0 , x ^ i ( t k ) x i + ( t k ) , x i ( t k ) x ^ i ( t k ) } ,
where ϱ i ( k ) is the nonnegative boundary-excess residual for channel i, x ^ i ( t k ) is the measured or reconstructed state component, and x i ( t k ) and x i + ( t k ) are the lower and upper interval bounds. It satisfies ϱ i ( k ) = 0 under fault-free interval inclusion. Persistent actuator faults may induce sustained growth of ϱ i ( k ) beyond zero boundary excess.
The detector evaluates only the diagnostic axes I d = { 1 , 6 , 5 } in manuscript notation, corresponding to surge velocity u, yaw rate r, and pitch rate q. In the implementation these same channels are stored with zero-based indices { 0 , 5 , 4 } . The stern-plane fault is injected in the physical vertical-plane channel and diagnosed through the pitch-rate axis q.
To discriminate persistent faults from transient disturbances, the detector uses a leaky CUSUM statistic with a clearance reset and a recovery-clear rule. Let δ i clr denote the hard-flush clearance level and δ i floor the fixed per-channel noise floor. The statistic is
S i ( k ) = 0 , ϱ i ( k ) < δ i clr , σ i S i ( k 1 ) , δ i clr ϱ i ( k ) < δ i floor , max 0 , σ i S i ( k 1 ) + ϱ i ( k ) δ i floor , ϱ i ( k ) δ i floor ,
with S i ( 0 ) = 0 . Here δ i floor > 0 is a noise floor ensuring ϱ i ( k ) δ i floor S i ( k ) does not grow, and σ i ( 0 , 1 ) provides exponential forgetting of isolated spikes. The clearance reset makes the tabulated parameter δ i clr explicit. The residuals below the clearance level flush the accumulator, while residuals between the clearance and noise-floor levels decay without new evidence. This matches the implemented bound-crossing logic. After a confirmed alarm, the same channel enters a correction-freeze state with h i , k = 0 in (39). This freeze is held for a short minimum dwell time and is released when the boundary-excess residual has clearly peaked and is recovering by (66).
ϱ i ( k ) ρ i rec max A i ( k ) ϱ i ( ) or ϱ i ( k ) δ i rec ,
for N i rec consecutive samples, where A i ( k ) denotes the active alarm interval for channel i. The recovery ratio ρ i rec ( 0 , 1 ) tests whether the residual has fallen by a prescribed fraction of its alarm-interval peak, δ i rec > 0 is an absolute recovery level, and N i rec N is the required consecutive-sample count. During the recovery tail, a lockout suppresses re-accumulation only while ϱ i ( k ) continues to decrease; if the residual rises again, the lockout is removed and the detector can re-arm.
Remark 6.
The design parameters ( σ i , δ i floor , δ i clr , τ i base , N i on , T i hold , T i min , ρ i rec , N i rec ) govern the false-alarm–missed-detection–delay trade-off. The leak σ i sets the delay–spike-rejection balance, while τ i base is calibrated offline after the propagated tube, measurement margin and coupling-aware margins are fixed. N i on is the number of consecutive qualifying samples required to confirm a fault, T i hold is the alarm-hold duration, and T i min is the minimum correction-freeze dwell. The missed-detection condition σ i ε i / ( 1 σ i ) τ i base defines the minimum detectable persistent residual for the leaky update in (65). Here ε i : = ϱ i δ i floor > 0 denotes a persistent per-sample residual excess above the noise floor. The detector parameters are selected offline and fixed throughout evaluation, while their sensitivity is assessed using the scenario-sweep results.
The final alarm threshold is fixed offline for each diagnostic channel and is given by (67),
τ i dyn ( k ) = τ i base
where τ i dyn ( k ) is the online alarm gate and τ i base > 0 is the fixed value calibrated from validation data. A fault is declared in channel i when
S i ( k ) > τ i dyn ( k ) and ϱ i ( k ) > δ i clr ,
where S i ( k ) is the leaky accumulated residual for channel i, τ i dyn ( k ) is the fixed online alarm gate and δ i clr is the hard-clearance level. The second condition is a clearance gate. It prevents a stored accumulator value from producing an alarm after the instantaneous residual has already returned below the admissible clearance level. The detector then applies short persistence confirmation: N i on = 3 for surge, yaw and pitch rate. Once confirmed, the alarm is held for a short fixed interval to prevent immediate dropout during transient tube re-entry. In addition, the recovery-clear condition (66) resets the accumulator and releases the channel-wise correction freeze once the residual has decayed sufficiently from its active-alarm peak.
Proposition 3.
Suppose there exist k 0 N and ε i > 0 such that ϱ i ( k ) δ i floor and ϱ i ( k ) δ i floor ε i for all k k 0 . Then S i ( k ) is lower-bounded by the leaky accumulation sequence and a fault is declared in finite time whenever
σ i ε i 1 σ i > τ i base .
This is a conservative sufficient condition before the finite N i on confirmation count is applied.
Proof. 
Under the persistence condition, S i ( k ) σ i S i ( k 1 ) + σ i ε i . Iterating from k 0 with S i ( k 0 ) 0 yields (70):
S i ( k ) σ i ε i 1 σ i 1 σ i k k 0 ,
which converges monotonically to σ i ε i / ( 1 σ i ) . Since τ i dyn ( k ) = τ i base , threshold crossing occurs in finite time when (69) holds.    □
After the fixed threshold is crossed, N i on consecutive raw crossings are required before confirming an alarm, which is then held for T i hold . This discrete decision logic does not alter the width propagation, Metzler certificate, or fault-free inclusion result.
Remark 7.
The leaky accumulator in (65) gives an explicit delay–false-alarm trade-off. For a persistent residual excess ϱ i ( k ) δ i floor ε i > 0 , the accumulator crosses the fixed gate when (69) holds. The reported empirical delay also includes the confirmation count, alarm hold, recovery lockout, and coupling-margin logic under the fixed detector configuration.
The detected index set is defined in (71):
I f ( k ) = i S i ( k ) > τ i dyn ( k ) , ϱ i ( k ) > δ i clr .
where I f ( k ) is the set of detected faulty diagnostic channels at step k, and S i ( k ) , τ i dyn ( k ) , ϱ i ( k ) , and δ i clr are defined in (65)–(68).
For isolation, the cross-channel noise floor for non-target channels must be set to mask coupled residuals. Specifically, the nominal floor for channel i is defined as
δ i floor , 0 = max e ¯ i + ε buf , η i + max j : ( G d ) i j = 0 g i j | f ¯ j | max ,
where δ i floor , 0 is the nominal floor for channel i, e ¯ i is the estimation-error bound, ε buf is the positive buffer, η i is the disturbance-induced center-error bound, g i j is the fault-to-channel gain from (57), and | f ¯ j | max is an upper bound on the anticipated fault magnitude determined offline from the prescribed fault scenario set used for evaluation. The operative detector floor satisfies δ i floor δ i floor , 0 .
Proposition 4.
Assume a single actuator fault affects actuator j, f a ( t ) = f ¯ j e j . Let G d be the diagnostic support matrix in (17). If condition (60) holds for the direct diagnostic channel and the cross-channel floor condition (72) holds for the non-target channels, then for sufficiently persistent faults within the calibrated scenario class,
I f ( k ) = i I d ( G d ) i , j 0 .
where ( G d ) i , j is the diagnostic-support entry mapping actuator fault j into monitored diagnostic channel i.
Equation (73) is the diagnostic-channel statement used in the reported benchmark. The full physical distribution matrix F still appears in the plant and observer propagation, so coupled responses in the other states remain part of the 6-DOF dynamics, but they are not separate alarm channels.
Proof. 
Case 1: ( G d ) i , j = 0 . When ( G d ) i , j = 0 , the fault is not a direct diagnostic-channel fault for channel i. It can still influence channel i through model coupling, with gain g i j from (57). The operative floor (72) satisfies δ i floor δ i floor , 0 . The offline calibration behind Remark 5 selects the cross-channel margins so that non-target coupled residuals remain below the direct fault-channel separation used in (60). Consequently, ϱ i ( k ) δ i floor for all sufficiently large k, the persistence condition of Proposition 3 fails, and S i ( k ) remains below the fixed calibrated gate.
Case 2: ( G d ) i , j 0 . The persistent fault induces a steady-state bias in e m , i via (56). From (58) and the detectability condition (60), ϱ i ( k ) δ i floor ε i > 0 for all large k, and Proposition 3 guarantees threshold crossing.
Combining both cases, I f ( k ) coincides with the nonzero entries of column j of the diagnostic support matrix G d , establishing (73) for the monitored channels.    □

3.5. Reinforcement Learning Framework

The uncertainty-bound modulation problem is formulated as the discounted Markov decision process in (74):
M = ( S , A , P , R , γ SAC ) ,
where S is the policy state space, A R 11 is the compact SAC action space, P is the transition kernel, R ( s k , a k ) is the training reward, and γ SAC ( 0 , 1 ) is the discount factor (distinct from the center correction gain γ of Section 3.1). The learning module acts strictly as a supervisory layer; the scheduled-gain rule, system matrices, and center error dynamics remain independent of β k . Figure 4 summarizes the corresponding online propagation and decision sequence.
The policy state is
s k = e cte , k e ψ , k e z , k ν ^ 1 , k d ^ k R 10 ,
where s k is the SAC policy state, e cte , k is the cross-track error, e ψ , k is the heading error, e z , k is the depth error, ν ^ 1 , k is the estimated surge velocity, and d ^ k R 6 is the FTESO generalized disturbance estimate, corresponding to the first six entries of d ^ o ( t k ) ; hence, the state contains 3 + 1 + 6 = 10 components. The policy state excludes interval center and width states, preserving the separation between learning and observer certification.
The SAC policy outputs an action vector a k = [ a k c , β ˜ k ] R 11 . The subvector a k c R 10 contains the first ten bounded control/guidance tuning actions, used by the TNMPC environment, while the last component is mapped to the RAPIO modulation signal by
β k = ( 1 α β ) β k 1 + α β sat [ 1 , 1 ] ( β ˜ k ) + 1 2 [ 0 , 1 ] ,
with smoothing factor α β = 0.15 . The scalar operator is sat [ 1 , 1 ] ( z ) : = min { 1 , max { 1 , z } } , and β ˜ k R is the raw policy output associated with tube modulation. Since (76) is a convex low-pass update, 0 β k 1 and | β k β k 1 | α β for all raw policy outputs. The observer guarantees require only bounded measurability of β k [ 0 , 1 ] . Only this last component enters the RAPIO width dynamics; the remaining action components affect the closed-loop controller but do not enter the observer gains or center error dynamics.
SAC is used only to optimize the bounded supervisory signal over the episode-level trade-off among missed detections, false alarms, tracking performance, and beta chatter; it does not change the inclusion or stability certificate.
The training reward combines trajectory-tracking performance with diagnostic penalties:
R k = R k track 0.1 β k 2 100 I fault , k I miss , k I { β k > 0.6 } 50 I healthy , k I FA , k 0.5 | β k β k 1 | ,
where R k is the scalar SAC stage reward, R k track contains the TNMPC path-tracking reward, β k is the bounded confidence signal, I fault , k and I healthy , k indicate fault-active and healthy operation, respectively, I miss , k penalizes missed detections during training fault windows when the policy still declares high confidence, I FA , k penalizes false alarms during healthy operation, and the final term discourages beta chattering. For any proposition E, I { E } equals one when E is true and zero otherwise; all I terms in (77) are binary indicators.
The analytical RAPIO guarantees depend only on the bounded scalar β k ; the other action components influence tracking but not observer positivity, center prediction-error framing, or interval inclusion.
Discretizing (47) over T s with the matrix-exponential positive-system update gives
w k + 1 = e A o , k T s w k + 0 T s e A o , k ( T s τ ) 2 ϕ d ( k ) + 2 ( 1 β k ) d slack d τ .
where w k is the interval width vector at step k, A o , k is the frozen observer error matrix, ϕ d ( k ) is the bounded disturbance envelope, d slack is the maximum slack vector, β k is the bounded RL confidence signal, and τ is the integration variable. Since A o , k is Metzler and Hurwitz, e A o , k T s 0 and is Schur stable. Equation (78) can be evaluated with the augmented Van Loan matrix exponential, which computes the integral term exactly for the frozen matrix over the sampling interval. The continuous-time inclusion result is therefore preserved by the positive discrete propagation used in the simulations. Since ϕ d ( k ) and β k [ 0 , 1 ] are uniformly bounded, the positive stable recursion is uniformly bounded, giving sup k w k < .
The policy maximizes the objective in (79):
J ( π ϑ ) = E k = 0 γ SAC k R k , subject to a k A , β k [ 0 , 1 ] .
where J ( π ϑ ) is the discounted return objective of the SAC policy π ϑ , ϑ is its trainable parameter vector (distinct from the Euler pitch angle θ ), E [ · ] denotes expectation over closed-loop rollouts, γ SAC is the SAC discount factor, R k is the stage reward, a k is the policy action, and A is the compact action set.
Full reproducibility details—network architecture, training seed, SAC hyperparameters, action bounds, and policy selection criterion—are reported in Section 4.
Algorithm 1 summarizes the offline policy-training and online RAPIO deployment sequence. 
Algorithm 1 RL-augmented uncertainty bound modulation for the RAPIO observer
  • Require: Nonlinear plant f ( · ) ; fixed scheduled-gain rule for L k ; diagnostic update period T s ; ϕ d ( t ) ; channel-wise slack d slack ; trained policy π ϑ
     1:
Offline Training:
     2:
Train bounded policy π ϑ via SAC to maximize J ( π ϑ ) ; freeze ϑ
     3:
Online Deployment:
     4:
Initialize: x 0 x 0 x 0 + ; equivalently store m 0 and w 0 0 ; set S i ( 0 ) = 0 i
     5:
for  k = 0 , 1 , 2 ,   do
     6:
      Acquire the reconstructed state x ^ k from FTESO
     7:
      Construct policy state s k via (75)
     8:
      Evaluate a k = π ϑ ( s k ) and compute β k from its last component via (76)
     9:
      Center: propagate m k + 1 p via (38)
   10:
      Width: propagate w k + 1 once via (78) using β k
   11:
      Compute residual ϱ i ( k ) via (64) i and apply correction m k + 1 c via ()
   12:
      Update S i ( k ) using ϱ i ( k ) via (65) i
   13:
      Use the fixed calibrated gate τ i dyn ( k ) from (67) i
   14:
      Evaluate (68) and update I f ( k )
   15:
end for
The learning-based modulation preserves positivity and boundedness of the interval width (Lemma 1), the center prediction-error bound (Proposition 1), and interval inclusion (Theorem 2), while reducing conservatism relative to a fixed worst-case envelope.

4. Simulation Setup

All simulations are performed using Python 3.11 and CasADi 3.7.2. The plant integrator uses Δ t = 0.1 s, while the IO-FDI observer, CUSUM statistic, and alarm logic are updated once per agent step at the diagnostic period T s = 1.0 s. Confirmed alarms are inhibited during the first 25 s to allow the FTESO and RAPIO states to settle before the first fault at t = 50 s. The plant is a nonlinear 6-DOF AUV model based on standard marine vehicle dynamics [34], with hydrodynamic parameters consistent with [36]. Table 3 lists the principal coefficients. Disturbance bounds cover ± 50 % hydrodynamic and inertial perturbations and ocean currents up to 0.3 m/s. Ocean currents comprise a base component of up to 0.3 m/s, a time-varying sinusoidal component, a linear depth-shear term, and spatially varying sinusoidal components. Together, they enter the plant as an unknown bounded disturbance consistent with Assumption 3.
The fixed RAPIO interval-observer settings are summarized in Table 4.
The uncertainty-bound modulation policy is trained offline with SAC [38] for 5 × 10 5 environment steps using Stable-Baselines3 2.7.0. The actor maps s k R 10 to an 11-dimensional action. Table 5 gives the remaining SAC parameters. Since β k is bounded for all policy outputs, the frozen policy satisfies the bounded-input condition used in Theorem 2.
In Table 5, ρ Polyak denotes the soft target-network update coefficient; its subscript distinguishes it from the LPV scheduling vector ρ ( t ) .
The simulations apply LoE factors to the actuators and monitor the dominant diagnostic signatures summarized by (17). Surge, yaw-rate, and pitch-rate signatures correspond to columns 1, 2, and 3 of G d and rows 1, 6, and 5, respectively. The complete 6-DOF plant then propagates the degraded actuation through its coupled dynamics before the noisy FDI measurement is formed. Surge is evaluated as a partial LoE case rather than complete loss because the underactuated AUV has a single main thruster; a 100 % surge LoE would remove propulsion authority and make the rudder and stern-plane ineffective for controlled maneuvering. The yaw and pitch cases use complete LoE in their respective control surfaces. The injection schedule is summarized in Table 6.
The FDI module uses interval-consistency checks with a leaky cumulative statistic. All detection parameters are selected offline and fixed throughout evaluation (Table 7).
The detector uses a coupling-aware disturbance allocation. When one diagnostic channel leaves its base interval, a bounded temporary margin is allocated only to physically coupled target channels. This online allocation is driven by measured tube-exit magnitudes, not by fault-window labels, and the source channel itself is not widened by its own violation.
The additional schedule-blind campaigns vary LoE severity, temporal profile, and persistence and evaluate additive-bias and stuck/jam faults under nominal and stressed environments. Table 8 defines the campaign matrix. All detector parameters remain fixed, and no fault label, magnitude, onset, duration, or profile is supplied to the online FDI. LoE magnitude is defined as 100 ( 1 α i ) % .

5. Fault Detection Performance and Evaluation

RAPIO performance is evaluated under progressively more demanding conditions, from model uncertainty and isolated faults to overlapping faults, varied fault profiles, and benchmark comparisons. Performance is measured through detection rate (DR), mean delay, false alarms, and the logged Metzler/Hurwitz certificate status.

5.1. Scenario Grid and Disturbance-Envelope Sensitivity

Robustness to model mismatch is examined during 3D helical tracking across N = 72 cases spanning six hydrodynamic uncertainty levels (0– 50 % of nominal parameter values), four base ocean-current amplitudes, and three disturbance realizations per combination ( 6 × 4 × 3 = 72 ). Each uncertainty level creates a deliberate structural mismatch. The observer Jacobian A k is frozen using nominal parameters, while the true plant uses perturbed hydrodynamic coefficients. The ± 50 % perturbation envelope exceeds the ± 15 25 % identification uncertainty reported under standard sea-trial conditions [36]. Across all scenarios, the second-order linearization residual Δ ( t ) remained within the selected disturbance envelope ϕ d ( t ) , supporting the locally frozen LTI approximation over the tested operating set.
Sensitivity to the observer disturbance envelope is evaluated by scaling the RAPIO/IO disturbance floor by factors of 0.5 , 1.0 , and 2.0 while keeping all gains, D i , D i slack , measurement margins, cross-channel margins, and alarm parameters fixed. The full 72-case scenario grid is repeated at each scale. Table 9 shows that the proposed RAPIO detector maintains 100 % all-channel detection, zero event-level false alarms and a 100 % Metzler/Hurwitz certificate rate throughout the tested range. The delay variations are small relative to the 30 s fault windows, indicating that the proposed decision logic is dominated by persistent interval evidence rather than by the particular disturbance-floor scale.

5.2. Sequential Fault Detection Performance

Sequential-fault performance is assessed under high uncertainty by examining healthy-state enclosure and channel-selective alarms. Figure 5 shows that all states remain enclosed within admissible intervals before fault onset. Confirmed alarms occur in the intended surge, yaw-rate, and pitch-rate channels with delays of 7, 6, and 5 s, respectively. No new alarm is raised during healthy operation after fault clearance. The β ( t ) signal remains bounded while modulating the propagated tube slack.
Table 10 summarizes the aggregate performance of the sequential fault campaign.

5.3. Concurrent-Fault Isolation

RAPIO is evaluated under overlapping surge and yaw faults to determine whether channel selectivity is preserved despite dynamic coupling. A surge fault is injected during t = 50 –90 s and a yaw fault during t = 70 –120 s, with a 20 s overlap ( t = 70 –90 s). During the overlap, the cross-channel masking rule raises the yaw noise floor from 0.20 to 0.60 rad/s to suppress Coriolis-induced residuals. The yaw fault remains above the masked floor and therefore continues to increase the accumulator. Results over the full N = 72 scenario grid are shown in Figure 6 and Table 11.
Every fault is correctly isolated across all 72 scenarios, producing the correct FDI labels with zero false alarms. The mean yaw delay of 4.29 s, despite the concurrent surge fault, confirms that the masked floor suppresses coupling residuals while remaining permeable to the genuine yaw signal.

5.4. Fault-Severity and Temporal-Profile Characterization

RAPIO is further assessed through severity and timing sweeps to establish when a fault produces sufficient evidence for confirmation. The 180-case sweep identifies the lowest LoE level confirmed in all six paired cases. Table 12 gives boundaries of 50% for surge, 70% for rudder, and 100% for the stern plane. Every case retained the Metzler certificate with no false alarms or cross-detections.
Below the complete-detection boundary, partial stern-plane losses were not confirmed in every paired condition because TNMPC can compensate them before a persistent interval inconsistency develops.
Table 13 reports six paired cases for each channel and profile, with mean delay computed over detected cases. Gradual-fault delay starts with the ramp, while intermittent detection is credited only during active degraded intervals. All 54 cases had no false alarms or cross-detections.
The missed cases are limited to intermittent surge and gradual yaw faults. Surge bursts of 6–7 s can end before sufficient evidence accumulates, after which the accumulator decays during the healthy gap. Gradual yaw loss begins near healthy operation and is initially absorbed by the closed-loop response. It is confirmed once the developing effect produces a sustained interval inconsistency, separating persistent faults from short recovery transients.
For 50% surge LoE, Table 14 identifies 17 s as the shortest tested duration with 6/6 first-burst detections. The maximum delay was 16 s, and one case was not confirmed within a 16 s burst. Yaw faults were confirmed during 5–8 s bursts with a 3 s delay. Stern-plane faults used 7–8 s bursts and reached 100% eventual detection with a 7.5 s mean delay and a 5–16 s range. No false alarm or cross-detection occurred, and fault timing was not supplied to RAPIO.
The evaluation is extended from LoE to additive bias and stuck/jam faults without changing RAPIO. Using Equation (16), these faults are applied only to the plant. The 18 jam events cover three channels, nominal and stressed conditions, and three seeds. Table 15 reports 18/18 detections and paired 100%-detection bias boundaries of 1300 RPM, 15 ° , and 22.5 ° . The biases are additive actuator offsets, and nominal surge also produced one detection at 1137.5 RPM. Across all 60 runs, no false alarms occurred and the Metzler certificate was retained. LoE is more demanding because it scales with the command and can be partly compensated by TNMPC, whereas bias persists and a jam breaks command tracking. Representative schedule-blind timelines appear in Figure 7 and Figure 8.

5.5. Comparative Benchmark Evaluation

RAPIO is benchmarked against a sign-based adaptive interval observer [25] and P2P threshold FDI [20] under identical conditions. All three methods use the same AUV dynamics, TNMPC controller, sensor noise, fault-injection protocol, and 72-case scenario grid. Figure 9, Figure 10 and Figure 11 compare the channel responses, and Table 16 summarizes the aggregate metrics. The adaptive-IO baseline transfers the reference method’s sign-based adaptive-bound mechanism to the actuator-sensitive AUV diagnostic channels. It therefore serves as a mechanism-level baseline; the original sensor-fault, linear-system application is not treated as equivalent to the present nonlinear actuator-fault problem.
The unchanged detector therefore identifies these command-path faults through persistent interval inconsistencies and characterizes the magnitude and persistence needed for confirmation.
Table 10 reports RAPIO’s channel-wise event counts, false alarms, and delay distributions, whereas Table 16 organizes the same metrics by method and channel for direct comparison.
The benchmark highlights a reliability–speed trade-off. RAPIO is not the fastest detector in surge or pitch, but it is the only method achieving 100 % detection on all three channels with zero event-level false alarms; the increased surge delay is the direct cost of this false-alarm immunity. The confirmation gate suppresses the Coriolis-induced transient residuals responsible for the 12 yaw and 34 pitch false alarms in the adaptive IO [25]. The P2P threshold FDI [20] responds faster but misses 25 % of yaw faults and retains 9 surge and 4 pitch false-alarm events.
RAPIO addresses the coupled AUV case through a bounded propagated tube. The physics-based envelope ϕ d ( t ) and the learned β ( t ) modulate only the admissible width, while the scheduled gain and alarm thresholds remain fixed after calibration. The resulting detector favors persistent interval-consistency evidence over immediate threshold crossings; this increases the mean surge delay to 11.250 ± 2.161 s, but avoids nuisance alarms under hydrodynamic uncertainty and coupled surge–yaw–pitch motion. Across the completed sequential-fault campaign, RAPIO detected all 216 injected fault events without producing an alarm rising edge outside an active fault window.

6. Discussion

The central strength of RAPIO lies in allowing uncertainty bounds to adapt while keeping the diagnostic logic physically interpretable and verifiable. This separation places the method between conservative fixed-bound interval observers and fully adaptive decision rules. A fixed worst-case envelope can preserve inclusion, but it often produces an unnecessarily wide tube. Directly adapting the residual thresholds creates a different concern because coupled transients may then influence the fault decision. RAPIO balances these effects by restricting the learned signal β ( t ) to a bounded adjustment of the tube slack. The nominal center, scheduled gain, Metzler/Hurwitz certificate, and alarm gates remain model-based. Consequently, learning regulates the admissible uncertainty without changing either the physical interpretation of a fault or the conditions used to verify fault-free interval inclusion.
The common-grid benchmark in Table 16 supports this design choice. Among the compared methods, only RAPIO achieves complete detection in all three channels without event-level false alarms. P2P responds more quickly, but it misses yaw faults and produces false alarms in the surge and pitch channels. The adaptive IO detects all faults, although nuisance alarms remain in the coupled yaw and pitch channels. RAPIO’s longer surge delay has a different origin from a general computational delay. TNMPC initially attenuates the propulsion mismatch, and the persistence gate confirms the event after the tube violation becomes sustained. This behavior favors reliable isolation in long-duration AUV operation, where a transient departure may also accompany ordinary disturbance recovery.
The fault-characterization campaigns further show that detectability depends on fault magnitude, temporal development, and closed-loop sensitivity. The channel boundaries in Table 12 arise from differences in control authority and hydrodynamic coupling. Partial losses can remain compensated until their effect on the monitored state exceeds the admissible tube. This behavior also explains the influence of the fault profile. A gradual fault needs time to develop a persistent signature, whereas the healthy gaps in an intermittent profile allow accumulated evidence to decay. For the tested 50% LoE surge cases, Table 14 shows that a 17 s active-fault duration was sufficient for full first-burst detection. By contrast, additive bias and jammed actuation create sustained disagreement between the commanded and realized response. The unchanged detector therefore confirms every tested stuck/jam event and the full-detection bias boundaries reported in Table 15. These results show that the same interval-consistency logic responds to distinct command-path fault structures without receiving the fault schedule or class.
Taken together, the analytical and simulation results define a clear operating domain for RAPIO. The evaluation covers a nonlinear 6-DOF plant, bounded-error FTESO reconstruction, paired control-surface actuation, large hydrodynamic variation, current disturbances, and multiple seeded fault profiles. Within this setting, the model-based certificate establishes fault-free enclosure under the stated disturbance and local-freezing assumptions. The severity and persistence sweeps complement that result by identifying the demonstrated detection domain. Hardware-in-the-loop testing is the next stage for evaluating real-time implementation effects. The present fault scope covers LoE, additive-bias, and stuck/jammed command-path faults. Actuator-dynamics and hydrodynamic-parameter faults require extended actuator and parameter-estimation models.

7. Conclusions

This paper presented a robust adaptive propagated interval observer (RAPIO) for actuator fault detection and isolation in underactuated AUVs under bounded disturbances and hydrodynamic uncertainty. RAPIO combines a Metzler-certified interval observer, Van Loan width propagation, interval-consistency residuals, and a fixed-threshold leaky decision layer. Reinforcement learning supplies the bounded scalar signal β ( t ) that modulates the admissible uncertainty width.
Under the stated assumptions, separating learning from the observer core preserves positivity and interval inclusion. The scheduled observer matrix remains Metzler and Hurwitz, while the half-width dynamics remain positive and uniformly ultimately bounded. The true state also remains enclosed in the fault-free case. The certified tube supports channel-wise violations, persistence confirmation, and recovery logic for surge, yaw-rate, and pitch-rate actuator fault isolation under LoE, additive-bias, and stuck/jammed actuation.
Across 72 severe sequential-fault cases spanning ± 50 % hydrodynamic variation and 0.3 m / s base currents with time-varying, depth-shear, and spatial components, together with the additional severity, timing, bias, and stuck/jam campaigns, RAPIO produced no false alarms. It detected every fault in the 72-case sequential campaign and all 18 stuck/jam events. Mean surge, yaw-rate, and pitch-rate detection delays in the 72-case sweep were 11.25, 3.60, and 7.03 s, respectively. Against the adaptive IO [25] and P2P threshold FDI [20], RAPIO achieved the strongest reliability profile, with 100 % detection on all channels and zero event-level false alarms. In contrast, the baselines produced yaw/pitch false alarms or missed yaw faults.
These results support RAPIO as a safety-oriented diagnostic layer for long-duration AUV missions, where false alarms can be more disruptive than modest confirmation delays. Future work will address degraded sensing, joint actuator–sensor diagnosis, fault-tolerant control, and hardware-in-the-loop validation.

Author Contributions

Conceptualization, I.A., A.A. and J.L.; methodology, I.A., A.A. and J.L.; software, I.A.; validation, I.A., A.A. and J.L.; formal analysis, I.A., A.A. and J.L.; investigation, I.A.; resources, A.A., A.J., M.B. and J.L.; data curation, I.A.; writing—original draft preparation, I.A.; writing—review and editing, I.A., A.A., J.L., A.J. and M.B.; visualization, I.A.; supervision, J.L.; project administration, J.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research work was funded by Umm Al-Qura University, Saudi Arabia, under grant number 26UQU4290339GSSR02

Data Availability Statement

The full simulation code and implementation are part of an ongoing research program and are not publicly available at this time. Further inquiries may be directed to the corresponding author.

Acknowledgments

The authors extend their appreciation to Umm Al-Qura University, Saudi Arabia, for funding this research work through grant number 26UQU4290339GSSR02.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Liu, C.; Filaretov, V.; Zuev, A.; Protsenko, A.; Zhirabok, A. Fault tolerant control in underwater vehicles. J. Mar. Sci. Eng. 2024, 12, 1836. [Google Scholar] [CrossRef]
  2. Wang, Y.; Yu, G.; Xie, W.; Zhang, W.; Silvestre, C. Cooperative path following control of a team of Quadrotor-Slung-Load Systems under disturbances. IEEE Trans. Intell. Veh. 2023, 8, 4169–4179. [Google Scholar] [CrossRef]
  3. Shan, Y.; Zhu, F. Interval observer-based fault tolerant control strategy with fault estimation and compensation. Asian J. Control 2022, 24, 895–906. [Google Scholar] [CrossRef]
  4. Hao, L.Y.; Yang, X.; Zhang, Y.Q.; Atajan, H.; Liu, Y. Fault-tolerant control of unmanned marine vehicles with unknown parametric dynamics via integral sliding mode output feedback control. Asian J. Control 2025, 27, 840–851. [Google Scholar] [CrossRef]
  5. Zhou, M.; Li, H.; Wang, J.; Raïssi, T. Event-triggered H-fault-tolerant consistency control for multi-agent systems. Eur. J. Control 2024, 79, 101083. [Google Scholar] [CrossRef]
  6. Gao, Z.; Cecati, C.; Ding, S.X. A Survey of Fault Diagnosis and Fault-Tolerant Techniques—Part I: Fault Diagnosis with Model-Based and Signal-Based Approaches. IEEE Trans. Ind. Electron. 2015, 62, 3757–3767. [Google Scholar] [CrossRef]
  7. Lan, J.; Patton, R.J. Robust Integration of Model-Based Fault Estimation and Fault-Tolerant Control; Springer Nature: Cham, Switzerland, 2021. [Google Scholar] [CrossRef]
  8. Chanthery, E.; Travé-Massuyès, L.; Indra, S. Fault isolation on request based on decentralized residual generation. IEEE Trans. Syst. Man Cybern. Syst. 2016, 46, 598–610. [Google Scholar] [CrossRef]
  9. Blanke, M.; Kinnaert, M.; Lunze, J.; Staroswiecki, M. Diagnosis and Fault-Tolerant Control; Springer Series in Control Engineering; Springer: Berlin/Heidelberg, Germany, 2006. [Google Scholar]
  10. Efimov, D.; Raïssi, T. Design of Interval Observers for Uncertain Dynamical Systems. Autom. Remote Control 2016, 77, 191–225. [Google Scholar] [CrossRef]
  11. Raïssi, T.; Ramdani, N.; Candau, Y. Set Membership State and Parameter Estimation for Systems Described by Nonlinear Differential Equations. Automatica 2004, 40, 1771–1777. [Google Scholar] [CrossRef]
  12. Chevet, T.; Dinh, T.N.; Marzat, J.; Raïssi, T. Robust Sensor Fault Detection for Linear Parameter-Varying Systems using Interval Observer. In Proceedings of the 31st European Safety and Reliability Conference, Angers, France, 19–23 September 2021; pp. 1486–1493. [Google Scholar] [CrossRef]
  13. Akremi, R.; Lamouchi, R.; Amairi, M.; Dinh, T.N.; Raïssi, T. Functional interval observer design for multivariable linear parameter-varying systems. Eur. J. Control 2023, 71, 100794. [Google Scholar] [CrossRef]
  14. Mizouri, H.; Lamouchi, R.; Amairi, M.; Raissi, T. Functional Interval Observers-Based Fault-Tolerant Control for Linear Parameter-Varying Systems. Int. J. Robust Nonlinear Control 2026, 36, 155–171. [Google Scholar] [CrossRef]
  15. Li, J.; Wang, Z.; Raïssi, T.; Shen, Y. Unknown Input Observer Design for Linear Parameter-Varying Systems in a Bounded Error Context. IEEE Trans. Autom. Control 2021, 66, 4246–4251. [Google Scholar] [CrossRef]
  16. Zhang, Z.H.; Yang, G.H. Distributed fault detection and isolation for multiagent systems: An interval observer approach. IEEE Trans. Syst. Man Cybern. Syst. 2020, 50, 2220–2230. [Google Scholar] [CrossRef]
  17. Zhang, R.; Wang, Z.; Meslem, N.; Raïssi, T.; Shen, Y. Fault Detection for Lipschitz Nonlinear Systems Using Robust Observer and Ellipsoidal Analysis. Int. J. Robust Nonlinear Control 2025, 35, 4523–4535. [Google Scholar] [CrossRef]
  18. Zhu, F.; Fu, Y.; Dinh, T.N. Asymptotic convergence unknown input observer design via interval observer. Automatica 2023, 147, 110744. [Google Scholar] [CrossRef]
  19. Zhu, F.; Li, M. Distributed Interval Observer and Distributed Unknown Input Observer Designs. IEEE Trans. Autom. Control 2024, 69, 8868–8875. [Google Scholar] [CrossRef]
  20. Ma, Y.; Wang, Z.; Meslem, N.; Raïssi, T.; Shen, Y. Fault diagnosis by interval-based adaptive thresholds and peak-to-peak observers. Int. J. Adapt. Control Signal Process. 2023, 37, 519–537. [Google Scholar] [CrossRef]
  21. Tian, M.; Guo, Z.; Wang, X.; Niu, B.; Wang, X.; Wang, D.; Wang, H. Fault Detection and Isolation for Multiagent Systems Under Event-Triggered Communication: A Zonotopic Joint Estimation Method. IEEE Internet Things J. 2025, 12, 29940–29947. [Google Scholar] [CrossRef]
  22. Zhu, X.; Li, Y.; Yin, G.; Patton, R.J. Interval Observer-Based Fault Detection and Isolation for Quadrotor UAV With Cable-Suspended Load. IEEE Trans. Syst. Man Cybern. Syst. 2024, 54, 5876–5888. [Google Scholar] [CrossRef]
  23. Guzmán-Rabasa, J.; Rodríguez, F.; Valencia-Palomo, G.; Santos-Ruiz, I.; Gómez-Peñate, S.; López-Estrada, F.R. Convex Fault Diagnosis of a Three-Degree-of-Freedom Mechanical Crane. Mathematics 2023, 11, 4258. [Google Scholar] [CrossRef]
  24. Song, J.; He, X. Robust state estimation and fault detection for Autonomous Underwater Vehicles considering hydrodynamic effects. Control Eng. Pract. 2023, 135, 105497. [Google Scholar] [CrossRef]
  25. Wang, X.; Pin Tan, C.; Wang, Y.; Zhang, Z. Active fault tolerant control based on adaptive interval observer for uncertain systems with sensor faults. Int. J. Robust Nonlinear Control 2021, 31, 2857–2881. [Google Scholar] [CrossRef]
  26. Yan, Z.; Xu, F.; Tan, J.; Liu, H.; Liang, B. Reinforcement learning-based integrated active fault diagnosis and tracking control. ISA Trans. 2023, 132, 364–376. [Google Scholar] [CrossRef] [PubMed]
  27. Ali, S.M.; Rabbani, T.; Bilal, M.; Alharbi, A.; Aldosari, F.M.; Amin, R.U. Dynamics Estimation and Deep Reinforcement Learning based Fractal Inspired Sliding Mode Control of Stewart Platform. Results Control Optim. 2026, 100779. [Google Scholar] [CrossRef]
  28. Li, Z.; Wang, M.; Ma, G.; Zou, T. Adaptive reinforcement learning fault-tolerant control for AUVs with thruster faults based on the integral extended state observer. Ocean Eng. 2023, 271, 113722. [Google Scholar] [CrossRef]
  29. Jmal, A.; Naifar, O.; Ben Makhlouf, A.; Derbel, N.; Hammami, M.A. Observer-based model reference control for linear fractional-order systems. Int. J. Digit. Signals Smart Syst. 2018, 2, 136–149. [Google Scholar] [CrossRef]
  30. Naifar, O.; Jmal, A.; Ben Makhlouf, A.; Derbel, N.; Hammami, M.A. Observer-based control for fractional-order systems. In Fractional Order Systems—Control Theory and Applications: Fundamentals and Applications; Springer: Cham, Switzerland, 2021; pp. 75–93. [Google Scholar] [CrossRef]
  31. Jmal, A.; Naifar, O.; Rhaima, M.; Ben Makhlouf, A.; Mchiri, L. On Observer and Controller Design for Nonlinear Hadamard Fractional-Order One-Sided Lipschitz Systems. Fractal Fract. 2024, 8, 606. [Google Scholar] [CrossRef]
  32. Alsharif, A.O.M.; Jmal, A.; Naifar, O.; Ben Makhlouf, A.; Rhaima, M.; Mchiri, L. Unknown Input Observer Scheme for a Class of Nonlinear Generalized Proportional Fractional Order Systems. Symmetry 2023, 15, 1233. [Google Scholar] [CrossRef]
  33. Ahmed, H.; Jmal, A.; Ben Makhlouf, A. A practical observer for state and sensor fault reconstruction of a class of fractional-order nonlinear systems. Eur. Phys. J. Spec. Top. 2023, 232, 2437–2443. [Google Scholar] [CrossRef]
  34. Fossen, T.I. Handbook of Marine Craft Hydrodynamics and Motion Control; John Wiley & Sons: Chichester, UK, 2011. [Google Scholar]
  35. Wang, Z.; Shen, Y. Model-Based Fault Diagnosis: Methods for State-Space Systems; Studies in Systems, Decision and Control; Springer: Singapore, 2023; Volume 221. [Google Scholar] [CrossRef]
  36. Prestero, T. Development of a six-degree of freedom simulation model for the REMUS autonomous underwater vehicle. In Proceedings of the MTS/IEEE Oceans 2001. An Ocean Odyssey. Conference Proceedings (IEEE Cat. No.01CH37295), Honolulu, HI, USA, 5–8 November 2001; Volume 1, pp. 450–455. [Google Scholar] [CrossRef]
  37. Mazenc, F.; Bernard, O. Interval observers for linear time-invariant systems with disturbances. Automatica 2011, 47, 140–147. [Google Scholar] [CrossRef]
  38. Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning; Dy, J., Krause, A., Eds.; PMLR: New York, NY, USA, 2018; Volume 80, pp. 1861–1870. [Google Scholar]
Figure 1. Underactuated AUV showing the NED and body-fixed frames, ocean-current velocity, body-motion variables, and actuator force/moment channels.
Figure 1. Underactuated AUV showing the NED and body-fixed frames, ocean-current velocity, body-motion variables, and actuator force/moment channels.
Jmse 14 01445 g001
Figure 2. Proposed RAPIO-based FDI architecture. The labeled shared-command branch shows that the same actuator command u ( t ) is supplied to the AUV plant, FTESO, and observer state propagation. The plant/reconstructor output y ( t ) , FTESO disturbance estimate d ^ ( t ) , interval bounds, and learned modulation β ( t ) are separate signals.
Figure 2. Proposed RAPIO-based FDI architecture. The labeled shared-command branch shows that the same actuator command u ( t ) is supplied to the AUV plant, FTESO, and observer state propagation. The plant/reconstructor output y ( t ) , FTESO disturbance estimate d ^ ( t ) , interval bounds, and learned modulation β ( t ) are separate signals.
Jmse 14 01445 g002
Figure 3. RL-augmented bound modulation architecture. The RAPIO observer receives a bounded scalar modulation signal β ( t ) from the learned supervisory policy. This signal modulates only the slack term in the interval width dynamics, while the scheduled-gain rule and nominal center predictor remain fixed.
Figure 3. RL-augmented bound modulation architecture. The RAPIO observer receives a bounded scalar modulation signal β ( t ) from the learned supervisory policy. This signal modulates only the slack term in the interval width dynamics, while the scheduled-gain rule and nominal center predictor remain fixed.
Jmse 14 01445 g003
Figure 4. Online RAPIO propagation and fault-decision process. The frozen SAC policy modulates only the interval-width slack; the locally frozen model, fixed scheduled-gain rule, and center dynamics remain independent of learning.
Figure 4. Online RAPIO propagation and fault-decision process. The frozen SAC policy modulates only the interval-width slack; the locally frozen model, fixed scheduled-gain rule, and center dynamics remain independent of learning.
Jmse 14 01445 g004
Figure 5. Sequential actuator-fault detection for (a) surge, (b) yaw-rate, and (c) pitch-rate channels under RL–RAPIO interval-consistency monitoring; (d) learned confidence β ( t ) .
Figure 5. Sequential actuator-fault detection for (a) surge, (b) yaw-rate, and (c) pitch-rate channels under RL–RAPIO interval-consistency monitoring; (d) learned confidence β ( t ) .
Jmse 14 01445 g005
Figure 6. Concurrent surge–yaw faults followed by a pitch fault, showing channel-wise isolation in (a) surge, (b) yaw rate, and (c) pitch rate; (d) learned confidence β ( t ) .
Figure 6. Concurrent surge–yaw faults followed by a pitch fault, showing channel-wise isolation in (a) surge, (b) yaw rate, and (c) pitch rate; (d) learned confidence β ( t ) .
Jmse 14 01445 g006
Figure 7. Schedule-blind RAPIO response to detected additive actuator-bias faults: (a) surge, (b) yaw, (c) pitch, and (d) learned confidence β ( t ) .
Figure 7. Schedule-blind RAPIO response to detected additive actuator-bias faults: (a) surge, (b) yaw, (c) pitch, and (d) learned confidence β ( t ) .
Jmse 14 01445 g007
Figure 8. Schedule-blind RAPIO response to detected stuck/jammed actuator faults: (a) surge, (b) yaw, (c) pitch, and (d) learned confidence β ( t ) .
Figure 8. Schedule-blind RAPIO response to detected stuck/jammed actuator faults: (a) surge, (b) yaw, (c) pitch, and (d) learned confidence β ( t ) .
Jmse 14 01445 g008
Figure 9. Surge-channel benchmark: (a) RAPIO, (b) adaptive IO [25], and (c) P2P threshold FDI [20].
Figure 9. Surge-channel benchmark: (a) RAPIO, (b) adaptive IO [25], and (c) P2P threshold FDI [20].
Jmse 14 01445 g009
Figure 10. Yaw-rate-channel benchmark: (a) RAPIO, (b) adaptive IO [25], and (c) P2P threshold FDI [20].
Figure 10. Yaw-rate-channel benchmark: (a) RAPIO, (b) adaptive IO [25], and (c) P2P threshold FDI [20].
Jmse 14 01445 g010
Figure 11. Pitch-channel benchmark: (a) RAPIO, (b) adaptive IO [25], and (c) P2P threshold FDI [20].
Figure 11. Pitch-channel benchmark: (a) RAPIO, (b) adaptive IO [25], and (c) P2P threshold FDI [20].
Jmse 14 01445 g011
Table 1. Qualitative effect of varying the disturbance-envelope parameters η and ε from the nominal design. Detectability margin references (60); false-alarm risk is assessed relative to the nominal zero-false-alarm benchmark result.
Table 1. Qualitative effect of varying the disturbance-envelope parameters η and ε from the nominal design. Detectability margin references (60); false-alarm risk is assessed relative to the nominal zero-false-alarm benchmark result.
η ε Inclusion MarginDetectability EffectFalse-Alarm Risk
0.250.25Reduced; tighter tubeLower minimum detectable fault; improved sensitivityIncreased if true disturbance exceeds the tighter envelope
0.500.50Nominal (scenario-sweep verified)Minimum detectable: 0.18 m/s, 0.09 rad/s, 0.42 m/sZero across all 72 scenario cases
0.750.75Enlarged; wider tubeHigher minimum detectable fault; reduced sensitivityNear zero; wider tube keeps healthy residuals well within bounds, but the same envelope also absorbs small fault signatures (see Detectability effect)
Table 2. Sign and diagnostic impact of the learned β k signal. Neither role modifies the scheduled observer gain L k , frozen Jacobian A k , nominal center predictor, or detector gate.
Table 2. Sign and diagnostic impact of the learned β k signal. Neither role modifies the scheduled observer gain L k , frozen Jacobian A k , nominal center predictor, or detector gate.
RL Signal/Observer QuantityTube EffectDiagnostic Impact
β k (Higher)Smaller confidence slack; propagated tube tightens toward the calibrated disturbance boundInterval exits are less masked by added slack and can reach the fixed accumulator threshold earlier
β k (Lower)Larger confidence slack; propagated tube widens under uncertaintyMore conservative enclosure under uncertain operation; stronger suppression of nuisance residuals before the fixed threshold is evaluated
L , ( A k L C ) UnchangedObserver stability and interval-inclusion proof remain independent of learning
Table 3. Principal AUV hydrodynamic and inertial parameters.
Table 3. Principal AUV hydrodynamic and inertial parameters.
ParameterValueParameterValue
m30.581 kg I x x 0.177 kg m2
I y y 3.45 kg m2 I z z 3.45 kg m2
X u ˙ 0.93 kg Y v ˙ 35.5 kg
Z w ˙ 35.5 kg K p ˙ 0.0704 kg m2
M q ˙ 4.88 kg m2 N r ˙ 4.88 kg m2
X u u 1.62 kg/m Y v v 1310 kg/m
Z w w 131 kg/m K p p 0.13 kg m2
M q q 188 kg m2 N r r 94 kg m2
Table 4. RAPIO interval observer parameters.
Table 4. RAPIO interval observer parameters.
GroupParameterValue
TimingPlant step Δ t 0.1 s
TimingDiagnostic period T s 1.0 s
Certificate A o , k conditionMetzler/Hurwitz at each update
Diagonal gain u , r , q 1.5 , 0.6 , 5.0 s−1
Diagonal gain z , other 1.0 , 5.0 s−1
Off-diagonal gainCooperative margin μ c 0.05
Center correction γ 0.18
Uncertainty envelope Δ max , ε 0.5 , 0.5
Width limits w ¯ , w ¯ r , w ̲ 1.0 , 0.8 , 10 3
Beta slack D i slack Channel-specific offline values
Table 5. Key SAC training parameters.
Table 5. Key SAC training parameters.
ParameterValue
Observation dimension/Action dimension10/11
Network/Normalization/Curriculum [ 256 , 256 ] MLP/none/none
Control-action bounds/beta-action bounds [ 0.75 , 1.5 ] / [ 1 , 1 ]
Hyperparameters (Learning rate, γ SAC , ρ Polyak ) 3 × 10 4 , 0.99, 0.005
Reward penalties (beta, missed fault, false alarm, chatter) 0.1 , 100, 50, 0.5
Steps/Buffer/Parallel envs/Base seed 5 × 10 5 / 10 5 /5/42
Training uncertainty/Safety margin 20 % / ε = 0.5
Table 6. Injected diagnostic fault signatures for the sequential schedule. Magnitudes are reported as actuator LoE levels.
Table 6. Injected diagnostic fault signatures for the sequential schedule. Magnitudes are reported as actuator LoE levels.
FaultInterval [s] G d EntryLoE Magnitude
Surge50–80row 1, column 1 (u)50%
Yaw130–160row 6, column 2 (r)100%
Pitch210–240row 5, column 3 (q)100%
LoE magnitude is 100 ( 1 α i ) % , where α i is the retained actuator effectiveness in the actuator LoE model of Section 2.1.
Table 7. Per-channel IO-FDI detection parameters.
Table 7. Per-channel IO-FDI detection parameters.
ParameterSurge uYaw rPitch q
Disturbance bound D i 0.210 m/s0.095 rad/s0.110 rad/s
Measurement margin0.125 m/s0.030 rad/s0.003 rad/s
Beta slack D i slack 0.500 m/s0.150 rad/s0.200 rad/s
Alarm threshold τ i base 0.080.020.35
Clearance level0.0020.0020.002
Trigger count333
Leak factor0.880.880.88
Alarm hold [s]585
Freeze dwell T i min [s]566
Recovery ratio ρ i rec 0.700.700.70
Recovery level δ i rec 0.0200.0100.015
Recovery count N i rec 222
Table 8. Fault-characterization campaigns.
Table 8. Fault-characterization campaigns.
CampaignVariationPaired SamplingCases
Severity10–100% LoE; three actuators2 environments × 3 seeds180
Fault profileRandomized abrupt, gradual, intermittent3 channels × 3 profiles × 2 environments × 3 seeds54
Surge persistence8–18 s bursts at 50% LoE7 durations × 2 environments × 3 seeds42
Bias and stuck/jamReference bias/stuck cases plus 0.25 2.0 bias-magnitude sweep6 reference-bias + 6 stuck/jam runs (18 jam events) + 48 bias-magnitude runs60
Table 9. RAPIO/IO disturbance-floor sensitivity sweep results ( N = 72 scenario cases per scale; all detector parameters fixed).
Table 9. RAPIO/IO disturbance-floor sensitivity sweep results ( N = 72 scenario cases per scale; all detector parameters fixed).
ScaleDR [%]FA/Case uFA/Case rFA/Case qDelay u
[s]
Delay r
[s]
Delay q
[s]
Metzler
OK [%]
0.5 × (Low)100.00.000.000.00 11.31 ± 2.13 3.75 ± 0.44 8.06 ± 2.14 100.0
1.0 × (Nominal)100.00.000.000.00 11.28 ± 2.20 3.74 ± 0.44 7.93 ± 2.00 100.0
2.0 × (High)100.00.000.000.00 11.24 ± 2.16 3.74 ± 0.44 8.33 ± 2.37 100.0
DR = detection rate; FA/Case = false-alarm events per scenario case (rising-edge, healthy time only); delays are in seconds and are reported as mean ± standard deviation over detected cases. The Metzler OK column reports the percentage of cases for which the logged frozen-matrix certificate held at every diagnostic update.
Table 10. Sequential actuator-fault campaign after detector recovery update ( N = 72 scenario cases; all parameters fixed).
Table 10. Sequential actuator-fault campaign after detector recovery update ( N = 72 scenario cases; all parameters fixed).
MetricSurge (u)Yaw (r)Pitch (q)
Detection rate [%]100.0100.0100.0
Detected events/total72/7272/7272/72
Rising-edge false alarms000
Delay [s] 11.25 ± 2.16 3.60 ± 0.49 7.03 ± 1.49
Delay range [s]7–163–45–12
Post-clearance alarm tail [s] 3.83 mean, 5 max 1.93 mean, 8 max 2.08 mean, 8 max
Delays are measured on the 1 s IO-FDI diagnostic update grid. Rising-edge false alarms count new alarms outside active fault windows. Post-clearance tails are continuing alarms after a detected fault has vanished, not new false re-arms.
Table 11. Concurrent fault isolation results under overlapping surge–yaw fault scenario ( N = 72 scenario cases; surge–yaw overlap window: t = 70 –90 s).
Table 11. Concurrent fault isolation results under overlapping surge–yaw fault scenario ( N = 72 scenario cases; surge–yaw overlap window: t = 70 –90 s).
MetricSurge (u)Yaw (r)Pitch (q)
Detection Rate [%]100.0100.0100.0
FA/Case0.000.000.00
Mean Delay [s] 9.50 ± 1.69 4.29 ± 0.57 5.25 ± 1.02
Yaw DR in overlap window [%]100.0
Table 12. Empirical LoE severity characterization (six paired cases per level).
Table 12. Empirical LoE severity characterization (six paired cases per level).
MetricSurge uYaw rPitch q
0% eventual-DR range10–20%10–60%10–70%
Transition range30–40%: 50% DR80%: 16.7%; 90%: 50% DR
Minimum 100% eventual-DR level50% LoE70% LoE100% LoE
DR at boundary6/66/66/6
Mean delay [s]11.334.006.83
False alarms000
Table 13. Detection statistics for randomized and time-varying LoE profiles.
Table 13. Detection statistics for randomized and time-varying LoE profiles.
Channel and ProfileEventual DRMean Delay [s]FA/Cross
Surge—randomized abrupt100.0%11.500/0
Surge—gradual100.0%22.830/0
Surge—intermittent50.0%16.670/0
Yaw—randomized abrupt100.0%4.500/0
Yaw—gradual83.3%22.000/0
Yaw—intermittent100.0%3.000/0
Pitch—randomized abrupt100.0%6.000/0
Pitch—gradual100.0%24.000/0
Pitch—intermittent100.0%7.500/0
All channels and profiles92.6%0/0
Table 14. First-burst persistence characterization for intermittent 50% surge LoE.
Table 14. First-burst persistence characterization for intermittent 50% surge LoE.
Duration [s]First-Burst DetectionEventual DetectionMean Delay [s]Maximum Delay [s]
81/63/67.007
102/63/68.009
123/64/69.0011
143/65/69.0011
165/65/611.4015
176/66/612.1716
186/66/612.1716
Table 15. Schedule-blind additive-bias and stuck/jam detection results (nominal and stressed environments, three seeds each; detector parameters unchanged).
Table 15. Schedule-blind additive-bias and stuck/jam detection results (nominal and stressed environments, three seeds each; detector parameters unchanged).
MetricSurge (u)Yaw (r)Pitch (q)
Nominal: last 0% bias DR975 RPM 7.5 ° 18 °
Nominal: first 100% bias DR1300 RPM 10 ° 22.5 °
Stressed: last 0% bias DR325 RPM 12.5 ° 18 °
Stressed: first 100% bias DR487.5 RPM 15 ° 22.5 °
Stuck/jam detections6/66/66/6
Stuck/jam mean delay [s]9.02.03.0
Table 16. Three-method FDI benchmark over the common N = 72 scenario cases. Each row gives the event count, detection rate, total healthy-time false-alarm events, and detection-delay distribution for one method–channel pair.
Table 16. Three-method FDI benchmark over the common N = 72 scenario cases. Each row gives the event count, detection rate, total healthy-time false-alarm events, and detection-delay distribution for one method–channel pair.
FDI TechniqueChannelEventsDR [%]FADelay [s]
Proposed RAPIOSurge (u)72/72100.00 11.250 ± 2.161
Yaw (r)72/72100.00 3.597 ± 0.494
Pitch (q)72/72100.00 7.028 ± 1.491
Adaptive IO [25]Surge (u)72/72100.00 3.000 ± 0.000
Yaw (r)72/72100.012 6.486 ± 0.692
Pitch (q)72/72100.034 9.153 ± 3.531
P2P threshold FDI [20]Surge (u)72/72100.09 1.208 ± 1.198
Yaw (r)54/7275.00 1.667 ± 0.476
Pitch (q)72/72100.04 2.403 ± 0.781
DR = event-level detection rate; FA = healthy-time alarm rising edges; delay is mean ± standard deviation over detected events.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ahmed, I.; Alharbi, A.; Lu, J.; Jaffar, A.; Bilal, M. Robust Adaptive Propagated Interval Observer for Actuator Fault Diagnosis in Underactuated AUVs. J. Mar. Sci. Eng. 2026, 14, 1445. https://doi.org/10.3390/jmse14151445

AMA Style

Ahmed I, Alharbi A, Lu J, Jaffar A, Bilal M. Robust Adaptive Propagated Interval Observer for Actuator Fault Diagnosis in Underactuated AUVs. Journal of Marine Science and Engineering. 2026; 14(15):1445. https://doi.org/10.3390/jmse14151445

Chicago/Turabian Style

Ahmed, Ishaq, Ayman Alharbi, Jun Lu, Amar Jaffar, and Muhammad Bilal. 2026. "Robust Adaptive Propagated Interval Observer for Actuator Fault Diagnosis in Underactuated AUVs" Journal of Marine Science and Engineering 14, no. 15: 1445. https://doi.org/10.3390/jmse14151445

APA Style

Ahmed, I., Alharbi, A., Lu, J., Jaffar, A., & Bilal, M. (2026). Robust Adaptive Propagated Interval Observer for Actuator Fault Diagnosis in Underactuated AUVs. Journal of Marine Science and Engineering, 14(15), 1445. https://doi.org/10.3390/jmse14151445

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop