Next Article in Journal
Evaluating the Periodic Sustainability of Cislunar Logistics Architectures: A Reproducible Methodology with an Artemis III–Derived Case Study
Previous Article in Journal
Preview-Aware LSTM-Assisted Predictive Control for Turboshaft Engines Under Tiltrotor Conversion-Flight Power Demand
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Multi-History-Weighted Projection-Adaptive Jacobian Control for Close-Range UAV Visual Servoing near Overhead Ground Wires

School of Electrical Engineering and Automation, Nantong University, Seyuan Road 9, Nantong 226019, China
*
Author to whom correspondence should be addressed.
Aerospace 2026, 13(9), 844; https://doi.org/10.3390/aerospace13090844
Submission received: 23 July 2026 / Revised: 11 September 2026 / Accepted: 14 September 2026 / Published: 16 September 2026
(This article belongs to the Section Aeronautics)

Abstract

Close-range UAV visual servoing near overhead ground wires requires simultaneous regulation of the target’s lateral position, apparent width, and orientation angle in the image. Owing to flight-control response lag and visual-processing delay, current image changes may reflect the combined effects of multiple historical control inputs. To address this issue, this paper proposes a multi-history-weighted projection-adaptive Jacobian control method (MHW-PAJ-CLF). For each visual channel, the method takes a weighted sum of the parameter corrections associated with different historical inputs based on response prediction errors and input magnitudes and updates the local Jacobian matrix online under projection bounds. The online estimate is then blended with the nominal model. Control Lyapunov function-based quadratic programming (CLF-QP) uses the blended model to generate velocity and yaw-rate commands, coordinating the regulation of the three visual errors under input constraints. Comparative experiments were conducted on the RflySim–PX4 hardware-in-the-loop platform. The proposed method achieves a mean success rate of 75.50%, exceeding that of the best baseline by 11.05 percentage points. In tests with additional visual-feedback delay and command–response lag, the proposed method achieves lower overall tracking error than single-history PAJ-CLF, indicating that multi-history weighting helps improve visual tracking under time delays.

1. Introduction

Overhead ground wires are key components of lightning-protection systems for transmission lines, and their operating conditions directly affect transmission-line safety and power-supply reliability. Compared with manual inspection, UAV-based inspection can reduce risks to personnel during high-voltage and long-distance operations and improve inspection efficiency. Existing work on UAV-based power-line inspection covers inspection methods and engineering challenges [1], visual assistance for inspection and landing [2], UAV-LiDAR-based intelligent inspection [3], perception-aware control for line inspection [4], and multi-UAV cooperative inspection and maintenance [5]. However, close-range visual servoing near an overhead ground wire differs from conventional waypoint tracking. The UAV must maintain a line-shaped target within the camera field of view while simultaneously regulating its lateral position, apparent width, and orientation angle. In ultra-high-voltage electromagnetic environments, external positioning information or relative-range measurements may be unreliable, making close-range visual servoing more dependent on onboard visual measurements. Image sampling, target detection, filtering, and actuation jointly introduce a temporal lag between control input and visual response. This paper investigates coordinated regulation of multidimensional image errors under this temporal lag in close-range visual servoing near overhead ground wires.
Reliable extraction of the line target is the prerequisite for this task. Zhang et al. proposed an attention-guided multitask network for power-line component detection [6]. Liu et al. developed a deep-learning framework for key-target and defect detection on high-voltage transmission lines [7]. Choudhary et al. designed a lightweight detector for power-line components in UAV images [8]. Xu et al. improved orientation representation for rotated targets through instance-level orientation enhancement [9]. These advances provide the image measurements required for close-range control, but the detected line features must still be converted into coordinated flight commands. Beyond perception alone, Lin and Lai integrated YOLO-based runway detection and CNN-based localization with guidance and flight-control laws for fixed-wing UAV landing in GNSS-denied environments [10].
Image-based visual servoing (IBVS) generates motion commands directly from image-feature errors and therefore reduces dependence on complete three-dimensional reconstruction. Wang et al. extended IBVS to quadrotors tracking arbitrary flight targets [11]. Li et al. developed a robocentric model-based visual-servoing framework for quadrotor flight [12]. Xie et al. constructed a UAV visual-servoing controller using a virtual camera to accommodate changes in system parameters [13]. Lin et al. introduced robust observer-based visual control for tracking unknown moving targets [14]. Image-based feedback has also been applied to moving-platform landing and planar-target tracking [15,16]. Yang et al. combined IBVS with delayed state estimation for high-speed multicopter interception [17]. For measurement delay caused by image acquisition and processing, Yang et al. developed a sampled-data visual-servoing method using a time-delay disturbance observer [18]. Chen et al. addressed unmeasured states and disturbances through robust observation and output feedback [19]. Sepahvand et al. employed a self-organizing neural network to compensate for uncertainty in aerial-robot IBVS [20]. These studies cover target-relative flight in tracking, interception, and landing tasks. Close-range line tracking also requires the local mapping error from control commands to image variation to be addressed.
Constraint-aware, predictive, and learning-based methods have been introduced to improve visual-control performance. Yang and Li incorporated visibility constraints into robust model predictive visual servoing [21]. Zhou et al. combined constrained IBVS with sampling-based planning to handle image and physical constraints [22]. Liu et al. developed a dual-rate predictive observer for low camera-sampling rates and view constraints [23]. Yi et al. used a neural network to compensate for model uncertainty in quadrotor IBVS [24]. More recently, Byun et al. proposed vector-field-based robust landing on a moving platform [25]. Zhang et al. developed data-driven generalized iterative predictive control for nonlinear and coupled UAV dynamics [26]. Guo et al. addressed image-based formation encirclement and tracking of unknown targets [27]. These methods improve visibility maintenance, predictive regulation, and uncertainty compensation. Updating the local image-response model also requires relating the observed image changes to the earlier commands that produced them.
Online estimation offers a complementary means of reducing local model mismatch. Pan and Shi incorporated online data memory into adaptive estimation and control [28]. Li et al. used stored data in composite learning for uncalibrated visual servoing [29]. Kamath and Feroskhan mapped image variations directly to multirotor input commands through augmented dynamics [30]. Kamath et al. subsequently introduced a physics-informed Koopman neural operator for augmented-dynamics visual servoing [31]. These methods use closed-loop data to update the mapping between control input and image response. A projection-adaptive Jacobian (PAJ) update provides a lightweight online correction of the local image-response matrix. The PAJ baseline uses the current image-state change and a single recent control input to update the model. Flight-control response, actuation, and visual processing take time. The current image change may therefore reflect several earlier commands. The relevant input history can vary with the command response and visual channel.
The proposed multi-history-weighted projection-adaptive Jacobian method uses CLF-QP command generation and is termed MHW-PAJ-CLF. For each visual channel, the current model predicts the response to each historical input. Comparing each prediction with the observed state-change rate gives its response error. These errors and the input magnitudes determine the weights used to sum the PAJ parameter corrections. The update includes a correction toward the nominal value, and projection bounds the result. The updated estimate and the nominal model are combined in a prescribed proportion. The CLF-QP uses this model to regulate the three visual errors within the velocity and yaw-rate limits. The main contributions are as follows:
  • A multi-history-weighted projection-adaptive Jacobian update method is proposed to address flight-control response lag and visual-processing delay. For each visual channel, parameter corrections associated with multiple historical inputs are weighted and summed according to response prediction errors and input magnitudes to update the local image-response model online.
  • A three-channel visual-servoing framework is developed for close-range operations near overhead ground wires. The online Jacobian estimate is blended with the nominal model and used by the CLF-QP to coordinate regulation of the target’s lateral position, apparent width, and orientation angle within the velocity and yaw-rate limits.
  • Comparative experiments are conducted on the RflySim–PX4 hardware-in-the-loop platform. The results show that the proposed method achieves better overall visual tracking performance than the comparison baselines, with lower mean absolute visual errors and a higher mean success rate.

2. System Modeling

This section first establishes the UAV dynamic model, then defines the visual errors of the overhead ground wire, and, finally, describes the local relationship between the control input and visual-error variation.

2.1. UAV Modeling

The UAV dynamics are described by a six-degree-of-freedom rigid-body model. The NED frame is used as the inertial frame I , and the FRD frame is used as the body frame B , where x b , y b , and  z b point forward, rightward, and downward, respectively. The UAV position and velocity are expressed in the NED frame as
p u = [ x , y , z ] T ,
v = [ v x , v y , v z ] T .
The attitude is described by the rotation matrix R S O ( 3 ) , which transforms vectors from the body frame B to the NED frame I . The body-frame angular velocity is
ω = [ p , q , r ] T .
The translational and rotational dynamics are
p ˙ u = v ,
m v ˙ = 0 0 m g + R 0 0 T + f d ,
R ˙ = R [ ω ] × ,
J ω ˙ = τ ω × J ω + τ d ,
where m is the UAV mass, J is the inertia matrix, g is the gravitational acceleration, and  f d and τ d denote aerodynamic drag, modeling errors, and external disturbances. The total thrust T and body-frame control torque τ = [ τ x , τ y , τ z ] T are obtained by summing the contributions of the four rotors:
T = i = 1 4 T i ,
τ = i = 1 4 r i × 0 0 T i + 0 0 M i ,
where T i and M i are the thrust and reaction torque generated by the ith rotor, respectively, and  r i is its body-frame position vector relative to the vehicle center of mass. For the ith rotor with angular speed Ω i , T i and M i satisfy the quadratic relations
T i = k T Ω i 2 ,
M i = σ i k M Ω i 2 ,
where k T and k M are the thrust and reaction-torque coefficients, respectively, and  σ i { + 1 , 1 } denotes the rotor rotation direction. The motor response is approximated by a first-order lag:
Ω ˙ i = 1 τ m ( Ω i , ss Ω i ) , Ω i , ss = a Ω u m , i + b Ω ,
where Ω i , ss is the steady-state motor speed corresponding to u m , i , u m , i is the normalized motor-throttle command generated by the mixer, a Ω and b Ω are the slope and intercept of the steady-state throttle–speed relation, and  τ m is the motor response time constant.

2.2. Visual-Error Modeling

The YOLO11s-OBB detector provides the rotated bounding box of the overhead ground wire. Its centerline position, short-side width, and orientation angle define the three image-space quantities used for visual regulation, as illustrated in Figure 1b.
As shown in Figure 1b, the green solid box denotes the current rotated bounding box of the overhead ground wire detected by YOLO11s-OBB, and the green dashed box denotes the desired target box. The two dark-gray parallel lines denote the image boundaries of the same overhead ground wire. Let the image center be c img , the centerline of the current detection box be 𝓁 box , the short-side width of the detection box be w box , and the box orientation angle be ψ box . With desired box width w ref and orientation angle ψ ref , the raw visual-error vector is
e = e x e w e ψ T ,
where e x is the lateral-position error between the image center and the target-box centerline, e w = w box w ref is the apparent-width error, and  e ψ is the angle error between the detected and desired box orientation angles.
Before e ψ is calculated, the detected box orientation is mapped to [ π / 2 , π / 2 ) . When the angle crosses an interval boundary, equivalent wire orientations separated by π are used to correct angle jumps between successive frames.
Within the local regulation region, when the UAV is approximately level and the wire is near the center of the downward-looking image, the apparent wire width increases as the camera–wire distance decreases. It therefore provides visual feedback for regulating the UAV’s vertical position relative to the wire.
As shown in Figure 1a, flight height h and wire height h w share the same height reference. When the UAV is approximately level and the wire is near the center of the downward-looking image, the camera–wire distance is approximately h h w . Camera attitude and viewing direction also affect the apparent width. Calibration was performed during descent at the reference viewpoint, relative yaw offsets of ± 10 , and lateral offsets of ± 0.5 m, with flight height and detected short-side width recorded together (Figure 2). The desired width is set to w ref = 35 px .

2.3. Three-Channel Visual State and Control Input

Because the lateral-position, apparent-width, and angle errors have different physical dimensions and numerical scales, the normalized three-channel visual state is defined to balance the relative contributions of the three channels:
q = e x / s x e w / s w e ψ / s ψ T ,
where s x , s w , and  s ψ are the normalization scales for the lateral-position, apparent-width, and angle errors, respectively. Let
S = diag ( s x , s w , s ψ ) .
The normalized visual state can then be written as
q = S 1 e .
If B e ( t ) denotes the local image-response matrix in the raw-error coordinates, its expression in the normalized visual-state coordinates is
B ( t ) = S 1 B e ( t ) .
Therefore, the local image-response matrix B ( t ) , the online estimate B ^ , and the nominal matrix B 0 are all defined in the same normalized visual-state coordinates.
The UAV control input is defined as
u = v y ref v z ref r ref T ,
where v y ref , v z ref , and  r ref denote the lateral-velocity, vertical-velocity, and yaw-rate reference commands, respectively. The velocity reference commands use the body FRD frame. The three command channels predominantly regulate lateral image alignment, camera–wire distance represented by apparent width, and relative heading represented by the angle.

2.4. Local Image-Response Model

Within the current visual-regulation region, the relationship between the control input and the normalized visual-state rate is locally approximated as
q ˙ = B ( t ) u + η ( t ) ,
where B ( t ) is the local image-response matrix in normalized visual-state coordinates and describes the local effect of the control input u on the visual-state rate q ˙ . The term η ( t ) collects the effects of detection noise, visual filtering, unmodeled dynamics, and other disturbances.
According to the dominant correspondences between lateral velocity, vertical velocity, and yaw rate and the three visual states, the local image-response matrix is decomposed as
B ( t ) = B d ( t ) + B o ( t ) ,
where the diagonal component is defined as
B d ( t ) = diag B 11 ( t ) , B 22 ( t ) , B 33 ( t )
The matrix B d ( t ) contains the three dominant channels, whereas B o ( t ) contains the off-diagonal coupling terms between channels. Equation (19) can therefore be rewritten as
q ˙ = B d ( t ) u + η o ( t ) ,
where η o ( t ) = B o ( t ) u + η ( t ) represents the off-diagonal channel response and other unmodeled terms.
Command transmission, flight-control response, actuation, and visual processing take time. The change between two visual frames may therefore reflect several earlier commands. This effect is expressed as
q k q k 1 = T img , k j = 0 J k H k ( j ) u k 1 j + ζ k ,
where T img , k = t k t k 1 is the interval between adjacent valid visual frames, J k is the currently available history depth, H k ( j ) denotes the unknown linear contribution of the jth historical command to the current visual-state increment, and  ζ k represents the error outside the finite-history approximation.
Within the local visual-regulation region, each channel’s response gain is treated as approximately constant over the short history window. Each historical input in that channel is evaluated using the same current parameter estimate.
Equation (23) describes how several historical commands contribute to the current visual change. For online updating, each historical input is multiplied by the current gain estimate of its channel to obtain the corresponding prediction of the visual-state rate. The difference between each prediction and the observed rate is used to compute a PAJ parameter correction. These corrections are weighted and summed separately for each channel. Inter-channel effects, timing differences, and other dynamics are collected in the local-model error.
The model is initialized using the nominal local image-response matrix B 0 :
B 0 = b 11 0 0 0 b 22 0 0 0 b 33 ,
where b 11 , b 22 , and  b 33 describe the locally dominant responses from the lateral-velocity, vertical-velocity, and yaw-rate commands to the three normalized visual-state channels, respectively. Each online update includes a correction toward B 0 .The control model combines the updated estimate and B 0 in a prescribed proportion.

3. MHW-PAJ-CLF Control Method

3.1. Control Architecture

Figure 3 shows the MHW-PAJ-CLF control loop. At each valid visual update, the perception module forms q k from the three filtered visual errors. It estimates q ˙ k from two consecutive valid states. MHW-PAJ uses each historical input in U k to compute a response error and a PAJ parameter correction. For each channel, the corrections are weighted using the response errors and input magnitudes, then summed. The update includes a correction toward the nominal value, and projection bounds the result. The updated estimate and B 0 are combined in a prescribed proportion to obtain B c , k . The CLF-QP uses q k and B c , k to compute u k within the input limits. The command is sent to the UAV; at each valid visual update, the previous control input u k 1 is used together with earlier inputs to update the model. A new valid visual frame triggers a parameter update. Between visual frames, the control loop computes commands using the latest visual state and model.

3.2. Conventional Projection-Adaptive Jacobian Update

The PAJ update uses the visual-state rate computed from two consecutive valid visual frames. For channel i, this rate is approximated as
q ˙ i , k = q i , k q i , k 1 t k t k 1 = q i , k q i , k 1 T img , k ,
where T img , k = t k t k 1 is the interval between adjacent valid visual frames and is used directly in the finite difference.
At each valid visual update, PAJ pairs q ˙ i , k with the most recent historical control input u i , k 1 and updates the local image-response parameter as
θ ^ i , k = Π i θ ^ i , k 1 + γ u i , k 1 ϵ + u i , k 1 2 q ˙ i , k θ ^ i , k 1 u i , k 1 + λ T img , k ( θ 0 , i θ ^ i , k 1 ) ,
where γ > 0 is the adaptive gain that regulates online correction speed and sensitivity to measurement noise, ϵ > 0 is the update regularizer that limits numerical amplification by small inputs, λ 0 slowly moves the estimate toward its nominal value under weak excitation, and  θ 0 , i is the nominal response parameter of channel i. The projection operator Π i restricts the updated estimate to [ θ ̲ i , θ ¯ i ] , thereby ensuring bounded online estimates. The denominator ϵ + u i , k 1 2 scales the correction of channel i using the corresponding channel-input energy.

3.3. Multi-History-Weighted Projection-Adaptive Jacobian Update

The current visual change may reflect several historical commands. MHW-PAJ uses the current parameter estimate to predict the state-change rate for each historical input. It compares these predictions with the observed rate and assigns weights using the response errors and input magnitudes.
At the kth valid visual update, the control-input history is written as
U k = { u k 1 , u k 2 , , u k 1 N d } ,
where N d is the maximum history-buffer depth, J k N d is the largest currently available history lag, and  j = 0 , , J k denotes the lag, in visual-update steps, of a historical input relative to the current update.
For channel i { 1 , 2 , 3 } , the input magnitudes within the history window are checked first. Let u th > 0 be the prescribed excitation-level parameter, with the channel threshold set to u th / 3 . If
max 0 j J k | u i , k 1 j | < u th 3 ,
the response-error correction of channel i is set to zero, and the λ term in Equation (34) moves the estimate toward the nominal value before projection is applied. Otherwise, the response error and weight associated with each historical input are calculated. The response error for the jth historical input is defined as
r i , k ( j ) = q ˙ i , k θ ^ i , k 1 u i , k 1 j ,
where r i , k ( j ) is the difference between the observed state-change rate and the prediction for the jth historical input. All predictions for channel i use the same estimate θ ^ i , k 1 . The response errors are scaled within each channel using
σ i , k 2 = 1 J k + 1 j = 0 J k r i , k ( j ) 2 + ϵ r ,
where ϵ r > 0 is the error-scale regularizer. The input-magnitude factor is defined as
χ i , k ( j ) = u i , k 1 j 2 ϵ u + u i , k 1 j 2 ,
where ϵ u > 0 is a regularizer; χ i , k ( j ) approaches unity as the input magnitude increases.
The historical-input weights are defined by the response error and input magnitude as
w ˜ i , k ( j ) = χ i , k ( j ) exp ( r i , k ( j ) ) 2 σ i , k 2 ,
w i , k ( j ) = w ˜ i , k ( j ) 𝓁 = 0 J k w ˜ i , k ( 𝓁 ) ,
where w ˜ i , k ( j ) and w i , k ( j ) are the unnormalized and normalized weights, respectively, with  w i , k ( j ) 0 and j = 0 J k w i , k ( j ) = 1 . A historical input receives a larger weight when its response error is smaller and its magnitude is larger. The weights are held fixed during the current parameter correction.
When the channel passes the input-magnitude check, MHW-PAJ computes a PAJ parameter correction for each historical input. The corrections are weighted and summed to obtain g i , k :
g i , k = j = 0 J k w i , k ( j ) u i , k 1 j ϵ + u i , k 1 j 2 r i , k ( j ) ,
θ ^ i , k = Π i θ ^ i , k 1 + γ g i , k + λ T img , k θ 0 , i θ ^ i , k 1 ,
Here, g i , k is the weighted sum of the parameter corrections for channel i. The coefficients γ , ϵ , λ , θ 0 , i , and the projection operator Π i follow the definitions in Section 3.2. The weighted correction and the correction toward the nominal value are added to the previous estimate. Projection places the result within its prescribed interval.

3.4. CLF-QP Command Generation

The three channel estimates form B ^ k = diag ( θ ^ 1 , k , θ ^ 2 , k , θ ^ 3 , k ) . For command generation, this online estimate is combined with the nominal model in the following proportion:
B c , k = ( 1 β ) B 0 + β B ^ k , 0 β 1 ,
where β sets the proportion of the online estimate in the control model.
The λ term moves the parameter toward its nominal value, while β sets the proportion of the online estimate used for command generation.
The nominal command is computed using the control model B c , k . It aims to make the visual errors decrease according to q ˙ = κ q , where κ > 0 sets the desired convergence rate. The calculation also penalizes command magnitude:
u k nom = arg min u 1 2 B c , k u + κ q k 2 + ρ a 2 u 2 ,
where ρ a > 0 is the command regularization coefficient. The first term measures how closely the model prediction follows the desired decrease in visual error. The second term penalizes command magnitude. The diagonal model allows the three command components to be computed separately.
For channel i, the nominal command has the form u i , k nom = κ b c , i , k q i , k / ( b c , i , k 2 + ρ a ) , where b c , i , k is the corresponding diagonal entry of B c , k . The response coefficient thus determines the nominal feedback gain for that channel. Online updating adjusts this gain as the observed response changes.
To impose an explicit local descent condition under the command constraints, consider the control Lyapunov function
V ( q ) = 1 2 q T q .
Under the local image-response model used for control, q ˙ B c , k u , the Lyapunov-function derivative is approximated by V ˙ q k T B c , k u . The following CLF condition is imposed to drive the visual state toward the desired configuration:
q k T B c , k u κ q k T q k ,
A non-negative slack variable δ softens the CLF descent condition to accommodate the available command range. The following quadratic program balances the desired descent with the command limits:
( u k , δ k ) = arg min u , δ 0 1 2 ( u u k nom ) T W ( u u k nom ) + ρ δ 2 δ 2 , s . t . q k T B c , k u κ q k 2 + δ ,
u min u u max ,
where W 0 is the positive diagonal weighting matrix for the command deviation, ρ δ > 0 is the slack penalty, and  δ k is the optimal CLF slack. For a fixed command, the minimizing slack is δ = [ q k T B c , k u + κ q k 2 ] + , where [ x ] + = max { x , 0 } .For any control command satisfying the input limits, the slack given by this expression satisfies the CLF constraint; hence, the soft-constrained QP has a feasible solution. Eliminating δ reduces Equations (39) and (40) to
min u min u u max 1 2 ( u u k nom ) T W ( u u k nom ) + ρ δ 2 [ q k T B c , k u + κ q k 2 ] + 2 .
Because W is diagonal, the KKT conditions of Equation (41) reduce to a monotone scalar equation whose root is computed by bisection. The resulting command satisfies the input box, and  δ k is recovered from the positive-part expression above.

3.5. Local Ultimate Boundedness

The projection operator in Equation (34) restricts each channel estimate to its prescribed interval:
θ ^ i , k [ θ ̲ i , θ ¯ i ] , i { 1 , 2 , 3 } .
Consequently, the estimated local image-response matrix is assembled as
B ^ k = diag ( θ ^ 1 , k , θ ^ 2 , k , θ ^ 3 , k )
The projection bounds keep B ^ k bounded throughout the online update. Because the nominal local image-response matrix B 0 is fixed and bounded and the combination coefficient satisfies 0 β 1 , the corresponding control model is
B c , k = ( 1 β ) B 0 + β B ^ k ,
The convex combination in Equation (44) is therefore also bounded.
For the closed-loop analysis, q ( t ) denotes the current visual state in the local continuous-time model. The matrix B c ( t ) is the model used to generate the currently held command and is held over the corresponding control interval. According to Equation (19), the realized local visual-state dynamics can be expressed as
q ˙ = B c u + η c ,
where the local-model error term is defined as
η c = ( B B c ) u + η
The term η c includes local image-response estimation error, inter-channel effects, visual measurement error, and other disturbances.
The visual loop runs at 20 Hz, and the control loop computes commands at 100 Hz. Between valid visual updates, the controller uses the latest observation and model. Let q h ( t ) denote the observed state used to compute the current command. The state-holding error is the difference between the current state q ( t ) and this observation:
e h ( t ) = q ( t ) q h ( t ) .
This difference reflects the effects of sampling, feedback lag, measurement, and filtering. Let δ ( t ) denote the optimal slack held over the corresponding control interval. The QP condition in Equation (40) applies to q h and gives
q h T B c u κ q h 2 2 + δ .
The closed-loop trajectory is assumed to remain in a compact local regulation region. Within this region, the target stays visible and the local image-response model remains valid. The input box contains feasible commands. Let finite constants η ¯ , e ¯ , and  δ ¯ satisfy
η c ( t ) 2 η ¯ , e h ( t ) 2 e ¯ , 0 δ ( t ) δ ¯ .
The bounds on B c and the input give a finite constant M ¯ such that B c u 2 M ¯ . Using q h = q e h , the derivative of the Lyapunov function in Equation (37) satisfies
V ˙ = q T ( B c u + η c ) = q h T B c u + e h T B c u + q T η c κ q h 2 2 + M ¯ e ¯ + η ¯ q 2 + δ ¯ κ q 2 2 + ( η ¯ + 2 κ e ¯ ) q 2 + δ ¯ + M ¯ e ¯ .
Let s = q 2 . The right-hand side of Equation (50) is
κ s 2 + ( η ¯ + 2 κ e ¯ ) s + δ ¯ + M ¯ e ¯ ,
which is strictly negative whenever
s > s = η ¯ + 2 κ e ¯ + ( η ¯ + 2 κ e ¯ ) 2 + 4 κ ( δ ¯ + M ¯ e ¯ ) 2 κ .
Thus, V decreases outside the ball of radius s , and the normalized visual state is locally ultimately bounded with
lim   sup t q ( t ) 2 s .
The radius accounts for local-model error, state-holding error, and QP slack. If e ¯ = 0 and δ ( t ) = 0 throughout the region, then δ ¯ = 0 and the radius reduces to
s = η ¯ κ .
If η ¯ = 0 also holds, then
V ˙ κ q 2 2 = 2 κ V ,
which gives local exponential convergence of the visual-state origin.

3.6. Online Procedure

The visual loop runs at 20 Hz , updating the visual state and MHW-PAJ parameters when a valid frame arrives. At 100 Hz , the control loop computes the nominal command and solves the CLF-QP using the latest visual state and model. During detection loss, the preceding lateral-velocity, vertical-velocity, and yaw-rate commands are progressively attenuated. When the loss threshold is exceeded, the response-gain estimate is reset and a yaw-search command is issued. Algorithm 1 summarizes this procedure.
Let n q be the number of visual channels. The conventional PAJ update has complexity O ( n q ) . At each valid visual update, the MHW-PAJ computation has complexity O n q ( J k + 1 ) and worst-case complexity O n q ( N d + 1 ) ; its history requires O n q ( N d + 1 ) memory. Here, n q = 3 . Each bisection evaluation of the reduced soft CLF-QP costs O ( n q ) , and  N bis evaluations give a solution complexity of O ( N bis n q ) . Thus, the principal additional cost of the multi-history update grows linearly with the history depth.
Algorithm 1 MHW-PAJ-CLF Dual-Rate Online Control
Require: Latest filtered visual state q k , detection status, new-frame flag, control-input history U k , B 0 , and  B ^ k 1
Ensure: Command u k
  1: if the current detection is the first valid OBB then
  2:     Set q k 1 = q k and B ^ k = B 0 , and clear U k .
  3:     Issue a zero command and return.
  4: end if
  5: if a new valid visual frame has arrived then
  6:     Append the most recently applied command to U k , estimate q ˙ k , and set J k .
  7:     for  i = 1 , 2 , 3  do
  8:          if  max 0 j J k | u i , k 1 j | < u th / 3  then
  9:             Set g i , k = 0 , move the estimate toward the nominal value, and apply projection in Equation (34).
10:          else
11:             Compute r i , k ( j ) using Equation (28), followed by σ i , k 2 , χ i , k ( j ) , and  w i , k ( j ) .
12:             Update θ ^ i , k using Equations (33) and (34) and project the result.
13:          end if
14:     end for
15: end if
16: Construct B c , k using Equation (35).
17: Compute u k nom and solve the reduced soft CLF-QP in Equation (41) by scalar bisection.
18: Compute the optimal slack δ k from the CLF constraint residual.
19: Issue u k ; the command computation above is executed at every control cycle.

4. Experimental Validation and Results

Experiments were conducted on the RflySim–PX4 hardware-in-the-loop platform shown in Figure 4, using a fixed overhead-ground-wire scene. The experiments evaluate visual tracking, robustness to perturbations, and the contribution of historical-input weighting. Local-model comparisons, parameter sensitivity, and computation time are also examined.

4.1. Experimental Platform and Evaluation Metrics

The platform consists of flight-control hardware, the RflySim simulation environment, and the visual-control program [32]. The flight-control inner loop runs on the hardware, while RflySim supplies the UAV dynamics, sensor feedback, three-dimensional scene, and onboard-camera images. The visual-control program converts YOLO11s-OBB detections into velocity and yaw-rate commands. Each visual-error component was processed by an independent constant-velocity Kalman filter before being used for control. Table 1 lists the platform and visual-detection settings, and Table 2 gives the MHW-PAJ-CLF parameters.
The controller parameters were adjusted progressively according to the closed-loop responses, considering success rate, transient overshoot, and oscillations in the tracking errors and control commands. Table 2 gives the resulting parameter values. The parameter settings of the baseline methods are listed in Table 3.
PID and ADRC use independent three-channel feedback structures, with ADRC additionally using extended-state observation. NVPC balances predicted tracking error, command magnitude, and command changes, with a visibility penalty that discourages motion toward the image boundary. RBF-IBVS uses Gaussian basis functions to estimate and compensate for local model error. The baseline parameters were adjusted progressively according to the observed closed-loop responses. The selected baseline parameters were kept fixed across the five evaluation runs. PID uses bounded integral accumulation and command clipping; ADRC commands are clipped to their prescribed bounds; NVPC projects candidate inputs into the admissible range; and the RBF-IBVS output is limited before command application. These controller-generated commands are subject to the common bounds listed in Table 1. All five methods use the same detector and visual-filter settings, initial conditions, sampling rates, bounds on the controller-generated commands, and success criterion. Results presented as the mean ± sample standard deviation are calculated from five runs. The numerical results in the tables are rounded to the precision shown, and the reported percentage changes in tracking error are calculated from the displayed mean values. The best mean values are shown in bold in the comparison tables.
For the initial-pose, noise-and-dropout, and command-bias conditions in Section 4.3, PAJ-CLF and MHW-PAJ-CLF were each evaluated in five runs using run seeds 12 , 001 12 , 005 . The same seed was assigned to the two methods for each repetition.
The mean absolute error measures tracking accuracy in each visual channel. The overall tracking error E q combines the three normalized errors:
E q = 1 N k = 1 N q k 2 2 ,
where N is the number of observations. A smaller E q indicates a smaller overall tracking error.
For each run, the success rate is the percentage of the total run duration during which all three visual errors simultaneously satisfy the following tolerances:
| e x |     8 px , | e w |     3 px , | e ψ |     0.10 rad .
For results reported as a mean, the success rates are averaged over five runs.
The local-response residual E B is computed consistently across the local-model update methods, using the preceding control model and the most recent historical input for prediction:
E B = 1 3 N h k H q ˙ k B c , k 1 u k 1 2 2 ,
where H contains the consecutive valid visual observations used to calculate the response errors throughout the run and N h = | H | . A smaller E B indicates closer agreement between the model prediction and the observed visual response.

4.2. Visual-Tracking Performance

The visual-tracking comparison included MHW-PAJ-CLF, PID, ADRC, an NVPC baseline adapted from [33], and RBF-IBVS [20]. Figure 5 shows the lateral-position, apparent-width, and angle-error responses. Table 4 compares the tracking errors and success rates.
MHW-PAJ-CLF achieves the lowest mean absolute error in all three visual channels. Its lateral-position error is 2.51% lower than that of PID, the best baseline for this channel. The improvements in apparent width and angle are larger: relative to RBF-IBVS, the best baseline for these two channels, the errors decrease by 8.36% and 18.77%, respectively.
MHW-PAJ-CLF achieves a mean success rate of 75.50%, exceeding the 64.45% of RBF-IBVS by 11.05 percentage points. Together with the reductions in width and angle errors, the higher success rate shows that the proposed controller improves coordinated regulation of the three visual variables while maintaining lateral alignment accuracy.
Figure 6 shows the YOLO11s-OBB detections during the initial closed-loop phase.

4.3. Performance Under Perturbations and Temporal Lag

To assess tracking robustness, PAJ-CLF and MHW-PAJ-CLF are compared under changes in initial pose, measurement noise and frame dropout, command bias, and additional visual-feedback delay. Table 5 summarizes the results.
The initial-pose condition introduces a lateral displacement of 0.70 m, a height change of 1.0 m, and a yaw offset of 10 ° relative to the reference condition. The noise-and-dropout condition uses measurement-noise standard deviations of 4 px, 1.5 px, and 0.02 rad for the lateral-position, apparent-width, and orientation-angle errors, respectively, with a 0.5 s frame dropout beginning at 12 s.
The command-bias condition applies offsets of 0.08 m/s, 0.06 m/s, and 0.06 rad/s to the lateral-velocity, vertical-velocity, and yaw-rate commands, respectively, during 10–25 s, with 1 s linear onset and removal ramps. These offsets are added after controller output limiting. Additional visual-feedback delays of 100 and 200 ms are evaluated separately.
MHW-PAJ-CLF achieves lower mean tracking error E q and local-response prediction error E B in every condition in Table 5. The largest reduction in tracking error, 11.08%, occurs under initial-pose variation, where the standard deviation also decreases from 0.222 to 0.051. This combination indicates more accurate and more consistent tracking as the initial pose changes. Under noise and frame dropout and under command bias, the mean tracking error decreases by 5.10% and 4.76%, respectively.
With additional visual-feedback delays of 100 and 200 ms, MHW-PAJ-CLF retains lower mean tracking errors than PAJ-CLF. Its E q values are 0.728 and 0.737, compared with 0.735 and 0.758 for PAJ-CLF, respectively.
The relationship between the control-method output u c and the command applied through the flight-control interface, u a , is represented by the first-order dynamics
τ a u ˙ a + u a = u c ,
where τ a is the time constant of the first-order element imposed at the command interface; when τ a = 0 , this element satisfies u a = u c . Figure 7 and Table 6 compare the tracking performance at τ a = 0 , 50 ms, and 100 ms.
MHW-PAJ-CLF achieves smaller mean absolute errors in all three channels at each command-response time constant. As τ a increases from 0 to 100 ms, its E q rises from 0.7287 to 0.7391 and remains below that of PAJ-CLF. At 100 ms, the success rate is 73.06%, compared with 58.49% for PAJ-CLF. Together with the visual-feedback-delay results, these findings show that multi-history updating retains its tracking benefit under both delayed visual feedback and slower command response.
Across the 60 PAJ-CLF and MHW-PAJ-CLF runs, the Pearson and Spearman correlations between E B and E q are 0.282 and 0.420 , respectively. These modest positive correlations indicate that lower prediction errors tend to accompany lower tracking errors. In closed-loop operation, the estimated response gains also shape the feedback commands, so tracking performance depends on both prediction accuracy and the resulting control action.

4.4. Analysis of Historical-Input Weighting

Table 7 compares the fixed nominal model, single-history PAJ update, and multi-history-weighted PAJ update within the same CLF-QP control framework. The time T first measures when all three errors first satisfy their tolerances, and T 3 s marks the start of the first interval in which they remain within the tolerances for at least 3 s.
MHW-PAJ-CLF achieves the lowest mean tracking error and reaches the three error tolerances 3.22 s earlier than the fixed model and 1.72 s earlier than PAJ-CLF, on average. The mean T 3 s values are close, ranging from 10.00 to 10.44 s. MHW-PAJ-CLF and the fixed model also have similar mean E B values. The main improvements therefore concern the initial reduction of the visual errors and overall closed-loop tracking.
During the first 10 s, the mean E q values are 1.493 ± 0.067 , 1.550 ± 0.232 , and 1.425 ± 0.073 for the fixed model, PAJ-CLF, and MHW-PAJ-CLF, respectively. The corresponding mean lateral feedback gains are 0.108 , 0.121 , and 0.168 , and the mean RMS lateral-velocity commands are 0.110 , 0.124 , and 0.134 m/s. As described in Section 3.4, the estimated response gain determines the feedback gain used to generate the command. The larger initial feedback gain of MHW-PAJ-CLF produces stronger lateral corrections, which are consistent with its earlier first entry and lower initial tracking error.
Table 8 examines the weighting rule through three alternatives. Uniform weighting assigns equal weights to all historical inputs; response-error-only weighting uses their prediction errors; shared weighting applies the same weights to all three visual channels.
The full configuration determines the weights separately for each channel using both prediction error and input magnitude. It achieves the lowest mean E q among the tested configurations. Uniform weighting, response-error-only weighting, and shared weighting give mean errors that are 1.75%, 1.48%, and 1.90% higher, respectively. These results support the combined weighting rule used by MHW-PAJ.
The weighted mean history position describes whether the weights favor more recent or earlier inputs:
d ¯ i , k = j = 0 J k j w i , k ( j ) ,
The equivalent number of historical inputs measures how widely the weights are distributed:
N eff , i , k = j = 0 J k w i , k ( j ) 2 1 .
A larger d ¯ i , k indicates greater weighting toward earlier inputs, while a larger N eff , i , k indicates that more historical inputs share the weights.
Figure 8 shows the historical-input weights in the lateral-position, apparent-width, and angle channels. Each bar corresponds to one history position, ordered from the newest input on the left to the oldest on the right. With a nominal visual-update interval of 50 ms, 0 ms denotes the most recent input and 500 ms denotes the input ten updates earlier.
The mean weights range from approximately 7.1 % to 10.4 % , showing that multiple historical inputs contribute to the update. The lateral-position and apparent-width channels give slightly higher weights to older inputs, while the angle channel gives higher weights to inputs at 300–400 ms. Each channel uses prediction error and input magnitude to weight the parameter corrections associated with these historical inputs, then combines the corrections to update its local response gain.
Table 9 gives mean equivalent numbers of historical inputs between 9.35 and 10.26 within the eleven-input window, further indicating that the update draws on a broad part of the available history.
With additional visual-feedback delays of 0, 100, and 200 ms, the weighted mean history positions in the three channels, expressed in time units, are ( 265 , 263 , 265 ) , ( 270 , 263 , 278 ) , and ( 263 , 262 , 282 ) ms, respectively. Increasing the additional delay from 0 to 200 ms changes these mean positions by 1.7 , 1.6 , and 17.1 ms. The weighted mean history position characterizes the center of the gain-update weight distribution over the input history. Its value depends on the input amplitudes, prediction errors, and available history length.

4.5. Comparison of Diagonal and Full Local Image-Response Models

The controller uses B 0 as a preset nominal matrix, while B ^ diag and B ^ full are diagonal and full local constant models identified from experimental data. Both models use the input–response relation
q ˙ k B u k 1 .
Pseudo-random binary inputs are used to obtain the identification data at the reference viewpoint and within the command limits listed in Table 1. The full and diagonal models are identified by multivariable and channel-wise least squares, respectively, and evaluated using validation data. The rows correspond to q x , q w , and q ψ , and the columns correspond to v y , v z , and r. The identified matrices are
B ^ full = 0.6550 0.0534 0.0981 0.0391 0.5213 0.0249 0.0402 0.0269 1.0741 , B ^ diag = diag ( 0.6720 , 0.5332 , 1.0763 ) .
The corresponding diagonal elements of B ^ full and B ^ diag differ by 2.52%, 2.23%, and 0.20%.The ratio between the Frobenius norms of the off-diagonal and diagonal parts of B ^ full is 9.57%. The largest off-diagonal element is B 13 = 0.0981 , whose magnitude is 14.98% of the main diagonal element in the same row. The validation RMSE values of the two models are summarized in Table 10.
For predictions of the normalized visual-state rate, the overall validation RMSE values of the diagonal and full models are 0.2150 and 0.2161, respectively. At the reference viewpoint and within the tested command range, the three corresponding input–output channels provide the dominant responses, supporting the use of a diagonal local image-response structure for control.

4.6. Parameter Sensitivity and Computation Time

Table 11 examines how the CLF-QP slack penalty ρ δ affects tracking performance at values of 50, 100, and 200.
Across this range, the mean E q varies by 0.53%, and the mean slack remains between 0.0340 and 0.0360. The small change in mean error indicates stable tracking performance over this parameter range, which includes the reference value ρ δ = 100 .
The adaptive gain γ sets the strength of the response-error correction, while the history depth N d sets the extent of the input history used for updating. Table 12 examines their effects on tracking performance.
With N d = 10 , increasing γ from 0.01 to 0.02 reduces the mean E q from 0.768 to 0.716. A further increase to 0.04 gives a smaller reduction to 0.711. Thus, strengthening the correction has a larger benefit at the lower gains, and the reference value of 0.02 achieves mean tracking performance close to that at 0.04.
With γ = 0.02 , history depths of 5 and 10 give similar mean E q values of 0.721 and 0.716, while a depth of 15 gives a higher value of 0.752. A longer window incorporates more distant commands into the local-response update. The lower errors at depths 5 and 10 support the reference history depth of 10 under the tested conditions.
The controller computation times are summarized in Table 13. The mean computation times of PAJ-CLF and MHW-PAJ-CLF are 0.037 and 0.038 ms, respectively, and their 95th-percentile times are 0.107 and 0.113 ms. For MHW-PAJ-CLF, the 95th-percentile time occupies 1.13% of the 10 ms control period, supporting the dual-rate configuration with channel-gain updates at 20 Hz and CLF-QP command generation at 100 Hz.

5. Discussion

The five-method comparison highlights the contribution of MHW-PAJ-CLF to coordinated visual regulation. It maintains lateral alignment accuracy while improving apparent-width and angle regulation, increasing the proportion of time during which all three visual errors simultaneously satisfy their tolerances.
MHW-PAJ-CLF and the fixed model have similar mean prediction errors, and MHW-PAJ-CLF reaches all three error tolerances earlier. Online updates of the local model adjust the feedback gains through the nominal control law and influence the commands generated by the CLF-QP. During the first 10 s, MHW-PAJ-CLF exhibits a larger lateral feedback gain and lateral-velocity command magnitude, consistent with its lower initial tracking error and earlier simultaneous satisfaction of all three error tolerances.
MHW-PAJ combines parameter corrections associated with several preceding inputs. Each visual channel weights these corrections according to prediction error and input magnitude. The full weighting rule gives the lowest mean tracking error among the tested configurations, and the comparisons with PAJ-CLF show that its tracking benefit extends to additional visual-feedback delay and command-response lag.
The three-channel local model is used near the reference pose, and closed-loop performance is evaluated through HIL experiments in a fixed overhead-ground-wire scene. The experiments did not include controlled variations in illumination or scene background, so robustness to these changes remains unverified. Future work will evaluate the effects of illumination and background variation on visual tracking, alongside adaptive history-window selection and real-flight validation.

6. Conclusions

This paper develops MHW-PAJ-CLF for three-channel close-range UAV visual servoing near overhead ground wires. The method uses response error and input magnitude to assign weights to historical inputs separately for each channel. It then takes the weighted sum of the corresponding PAJ parameter corrections. The update includes a correction toward the nominal value, and projection restricts the estimated response parameter of each channel to its prescribed interval. The updated estimate and the nominal model are combined in a prescribed proportion. The CLF-QP uses this model to regulate the three visual errors within the input limits.
Across five HIL runs, MHW-PAJ-CLF achieves mean absolute lateral-position, apparent-width, and angle errors of 15.201 px, 2.005 px, and 0.07561 rad, respectively, and a mean success rate of 75.50 % . Compared with the best baseline for each visual-error metric, the three mean absolute errors decrease by 2.51 % , 8.36 % , and 18.77 % , respectively. The mean success rate exceeds that of RBF-IBVS by 11.05 percentage points. In the reference comparison with PAJ-CLF, the mean time to first satisfy all three error tolerances decreases from 7.63 to 5.91 s. Under initial-pose variation, measurement noise and frame dropout, control-command bias, additional visual-feedback delay, and different command-response time constants, MHW-PAJ-CLF achieves a lower overall tracking error E q than PAJ-CLF. The three simplified weighting configurations have mean tracking errors 1.48–1.90% above that of the full configuration, and the 95th-percentile computation time of MHW-PAJ-CLF is 0.113 ms. These HIL results support the use of multi-history weighting to improve close-range UAV visual tracking under visual-feedback delay and command-response lag.

Author Contributions

Conceptualization, L.H., B.W., Z.Z., Y.C. and Y.Y.; methodology, L.H., B.W., Z.Z., Y.C. and Y.Y.; resources, L.H. and B.W.; investigation, L.H. and B.W.; formal analysis, L.H. and B.W.; data curation, L.H. and B.W.; writing—original draft preparation, L.H. and B.W.; writing—review and editing, Z.Z. and Y.C.; supervision, Y.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China, grant number 62473216, and the Natural Science Foundation of Nantong City, grant numbers JC2023006 and JC2025059.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data, controller configurations, and evaluation results supporting the findings of this study are available from the corresponding author upon reasonable request.

Acknowledgments

The authors used AI tools solely for English-language editing. The authors reviewed and edited all AI-generated outputs and take full responsibility for the content of this manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ADRCActive disturbance rejection control
CLF-QPControl Lyapunov function-based quadratic program
IBVSImage-based visual servoing
MHW-PAJMulti-history-weighted projection-adaptive Jacobian update
NVPCNonlinear visual predictive control
PAJProjection-adaptive Jacobian
UAVUnmanned aerial vehicle
YOLO11s-OBBYou Only Look Once version 11 small oriented bounding box

References

  1. Ahmed, F.; Mohanta, J.C.; Keshari, A. Power transmission line inspections: Methods, challenges, current status and usage of unmanned aerial systems. J. Intell. Robot. Syst. 2024, 110, 54. [Google Scholar] [CrossRef] [Scilit]
  2. Diniz, L.F.; Pinto, M.F.; Melo, A.G.; Honorio, L.M. Visual-based assistive method for UAV power line inspection and landing. J. Intell. Robot. Syst. 2022, 106, 41. [Google Scholar] [CrossRef] [Scilit]
  3. Guan, H.; Sun, X.; Su, Y.; Hu, T.; Wang, H.; Wang, H.; Peng, C.; Guo, Q. UAV-LiDAR aids automatic intelligent powerline inspection. Int. J. Electr. Power Energy Syst. 2021, 130, 106987. [Google Scholar] [CrossRef] [Scilit]
  4. Xing, J.; Cioffi, G.; Hidalgo-Carrio, J.; Scaramuzza, D. Autonomous power line inspection with drones via perception-aware MPC. In Proceedings of the 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Detroit, MI, USA, 1–5 October 2023; pp. 1086–1093. [Google Scholar] [CrossRef] [Scilit]
  5. Ollero, A.; Suarez, A.; Papaioannidis, C.; Pitas, I.; Marredo, J.M.; Hoang, V.D.; Ebeid, E.; Kratky, V.; Saska, M.; Hanoune, C.; et al. Multi-aerial robotic system for power line inspection and maintenance: Comparative analysis from the AERIAL-CORE final experiments. IEEE Trans. Field Robot. 2025, 2, 549–573. [Google Scholar] [CrossRef] [Scilit]
  6. Zhang, H.; Wu, L.; Chen, Y.; Chen, R.; Kong, S.; Wang, Y.; Hu, J.; Wu, J. Attention-guided multitask convolutional neural network for power line parts detection. IEEE Trans. Instrum. Meas. 2022, 71, 5008213. [Google Scholar] [CrossRef] [Scilit]
  7. Liu, Z.; Wu, G.; He, W.; Fan, F.; Ye, X. Key target and defect detection of high-voltage power transmission lines with deep learning. Int. J. Electr. Power Energy Syst. 2022, 142, 108277. [Google Scholar] [CrossRef] [Scilit]
  8. Choudhary, S.; Saurav, S.; Gidde, P.; Saini, R.; Singh, S. LPC-Det: Attention-based lightweight object detector for power line component detection in UAV images. Comput. Electr. Eng. 2025, 126, 110476. [Google Scholar] [CrossRef] [Scilit]
  9. Xu, Y.; Xu, Z.; Wang, H.; Wei, Z.; Wu, Z. Instance-level orientation enhancement for horizontal box supervised oriented object detection in remote sensing images. IEEE Trans. Image Process. 2025, 34, 7613–7626. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. Lin, Y.X.; Lai, Y.C. Deep Learning-Based Navigation System for Automatic Landing Approach of Fixed-Wing UAVs in GNSS-Denied Environments. Aerospace 2025, 12, 324. [Google Scholar] [CrossRef] [Scilit]
  11. Wang, G.; Qin, J.; Liu, Q.; Ma, Q.; Zhang, C. Image-based visual servoing of quadrotors to arbitrary flight targets. IEEE Robot. Autom. Lett. 2023, 8, 2022–2029. [Google Scholar] [CrossRef] [Scilit]
  12. Li, Y.; Lu, G.; He, D.; Zhang, F. Robocentric model-based visual servoing for quadrotor flights. IEEE/ASME Trans. Mechatron. 2023, 28, 2155–2166. [Google Scholar] [CrossRef] [Scilit]
  13. Xie, H.; Fink, G.; Lynch, A.F.; Jagersand, M. Adaptive visual servoing of UAVs using a virtual camera. IEEE Trans. Aerosp. Electron. Syst. 2016, 52, 2529–2538. [Google Scholar] [CrossRef] [Scilit]
  14. Lin, J.; Wang, Y.; Miao, Z.; Fan, S.; Wang, H. Robust observer-based visual servo control for quadrotors tracking unknown moving targets. IEEE/ASME Trans. Mechatron. 2023, 28, 1268–1279. [Google Scholar] [CrossRef] [Scilit]
  15. Lin, J.; Wang, Y.; Miao, Z.; Wang, H.; Fierro, R. Robust image-based landing control of a quadrotor on an unpredictable moving vehicle using circle features. IEEE Trans. Autom. Sci. Eng. 2023, 20, 1429–1440. [Google Scholar] [CrossRef] [Scilit]
  16. Kumar, Y.; Shamsi, B.P.; Roy, S.B.; Sujit, P.B. Tracking a planar target using image-based visual servoing technique. IEEE Trans. Intell. Veh. 2024, 9, 4362–4372. [Google Scholar] [CrossRef] [Scilit]
  17. Yang, K.; Bai, C.; She, Z.; Quan, Q. High-speed interception multicopter control by image-based visual servoing. IEEE Trans. Control Syst. Technol. 2025, 33, 119–135. [Google Scholar] [CrossRef] [Scilit]
  18. Yang, J.; Liu, X.; Sun, J.; Li, S. Sampled-data robust visual servoing control for moving target tracking of an inertially stabilized platform with a measurement delay. Automatica 2022, 137, 110105. [Google Scholar] [CrossRef] [Scilit]
  19. Chen, Y.; Li, B.; Shi, W.; Zhang, X. Image-based visual servoing of micro aerial vehicles with robust observation and output feedback. IEEE Trans. Control Syst. Technol. 2025, 33, 1995–2004. [Google Scholar] [CrossRef] [Scilit]
  20. Sepahvand, S.; Janabi-Sharifi, F.; Masnavi, H.; Aghili, F.; Amiri, N. Robust image-based visual servoing of an aerial robot using self-organizing neural networks. Int. J. Control Autom. Syst. 2024, 22, 3762–3776. [Google Scholar] [CrossRef] [Scilit]
  21. Yang, Q.; Li, H. RMPC-based visual servoing for trajectory tracking of quadrotor UAVs with visibility constraints. IEEE/CAA J. Autom. Sin. 2024, 11, 2027–2029. [Google Scholar] [CrossRef] [Scilit]
  22. Zhou, Y.; Xu, F.; Zhou, Y.; Zhang, Z.; Wang, H. Constrained image-based visual servoing with a sampling-based planning framework. IEEE/ASME Trans. Mechatron. 2025, 30, 4899–4909. [Google Scholar] [CrossRef] [Scilit]
  23. Liu, Q.; Mao, J.; Han, L.; Zhang, C.; Yang, J. Predictive observer-based dual-rate prescribed performance control for visual servoing of robot manipulators with view constraints. IEEE Trans. Cybern. 2025, 55, 2424–2436. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Yi, X.; Luo, B.; Zhao, Y. Neural network-based robust guaranteed cost control for image-based visual servoing of quadrotor. IEEE Trans. Neural Netw. Learn. Syst. 2024, 35, 12693–12705. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  25. Byun, W.; Huh, S.; Jang, H.; Yu, S.; Lim, S.; Lee, S.; Nam, W. Vector Field-Based Robust Quadrotor Landing on a Moving Ground Platform. Aerospace 2025, 12, 590. [Google Scholar] [CrossRef] [Scilit]
  26. Zhang, L.; Lin, Y.; Chen, Y.; Wang, A.; Ren, Z. Data-driven generalized iterative predictive control with IESO for UAV trajectory tracking. Int. J. Control Autom. Syst. 2026, 24, 1156–1170. [Google Scholar] [CrossRef] [Scilit]
  27. Guo, H.; Song, T.; Ye, J.; Abdulrahman, Y.; Gu, X.; Jiang, T.; Dong, Y. Image-Based Visual Servoing for Quadrotor Formation Encirclement and Tracking of Unknown Targets. Aerospace 2026, 13, 138. [Google Scholar] [CrossRef] [Scilit]
  28. Pan, Y.; Shi, T. Adaptive estimation and control with online data memory: A historical perspective. IEEE Control Syst. Lett. 2024, 8, 267–278. [Google Scholar] [CrossRef] [Scilit]
  29. Li, Z.; Lai, B.; Pan, Y. Image-based composite learning robot visual servoing with an uncalibrated eye-to-hand camera. IEEE/ASME Trans. Mechatron. 2024, 29, 2499–2509. [Google Scholar] [CrossRef] [Scilit]
  30. Kamath, A.K.; Feroskhan, M. Augmented dynamics visual servoing: Mapping image variations to multirotor’s input commands. IEEE Trans. Aerosp. Electron. Syst. 2025, 61, 14928–14942. [Google Scholar] [CrossRef] [Scilit]
  31. Kamath, A.K.; Yan, B.; Shi, P.; Feroskhan, M. Physics-informed Koopman neural operator for augmented dynamics visual servoing of multirotors. IEEE Trans. Autom. Sci. Eng. 2026, 23, 4639–4653. [Google Scholar] [CrossRef] [Scilit]
  32. Wang, S.; Dai, X.; Ke, C.; Quan, Q. RflySim: A rapid multicopter development platform for education and research based on Pixhawk and MATLAB. In Proceedings of the 2021 International Conference on Unmanned Aircraft Systems (ICUAS), Athens, Greece, 15–18 June 2021; pp. 1587–1594. [Google Scholar] [CrossRef] [Scilit]
  33. Zhang, Y.; Yao, J.; Qian, C. Event-triggered nonlinear visual predictive control strategy for robots. J. Intell. Robot. Syst. 2025, 111, 80. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Physical task geometry and visual-error definition: (a) physical geometry and reference-command directions; (b) image-space visual errors.
Figure 1. Physical task geometry and visual-error definition: (a) physical geometry and reference-command directions; (b) image-space visual errors.
Aerospace 13 00844 g001
Figure 2. Height–width calibration from five viewpoints.
Figure 2. Height–width calibration from five viewpoints.
Aerospace 13 00844 g002
Figure 3. Closed-loop structure and signal flow of the proposed MHW-PAJ-CLF control method. Arrows indicate the direction of signal flow; the superscript asterisk denotes the optimal control command obtained by solving the CLF-QP.
Figure 3. Closed-loop structure and signal flow of the proposed MHW-PAJ-CLF control method. Arrows indicate the direction of signal flow; the superscript asterisk denotes the optimal control command obtained by solving the CLF-QP.
Aerospace 13 00844 g003
Figure 4. RflySim–PX4 HIL platform and visual-control program.
Figure 4. RflySim–PX4 HIL platform and visual-control program.
Aerospace 13 00844 g004
Figure 5. Visual-error responses of the five methods: (a) lateral-position error; (b) apparent-width error; (c) angle error. The gray dashed lines indicate the error tolerance limits: ± 8 px in (a), ± 3 px in (b), and ± 0.10 rad in (c).
Figure 5. Visual-error responses of the five methods: (a) lateral-position error; (b) apparent-width error; (c) angle error. The gray dashed lines indicate the error tolerance limits: ± 8 px in (a), ± 3 px in (b), and ± 0.10 rad in (c).
Aerospace 13 00844 g005
Figure 6. YOLO11s-OBB detection results during the initial closed-loop phase: (af) successive visual observations from 0 to 10 s. The green boxes indicate the oriented bounding boxes of the overhead ground wire detected by YOLO11s-OBB.
Figure 6. YOLO11s-OBB detection results during the initial closed-loop phase: (af) successive visual observations from 0 to 10 s. The green boxes indicate the oriented bounding boxes of the overhead ground wire detected by YOLO11s-OBB.
Aerospace 13 00844 g006
Figure 7. Visual-tracking performance at different command-response time constants: (a) success rate; (b) normalized tracking root-mean-square error E q .
Figure 7. Visual-tracking performance at different command-response time constants: (a) success rate; (b) normalized tracking root-mean-square error E q .
Aerospace 13 00844 g007
Figure 8. Historical-input weights in the (a) lateral-position, (b) apparent-width, and (c) angle channels. Bar heights and error bars show the mean and sample standard deviation of the run-averaged weights across five reference runs.
Figure 8. Historical-input weights in the (a) lateral-position, (b) apparent-width, and (c) angle channels. Bar heights and error bars show the mean and sample standard deviation of the run-averaged weights across five reference runs.
Aerospace 13 00844 g008
Table 1. Common HIL platform, visual-perception, and experimental settings.
Table 1. Common HIL platform, visual-perception, and experimental settings.
ItemValueItemValue
Inertia matrix J 0.0211 , 0.0219 , 0.0366 kg m 2 Mass m 1.4 kg
Arm length 0.225 mMotor time constant τ m 0.05 s
Camera offset [ 0.1 , 0 , 0 ] mCamera FOV 90
Image resolution 640 × 480 Reference width w ref 35 px
Visual rate20 HzControl rate100 Hz
Run duration40 sInference mean/p95 11.39 / 28.89 ms
Scales ( s x , s w , s ψ ) ( 80 , 20 , 0.5 ) Command bounds ( 0.20 , 0.30 , 0.60 )
DetectorYOLO11s-OBBSelected OBBHighest confidence
Object classWire (one class)Dataset sourceRflySim
Training images800Validation images200
Input size640Batch size16
Training epochs100Detector training seed0
Confidence 0.25 NMS IoU 0.70
Precision 1.000 Recall 0.994
mAP 50 0.995 mAP 50 : 95 0.779
The command bounds are symmetric; the entries for ( v y , v z , r ) are given in m / s , m / s , and rad / s , respectively.
Table 2. Default parameters of the proposed MHW-PAJ-CLF controller.
Table 2. Default parameters of the proposed MHW-PAJ-CLF controller.
ParameterValueParameterValue
diag ( B 0 ) ( 5.55 , 1.00 , 2.30 ) β 0.80
θ min ( 11.25 , 1.00 , 8.00 ) θ max ( 0.50 , 12.50 , 0.30 )
N d 10 γ 0.02
λ 0.04 u th 0.03
( ϵ , ϵ r , ϵ u ) ( 10 2 , 10 6 , 10 4 ) ( ρ a , ρ δ ) ( 0.01 , 100 )
κ 0.60 diag ( W ) ( 1 , 2 , 8 )
Table 3. Parameter settings of the baseline methods.
Table 3. Parameter settings of the baseline methods.
MethodParameter GroupSetting
PIDChannel gains K p = ( 0.002 , 0.010 , 0.40 ) ; K i = ( 0.00005 , 0.0002 , 0 ) ; K d = ( 0 , 0 , 0.005 ) .
ADRCx channel ( ω c , ω o , b 0 ) x = ( 0.13 , 1.4 , 360 ) .
w channel ( ω c , ω o , b 0 ) w = ( 0.5 , 2.5 , 100 ) .
ψ channel ( ω c , ω o , b 0 ) ψ = ( 1.5 , 5.0 , 6.0 ) .
NVPCPrediction/visibility N p = 16 ; Δ t = 0.05 s; ρ v = 220 ; minimum normalized edge margin = 0.18 .
Cost weights Q = ( 0.45 , 0.80 , 0.40 ) ; R = ( 0.70 , 0.25 , 0.60 ) ; R Δ = ( 6.00 , 1.50 , 2.00 ) .
Command dampingSmoothing coefficient = 0.88 ; damping blend = 0.95 ; proportional gain = 0.25 ; error-rate gains = ( 1.20 , 0.20 , 0.15 ) .
RBF-IBVSBasis functions27 centers on { 1 , 0 , 1 } 3 ; Gaussian width = 0.85 .
Adaptive updateLearning rate = 0.025 ; leakage = 0.02 ; compensation gain = 0.55 .
Control parametersConvergence rate = 0.48 ; f ^ max = ( 1.20 , 1.00 , 0.45 ) .
Table 4. Visual-tracking performance of the five methods.
Table 4. Visual-tracking performance of the five methods.
Method | e x | ¯ (px) | e w | ¯ (px) | e ψ | ¯ (rad)Success Rate (%)
PID 15.592 ± 0.843 3.526 ± 0.137 0.09993 ± 0.00168 37.82 ± 11.43
ADRC 25.038 ± 0.700 2.891 ± 0.083 0.11644 ± 0.00103 46.44 ± 3.59
NVPC 36.380 ± 0.492 2.983 ± 0.041 0.18845 ± 0.00070 40.91 ± 2.40
RBF-IBVS 20.878 ± 3.323 2.188 ± 0.058 0.09308 ± 0.00026 64.45 ± 7.14
MHW-PAJ-CLF 15 . 201 ± 3.564 2 . 005 ± 0.222 0 . 07561 ± 0.00446 75 . 50 ± 6.10
Bold values indicate the best mean performance in each column.
Table 5. Tracking and local-response prediction under different perturbations.
Table 5. Tracking and local-response prediction under different perturbations.
TestTracking RMSE E q Local-Response Residual E B
PAJMHWPAJMHW
Reference 0.780 ± 0.119 0 . 716 ± 0 . 037 0.270 ± 0.101 0 . 237 ± 0 . 057
Initial pose 0.803 ± 0.222 0 . 714 ± 0 . 051 0.256 ± 0.054 0 . 250 ± 0 . 014
Noise + dropout 0.745 ± 0.039 0 . 707 ± 0 . 022 0.432 ± 0.014 0 . 413 ± 0 . 015
Command bias 0.799 ± 0.038 0 . 761 ± 0 . 039 0.278 ± 0.020 0 . 264 ± 0 . 019
Delay, 100 ms 0.735 ± 0.006 0 . 728 ± 0 . 016 0.234 ± 0.014 0 . 226 ± 0 . 007
Delay, 200 ms 0.758 ± 0.018 0 . 737 ± 0 . 005 0.248 ± 0.019 0 . 242 ± 0 . 006
Bold values indicate the lower mean value for each metric under each test condition.
Table 6. Visual-tracking metrics at different command-response time constants.
Table 6. Visual-tracking metrics at different command-response time constants.
τ a (ms)Method | e x | ¯ (px) | e w | ¯ (px) | e ψ | ¯ (rad)Success
Rate (%)
E q
0PAJ-CLF16.6081.9120.0824971.510.7451
0MHW-PAJ-CLF14.4401.8060.0712681.830.7287
50PAJ-CLF16.9632.1640.0836262.610.7503
50MHW-PAJ-CLF16.2041.8290.0735969.860.7328
100PAJ-CLF17.4142.0990.0835058.490.7508
100MHW-PAJ-CLF16.0401.8400.0767073.060.7391
Table 7. Control performance of the local-model update methods.
Table 7. Control performance of the local-model update methods.
Method E q E B T first (s) T 3 s (s)
Fixed- B 0 -CLF 0.7500 ± 0.0345 0.2379 ± 0.0229 9.13 ± 1.79 10.44 ± 2.60
PAJ-CLF 0.7798 ± 0.1187 0.2703 ± 0.1005 7.63 ± 2.06 10.00 ± 5.55
MHW-PAJ-CLF 0.7162 ± 0.0368 0.2372 ± 0.0572 5.91 ± 1.68 10.13 ± 3.14
Table 8. Ablation of historical-input weighting.
Table 8. Ablation of historical-input weighting.
Weighting E q Δ E q (%)
Full0.7162 ± 0.0368
Uniform0.7287 ± 0.0297+1.75
Error only0.7268 ± 0.0293+1.48
Shared0.7298 ± 0.0330+1.90
Bold entries indicate the method with the lowest mean tracking error.
Table 9. Historical-input weight statistics.
Table 9. Historical-input weight statistics.
ChannelMean d ¯ i (Steps)Mean N eff , i
Lateral 5.294 ± 0.106 10.258 ± 0.258
Width 5.268 ± 0.058 9.346 ± 0.173
Angle 5.297 ± 0.245 9.526 ± 0.364
Table 10. Validation RMSE values of the diagonal and full local image-response models.
Table 10. Validation RMSE values of the diagonal and full local image-response models.
Model q x q w q ψ Overall
Diagonal0.18050.28770.15250.2150
Full0.18370.28830.15250.2161
Table 11. Effect of the slack-penalty coefficient ρ δ on control performance.
Table 11. Effect of the slack-penalty coefficient ρ δ on control performance.
ρ δ E q Mean Slack
50 0.7188 ± 0.0232 0.0360
100 0.7162 ± 0.0368 0.0340
200 0 . 7150 ± 0 . 0207 0.0350
Bold values indicate the lowest mean value in each performance column.
Table 12. Effects of adaptive gain and history depth on control performance.
Table 12. Effects of adaptive gain and history depth on control performance.
γ N d E q
0.01 10 0.768 ± 0.071
0.02 10 0.716 ± 0.037
0.04 10 0 . 711 ± 0 . 030
0.02 5 0.721 ± 0.020
0.02 15 0.752 ± 0.030
Bold values indicate the lowest mean tracking error among the tested parameter settings.
Table 13. Controller computation time under the repeated reference setting.
Table 13. Controller computation time under the repeated reference setting.
MethodMean (ms)p95 (ms)
PAJ-CLF0.0370.107
MHW-PAJ-CLF0.0380.113
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Hua, L.; Wang, B.; Zhang, Z.; Cheng, Y.; Yuan, Y. Multi-History-Weighted Projection-Adaptive Jacobian Control for Close-Range UAV Visual Servoing near Overhead Ground Wires. Aerospace 2026, 13, 844. https://doi.org/10.3390/aerospace13090844

AMA Style

Hua L, Wang B, Zhang Z, Cheng Y, Yuan Y. Multi-History-Weighted Projection-Adaptive Jacobian Control for Close-Range UAV Visual Servoing near Overhead Ground Wires. Aerospace. 2026; 13(9):844. https://doi.org/10.3390/aerospace13090844

Chicago/Turabian Style

Hua, Liang, Bowen Wang, Zhen Zhang, Yun Cheng, and Yinlong Yuan. 2026. "Multi-History-Weighted Projection-Adaptive Jacobian Control for Close-Range UAV Visual Servoing near Overhead Ground Wires" Aerospace 13, no. 9: 844. https://doi.org/10.3390/aerospace13090844

APA Style

Hua, L., Wang, B., Zhang, Z., Cheng, Y., & Yuan, Y. (2026). Multi-History-Weighted Projection-Adaptive Jacobian Control for Close-Range UAV Visual Servoing near Overhead Ground Wires. Aerospace, 13(9), 844. https://doi.org/10.3390/aerospace13090844

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop