Next Article in Journal
Physics-Informed Residual Convolutional Network Model for Depth-Averaged Landslide Dynamics
Previous Article in Journal
A Learning-Free Noise-Adaptive Framework for Feature-Preserving Point Cloud Denoising
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Graph Attention-Based Distillation for Self-Alignment Localization of UAV Wireless Charging

College of Computer and Control Engineering, Northeast Forestry University, Harbin 150040, China
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(13), 6636; https://doi.org/10.3390/app16136636
Submission received: 13 May 2026 / Revised: 21 June 2026 / Accepted: 1 July 2026 / Published: 2 July 2026
(This article belongs to the Section Electrical, Electronics and Communications Engineering)

Abstract

To address the residual lateral coil misalignment after an unmanned aerial vehicle (UAV) lands on a fixed wireless-charging platform, this study proposes a graph-attention-based knowledge distillation method for embedded self-alignment localization. Four detection-coil voltages form an induced-voltage fingerprint database organized as a multi-scale spatial graph. A graph attention network (GAT) teacher model is trained offline to learn neighborhood correlations in the voltage–position mapping, and its spatial knowledge is distilled into a lightweight Tiny-MLP student model for microcontroller unit (MCU)-based online inference. Experimental results show that the GAT teacher achieves a mean absolute error (MAE) of 0.589 cm, while the distilled Tiny-MLP reduces the MAE of the directly trained Tiny-MLP from 1.548 cm to 1.148 cm (a 25.8% reduction under a fixed seed). In 2000 closed-loop alignment trials with random initial positions, the system achieves an 85.5% success rate under a 0.5 cm threshold, indicating that the method supports low-complexity closed-loop self-alignment for UAV wireless charging.

1. Introduction

Unmanned aerial vehicles (UAVs) have been widely adopted in inspection, monitoring, logistics, and emergency response because of their high mobility and flexible deployment [1,2,3]. However, their endurance is still strongly constrained by the limited capacity and weight of onboard batteries [1]. Wireless power transfer (WPT) provides a promising solution for autonomous return-to-base charging because it enables contactless energy replenishment without manual battery replacement or plug-in operation [2,4]. Fixed docking and charging platforms represent an important implementation form of autonomous UAV energy replenishment [5,6,7,8]. After the UAV is supported by such a platform, the landing structure establishes the nominal relative height and attitude of the charging coils, while a residual lateral offset may remain because of landing-position error and installation tolerance. This lateral offset changes the mutual coupling between the transmitting and receiving coils, reduces the coupling coefficient and transferred power, and may prevent the system from entering a stable high-efficiency charging state [9,10]. Therefore, accurate and low-complexity estimation of the residual lateral offset is an important requirement for fixed UAV wireless-charging platforms.
External vision- or radio-based localization methods can assist UAV landing and positioning [11,12,13]. However, vision-based approaches rely on favorable imaging conditions and a clear line of sight, making them sensitive to illumination changes and occlusion. Radio-based approaches often require additional hardware or infrastructure and may be affected by multipath propagation and electromagnetic interference. These dependencies increase system complexity and deployment cost. Therefore, localization approaches that can be integrated directly into wireless charging systems have attracted increasing attention as a means of reducing reliance on external sensing infrastructure.
From the algorithmic perspective, existing misalignment detection and localization methods can be roughly divided into three categories. (1) The first category is model-driven or analytical estimation methods, in which coil offset is inferred from mutual-inductance models, magnetic-field distribution, or auxiliary sensing coils. These methods have good physical interpretability and usually require limited training data [10,14], but their accuracy strongly depends on coil parameters, vertical distance, attitude variation, and calibration conditions. (2) The second category is search- or fingerprint-matching-based methods, such as array-coil scanning, lookup tables, weighted nearest-neighbor matching, and coil-selection strategies [15,16]. These methods can adapt to nonlinear spatial magnetic-field distributions, but they often require online distance calculation, neighborhood search, or repeated array activation, which increases inference latency and limits their use in fast embedded closed-loop alignment. (3) The third category is data-driven regression methods, where machine learning models directly map multi-channel sensing signals to spatial coordinates [16,17,18]. Such methods can learn nonlinear voltage–position relationships and reduce the need for explicit modeling. However, conventional lightweight models such as multilayer perceptrons usually treat each sampling point as an independent sample and cannot fully exploit the spatial continuity and neighborhood correlations contained in the induced-voltage fingerprint map.
Recent embedded-AI architectures, including depthwise-separable convolutional networks, compact Transformer models, and prototype-based EdgeML methods, have demonstrated promising performance in resource-constrained intelligent systems [19,20,21]. Nevertheless, these methods are typically designed to exploit specific data structures, such as local spatial correlations, long-range dependencies, or prototype-based representations. Whether such architectural characteristics are advantageous for low-dimensional induced-voltage localization remains unclear and has not been systematically investigated in UAV wireless-charging alignment applications.
Graph neural networks (GNNs) provide a potential way to model such spatial correlations because sampling points in the localization area can be naturally organized as graph nodes [22,23]. In particular, graph attention networks can adaptively aggregate information from neighboring nodes and enhance the representation of local and regional spatial features [22]. However, direct deployment of a graph model on a resource-constrained microcontroller is difficult because online graph construction and message passing introduce additional memory consumption and computational latency [20,21,24]. This leads to a contradiction between localization accuracy and embedded deployability: high-capacity graph models are suitable for offline spatial feature learning, whereas compact feedforward networks are more suitable for real-time microcontroller unit (MCU) inference.
To address such issues, this study proposes a graph-attention-based knowledge distillation method for self-alignment localization in UAV wireless charging. In the offline stage, four-channel induced-voltage fingerprints are collected and organized into a multi-scale spatial graph, enabling the graph attention network (GAT) teacher model to learn neighborhood correlations in the voltage–position mapping. The learned spatial knowledge is then distilled into a lightweight Tiny-multilayer perceptron (Tiny-MLP), which is deployed on the MCU for online localization. In this way, the proposed method preserves the spatial modeling ability of graph learning while avoiding online graph construction and message passing, making it suitable for real-time closed-loop alignment control.
Key contributions of this paper are summarized as follows: (1) For the practical coil-misalignment problem in UAV self-aligning wireless charging, a four-channel induced-voltage localization scheme is developed, which realizes lateral offset estimation using the charging-platform sensing coils rather than external visual or ranging sensors. (2) To overcome the limitation of existing voltage-fingerprint localization methods that treat sampling points independently, a graph-structured fingerprint representation is constructed to introduce spatial neighborhood correlations into the voltage–position mapping. (3) To solve the deployment conflict between graph-based spatial modeling and MCU-based real-time control, a GAT-assisted knowledge distillation framework is proposed, in which spatial knowledge is learned offline by the GAT teacher and transferred to a lightweight Tiny-MLP for online closed-loop self-alignment.
The remainder of this paper is organized as follows. Section 2 describes the system architecture and working principle, including the localization circuit, fingerprint-database construction, teacher–student models, and MCU deployment. Section 3 presents the localization, ablation, and closed-loop alignment results. Section 4 discusses the main findings, applicable operating conditions, and current limitations. Section 5 concludes the paper and summarizes its engineering significance.

2. Materials and Methods

2.1. Structure of the UAV Wireless Charging Alignment Platform

The system described in this study primarily comprises a wireless charging main circuit, a detection coil array, a controller, and a two-dimensional mobile platform, as shown in Figure 1. The receiving coil is carried by the UAV, whereas the transmitting coil, detection coils, and alignment mechanism are integrated into the charging platform.
The alignment process begins after the UAV has landed and is stably supported by the platform. At this stage, the landing gear and supporting structure establish a repeatable relative geometry between the transmitting and receiving coils, including their nominal vertical separation and approximately parallel orientation. Landing-position error may nevertheless leave a lateral displacement between the coil centers, which is the alignment variable estimated and compensated in this study.
Before entering the energy-transfer mode, the transmitter generates a high-frequency magnetic field for alignment detection. The lateral position of the receiving coil changes the local coupling relationships, and the four detection coils arranged around the transmitting coil produce induced-voltage signals corresponding to the current lateral misalignment.
These induced-voltage signals are conditioned and then sampled by the controller. Based on the trained lightweight localization model, the controller estimates the current two-dimensional position of the receiving coil and calculates the corresponding offset relative to the target center. The XY platform is then driven to compensate for this offset, so that the receiving coil gradually approaches the center of the transmitting coil. After each movement, the system performs voltage sampling and position estimation again, forming an iterative closed-loop alignment process. Therefore, the overall alignment mechanism can be summarized as “sampling—inference—compensation”.

2.2. Localization Circuit Design and Localization Fingerprint

To meet the weight reduction requirements of UAV and the characteristics of the compensation network, an LCC-S resonant compensation structure is adopted [25]. As shown in Figure 2, the localization auxiliary array consisting of four coils arranged symmetrically is placed above the transmit coil. Only one detection coil is shown in the figure.
This circuit contains a power main circuit and a detection and control circuit. The primary side of the main circuit includes a full-bridge inverter, an LCC-S compensation network ( L t , C p , C t ), and a transmitting coil (self-inductance L p , parasitic resistance R p ); the secondary side comprises a receiving coil (self-inductance L s , parasitic resistance R s ), a resonant capacitor C s , and a rectifier bridge (with a parallel filter capacitor C o ); the load side is controlled by S W 1 and S W 2 to switch between operating modes. When S W 1 is open and S W 2 is closed, the system operates in alignment detection mode, with a virtual load R e connected to establish a stable excitation and acquire measurements from the detection coil; when S W 1 is closed and S W 2 is open, the system switches to energy transfer mode for high-efficiency charging after alignment. The detection loop consists of an auxiliary coil L x magnetically coupled to the main power link, and a current-limiting resistor R x , which outputs an induced voltage U x ( x = 1,2 , 3,4 ) that characterizes the coupling state.
Since the resonant network suppresses higher-order harmonics, the fundamental-wave equivalent mutual inductance model (established in Figure 3) is used to analyze the coupling relationship between the coils and the induced voltage:
Let the system operating frequency be ω . When the transmitting coil and the detection coil are completely decoupled ( M p x = 0 ), Kirchhoff’s voltage law and the equivalent virtual load at the receiving side yield:
U x = j ω M s x M s p 8 R e π 2 + R s L t
When the receiving coil is displaced relative to the transmitting coil, the mutual-coupling relationships among the transmitting, receiving, and detection coils change, resulting in corresponding variations in the four induced voltages. Vertical separation and roll/pitch deviations can also modify this voltage–position relationship. In the fixed post-landing platform considered in this study, however, these quantities are constrained by the landing and supporting structure and remain consistent during fingerprint acquisition and online localization. Their systematic effects are therefore incorporated into the calibrated fingerprint database rather than treated as unknown localization variables. If the nominal platform geometry changes, the fingerprint database should be recalibrated accordingly. Moreover, the transmitting and receiving coils are circular planar coils arranged in a stacked configuration and are rotationally symmetric about their normal axes; thus, yaw rotation does not change their ideal geometric overlap and is not independently estimated. The localization task is therefore formulated as a regression from the four induced voltages to the two-dimensional lateral coordinates.

2.3. System Localization Implementation

According to Figure 4, a complete alignment control chain is established and partitioned into offline and online stages. In the offline stage, static sampling is performed within the target positioning area based on a preset grid. The four induced voltages and true coordinates corresponding to each sampling point are recorded to construct a static induced voltage fingerprint database. Subsequently, the GAT teacher model is trained based on this graph structure, and a lightweight Tiny-MLP student model suitable for online deployment is obtained through knowledge distillation. In the online stage, the MCU acquires four induced voltages and preprocesses the measurements using the normalization parameters saved offline. Then the Tiny-MLP model is invoked to output an estimated coordinate for the current position; the XY platform is controlled based on the offset between the predicted coordinates and the target center so as to perform positional compensation. Iterative corrections are made using a preset threshold until the charging state is triggered after meeting the alignment requirements.

2.4. Building of the Static Induced Voltage Fingerprint Database

Static grid sampling is performed within the predefined position area, as shown in Figure 5. The positioning area measures 16 cm × 16 cm, with a coordinate range of [−8, 8] cm and a sampling interval of 0.5 cm, resulting in a 33 × 33 grid size comprising 1089 sampling nodes. At each of the 1089 grid points, ten consecutive measurements were collected under the same excitation condition, and their mean was used as the fingerprint feature to reduce random measurement noise, analog-to-digital converter (ADC) fluctuation, and short-term excitation jitter. The dataset was randomly divided into 762 training nodes, 163 validation nodes, and 164 test nodes.
The static fingerprint database serves as the calibration of the complete physical charging platform rather than as a universal voltage lookup table. In addition to lateral position, the measured fingerprints inherently include the systematic effects of the actual coil parameters, nominal air gap, installation geometry, excitation circuit, signal-conditioning channels, and sensor gains of the prototype. The same platform configuration and signal-acquisition chain are used during online localization, ensuring consistency between calibration and inference. For another charging platform or a substantially modified hardware configuration, the same grid-sampling procedure can be used to establish a new fingerprint database, after which the proposed graph-learning and distillation framework is applied without changing its overall methodology.
Let the induced voltages output by the four detection coils at the i-th sampling point be U 1 i , U 2 i , U 3 i and U 4 i , respectively. Then, the feature vector at this point is expressed as:
V i = U 1 i , U 2 i , U 3 i , U 4 i T
Its true coordinate label is denoted by P i = ( x i , y i ) T , then the entire dataset can be represented as:
D = V i , P i i = 1 N ,   N = 1089
Taking into account the local continuity of the induced voltage distribution in adjacent spatial regions, the sampling points are further organized into a graph-structured sample set:
G = V , E , X , P
where V represents the set of nodes, corresponding to each sampling point; E refers to the set of edges; X R N × 4 is the node feature matrix, where each row consists of four channels of induced voltages; and P R N × 2 is the node coordinate label matrix.
Inspired by grid-neighborhood graph construction for spatial graph learning [23], this study first establishes single-scale 8-neighborhood connections at a distance of one grid step and further introduces a multi-scale graph structure by adding 16-neighborhood connections at a distance of two grid steps and self-loops for each node, thereby enabling the model to simultaneously capture both local details and broader spatial trends during message propagation. Let the planar coordinates of nodes i and j be p i and p j , respectively. The edge weights are then modeled using a Gaussian decay function based on Euclidean distance:
w i j = exp p i p j 2 2 σ 2
where = 0.5   c m represents the grid spacing.
To determine an appropriate decay scale, σ was swept over {0.5 , 1.0 , 1.5 , 2.0 , 3.0 } while keeping all other settings fixed, and the teacher mean absolute error (MAE) was evaluated on the validation set (Table 1).
Although the validation MAE continued to decrease as σ increased, σ = 1.5 (i.e., σ = 0.75   c m ) was selected to preserve the locality of neighborhood aggregation and avoid making the multi-scale graph close to uniform weighting. This setting is used consistently in the following experiments.
To mitigate the impact of dimension differences in the samples, the Min-Max normalization is performed on the input voltage of each channel. Let the normalized result of the k-th channel voltage ( k = 1,2 , 3,4 ) be:
U ~ i , k = U i , k U k m i n U k m a x U k m i n
where U k m i n and U k m a x denote the minimum and maximum values of the k-th channel voltage in the training set, respectively. This normalization parameter will be deployed on the MCU alongside the student model parameters for online inference.
To address the training instability of lightweight student models during large-scale coordinate regression, normalization is also applied to the coordinate labels. Let the original coordinates be p i = [ x i , y i ] T 8,8 2   c m ; the normalized representation is given by:
p ~ i = p i p m i n p m a x p m i n
where p m i n = [ 8 , 8 ] T , p m a x = [ 8,8 ] T . This transformation uniformly maps the target space to the range [0, 1], which contributes to reducing gradient fluctuations during the training of lightweight networks and aligns well with the Sigmoid constraint on the output layer of student model, thereby enhancing convergence stability. The key static-sampling and graph-construction parameters used in the fingerprint database are summarized in Table 2.

2.5. GAT-Based Teacher Localization Model

Since conventional MLPs struggle to fully utilize the spatial correlation information between nodes within the localization area, a GAT is introduced as the teacher model to extract local topological features from graph structure fingerprints through a neighborhood attention aggregation mechanism.
Let the input features of node i be h i , which, after a linear transformation, are obtained as:
z i = W h i
For node i and its neighboring node j Ɲ ( i ) , the attention weight is defined as:
e i j = L e a k y R e L U a T z i z j
In the equation, W and a are learnable parameters, and [ · · ] indicates vector concatenation. Given that the graph constructed in this study includes multi-scale edges and Gaussian distance weight w i j , the edge weight modulation term is incorporated into the attention calculation during actual implementation, so that the contributions of distant neighbors are subject to attenuation constraints prior to attention normalization. The normalized neighborhood weights are expressed as:
α i j = exp e i j w i j k ϵ Ɲ i exp e i k w i k
Then the aggregate output of node i is:
h i = σ j ϵ Ɲ i α i j z j
where σ ( · ) represents a nonlinear activation function. After propagating through multi-layer graph attention, the node features are fed into a regression head to output coordinate prediction P ^ i = ( x ^ i , y ^ i ) .
In practice, the teacher model adopts a 3-layer graph attention convolutional structure, with each layer using a 4-head attention mechanism. To improve training stability and feature representation, the following engineering optimizations are introduced:
(1)
Layer Normalization is applied instead of BN to achieve more stable normalization of node-level features;
(2)
Residual connections are introduced from the second layer onward to mitigate the gradient vanishing problem in deep graph networks;
(3)
The rectified linear unit (ReLU) is replaced by exponential linear unit (ELU) activation function to mitigate the dying-ReLU problem;
(4)
The multi-scale graph edge weights are incorporated into the attention coefficient calculation, allowing information from distant neighbors to be modulated by distance decay during propagation.
Following the graph attention layer, a three-layer fully connected regression head (64→32→2) is used, as shown in Figure 6, which outputs the final 2D coordinate estimates in combination with Dropout regularization, thereby strengthening the teacher model’s ability to model spatial relationships.
When training the teacher model using a supervised regression objective, considering the possible small number of outlier samples in the position data, this study employs a weighted combination of the mean squared error (MSE) loss and the Huber loss:
L G A T = 0.7 L M S E + 0.3 L H u b e r
where the MSE loss helps strengthen the model’s penalization of large prediction errors, while the Huber loss is more robust to outliers.
The AdamW optimizer is used during training, with weight decay set to 5 × 10 4 , and initial learning rate set to 5 × 10 4 ; and a learning rate scheduling strategy of “Warm-up for the first 30 epochs + follow-up cosine-smoothed decay” is adopted; gradient clipping is applied to prevent gradient explosion, and an early stopping mechanism is implemented; patience is set to 120 epochs, with a maximum training duration of 800 epochs. This training strategy facilitates stable training and superior validation in scenarios with small sample sizes and regular grids.

2.6. Student Model Building and MCU Deployment

2.6.1. Tiny-MLP Student Model and Knowledge Distillation Strategy

To achieve lightweight online localization, the Tiny-MLP student model is built in Figure 7, which takes 4-dimensional induced voltages as input and outputs 2-dimensional coordinate estimates.
The deployed Tiny-MLP adopts a 4→48→24→2 architecture, with BN layers used after the two hidden layers during training. After BN fusion, the model contains 1466 parameters and retains a standard feedforward inference path.
Considering the instability of lightweight networks in large-scale coordinate regression, the following improvements are introduced to the student model.
First, coordinate normalization combined with a Sigmoid output constraint is applied. In this study, the coordinate labels are normalized within [0, 1] and a Sigmoid activation function is used on the student model output layer, thereby constraining the network output to this range to prevent out-of-bounds predictions. During online inference, the MCU performs de-normalization on the Sigmoid output to restore the actual coordinate values, thus significantly improving the trainability of lightweight networks.
Second, in order to accelerate the training convergence of Tiny-MLP, a BN layer is introduced after each of the two hidden layers. Since BN layers can be folded into adjacent fully connected layers during deployment [26], BN is used only as a training aid in this study. To avoid storing additional means and variances on the MCU side, the BN parameters are fused into the preceding fully connected layer before deployment. Let the weights and biases of the fully connected layer be W and b, respectively, and the BN parameters be γ , β , μ , σ 2 , ε , then the fused weights and biases are:
W f u s e d = γ σ 2 + ε W
b f u s e d = γ σ 2 + ε b μ + β
After fusion, the inference path of the deployed model remains a standard fully connected network, incurring no additional computational or memory overhead.
Subsequently, during the knowledge distillation, let the student model output be p ^ i S , the teacher model output be p ^ i T , and the ground truth label be p ~ . This study adopts a joint distillation loss function:
L K D = 1 λ L h a r d + λ L s o f t
where
L h a r d = 1 N t r a i n i D t r a i n p ^ i S p ~ i 2 2
L s o f t = 1 N a l l i = 1 N a l l p ^ i S p ^ i T 2 2
λ denotes the distillation weight coefficients. The hard-label weight λ balances supervision from ground-truth labels against the teacher’s soft targets. As shown in Table 3, a smaller λ yields lower validation error because the student leans more on the teacher’s smoother mapping; however, relying solely on soft targets (λ→0) makes the student fully dependent on the teacher and inherits any teacher bias, reducing robustness to teacher error. A moderate λ = 0.3 was therefore chosen to retain a meaningful contribution from the true labels while still benefiting from distillation.
Furthermore, this study introduces a warm-start initialization strategy: the Tiny-MLP is first trained under direct supervision to acquire preliminary localization capabilities; then its model parameters are used as the initial weights for the distillation student. This prevents the student model from getting stuck in poor local optima when starting from random points, thereby enhancing the stability of the distillation training and the final accuracy.
Other training settings for the student model are as follows: the Adam optimizer (Adaptive Moment Estimation) is used, with the initial learning rate at 5 × 10 3 ; and the learning rate scheduling employs the ReduceLROnPlateau strategy with parameters set to factor = 0.5 and patience = 150.
To evaluate the robustness of the distillation improvement against random initialization, a paired multi-seed experiment was conducted using 20 student-model seeds. For each seed, a Tiny-MLP was first trained under hard-label supervision to obtain a common warm-start checkpoint, which was then duplicated into two branches with identical random states. The direct-continuation branch was further optimized using only the ground-truth labels, whereas the distillation branch was trained using the joint hard- and soft-label loss defined in Equation (15). Both branches used the same data split, initialization, learning rate, training budget, and early-stopping settings, while a fixed pretrained GAT teacher was maintained across all runs. Results are reported as the mean ± standard deviation over the 20 paired runs. The 95% confidence intervals of the paired reductions were estimated using 10,000 bootstrap resamples, and statistical significance was assessed using two-sided paired Wilcoxon signed-rank tests with Holm correction across MAE, RMSE, and P90.

2.6.2. MCU Deployment and Closed-Loop Control

The trained Tiny-MLP parameters, fully connected parameters after BN fusion, and voltage and coordinate normalization parameters are deployed to the MCU, so that the system can independently perform forward inference from 4-dimensional voltage inputs to 2-dimensional coordinate outputs at the edge, without relying on a host computer or complex graph operations. Compared to methods requiring online graph propagation, neighborhood search, or distance matching, the deployed Tiny-MLP involves only a few matrix-vector multiplications and activation operations, with a fixed computation path and stable runtime latency, making it more suitable for online control on general MCU platforms, as shown in Table 4. For comparison, a smaller 4→32→16→2 configuration is also listed in Table 4.
To ensure smooth charging, a closed-loop control strategy is introduced. Let the predicted current position be P ^ = [ x ^ , y ^ ] T , and the target center coordinates be P 0 = [ x 0 , y 0 ] T ; then, the current alignment error can be expressed as:
e = x ^ x 0 2 + y ^ y 0 2
At each iteration, the MCU drives the XY platform in the compensation direction of the predicted offset. After the movement is completed, the four induced voltages are sampled again and fed into the Tiny-MLP model to update the position estimate. The closed-loop process is repeated until the residual alignment error is less than the preset threshold of 0.5 cm or the maximum number of iterations is reached. In this study, the maximum number of iterations is set to five. If the residual error remains larger than the threshold after five iterations, the trial is regarded as unsuccessful; otherwise, the system switches from alignment detection mode to wireless charging mode.

2.7. Experimental Platform and Settings

To verify the effectiveness of the proposed method, an alignment experimental platform for UAV wireless charging is constructed as shown in Figure 8. All settings are consistent with those described earlier, and the system hardware parameters are listed in Table 5.
In this prototype, the nominal vertical separation between the transmitting and receiving coils was set to 4.8 cm by the landing and mechanical support structure. The same physical configuration was used during fingerprint acquisition, model evaluation, and the 2000 closed-loop alignment trials. The experiments therefore evaluate the complete sensing-inference-compensation process on an actual hardware platform rather than using simulated voltage data or an idealized numerical model.
The deployment-oriented comparison included a MobileNet-style DS-CNN, a Tiny Transformer, ProtoNN, and ResMLP. The selected DS-CNN used 24 stem channels and 36 pointwise-convolution channels. The Tiny Transformer contained two encoder layers with a model dimension of 32, two attention heads, and a feedforward dimension of 64. ProtoNN used a projection dimension of 4 and 32 prototypes. All models used the same data split and normalization procedure, and their configurations were selected according to validation-set MAE before final evaluation on the test set.

3. Results

3.1. Localization Performance Comparison

To comprehensively evaluate the proposed method, thirteen models were trained and tested on the same fingerprint dataset using an identical 762/163/164 training, validation, and test split. All localization metrics were calculated on the same 164-node test set. A fixed random seed of 42 was used for the cross-model comparison. According to their roles in the localization framework, the models were divided into three groups: graph-based teacher models, deployment-oriented localization models, and conventional baselines. The teacher group included GAT, GraphSAGE, and GCN. The deployment-oriented group included the directly trained and distilled Tiny-MLP models, MobileNet-style DS-CNN, Tiny Transformer, ProtoNN, and ResMLP. The conventional baseline group consisted of GPR, WKNN, SVR, and XGBoost. The results are summarized in Table 6 and Figure 9.
Table 6 shows that GAT achieved the best performance among the teacher candidates, obtaining the lowest MAE, RMSE, maximum error, and P90, and was therefore selected as the teacher model for knowledge distillation.
For the deployment-oriented models, knowledge distillation improved the performance of Tiny-MLP across all evaluation metrics. ResMLP achieved the lowest MAE, whereas the distilled Tiny-MLP obtained the lowest RMSE, maximum error, and P90. The distilled Tiny-MLP also outperformed DS-CNN, Tiny Transformer, ProtoNN, and the directly trained Tiny-MLP, indicating that it provides a favorable balance between localization accuracy and error stability.
Compared with the conventional baselines, including GPR, WKNN, SVR, and XGBoost, the distilled Tiny-MLP consistently achieved lower localization errors. These results demonstrate the effectiveness of combining graph-based spatial learning with a compact student network for embedded localization.
Figure 9 provides a role-based comparison of the evaluated models. Within the teacher group, GAT shows a clear advantage over GraphSAGE and GCN. Within the deployment-oriented group, ResMLP achieves the lowest MAE, whereas the distilled Tiny-MLP achieves the lowest RMSE and substantially lower error than the directly trained Tiny-MLP, DS-CNN, Tiny Transformer, and ProtoNN. The conventional baseline group exhibits generally higher errors, particularly in RMSE, indicating weaker control of large localization deviations.
To further evaluate the balance between localization accuracy and implementation complexity, Figure 10 plots the test MAE against a model-size proxy on a logarithmic scale. Because different model families adopt different representations, the horizontal coordinate denotes the principal stored model elements of each method, including network parameters, stored fingerprints, support vectors, or tree nodes.
As shown in Figure 10, the GAT teacher achieves the lowest MAE but requires the largest model scale. ResMLP obtains a slightly lower MAE than the distilled Tiny-MLP, but uses substantially more parameters and exhibits higher RMSE, maximum error, and P90. Among the compact deployment-oriented models, the distilled Tiny-MLP achieves lower localization errors than the DS-CNN, Tiny Transformer, ProtoNN, and directly trained Tiny-MLP. For the present four-channel voltage-regression task, transferring the spatial mapping learned by the graph teacher to a compact feedforward network therefore provides a favorable balance between localization accuracy, error stability, and implementation complexity.

3.2. Training Convergence and Statistical Validation

To examine both representative training convergence and the robustness of the knowledge-distillation improvement, Figure 11 presents the training and validation MAE curves of the GAT teacher and distilled Tiny-MLP. Table 7 summarizes their fixed-seed test performance together with that of the directly trained Tiny-MLP, while Table 8 reports the paired multi-seed statistical results.
As shown in Figure 11, the GAT teacher and distilled Tiny-MLP both exhibit stable convergence, with their validation curves remaining close to the corresponding training curves. The GAT teacher reaches a lower final error because of its stronger graph-based spatial representation, whereas the distilled Tiny-MLP retains a compact feedforward structure suitable for embedded inference. Since representative convergence curves alone cannot determine whether the improvement is robust to random initialization, the direct and distilled student models are further evaluated through the paired multi-seed analysis presented below.
Under the fixed-seed setting, knowledge distillation reduced the Tiny-MLP MAE from 1.548 cm to 1.148 cm. The RMSE, maximum error, and P90 were also reduced from 2.173, 8.973, and 3.424 cm to 1.426, 5.502, and 2.250 cm, respectively. These results correspond to reductions of 25.8% in MAE, 34.4% in RMSE, 38.7% in maximum error, and 34.3% in P90. To determine whether these improvements persist across different student-model initializations, a paired 20-seed evaluation was further conducted, as reported in Table 8.
Across the 20 paired runs, knowledge distillation reduced the mean MAE from 1.531 cm to 1.142 cm, corresponding to a paired reduction of 0.389 cm (95% CI: 0.349–0.433 cm) and a mean relative reduction of 25.18%. Consistent reductions were also observed in RMSE and P90, and all differences remained statistically significant after Holm correction (p = 5.72 × 10−6). These results confirm that the improvement is robust across different student-model initializations.

3.3. Analysis of Test Set Error Distribution and Spatial Positioning Results

To further analyze the positioning performance on the test set, the predicted-point distribution and the spatial distribution of localization error are examined, as shown in Figure 12.
As shown in Figure 12a, the predicted points of the distilled Tiny-MLP follow the ground truth closely across the positioning area, indicating that the proposed optimizations effectively strengthen the student’s spatial representation. The error map in Figure 12b further shows that the error remains low over most of the area, with only a few isolated hotspots near the boundaries and corners. These hotspots may be associated with reduced fingerprint discrimination in these regions. Overall, the distilled student maintains accurate and spatially consistent localization across the working area.
The error distributions of the two key models—the GAT teacher and the distilled Tiny-MLP—are compared in Figure 13. The distilled student shows a narrow interquartile box with a low median error, and although its distribution sits slightly higher than the teacher’s, both its central tendency and upper tail (P90 = 2.250 cm) remain tightly bounded, indicating well-suppressed large errors and stable output. The teacher exhibits an even tighter distribution (P90 = 1.046 cm), consistent with its role as the higher-accuracy teacher reference that the lightweight student approaches.

3.4. Ablation Study

To assess the contribution of each design component, an ablation study was conducted by removing one component at a time from the full model while keeping all other settings fixed. The graph-structure component (16-neighborhood connectivity) was evaluated on the GAT teacher, and the training components (warm-start, Sigmoid output constraint, and BN-assisted training) on the distilled Tiny-MLP. All variants were evaluated on the same 164-node test set, as summarized in Table 9.
For the teacher, removing the 16-neighborhood connections and retaining only the 8-neighborhood graph raises the MAE from 0.589 cm to 0.768 cm, confirming that the multi-scale structure provides useful regional context beyond immediate neighbors.
For the student, removing warm-start or BN-assisted training increases both MAE and the tail error (P90), indicating that these components improve convergence and suppress large deviations. Removing the Sigmoid constraint slightly increases the mean error and markedly increases the P90 error (from 2.250 cm to 2.364 cm); since the Sigmoid also bounds the output to a valid coordinate range and prevents out-of-range predictions, it is retained for its stability and tail-error benefits. Overall, each component contributes positively to either accuracy or robustness, supporting the proposed design.

3.5. Analysis of Closed-Loop Alignment Experiment Results

To verify the practical performance of the distilled Tiny-MLP in closed-loop alignment, an iterative “sampling—inference—compensation” process was established based on the two-dimensional offset predicted by the model, with the alignment threshold set to 0.5 cm and a maximum of five iterations per trial.
To statistically evaluate the overall alignment performance, 2000 closed-loop trials were conducted with randomly distributed initial positions across the positioning area. Under the 0.5 cm threshold, the system achieved an overall success rate of 85.5%, indicating that the proposed closed-loop strategy is effective over a wide range of random initial offsets. From these trials, a representative subset was further examined to analyze the convergence behavior in more detail, and 12 typical cases covering different initial-error ranges were selected for case-by-case illustration, as listed in Table 10.
Among these 12 representative cases, 9 converged within the threshold; the successful ones required 2.9 iterations, reached a final error of 0.408 cm, and completed alignment in 0.42 s on average, showing that the deployed lightweight model can achieve effective correction within a small number of iterations. For the three cases that did not meet the 0.5 cm criterion within five iterations, the residual errors (0.738 cm, 0.890 cm, and 0.668 cm) were nonetheless substantially reduced from their initial offsets. These selected cases are intended to illustrate the per-trial convergence process rather than to re-estimate the overall success rate, which is given by the 2000-trial statistic above.
Figure 14 shows that, for the illustrated successful cases, the platform progressively approaches the target along the model-predicted compensation direction. This demonstrates that the distilled Tiny-MLP can provide effective directional and displacement information for closed-loop correction.
Figure 15 shows the curves of residual error versus the iteration number for 12 sets of samples, in which the error curves for successful samples mostly cross the 0.5 cm threshold line within a small number of iterations, while those for failed samples, although showing a downward trend, still fail to fall below the threshold within the maximum number of iterations. This indicates that the adopted closed-loop compensation method exhibits generally convergent behavior under the tested conditions. However, the accumulation of model prediction errors may still affect the final closed-loop accuracy for some samples with large offsets or in locally complex areas.

4. Discussion

The experimental validation targets a representative fixed UAV wireless-charging platform. Localization begins after the UAV is mechanically supported, so the nominal air gap and coil attitude are established before lateral alignment. The proposed method therefore estimates the residual lateral offset rather than the full pose of an unconstrained UAV during descent.
The induced-voltage mapping is platform-dependent and is calibrated through the static fingerprint database, which captures the electromagnetic and sensing characteristics of the assembled device. The trained model is therefore specific to a given hardware configuration, whereas the localization framework can be transferred to another platform by reconstructing its fingerprint database.
The experiments validate the complete sensing, localization, and closed-loop compensation process on physical hardware. Operation before stable mechanical support or under severe external disturbances remains outside the scope of the present study and will be investigated in future work.

5. Conclusions

This study developed a graph-attention-based knowledge-distillation framework for post-landing lateral alignment on fixed UAV wireless-charging platforms. The offline GAT teacher captures spatial correlations in the induced-voltage fingerprint map, while the distilled Tiny-MLP performs lightweight coordinate inference on the MCU. The distilled model reduced the fixed-seed MAE from 1.548 cm to 1.148 cm, and the paired multi-seed evaluation confirmed that this improvement was statistically robust. The complete system further achieved an 85.5% success rate in 2000 closed-loop alignment trials.
The significance of the proposed framework lies in separating platform calibration and high-capacity spatial learning from real-time embedded control. Complex electromagnetic mappings can be learned offline, while online alignment is implemented through a compact and deterministic inference path. This provides a practical approach for integrating tiny machine learning (TinyML) into autonomous charging platforms and other robotic energy systems subject to strict computational, memory, and power constraints.

Author Contributions

Conceptualization, B.A. and J.L.; methodology, B.A. and C.Z.; software, B.A.; validation, B.A., J.L. and P.S.; formal analysis, C.Z.; data curation, B.A.; writing—original draft preparation, B.A. and J.L.; writing—review and editing, B.A. and D.Y.; visualization, P.S. and C.Z.; supervision, D.Y.; project administration, D.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
UAVunmanned aerial vehicle
WPTwireless power transfer
MAEmean absolute error
GNNgraph neural network
MCUmicrocontroller unit
GATgraph attention network
MLPmultilayer perceptron
ADCanalog-to-digital converter
BNbatch normalization
ReLUrectified linear unit
ELUexponential linear unit
MSEmean squared error
RMSEroot mean square error
TinyMLtiny machine learning

References

  1. Gubran, A.Q.A.; Qasem, N.A.A.; Abdelrahman, W.G.; Abdallah, A.M. A review of powering unmanned aerial vehicles by clean and renewable energy technologies. Sustain. Energy Technol. Assess. 2025, 73, 104150. [Google Scholar] [CrossRef] [Scilit]
  2. Šćuric, A.; Krznar, N.; Penđer, A.; Štedul, I.; Kotarski, D. Autonomous multirotor UAV docking and charging: A comprehensive review of systems, mechanisms, and emerging technologies. Symmetry 2025, 17, 1988. [Google Scholar] [CrossRef] [Scilit]
  3. Sun, H.; Xu, R.; Luo, J.; Cheng, H. Review of the application of UAV edge computing in fire rescue. Sensors 2025, 25, 3304. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Cheng, Y.; Yu, S.; Zhang, X.; Zhang, R.; Liu, P.; Yu, S. An efficient magnetic coupler with tight coupling, precise alignment, and low leakage shielding for UAV wireless charging. Electronics 2025, 14, 4358. [Google Scholar] [CrossRef] [Scilit]
  5. Jain, S.; Bharadwaj, A.; Sharma, A. A mortarboard-shaped receiver antenna for drone wireless charging to achieve an outspread misalignment tolerance towards imperfect landing. IEEE Trans. Veh. Technol. 2025, 74, 61–73. [Google Scholar] [CrossRef] [Scilit]
  6. Aydin, E. A misalignment tolerant magnetic coupler design and implementation for wireless charging of unmanned aerial vehicles (UAVs). IEEE Access 2025, 13, 214618–214626. [Google Scholar] [CrossRef] [Scilit]
  7. Zhao, Y.; Shen, S.; Yin, F.; Wang, L. A high misalignment-tolerant hybrid coupler for unmanned aerial vehicle WPT charging systems. IEEE Trans. Transp. Electrif. 2025, 11, 1570–1581. [Google Scholar] [CrossRef] [Scilit]
  8. Ağçal, A.; Doğan, T.H. A novel folding wireless charging station design for drones. Drones 2024, 8, 289. [Google Scholar] [CrossRef] [Scilit]
  9. Wang, C.; Ren, W.; Chen, Y.; Li, X. Review of high-misalignment tolerance techniques in wireless power transfer systems. Energies 2026, 19, 713. [Google Scholar] [CrossRef] [Scilit]
  10. Gordhan, U.; Jayalath, S. Comparative analysis of wireless power transfer couplers for unmanned aerial vehicles and drones. IEEE Open J. Power Electron. 2024, 5, 618–633. [Google Scholar] [CrossRef] [Scilit]
  11. Semerikov, S.O.; Nechypurenko, P.P.; Vakaliuk, T.A.; Mintii, I.S.; Kolhatin, A.O. Vision-based autonomous UAV landing: A comprehensive review of technologies, techniques, and applications. J. Intell. Robot. Syst. 2025, 111, 115. [Google Scholar] [CrossRef] [Scilit]
  12. Aziz, T.; Koo, I. A comprehensive review of indoor localization techniques and applications in various sectors. Appl. Sci. 2025, 15, 1544. [Google Scholar] [CrossRef] [Scilit]
  13. Stefanoni, M.; Kovács, I.; Sarcevic, P.; Odry, Á. A survey on the main techniques adopted in indoor and outdoor localization. Electronics 2025, 14, 2069. [Google Scholar] [CrossRef] [Scilit]
  14. Abderrazak, F.; Antonino-Daviu, E.; Talbi, L.; Ferrando-Bataller, M. Characteristic modes analyses for misalignment in wireless power transfer system. IEEE Access 2024, 12, 65007–65023. [Google Scholar] [CrossRef] [Scilit]
  15. Cai, C.; Yang, J.; Wu, S.; Zhang, H.; Chai, W. Landing position detection and array coil matching of multi-UAV wireless power transfer system. IEEE Trans. Transp. Electrif. 2025, 11, 11054–11064. [Google Scholar] [CrossRef] [Scilit]
  16. Yuan, D.; Li, L.; Han, Z.; Liu, J.; Zhao, C. Position identification for UAV wireless charging coupler using neural network and voltage fingerprint. Appl. Sci. 2026, 16, 3318. [Google Scholar] [CrossRef] [Scilit]
  17. Issi, F. Modeling and implementation of a machine learning-based wireless charging system with high misalignment tolerance. Ain Shams Eng. J. 2024, 15, 102970. [Google Scholar] [CrossRef] [Scilit]
  18. Li, M.; Li, J.; Xiao, W.; Zhou, C. Design method of array-type coupler for UAV wireless power transmission system based on the deep neural network. Drones 2025, 9, 532. [Google Scholar] [CrossRef] [Scilit]
  19. Jung, V.J.-B.; Burrello, A.; Scherer, M.; Conti, F.; Benini, L. Optimizing the deployment of Tiny Transformers on low-power MCUs. IEEE Trans. Comput. 2025, 74, 526–541. [Google Scholar] [CrossRef] [Scilit]
  20. Cordova-Cardenas, R.; Amor, D.; Gutiérrez, Á. Edge AI in practice: A survey and deployment framework for neural networks on embedded systems. Electronics 2025, 14, 4877. [Google Scholar] [CrossRef] [Scilit]
  21. Somvanshi, S.; Islam, M.M.; Chhetri, G.; Chakraborty, R.; Mimi, M.S.; Shuvo, S.A.; Islam, K.S.; Javed, S.A.; Rafat, S.A.; Dutta, A.; et al. From tiny machine learning to tiny deep learning: A survey. ACM Comput. Surv. 2025, 58, 168. [Google Scholar] [CrossRef] [Scilit]
  22. Sun, C.; Li, C.; Lin, X.; Zheng, T.; Meng, F.; Rui, X.; Wang, Z. Attention-based graph neural networks: A survey. Artif. Intell. Rev. 2023, 56, 2263–2310. [Google Scholar] [CrossRef] [Scilit]
  23. Han, B.; Qu, T.; Jiang, J. GN-GCN: Grid neighborhood-based graph convolutional network for spatio-temporal knowledge graph reasoning. ISPRS J. Photogramm. Remote Sens. 2025, 220, 68–85. [Google Scholar] [CrossRef] [Scilit]
  24. Moslemi, A.; Briskina, A.; Dang, Z.; Li, J. A survey on knowledge distillation: Recent advancements. Mach. Learn. Appl. 2024, 18, 100605. [Google Scholar] [CrossRef] [Scilit]
  25. Rong, C.; Wang, H.; Guo, Y.; Wu, J.; Gao, H.; Cai, W. An IPT-CPT hybrid wireless power transfer system for uncrewed aerial vehicles. IEEE Trans. Ind. Electron. 2026, 73, 8125–8135. [Google Scholar] [CrossRef] [Scilit]
  26. Yvinec, E.; Dapogny, A.; Bailly, K. To fold or not to fold: A necessary and sufficient condition on batch-normalization layers folding. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, Vienna, Austria, 23–29 July 2022; pp. 1601–1607. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Overall system structure: The platform includes the wireless power circuit, detection-coil array, MCU controller, and XY motion stage. White arrows indicate the main components, and red arrows denote the control flow.
Figure 1. Overall system structure: The platform includes the wireless power circuit, detection-coil array, MCU controller, and XY motion stage. White arrows indicate the main components, and red arrows denote the control flow.
Applsci 16 06636 g001
Figure 2. Structure of system localization circuit: The LCC-S resonant circuit with four symmetrically arranged detection coils.
Figure 2. Structure of system localization circuit: The LCC-S resonant circuit with four symmetrically arranged detection coils.
Applsci 16 06636 g002
Figure 3. Equivalent mutual-inductance circuit: The colored regions correspond to the detection coil, transmitter, and receiver in Figure 2, and the black arrows indicate the reference directions of currents and voltages.
Figure 3. Equivalent mutual-inductance circuit: The colored regions correspond to the detection coil, transmitter, and receiver in Figure 2, and the black arrows indicate the reference directions of currents and voltages.
Applsci 16 06636 g003
Figure 4. Process of system positioning implementation: The offline fingerprint construction and online closed-loop alignment workflow.
Figure 4. Process of system positioning implementation: The offline fingerprint construction and online closed-loop alignment workflow.
Applsci 16 06636 g004
Figure 5. Schematic illustration of the static grid-sampling procedure. The actual fingerprint database contains 33 × 33 nodes with a spacing of 0.5 cm.
Figure 5. Schematic illustration of the static grid-sampling procedure. The actual fingerprint database contains 33 × 33 nodes with a spacing of 0.5 cm.
Applsci 16 06636 g005
Figure 6. Diagram of GAT teacher model structure: The 3-layer GAT architecture with 4-head attention, residual connections, and a 64→32→2 regression head.
Figure 6. Diagram of GAT teacher model structure: The 3-layer GAT architecture with 4-head attention, residual connections, and a 64→32→2 regression head.
Applsci 16 06636 g006
Figure 7. Teacher-Student knowledge distillation framework: The GAT teacher transfers knowledge to the lightweight Tiny-MLP student via joint hard and soft loss.
Figure 7. Teacher-Student knowledge distillation framework: The GAT teacher transfers knowledge to the lightweight Tiny-MLP student via joint hard and soft loss.
Applsci 16 06636 g007
Figure 8. Physical prototype of the fixed UAV wireless-charging and alignment platform: The physical experimental setup for UAV wireless charging alignment.
Figure 8. Physical prototype of the fixed UAV wireless-charging and alignment platform: The physical experimental setup for UAV wireless charging alignment.
Applsci 16 06636 g008
Figure 9. Localization-error comparison by model role: MAE and RMSE of graph-based teacher models, deployment-oriented localization models, and conventional baselines.
Figure 9. Localization-error comparison by model role: MAE and RMSE of graph-based teacher models, deployment-oriented localization models, and conventional baselines.
Applsci 16 06636 g009
Figure 10. Accuracy–complexity trade-off of all evaluated models: test MAE versus model-size proxy (log scale). The proxy denotes parameter count for neural and ProtoNN models, stored training elements for GPR, fingerprints for WKNN, support vectors for SVR, and tree nodes for XGBoost.
Figure 10. Accuracy–complexity trade-off of all evaluated models: test MAE versus model-size proxy (log scale). The proxy denotes parameter count for neural and ProtoNN models, stored training elements for GPR, fingerprints for WKNN, support vectors for SVR, and tree nodes for XGBoost.
Applsci 16 06636 g010
Figure 11. Model convergence: Training and validation MAE curves of the GAT teacher and distilled Tiny-MLP.
Figure 11. Model convergence: Training and validation MAE curves of the GAT teacher and distilled Tiny-MLP.
Applsci 16 06636 g011
Figure 12. Spatial positioning results of the distilled Tiny-MLP on the test set: (a) predicted versus ground-truth coordinates; (b) spatial distribution of localization error across the positioning area.
Figure 12. Spatial positioning results of the distilled Tiny-MLP on the test set: (a) predicted versus ground-truth coordinates; (b) spatial distribution of localization error across the positioning area.
Applsci 16 06636 g012
Figure 13. Box-and-whisker plot of the localization error distribution: Error distribution comparison for the GAT teacher and the distilled Tiny-MLP.
Figure 13. Box-and-whisker plot of the localization error distribution: Error distribution comparison for the GAT teacher and the distilled Tiny-MLP.
Applsci 16 06636 g013
Figure 14. Closed-loop alignment trajectories for typical successful samples: Convergence trajectories of four typical successful samples toward the target center.
Figure 14. Closed-loop alignment trajectories for typical successful samples: Convergence trajectories of four typical successful samples toward the target center.
Applsci 16 06636 g014
Figure 15. Variation curves of residual error with the iteration number under different initial offset conditions: The legend denotes the initial offset coordinates, and “Failed” indicates cases that did not converge within five iterations.
Figure 15. Variation curves of residual error with the iteration number under different initial offset conditions: The legend denotes the initial offset coordinates, and “Failed” indicates cases that did not converge within five iterations.
Applsci 16 06636 g015
Table 1. Effect of the Gaussian decay scale σ on teacher validation error.
Table 1. Effect of the Gaussian decay scale σ on teacher validation error.
σ 0.5∆1.0∆1.5∆2.0∆3.0∆
MAE (cm)1.0320.8190.5780.5300.481
Table 2. Static sampling and graph-construction parameters.
Table 2. Static sampling and graph-construction parameters.
ParametersSettings
Position area16 cm × 16 cm
Coordinate range[−8, 8] cm
Sampling interval0.5 cm
Grid size33 × 33
Number of sampling nodes1089
Number of repeated samples per point10 measurements, averaged
Dataset partitioning70%/15%/15% (762/163/164 nodes)
Graph connection method8-neighborhood + 16-neighborhood + self-loop
Edge weightGaussian distance decay, σ = 0.75   c m
Number of edges25,281 (algorithm output)
Table 3. Effect of the hard-label weight λ on student validation error.
Table 3. Effect of the hard-label weight λ on student validation error.
λ 0.10.30.50.71.0
MAE (cm)1.1511.1451.1891.3141.453
Table 4. Model-complexity comparison.
Table 4. Model-complexity comparison.
ModelNetwork StructureNumber of ParametersSingle-Inference Floating-Point Operations (kFLOPs)Storage Space/KBGraph Structure RequiredSuitable for MCU Deployment
GAT Teacher4→GAT × 3→256→64→32→2153,76222,300600.6YN
Standard Tiny-MLP4→32→16→27221.42.8NY
Widened Tiny-MLP4→48→24→214662.95.7NY
Note: The number of parameters refers to the core parameters at deployment; the Tiny-MLP widened configuration is the actual distilled student model used; GAT inference requires all 1089 nodes in the graph to participate in message propagation.
Table 5. Main hardware specifications of the experimental platform.
Table 5. Main hardware specifications of the experimental platform.
ParametersValuesParametersValues
Main control chipSTM32F405RGT6Primary-side topology inductor L t 23.6 µH
DC bus voltage U I 6 VPrimary-side topology capacitor C t 135 nF
Transmitting coil self-inductance L P 61.34 µHPrimary-side resonant capacitor C p 92.1 nF
Receiving coil self-inductance L S 60.55 µHSecondary-side resonant capacitor C s 55.9 nF
Transmitting coil internal resistance R P 315.1 mΩResonant frequency85 kHz
Receiving coil internal resistance R S 137.5 mΩVertical separation4.8 cm
Load R e 5 ΩFlash memory1 MB
Table 6. Localization performance comparison of different models.
Table 6. Localization performance comparison of different models.
GroupModelTypeMAE
/cm
RMSE
/cm
Max Error
/cm
P90
/cm
TeacherGAT TeacherGraph attention0.5890.7322.5501.046
TeacherGraphSAGE TeacherGraph aggregation1.5061.7996.0012.705
TeacherGCN TeacherGraph convolution1.9972.44910.0164.069
DeploymentTiny-MLP DistilledDistilled lightweight MLP1.1481.4265.5022.250
DeploymentTiny-MLP DirectDirectly trained lightweight MLP1.5482.1738.9733.424
DeploymentMobileNet-style DS-CNNDepthwise-separable CNN1.4292.0179.6462.896
DeploymentTiny TransformerCompact self-attention model1.3641.7506.0492.941
DeploymentProtoNNPrototype-based EdgeML2.1372.64910.2054.073
DeploymentResMLPHigh-capacity MLP reference1.0761.7308.5942.305
BaselineGPRKernel regression1.6572.3008.1963.918
BaselineWKNNFingerprint1.7472.83114.7073.759
BaselineSVRKernel regression1.9742.77510.5324.877
BaselineXGBoostTree ensemble baseline2.3512.97411.3334.462
Table 7. Fixed-seed test performance of the main models.
Table 7. Fixed-seed test performance of the main models.
ModelsMAE/cmRMSE/cmMax Error/cmP90/cm
GAT Teacher0.5890.7322.5501.046
Tiny-MLP (Direct)1.5482.1738.9733.424
Tiny-MLP (Distilled)1.1481.4265.5022.250
Table 8. Statistical robustness of knowledge distillation: Paired comparison of direct and distilled Tiny-MLP over 20 seeds.
Table 8. Statistical robustness of knowledge distillation: Paired comparison of direct and distilled Tiny-MLP over 20 seeds.
MetricDirect/cmDistilled/cmPaired Reduction (95% CI)/cmHolm-Adjusted p-Value
MAE1.531 ± 0.0871.142 ± 0.0440.389
[0.349, 0.433]
5.72 × 10−6
RMSE2.176 ± 0.2041.461 ± 0.0780.714
[0.637, 0.801]
5.72 × 10−6
P903.287 ± 0.2772.215 ± 0.1481.072
[0.937, 1.206]
5.72 × 10−6
Table 9. Ablation study of the proposed method.
Table 9. Ablation study of the proposed method.
ModelVariantMAE/cmRMSE/cmP90/cmΔMAE/cm
GAT TeacherFull0.5890.7321.046
GAT Teacherw/o 16-neighborhood0.7680.9451.497+0.179
Tiny-MLP DistilledFull1.1481.4262.250
Tiny-MLP Distilledw/o warm-start1.2081.5872.267+0.060
Tiny-MLP Distilledw/o sigmoid1.1711.4832.364+0.023
Tiny-MLP Distilledw/o BN1.2221.6232.764+0.074
Table 10. Closed-loop alignment results.
Table 10. Closed-loop alignment results.
Starting Position (x, y)/cmInitial Error/cmNumber of IterationsFinal Error/cmTime/sSuccess
(−1.7, 0.6)1.80350.7380.61N
(0.4, 1.8)1.84450.4280.61Y
(1.7, 1.1)2.02520.4820.17Y
(−1.0, −2.2)2.41720.4820.28Y
(0.9, −3.0)3.13240.4360.54Y
(−2.3, 2.5)3.39730.3640.36Y
(−2.9, −2.4)3.76450.8900.55N
(−1.9, −4.1)4.51950.6680.60N
(4.1, 2.1)4.60710.4480.26Y
(−1.1, −4.7)4.82740.4990.70Y
(2.6, 4.2)4.94010.1530.26Y
(0.0, −6.5)6.50040.3770.60Y
Note: Y indicates that the final residual error is less than the threshold, and N indicates failure.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ai, B.; Liu, J.; Yuan, D.; Zhao, C.; Shen, P. Graph Attention-Based Distillation for Self-Alignment Localization of UAV Wireless Charging. Appl. Sci. 2026, 16, 6636. https://doi.org/10.3390/app16136636

AMA Style

Ai B, Liu J, Yuan D, Zhao C, Shen P. Graph Attention-Based Distillation for Self-Alignment Localization of UAV Wireless Charging. Applied Sciences. 2026; 16(13):6636. https://doi.org/10.3390/app16136636

Chicago/Turabian Style

Ai, Binghong, Jiali Liu, Dechun Yuan, Chaoyue Zhao, and Pange Shen. 2026. "Graph Attention-Based Distillation for Self-Alignment Localization of UAV Wireless Charging" Applied Sciences 16, no. 13: 6636. https://doi.org/10.3390/app16136636

APA Style

Ai, B., Liu, J., Yuan, D., Zhao, C., & Shen, P. (2026). Graph Attention-Based Distillation for Self-Alignment Localization of UAV Wireless Charging. Applied Sciences, 16(13), 6636. https://doi.org/10.3390/app16136636

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop