AutoUAVFormer: Neural Architecture Search with Implicit Super-Resolution for Real-Time UAV Aerial Object Detection
Highlights
- This paper introduces AutoUAVFormer, an NAS framework that incorporates super-resolution guidance into Transformer-based detection. By co-optimizing the detection architecture and the super-resolution reconstruction mechanisms in a unified search space, the method enables adaptive and computationally efficient Anti-UAV detection across diverse low-altitude environments.
- AutoUAVFormer demonstrates state-of-the-art performance with unprecedented mAP@0.5 scores of 98.6% (DetFly), 95.5% (DUT Anti-UAV), and 89.9% (UAV Swarm), while maintaining real-time inference. Its superiority over existing methods is attributed to multi-scale feature fusion via evolutionary search strategies.
- AutoUAVFormer integrates neural architecture search (NAS) with a super-resolution auxiliary branch to deliver a practical, resource-efficient and robust solution for Anti-UAV detection, enabling safety in civilian airspace, airports, and urban management.
- AutoUAVFormer validates the potential of an NAS-integrated Transformer for small target detection, offering a scalable framework applicable to diverse low-altitude vision scenarios.
Abstract
1. Introduction
- We propose AutoUAVFormer, a unified NAS framework that jointly optimizes Transformer encoder hyperparameters and the structural parameters of the FPN within a single search space, enabling automated discovery of task-specific architectures for Anti-UAV detection. Unlike prior NAS approaches that optimize backbone or encoder in isolation, our joint co-evolution eliminates suboptimal solutions arising from a sequential design pipeline.
- We design a Multi-Kernel Center-Decoupled FPN (MKCD-FPN) that integrates multi-scale parallel convolutions with progressive depth fusion and center-aggregation at the P4 scale. This module is an incremental improvement over conventional FPN topologies, specifically engineered to mitigate small target feature suppression in cluttered low-altitude backgrounds. Its structural parameters are incorporated into the joint NAS search space for co-optimization with the Transformer encoder.
- Inspired by the training-only auxiliary branch paradigm [32], we adapt and integrate a super-resolution auxiliary branch into the AutoUAVFormer framework. This branch reconstructs high-resolution spatial details from shallow backbone features during training, thereby enhancing fine-grained texture and edge representations critical for small UAV target discrimination, while introducing zero inference overhead. Within the joint NAS framework, the SR-augmented training objective acts as a structural regularizer that encourages the evolutionary search to discover architectures with improved feature discriminability under the given parameter budget constraints.
2. Related Works
2.1. Anti-UAV Methods Based on Deep Learning
2.2. Super-Resolution Reconstruction Methods
2.3. Neural Architecture Search
3. Methodology
3.1. Overall Structure
3.2. Multikernel Center-Decoupled FPN
3.3. Neural Architecture Search
3.3.1. Search Space
3.3.2. Supernet Training
3.3.3. Evolutionary Search
3.4. Super Resolution
4. Experiments
4.1. Datasets
4.2. Training Settings and Evaluation Metrics
4.3. Ablation Experiments
4.3.1. Ablation on the Overall Architecture
4.3.2. MKCD-FPN Module
Multi-Kernel Size Selection
Feature Aggregation Layer Selection
4.3.3. Weight Ratio of the Multi-Task Loss Function
4.3.4. Super Resolution Branch
4.4. Comparison Experiments
4.4.1. Comparison and Analysis of Experiments with General Object Detection Methods
4.4.2. Comparison and Analysis of Experiments with Anti-UAV Detection Methods
| Dataset | Model | Publisher & Year | mAP@0.5 | mAP@0.5:0.95 | P | R |
|---|---|---|---|---|---|---|
| DetFly | ATA-YOLOv8 [27] | Drones 2025 | 96.4 | - | 98.0 | 92.4 |
| EDGS-YOLOv8 [33] | Drones 2024 | 93.4 | - | 91.4 | 91.9 | |
| DRF-YOLO [35] | Appl. Syst. Innov. 2025 | 91.1 | 55.4 | 95.0 | 86.5 | |
| ALDNet [39] | Meas. Sci. Technol. 2025 | 98.3 | 67.3 | 98.3 | 97.9 | |
| VDTNet [28] | T-ITS 2024 | 94.8 | - | 92.1 | 94.9 | |
| AutoUAVFormer | - | 98.6 | 68.1 | 99.1 | 97.2 | |
| DUT Anti-UAV | DCR-YOLO [34] | Aerosp. Sci. Technol. 2025 | 88.8 | 58.6 | 95.0 | 81.3 |
| DRF-YOLO [35] | Appl. Syst. Innov. 2025 | 86.9 | 54.8 | 93.9 | 80.3 | |
| YOLOv7-GS [36] | Drones 2024 | 93.2 | - | 96.8 | 90.3 | |
| Lightweight YOLO11 [37] | Drones 2024 | 95.2 | 65.7 | 96.9 | 89.8 | |
| DACG-Net [40] | T-AES 2025 | 92.8 | 62.6 | 96.3 | 85.1 | |
| AutoUAVFormer | - | 95.5 | 67.0 | 98.4 | 90.4 |
4.5. Feature Visualization Analysis
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Alzahrani, B.; Oubbati, O.S.; Barnawi, A.; Atiquzzaman, M.; Alghazzawi, D. UAV assistance paradigm: State-of-the-art in applications and challenges. J. Netw. Comput. Appl. 2020, 166, 102706. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Y.; Rao, B.; Wang, W. UAV swarm intelligence: Recent advances and future trends. IEEE Access 2020, 8, 183856–183878. [Google Scholar] [CrossRef] [Scilit]
- Hu, C.; Wang, Y.; Wang, R.; Zhang, T.; Cai, J.; Liu, M. An improved radar detection and tracking method for small UAV under clutter environment. Sci. China Inf. Sci. 2019, 62, 29306. [Google Scholar] [CrossRef] [Scilit]
- Xie, Y.; Jiang, P.; Gu, Y.; Xiao, X. Dual-source detection and identification system based on UAV radio frequency signal. IEEE Trans. Instrum. Meas. 2021, 70, 2006215. [Google Scholar] [CrossRef] [Scilit]
- Fang, J.; Finn, A.; Wyber, R.; Brinkworth, R.S. Acoustic detection of unmanned aerial vehicles using biologically inspired vision processing. J. Acoust. Soc. Am. 2022, 151, 968–981. [Google Scholar] [CrossRef] [Scilit]
- Medaiyese, O.O.; Ezuma, M.; Lauf, A.P.; Guvenc, I. Wavelet transform analytics for RF-based UAV detection and identification system using machine learning. Pervasive Mob. Comput. 2022, 82, 101569. [Google Scholar] [CrossRef] [Scilit]
- Yan, X.; Fu, T.; Lin, H.; Xuan, F.; Huang, Y.; Cao, Y.; Hu, H.; Liu, P. UAV detection and tracking in urban environments using passive sensors: A survey. Appl. Sci. 2023, 13, 11320. [Google Scholar] [CrossRef] [Scilit]
- Chen, C.; Zheng, Z.; Xu, T.; Guo, S.; Feng, S.; Yao, W.; Lan, Y. Yolo-based UAV technology: A review of the research and its applications. Drones 2023, 7, 190. [Google Scholar] [CrossRef] [Scilit]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, 11–14 October 2016; pp. 21–37. [Google Scholar]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2980–2988. [Google Scholar]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 16965–16974. [Google Scholar]
- Chen, C.; Liu, M.-Y.; Tuzel, O.; Xiao, J. R-CNN for small object detection. In Proceedings of the Asian Conference on Computer Vision, Taipei, Taiwan, 20–24 November 2016; pp. 214–230. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. Adv. Neural Inf. Process. Syst. 2015, 28, 1137–1149. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2961–2969. [Google Scholar]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable DETR: Deformable transformers for end-to-end object detection. arXiv 2020, arXiv:2010.04159. [Google Scholar]
- Zoph, B.; Le, Q.V. Neural architecture search with reinforcement learning. arXiv 2016, arXiv:1611.01578. [Google Scholar]
- Real, E.; Aggarwal, A.; Huang, Y.; Le, Q.V. Regularized evolution for image classifier architecture search. In Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA, 27 January–1 February 2019; pp. 4780–4789. [Google Scholar]
- Liu, H.; Simonyan, K.; Yang, Y. Darts: Differentiable architecture search. arXiv 2018, arXiv:1806.09055. [Google Scholar]
- Chen, Y.; Yang, T.; Zhang, X.; Meng, G.; Xiao, X.; Sun, J. DetNAS: Backbone search for object detection. Adv. Neural Inf. Process. Syst. 2019, 32, 6642–6652. [Google Scholar]
- Stamoulis, D.; Ding, R.; Wang, D.; Lymberopoulos, D.; Priyantha, B.; Liu, J.; Marculescu, D. Single-path NAS: Designing hardware-efficient convnets in less than 4 hours. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Würzburg, Germany, 16–20 September 2019; pp. 481–497. [Google Scholar]
- Chen, M.; Peng, H.; Fu, J.; Ling, H. Autoformer: Searching transformers for visual recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Virtual, 11–17 October 2021; pp. 12270–12280. [Google Scholar]
- Ghiasi, G.; Lin, T.-Y.; Le, Q.V. NAS-FPN: Learning scalable feature pyramid architecture for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–17 June 2019; pp. 7036–7045. [Google Scholar]
- Tan, M.; Pang, R.; Le, Q.V. Efficientdet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 13–19 June 2020; pp. 10781–10790. [Google Scholar]
- Zheng, Y.; Chen, Z.; Lv, D.; Li, Z.; Lan, Z.; Zhao, S. Air-to-air visual detection of micro-UAVs: An experimental evaluation of deep learning. IEEE Robot. Autom. Lett. 2021, 6, 1020–1027. [Google Scholar] [CrossRef] [Scilit]
- Zhao, J.; Zhang, J.; Li, D.; Wang, D. Vision-based anti-UAV detection and tracking. IEEE Trans. Intell. Transp. Syst. 2022, 23, 25323–25334. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Su, Y.; Wang, J.; Wang, T.; Gao, Q. UAVSwarm dataset: An unmanned aerial vehicle swarm dataset for multiple object tracking. Remote Sens. 2022, 14, 2601. [Google Scholar] [CrossRef] [Scilit]
- Hao, H.; Peng, Y.; Ye, Z.; Han, B.; Zhang, X.; Tang, W.; Kang, W.; Li, Q. A high performance air-to-air unmanned aerial vehicle target detection model. Drones 2025, 9, 154. [Google Scholar] [CrossRef] [Scilit]
- Zhou, X.; Yang, G.; Chen, Y.; Li, L.; Chen, B.M. VDTNet: A high-performance visual network for detecting and tracking of intruding drones. IEEE Trans. Intell. Transp. Syst. 2024, 25, 9828–9839. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Wang, X.; Zhou, C.; Meng, W.; Shi, Z. Low in resolution, high in precision: UAV detection with super-resolution and motion information extraction. In Proceedings of the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; pp. 1–5. [Google Scholar]
- Luo, A.; Hu, K.; Jiang, K. Rethinking Dual-Stream Super-Resolution for Enhancing Remote Sensing Object Detection. In Proceedings of the ICASSP 2025—2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Hyderabad, India, 6–11 April 2025; pp. 1–5. [Google Scholar]
- Du, Y.; Wu, T.; Dai, Z.; Xie, H.; Hu, C.; Wei, S. F-yolov7: Fast and robust real-time UAV detection. Computing 2025, 107, 50. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Li, D.; Zhu, Y.; Tian, L.; Shan, Y. Dual super-resolution learning for semantic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 13–19 June 2020; pp. 3774–3783. [Google Scholar]
- Huang, M.; Mi, W.; Wang, Y. EDGS-YOLOv8: An improved YOLOv8 lightweight UAV detection model. Drones 2024, 8, 337. [Google Scholar] [CrossRef] [Scilit]
- Ding, S.; Zhang, M.; Liu, D.; Liang, J. DCR-YOLO: An enhanced anti-UAV detection method based on triple collaborative optimization strategy. Aerosp. Sci. Technol. 2025, 168, 111319. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Shuai, J.; Yang, Y.; Hu, X.; Yang, C.; Zhou, Y. DRF-YOLO Model for Small UAV Detection Through Multi-Scale Residual Enhancement and Progressive Feature Fusion. Appl. Syst. Innov. 2025, 8, 179. [Google Scholar] [CrossRef] [Scilit]
- Bo, C.; Wei, Y.; Wang, X.; Shi, Z.; Xiao, Y. Vision-based anti-UAV detection based on YOLOv7-GS in complex backgrounds. Drones 2024, 8, 331. [Google Scholar] [CrossRef] [Scilit]
- Gao, Y.; Xin, Y.; Yang, H.; Wang, Y. A lightweight anti-unmanned aerial vehicle detection method based on improved YOLOv11. Drones 2024, 9, 11. [Google Scholar] [CrossRef] [Scilit]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In Proceedings of the European Conference on Computer Vision, Virtual, 23–28 August 2020; pp. 213–229. [Google Scholar]
- Cai, H.; Zhang, J.; Xu, J. ALDNet: A lightweight and efficient drone detection network. Meas. Sci. Technol. 2025, 36, 025402. [Google Scholar] [CrossRef] [Scilit]
- Li, W.; Xiao, L.; Li, H.; Yao, S.; Wan, B.; Ren, D. DACG-Net: A dual-backbone and context-guided fusion network for aerial UAV detection. IEEE Trans. Aerosp. Electron. Syst. 2025, 61, 17109–17123. [Google Scholar] [CrossRef] [Scilit]
- Peng, Y.; Li, H.; Wu, P.; Zhang, Y.; Sun, X.; Wu, F. D-FINE: Redefine regression task in DETRs as fine-grained distribution refinement. arXiv 2024, arXiv:2410.13842. [Google Scholar]
- Li, Y.; Fan, Q.; Huang, H.; Han, Z.; Gu, Q. A modified YOLOv8 detection network for UAV aerial image recognition. Drones 2023, 7, 304. [Google Scholar] [CrossRef] [Scilit]
- Kong, Y.; Shang, X.; Jia, S. Drone-DETR: Efficient small object detection for remote sensing image using enhanced RT-DETR model. Sensors 2024, 24, 5496. [Google Scholar] [CrossRef] [Scilit]
- Cao, X.; Wang, H.; Wang, X.; Hu, B. DFS-DETR: Detailed-feature-sensitive detector for small object detection in aerial images using transformer. Electronics 2024, 13, 3404. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Jiang, H.; Li, Z.; Yang, J.; Ma, X.; Chen, J.; Tang, X. PHSI-RTDETR: A lightweight infrared small target detection algorithm based on UAV aerial photography. Drones 2024, 8, 240. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Q.; Sheng, K.; Zheng, X.; Li, K.; Sun, X.; Tian, Y.; Chen, J.; Ji, R. Training-free transformer architecture search. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 10894–10903. [Google Scholar]
- Lee, J.; Ham, B. AZ-NAS: Assembling zero-cost proxies for network architecture search. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 5893–5903. [Google Scholar]
- Wang, N.; Gao, Y.; Chen, H.; Wang, P.; Tian, Z.; Shen, C.; Zhang, Y. NAS-FCOS: Fast neural architecture search for object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 14–19 June 2020; pp. 11943–11951. [Google Scholar]
- Yu, X.; Lin, Z.; Wang, Y. SAR-NAS: Lightweight SAR Object Detection with Neural Architecture Search. In Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision (PRCV), Shanghai, China, 15–18 October 2025; pp. 232–244. [Google Scholar]
- Gupta, C.; Gill, N.S.; Gulia, P.; Kumar, A.; Karamti, H.; Moges, D.M. An optimized YOLO NAS based framework for realtime object detection. Sci. Rep. 2025, 15, 32903. [Google Scholar] [CrossRef] [Scilit]










| Search Space Configuration | ||||
|---|---|---|---|---|
| mlp_ratio | num_heads | depth | embed_dim | phase1_number |
| [1, 2, 3, 4, 5] | [1, 2, 4, 8, 16, 32] | [1, 2, 3, 4, 5, 6] | [256, 288, 320, 352, 384, 416] | [1, 2] |
| Model (ID) | NAS | MKCD FPN | SR | Search Results | mAP@0.5 | mAP@0.5:0.95 | R | P | Params | GFLOPS | FPS |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | - | 93.1 | 63.0 | 92.8 | 93.7 | 19.8 | 59.8 | 241 | |||
| 1 | √ | {8;3.0,4.0,4.0,1.0,1.0,1.0,2.0,4.0; 16,32,8,16,16,8,16,4;288;1} | 94.1 | 64.8 | 92.1 | 95.3 | 20.42 | 35.9 | 221 | ||
| 2 | √ | √ | {8;3.0,4.0,4.0,1.0,1.0,1.0,2.0,4.0; 16,32,8,16,16,8,16,4;288;1} | 94.2 | 65.6 | 93.3 | 96.5 | 21.87 | 36.8 | 188 | |
| 3 | √ | √ | √ | {5;5.0,4.0,4.0,3.0,5.0; 8,8,8,4,4;256;1} | 95.5 | 67.0 | 90.4 | 98.4 | 20.80 | 36.6 | 193 |
| Model (ID) | Kernel | Search Results | mAP@0.5 | mAP@0.5:0.95 | R | P | Params | GFLOPS | FPS | |||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 3 × 3 | [3;4.0,3.0,1.0;4,8,4;256;2] | 92.6 | 63.0 | 92.2 | 94.8 | 22.07 | 38.3 | 177 | |||
| 2 | 3 × 3 | 5 × 5 | [8;5.0,3.0,2.0,4.0,5.0,1.0,5.0,5.0; 16,8,4,4,8,4,8,4;256;1] | 94.0 | 65.6 | 93.1 | 96.8 | 21.40 | 37.2 | 181 | ||
| 3 | 3 × 3 | 5 × 5 | 7 × 7 | [5;5.0,4.0,4.0,3.0,5.0; 8,8,8,4,4;256;1] | 95.5 | 67.0 | 90.4 | 98.4 | 20.80 | 36.6 | 193 | |
| 4 | 3 × 3 | 5 × 5 | 7 × 7 | 9 × 9 | [5;5.0,4.0,4.0,3.0,5.0; 8,8,8,4,4;256;1] | 93.8 | 64.6 | 93.0 | 95.7 | 21.10 | 36.8 | 188 |
| Model (ID) | Aggregation Layer | Search Results | mAP@0.5 | mAP@0.5:0.95 | R | P | Params | GFLOPS | FPS |
|---|---|---|---|---|---|---|---|---|---|
| 1 | P3 | [6;1.0,3.0,2.0,4.0,1.0,1.0; 8,8,4,1,2,16;256;1] | 95.2 | 66.7 | 92.2 | 97.5 | 19.98 | 36.1 | 214 |
| 2 | P4 | [5;5.0,4.0,4.0,3.0,5.0; 8,8,8,4,4;256;1] | 95.5 | 67.0 | 90.4 | 98.4 | 20.80 | 36.6 | 193 |
| 3 | P5 | [8;4.0,4.0,3.0,2.0,1.0,1.0,1.0,1.0; 4,8,8,4,4,4,16,1;256;1] | 94.3 | 65.4 | 93.6 | 96.1 | 20.97 | 36.6 | 210 |
| Model (ID) | Loss Ratio (Detection:SR) | Search Results | mAP@0.5 | mAP@0.5:0.95 | R | P | Params | GFLOPS | FPS |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 1:1 | [8;5.0,5.0,3.0,3.0,2.0,1.0,2.0,5.0; 16,4,4,4,4,8,32,32;288;1] | 94.1 | 63.1 | 92.4 | 94.6 | 22.06 | 37.7 | 184 |
| 2 | 10:1 | [5;4.0,4.0,2.0,2.0,4.0; 4,4,4,8,16;288;1] | 94.3 | 65.6 | 93.5 | 96.8 | 21.26 | 37.1 | 196 |
| 3 | 5:1 | [8;3.0,4.0,4.0,1.0,4.0,2.0,1.0,4.0;8,8,4,16,4,2,4,4;288;1] | 94.1 | 65.1 | 92.8 | 96.0 | 22.04 | 37.5 | 188 |
| 4 | 7.5:1 | [5;5.0,4.0,4.0,3.0,5.0; 8,8,8,4,4;256;1] | 95.5 | 67.0 | 90.4 | 98.4 | 20.80 | 36.6 | 193 |
| Dataset | Model | Search Results | mAP@0.5 | mAP@0.5:0.95 | R | P | Params | GFLOPS | FPS |
|---|---|---|---|---|---|---|---|---|---|
| DetFly | AutoUAVFormer without SR | [4;4.0,1.0,3.0,5.0;8,8,4,8;320;1] | 97.7 | 67.2 | 97.5 | 98.5 | 20.05 | 36.1 | 212 |
| AutoUAVFormer | [8;3.0,5.0,3.0,3.0,1.0,1.0,5.0,5.0; 16,4,4,4,4,2,8,4;320;1] | 98.6 | 68.1 | 97.2 | 99.1 | 21.17 | 36.6 | 196 | |
| DUT Anti-UAV | AutoUAVFormer without SR | [8;1.0,4.0,4.0,4.0,5.0,5.0,2.0,4.0; 8,8,8,8,2,8,8,32;256;1] | 94.2 | 65.6 | 93.3 | 94.5 | 20.87 | 36.8 | 188 |
| AutoUAVFormer | [5;5.0,4.0,4.0,3.0,5.0; 8,8,8,4,4;256;1] | 95.5 | 67.0 | 90.4 | 98.4 | 20.80 | 36.6 | 193 | |
| UAV Swarm | AutoUAVFormer without SR | [2;5.0,1.0;4,1;352;2] | 87.7 | 43.9 | 87.2 | 89.6 | 21.00 | 36.7 | 180 |
| AutoUAVFormer | [6;2.0,2.0,2.0,1.0,1.0,2.0; 16,4,4,16,2,4;256;1] | 89.9 | 45.4 | 86.4 | 91.8 | 19.43 | 36.2 | 204 |
| Dataset | Model | Backbone | mAP@0.5 | mAP@0.5:0.95 | R | P | Params | GFLOPS | FPS |
|---|---|---|---|---|---|---|---|---|---|
| DetFly | YOLOv3 | Darknet53 | 87.8 | 58.9 | 74.6 | 88.4 | 61.5 | 154.7 | 100 |
| YOLOv5 | CSPNet | 94.2 | 63.6 | 76.7 | 91.7 | 25.0 | 64.0 | 249 | |
| YOLOv7 | EfficientRep | 94.1 | 63.4 | 76.3 | 92.1 | 37.0 | 104.8 | 142 | |
| YOLOv8 | CSPNet | 96.2 | 65.7 | 77.9 | 93.2 | 25.8 | 78.7 | 228 | |
| YOLOv9 | CSPNet | 96.4 | 66.2 | 78.5 | 93.5 | 20.0 | 76.5 | 198 | |
| YOLOv10 | CSPNet | 95.4 | 62.9 | 79.1 | 92.3 | 16.5 | 63.4 | 262 | |
| YOLOv11 | CSPNet | 96.8 | 65.5 | 79.9 | 95.4 | 20.1 | 68.2 | 221 | |
| RT-DETR | HGNetv2 | 95.3 | 64.0 | 83.7 | 94.8 | 34.3 | 108.3 | 174 | |
| RT-DETR | ResNet34 | 95.5 | 63.8 | 81.1 | 94.3 | 31.4 | 91.8 | 165 | |
| RT-DETR | ResNet50 | 96.8 | 65.9 | 85.2 | 95.6 | 43.0 | 134.8 | 116 | |
| D-FINE | HGNetv2 | 97.1 | 66.1 | 96.1 | 96.6 | 19.8 | 59.8 | 236 | |
| AutoUAVFormer | HGNetv2 | 98.6 | 68.1 | 97.2 | 99.1 | 21.2 | 36.6 | 196 | |
| DUT Anti-UAV | YOLOv3 | Darknet53 | 78.9 | 47.6 | 70.3 | 80.8 | 61.5 | 154.7 | 113 |
| YOLOv5 | CSPNet | 84.3 | 52.7 | 73.6 | 91.3 | 25.0 | 64.0 | 252 | |
| YOLOv7 | EfficientRep | 87.1 | 56.8 | 79.6 | 89.2 | 37.0 | 104.8 | 148 | |
| YOLOv8 | CSPNet | 86.4 | 55.3 | 76.7 | 93.5 | 25.8 | 78.7 | 232 | |
| YOLOv9 | CSPNet | 87.2 | 56.8 | 79.4 | 93.7 | 20.0 | 76.5 | 207 | |
| YOLOv10 | CSPNet | 88.6 | 57.7 | 81.9 | 93.4 | 16.5 | 63.4 | 265 | |
| YOLOv11 | CSPNet | 88.3 | 57.5 | 81.2 | 93.5 | 20.1 | 68.2 | 218 | |
| RT-DETR | HGNetv2 | 90.5 | 60.9 | 82.2 | 94.5 | 34.3 | 108.3 | 180 | |
| RT-DETR | ResNet34 | 88.7 | 57.6 | 81.7 | 93.5 | 31.4 | 91.8 | 166 | |
| RT-DETR | ResNet50 | 92.5 | 62.5 | 84.5 | 96.1 | 43.0 | 134.8 | 112 | |
| D-FINE | HGNetv2 | 93.1 | 63.0 | 92.8 | 93.7 | 19.8 | 59.8 | 241 | |
| AutoUAVFormer | HGNetv2 | 95.5 | 67.0 | 90.4 | 98.4 | 20.8 | 36.6 | 193 | |
| UAV Swarm | YOLOv3 | Darknet53 | 74.0 | 30.6 | 67.3 | 76.5 | 61.5 | 154.7 | 99 |
| YOLOv5 | CSPNet | 76.2 | 30.4 | 68.1 | 77.9 | 25.0 | 64.0 | 244 | |
| YOLOv7 | EfficientRep | 80.6 | 34.4 | 74.6 | 82.1 | 37.0 | 104.8 | 133 | |
| YOLOv8 | CSPNet | 81.8 | 36.3 | 77.3 | 87.7 | 25.8 | 78.7 | 220 | |
| YOLOv9 | CSPNet | 82.6 | 37.4 | 78.4 | 87.2 | 20.0 | 76.5 | 196 | |
| YOLOv10 | CSPNet | 82.1 | 37.8 | 77.9 | 88.6 | 16.5 | 63.4 | 252 | |
| YOLOv11 | CSPNet | 83.4 | 38.5 | 79.7 | 89.1 | 20.1 | 68.2 | 213 | |
| RT-DETR | HGNetv2 | 86.9 | 39.9 | 79.8 | 90.3 | 34.3 | 108.3 | 172 | |
| RT-DETR | ResNet34 | 83.7 | 39.1 | 79.1 | 88.2 | 31.4 | 91.8 | 150 | |
| RT-DETR | ResNet50 | 87.6 | 41.7 | 79.5 | 89.4 | 43.0 | 134.8 | 92 | |
| D-FINE | HGNetv2 | 88.5 | 43.3 | 86.3 | 91.0 | 19.8 | 59.8 | 230 | |
| AutoUAVFormer | HGNetv2 | 89.9 | 45.4 | 86.4 | 91.8 | 19.4 | 36.2 | 204 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Pan, L.; Wan, H.; Nurmamat, P.; Chen, J.; Sun, L.; Cao, Y.; Wang, S.; Li, Y.; Huang, Z. AutoUAVFormer: Neural Architecture Search with Implicit Super-Resolution for Real-Time UAV Aerial Object Detection. Remote Sens. 2026, 18, 1268. https://doi.org/10.3390/rs18091268
Pan L, Wan H, Nurmamat P, Chen J, Sun L, Cao Y, Wang S, Li Y, Huang Z. AutoUAVFormer: Neural Architecture Search with Implicit Super-Resolution for Real-Time UAV Aerial Object Detection. Remote Sensing. 2026; 18(9):1268. https://doi.org/10.3390/rs18091268
Chicago/Turabian StylePan, Li, Huiyao Wan, Pazlat Nurmamat, Jie Chen, Long Sun, Yice Cao, Shuai Wang, Yingsong Li, and Zhixiang Huang. 2026. "AutoUAVFormer: Neural Architecture Search with Implicit Super-Resolution for Real-Time UAV Aerial Object Detection" Remote Sensing 18, no. 9: 1268. https://doi.org/10.3390/rs18091268
APA StylePan, L., Wan, H., Nurmamat, P., Chen, J., Sun, L., Cao, Y., Wang, S., Li, Y., & Huang, Z. (2026). AutoUAVFormer: Neural Architecture Search with Implicit Super-Resolution for Real-Time UAV Aerial Object Detection. Remote Sensing, 18(9), 1268. https://doi.org/10.3390/rs18091268

