ADA-YOLO: An Adaptive Dynamic Aggregation Network for Small Object Detection in UAV Imagery
Abstract
1. Introduction
- High-resolution spatial representation via the P2 detection head: By incorporating a feature branch with a stride of 4 into the detection head, we explicitly compensate for the spatial information loss inherent in deep downsampling, significantly strengthening the localization capability for extremely small objects.
- Content-aware feature reconstruction via DySample: We replace traditional static interpolation with a lightweight dynamic upsampling mechanism. Within our multi-scale architecture, this adaptively enhances target contour sharpness without imposing additional computational burden.
- Cross-scale semantic alignment via ASFF: To resolve the semantic conflicts exacerbated by the addition of the P2 branch, we utilize pixel-level adaptive weighting for multi-scale feature fusion. This explicitly suppresses cross-scale background noise interference under complex UAV perspectives.
2. Related Work
2.1. Research on UAV Object Detection
2.2. Related Methods for Small Object Detection
3. Methodology
3.1. Overall Model Architecture
3.2. High-Resolution P2 Detection Head
3.3. DySample Dynamic Upsampling Mechanism
3.4. ASFF Adaptive Spatial Feature Fusion Strategy
- (1) Feature Scaling. For a target output layer , the corresponding ASFF module receives features from four scales and maps them uniformly to scale l:
- When the input feature resolution is higher than the target layer (), downsampling is performed using a stride-2 convolution or pooling operation;
- When the input feature resolution is lower than the target layer (), channel alignment is first conducted via a convolution, followed by DySample dynamic upsampling to restore the spatial resolution.
- (2) Adaptive Fusion. At each spatial location , ASFF generates weights corresponding to the four scales through a lightweight convolutional network:
4. Experiments
4.1. Datasets and Experimental Setup
- VisDrone2019 [20]: These dataset features highly diverse urban aerial perspectives and defines 10 object categories with distinct characteristics under UAV top-down viewpoints.
- AI-TOD [21]: A specialized dataset designed for tiny object detection in aerial imagery, where the average object size is only approximately 12.8 pixels. It is specifically constructed to address the challenges of extremely small object detection in remote sensing scenarios.
4.2. Experimental Results
4.2.1. Experimental Results on the VisDrone2019 Dataset
4.2.2. Experimental Results on the AI-TOD Dataset
4.3. Ablation Study
4.4. Comparison of Detection Results
5. Conclusions and Future Work
5.1. Conclusions
- Introduction of a high-resolution P2 detection head to construct a P2–P5 four-scale prediction structure, thereby enhancing the fine-grained spatial information representation for extremely small objects;
- Adoption of DySample dynamic upsampling to replace traditional static interpolation, improving feature reconstruction quality;
- The incorporation of the ASFF adaptive spatial feature fusion mechanism during the multi-scale integration phase to mitigate semantic discrepancies and suppress background noise interference.
5.2. Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Tang, G.; Ni, J.; Zhao, Y.; Gu, Y.; Cao, W. A survey of object detection for UAVs based on deep learning. Remote Sens. 2024, 16, 149. [Google Scholar] [CrossRef] [Scilit]
- Nikouei, M.; Baroutian, B.; Nabavi, S.; Taraghi, F.; Aghaei, A.; Sajedi, A.; Moghaddam, M.E. Small object detection: A comprehensive survey on challenges, techniques and real-world applications. Intell. Syst. Appl. 2025, 27, 200561. [Google Scholar] [CrossRef] [Scilit]
- Aldubaikhi, A.; Patel, S. Advancements in small-object detection (2023–2025): Approaches, datasets, benchmarks, applications, and practical guidance. Appl. Sci. 2025, 15, 11882. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2016; pp. 779–788. [Google Scholar] [CrossRef] [Scilit]
- Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLOv8. Version 8.0.180, License: AGPL-3.0. 2023. Available online: https://pypi.org/project/ultralytics/8.0.180/ (accessed on 2 June 2026).
- Xu, C.; Wang, J.; Yang, W.; Yu, H.; Yu, L.; Xia, G.S. Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark. ISPRS J. Photogramm. Remote Sens. 2022, 190, 79–93. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Hu, Y.; Liang, Q.; He, Y.; Zhou, L. An improved YOLOv8s-based UAV target detection algorithm. PLoS ONE 2025, 20, e0327732. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2017; pp. 2117–2125. [Google Scholar] [CrossRef] [Scilit]
- Du, Z.; Hu, Z.; Zhao, G.; Jin, Y.; Ma, H. Cross-layer feature pyramid transformer for small object detection in aerial images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5625714. [Google Scholar] [CrossRef] [Scilit]
- Lee, Y.W.; Kim, B.G. Attention-based scale sequence network for small object detection. Heliyon 2024, 10, e32931. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, W.; Shao, X.; Mei, C.; Pan, X.; Lu, X. Multiscale Adaptively Spatial Feature Fusion Network for Spacecraft Component Recognition. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 3501–3513. [Google Scholar] [CrossRef] [Scilit]
- Qu, S.; Dang, C.; Chen, W.; Liu, Y. Sma-yolo: An improved yolov8 algorithm based on parameter-free attention mechanism and multi-scale feature fusion for small object detection in uav images. Remote Sens. 2025, 17, 2421. [Google Scholar] [CrossRef] [Scilit]
- Yang, C.; Shen, Y.; Wang, L. EMFE-YOLO: A Lightweight Small Object Detection Model for UAVs. Sensors 2025, 25, 5200. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jia, Z.; Zhang, M.; Yuan, C.; Liu, Q.; Liu, H.; Qiu, X.; Zhao, W.; Shi, J. ADL-YOLOv8: A field crop weed detection model based on improved YOLOv8. Agronomy 2024, 14, 2355. [Google Scholar] [CrossRef] [Scilit]
- Li, T.; Chen, X.; Xie, X.; Liu, F.; Pan, Y. Small-object detection based on its-yolov5s-p2. In Proceedings of the 2024 5th International Conference on Computer Engineering and Intelligent Control (ICCEIC); IEEE: Piscataway, NJ, USA, 2024; pp. 17–20. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Chen, K.; Xu, R.; Liu, Z.; Loy, C.C.; Lin, D. Carafe: Content-aware reassembly of features. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2019; pp. 3007–3016. [Google Scholar] [CrossRef] [Scilit]
- Lu, H.; Liu, W.; Fu, H.; Cao, Z. FADE: Fusing the assets of decoder and encoder for task-agnostic upsampling. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2022; pp. 231–247. [Google Scholar] [CrossRef] [Scilit]
- Liu, W.; Lu, H.; Fu, H.; Cao, Z. Learning to upsample by learning to sample. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: Piscataway, NJ, USA, 2023; pp. 6027–6037. [Google Scholar] [CrossRef] [Scilit]
- Liu, S.; Huang, D.; Wang, Y. Learning spatial fusion for single-shot object detection. arXiv 2019, arXiv:1911.09516. [Google Scholar] [CrossRef] [Scilit]
- Du, D.; Zhu, P.; Wen, L.; Bian, X.; Lin, H.; Hu, Q.; Peng, T.; Zheng, J.; Wang, X.; Zhang, Y.; et al. VisDrone-DET2019: The vision meets drone object detection in image challenge results. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops; IEEE: Piscataway, NJ, USA, 2019; pp. 213–226. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Yang, W.; Guo, H.; Zhang, R.; Xia, G.S. Tiny object detection in aerial images. In Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR); IEEE: Piscataway, NJ, USA, 2021; pp. 3791–3798. [Google Scholar] [CrossRef] [Scilit]
- Jocher, G.; Qiu, J.; Liu, M.; Lyu, S.; Akyon, F.C.; Kalfaoglu, M.E. Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models. arXiv 2026, arXiv:2606.03748. [Google Scholar] [CrossRef] [Scilit]
- Liu, H.; Sun, F.; Gu, J.; Deng, L. Sf-yolov5: A lightweight small object detection algorithm based on improved feature fusion mode. Sensors 2022, 22, 5817. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tang, S.; Zhang, S.; Fang, Y. HIC-YOLOv5: Improved YOLOv5 for small object detection. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2024; pp. 6614–6619. [Google Scholar] [CrossRef] [Scilit]
- He, Z.; She, R.; Tan, B.; Li, J.; Lei, X. SSCW-YOLO: A Lightweight and High-Precision Model for Small Object Detection in UAV Scenarios. Drones 2026, 10, 41. [Google Scholar] [CrossRef] [Scilit]
- Zhang, G.; Peng, Y.; Li, J. YOLO-MARS: An enhanced YOLOV8N for small object detection in UAV aerial imagery. Sensors 2025, 25, 2534. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Luo, J.; Chang, K.; Huang, J.; Sun, X.; Ji, Y. A UAV aerial image small object detection algorithm based on fine-grained feature preservation and multi-scale feature pyramid balancing. Complex Intell. Syst. 2026, 12, 12. [Google Scholar] [CrossRef] [Scilit]
- Su, J.; Qin, Y.; Jia, Z.; Liang, B. MPE-YOLO: Enhanced small target detection in aerial imaging. Sci. Rep. 2024, 14, 17799. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xiao, Y.; Xu, T.; Xin, Y.; Li, J. Fbrt-yolo: Faster and better for real-time aerial image detection. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Washington, DC, USA, 2025; Volume 39, pp. 8673–8681. [Google Scholar] [CrossRef] [Scilit]










| Model | Param (M) | FLOPs (G) | FPS | ||
|---|---|---|---|---|---|
| YOLOv8n | 3.01 | 8.2 | 147.24 | 0.292 | 0.167 |
| YOLO26n | 2.51 | 5.8 | 66.71 | 0.283 | 0.164 |
| SF-YOLOv5 [23] | 2.24 | 13.8 | 86.96 | 0.343 † | 0.182 † |
| HIC-YOLOv5 [24] | 8.39 | – | – | 0.369 † | 0.208 † |
| SSCW-YOLO [25] | 2.73 | 8.7 | 115.10 | 0.378 † | 0.223 † |
| YOLO-MARS [26] | 2.93 | – | – | 0.409 † | 0.234 † |
| EMFE-YOLO [13] | 3.00 | 33.1 | 121.0 | 0.376 † | – |
| ADA-YOLO (Ours) | 3.17 | 13.2 | 100.23 | 0.405 | 0.249 |
| Model | Param (M) | FLOPs (G) | FPS | ||
|---|---|---|---|---|---|
| YOLOv8n | 3.01 | 8.2 | 147.24 | 0.410 | 0.176 |
| FMFN-YOLO [27] | 8.6 | 47.1 | 158.73 | 0.436 | 0.201 |
| MPE-YOLO [28] | 8.7 | 4.4 | 71 | 0.471 | 0.218 |
| FBRT-YOLO-S [29] | 2.9 | 22.9 | 142 | 0.458 | 0.202 |
| ADA-YOLO (Ours) | 3.17 | 13.2 | 100.23 | 0.444 | 0.221 |
| Model | A | B | C | Param (M) | FLOPs (G) | ||
|---|---|---|---|---|---|---|---|
| Baseline | × | × | × | 3.01 | 8.2 | 0.292 | 0.167 |
| + A | ✓ | × | × | 3.05 | 12.4 | 0.345 | 0.198 |
| + B | × | ✓ | × | 3.03 | 8.7 | 0.310 | 0.183 |
| + C | × | × | ✓ | 3.10 | 8.4 | 0.325 | 0.191 |
| + A + B | ✓ | ✓ | × | 3.08 | 13.1 | 0.368 | 0.216 |
| + A + C | ✓ | × | ✓ | 3.14 | 12.7 | 0.381 | 0.228 |
| + A + B + C | ✓ | ✓ | ✓ | 3.17 | 13.2 | 0.405 | 0.249 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Chen, J.; Jiang, S.; Li, Y.; Tuersunayi, S.; Liu, Y. ADA-YOLO: An Adaptive Dynamic Aggregation Network for Small Object Detection in UAV Imagery. Sensors 2026, 26, 3908. https://doi.org/10.3390/s26123908
Chen J, Jiang S, Li Y, Tuersunayi S, Liu Y. ADA-YOLO: An Adaptive Dynamic Aggregation Network for Small Object Detection in UAV Imagery. Sensors. 2026; 26(12):3908. https://doi.org/10.3390/s26123908
Chicago/Turabian StyleChen, Jiajun, Shaochen Jiang, Yongming Li, Sulaiman Tuersunayi, and Yong Liu. 2026. "ADA-YOLO: An Adaptive Dynamic Aggregation Network for Small Object Detection in UAV Imagery" Sensors 26, no. 12: 3908. https://doi.org/10.3390/s26123908
APA StyleChen, J., Jiang, S., Li, Y., Tuersunayi, S., & Liu, Y. (2026). ADA-YOLO: An Adaptive Dynamic Aggregation Network for Small Object Detection in UAV Imagery. Sensors, 26(12), 3908. https://doi.org/10.3390/s26123908

