WaveUAV-YOLO: A Lightweight Architecture for UAV Object Detection with Frequency-Domain Edge Preservation and Stable Gradient Fusion
Highlights
- A lightweight architecture, WaveUAV-YOLO, is proposed to effectively preserve sub-pixel edges and stabilize gradient fusion for UAV object detection.
- By integrating the WHFD, C2f_AMSB, and FFM_Concat modules, the model’s detection accuracy for small targets is significantly improved against complex background interference.
- Extensive experiments on the VisDrone2019, DIOR, NWPU VHR-10, and HIT-UAV datasets confirm that the model achieves advanced overall performance and strong cross-modal generalization capabilities.
- WaveUAV-YOLO provides an effective, real-time solution for wide-area UAV surveillance, demonstrating great potential for practical applications such as 24/7 urban traffic monitoring and emergency rescue.
Abstract
1. Introduction
- Unresolved Limitation 1 (Loss of high-frequency edges): Recent baselines rely on stride-2 convolutions or standard pooling for spatial downsampling. These operations act as lossy low-pass filters, tending to smooth out the delicate, sub-pixel high-frequency edges of tiny targets. Our Solution: We propose the Wavelet High-Frequency Downsampling (WHFD) module, which introduces a parallel Haar wavelet gating path to compensate for the information loss in stride-2 downsampling, effectively preserving boundary fidelity without heavy computation.
- Unresolved Limitation 2 (Morphological distortion): Existing detectors typically utilize isotropic square convolutions and impose mandatory channel compression (e.g., expansion ratio ) within bottlenecks. This forced compression tends to distort the anisotropic shapes of elongated targets (e.g., vehicles) common in UAV views. Our Solution: We design the Asymmetric Multi-scale Bottleneck (C2f_AMSB), which removes channel dimensionality compression () and deploys asymmetric convolutions ( and ) to accurately model extreme aspect ratios.
- Unresolved Limitation 3 (Gradient suppression during fusion): During cross-level feature aggregation, current architectures mainly use standard channel concatenation. This direct fusion causes the strong activation gradients of deep backgrounds to mathematically suppress the weak gradients of shallow small targets during backpropagation. Our Solution: We introduce the Mean-Normalized Feature Aggregation (FFM_Concat) mechanism, a structural weighting scheme that maintains the magnitude of weak feature streams and defends their gradient space against background dominance.
- A Wavelet High-Frequency Downsampling (WHFD) module: WHFD uses Haar wavelet transform to extract high-frequency edge masks, preserving boundaries during downsampling with low computational cost.
- An Asymmetric Multi-scale Bottleneck without Dimensionality Compression (C2f_AMSB): C2f_AMSB removes channel compression and uses asymmetric convolution kernels to better represent elongated targets under UAV views.
- A Mean-Normalized Feature Aggregation mechanism (FFM_Concat): FFM_Concat uses a mean-constrained weighting scheme to prevent strong features from suppressing weak ones during cross-level concatenation.
2. Related Work
2.1. Object Detection in Remote Sensing and UAV Vision
2.2. Frequency-Domain Representation for Small Targets
2.3. Morphological Perception and Multi-Scale Feature Fusion
- Shape distortion under channel compression: Targets in UAV views (e.g., elongated vehicles, densely arranged ships) often exhibit significant anisotropy. Existing models typically rely on isotropic square convolution kernels and impose channel compression inside bottleneck layers (e.g., expansion ratio ). This forced compression distorts anisotropic shapes and can cause features of nearby targets to blend together.
- Gradient suppression during feature concatenation: When aggregating features from different levels, the high activation responses of deep backgrounds often coexist with the weak features of shallow sub-pixel targets. Standard channel concatenation [28] can cause the gradients of weak signals to be suppressed by strong backgrounds during backpropagation.
3. Proposed Method
3.1. Overall Architecture of the Proposed Model
- Backbone downsampling reconstruction: At the , , and stages, the standard spatial downsampling convolution is replaced by our proposed Wavelet High-Frequency Downsampling (WHFD), aiming to preserve the weak high-frequency edges of sub-pixel targets.
- Morphology-aware bottleneck: The default bottleneck in feature extraction is upgraded to C2f_AMSB. Specifically, its internal expansion ratio is set to to avoid shape distortion caused by channel compression and enhance the geometric fitting of anisotropic targets.
- Numerical-stable fusion nodes: At critical positions of the feature pyramid, the native Concat and FullPAD_Tunnel operations are replaced by the Mean-Normalized Feature Aggregation (FFM_Concat) and Channel-Spatial Fusion Module (CSFM_Fusion), protecting the gradient space of weak targets from both numerical stability and semantic noise suppression.
3.2. Frequency-Domain Reconstruction: Spatial-High-Frequency Dual-Track Downsampling (WHFD)
3.2.1. Mathematical Priors of Discrete Wavelet Transform
3.2.2. Progressive Derivation of the Dual-Track Fusion Mechanism
3.3. Morphology-Aware Reconstruction: Asymmetric Multi-Scale Bottleneck Without Dimensionality Compression
3.4. Alleviating Feature Suppression: Mean-Normalized Feature Aggregation (FFM_Concat)
3.4.1. Dynamic Weighting with Mean Constraint
3.4.2. Why Mean Constraint?
3.5. Cross-Semantic Noise Suppression: Channel-Spatial Fusion Module (CSFM_Fusion)
- Channel attention: Adaptive global average pooling (GAP) extracts statistics, which pass through an MLP to evaluate channel importance:
- Spatial attention: Unlike traditional CBAM, which generates spatial masks via channel pooling on the fused feature map, our module independently applies convolutions to project both the low-level and high-level features into a single-channel spatial map before summation:This dual-source design explicitly preserves fine-grained localization cues from shallow layers while leveraging semantic guidance from deep layers, preventing early blurring.
4. Experiments and Results
4.1. Experimental Setup and Implementation Details
4.2. Evaluation Metrics
4.3. Benchmark Comparison on VisDrone2019
4.4. Cross-Scenario Generalization Analysis
4.4.1. Comparison on DIOR Large-Scale Remote Sensing Dataset
4.4.2. Validation on NWPU VHR-10 and HIT-UAV
4.5. Ablation Study and Mechanism Coupling Analysis
4.5.1. Mechanism Coupling Analysis
4.5.2. Internal Mechanism Analysis
4.6. Visual Analysis
4.6.1. Feature De-Adhesion in Dense Scenarios
4.6.2. Real-World Performance in Extreme Environments
4.6.3. Direct Visual Evidence of Edge Preservation
4.6.4. Real-World Unconstrained Flight Testing
5. Discussion
5.1. Discussion on the Integration of Smooth Bounding Box Losses
5.2. Analysis of Limitations and Failure Cases in Extreme Scenarios
- False Negatives (Missed Detections) under Severe Shadowing: In the heavily occluded regions (e.g., targets concealed by thick tree canopies or building shadows), the model fails to recall certain sub-pixel vehicles. At a input resolution, the high-frequency edge gradients of these heavily shadowed targets decay to near-zero levels. Consequently, the WHFD module lacks sufficient signal strength to extract structural boundaries, reaching the physical perception limit of the sensor and the spatial resolution constraint.
- Bounding Box Redundancy (False Positives) in Dense Clusters: In regions of extreme density, we observe occasional instances of redundant or overlapping bounding boxes that survive the Non-Maximum Suppression (NMS) post-processing. This is an unintended side effect of our frequency-domain enhancement: the WHFD module is highly sensitive to high-frequency signals. In chaotic, dense environments, strong specular reflections (e.g., from windshields or complex lighting) can produce spurious, sharp geometric contours. The network occasionally misinterprets these high-frequency artifacts as independent vehicle boundaries.
6. Conclusions
- Superiority of WHFD: It preserves high-frequency edges better than standard downsampling, which benefits later morphological processing.
- Efficacy of FFM_Concat: It mathematically prevents weak features from being suppressed during fusion, which is more cost-effective than adding complex spatial modules.
- Interplay with smoothness losses: We observed that smoothness-based losses like NWD can inadvertently hurt performance when the network architecture already preserves sharp, high-frequency edges.
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Srivastava, S.; Narayan, S.; Mittal, S. A Survey of Deep Learning Techniques for Vehicle Detection from UAV Images. J. Syst. Archit. 2021, 117, 102152. [Google Scholar] [CrossRef] [Scilit]
- Manaswini, V.N.S.; Kamatchi, K.; Nigam, C.; Ali, S.S.; Niranjana, R.; Suman. Real-Time Object Detection in Drone Surveillance Using YOLOv5. In Proceedings of the 2025 3rd International Conference on IoT, Communication and Automation Technology (ICICAT); IEEE: Gorakhpur, India, 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Hasan, M.M.; Wang, Z.; Fan, H.; Fatima, K.; Hussain, M.A.I.; Shaha, R.; Habib, T.M.A. BDNet: A lightweight YOLOv12-based vehicle detection framework for smart urban traffic monitoring. Smart Cities 2026, 9, 33. [Google Scholar] [CrossRef] [Scilit]
- Bashir, M.H.; Ahmad, M.; Rizvi, D.R.; El-Latif, A.A.A. Efficient CNN-based disaster events classification using UAV-aided images for emergency response application. Neural Comput. Appl. 2024, 36, 10599–10612. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Li, Q.; Yuan, Y.; Du, Q.; Wang, Q. ABNet: Adaptive Balanced Network for Multiscale Object Detection in Remote Sensing Imagery. IEEE Trans. Geosci. Remote Sens. 2022, 60, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Hua, W.; Chen, Q. A survey of small object detection based on deep learning in aerial images. Artif. Intell. Rev. 2025, 58, 162. [Google Scholar] [CrossRef] [Scilit]
- Wang, Q.; Cang, M.; Chen, Y. ECP-YOLO: Integrating Edge-Aware Attention and Contextual Refinement for UAV Object Detection. Electronics 2026, 15, 2067. [Google Scholar] [CrossRef] [Scilit]
- Lei, M.; Li, S.; Wu, Y.; Hu, H.; Zhou, Y.; Zheng, X.; Ding, G.; Du, S.; Wu, Z.; Gao, Y. YOLOv13: Real-time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception. arXiv 2025, arXiv:2506.17733. [Google Scholar]
- Yan, H.; Kong, X.; Wang, J.; Tomiyama, H. ST-YOLO: An Enhanced Detector of Small Objects in Unmanned Aerial Vehicle Imagery. Drones 2025, 9, 338. [Google Scholar] [CrossRef] [Scilit]
- Xu, B.; Cai, D.; Sui, K.; Wang, Z.; Liu, C.; Pei, X. MBD-YOLO: An Improved Lightweight Multi-Scale Small-Object Detection Model for UAVs Based on YOLOv8. Appl. Sci. 2025, 15, 10877. [Google Scholar] [CrossRef] [Scilit]
- Liu, H.; Dong, H.; Shi, H.; Li, F. CCAI-YOLO: A High-Precision Synthetic Aperture Radar Ship Detection Model Based on YOLOv8n Algorithm. Remote Sens. 2026, 18, 145. [Google Scholar] [CrossRef] [Scilit]
- Deng, C.; Zhao, Z.; Xu, X.; Xia, Y.; Li, J.; Plaza, A. GSFANet: Global Spatial-Frequency Attention Network for Infrared Small Target Detection. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5007017. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.; Zhou, P.; Yao, X.; Yao, C.; Zhang, Y.; Han, J. Object detection in VHR optical remote sensing images via learning rotation-invariant HOG feature. In Proceedings of the 2016 4th International Workshop on Earth Observation and Remote Sensing Applications (EORSA), Guangzhou, China, 4–6 July 2016; pp. 433–436. [Google Scholar]
- Ren, Y.; Zhu, C.; Xiao, S. Small object detection in optical remote sensing images via modified faster R-CNN. Appl. Sci. 2018, 8, 813. [Google Scholar] [CrossRef] [Scilit]
- Ding, J.; Xue, N.; Xia, G.S.; Bai, X.; Yang, W.; Yang, M.Y.; Belongie, S.; Luo, J.; Datcu, M.; Pelillo, M.; et al. Object detection in aerial images: A large-scale benchmark and challenges. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 44, 7778–7796. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
- Wang, C.Y.; Liao, H.Y.M. YOLOv9: Learning Clean Filters and What to Learn Using Programmable Gradient Information. arXiv 2024, arXiv:2402.13616. [Google Scholar]
- Han, W.; Dong, S.; Zhang, Y. Multispectral Small Object Detection for UAV Remote Sensing: A Comprehensive Review. Remote Sens. 2025, 17, 249. [Google Scholar] [CrossRef] [Scilit]
- Yang, F.; Fan, H.; Chu, P.; Blasch, E.; Ling, H. Clustered object detection in aerial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 8311–8320. [Google Scholar]
- Ye, T.; Qin, W.; Li, Y.; Wang, S.; Zhang, J.; Zhao, Z. Dense and small object detection in UAV-vision based on a global-local feature enhanced network. IEEE Trans. Instrum. Meas. 2022, 71, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Jin, B.; Yin, F.; Cai, W.; Li, H.; Zhu, H.; Huang, W.; Wu, Q.; Chen, H.; Sun, Z. HWANet: A Haar Wavelet-based Attention Network for remote sensing object detection. PLoS ONE 2025, 20, e0330759. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, J.; Liu, N.; Sun, H.; Wang, Y. Freq-DETR: Frequency-aware transformer for real-time small object detection in unmanned aerial vehicle imagery. Expert Syst. Appl. 2025, 298, 129710. [Google Scholar]
- Yang, Y.; Yuan, G.; Li, J. SFFNet: A Wavelet-Based Spatial and Frequency Domain Fusion Network for Remote Sensing Segmentation. IEEE Trans. Geosci. Remote Sens. 2024, 62, 1–17. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Bao, W.; Yang, Y.; Wan, W.; Xiao, Q.; Zou, X. HAFNet: Hierarchical Attention Fusion Network for Infrared Small Target Detection. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5007316. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Wang, C.; Li, X.; Xia, C.; Xu, J. MLP-Net: Multilayer Perceptron Fusion Network for Infrared Small Target Detection. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5601313. [Google Scholar] [CrossRef] [Scilit]
- Tan, M.; Pang, R.; Le, Q.V. EfficientDet: Scalable and Efficient Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2020; pp. 10778–10787. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Hou, Q.; Zheng, Z.; Cheng, M.M.; Yang, J.; Li, X. Large Selective Kernel Network for Remote Sensing Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 16740–16750. [Google Scholar]
- Huang, G.; Liu, Z.; Weinberger, K.Q. Densely Connected Convolutional Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 2261–2269. [Google Scholar]
- Rahman, M.M.; Marculescu, R. MK-UNet: Multi-kernel Lightweight CNN for Medical Image Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Honolulu, HI, USA, 19–20 October 2025; pp. 979–988. [Google Scholar]
- Li, Q.; Shen, L.; Guo, S.; Lai, Z. WaveCNet: Wavelet Integrated CNNs to Suppress Aliasing Effect for Noise-Robust Image Classification. IEEE Trans. Image Process. 2021, 30, 7074–7089. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Williams, T.; Li, R.Y. Wavelet Pooling for Convolutional Neural Networks. In Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar]
- Ding, X.; Guo, Y.; Ding, G.; Han, J. ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution Blocks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1911–1920. [Google Scholar]
- Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adam, H. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv 2017, arXiv:1704.04861. [Google Scholar]
- Lin, T.Y.; Dollár, P.; Girshick, R.B.; He, K.; Hariharan, B.; Belongie, S.J. Feature Pyramid Networks for Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 936–944. [Google Scholar]
- Ba, J.L.; Kiros, J.R.; Hinton, G.E. Layer Normalization. arXiv 2016, arXiv:1607.06450. [Google Scholar]
- Zhang, B.; Sennrich, R. Root Mean Square Layer Normalization. Adv. Neural Inf. Process. Syst. 2019, 32, 12360–12371. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
- Du, D.; Zhu, P.; Wen, L.; Bian, X.; Ling, H.; Hu, Q.; Zheng, J.; Peng, T.; Wang, X.; Zhang, Y.; et al. VisDrone-DET2019: The Vision Meets Drone Object Detection in Image Challenge Results. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshop (ICCVW); IEEE: Piscataway, NJ, USA, 2019; pp. 213–226. [Google Scholar] [CrossRef] [Scilit]
- Li, K.; Wan, G.; Cheng, G.; Cao, L.; Han, J. DIOR: A Large-Scale Benchmark Dataset for Object Detection in Optical Remote Sensing Images. ISPRS J. Photogramm. Remote Sens. 2020, 164, 13–23. [Google Scholar]
- Liu, J.; Huang, B.; Lv, J.Y. YOLO-PKFF: Remote Sensing Object Detection Enhanced with Poly Kernel Inception and Attentional Cross-Level Feature Fusion. IEEE Geosci. Remote Sens. Lett. 2025, 22, 6010805. [Google Scholar] [CrossRef] [Scilit]
- Gu, Q.; Huang, H.; Han, Z.; Fan, Q.; Li, Y. GLFE-YOLOX: Global and Local Feature Enhanced YOLOX for Remote Sensing Images. IEEE Trans. Instrum. Meas. 2024, 73, 2516112. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.; Han, J.; Lu, X. Remote Sensing Image Scene Classification: Benchmark and State of the Art. Proc. IEEE 2017, 105, 1865–1883. [Google Scholar] [CrossRef] [Scilit]
- Suo, J.; Wang, T.; Zhang, X.; Chen, H.; Zhou, W.; Shi, W. HIT-UAV: A High-altitude Infrared Thermal Dataset for Unmanned Aerial Vehicle-based Object Detection. Sci. Data 2023, 10, 227. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, J.; Xu, C.; Yang, W.; Yu, L. A Normalized Gaussian Wasserstein Distance for Tiny Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 764–773. [Google Scholar]
- Xu, C.; Wang, J.; Yang, W.; Yu, H.; Yu, L.; Xia, G.S. Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark. ISPRS J. Photogramm. Remote Sens. 2022, 190, 79–93. [Google Scholar] [CrossRef] [Scilit]
- Wang, G.; Zhao, H.; Lyu, S.; Cheng, G.; Chang, Q.; Feng, W.; Zhao, Q.; Shi, Z. SWIN-TOD: Smooth wasserstein distance and instance-level neighboring enhancement for remote sensing tiny object detection. IEEE Trans. Geosci. Remote Sens. 2024, 62, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Zhang, F.; Zhou, S.; Wang, Y.; Wang, X.; Hou, Y. Label assignment matters: A gaussian assignment strategy for tiny object detection. IEEE Trans. Geosci. Remote Sens. 2024, 62, 1–12. [Google Scholar] [CrossRef] [Scilit]















| Model | Params (M) | GFLOPs | Precision (%) | Recall (%) | mAP50 (%) | mAP50–95 (%) |
|---|---|---|---|---|---|---|
| YOLOv8s [7] | 11.20 | 28.70 | 43.50 | 34.30 | 32.00 | 18.10 |
| YOLOv10s [7] | 8.00 | 24.80 | 43.50 | 33.60 | 31.60 | 18.00 |
| YOLOv11s [7] | 9.40 | 23.50 | 44.40 | 34.50 | 32.30 | 18.20 |
| YOLOv12s [7] | 9.23 | 21.20 | 45.10 | 33.40 | 31.80 | 18.50 |
| RT-DETR-R18 [10] | 19.90 | 57.00 | 47.90 | 32.40 | 32.70 | 18.30 |
| MBD-YOLO [10] | 12.10 | 24.30 | 51.80 | 31.30 | 33.20 | 21.90 |
| ST-YOLO [9] | 8.96 | 20.07 | 44.80 | 33.10 | 33.20 | 18.20 |
| YOLOv13s (Baseline) | 9.04 | 20.70 | 42.80 | 31.80 | 29.60 | 17.00 |
| WaveUAV-YOLO (Ours) | 11.70 | 30.60 | 46.70 | 35.20 | 33.50 | 19.30 |
| Detector | Pedestrian | People | Bicycle | Car | Van | Truck | Tricycle | Awning-Tricycle | Bus | Motor | mAP50 (%) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| YOLOv13s (Baseline) | 25.4 | 12.4 | 7.57 | 70.1 | 34.1 | 34.1 | 16.1 | 15.8 | 53.3 | 27.3 | 29.6 |
| WaveUAV-YOLO (Ours) | 29.3 | 15.6 | 9.63 | 73.7 | 40.2 | 40.4 | 18.5 | 19.3 | 57.3 | 31.0 | 33.5 |
| Absolute gain () | +3.9 | +3.2 | +2.06 | +3.6 | +6.1 | +6.3 | +2.4 | +3.5 | +4.0 | +3.7 | +3.9 |
| Model | Params (M) | GFLOPs | mAP50 (%) | mAP50–95 (%) |
|---|---|---|---|---|
| YOLO11 [40] | 9.40 | 21.5 | 77.70 | 56.50 |
| YOLO-PKFF [40] | – | – | 80.10 | – |
| GLFE-YOLOX [41] | 42.45 | 84.2 | 78.90 | – |
| YOLOv13s (Baseline) | 9.04 | 20.8 | 80.65 | 59.70 |
| WaveUAV-YOLO (Ours) | 11.70 | 30.6 | 81.37 | 61.51 |
| Dataset | Detector | Params (M) | GFLOPs | mAP50 (%) | mAP50–95 (%) |
|---|---|---|---|---|---|
| NWPU | YOLOv13s (Baseline) | 9.04 | 20.7 | 88.00 | 53.05 |
| WaveUAV-YOLO (Ours) | 11.70 | 30.6 | 90.46 | 55.36 | |
| HIT-UAV | YOLOv13s (Baseline) | 9.04 | 20.7 | 85.80 | 57.20 |
| WaveUAV-YOLO (Ours) | 11.70 | 30.6 | 86.62 | 57.60 |
| Exp. | Base | WHFD | AMSB | FFM | CSFM | mAP50 (%) | mAP50–95 (%) | Params (M) | GFLOPs |
|---|---|---|---|---|---|---|---|---|---|
| 1 | ✓ | - | - | - | - | 29.6 | 17.9 | 9.04 | 20.7 |
| 2 | ✓ | ✓ | - | - | - | 31.4 | 18.3 | 11.53 | 26.9 |
| 3 | ✓ | - | ✓ | - | - | 31.1 | 18.1 | 9.95 | 24.2 |
| 4 | ✓ | - | - | ✓ | - | 31.1 | 18.1 | 9.06 | 21.1 |
| 5 | ✓ | - | - | - | ✓ | 31.0 | 18.0 | 9.19 | 21.9 |
| 6 | ✓ | ✓ | ✓ | - | - | 32.8 | 18.9 | 11.57 | 30.1 |
| 7 | ✓ | ✓ | ✓ | ✓ | - | 33.2 | 19.1 | 11.59 | 30.1 |
| 8 | ✓ | ✓ | ✓ | ✓ | ✓ | 33.5 | 19.3 | 11.70 | 30.6 |
| Wavelet Basis | Filter Taps | Params (M) | GFLOPs | mAP50 (%) |
|---|---|---|---|---|
| Daubechies (db4) | 8 | 11.70 | 30.6 | 32.98 |
| Symlets (sym2) | 4 | 11.70 | 30.6 | 32.97 |
| Haar (db1) [Ours] | 2 | 11.70 | 30.6 | 33.50 |
| Base Spatial () | Anisotropic ( & ) | Params (M) | GFLOPs | mAP50 (%) |
|---|---|---|---|---|
| × | ✓ | 11.30 | 28.5 | 32.97 |
| ✓ | × | 11.45 | 29.2 | 33.08 |
| ✓ | ✓ | 11.70 | 30.6 | 33.50 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, Q.; Xu, S.; Chen, Y. WaveUAV-YOLO: A Lightweight Architecture for UAV Object Detection with Frequency-Domain Edge Preservation and Stable Gradient Fusion. Remote Sens. 2026, 18, 2404. https://doi.org/10.3390/rs18142404
Wang Q, Xu S, Chen Y. WaveUAV-YOLO: A Lightweight Architecture for UAV Object Detection with Frequency-Domain Edge Preservation and Stable Gradient Fusion. Remote Sensing. 2026; 18(14):2404. https://doi.org/10.3390/rs18142404
Chicago/Turabian StyleWang, Qi, Shengqi Xu, and Yongji Chen. 2026. "WaveUAV-YOLO: A Lightweight Architecture for UAV Object Detection with Frequency-Domain Edge Preservation and Stable Gradient Fusion" Remote Sensing 18, no. 14: 2404. https://doi.org/10.3390/rs18142404
APA StyleWang, Q., Xu, S., & Chen, Y. (2026). WaveUAV-YOLO: A Lightweight Architecture for UAV Object Detection with Frequency-Domain Edge Preservation and Stable Gradient Fusion. Remote Sensing, 18(14), 2404. https://doi.org/10.3390/rs18142404

