Rethinking Adaptive Contextual Information and Multi-Scale Feature Fusion for Small-Object Detection in UAV Imagery
Abstract
1. Introduction
- (1)
- A DSAF module that enhances contextual information representation by effectively integrating detailed features with semantic information, thereby improving the feature representation of small objects.
- (2)
- An MSRSA mechanism that overcomes the limitations of traditional spatial attention in multi-scale modeling and adaptability, effectively mitigating interference from complex backgrounds and scale variations.
- (3)
- An FPRFN that incorporates a dedicated detection head to enhance tiny object perception, along with a multi-level feature reuse strategy and cross-scale fusion of high-level semantic information, significantly improving performance in small object detection tasks.
2. Related Work
2.1. Contextual Information Modeling
2.2. Multi-Scale Feature Fusion
3. Proposed Method
3.1. Overall Architecture
3.2. DSAF Module
3.3. MSRSA Mechanism
- (a)
- Pooling concatenation: Combine average and max pooling outputs.
- (b)
- Multi-scale convolution: Apply parallel convolutions to .
- (c)
- Attention fusion: Generate attention weights via adaptive average pooling (AAP) and Sigmoid activation.
3.4. FPRFN
4. Experiments and Results Analysis
4.1. Dataset
4.2. Experimental Setup and Evaluation Metrics
4.3. Comparison and Analysis of Key Improvements
4.3.1. Comparison Experiments of Different Attention Mechanisms
4.3.2. Comparative Experiments on Shallow Feature Fusion Methods
4.4. Ablation Experiments
4.5. Visual Analysis
4.6. Comparative Study
4.6.1. Performance Comparison on the VisDrone2019 Dataset
4.6.2. Performance Comparison on the WAID Dataset
4.6.3. Performance Comparison on the AI-TOD Dataset
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Muchiri, G.; Kimathi, S. A Review of Applications and Potential Applications of UAV. In Proceedings of the Sustainable Research and Innovation Conference, Pretoria, South Africa, 20–24 June 2022; pp. 280–283. [Google Scholar]
- Tang, G.; Ni, J.; Zhao, Y.; Gu, Y.; Cao, W. A Survey of Object Detection for UAVs Based on Deep Learning. Remote Sens. 2023, 16, 149. [Google Scholar] [CrossRef]
- Dalal, N.; Triggs, B. Histograms of Oriented Gradients for Human Detection. In Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), San Diego, CA, USA, 20–25 June 2005; Volume 1, pp. 886–893. [Google Scholar]
- Lowe, D.G. Object Recognition from Local Scale-Invariant Features. In Proceedings of the Seventh IEEE International Conference on Computer Vision, Corfu, Greece, 20–27 September 1999; Volume 2, pp. 1150–1157. [Google Scholar]
- Hearst, M.A.; Dumais, S.T.; Osuna, E.; Platt, J.; Scholkopf, B. Support Vector Machines. IEEE Intell. Syst. Their Appl. 1998, 13, 18–28. [Google Scholar] [CrossRef]
- Ali, M.L.; Zhang, Z. The YOLO Framework: A Comprehensive Review of Evolution, Applications, and Benchmarks in Object Detection. Computers 2024, 13, 336. [Google Scholar] [CrossRef]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2980–2988. [Google Scholar]
- Liu, C.; Gao, G.; Huang, Z.; Hu, Z.; Liu, Q.; Wang, Y. Yolc: You Only Look Clusters for Tiny Object Detection in Aerial Images. IEEE Trans. Intell. Transp. Syst. 2024, 25, 13863–13875. [Google Scholar] [CrossRef]
- Song, W.; Zhou, X.; Zhang, S.; Wu, Y.; Zhang, P. Glf-Net: A Semantic Segmentation Model Fusing Global and Local Features for High-Resolution Remote Sensing Images. Remote Sens. 2023, 15, 4649. [Google Scholar] [CrossRef]
- Lin, H.; Zhou, J.; Gan, Y.; Vong, C.-M.; Liu, Q. Novel Up-Scale Feature Aggregation for Object Detection in Aerial Images. Neurocomputing 2020, 411, 364–374. [Google Scholar] [CrossRef]
- Ye, T.; Qin, W.; Zhao, Z.; Gao, X.; Deng, X.; Ouyang, Y. Real-Time Object Detection Network in UAV-Vision Based on CNN and Transformer. IEEE Trans. Instrum. Meas. 2023, 72, 2505713. [Google Scholar] [CrossRef]
- Li, J.; Wei, Y.; Liang, X.; Dong, J.; Xu, T.; Feng, J.; Yan, S. Attentive Contexts for Object Detection. IEEE Trans. Multimed. 2016, 19, 944–954. [Google Scholar] [CrossRef]
- Yu, F.; Koltun, V. Multi-Scale Context Aggregation by Dilated Convolutions. arXiv 2015, arXiv:1511.07122. [Google Scholar]
- Wang, X.; Girshick, R.; Gupta, A.; He, K. Non-Local Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 7794–7803. [Google Scholar]
- Liu, Y.; Shao, Z.; Hoffmann, N. Global Attention Mechanism: Retain Information to Enhance Channel-Spatial Interactions. arXiv 2021, arXiv:2112.05561. [Google Scholar] [CrossRef]
- Scarselli, F.; Gori, M.; Tsoi, A.C.; Hagenbuchner, M.; Monfardini, G. The Graph Neural Network Model. IEEE Trans. Neural Networks 2008, 20, 61–80. [Google Scholar] [CrossRef]
- Liang, X.; Zhang, J.; Zhuo, L.; Li, Y.; Tian, Q. Small Object Detection in Unmanned Aerial Vehicle Images Using Feature Fusion and Scaling-Based Single Shot Detector with Spatial Context Analysis. IEEE Trans. Circuits Syst. Video Technol. 2019, 30, 1758–1770. [Google Scholar] [CrossRef]
- Han, W.; Li, J.; Wang, S.; Wang, Y.; Yan, J.; Fan, R.; Zhang, X.; Wang, L. A Context-Scale-Aware Detector and a New Benchmark for Remote Sensing Small Weak Object Detection in Unmanned Aerial Vehicle Images. Int. J. Appl. Earth Obs. Geoinf. 2022, 112, 102966. [Google Scholar] [CrossRef]
- Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature Pyramid Networks for Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2117–2125. [Google Scholar]
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path Aggregation Network for Instance Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 8759–8768. [Google Scholar]
- Zhang, W.; Liu, C.; Chang, F.; Song, Y. Multi-Scale and Occlusion Aware Network for Vehicle Detection and Segmentation on UAV Aerial Images. Remote Sens. 2020, 12, 1760. [Google Scholar] [CrossRef]
- Qiu, M.; Huang, L.; Tang, B.-H. Asff-Yolov5: Multielement Detection Method for Road Traffic in UAV Images Based on Multiscale Feature Fusion. Remote Sens. 2022, 14, 3498. [Google Scholar] [CrossRef]
- Zunair, H.; Khan, S.; Hamza, A.B. RSUD20K: A Dataset for Road Scene Understanding in Autonomous Driving. In Proceedings of the IEEE International Conference on Image Processing, Abu Dhabi, United Arab Emirates, 27–30 October 2024; pp. 708–714. [Google Scholar]
- Sunkara, R.; Luo, T. No More Strided Convolutions or Pooling: A New CNN Building Block for Low-Resolution Images and Small Objects. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases, Grenoble, France, 19–23 September 2022; pp. 443–459. [Google Scholar]
- Wang, Q.; Wu, B.; Zhu, P.; Li, P.; Zuo, W.; Hu, Q. Eca-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 11534–11542. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.-Y.; Kweon, I.S. Cbam: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
- Roy, A.G.; Navab, N.; Wachinger, C. Concurrent Spatial and Channel ’Squeeze & Excitation’ in Fully Convolutional Networks. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Granada, Spain, 16–20 September 2018; pp. 421–429. [Google Scholar]
- Du, D.; Zhu, P.; Wen, L.; Bian, X.; Lin, H.; Hu, Q.; Peng, T.; Zheng, J.; Wang, X.; Zhang, Y.; et al. Visdrone-Det2019: The Vision Meets Drone Object Detection in Image Challenge Results. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, Seoul, Republic of Korea, 27–28 October 2019. [Google Scholar]
- Mou, C.; Liu, T.; Zhu, C.; Cui, X. Waid: A Large-Scale Dataset for Wildlife Detection with Drones. Appl. Sci. 2023, 13, 10397. [Google Scholar] [CrossRef]
- Wang, J.; Yang, W.; Guo, H.; Zhang, R.; Xia, G.-S. Tiny Object Detection in Aerial Images. In Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, 10–15 January 2021; pp. 3791–3798. [Google Scholar]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 7132–7141. [Google Scholar]
- Narayanan, M. Senetv2: Aggregated Dense Layer for Channelwise and Global Representations. arXiv 2023, arXiv:2311.10807. [Google Scholar] [CrossRef]
- Yang, L.; Zhang, R.; Li, L.S.; Simam, X. A Simple, Parameter-Free Attention Module for Convolutional Neural Networks. In Proceedings of the International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; pp. 21–24. [Google Scholar]
- Hou, Q.; Zhou, D.; Feng, J. Coordinate Attention for Efficient Mobile Network Design. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 13713–13722. [Google Scholar]
- Ouyang, D.; He, S.; Zhang, G.; Luo, M.; Guo, H.; Zhan, J.; Huang, Z. Efficient Multi-Scale Attention Module with Cross-Spatial Learning. In Proceedings of the ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; pp. 1–5. [Google Scholar]
- Li, X.; Wang, W.; Hu, X.; Yang, J. Selective Kernel Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 510–519. [Google Scholar]
- Redmon, J.; Farhadi, A. Yolov3: An Incremental Improvement. arXiv 2018, arXiv:1804.02767. [Google Scholar] [CrossRef]
- Jocher, G.; Ultralytics. Yolov5 by Ultralytics. 2020. [Online]. Available online: https://github.com/ultralytics/yolov5 (accessed on 25 August 2025).
- Li, C.; Li, L.; Jiang, H.; Weng, K.; Geng, Y.; Li, L.; Ke, Z.; Li, Q.; Cheng, M.; Nie, W.; et al. Yolov6: A Single-Stage Object Detection Framework for Industrial Applications. arXiv 2022, arXiv:2209.02976. [Google Scholar] [CrossRef]
- Wang, C.-Y.; Bochkovskiy, A.; Liao, H.-Y.M. Yolov7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 7464–7475. [Google Scholar]
- Jocher, G.; Ultralytics. Yolov8. 2023. [Online]. Available online: https://github.com/ultralytics/ultralytics (accessed on 25 August 2025).
- Wang, C.-Y.; Yeh, I.-H.; Liao, H.-Y.M. Yolov9: Learning What You Want to Learn Using Programmable Gradient Information. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; pp. 1–21. [Google Scholar]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J. Yolov10: Real-Time End-to-End Object Detection. Adv. Neural Inf. Process. Syst. 2024, 37, 107984–108011. [Google Scholar]
- Khanam, R.; Hussain, M. Yolov11: An Overview of the Key Architectural Enhancements. arXiv 2024, arXiv:2410.17725. [Google Scholar] [CrossRef]
- Tian, Y.; Ye, Q.; Doermann, D. Yolov12: Attention-Centric Real-Time Object Detectors. arXiv 2025, arXiv:2502.12524. [Google Scholar]
- Chen, Z.-H.; Luo, A.-W.; Ding, L.; Zheng, J.-L.; Huang, Z.-K. A Robust and Efficient Multi-Scenario Object Detection Network for Edge Devices. IEEE Geosci. Remote Sens. Lett. 2025, 22, 6003205. [Google Scholar]
- Lan, Z.; Zhuang, F.; Lin, Z.; Chen, R.; Wei, L.; Lai, T.; Yang, C. MFO-Net: A Multiscale Feature Optimization Network for UAV Image Object Detection. IEEE Geosci. Remote Sens. Lett. 2024, 21, 6006605. [Google Scholar] [CrossRef]
- Jiang, L.; Yuan, B.; Du, J.; Chen, B.; Xie, H.; Tian, J.; Yuan, Z. MFFSODNet: A Multiscale Feature Fusion Small Object Detection Network for UAV Aerial Images. IEEE Trans. Instrum. Meas. 2024, 73, 5015214. [Google Scholar] [CrossRef]














| Components | Layers | From | N | Params | Module | Arguments |
|---|---|---|---|---|---|---|
| Backbone | 0 | −1 | 1 | 12,496 | DSAF | [3, 32] |
| 1 | −1 | 1 | 704 | GCBS | [32, 64, 3, 2] | |
| 2 | −1 | 1 | 29,056 | C2F | [64, 64, 1] | |
| 3 | −1 | 1 | 1408 | GCBS | [64, 128, 3, 2] | |
| 4 | −1 | 1 | 115,456 | C2F | [128, 128, 1] | |
| 5 | −1 | 1 | 2816 | GCBS | [128, 256, 3, 2] | |
| 6 | −1 | 1 | 460,288 | C2F | [256, 256, 1] | |
| 7 | −1 | 1 | 2816 | GCBS | [256, 256, 3, 2] | |
| 8 | −1 | 1 | 460,288 | C2F | [256, 256, 1] | |
| 9 | −1 | 1 | 164,608 | SPPF | [256, 256, 5] | |
| 10 | −1 | 1 | 143,032 | MSRSA | [256] | |
| Neck | 11 | −1 | 1 | 0 | Upsample | [None, 2, ‘nearest’] |
| 12 | [−1, 6] | 1 | 0 | Concat | [1] | |
| 13 | −1 | 1 | 525,824 | C2F | [512, 256, 1] | |
| 14 | −1 | 1 | 0 | Upsample | [None, 2, ‘nearest’] | |
| 15 | [−1, 4] | 1 | 0 | Concat | [1] | |
| 16 | −1 | 1 | 148,224 | C2F | [384, 128, 1] | |
| 17 | −1 | 1 | 0 | Upsample | [None, 2, ‘nearest’] | |
| 18 | [−1, 2] | 1 | 0 | Concat | [1] | |
| 19 | −1 | 1 | 37,248 | C2F | [192, 64, 1] | |
| 20 | −1 | 1 | 36,992 | CBS | [64, 64, 3, 2] | |
| 21 | [−1, 4, 16] | 1 | 0 | Concat | [1] | |
| 22 | −1 | 1 | 140,032 | C2F | [320, 128, 1] | |
| 23 | −1 | 1 | 147,712 | CBS | [128, 128, 3, 2] | |
| 24 | [−1, 6, 13] | 1 | 0 | Concat | [1] | |
| 25 | −1 | 1 | 558,592 | C2F | [640, 256, 1] | |
| Head | 26 | −1 | 1 | 753,262 | Detect | [10, [64, 128, 256]] |
| summary: 156 layers, 3,740,854 parameters, 3,740,838 gradients, 29.6 GFLOPs | ||||||
| Types | Label | Number | |||
|---|---|---|---|---|---|
| Training Set | Validation Set | Testing Set | Total | ||
| Pedestrian | 0 | 79,336 | 8843 | 21,005 | 109,184 |
| People | 1 | 27,058 | 5124 | 6375 | 38,557 |
| Bicycle | 2 | 10,479 | 1286 | 1301 | 13,066 |
| Car | 3 | 144,866 | 14,063 | 28,073 | 187,002 |
| Van | 4 | 24,955 | 1974 | 5770 | 32,699 |
| Truck | 5 | 12,874 | 749 | 2658 | 16,281 |
| Tricycle | 6 | 4811 | 1044 | 529 | 6384 |
| Awning-tricycle | 7 | 3245 | 531 | 598 | 4374 |
| Bus | 8 | 5952 | 250 | 2939 | 9114 |
| Motor | 9 | 29,646 | 4885 | 5844 | 40,375 |
| Types | Label | Number | |||
|---|---|---|---|---|---|
| Training Set | Validation Set | Testing Set | Total | ||
| Sheep | 0 | 91,495 | 26,062 | 13,322 | 130,879 |
| Cattle | 1 | 44,244 | 12,604 | 6239 | 63,087 |
| Seal | 2 | 15,761 | 4925 | 2688 | 23,374 |
| Camelus | 3 | 4675 | 1270 | 675 | 6620 |
| Kiang | 4 | 3311 | 836 | 459 | 4606 |
| Zebra | 5 | 3791 | 1000 | 431 | 5222 |
| Types | Label | Number | |||
|---|---|---|---|---|---|
| Training Set | Validation Set | Testing Set | Total | ||
| Airplane | 0 | 622 | 169 | 744 | 1535 |
| Bridge | 1 | 511 | 139 | 688 | 1338 |
| Storage-tank | 2 | 5268 | 2476 | 5859 | 13,603 |
| Ship | 3 | 13,538 | 3790 | 17,632 | 34,960 |
| Swimming-pool | 4 | 292 | 33 | 291 | 616 |
| Vehicle | 5 | 248,041 | 59,903 | 306,664 | 614,608 |
| Person | 6 | 14,125 | 3840 | 15,442 | 33,407 |
| Windmill | 7 | 175 | 66 | 289 | 530 |
| Environment Item | Configuration Parameters |
|---|---|
| Operating System | CentOS 8.5.2 |
| GPU NVIDIA GeForce RTX 5090 | 32 G |
| CUDA | 12.0 |
| CPU | Intel(R) Xeon(R) Platinum 8336C |
| Memory | 314G |
| Deep Learning Framework | PyTorch |
| Programming Language | Python3.9 |
| Hyperparameter | Value |
|---|---|
| Image size | 640 × 640 |
| Batch size | 64 |
| Epoch | 200 |
| Learning rate | 0.01 |
| Optimizer | SGD |
| Momentum | 0.937 |
| Weight decay | 0.0005 |
| Model | P (%) | R (%) | mAP@0.5 (%) | mAP@0.5:0.95 (%) | Params () | GFLOPs |
|---|---|---|---|---|---|---|
| YOLOv8s | 43.3 | 33.2 | 31.4 | 18.1 | 11.1 | 28.5 |
| +SEv1 [31] | 45.3 | 32.7 | 31.3 | 18.1 | 11.2 | 28.5 |
| +SEv2 [32] | 42.6 | 33.0 | 31.2 | 18.0 | 11.2 | 28.5 |
| +ECA [25] | 44.0 | 32.5 | 31.0 | 17.9 | 11.1 | 28.5 |
| +SimAM [33] | 43.3 | 32.7 | 31.1 | 18.0 | 11.1 | 28.5 |
| +CBAM [26] | 44.4 | 32.7 | 31.5 | 18.1 | 11.4 | 28.7 |
| +CA [34] | 44.9 | 32.9 | 31.5 | 18.2 | 11.2 | 28.5 |
| +GAM [16] | 42.7 | 33.1 | 31.3 | 18.1 | 17.7 | 33.7 |
| +EMA [35] | 43.1 | 33.0 | 31.0 | 17.9 | 11.1 | 28.5 |
| +SKA [36] | 44.3 | 32.6 | 31.2 | 18.1 | 33.2 | 46.1 |
| +MSRSA (ours) | 44.1 | 33.5 | 31.7 | 18.3 | 11.7 | 28.9 |
| Model | P (%) | R (%) | mAP@0.5 (%) | mAP@0.5:0.95 (%) | Params () | GFLOPs |
|---|---|---|---|---|---|---|
| PAN- P3+P4+P5 | 43.3 | 33.2 | 31.4 | 18.1 | 11.1 | 28.5 |
| PAN- P2+P3+P4 | 45.4 | 36.2 | 34.5 | 19.9 | 7.4 | 34.1 |
| PAN- P2+P3+P4+P5 | 45.6 | 35.8 | 34.2 | 19.8 | 10.6 | 36.7 |
| FPRFN- P2+P3+P4 (ours) | 46.7 | 35.8 | 34.8 | 20.0 | 7.5 | 34.5 |
| Model | (%) | (%) | (%) |
|---|---|---|---|
| PAN-P3+P4+P5 | 7.3 | 26.0 | 37.5 |
| PAN-P2+P3+P4 | 9.2 (+1.9) | 27.4 (+1.4) | 36.9 (−0.6) |
| PAN-P2+P3+P4+P5 | 9.2 (+1.9) | 27.4 (+1.4) | 38.3 (+0.8) |
| FPRFN-P2+P3+P4 (ours) | 9.4 (+2.1) | 27.4 (+1.4) | 38.2 (+0.7) |
| Model | DSAF | MSRSA | FPRFN | GConv |
|---|---|---|---|---|
| Baseline | × | × | × | × |
| Model1 | √ | × | × | × |
| Model2 | × | √ | × | × |
| Model3 | × | × | √ | × |
| Model4 | √ | √ | × | × |
| Model5 | √ | × | √ | × |
| Model6 | × | √ | √ | × |
| Model7 | √ | √ | √ | × |
| YOLO-DMF | √ | √ | √ | √ |
| Model | P (%) | R (%) | mAP@0.5 (%) | mAP@0.5:0.95 (%) | Params () | GFLOPs |
|---|---|---|---|---|---|---|
| Baseline | 43.3 | 33.2 | 31.4 | 18.1 | 11.1 | 28.5 |
| Model1 | 44.8 | 34.1 | 32.0 | 18.6 | 11.1 | 30.6 |
| Model2 | 44.3 | 33.4 | 31.7 | 18.3 | 11.7 | 28.9 |
| Model3 | 46.7 | 35.8 | 34.8 | 20.0 | 7.5 | 34.5 |
| Model4 | 43.4 | 34.0 | 32.3 | 18.7 | 11.7 | 31.0 |
| Model5 | 46.7 | 35.9 | 35.1 | 20.4 | 7.5 | 36.7 |
| Model6 | 46.5 | 35.5 | 34.9 | 20.2 | 8.0 | 35.0 |
| Model7 | 47.6 | 36.5 | 35.5 | 20.4 | 8.0 | 37.1 |
| YOLO-DMF | 47.3 | 36.0 | 35.3 | 20.6 | 3.7 | 29.3 |
| Model | P (%) | R (%) | mAP@0.5 (%) | mAP@0.5:0.95 (%) | Params () | GFLOPs | FPS |
|---|---|---|---|---|---|---|---|
| YOLOv3t [37] | 34.3 | 20.6 | 18.8 | 10.2 | 12.1 | 18.9 | 57.2 |
| YOLOv5s [38] | 42.2 | 32.6 | 30.7 | 17.7 | 9.1 | 23.8 | 48.4 |
| YOLOv6s [39] | 42.4 | 31.3 | 30.0 | 17.5 | 16.3 | 44.0 | 36.4 |
| YOLOv7t [40] | 41.7 | 33.8 | 29.4 | 14.5 | 6.0 | 13.1 | 72.7 |
| YOLOv8s [41] | 43.3 | 33.2 | 31.4 | 18.1 | 11.1 | 28.5 | 38.0 |
| YOLOv9s [42] | 45.8 | 33.1 | 31.8 | 18.4 | 7.2 | 26.7 | 29.4 |
| YOLOv10s [43] | 43.9 | 32.9 | 31.4 | 18.0 | 8.0 | 24.5 | 33.8 |
| YOLOv11s [44] | 42.8 | 33.2 | 31.0 | 17.9 | 9.4 | 21.3 | 33.0 |
| YOLOv12s [45] | 43.5 | 33.0 | 31.4 | 18.1 | 9.2 | 21.2 | 31.3 |
| MCA-YOLO [46] | 41.1 | 34.4 | 30.8 | 16.3 | 6.3 | 12.4 | 73.6 |
| MFO-Net [47] | 43.2 | 34.5 | 31.4 | 17.8 | 6.4 | 12.5 | 72.4 |
| MFFSODNet [48] | 39.6 | 33.2 | 30.1 | 16.0 | 1.6 | 21.5 | 50.2 |
| YOLO-DMF | 47.3 | 36.0 | 35.3 | 20.6 | 3.7 | 29.3 | 34.1 |
| Model | P (%) | R (%) | mAP@0.5 (%) | mAP@0.5:0.95 (%) | Params () | GFLOPs | FPS |
|---|---|---|---|---|---|---|---|
| YOLOv3t [37] | 90.3 | 77.9 | 85.4 | 49.2 | 12.1 | 18.9 | 57.2 |
| YOLOv5s [38] | 92.8 | 92.0 | 95.1 | 60.3 | 9.1 | 23.8 | 48.4 |
| YOLOv6s [39] | 93.1 | 90.3 | 94.8 | 59.8 | 16.3 | 44.0 | 36.4 |
| YOLOv7t [40] | 94.4 | 90.7 | 95.3 | 56.1 | 6.0 | 13.1 | 72.7 |
| YOLOv8s [41] | 94.0 | 92.2 | 95.3 | 61.0 | 11.1 | 28.4 | 38.0 |
| YOLOv9s [42] | 94.3 | 92.8 | 96.0 | 61.8 | 7.2 | 26.7 | 29.4 |
| YOLOv10s [43] | 93.9 | 92.0 | 95.9 | 61.4 | 8.0 | 24.5 | 33.8 |
| YOLOv11s [44] | 94.2 | 91.5 | 95.3 | 61.2 | 9.4 | 21.3 | 33.0 |
| YOLOv12s [45] | 93.6 | 92.0 | 95.5 | 61.1 | 9.2 | 21.2 | 31.3 |
| YOLO-DMF | 94.3 | 93.3 | 96.4 | 62.8 | 3.7 | 29.3 | 34.1 |
| Model | P (%) | R (%) | mAP@0.5 (%) | mAP@0.5:0.95 (%) | Params () | GFLOPs | FPS |
|---|---|---|---|---|---|---|---|
| YOLOv3t [37] | 37.9 | 20.9 | 20.9 | 8.6 | 12.1 | 18.9 | 57.2 |
| YOLOv5s [38] | 52.5 | 38.1 | 38.7 | 17.1 | 9.1 | 23.8 | 48.4 |
| YOLOv6s [39] | 70.3 | 32.7 | 34.0 | 15.2 | 16.3 | 44.0 | 36.4 |
| YOLOv7t [40] | 73.0 | 28.1 | 27.1 | 10.6 | 6.0 | 13.2 | 72.7 |
| YOLOv8s [41] | 53.0 | 38.7 | 39.6 | 17.5 | 11.1 | 28.5 | 38.0 |
| YOLOv9s [42] | 58.4 | 38.4 | 39.1 | 17.5 | 7.2 | 26.7 | 29.4 |
| YOLOv10s [43] | 65.3 | 36.6 | 38.8 | 17.4 | 8.0 | 24.5 | 33.8 |
| YOLOv11s [44] | 52.4 | 37.0 | 38.7 | 17.2 | 9.4 | 21.3 | 33.0 |
| YOLOv12s [45] | 48.3 | 27.0 | 28.4 | 12.4 | 9.2 | 21.2 | 31.3 |
| YOLO-DMF | 63.0 | 41.5 | 43.6 | 19.5 | 3.7 | 29.2 | 34.1 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Liu, C.; Wang, Y.; Cao, Q.; Zhang, C.; Cheng, A. Rethinking Adaptive Contextual Information and Multi-Scale Feature Fusion for Small-Object Detection in UAV Imagery. Sensors 2025, 25, 7312. https://doi.org/10.3390/s25237312
Liu C, Wang Y, Cao Q, Zhang C, Cheng A. Rethinking Adaptive Contextual Information and Multi-Scale Feature Fusion for Small-Object Detection in UAV Imagery. Sensors. 2025; 25(23):7312. https://doi.org/10.3390/s25237312
Chicago/Turabian StyleLiu, Chang, Yong Wang, Qiang Cao, Changlei Zhang, and Anyu Cheng. 2025. "Rethinking Adaptive Contextual Information and Multi-Scale Feature Fusion for Small-Object Detection in UAV Imagery" Sensors 25, no. 23: 7312. https://doi.org/10.3390/s25237312
APA StyleLiu, C., Wang, Y., Cao, Q., Zhang, C., & Cheng, A. (2025). Rethinking Adaptive Contextual Information and Multi-Scale Feature Fusion for Small-Object Detection in UAV Imagery. Sensors, 25(23), 7312. https://doi.org/10.3390/s25237312
