DMAC-Net: Direction-Aware Multi-Granularity Enhancement with Asymmetric Context Guidance for Multimodal UAV-Based Small Object Detection
Abstract
1. Introduction
- To address the unique physical challenges inherent in complex aerial environments, we construct a multimodal detection network that integrates direction-aware multi-granularity and asymmetric context guidance, termed DMAC-Net, which provides a deep coupling paradigm for small object detection in all-weather scenarios.
- A Direction-Aware Multi-Granularity Enhancement (DAGE) module is designed for unified backbone feature extraction. By capturing local orientations and edge contours of targets to suppress background false alarms, and enlarging the receptive field via a multi-granularity mechanism, it effectively improves the detection performance of dense and feature-sparse small targets.
- The Asymmetric Context Guided Fusion (ACGF) module is constructed to achieve semantic alignment and complementarity between infrared thermal radiation priors and visible geometric details via asymmetric feature guidance and dynamic weight assignment, thereby enabling efficient cross-modal feature interaction.
- The effectiveness of the proposed model is validated across multiple multimodal datasets. Extensive experiments show that DMAC-Net achieves superior performance not only on UAV-view tasks including RGBTDronePerson and AVMS, but also maintains high-precision localization and robustness in extreme lighting and low-altitude near-ground wide-area surveillance scenarios, as demonstrated on the LLVIP dataset.
2. Related Work
3. Methods
3.1. Overall Model Architecture
3.2. Direction-Aware Granularity Enhancement Block (DAGE)
3.2.1. Angular Pinwheel Convolution Mechanism (APConv)
3.2.2. Multi-Granularity Dilation Aggregator (MGD)
3.3. Asymmetric Context Guided Fusion Module (ACGF)
4. Experiments
4.1. Datasets and Preprocessing
4.2. Experimental Settings
4.3. Ablation Studies
4.3.1. Synergistic Ablation Analysis of Core Components Within the DAGE Module
4.3.2. Ablation Study on the Overall Architecture of the Multimodal Detection Network
4.3.3. False Positive Error Evolution Analysis
4.3.4. Visualization Analysis
4.3.5. Validation of the Direction-Aware Front-End Strategy
4.4. Comparative Experiments
4.5. Cross-Scale Scenario Comparison
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- White, E.C.; Seymour, A.C.; Dale, J.; Newton, E.; Johnston, D.W. Mapping the ‘Ghost Fleet’ of Mallows Bay, Maryland Using Drone Remote Sensing. Sci. Data 2025, 12, 1547. [Google Scholar] [CrossRef] [PubMed]
- Deng, Q.; Zhang, Y.; Lin, Z.; Gao, X.; Weng, Z. The Impact of Digital Technology Application on Agricultural Low-Carbon Transformation—A Case Study of the Pesticide Reduction Effect of Plant Protection Unmanned Aerial Vehicles (UAVs). Sustainability 2024, 16, 10920. [Google Scholar] [CrossRef]
- Liu, H.; Yu, Y.; Liu, S.; Wang, W. A Military Object Detection Model of UAV Reconnaissance Image and Feature Visualization. Appl. Sci. 2022, 12, 12236. [Google Scholar] [CrossRef]
- Wang, W.; Lu, B.; Wu, C.H. Cost-effective drone monitoring and evaluating toolkits for stream habitat health: Development and application. Environ. Monit. Assess. 2026, 198, 10. [Google Scholar]
- Lyu, Z.; Gao, Y.; Chen, J.; Du, H.; Xu, J.; Huang, K.; Kim, D.I. Empowering Intelligent Low-Altitude Economy with Large AI Model Deployment. IEEE Wirel. Commun. 2026, 33, 64–72. [Google Scholar] [CrossRef]
- Matos-Carvalho, J.P.; Seman, L.O.; Stefanon, S.F.; Khreast, M.K.M.; Villarrubia González, G. A Novel YOLO26-MoE Optimized by an LLM Agent for Insulator Fault Detection Considering UAV Images. arXiv 2026, arXiv:2605.19595. [Google Scholar] [CrossRef]
- Han, J.; Sun, F.; Xu, Z.; Song, L.; Fang, J. An Enhanced Algorithm Integrating YOLOv11 and ByteTrack for Small-Object Detection and Tracking in Low-Altitude Remote Sensing Imagery. Remote Sens. 2026, 18, 1547. [Google Scholar] [CrossRef]
- Lei, H.; Shang, L.; Zhao, H.; Yang, W. Polarity aware detection transformer with hierarchical cross attention for unmanned aerial vehicle small object detection. Pattern Recognit. 2026, 179, 113812. [Google Scholar] [CrossRef]
- Liu, J.; Liu, Q.; Wang, R.; Qu, M. HB-YOLOv11: A model focused on enhancing the detection of remote sensing small targets in complex backgrounds. Signal Image Video Process. 2026, 20, 249. [Google Scholar] [CrossRef]
- Zhou, M.; He, S.; Wang, C.; Wang, J. EFSL-YOLO: An Improved Model for Small Object Detection in UAV Vision. Drones 2026, 10, 243. [Google Scholar] [CrossRef]
- Yang, J.; Liu, S.; Wu, J.; Su, X.; Hai, N.; Huang, X. Pinwheel-shaped convolution and scale-based dynamic loss for infrared small target detection. arXiv 2024, arXiv:2412.16986. [Google Scholar] [CrossRef]
- Wang, Y.; Jiang, Y.; Zeng, W.; Cao, S. Research on detection methods for dynamic ship targets in complex marine environment from visible light images. Eng. Rep. 2025, 7, e70000. [Google Scholar] [CrossRef]
- Li, Z.; Wang, Q.; Zhao, Z. Research on small target detection in multispectral remote sensing images based on multimodal deep learning. J. Phys. Conf. Ser. 2024, 2917, 012030. [Google Scholar] [CrossRef]
- Tang, S.; Cao, C.; Lu, J.; Fan, Q.; Hu, S.; Ding, G. Resolving reliability-generalization dilemma in computational fluid dynamics surrogate models for residential environments: A multimodal feature fusion and meta-learning approach. J. Build. Eng. 2026, 127, 116336. [Google Scholar] [CrossRef]
- Jiao, T.; Guo, C.; Feng, X.; Chen, Y.; Song, J. A Comprehensive Survey on Deep Learning Multi-Modal Fusion: Methods, Technologies and Applications. Comput. Mater. Contin. 2024, 80, 1–35. [Google Scholar] [CrossRef]
- Liu, Y.; Yang, Y.; Li, X.; Yang, F.; Xie, H.; Wang, W.; Dong, C. A Deep Learning-Based Pipeline for Detecting Rip Currents from Satellite Imagery. Remote Sens. 2026, 18, 368. [Google Scholar] [CrossRef]
- Deng, Y.; Hu, Y.; Ye, Y.; Xu, P. AD-YOLO: A Unified Method for Traffic-Dense and Small Object Detection in UAV Images. Drones 2026, 10, 338. [Google Scholar] [CrossRef]
- Yang, Y.; Guo, F.; Niu, P. UAVDet: A CNN–Mamba hybrid network for efficient small object detection in UAV imagery. Comput. Vis. Image Underst. 2026, 264, 104637. [Google Scholar] [CrossRef]
- Mahaur, B.; Mishra, K.K. Small-object detection based on YOLOv5 in autonomous driving systems. Pattern Recognit. Lett. 2023, 168, 115–122. [Google Scholar] [CrossRef]
- Kisantal, M.; Wojna, Z.; Murawski, J.; Naruniec, J.; Cho, K. Augmentation for small object detection. arXiv 2019, arXiv:1902.07296. [Google Scholar] [CrossRef]
- Chen, Y.; Zhang, P.; Li, Z.; Li, Y.; Zhang, X.; Meng, G.; Xiang, S.; Sun, J.; Jia, J. Stitcher: Feedback-driven Data Provider for Object Detection. arXiv 2020, arXiv:2004.12432. [Google Scholar] [CrossRef]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 936–944. [Google Scholar] [CrossRef]
- Yang, C.; Huang, Z.; Wang, N. QueryDet: Cascaded sparse query for accelerating high-resolution small object detection. arXiv 2021, arXiv:2103.09136. [Google Scholar] [CrossRef]
- Zhu, X.; Lyu, S.; Wang, X.; Zhao, Q. TPH-YOLOv5: Improved YOLOv5 based on transformer prediction head for object detection on drone-captured scenarios. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, BC, Canada, 11–17 October 2021; pp. 2778–2788. [Google Scholar] [CrossRef]
- Zhou, S.; Zhou, H.; Qian, L. A multi-scale small object detection algorithm SMA-YOLO for UAV remote sensing images. Sci. Rep. 2025, 15, 9255. [Google Scholar] [CrossRef] [PubMed]
- Qi, H.; Qin, H.; Xiang, X.; Yang, C.; Tan, Y. LF-DETR: A Laplacian frequency enhanced DETR for aerial RGB-infrared pedestrian detection. Remote Sens. 2026, 18, 531. [Google Scholar] [CrossRef]
- Chen, Y.; Wang, B.; Guo, X.; Zhu, W.; He, J.; Liu, X.; Yuan, J. DEYOLO: Dual-feature-enhancement YOLO for cross-modality object detection. arXiv 2024, arXiv:2412.04931. [Google Scholar] [CrossRef]
- Shen, J.; Chen, Y.; Liu, Y.; Zuo, X.; Fan, H.; Yang, W. ICAFusion: Iterative cross-attention guided feature fusion for multispectral object detection. arXiv 2023, arXiv:2308.07504. [Google Scholar] [CrossRef]
- Zhang, J.; Lei, J.; Xie, W.; Fang, Z.; Li, Y.; Du, Q. SuperYOLO: Super resolution assisted object detection in multimodal remote sensing imagery. arXiv 2022, arXiv:2209.13351. [Google Scholar] [CrossRef]
- Tang, L.; Zhang, H.; Xu, H.; Ma, J. Rethinking the necessity of image fusion in high-level vision tasks: A practical infrared and visible image fusion network based on progressive semantic injection and scene fidelity. Inf. Fusion 2023, 99, 101870. [Google Scholar] [CrossRef]
- Moesl, B.; Schaffernak, H.; Vorraber, W.; Braunstingl, R.; Koglbauer, I.V. Multimodal augmented reality applications for training of traffic procedures in aviation. Multimodal Technol. Interact. 2022, 7, 3. [Google Scholar] [CrossRef]
- Roy, D.; Fragulis, G.; Cantu Campos, H.A.; Martens, W.L.; Cohen, M. Spatial navigation by seated users of multimodal augmented reality systems. SHS Web Conf. 2021, 102, 04022. [Google Scholar] [CrossRef]
- Hovhannisyan, S. Robust perception in degraded visual environments: A multimodal enhancement framework. Pattern Recognit. Image Anal. 2026, 35, 1003–1014. [Google Scholar] [CrossRef]
- Yu, X.; Wei, Z.; Wu, H.; Lu, C.; Wu, Z.; Zhan, T. IMENet: Infrared-guided multimodal enhancement network for low-light vision. Multimed. Syst. 2025, 32, 20. [Google Scholar] [CrossRef]
- Zheng, Z.; Wu, H.; Lv, L.; Bardou, D.; Niu, S.; Yu, G. MERGE: Multimodal-enhanced representation and guided ensemble for pneumonia recognition in chest X-ray images. J. Supercomput. 2025, 81, 907. [Google Scholar] [CrossRef]
- Liu, Q.; Zhang, D.; Li, S. FedVPN: A Federated Multi-Modal Perception Framework for Multi-UAV in Mountain Search and Rescue. Electronics 2026, 15, 2678. [Google Scholar] [CrossRef]
- Wang, B.; Ge, F.; Xu, Y.; Li, Z. CFISRO: Cross-modal feature interaction and similarity ranking optimization for image-text retrieval. Concurr. Comput. Pract. Exp. 2025, 37, e70427. [Google Scholar] [CrossRef]
- Wu, J.; Zeng, J.; Dong, W.; Shi, G.; Lin, W. Blind image quality assessment with hierarchy: Degradation from local structure to deep semantics. J. Vis. Commun. Image Represent. 2018, 58, 353–362. [Google Scholar] [CrossRef]
- Wei, H.; Liu, X.; Xu, S.; Dai, Z.; Dai, Y.; Xu, X. DWRSeg: Rethinking efficient acquisition of multi-scale contextual information for real-time semantic segmentation. arXiv 2022, arXiv:2212.01173. [Google Scholar] [CrossRef]
- Wu, T.; Tang, S.; Zhang, R.; Zhang, Y. CGNet: A light-weight context guided network for semantic segmentation. IEEE Trans. Image Process. 2018, 30, 1169–1179. [Google Scholar]
- Khanam, R.; Hussain, M. YOLOv11: An overview of the key architectural enhancements. arXiv 2024, arXiv:2410.17725. [Google Scholar] [CrossRef]
- Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-centric real-time object detectors. arXiv 2025, arXiv:2502.12524. [Google Scholar] [CrossRef]
- Ghorbani Kolahi, S.; Chaharsooghi, S.K.; Khatibi, T.; Bozorgpour, A.; Azad, R.; Heidari, M.; Hacihaliloglu, I.; Merhof, D. MSA^2Net: Multi-scale adaptive attention-guided network for medical image segmentation. arXiv 2024, arXiv:2407.21640. [Google Scholar] [CrossRef]
- Yang, Y.; Yuan, G.; Li, J. SFFNet: A wavelet-based spatial and frequency domain fusion network for remote sensing segmentation. arXiv 2024, arXiv:2405.01992. [Google Scholar] [CrossRef]
- Fang, Q.; Han, D.; Wang, Z. Cross-modality fusion transformer for multispectral object detection. arXiv 2021, arXiv:2111.00273. [Google Scholar] [CrossRef]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. In Proceedings of the 38th Conference on Neural Information Processing Systems (NeurIPS 2024), Vancouver, BC, Canada, 10–15 December 2024. [Google Scholar]
- Lei, M.; Li, S.; Wu, Y.; Hu, H.; Zhou, Y.; Zheng, X.; Ding, G.; Du, S.; Wu, Z.; Gao, Y. YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception. arXiv 2025, arXiv:2506.17733. [Google Scholar] [CrossRef]
- Chen, H.; Li, Y.; Jia, Y.; Yao, G.; Zhu, R. RKF-YOLO: A Lightweight Dual-Task Model for Illegal Parking Detection and License Plate Recognition on Edge Devices. Electronics 2026, 15, 2638. [Google Scholar] [CrossRef]
- Zhao, F.; Lou, W.; Feng, H.; Ding, N.; Li, C. MFMG-Net: Multispectral Feature Mutual Guidance Network for Visible–Infrared Object Detection. Drones 2024, 8, 112. [Google Scholar] [CrossRef]
- Wang, Y.; Wang, Y.; Rohra, A.; Yin, B. End-to-end model compression via pruning and knowledge distillation for lightweight image super resolution. Pattern Anal. Appl. 2025, 28, 94. [Google Scholar] [CrossRef]










| Parameter | Value |
|---|---|
| GPU | NVIDIA GeForce RTX 5070 Ti (NVIDIA Corp., Santa Clara, USA) |
| Python version | 3.11 (Python Software Foundation, Beaverton, OR, USA) |
| CUDA version | 12.8 (NVIDIA Corp., Santa Clara, CA, USA). |
| PyTorch version | 2.9.0 (PyTorch Foundation, San Francisco, CA, USA) |
| Momentum | 0.937 |
| Weight Decay | 0.0005 |
| Methods | APConv | MGD | Precision (%) | Recall (%) | F1 | Map50 (%) | Map50-95 (%) | Map75 (%) | GFLOPs |
|---|---|---|---|---|---|---|---|---|---|
| Base | × | × | 58.0 | 49.9 | 53.6 | 50.7 | 27.3 | 26.0 | 6.3 |
| Only APConv | √ | × | 63.7 | 50.6 | 56.4 | 53.9 | 29.6 | 28.6 | 6.8 |
| Only MGD | × | √ | 63.3 | 53.0 | 57.6 | 55.6 | 30.3 | 29.6 | 7.3 |
| DAGE | √ | √ | 63.5 | 53.4 | 58.0 | 56.1 | 30.6 | 29.7 | 7.8 |
| Methods | DAGE | ACGF | MGD | Modality | Precision (%) | Recall (%) | Map50 (%) | Map50-95 (%) | Map75 (%) |
|---|---|---|---|---|---|---|---|---|---|
| Base | × | × | × | single | 48.3 | 45.6 | 40.2 | 14.9 | 6.6 |
| Dual-Base | × | × | × | dual | 52.3 | 48.8 | 44.9 | 16.6 | 7.3 |
| Only DAGE | √ | × | × | dual | 54.3 | 50.8 | 49.0 | 18.4 | 8.0 |
| Only ACGF | × | √ | × | dual | 53.8 | 50.1 | 46.5 | 17.1 | 7.5 |
| ACGF+DAGE | √ | √ | × | dual | 55.9 | 53.4 | 50.0 | 18.8 | 8.3 |
| Only MGD | × | × | √ | dual | 55.6 | 49.8 | 48.0 | 17.7 | 7.3 |
| DMAC-Net | √ | √ | √ | dual | 56.4 | 50.9 | 50.2 | 19.1 | 8.6 |
| Class | Model | Bg-FP | Cls-FP |
|---|---|---|---|
| Crowd | Base | 755 | 260 |
| Dual-Base | 894 | 397 | |
| DMAC-Net | 663 | 252 | |
| Person | Base | 2164 | 322 |
| Dual-Base | 2203 | 287 | |
| DMAC-Net | 1275 | 315 |
| Methods | Modality | Precision | Recall | Map50 | Map50-95 | Parameters(M) | GFLOPs(G) |
|---|---|---|---|---|---|---|---|
| Conv | dual | 0.523 | 0.488 | 0.449 | 0.166 | 3.72 | 9.4 |
| APConv-pre | dual | 0.567 | 0.510 | 0.487 | 0.178 | 3.67 | 9.2 |
| APConv-post | dual | 0.548 | 0.514 | 0.482 | 0.181 | 4.80 | 16.8 |
| Methods | Precision (%) | Recall (%) | F1 (%) | Map50 (%) | Map50-95 (%) |
|---|---|---|---|---|---|
| YOLOv12 [42] | 52.6 | 48.1 | 50.2 | 44.7 | 16.1 |
| MASAG [43] | 51.1 | 47.5 | 49.2 | 42.8 | 15.7 |
| ICAFusion [28] | 52.9 | 47.7 | 50.2 | 44.4 | 15.9 |
| MDAF [44] | 55.7 | 50.1 | 52.7 | 47.3 | 17.2 |
| DEYOLO [27] | 55.8 | 51.2 | 53.4 | 47.5 | 17.2 |
| DMAC-Net | 56.4 | 50.9 | 53.5 | 50.2 | 19.1 |
| Methods | Precision (%) | Recall (%) | F1 (%) | Map50 (%) | Map50-95 (%) |
|---|---|---|---|---|---|
| YOLOv12 [42] | 55.0 | 51.7 | 53.3 | 51.2 | 29.7 |
| MASAG [43] | 45.5 | 50.1 | 47.7 | 45.1 | 25.2 |
| ICAFusion [28] | 68.1 | 45.8 | 54.8 | 53.5 | 29.3 |
| MDAF [44] | 60.1 | 49.4 | 54.2 | 52.8 | 29.3 |
| DEYOLO [27] | 56.8 | 47.1 | 51.5 | 56.5 | 24.7 |
| DMAC-Net | 60.6 | 51.7 | 55.8 | 57.6 | 33.2 |
| Methods | Precision | Recall | F1 | Map50 |
|---|---|---|---|---|
| YOLOv12 [42] | 0.955 | 0.861 | 0.906 | 0.942 |
| CFT [45] | 0.934 | 0.870 | 0.901 | 0.939 |
| MASAG [43] | 0.946 | 0.880 | 0.912 | 0.948 |
| ICAFusion [28] | 0.952 | 0.891 | 0.920 | 0.954 |
| MDAF [44] | 0.951 | 0.883 | 0.916 | 0.956 |
| DEYOLO [27] | 0.932 | 0.890 | 0.911 | 0.948 |
| YOLOv10 [46] | 0.946 | 0.880 | 0.912 | 0.940 |
| YOLOv13 [47] | 0.950 | 0.884 | 0.916 | 0.948 |
| DMAC-Net | 0.953 | 0.892 | 0.922 | 0.952 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Cheng, Q.; Jiang, Y.; Gao, Y.; Gao, Z.; Liu, S.; Tu, X. DMAC-Net: Direction-Aware Multi-Granularity Enhancement with Asymmetric Context Guidance for Multimodal UAV-Based Small Object Detection. Electronics 2026, 15, 3384. https://doi.org/10.3390/electronics15153384
Cheng Q, Jiang Y, Gao Y, Gao Z, Liu S, Tu X. DMAC-Net: Direction-Aware Multi-Granularity Enhancement with Asymmetric Context Guidance for Multimodal UAV-Based Small Object Detection. Electronics. 2026; 15(15):3384. https://doi.org/10.3390/electronics15153384
Chicago/Turabian StyleCheng, Qing, Yan Jiang, Yuan Gao, Zeng Gao, Su Liu, and Xiaoguang Tu. 2026. "DMAC-Net: Direction-Aware Multi-Granularity Enhancement with Asymmetric Context Guidance for Multimodal UAV-Based Small Object Detection" Electronics 15, no. 15: 3384. https://doi.org/10.3390/electronics15153384
APA StyleCheng, Q., Jiang, Y., Gao, Y., Gao, Z., Liu, S., & Tu, X. (2026). DMAC-Net: Direction-Aware Multi-Granularity Enhancement with Asymmetric Context Guidance for Multimodal UAV-Based Small Object Detection. Electronics, 15(15), 3384. https://doi.org/10.3390/electronics15153384
