LMGANet: A Multi-Scale Guided Aggregation Network for Small-Object Detection in Urban Remote Sensing
Highlights
- We develop LMGANet, a parameter-efficient multi-scale guided network for small-object detection in urban remote sensing imagery. The proposed architecture integrates C3K2-GDF, AMFAN, and LES-head to strengthen local detail extraction, cross-scale feature fusion, and precise localization of small targets in complex urban scenes.
- Experiments on the AI-TOD and VisDrone2019 benchmarks show that LMGANet consistently outperforms the YOLOv11S baseline, delivering mAP50 improvements of 3.2% and 4.8%, respectively, while maintaining a compact model size of 3.63 M parameters and real-time inference capability.
- The results demonstrate that accurate small-object detection in cluttered urban remote sensing imagery can be improved with a parameter-efficient and compact model design, highlighting the value of multi-scale feature enhancement and shared prediction for compact and real-time aerial perception.
- Owing to its favorable balance between accuracy, parameter efficiency, and real-time inference speed, LMGANet shows potential for real-time urban aerial perception tasks such as traffic inspection, public safety monitoring, and infrastructure management, where timely and reliable object screening is required.
Abstract
1. Introduction
- A parameter-efficient multi-scale guided aggregation network, termed LMGANet, is proposed for small-object detection in remote sensing imagery, achieving a favorable balance among detection accuracy, compact model size, and real-time inference capability.
- A C3K2-GDF module is designed to adaptively adjust receptive fields and enhance fine-grained feature representation, thereby strengthening the discriminative expression of small objects in complex aerial scenes.
- An Adaptive Multi-scale Feature Aggregation Network (AMFAN) is constructed to perform bidirectional adaptive multi-scale fusion, effectively bridging the semantic gap between low-level detailed features and high-level semantic features.
- A Lightweight Enhanced Shared (LES) detection head is developed to significantly reduce model parameters through shared decoding while preserving localization accuracy for small objects.
2. Related Work
2.1. Object Detection
2.2. Small Object Detection in Aerial Images
2.3. Efficient and Compact Network Design
3. Methods
3.1. Overall Architecture
3.2. C3K2-GDF Dynamic Multi-Scale Feature Refinement Module
3.3. Adaptive Multi-Scale Feature Aggregation Network (AMFAN)
3.4. Lightweight Enhanced Shared (LES) Detection Head
4. Experiments and Results
4.1. Experimental Setup
4.2. Datasets
4.3. Evaluation Metrics
4.4. Results on VisDrone2019 Dataset
4.4.1. Comparison with State-of-the-Art Detectors
4.4.2. Visual Study
4.5. Results on AI-TOD Dataset
4.5.1. Comparison with State-of-the-Art Detectors
4.5.2. Visual Study
4.6. Ablation Study
4.6.1. Overall Ablation on the Proposed Components
4.6.2. Comparison of Different Feature Fusion Strategies
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Khan, M.R.K.; Rishe, N. A unified framework for vehicle detection, tracking, and counting across ground and aerial views using knowledge distillation with YOLOv10-S. Remote Sens. 2026, 18, 842. [Google Scholar] [CrossRef]
- Yang, G.; Zhao, B.; Zhang, J.; Wen, J.; Li, Q.; Lei, L.; Chen, X.; Chen, B.M. Det-Recon-Reg: An intelligent framework toward automated UAV-based large-scale infrastructure inspection. IEEE Trans. Instrum. Meas. 2025, 74, 3539516. [Google Scholar] [CrossRef]
- Paulraj, S.; Vairavasundaram, S. Transformer-enabled weakly supervised abnormal event detection in intelligent video surveillance systems. Eng. Appl. Artif. Intell. 2025, 139, 109496. [Google Scholar] [CrossRef]
- Yang, F.; Chen, L.; Wang, X.; Zhang, Y.; Li, H.; He, M.; Shen, L. FKIFM-DETR: A multi-domain fusion-based transformer framework for small-target detection in UAV remote sensing imagery. Remote Sens. 2026, 18, 700. [Google Scholar] [CrossRef]
- Yu, D. Toward integrated urban observatories: Synthesizing remote and social sensing in urban science. Remote Sens. 2025, 17, 2041. [Google Scholar] [CrossRef]
- Wang, Q.; Kavhiza, N.J.; Islam, F.; Huqqani, I.A.; Abbas, M.; Barman, S. Multisensor data fusion for coastal boundary detection by Res-U-Net implementation using high-resolution UAV imagery. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 16722–16732. [Google Scholar] [CrossRef]
- Wang, Z.; Yi, J.; Chen, A.; Chen, L.; Lin, H.; Xu, K. Accurate semantic segmentation of very high-resolution remote sensing images considering feature state sequences: From benchmark datasets to urban applications. ISPRS J. Photogramm. Remote Sens. 2025, 220, 824–840. [Google Scholar] [CrossRef]
- Li, Y.; Zhang, C.; Su, W.; Jiang, S.; Nie, D.; Wang, Y.; Wang, Y.; He, H.; Chen, Q.; Martin, S.T.; et al. Copter-type UAV-based sensing in atmospheric chemistry: Recent advances, applications, and future perspectives. Environ. Sci. Technol. 2025, 59, 13532–13550. [Google Scholar] [CrossRef] [PubMed]
- Ding, J.; Xue, N.; Xia, G.-S.; Bai, X.; Yang, W.; Yang, M.Y.; Belongie, S.; Luo, J.; Datcu, M.; Pelillo, M.; et al. Object detection in aerial images: A large-scale benchmark and challenges. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 7778–7796. [Google Scholar] [CrossRef]
- Leng, J.; Ye, Y.; Mo, M.; Gao, C.; Gan, J.; Xiao, B.; Gao, X. Recent advances for aerial object detection: A survey. ACM Comput. Surv. 2024, 56, 296. [Google Scholar] [CrossRef]
- Zhan, Y.; Xiong, Z.; Yuan, Y. RSVG: Exploring data and models for visual grounding on remote sensing data. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5604513. [Google Scholar] [CrossRef]
- Hua, W.; Chen, Q. A survey of small object detection based on deep learning in aerial images. Artif. Intell. Rev. 2025, 58, 162. [Google Scholar] [CrossRef]
- Bakirci, M. Performance evaluation of low-power and lightweight object detectors for real-time monitoring in resource-constrained drone systems. Eng. Appl. Artif. Intell. 2025, 159, 111775. [Google Scholar] [CrossRef]
- Li, Q.; Chen, Y.; Zeng, Y. Transformer with transfer CNN for remote-sensing-image object detection. Remote Sens. 2022, 14, 984. [Google Scholar] [CrossRef]
- Zhang, C.; Jiang, W.; Zhang, Y.; Wang, W.; Zhao, Q.; Wang, C. Transformer and CNN hybrid deep neural network for semantic segmentation of very-high-resolution remote sensing imagery. IEEE Trans. Geosci. Remote Sens. 2022, 60, 4408820. [Google Scholar] [CrossRef]
- Mehta, S.; Rastegari, M. MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer. In Proceedings of the ICLR, Virtual, 25 April 2022. [Google Scholar]
- Zhao, X.; Xia, Y.; Zhang, W.; Zheng, C.; Zhang, Z. YOLO-ViT-Based Method for Unmanned Aerial Vehicle Infrared Vehicle Target Detection. Remote Sens. 2023, 15, 3778. [Google Scholar] [CrossRef]
- Cheng, Q.; Li, X.; Zhu, B.; Shi, Y.; Xie, B. Drone detection method based on MobileViT and CA-PANet. Electronics 2023, 12, 223. [Google Scholar]
- Qin, H.; Zhou, D.; Xu, T.; Bian, Z.; Li, J. Factorization vision transformer: Modeling long-range dependency with local window cost. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 3151–3164. [Google Scholar] [CrossRef] [PubMed]
- Khan, S.; Naseer, M.; Hayat, M.; Zamir, S.W.; Khan, F.S.; Shah, M. Transformers in vision: A survey. ACM Comput. Surv. 2022, 54, 200. [Google Scholar] [CrossRef]
- Li, Z.; Wang, Y.; Zhang, N.; Zhang, Y.; Zhao, Z.; Xu, D.; Ben, G.; Gao, Y. Deep learning-based object detection techniques for remote sensing images: A survey. Remote Sens. 2022, 14, 2385. [Google Scholar] [CrossRef]
- Zhang, X.; Cheng, S.; Wang, L.; Li, H. Asymmetric cross-attention hierarchical network based on CNN and transformer for bitemporal remote sensing images change detection. IEEE Trans. Geosci. Remote Sens. 2023, 61, 2000415. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar] [CrossRef]
- Zhu, L.; Liao, B.; Zhang, Q.; Wang, X.; Liu, W.; Wang, X. Vision Mamba: Efficient visual representation learning with bidirectional state space model. arXiv 2024, arXiv:2401.09417. [Google Scholar] [CrossRef]
- Jiao, J.; Liu, Y.; Liu, Y.; Tian, Y.; Wang, Y.; Xie, L.; Ye, Q.; Yu, H.; Zhao, Y. VMamba: Visual state space model. arXiv 2024, arXiv:2401.10166. [Google Scholar]
- Zhu, X.; Lyu, S.; Wang, X.; Zhao, Q. TPH-YOLOv5: Improved YOLOv5 Based on Transformer Prediction Head for Object Detection on Drone-Captured Scenarios. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, Montreal, QC, Canada, 11–17 October 2021; pp. 2778–2788. [Google Scholar]
- Wang, H.; Liu, J.; Zhao, J.; Zhang, J.; Zhao, D. Precision and speed: LSOD-YOLO for lightweight small object detection. Expert Syst. Appl. 2025, 269, 126440. [Google Scholar] [CrossRef]
- Wang, T.; Ma, Z.; Yang, T.; Zou, S. PETNet: A YOLO-based prior enhanced transformer network for aerial image detection. Neurocomputing 2023, 547, 126384. [Google Scholar] [CrossRef]
- Li, B.; Kang, Y.; Ding, Y.; Li, S.; Zhang, Z.; Ma, D. DAE-YOLO: Remote sensing small object detection method integrating YOLO and state space models. Remote Sens. 2026, 18, 109. [Google Scholar] [CrossRef]
- Xie, H.; Wang, M.; Cao, R.; Wang, J.; Jiang, Y.; Huang, Q.; Jiang, L. MTD-YOLO: A multi-scale perception framework with task decoupling and dynamic alignment for UAV small object detection. Remote Sens. 2025, 17, 3823. [Google Scholar] [CrossRef]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; IEEE: Piscataway, NJ, USA, 2014; pp. 580–587. [Google Scholar]
- Girshick, R. Fast R-CNN. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 7–13 December 2015; IEEE: Piscataway, NJ, USA, 2015; pp. 1440–1448. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef]
- Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature Pyramid Networks for Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 2117–2125. [Google Scholar]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision, Berlin/Heidelberg, Germany, 2016; Springer: Berlin/Heidelberg, Germany, 2016; pp. 21–37. [Google Scholar]
- Liu, X.; Yang, X.; Shao, L.; Wang, X.; Gao, Q.; Shi, H. GM-DETR: Research on a defect detection method based on improved DETR. Sensors 2024, 24, 3610. [Google Scholar] [CrossRef] [PubMed]
- Yuan, H.; Zhang, B. FSSC-Net: A frequency-spatial self-calibrated network for task-adaptive remote sensing image understanding. Remote Sens. 2026, 18, 824. [Google Scholar]
- Zhou, P.; Guo, X.; Sun, X.; Sun, B.; Su, S.; Jiang, W.; Guo, R.; Dang, Z.; Huang, S. RAPT-Net: Reliability-aware precision-preserving tolerance-enhanced network for tiny target detection in wide-area coverage aerial remote sensing. Remote Sens. 2026, 18, 449. [Google Scholar]
- Ma, J.; Bian, M.; Fan, F.; Kuang, H.; Liu, L.; Wang, Z.; Li, T.; Zhang, R. Vision-language guided semantic diffusion sampling for small object detection in remote sensing imagery. Remote Sens. 2025, 17, 3203. [Google Scholar]
- Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adam, H. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv 2017, arXiv:1704.04861. [Google Scholar] [CrossRef]
- Sandler, M.; Howard, A.; Zhu, M.; Zhmoginov, A.; Chen, L.-C. MobileNetV2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 4510–4520. [Google Scholar]
- Howard, A.; Sandler, M.; Chen, B.; Wang, W.; Chen, L.-C.; Tan, M.; Chu, G.; Vasudevan, V.; Zhu, Y.; Pang, R.; et al. Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1314–1324. [Google Scholar]
- Xue, C.; Xia, Y.; Wu, M.; Chen, Z.; Cheng, F.; Yun, L. EL-YOLO: An Efficient and Lightweight Low-Altitude Aerial Objects Detector for Onboard Applications. Expert Syst. Appl. 2024, 256, 124848. [Google Scholar] [CrossRef]
- Jiang, C.; Ren, H.; Ye, X.; Zhu, J.; Zeng, H.; Nan, Y.; Sun, M.; Ren, X.; Huo, H. Object Detection from UAV Thermal Infrared Images and Videos Using YOLO Models. Int. J. Appl. Earth Obs. Geoinf. 2022, 112, 102912. [Google Scholar] [CrossRef]
- Fan, Q.; Li, Y.; Deveci, M.; Zhong, K.; Kadry, S. LUD-YOLO: A novel lightweight object detection network for unmanned aerial vehicle. Inf. Sci. 2025, 686, 121366. [Google Scholar]
- Min, X.; Zhou, W.; Hu, R.; Wu, Y.; Pang, Y.; Yi, J. LWUAVDet: A lightweight UAV object detection network on edge devices. IEEE Internet Things J. 2024, 11, 24013–24023. [Google Scholar] [CrossRef]
- Cheng, S.; Song, J.; Zhou, M.; Wei, X.; Pu, H.; Luo, J.; Jia, W. EF-DETR: A lightweight transformer-based object detector with an encoder-free neck. IEEE Trans. Ind. Inform. 2024, 20, 12994–13002. [Google Scholar]
- Wang, R.; Lin, C.; Li, Y. RPLFDet: A lightweight small object detection network for UAV aerial images with rational preservation of low-level features. IEEE Trans. Instrum. Meas. 2025, 74, 5013514. [Google Scholar] [CrossRef]
- Yu, Y.; Zhang, Y.; Cheng, Z.; Song, Z.; Tang, C. MCA: Multidimensional collaborative attention in deep convolutional neural networks for image recognition. Eng. Appl. Artif. Intell. 2023, 126, 107079. [Google Scholar] [CrossRef]
- Si, Y.; Xu, H.; Zhu, X.; Zhang, W.; Dong, Y.; Chen, Y.; Li, H. SCSA: Exploring the synergistic effects between spatial and channel attention. Neurocomputing 2025, 634, 129866. [Google Scholar] [CrossRef]
- Zhu, P.; Wen, L.; Du, D.; Bian, X.; Fan, H.; Hu, Q.; Ling, H. Detection and Tracking Meet Drones Challenge. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 7380–7399. [Google Scholar] [PubMed]
- Wang, J.; Yang, W.; Guo, H.; Zhang, R.; Xia, G.-S. Tiny object detection in aerial images. In Proceedings of the 25th International Conference on Pattern Recognition, Milan, Italy, 10–15 January 2021; pp. 3791–3798. [Google Scholar]
- Redmon, J.; Farhadi, A. YOLOv3: An incremental improvement. arXiv 2018, arXiv:1804.02767. [Google Scholar] [CrossRef]
- Wang, C.-Y.; Bochkovskiy, A.; Liao, H.-Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 7464–7475. [Google Scholar]
- Wang, C.-Y.; Yeh, I.-H.; Liao, H.-Y.M. YOLOv9: Learning what you want to learn using programmable gradient information. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; pp. 1–21. [Google Scholar]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. arXiv 2024, arXiv:2405.14458. [Google Scholar]
- Lei, M.; Li, S.; Wu, Y.; Hu, H.; Zhou, Y.; Zheng, X.; Ding, G.; Du, S.; Wu, Z.; Gao, Y. YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception. arXiv 2025, arXiv:2506.17733. [Google Scholar]
- Xiao, Y.; Xu, T.; Xin, Y.; Li, J. FBRT-YOLO: Faster and Better for Real-Time Aerial Image Detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025; Volume 39, pp. 8673–8681. [Google Scholar]
- Zhao, Y.; Chen, W.; Wang, X.; Li, Y.; Luo, Z.; Yuan, J. RT-DETR: Real-time detection transformer. arXiv 2023, arXiv:2304.08069. [Google Scholar]
- Liu, Y.; He, M.; Hui, B. ESO-DETR: An Improved Real-Time Detection Transformer Model for Enhanced Small Object Detection in UAV Imagery. Drones 2025, 9, 143. [Google Scholar] [CrossRef]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 11–17 October 2021; pp. 10012–10022. [Google Scholar]
- Tan, M.; Pang, R.; Le, Q.V. EfficientDet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 14–19 June 2020; pp. 10781–10790. [Google Scholar]
- Yang, G.; Lei, J.; Zhu, Z.; Cheng, S.; Feng, Z.; Liang, R. AFPN: Asymptotic Feature Pyramid Network for Object Detection. arXiv 2023, arXiv:2306.15988. [Google Scholar] [CrossRef]
- Wang, C.; Han, Y.; Yang, C.; Wu, M.; Chen, Z.; Yun, L.; Jin, X. CF-YOLO for Small Target Detection in Drone Imagery Based on YOLOv11 Algorithm. Sci. Rep. 2025, 15, 16741. [Google Scholar] [CrossRef]








| Parameter | Configuration |
|---|---|
| CPU | Intel Core i9-10920X |
| GPU | NVIDIA RTX A4000 |
| CUDA Version | 11.8 |
| Python Version | 3.11.11 |
| PyTorch Version | 2.2.2 |
| Input Size | 640 × 640 |
| Epochs | 250 |
| Batch Size | 8 |
| Optimizer | SGD |
| Methods | P (%) | R (%) | mAP50 (%) | mAP50–95 (%) | APs | Params (M) | GFLOPs | FPS |
|---|---|---|---|---|---|---|---|---|
| Faster R-CNN [34] | 45.1 | 33 | 32.2 | 18.8 | 10.2 | 41.40 | 208.0 | 35.4 |
| YOLOv3-tiny [54] | 38.2 | 28.4 | 21.5 | 12.3 | 6.8 | 8.70 | 12.9 | 109.7 |
| YOLOv5S | 44.1 | 33.6 | 34.3 | 20.1 | 11.8 | 9.10 | 24.1 | 93.3 |
| YOLOv5M | 47.2 | 36.5 | 39.5 | 23.2 | 14.3 | 20.90 | 48.0 | 80.6 |
| YOLOv5L | 50.5 | 38.5 | 41.4 | 24.7 | 16.1 | 46.20 | 108.0 | 69.8 |
| YOLOv7-tiny [55] | 47.5 | 37.4 | 35.7 | 18.9 | 11.2 | 6.00 | 10.6 | 113.2 |
| YOLOv8S | 48.7 | 35.3 | 38.2 | 23.1 | 9.8 | 11.10 | 28.7 | 89.2 |
| YOLOv8M | 51.2 | 38.7 | 42.2 | 26.2 | 12.6 | 25.60 | 78.5 | 71.9 |
| YOLOv9S [56] | 52.5 | 30.1 | 39.1 | 23.4 | 13.8 | 7.30 | 27.4 | 90.3 |
| YOLOv10S [57] | 49.3 | 35.9 | 38.7 | 24.5 | 13.2 | 8.07 | 24.8 | 92.6 |
| YOLOv10M | 51.4 | 39 | 42.4 | 26.3 | 18.4 | 16.50 | 64.0 | 75.9 |
| YOLOv11S | 49.7 | 38.6 | 39.7 | 23.4 | 13.8 | 9.40 | 21.6 | 96.3 |
| YOLOv11M | 54.3 | 42.3 | 44 | 26.9 | 19.3 | 20.00 | 68.2 | 75.3 |
| YOLOv13S [58] | 48.3 | 35.7 | 36.8 | 21.5 | 11.5 | 9.00 | 21 | 97.1 |
| FBRT-YOLO-S [59] | 50.4 | 38.2 | 41.3 | 25.6 | 15.8 | 2.90 | 22.9 | 89.9 |
| FBRT-YOLO-M [59] | 53.7 | 42.4 | 44.1 | 26.9 | 17.3 | 7.20 | 58.7 | 70.5 |
| TPH-YOLOv5 [27] | 44.5 | 33.9 | 34.7 | 20.7 | 11.3 | 7.50 | 13.8 | 101.2 |
| RT-DETR-R18 [60] | 51.5 | 38.9 | 42.5 | 26.4 | 16.5 | 20.00 | 57.1 | 55.0 |
| ESO-DETR [61] | 57.6 | 43.1 | 47.1 | 28.9 | 20.1 | 14.9 | 66.0 | 58.3 |
| Swin Transformer [62] | 46.5 | 33.5 | 37.1 | 21.8 | 8.7 | 34.2 | 44.6 | 42.1 |
| YOLO-ViT [17] | 48.0 | 35.1 | 38.5 | 23.3 | 9.6 | 17.3 | 33.1 | 80.3 |
| Ours | 54.8 | 42.6 | 44.5 | 27.2 | 19.7 | 3.63 | 28.9 | 92.1 |
| Methods | P (%) | R (%) | mAP50 (%) | mAP50–95 (%) | APs | Params (M) | GFLOPs | FPS |
|---|---|---|---|---|---|---|---|---|
| YOLOv5S | 26.3 | 19.8 | 28.7 | 9.2 | 9.2 | 9.10 | 24.1 | 94.1 |
| YOLOv8S | 30.2 | 22.7 | 33.2 | 11.3 | 11.3 | 11.10 | 28.7 | 89.5 |
| YOLOv10S | 29.8 | 22.3 | 32.8 | 10.9 | 10.9 | 8.07 | 24.8 | 91.3 |
| YOLOv11S | 30.6 | 23 | 33.6 | 11 | 11.6 | 9.40 | 21.6 | 96.1 |
| YOLOv13S | 30.1 | 22.4 | 32.5 | 10.8 | 10.6 | 9.00 | 21 | 97.0 |
| TPH-YOLOv5 | 29.1 | 21.8 | 31.6 | 10.4 | 10.2 | 7.50 | 13.8 | 102.5 |
| FBRT-YOLO-S | 28.9 | 21.6 | 31.2 | 10.2 | 10.1 | 2.90 | 22.9 | 91.7 |
| Ours | 33.2 | 25.4 | 36.7 | 12.4 | 12.2 | 3.63 | 28.9 | 90.7 |
| Model | C3K2-GDF | AMFAN | LES-Head | P (%) | R (%) | mAP50 (%) | APs | Params (M) | GFLOPs |
|---|---|---|---|---|---|---|---|---|---|
| Baseline | × | × | × | 49.7 | 38.6 | 39.7 | 13.8 | 9.40 | 21.6 |
| Model 1 | √ | × | × | 51.6 | 39.8 | 41.3 | 15.2 | 8.15 | 23.5 |
| Model 2 | × | √ | × | 51.3 | 40.4 | 41.7 | 15.6 | 6.27 | 25.1 |
| Model 3 | × | × | √ | 50.8 | 39.2 | 40.6 | 14.3 | 5.54 | 22.1 |
| Model 4 | √ | √ | × | 53.2 | 41.5 | 43.2 | 17.4 | 5.18 | 27.0 |
| Model 5 | √ | × | √ | 52.4 | 40.6 | 42.1 | 16.1 | 4.76 | 24.6 |
| Model 6 | × | √ | √ | 52.1 | 41.1 | 42.5 | 16.5 | 4.08 | 26.2 |
| Ours | √ | √ | √ | 54.8 | 42.6 | 44.5 | 19.7 | 3.63 | 28.9 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhu, H.; Jiang, C.; Zhu, X. LMGANet: A Multi-Scale Guided Aggregation Network for Small-Object Detection in Urban Remote Sensing. Remote Sens. 2026, 18, 1578. https://doi.org/10.3390/rs18101578
Zhu H, Jiang C, Zhu X. LMGANet: A Multi-Scale Guided Aggregation Network for Small-Object Detection in Urban Remote Sensing. Remote Sensing. 2026; 18(10):1578. https://doi.org/10.3390/rs18101578
Chicago/Turabian StyleZhu, Haoliang, Chunli Jiang, and Xiuli Zhu. 2026. "LMGANet: A Multi-Scale Guided Aggregation Network for Small-Object Detection in Urban Remote Sensing" Remote Sensing 18, no. 10: 1578. https://doi.org/10.3390/rs18101578
APA StyleZhu, H., Jiang, C., & Zhu, X. (2026). LMGANet: A Multi-Scale Guided Aggregation Network for Small-Object Detection in Urban Remote Sensing. Remote Sensing, 18(10), 1578. https://doi.org/10.3390/rs18101578

