LE-DETR: A Lightweight and Efficient Model for Small-Object Detection in Remote Sensing Images
Highlights
- This paper proposes a lightweight and efficient detection model, LE-DETR, designed to balance accuracy and computational efficiency in order to effectively detect small targets in remote sensing images.
- By introducing the LFEM, ESFFM and EFAM-FPN modules, the model’s detection accuracy for small objects has been improved whilst suppressing complex background noise.
- Extensive experiments on the VisDrone2019, NWPU VHR-10 and DIOR datasets have confirmed the model’s leading position in terms of overall performance, demonstrating strong robustness and generalisation capabilities.
- LE-DETR provides an effective, lightweight solution for small-object detection in remote sensing scenarios and demonstrates significant potential in practical Earth observation applications such as agricultural monitoring and urban planning.
Abstract
1. Introduction
- To address the difficulty standard convolutional modules face in capturing multi-scale feature information of objects, a multi-branch architecture is constructed and a lightweight feature extraction module (LFEM) is introduced. This module is designed to extract rich object features across multiple scales, thereby preventing the loss of feature information during downsampling of feature maps and improving detection accuracy.
- Addressing the limitation of traditional object detection methods in remote sensing imagery, which rely excessively on spatial domain features whilst neglecting frequency domain information, this study proposes an Efficient Spatio-Frequency Fusion Module (ESFFM). This module aims to fully exploit the frequency information contained within shallow-level features and minimise the loss of useful information, thereby significantly enhancing the ability of lightweight detectors to detect small objects against complex backgrounds.
- To address the challenge faced by traditional Feature Pyramid Networks (FPNs) in integrating fine-grained information from shallow layers with semantic information from deep layers, an efficient Frequency-Aware Merge Feature Pyramid Network (EFAM-FPN) has been designed. By combining large-scale kernel awareness with small-scale kernel aggregation, the method simulates the dynamic multi-scale visual capabilities of the human visual system. Furthermore, by utilising a frequency-domain attention mechanism to filter out high-frequency noise, the model’s ability to understand complex visual scenes is enhanced, significantly improving the detection accuracy of minute objects in remote sensing images.
2. Materials and Methods
2.1. Lightweight Feature Extraction Module
2.2. Efficient Spatial-Frequency Fusion Module
2.3. Efficent Frequency Domain Awareness Fusion Module
3. Result
3.1. Datasets
- 1.
- VisDrone2019 [33]: This dataset was constructed and released by the AISKYEYE research team at Tianjin University, and is primarily designed for small-object detection tasks from a drone’s perspective. It comprises 288 aerial video clips (totalling 261,908 frames) and 10,209 static images, covering complex scenes such as urban roads and transport hubs, and is characterised by a high density of small and minute objects. The dataset is divided into a training set (6471 images), a validation set (548 images) and a test set (1610 images). Annotated objects include 10 target categories such as pedestrians, crowds, bicycles and cars (“pedestrian”, “people”, “bicycle”, “car”, “van”, “truck”, “tricycle”, “awning-tricycle”, “bus”, “motor”).
- 2.
- NWPU VHR-10 [34]: This dataset was compiled by Northwestern Polytechnical University and comprises a total of 800 high-resolution remote sensing images, with image dimensions ranging from 500 × 500 pixels to 1100 × 1100 pixels. Of these, 650 images contain targets to be detected, whilst the remaining 150 are background images without targets. The dataset features annotations for 10 typical object classes, namely: aeroplane (AP), ship (SP), storage tank (SK), baseball diamond (BD), tennis court (TC), basketball court (BC), ground track field (GTF), harbour (HB), bridge (BR) and vehicle (VE).
- 3.
- DIOR [35]: This dataset is a large-scale benchmark for object detection in optical remote sensing images, comprising 23,463 images and 192,472 object instances annotated with horizontal bounding boxes. The images in this dataset are split into a training set (16,424 images), a validation set (2346 images) and a test set (4693 images) in a 7:1:2 ratio. It covers 20 object classes: airplane (AP), airport (AT), baseball field (BF), basketball court (BC), bridge (BG), chimney (CM), dam (DM), expressway service area (EA), expressway toll station (ES), golf field (GF), ground track field (GD), harbour (HB), overpass (OP), ship (SP), stadium (SD), storage tank (ST), tennis court (TC), train station (TS), vehicle (VE), and windmill (WM). The images in the dataset are sourced from Google Earth, with a resolution of 800 × 800 and a spatial resolution ranging from 0.5 m to 30 m.
3.2. Test Environment and Evaluation Criteria
3.3. Analysis of Ablation Experiment Results
3.4. Visual Analysis of the Le-Detr Performance
3.4.1. Heatmap Comparison Analysis
3.4.2. Comparison of Detection Results in Different Scenarios
3.4.3. Comparison of Feature Map Visualisation Results
3.5. Experimental Results and Analysis Comparing LE-DETR with Other Models
3.5.1. Quantitative Analysis of Tiny Object Detection
3.5.2. Comparison on the Visdrone2019 Dataset
3.5.3. Comparison on the Nwpu Vhr-10 Dataset
3.5.4. Comparison on the Dior Dataset
4. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Aasen, H.; Honkavaara, E.; Lucieer, A.; Zarco-Tejada, P.J. Quantitative Remote Sensing at Ultra-High Resolution with UAV Spectroscopy: A Review of Sensor Technology, Measurement Procedures, and Data Correction Workflows. Remote Sens. 2018, 10, 1091. [Google Scholar] [CrossRef]
- Khanal, S.; KC, K.; Fulton, J.P.; Shearer, S.; Ozkan, E. Remote Sensing in Agriculture—Accomplishments, Limitations, and Opportunities. Remote Sens. 2020, 12, 3783. [Google Scholar] [CrossRef]
- Halder, B.; Bandyopadhyay, J.; Banik, P. Monitoring the effect of urban development on urban heat island based on remote sensing and geo-spatial approach in Kolkata and adjacent areas, India. Sustain. Cities Soc. 2021, 74, 103186. [Google Scholar] [CrossRef]
- Casagli, N.; Intrieri, E.; Tofani, V.; Gigli, G.; Raspini, F. Landslide detection, monitoring and prediction with remote-sensing techniques. Nat. Rev. Earth Environ. 2023, 4, 51–64. [Google Scholar] [CrossRef]
- Sun, Y.; Wang, D.; Li, L.; Ning, R.; Yu, S.; Gao, N. Application of remote sensing technology in water quality monitoring: From traditional approaches to artificial intelligence. Water Res. 2024, 267, 122546. [Google Scholar] [CrossRef] [PubMed]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 39, 1137–1149. [Google Scholar] [CrossRef] [PubMed]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2961–2969. [Google Scholar] [CrossRef]
- Cai, Z.; Vasconcelos, N. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 6154–6162. [Google Scholar] [CrossRef]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. Ssd: Single shot multibox detector. In Proceedings of the Computer Vision—ECCV 2016, Amsterdam, The Netherlands, 11–14 October 2016; pp. 21–37. [Google Scholar] [CrossRef]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2980–2988. [Google Scholar] [CrossRef]
- Jocher, G.; Chaurasia, A.; Stoken, A.; Borovec, J.; Kwon, Y.; Michael, K.; Fang, J.; Yifu, Z.; Wong, C.; Montes, D.; et al. ultralytics/yolov5: v7.0—YOLOv5 SOTA Realtime Instance Segmentation. Zenodo 2022. [Google Scholar] [CrossRef]
- Li, C.; Li, L.; Jiang, H.; Weng, K.; Geng, Y.; Li, L.; Ke, Z.; Li, Q.; Cheng, M.; Nie, W.; et al. YOLOv6: A single-stage object detection framework for industrial applications. arXiv 2022, arXiv:2209.02976. [Google Scholar] [CrossRef]
- Wang, C.Y.; Bochkovskiy, A.; Liao, H.Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 7464–7475. [Google Scholar] [CrossRef]
- Jocher, G.; Chaurasia, A.; Qiu, J. YOLO by Ultralytics. 2023. Available online: https://github.com/ultralytics/ultralytics (accessed on 2 April 2026).
- Wang, C.Y.; Yeh, I.H.; Liao, H.Y.M. Yolov9: Learning what you want to learn using programmable gradient information. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 1–21. [Google Scholar] [CrossRef]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. Yolov10: Real-time end-to-end object detection. arXiv 2024, arXiv:2405.14458. [Google Scholar] [CrossRef]
- Khanam, R.; Hussain, M. Yolov11: An overview of the key architectural enhancements. arXiv 2024, arXiv:2410.17725. [Google Scholar] [CrossRef]
- Tian, Y.; Ye, Q.; Doermann, D. Yolov12: Attention-centric real-time object detectors. arXiv 2025, arXiv:2502.12524. [Google Scholar] [CrossRef]
- Ge, Z.; Liu, S.; Wang, F.; Li, Z.; Sun, J. Yolox: Exceeding yolo series in 2021. arXiv 2021, arXiv:2107.08430. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Long Beach, CA, USA, 4–9 December 2017; pp. 5998–6008. [Google Scholar] [CrossRef]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 213–229. [Google Scholar] [CrossRef]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. Detrs beat yolos on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 16965–16974. [Google Scholar] [CrossRef]
- Wang, X.; Chen, H. HPS-DETR: Enhancing small object detection with lightweight feature extraction and transformer integration. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5602115. [Google Scholar] [CrossRef]
- Zhi, Y.; Zhao, J.; Song, C.; Ma, M.; Mei, S. SAFF-DETR: An End-to-End Object Detection Network for Remote Sensing Images With Targets of Varying Sizes Based on Scale Adaptation and Frequency Fusion. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5601216. [Google Scholar] [CrossRef]
- Xu, Y.; Qi, Q.; He, W.; Zhang, G.; Chen, S.; Tu, B. GSINet: Gradual semantic interaction network for remote sensing object detection based on dual attention mechanism. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 2500045. [Google Scholar] [CrossRef]
- Liao, P.; Zhang, X.; Chen, G.; Wang, T.; Li, X.; Yang, H.; Zhou, W.; He, C.; Wang, Q. S 2 Net: A multitask learning network for semantic stereo of satellite image pairs. IEEE Trans. Geosci. Remote Sens. 2023, 62, 5601313. [Google Scholar] [CrossRef]
- Yang, M.; Xu, R.; Yang, C.; Wu, H.; Wang, A. Hybrid-DETR: A differentiated module-based model for object detection in remote sensing images. Electronics 2024, 13, 5014. [Google Scholar] [CrossRef]
- Chen, J.; Kao, S.H.; He, H.; Zhuo, W.; Wen, S.; Lee, C.H.; Chan, S.H.G. Run, don’t walk: Chasing higher FLOPS for faster neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; pp. 12021–12031. [Google Scholar]
- Chollet, F. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1251–1258. [Google Scholar] [CrossRef]
- Lin, T.Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 2117–2125. [Google Scholar] [CrossRef]
- Sunkara, R.; Luo, T. No more strided convolutions or pooling: A new CNN building block for low-resolution images and small objects. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases (ECML PKDD), Grenoble, France, 19–23 September 2022; pp. 443–459. [Google Scholar] [CrossRef]
- Wang, A.; Chen, H.; Lin, Z.; Han, J.; Ding, G. Lsnet: See large, focus small. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 11–15 June 2025; pp. 9718–9729. [Google Scholar] [CrossRef]
- Du, D.; Zhu, P.; Wen, L.; Bian, X.; Lin, H.; Hu, Q.; Peng, T.; Zheng, J.; Wang, X.; Zhang, Y.; et al. VisDrone-DET2019: The vision meets drone object detection in image challenge results. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) Workshops, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 213–226. [Google Scholar] [CrossRef]
- Cheng, G.; Zhou, P.; Han, J. Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2016, 54, 7405–7415. [Google Scholar] [CrossRef]
- Li, K.; Wan, G.; Cheng, G.; Meng, L.; Han, J. Object detection in optical remote sensing images: A survey and a new benchmark. ISPRS J. Photogramm. Remote Sens. 2020, 159, 296–307. [Google Scholar] [CrossRef]
- Chattopadhay, A.; Sarkar, A.; Howlader, P.; Balasubramanian, V.N. Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks. In Proceedings of the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA, 12–15 March 2018; pp. 839–847. [Google Scholar] [CrossRef]
- Feng, C.; Zhong, Y.; Gao, Y.; Scott, M.R.; Huang, W. Tood: Task-aligned one-stage object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 3490–3499. [Google Scholar] [CrossRef]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable detr: Deformable transformers for end-to-end object detection. arXiv 2020, arXiv:2010.04159. [Google Scholar] [CrossRef]
- Tang, S.; Zhang, L.; Liu, X.; Lv, R.; Qin, R. RFHS-RTDETR: Multi-Domain Collaborative Network with Hierarchical Feature Integration for UAV-Based Object Detection. IEEE Access 2025, 13, 12450–12465. [Google Scholar] [CrossRef]
- Zhang, H.; Zhang, H.; Liu, K.; Gan, Z.; Zhu, G.N. UAV-DETR: Efficient end-to-end object detection for unmanned aerial vehicle imagery. In Proceedings of the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Hangzhou, China, 19–23 October 2025; pp. 15143–15149. [Google Scholar] [CrossRef]
- Xu, S.; Wu, Z.; Ke, Z.; Xue, Z.; Shen, M.; Xiao, W. SUPERLIGHT-DETR: A Lightweight DETR Model for Small Object Detection in Remote Sensing. In Proceedings of the 2025 International Conference on Virtual Reality and Visualization (ICVRV), Bogota, Colombia, 19–21 December 2025; pp. 661–666. [Google Scholar]
- Yang, H.; Chen, J.; Li, Z. CSD-DETR: Efficient Prompt-Aware Representation and High-Resolution Fusion Pyramid for Aerial Small Object Detection. In Proceedings of the 2025 6th International Conference on Computer Science and Management Technology, Xiamen, China, 26–28 December 2025; pp. 810–815. [Google Scholar]
- Wang, A.; Xu, Y.; Wang, H.; Wu, Z.; Wei, Z. CDE-DETR: A Real-Time End-To-End High-Resolution Remote Sensing Object Detection Method Based on RT-DETR. In Proceedings of the IGARSS 2024—2024 IEEE International Geoscience and Remote Sensing Symposium, Athens, Greece, 7–12 July 2024; pp. 8090–8094. [Google Scholar] [CrossRef]
- Shi, Y.; Li, J.; Jia, Y.; Hong, Q. LDA-DETR: A lightweight dynamic attention-enhanced DETR for small object detection. PLoS ONE 2026, 21, e0340977. [Google Scholar] [CrossRef] [PubMed]
- Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L.M.; Shum, H.Y. Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv 2022, arXiv:2203.03605. [Google Scholar] [CrossRef]
- Huang, Y.X.; Liu, H.I.; Shuai, H.H.; Cheng, W.H. Dq-detr: Detr with dynamic query for tiny object detection. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 290–305. [Google Scholar] [CrossRef]
- Siddique, A.; Azeem, A.; Yuting, Z.; Li, Y. Dynamic Adaptive Region Transformer for Tiny Object Detection in Remote Sensing. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5602410. [Google Scholar] [CrossRef]











| Dataset | Experiments | LFEM | ESFFM | EFAM-FPN | P (%) | R (%) | (%) | (%) | (M) | (G) |
|---|---|---|---|---|---|---|---|---|---|---|
| Test | A.Baseline | 54.7 | 38.4 | 37.1 | 21.6 | 19.9 | 57.0 | |||
| B | ✓ | 55.8 | 37.8 | 36.6 | 20.9 | 10.5 | 33.5 | |||
| C | ✓ | 55.1 | 38.9 | 37.6 | 21.6 | 19.7 | 56.9 | |||
| D | ✓ | 55.3 | 40.3 | 38.7 | 22.1 | 20.4 | 64.2 | |||
| E | ✓ | ✓ | 55.9 | 39.3 | 38.2 | 21.8 | 10.3 | 33.4 | ||
| F | ✓ | ✓ | ✓ | 56.7 | 41.0 | 39.7 | 22.7 | 11.6 | 48.4 | |
| Val | A.Baseline | 61.8 | 46.1 | 47.4 | 28.6 | 19.9 | 57.0 | |||
| B | ✓ | 61.6 | 45.1 | 47.2 | 28.6 | 10.5 | 33.5 | |||
| C | ✓ | 62.4 | 46.4 | 48.1 | 29.1 | 19.7 | 56.9 | |||
| D | ✓ | 62.8 | 48.0 | 49.4 | 30.1 | 20.4 | 64.2 | |||
| E | ✓ | ✓ | 62.6 | 47.3 | 48.8 | 28.9 | 10.3 | 33.4 | ||
| F | ✓ | ✓ | ✓ | 63.8 | 48.1 | 50.5 | 31.0 | 11.6 | 48.4 |
| Dataset | Model | (%) | (%) | (%) | Params (M) | Gflops (G) | |
|---|---|---|---|---|---|---|---|
| VisDrone2019-val [33] | RT-DETR | 19.0 | 36.5 | 41.4 | 19.9 | 57.0 | 114 |
| YOLOv12-s | 12.5 | 33.5 | 42.0 | 9.2 | 21.2 | 268 | |
| LE-DETR | 20.9 | 37.9 | 42.3 | 11.6 | 48.4 | 149 | |
| NWPU VHR-10-val [34] | RT-DETR | 22.9 | 59.2 | 56.9 | 19.9 | 57.0 | 42 |
| YOLOv12-s | 27.8 | 58.9 | 57.1 | 9.2 | 21.2 | 32 | |
| LE-DETR | 29.0 | 62.1 | 62.3 | 11.6 | 48.4 | 49 | |
| DIOR-test [35] | RT-DETR | 25.1 | 44.4 | 77.9 | 19.9 | 57.0 | 160 |
| YOLOv12-s | 22.6 | 43.3 | 79.2 | 9.2 | 21.2 | 175 | |
| LE-DETR | 28.0 | 45.6 | 79.8 | 11.6 | 48.4 | 179 |
| Models | (%) | (%) | (%) | (%) | (M) | (G) | |
|---|---|---|---|---|---|---|---|
| One-stage models | |||||||
| SSD [9] | 21.3 | 35.4 | 24.1 | 10.7 | 13.3 | 22.8 | - |
| RetinaNet [10] | 31.3 | 26.5 | 29.6 | 19.0 | 36.1 | 17.2 | 57 |
| TOOD [37] | 45.8 | 33.4 | 34.8 | 20.5 | 32.0 | 199.0 | 44 |
| YOLOv6-s [12] | 43.1 | 33.0 | 30.9 | 18.0 | 16.3 | 44.2 | 357 |
| YOLOv8-s [14] | 45.0 | 33.8 | 32.0 | 18.6 | 11.1 | 28.8 | 371 |
| YOLOv10-s [16] | 44.8 | 33.7 | 32.2 | 18.7 | 8.07 | 24.8 | 329 |
| YOLOv11-s [17] | 44.9 | 34.4 | 32.6 | 18.9 | 9.4 | 21.6 | 304 |
| YOLOv12-s [18] | 44.8 | 34.3 | 32.9 | 19.2 | 9.2 | 21.2 | 268 |
| YOLOvX [19] | 42.1 | 33.1 | 33.9 | 19.8 | 9.0 | 26.8 | 249 |
| Two-stage models | |||||||
| Faster-R-CNN [6] | 54.6 | 33.9 | 33.4 | 12.2 | 17.3 | 28.0 | 26 |
| End-to-end models | |||||||
| RT-DETR-R18 [22] | 54.7 | 38.4 | 37.1 | 21.6 | 19.9 | 57.0 | 114 |
| RT-DETR-R34 [22] | 56.8 | 40.1 | 38.5 | 22.3 | 31.4 | 90.3 | 103 |
| RT-DETR-R50 [22] | 57.7 | 40.3 | 39 | 22.4 | 41.9 | 129 | 59 |
| Deformable-DETR [38] | 54.9 | 41.0 | 34.9 | 20.2 | 40.0 | 173.1 | 87 |
| RFHS-RTDETR [39] | 56.5 | 39.8 | 39.1 | 22.9 | 13.0 | 48.0 | 76 |
| UAV-DETR [40] | 57.0 | 40.8 | 39.5 | 22.9 | 21.26 | 72.5 | 62 |
| SUPERLIGHT-DETR [41] | 46.8 | 36.1 | 34.7 | - | 7.5 | 15.2 | - |
| CSD-DETR [42] | - | - | 40.7 | 24.0 | 14.82 | 65.8 | 65.8 |
| LE-DETR(Ours) | 56.7 | 41.0 | 39.7 | 22.7 | 11.6 | 48.4 | 149 |
| Models | AP (%) | (%) | (%) | (M) | (G) | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AP | SP | ST | BD | TC | BC | GTF | HB | BR | VE | ||||||
| SSD [9] | 98.1 | 71.3 | 80.3 | 88.6 | 89.1 | 70.9 | 99.5 | 85.8 | 63.7 | 54.9 | 80.3 | 49.7 | 13.3 | 22.8 | - |
| YOLOv6-s [12] | 99.3 | 75.4 | 96.5 | 99.0 | 99.5 | 77.0 | 99.2 | 94.6 | 99.5 | 88.0 | 92.8 | 60.5 | 16.3 | 44.2 | 41 |
| YOLOv8-s [14] | 98.4 | 77.7 | 98.5 | 97.8 | 98.5 | 80.4 | 98.5 | 92.9 | 98.3 | 89.0 | 92.9 | 61.0 | 11.1 | 28.8 | 33 |
| YOLOv10-s [16] | 98.5 | 77.1 | 99.1 | 96.6 | 99.1 | 82.8 | 93.9 | 83.9 | 99.5 | 91.0 | 92.1 | 60.4 | 8.07 | 24.8 | 38 |
| YOLOv11-s [17] | 99.3 | 79.2 | 98.9 | 99.0 | 99.5 | 87.2 | 99.5 | 87.1 | 56.3 | 92.0 | 89.9 | 57.9 | 9.4 | 21.6 | 29 |
| YOLOv12-s [18] | 99.2 | 75.0 | 96.5 | 98.3 | 99.4 | 81.2 | 99.2 | 86.1 | 75.2 | 89.7 | 90.0 | 58.8 | 9.2 | 21.2 | 32 |
| SAFF-DETR [24] | 99.3 | 81.9 | 92.2 | 98.9 | 96.0 | 94.3 | 99.4 | 87.3 | 98.6 | 88.1 | 93.3 | 61.5 | 20.6 | 78.5 | - |
| CDE-DETR [43] | 99.2 | 92.5 | 72.5 | 99.3 | 94.6 | 91.5 | 99.5 | 92.2 | 87.2 | 91.0 | 92.0 | - | 18.1 | 49.2 | 69.2 |
| LDA-DETR [44] | 99.5 | 84.1 | 98.9 | 99.1 | 90.4 | 98.0 | 100 | 70.2 | 85.5 | 88.3 | 91.4 | - | 16.93 | 49.7 | 65.6 |
| DINO [45] | 99.2 | 78.3 | 92.9 | 95.6 | 89.8 | 92.1 | 99.9 | 85.4 | 94.6 | 88.5 | 92.2 | 59.1 | 47.5 | 265 | - |
| RT-DETR-R18 [22] | 98.4 | 75.8 | 95.4 | 97.2 | 99.4 | 78.6 | 96.5 | 91.1 | 99.5 | 86.4 | 91.8 | 59.2 | 19.9 | 57.0 | 42 |
| LE-DETR(Ours) | 99.5 | 73.1 | 97.9 | 98.2 | 99.5 | 85.8 | 99.5 | 89.6 | 99.5 | 92.7 | 93.5 | 63.1 | 11.6 | 48.4 | 49 |
| Models | AP (%) | (%) | (M) | (G) | ||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| AP | AT | BF | BC | BG | CM | DM | EA | ES | GF | GD | HB | OP | SP | SD | ST | TC | TS | VE | WM | |||||
| RetinaNet [10] | 61.5 | 70.3 | 76.0 | 86.5 | 40.0 | 75.6 | 61.6 | 75.8 | 68.2 | 77.4 | 78.1 | 36.1 | 54.3 | 73.0 | 59.1 | 61.8 | 87.4 | 39.3 | 43.3 | 86.2 | 64.7 | 36.1 | 17.2 | - |
| Faster-RCNN [6] | 54.9 | 81.4 | 70.1 | 88.5 | 39.3 | 79.3 | 60.6 | 77.8 | 59.8 | 78.2 | 83.1 | 52.9 | 57.1 | 68.4 | 74.5 | 52.6 | 84.6 | 59.7 | 36.0 | 78.9 | 68.9 | 17.1 | 27.8 | - |
| YOLOv6-s [12] | 96.9 | 91.6 | 93.7 | 88.4 | 57.2 | 84.3 | 85.8 | 95.8 | 81.6 | 85.9 | 88.7 | 76.1 | 69.1 | 92.2 | 95.4 | 88.9 | 96.7 | 72.2 | 65.5 | 91.4 | 84.9 | 16.3 | 44.1 | 133 |
| YOLOv8-s [14] | 96.6 | 91.5 | 94.4 | 89.2 | 59.8 | 87.8 | 83.7 | 95.7 | 83.7 | 84.5 | 89.4 | 76.1 | 72.1 | 92.6 | 95.7 | 90.9 | 95.9 | 71.4 | 68.2 | 92.8 | 85.6 | 11.1 | 28.5 | 137 |
| YOLOv10-s [16] | 96.4 | 91.6 | 93.9 | 88.8 | 59.1 | 86.6 | 88.3 | 95.6 | 84.6 | 85.9 | 88.7 | 73.2 | 70.2 | 91.9 | 94.7 | 90.4 | 95.5 | 68.4 | 67.9 | 90.3 | 85.1 | 8.1 | 24.5 | 236 |
| YOLOv11-s [17] | 97.9 | 93.3 | 94.8 | 90.2 | 60.7 | 87.8 | 85.8 | 97.3 | 83.2 | 86.5 | 90.2 | 75.6 | 71.1 | 93.2 | 97.1 | 91.3 | 96.9 | 68.5 | 69.0 | 92.9 | 86.2 | 9.4 | 21.3 | 144 |
| YOLOv12-s [18] | 97.7 | 93.0 | 95.0 | 89.5 | 61.2 | 86.9 | 89.5 | 96.9 | 83.4 | 88.2 | 90.0 | 76.6 | 72.2 | 93.0 | 96.7 | 91.1 | 96.8 | 72.7 | 68.8 | 92.5 | 86.6 | 9.2 | 21.3 | 175 |
| Net [26] | 92.7 | 89.9 | 90.1 | 91.0 | 58.5 | 82.1 | 76.3 | 93.1 | 84.4 | 84.6 | 87.5 | 67.6 | 68.6 | 79.1 | 83.4 | 78.5 | 92.9 | 74.5 | 61.0 | 93.3 | 81.4 | 19.6 | 59.1 | - |
| DQ-DETR [46] | 82.4 | 81.3 | 90.4 | 79.7 | 49.3 | 87.2 | 66.9 | 78.3 | 69.8 | 73.4 | 76.7 | 67.9 | 58.5 | 87.4 | 80.8 | 80.7 | 86.4 | 63.8 | 72.4 | 75.3 | 75.4 | 44.9 | - | - |
| DART [47] | 93.3 | 86.4 | 93.4 | 85.2 | 51.2 | 91.1 | 70.5 | 83.1 | 72.3 | 78.8 | 82.6 | 69.6 | 64.0 | 85.9 | 84.3 | 86.4 | 91.6 | 64.2 | 75.1 | 80.6 | 79.5 | 13.0 | 68.0 | - |
| RT-DETR-R18 [22] | 96.6 | 93.4 | 94.1 | 89.1 | 59.1 | 85.5 | 87.1 | 94.7 | 88.9 | 79.5 | 88.7 | 70.1 | 69.1 | 91.4 | 93.1 | 90.4 | 95.5 | 74.4 | 73.8 | 93.5 | 85.4 | 19.9 | 57.0 | 160 |
| LE-DETR(Ours) | 98.5 | 95.0 | 95.4 | 91.4 | 62.2 | 88.6 | 86.8 | 98.1 | 91.6 | 88.5 | 91.3 | 73.3 | 72.0 | 93.6 | 92.4 | 92.8 | 97.0 | 75.0 | 76.7 | 95.5 | 87.8 | 11.6 | 48.4 | 179 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, Q.; An, H.; Chen, Y. LE-DETR: A Lightweight and Efficient Model for Small-Object Detection in Remote Sensing Images. Remote Sens. 2026, 18, 2018. https://doi.org/10.3390/rs18122018
Wang Q, An H, Chen Y. LE-DETR: A Lightweight and Efficient Model for Small-Object Detection in Remote Sensing Images. Remote Sensing. 2026; 18(12):2018. https://doi.org/10.3390/rs18122018
Chicago/Turabian StyleWang, Qi, Hongyun An, and Yongji Chen. 2026. "LE-DETR: A Lightweight and Efficient Model for Small-Object Detection in Remote Sensing Images" Remote Sensing 18, no. 12: 2018. https://doi.org/10.3390/rs18122018
APA StyleWang, Q., An, H., & Chen, Y. (2026). LE-DETR: A Lightweight and Efficient Model for Small-Object Detection in Remote Sensing Images. Remote Sensing, 18(12), 2018. https://doi.org/10.3390/rs18122018

