HD-BSNet: A Plug-and-Play Dual-Mechanism Synergistic Enhancement Framework for Small Object Detection
Highlights
- A plug-and-play framework HD-BSNet for small object detection is proposed, which integrates three core modules: high-frequency differential perception, background semantic modeling, and parallel multi-scale focusing, to address the inherent limitations of small object detection in a targeted manner.
- Experimental results on three mainstream datasets (AI-TOD, VisDrone, and DUT Anti-UAV) demonstrate that the framework can significantly improve the detection accuracy of small objects in complex background environments, while effectively reducing the false detection rate.
- It precisely addresses the core pain points of existing small object detection methods—including insufficient detection accuracy for targets smaller than 20 pixels, high false detection rates of GAN-based enhancement methods, and poor framework adaptability—thus providing reliable technical support for high-precision detection in fields such as remote sensing and low-altitude security.
- The plug-and-play design reduces the costs of technical implementation. Its dual-mechanism collaborative enhancement and multi-scale focusing strategy provides a new paradigm for research on feature extraction and background suppression in the field of small object detection, facilitating the transformation and deployment of related technologies in practical scenarios.
Abstract
1. Introduction
- (1)
- We introduce HD-BSNet, a plug-and-play dual-mechanism collaborative object detection framework. This framework can be directly embedded not only into both one-stage and two-stage detection models but also is suitable for remote sensing and low-altitude UAV scenarios, and it can effectively address the issues of accuracy gaps, elevated false detection rates, and poor framework adaptability for small objects.
- (2)
- A dual-mechanism collaborative feature enhancement scheme for small objects is designed: it directly locates high-frequency information loss regions and captures features via differential operations to strengthen detailed representation; meanwhile, it integrates background semantic modeling to learn the dependencies between targets and backgrounds, thereby extracting salient features of small objects.
- (3)
- A parallel multi-scale focusing module is proposed. Through multi-scale window partitioning and normalized focusing, the module enhances the features of small objects, indirectly suppresses interference from complex backgrounds, and improves detection performance.
2. Related Work
2.1. Generic Object Detection
2.2. Small Object Detection
3. Method
3.1. Overview
3.2. High-Frequency Differential Perception Module
3.3. Background Semantic Modeling Module
3.4. Parallel Multi-Scale Focusing Module
4. Experiments
4.1. Experimental Setup
4.2. Comparative Experiments
4.3. Ablation Experiments
5. Discussion
5.1. Limitations and Causes of Detection Performance for Verytiny Targets
5.2. Conflict Between Computational Complexity and Real-Time Deployment and Its Causes
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Ren, Z.; He, L.; Lu, J. Context aware edge-enhanced GAN for remote sensing image super-resolution. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2023, 17, 1363–1376. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Y.; Huang, Y.; Yang, M.; Mao, D.; Zhang, Y.; Jiao, L.; Zhang, Y.; Yang, J. SAR Image Super-resolution based on Multi-scale Edge Texture-oriented GAN Approach. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 20359–20374. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Wei, S.; Sun, Y.; Shen, J.; Yang, Z.; Yan, J. EESAGAN: Edge-Enhanced and Structure-Aware GAN for Remote Sensing Image Super-Resolution. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 24947–24962. [Google Scholar] [CrossRef] [Scilit]
- Yi, H.; Liu, B.; Zhao, B.; Liu, E. Small object detection algorithm based on improved YOLOv8 for remote sensing. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2023, 17, 1734–1747. [Google Scholar] [CrossRef] [Scilit]
- Zheng, X.; Qiu, Y.; Zhang, G.; Lei, T.; Jiang, P. ESL-YOLO: Small object detection with effective feature enhancement and spatial-context-guided fusion network for remote sensing. Remote Sens. 2024, 16, 4374. [Google Scholar] [CrossRef] [Scilit]
- Fu, C.; Yuan, H.; Shen, L.; Hamzaoui, R.; Zhang, H. 3DAttGAN: A 3D attention-based generative adversarial network for joint space-time video super-resolution. IEEE Trans. Emerg. Top. Comput. Intell. 2024, 8, 3117–3128. [Google Scholar] [CrossRef] [Scilit]
- Cao, B.; Yao, H.; Zhu, P.; Hu, Q. Visible and clear: Finding tiny objects in difference map. In Computer Vision–ECCV 2024. ECCV 2024. Lecture Notes in Computer Science; Springer Nature: Cham, Switzerland, 2024; pp. 1–18. [Google Scholar]
- Li, Y.; Luo, J.; Zhang, Y.; Tan, Y.; Yu, J.-G.; Bai, S. Learning to holistically detect bridges from large-size VHR remote sensing imagery. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 11507–11523. [Google Scholar] [CrossRef] [Scilit]
- Xu, C.; Wang, J.; Yang, W.; Yu, L. Dot distance for tiny object detection in aerial images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 2021, Nashville, TN, USA, 20–25 June 2021; pp. 1192–1201. [Google Scholar]
- Xu, C.; Wang, J.; Yang, W.; Yu, H.; Yu, L.; Xia, G.-S. Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark. ISPRS J. Photogramm. Remote Sens. 2022, 190, 79–93. [Google Scholar] [CrossRef] [Scilit]
- Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2117–2125. [Google Scholar]
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path aggregation network for instance segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 8759–8768. [Google Scholar]
- Ge, L.; Wang, G.; Zhang, T.; Zhuang, Y.; Chen, H.; Dong, H.; Chen, L. Regression-guided refocusing learning with feature alignment for remote sensing tiny object detection. IEEE Trans. Geosci. Remote Sens. 2024, 62, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Liu, D.; Zhang, J.; Qi, Y.; Wu, Y.; Zhang, Y. Tiny object detection in remote sensing images based on object reconstruction and multiple receptive field adaptive feature enhancement. IEEE Trans. Geosci. Remote Sens. 2024, 62, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Qiao, S.; Chen, L.C.; Yuille, A. Detectors: Detecting objects with recursive feature pyramid and switchable atrous convolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 10213–10224. [Google Scholar]
- Li, Y.; Wang, L.; Wang, T.; Yang, X.; Luo, J.; Wang, Q.; Deng, Y.; Wang, W.; Sun, X.; Li, H.; et al. STAR: A first-ever dataset and a large-scale benchmark for scene graph generation in large-size satellite imagery. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 47, 1832–1849. [Google Scholar] [CrossRef] [Scilit]
- Deng, C.; Wang, M.; Liu, L.; Liu, Y.; Jiang, Y. Extended feature pyramid network for small object detection. IEEE Trans. Multimed. 2021, 24, 1968–1979. [Google Scholar] [CrossRef] [Scilit]
- Girshick, R. Fast r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 7–13 December 2015; pp. 1440–1448. [Google Scholar]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2961–2969. [Google Scholar]
- Cai, Z.; Vasconcelos, N. Cascade r-cnn: Delving into high quality object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 6154–6162. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 39, 1137–1149. [Google Scholar] [CrossRef] [Scilit]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 580–587. [Google Scholar]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, 11–14 October 2016; Springer International Publishing: Cham, Switzerland, 2016; pp. 21–37. [Google Scholar]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2980–2988. [Google Scholar]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
- Duan, K.; Bai, S.; Xie, L.; Qi, H.; Huang, Q.; Tian, Q. Centernet: Keypoint triplets for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6569–6578. [Google Scholar]
- Kong, T.; Sun, F.; Liu, H.; Jiang, Y.; Li, L.; Shi, J. Foveabox: Beyound anchor-based object detection. IEEE Trans. Image Process. 2020, 29, 7389–7398. [Google Scholar] [CrossRef] [Scilit]
- Tian, Z.; Shen, C.; Chen, H.; He, T. Fcos: Fully convolutional one-stage object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9627–9636. [Google Scholar]
- Zhou, Z.; Zhu, Y. KLDet: Detecting tiny objects in remote sensing images via Kullback–Leibler divergence. IEEE Trans. Geosci. Remote Sens. 2024, 62, 1–16. [Google Scholar] [CrossRef] [Scilit]
- Cheng, G.; Yuan, X.; Yao, X.; Yan, K.; Zeng, Q.; Xie, X.; Han, J. Towards large-scale small object detection: Survey and benchmarks. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 13467–13488. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Shim, S.-H.; Hyun, S.; Bae, D.; Heo, J.-P. Local attention pyramid for scene image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 7774–7782. [Google Scholar]
- Wang, J.; Yang, W.; Guo, H.; Zhang, R.; Xia, G.-S. Tiny object detection in aerial images. In Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, 10–15 January 2021; pp. 3791–3798. [Google Scholar]
- Du, D.; Zhu, P.; Wen, L.; Bian, X.; Lin, H.; Hu, Q.; Peng, T.; Zheng, J.; Wang, X.; Zhang, Y.; et al. VisDrone-DET2019: The vision meets drone object detection in image challenge results. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, Seoul, Republic of Korea, 27 October–2 November 2019. [Google Scholar]
- Zhao, J.; Zhang, J.; Li, D.; Wang, D. Vision-based anti-uav detection and tracking. IEEE Trans. Intell. Transp. Syst. 2022, 23, 25323–25334. [Google Scholar] [CrossRef] [Scilit]
- Xia, G.-S.; Bai, X.; Ding, J.; Zhu, Z.; Belongie, S.; Luo, J.; Datcu, M.; Pelillo, M.; Zhang, L. DOTA: A large-scale dataset for object detection in aerial images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 3974–3983. [Google Scholar]
- Li, K.; Wan, G.; Cheng, G.; Meng, L.; Han, J. Object detection in optical remote sensing images: A survey and a new benchmark. ISPRS J. Photogramm. Remote Sens. 2020, 159, 296–307. [Google Scholar] [CrossRef] [Scilit]
- Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision, Zurich, Switzerland, 6–12 September 2014; Springer International Publishing: Cham, Switzerland, 2014; pp. 740–755. [Google Scholar]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. Pytorch: An imperative style, high-performance deep learning library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2019; Available online: https://dl.acm.org/doi/10.5555/3454287.3455008 (accessed on 22 January 2026).
- Chen, K.; Wang, J.; Pang, J.; Cao, Y.; Xiong, Y.; Li, X.; Sun, S.; Feng, W.; Liu, Z.; Xu, J.; et al. MMDetection: Open mmlab detection toolbox and benchmark. arXiv 2019, arXiv:1906.07155. [Google Scholar] [CrossRef] [Scilit]
- Xu, C.; Wang, J.; Yang, W.; Yu, H.; Yu, L.; Xia, G.-S. RFLA: Gaussian receptive field based label assignment for tiny object detection. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–29 October 2022; Springer Nature: Cham, Switzerland, 2022; pp. 526–543. [Google Scholar]
- Shi, Z.; Hu, J.; Ren, J.; Ye, H.; Yuan, X.; Ouyang, Y.; He, J.; Ji, B.; Guo, J. HS-FPN: High frequency and spatial perception FPN for tiny object detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 22–28 February 2025; Volume 39, pp. 6896–6904. [Google Scholar]
- Liu, H.-I.; Tseng, Y.-W.; Chang, K.-C.; Wang, P.-J.; Shuai, H.-H.; Cheng, W.-H. A denoising fpn with transformer r-cnn for tiny object detection. IEEE Trans. Geosci. Remote Sens. 2024, 62, 4704415. [Google Scholar] [CrossRef] [Scilit]
- Hu, H.; Chen, S.B.; Tang, J. CFENet: Contextual Feature Enhancement Network for Tiny Object Detection in Aerial Images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4703113. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Farhadi, A. Yolov3: An incremental improvement. arXiv 2018, arXiv:1804.02767. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Chen, Y.; Wang, N.; Zhang, Z. Scale-aware trident networks for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6054–6063. [Google Scholar]
- Zhu, B.; Wang, J.; Jiang, Z.; Zong, F.; Liu, S.; Li, Z.; Sun, J. Autoassign: Differentiable label assignment for dense object detection. arXiv 2020, arXiv:2007.03496. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Chi, C.; Yao, Y.; Lei, Z.; Li, S.Z. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 9759–9768. [Google Scholar]
- Feng, C.; Zhong, Y.; Gao, Y.; Scott, M.R.; Huang, W. Tood: Task-aligned one-stage object detection. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; IEEE Computer Society: Washington, DC, USA, 2021; pp. 3490–3499. [Google Scholar]
- Wang, J.; Xu, C.; Yang, W.; Yu, L. A normalized Gaussian Wasserstein distance for tiny object detection. arXiv 2021, arXiv:2110.13389. [Google Scholar]
- Guo, G.; Chen, P.; Yu, X.; Han, Z.; Ye, Q.; Gao, S. Save the tiny, save the all: Hierarchical activation network for tiny object detection. IEEE Trans. Circuits Syst. Video Technol. 2023, 34, 221–234. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L.M.; Shum, H.-Y. Dino: Detr with improved denoising anchor boxes for end-to-end object detection. arXiv 2022, arXiv:2203.03605. [Google Scholar]
- Zhang, T.; Zhang, X.; Zhu, X.; Wang, G.; Han, X.; Tang, X.; Jiao, L. Multistage enhancement network for tiny object detection in remote sensing images. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5611512. [Google Scholar] [CrossRef] [Scilit]
- Chen, Q.; Wang, Y.; Yang, T.; Zhang, X.; Cheng, J.; Sun, J. You only look one-level feature. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 13039–13048. [Google Scholar]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. Detrs beat yolos on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–20 June 2024; pp. 16965–16974. [Google Scholar]
- Cheng, S.; Song, J.; Zhou, M.; Wei, X.; Pu, H.; Luo, J.; Jia, W. Ef-detr: A lightweight transformer-based object detector with an encoder-free neck. IEEE Trans. Ind. Inform. 2024, 20, 12994–13002. [Google Scholar] [CrossRef] [Scilit]
- Jocher, G.; Chaurasia, A.; Stoken, A.; Borovec, J.; Kwon, Y.; Fang, J.; Michael, K.; Montes, D.; Nadar, J.; Skalski, P.; et al. ultralytics/yolov5: V6. 1-Tensorrt, Tensorflow Edge Tpu and Openvino Export and Inference. Zenodo, 2022. Available online: https://zenodo.org/records/6222936 (accessed on 22 January 2026). [CrossRef]
- Li, C.; Li, L.; Jiang, H.; Weng, K.; Geng, Y.; Li, L.; Ke, Z.; Li, Q.; Cheng, M.; Nie, W.; et al. YOLOv6: A single-stage object detection framework for industrial applications. arXiv 2022, arXiv:2209.02976. [Google Scholar] [CrossRef] [Scilit]
- Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLO, Ver. 8.0.0, January 2023. Available online: https://github.com/ultralytics/ultralytics (accessed on 22 January 2026).
- Wang, C.Y.; Yeh, I.H.; Mark Liao, H.Y. Yolov9: Learning what you want to learn using programmable gradient information. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–29 October 2024; Springer Nature: Cham, Switzerland, 2024; pp. 1–21. [Google Scholar]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. Yolov10: Real-time end-to-end object detection. Adv. Neural Inf. Process. Syst. 2024, 37, 107984–108011. [Google Scholar]
- Khanam, R.; Hussain, M. Yolov11: An overview of the key architectural enhancements. arXiv 2024, arXiv:2410.17725. [Google Scholar] [CrossRef] [Scilit]









| Method | Venue | Backbone | AP (%) | AP50 (%) | AP75 (%) | (%) | (%) | (%) | (%) |
|---|---|---|---|---|---|---|---|---|---|
| SSD-512 [23] | ECCV2016 | VGG-16 | 6.5 | 21.0 | 2.4 | 1.0 | 5.3 | 11.9 | 21.2 |
| Faster R-CNN [21] | TPAMI2017 | ResNet-50 | 9.7 | 22.1 | 6.8 | 0.0 | 5.2 | 21.7 | 32.5 |
| RetinaNet [24] | ICCV2017 | ResNet-50 | 2.4 | 8.4 | 0.7 | 0.8 | 2.7 | 4.0 | 8.8 |
| Mask R-CNN [19] | ICCV2017 | ResNet-50 | 10.1 | 22.7 | 7.6 | 0.0 | 5.8 | 22.1 | 32.6 |
| PANet [12] | CVPR2018 | ResNet-50 | 9.7 | 22.2 | 6.8 | 0.0 | 5.4 | 21.6 | 32.5 |
| Cascade R-CNN [20] | CVPR2018 | ResNet-50 | 11.8 | 25.6 | 9.5 | 0.1 | 7.9 | 23.0 | 34.8 |
| YOLOv3 [45] | arXiv2018 | DarkNet-53 | 12.3 | 25.3 | 5.6 | 3.2 | 12.0 | 18.5 | 24.2 |
| TridentNet [46] | ICCV2019 | ResNet-50 | 4.9 | 13.1 | 2.7 | 0.0 | 2.8 | 10.8 | 19.2 |
| CenterNet [26] | ICCV2019 | ResNet-18 | 8.2 | 24.3 | 3.4 | 1.3 | 6.9 | 12.9 | 22.3 |
| AutoAssign [47] | arXiv2020 | ResNet-50 | 13.4 | 36.2 | 6.9 | 3.5 | 13.5 | 17.0 | 24.9 |
| ATSS [48] | CVPR2020 | ResNet-50 | 8.9 | 22.5 | 5.3 | 2.1 | 10.1 | 11.4 | 15.0 |
| TOOD [49] | ICCV2021 | ResNet-50 | 14.9 | 34.7 | 10.6 | 3.3 | 12.8 | 21.7 | 33.1 |
| M-CenterNet [33] | ICPR2021 | DLA-34 | 14.5 | 40.7 | 6.4 | 6.1 | 15.0 | 19.4 | 20.4 |
| NWD [50] | arXiv2021 | ResNet-50 | 20.8 | 49.3 | 14.3 | 6.4 | 19.7 | 29.6 | 38.3 |
| RFLA [41] | ECCV2022 | ResNet-50 | 24.8 | 55.2 | 18.5 | 9.3 | 24.8 | 30.3 | 38.2 |
| NWD-RKA [10] | ISPRS2022 | ResNet-50 | 23.4 | 53.5 | 16.8 | 8.7 | 23.8 | 28.5 | 36.0 |
| HANet [51] | TCSVT2023 | ResNet-50 | 22.1 | 53.7 | 14.4 | 10.9 | 22.2 | 27.3 | 36.8 |
| DINO [52] | ICLR2023 | ResNet-50 | 23.2 | 56.6 | 15.4 | 9.9 | 23.1 | 29.3 | 37.6 |
| Dotd [9] | TPAMI2024 | ResNet-50 | 16.1 | 39.2 | 10.6 | 8.3 | 17.6 | 18.1 | 22.1 |
| MENet [53] | TGRS2024 | Swin-T | 23.2 | 56.2 | 15.0 | 9.7 | 23.9 | 25.3 | 34.4 |
| ORFENet [14] | TGRS2024 | ResNet-50 | 16.6 | 38.3 | 11.6 | 5.9 | 17.6 | 29.9 | 23.0 |
| DNTR [43] | TGRS2024 | ResNet-50 | 24.9 | 54.3 | 18.4 | 11.5 | 25.0 | 29.0 | 35.3 |
| HS-FPN [42] | AAAI2025 | ResNet-50 | 25.1 | 55.7 | 19.1 | 12.1 | 25.3 | 29.9 | 36.9 |
| CFENet [44] | TGRS2025 | CSPDarkNet-s | 26.6 | 59.2 | 19.6 | 9.0 | 27.5 | 32.3 | 38.5 |
| Faster R-CNN w/HD-BSNet | ResNet-50 | 26.8 | 58.6 | 23.3 | 11.0 | 30.5 | 32.9 | 36.9 | |
| Cascade R-CNN w/HD-BSNet | ResNet-50 | 28.2 | 59.1 | 25.2 | 11.9 | 31.5 | 34.5 | 40.3 | |
| DetectoRS w/HD-BSNet | ResNet-50 | 29.8 | 61.7 | 26.9 | 11.5 | 32.7 | 36.9 | 42.7 |
| Method | AP (%) | AP50 (%) | AP75 (%) | (%) | (%) | (%) |
|---|---|---|---|---|---|---|
| YOLOF [54] | 15.1 | 26.3 | 15.4 | 6.1 | 25.2 | 32.4 |
| RT-DETR [55] | 15.9 | 36.5 | 10.7 | 13.8 | 28.4 | 23.1 |
| EF-DETR [56] | 13.2 | 24.2 | 12.3 | 8.1 | 18.4 | 35.7 |
| RetinaNet [24] | 17.6 | 29.6 | 18.1 | 8.1 | 29.4 | 36.8 |
| AutoAssign [47] | 23.2 | 43.5 | 21.8 | 15.2 | 33.3 | 40.9 |
| CenterNet [26] | 21.4 | 36.1 | 21.8 | 12.6 | 31.7 | 38.3 |
| PANet [12] | 23.0 | 38.1 | 24.2 | 13.9 | 34.6 | 39.0 |
| DINO [52] | 26.8 | 44.2 | 28.9 | 17.5 | 37.3 | 41.3 |
| Faster R-CNN [21] | 24.5 | 42.5 | 24.6 | 16.1 | 35.9 | 36.5 |
| Cascade R-CNN [20] | 25.6 | 43.1 | 26.2 | 16.4 | 37.4 | 41.4 |
| DetectoRS [15] | 27.0 | 45.1 | 27.5 | 17.7 | 39.2 | 44.5 |
| Faster R-CNN w/HD-BSNet | 27.0 | 47.9 | 27.1 | 18.4 | 37.9 | 42.2 |
| Cascade R-CNN w/HD-BSNet | 28.6 | 48.6 | 29.0 | 19.8 | 39.9 | 45.4 |
| DetectoRS w/HD-BSNet | 29.2 | 49.6 | 29.7 | 20.2 | 40.8 | 49.7 |
| Method | P | R | AP50 | AP50–95 | GFLOPs | Params (M) |
|---|---|---|---|---|---|---|
| YOLOv5 [57] | 93.4 | 79.6 | 88.1 | 57.1 | 7.1 | 2.5 |
| YOLOv6 [58] | 91.9 | 82.5 | 89.8 | 59.1 | 11.7 | 4.2 |
| YOLOv8 [59] | 94.4 | 82.3 | 90.1 | 60.0 | 8.1 | 3.0 |
| YOLOv9 [60] | 92.8 | 83.1 | 89.7 | 60.1 | 7.6 | 1.9 |
| YOLOv10 [61] | 93.8 | 82.6 | 90.5 | 59.7 | 6.5 | 2.2 |
| YOLOv11 [62] | 93.6 | 82.2 | 89.7 | 59.1 | 6.3 | 2.5 |
| YOLOv11 w/HD-BSNet | 94.5 | 91.8 | 95.1 | 65.8 | 14.6 | 2.7 |
| HDPM | BSMM | PMFM | AP (%) | AP50 (%) | AP75 (%) | (%) | (%) | (%) |
|---|---|---|---|---|---|---|---|---|
| 16.6 | 33.3 | 14.5 | 0.1 | 12.1 | 31.3 | |||
| ✓ | 27.2 | 58.5 | 23.6 | 10.3 | 29.8 | 34.1 | ||
| ✓ | 28.4 | 58.9 | 26.0 | 10.9 | 31.4 | 35.6 | ||
| ✓ | 26.6 | 57.1 | 23.2 | 8.7 | 28.8 | 33.8 | ||
| ✓ | ✓ | 28.9 | 60.2 | 25.9 | 11.1 | 31.4 | 36.4 | |
| ✓ | ✓ | 27.3 | 58.3 | 24.1 | 10.7 | 29.9 | 34.4 | |
| ✓ | ✓ | 29.0 | 60.3 | 26.0 | 11.0 | 31.8 | 36.8 | |
| ✓ | ✓ | ✓ | 29.8 | 61.7 | 26.9 | 11.5 | 32.7 | 36.9 |
| PMFM | LAPM | AP (%) | AP50 (%) | AP75 (%) | (%) | (%) | GFLOPs |
|---|---|---|---|---|---|---|---|
| ✓ | 26.6 | 57.1 | 23.2 | 28.8 | 33.8 | 152 | |
| ✓ | 26.3 | 56.8 | 23.1 | 28.6 | 33.7 | 152 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wen, J.; Zheng, X.; Pan, N.; Jia, D.; Wu, H.; Chen, T.; Zhou, J. HD-BSNet: A Plug-and-Play Dual-Mechanism Synergistic Enhancement Framework for Small Object Detection. Remote Sens. 2026, 18, 423. https://doi.org/10.3390/rs18030423
Wen J, Zheng X, Pan N, Jia D, Wu H, Chen T, Zhou J. HD-BSNet: A Plug-and-Play Dual-Mechanism Synergistic Enhancement Framework for Small Object Detection. Remote Sensing. 2026; 18(3):423. https://doi.org/10.3390/rs18030423
Chicago/Turabian StyleWen, Jianwei, Xiangyue Zheng, Nian Pan, Dan Jia, Haiying Wu, Tao Chen, and Jin Zhou. 2026. "HD-BSNet: A Plug-and-Play Dual-Mechanism Synergistic Enhancement Framework for Small Object Detection" Remote Sensing 18, no. 3: 423. https://doi.org/10.3390/rs18030423
APA StyleWen, J., Zheng, X., Pan, N., Jia, D., Wu, H., Chen, T., & Zhou, J. (2026). HD-BSNet: A Plug-and-Play Dual-Mechanism Synergistic Enhancement Framework for Small Object Detection. Remote Sensing, 18(3), 423. https://doi.org/10.3390/rs18030423

