Ship-DiffDet: A Lightweight Diffusion Model for Small-Object Ship Detection
Abstract
1. Introduction
- We propose an optimization method for diffusion models tailored to small ship detection, which significantly reduces the number of parameters and computational cost while improving detection accuracy for small-object ships. To the best of our knowledge, no prior work has explored lightweight diffusion models specifically designed for detecting small-object ships in visible-light images.
- To address the limitation that most publicly available ship datasets focus on close and medium range objects and lack sufficient samples of small-object ships, we constructed a dataset of long-range small-object ships that includes various weather conditions and interferences from flying objects such as seagulls.
- Extensive experiments on our custom-built small-object ship dataset demonstrate that Ship-DiffDet achieves superior performance in small-object detection accuracy compared to mainstream detectors. In addition, its inference speed can be adjusted according to application requirements, indicating strong potential for edge deployment.
2. Related Works
2.1. Small-Object Detection Methods
2.2. Diffusion Model for Object Detection
2.3. Lightweight Model Design
3. Proposed Model
3.1. Preliminary
3.2. Lightweight Backbone with Inception Depthwise Convolution—IDC-Net
3.3. Enhanced Feature Pyramid Network with Hybrid Pooling Attention—HP-FPN
3.4. Efficient Dynamic Head with Gating Mechanism
4. Experiment and Results
4.1. Experiment Platform
4.2. Metrics
4.3. Dataset
4.4. Ablation Experiments
4.5. Comparative Experiments
4.6. Parameter Sensitivity Analysis
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Pagire, V.; Chavali, M.; Kale, A. A comprehensive review of object detection with traditional and deep learning methods. Signal Process. 2025, 237, 110075. [Google Scholar] [CrossRef] [Scilit]
- Ramos, L.T.; Sappa, A.D. A decade of you only look once (yolo) for object detection: A review. IEEE Access 2025, 13, 192747–192794. [Google Scholar] [CrossRef] [Scilit]
- Muzammul, M.; Li, X. Comprehensive review of deep learning-based tiny object detection: Challenges, strategies, and future directions. Knowl. Inf. Syst. 2025, 67, 3825–3913. [Google Scholar] [CrossRef] [Scilit]
- Cheng, Y.; Lai, X.; Xia, Y.; Zhou, J. Infrared dim small target detection networks: A review. Sensors 2024, 24, 3885. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhao, X.; Zhao, Y.; Hu, S.; Wang, H.; Zhang, Y.; Ming, W. Progress in active infrared imaging for defect detection in the renewable and electronic industries. Sensors 2023, 23, 8780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kumar, N.; Singh, P. Small and dim target detection in infrared imagery: A review, current techniques and future directions. Neurocomputing 2025, 630, 129640. [Google Scholar] [CrossRef] [Scilit]
- Moreira, A.; Prats-Iraola, P.; Younis, M.; Krieger, G.; Hajnsek, I.; Papathanassiou, K.P. A tutorial on synthetic aperture radar. IEEE Geosci. Remote Sens. Mag. 2013, 1, 6–43. [Google Scholar] [CrossRef] [Scilit]
- Peter, E.; Ang, L.-M.; Seng, K.P.; Srivastava, S. Recent Advances in Deep Learning for SAR Images: Overview of Methods, Challenges, and Future Directions. Sensors 2026, 26, 1143. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lang, P.; Fu, X.; Dong, J.; Yang, H.; Yin, J.; Yang, J.; Martorella, M. Recent advances in deep learning based SAR image targets detection and recognition. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 6884–6915. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30. [Google Scholar] [CrossRef] [Scilit]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 580–587. [Google Scholar]
- Girshick, R. Fast r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 7–13 December 2015; pp. 1440–1448. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 39, 1137–1149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 24–27 October 2017; pp. 2961–2969. [Google Scholar]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. Ssd: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2016; pp. 21–37. [Google Scholar]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 24–27 October 2017; pp. 2980–2988. [Google Scholar]
- Shao, Z.; Yin, Y.; Lyu, H.; Soares, C.G.; Cheng, T.; Jing, Q.; Yang, Z. An efficient model for small object detection in the maritime environment. Appl. Ocean Res. 2024, 152, 104194. [Google Scholar] [CrossRef] [Scilit]
- Zhou, W.; Xu, B.; Wu, H. Dual-branch wavelet-CNN architecture for enhanced small ship detection. Ocean Eng. 2025, 339, 121967. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Wang, S. Efficient and lightweight deep learning model for enhanced ship detection in maritime surveillance. Ocean Eng. 2025, 328, 121085. [Google Scholar] [CrossRef] [Scilit]
- Kim, H.-C.; Lee, H.-T.; Cho, I.-S. Vessel detection for maritime traffic management using U-Net with backbone networks. Ocean Eng. 2025, 340, 121943. [Google Scholar] [CrossRef] [Scilit]
- Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature pyramid networks for object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 2117–2125. [Google Scholar]
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path aggregation network for instance segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 8759–8768. [Google Scholar]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 213–229. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Han, K.; Wang, Y.; Chen, H.; Chen, X.; Guo, J.; Liu, Z.; Tang, Y.; Xiao, A.; Xu, C.; Xu, Y. A survey on vision transformer. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 87–110. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wu, W.; Fan, X.; Hu, Z.; Zhao, Y. CGDU-DETR: An End-to-End Detection Model for Ship Detection in Day–Night Transition Environments. J. Mar. Sci. Eng. 2025, 13, 1155. [Google Scholar] [CrossRef] [Scilit]
- Liu, C.; Zhang, Y.; Shen, J.; Liu, F. Improved RT-DETR for infrared ship detection based on multi-attention and feature fusion. J. Mar. Sci. Eng. 2024, 12, 2130. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Li, X. Ship-DETR: A transformer-based model for efficient ship detection in complex maritime environments. IEEE Access 2025, 13, 66031–66039. [Google Scholar] [CrossRef] [Scilit]
- Song, J.; Meng, C.; Ermon, S. Denoising diffusion implicit models. arXiv 2020, arXiv:2010.02502. [Google Scholar]
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, 33, 6840–6851. [Google Scholar]
- Yang, L.; Zhang, Z.; Song, Y.; Hong, S.; Xu, R.; Zhao, Y.; Zhang, W.; Cui, B.; Yang, M.-H. Diffusion models: A comprehensive survey of methods and applications. ACM Comput. Surv. 2023, 56, 1–39. [Google Scholar] [CrossRef] [Scilit]
- Croitoru, F.-A.; Hondru, V.; Ionescu, R.T.; Shah, M. Diffusion models in vision: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 10850–10869. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, S.; Sun, P.; Song, Y.; Luo, P. Diffusiondet: Diffusion model for object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 4–6 October 2023; pp. 19830–19843. [Google Scholar]
- Zhou, H.; Cao, X.; Sun, K.; Wu, T.; Deng, B. Dense small object detection via multi-scale fusion and context information enhancement. J. Supercomput. 2025, 81, 955. [Google Scholar] [CrossRef] [Scilit]
- Peng, J.; Lv, K.; Wang, G.; Xiao, W.; Ran, T.; Yuan, L. MLSA-YOLO: A multi-level feature fusion and scale-adaptive framework for small object detection. J. Supercomput. 2025, 81, 528. [Google Scholar] [CrossRef] [Scilit]
- Guo, G.; Chen, P.; Yu, X.; Han, Z.; Ye, Q.; Gao, S. Save the tiny, save the all: Hierarchical activation network for tiny object detection. IEEE Trans. Circuits Syst. Video Technol. 2023, 34, 221–234. [Google Scholar] [CrossRef] [Scilit]
- Tan, M.; Pang, R.; Le, Q.V. Efficientdet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 10781–10790. [Google Scholar]
- Gong, Y.; Chen, Z.; Deng, W.; Tan, J.; Li, Y. Real-time long-distance ship detection architecture based on YOLOv8. IEEE Access 2024, 12, 116086–116104. [Google Scholar] [CrossRef] [Scilit]
- Shen, L.; Gao, T.; Yin, Q. Yolo-lpss: A lightweight and precise detection model for small sea ships. J. Mar. Sci. Eng. 2025, 13, 925. [Google Scholar] [CrossRef] [Scilit]
- Gong, Y.; Chen, Z.; Tan, J.; Yin, C.; Deng, W. Two-stage ship detection at long distances based on deep learning and slicing technique. PLoS ONE 2024, 19, e0313145. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, Q.; Chen, H.; Zhao, F. Maritime target detection algorithm based on fusion of visible and infrared images. J. Supercomput. 2025, 81, 22. [Google Scholar] [CrossRef] [Scilit]
- Fan, J.; Zhang, E.; Wei, Y.; Wang, Y.; Xia, J.; Liu, J.; Liu, X.; Ma, S. DDOWOD: DiffusionDet for open-world object detection. Pattern Recognit. Lett. 2024, 186, 170–177. [Google Scholar] [CrossRef] [Scilit]
- Orfaig, E.; Stainvas, I.; Bilik, I. RGBX-DiffusionDet: A framework for multi-modal RGB-X object detection using DiffusionDet. Pattern Recognit. 2025, 172, 112460. [Google Scholar] [CrossRef] [Scilit]
- Erabati, G.K.; Araujo, H. DDet3D: Embracing 3D object detector with diffusion: GK Erabati and H. Araujo. Appl. Intell. 2025, 55, 283. [Google Scholar]
- Han, J.; Sun, J.; Wang, F.; Sun, F.; Li, H. ORSIDiff: Diffusion model for salient object detection in optical remote sensing images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5627315. [Google Scholar] [CrossRef] [Scilit]
- Howard, A.G.; Zhu, M.; Chen, B.; Kalenichenko, D.; Wang, W.; Weyand, T.; Andreetto, M.; Adam, H. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv 2017, arXiv:1704.04861. [Google Scholar]
- Zhang, X.; Zhou, X.; Lin, M.; Sun, J. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 6848–6856. [Google Scholar]
- Liu, R.; Zhu, Z.; Ge, H.; Wang, J.; Shu, Y.; Ji, Q. Towards scale-adaptive and lightweight maritime ship object detection via dual-cross multi-scale knowledge distillation. Ocean Eng. 2026, 343, 123206. [Google Scholar] [CrossRef] [Scilit]
- Sang, H.; Lu, Q.; Sun, X.; Zhang, S.; Liu, F. A lightweight multi-scale ship detection framework for wave gliders with spatial-channel attention fusion. Measurement 2025, 258, 119280. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Jia, Z.; Zhang, X.; Yang, F.; Yang, X. A lightweight and efficient ship detection model for complex maritime environments. Appl. Ocean Res. 2025, 164, 104797. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, BC, Canada, 11–17 October 2021; pp. 10012–10022. [Google Scholar]
- Sun, P.; Zhang, R.; Jiang, Y.; Kong, T.; Xu, C.; Zhan, W.; Tomizuka, M.; Yuan, Z.; Luo, P. Sparse R-CNN: An end-to-end framework for object detection. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 15650–15664. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yu, W.; Zhou, P.; Yan, S.; Wang, X. Inceptionnext: When inception meets convnext. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 5672–5683. [Google Scholar]
- Szegedy, C.; Liu, W.; Jia, Y.; Sermanet, P.; Reed, S.; Anguelov, D.; Erhan, D.; Vanhoucke, V.; Rabinovich, A. Going deeper with convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; pp. 1–9. [Google Scholar]
- Chen, P.; He, W.; Qian, F.; Shi, G.; Yan, J. A synergistic CNN-transformer network with pooling attention fusion for hyperspectral image classification. Digit. Signal Process. 2025, 160, 105070. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Wang, Z.; Liu, Z.; Tan, C.; Lin, H.; Wu, D.; Chen, Z.; Zheng, J.; Li, S.Z. Moganet: Multi-order gated aggregation network. arXiv 2022, arXiv:2211.03295. [Google Scholar]
- Shao, Z.; Wu, W.; Wang, Z.; Du, W.; Li, C. Seaships: A large-scale precisely annotated dataset for ship detection. IEEE Trans. Multimed. 2018, 20, 2593–2604. [Google Scholar] [CrossRef] [Scilit]
- Prasad, D.K.; Rajan, D.; Rachmawati, L.; Rajabally, E.; Quek, C. Video processing from electro-optical sensors for object detection and tracking in a maritime environment: A survey. IEEE Trans. Intell. Transp. Syst. 2017, 18, 1993–2016. [Google Scholar] [CrossRef] [Scilit]
- Iancu, B.; Soloviev, V.; Zelioli, L.; Lilius, J. Aboships—An inshore and offshore maritime vessel detection dataset with precise annotations. Remote Sens. 2021, 13, 988. [Google Scholar] [CrossRef] [Scilit]











| Convolution Type | Parameter | FLOPs |
|---|---|---|
| Conventional convolution | k2C2 | 2k2C2HW |
| Depthwise convolution | k2C | 2k2CHW |
| Inception depthwise convolution | (2k + 9)C/8 | (2k + 9)CHW/4 |
| Configuration | Parameter |
|---|---|
| CPU | Inter core i7-10700 |
| GPU | NVIDIA RTX3090TI |
| Operating system | Ubuntu16.04 |
| CUDA | 11.1 |
| DIM_FEEDFORWARD | 1024 |
| HIDDEN_DIM | 128 |
| SAMPLE_STEP | 1 |
| NUM_PROPOSALS | 200 |
| DiffusionDet | IDC-Net | HP-FPN | MOGA | AP (%) | AP50 (%) | APs (%) | Parameter (M) | FLOPs (G) | GPU Memory Usage (G) |
|---|---|---|---|---|---|---|---|---|---|
| √ | 45.1 | 89.2 | 44.3 | 48.2 | 63.7 | 9.5 | |||
| √ | √ | 45.4 | 89.5 | 44.6 | 36.9 | 48.7 | 8.2 | ||
| √ | √ | 45.4 | 89.6 | 44.5 | 48.2 | 63.9 | 11.1 | ||
| √ | √ | 45.2 | 89.5 | 44.5 | 36.0 | 64.5 | 9.4 | ||
| √ | √ | √ | 45.7 | 90.4 | 45.0 | 36.9 | 48.9 | 10.0 | |
| √ | √ | √ | 45.8 | 89.8 | 44.9 | 24.7 | 49.5 | 8.3 | |
| √ | √ | √ | 46.0 | 90.5 | 45.2 | 36.0 | 64.7 | 11.3 | |
| √ | √ | √ | √ | 46.1 | 90.9 | 45.3 | 24.7 | 49.7 | 10.0 |
| Method | AP50 | FPS | Inference Time (ms) |
|---|---|---|---|
| Centernet | 64.6 | 48 | 20.8 |
| Faster R-CNN | 39.7 | 48 | 20.8 |
| SSD | 63.2 | 143 | 7.0 |
| TPH-YOLOv5 | 77.2 | 20 | 50.0 |
| YOLOv8 | 70.6 | 476 | 2.1 |
| YOLO-Fastestv2 | 81.0 | 3 | 333.3 |
| YOLOv11 | 75.0 | 455 | 2.2 |
| YOLOv26 | 87.9 | 468 | 2.1 |
| DiffusionDet | 89.2 | 45 | 22.2 |
| Ship-DiffDet | 90.9 | 42 | 23.8 |
| Proposal Boxes | 200 | 300 | 500 | 1000 | 2000 | 4000 | |
|---|---|---|---|---|---|---|---|
| Sample Steps | |||||||
| 1 | 45.3 | 45.3 | 45.3 | 45.6 | 45.4 | 45.3 | |
| 3 | 45.4 | 45.6 | 45.6 | 45.5 | 45.3 | 45.4 | |
| 5 | 45.8 | 46.0 | 46.2 | 45.6 | 45.3 | 45.2 | |
| 7 | 46.6 | 46.4 | 46.2 | 45.8 | 45.6 | 45.4 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Gong, Y.; Huang, J.; Zhang, D.; Sheng, J. Ship-DiffDet: A Lightweight Diffusion Model for Small-Object Ship Detection. J. Mar. Sci. Eng. 2026, 14, 1535. https://doi.org/10.3390/jmse14161535
Gong Y, Huang J, Zhang D, Sheng J. Ship-DiffDet: A Lightweight Diffusion Model for Small-Object Ship Detection. Journal of Marine Science and Engineering. 2026; 14(16):1535. https://doi.org/10.3390/jmse14161535
Chicago/Turabian StyleGong, Yanfeng, Jing Huang, Daiyong Zhang, and Jinlu Sheng. 2026. "Ship-DiffDet: A Lightweight Diffusion Model for Small-Object Ship Detection" Journal of Marine Science and Engineering 14, no. 16: 1535. https://doi.org/10.3390/jmse14161535
APA StyleGong, Y., Huang, J., Zhang, D., & Sheng, J. (2026). Ship-DiffDet: A Lightweight Diffusion Model for Small-Object Ship Detection. Journal of Marine Science and Engineering, 14(16), 1535. https://doi.org/10.3390/jmse14161535

