Underwater Robot Object Detection Algorithm Based on YOLOv11
Abstract
1. Introduction
- (1)
- A degradation-aware feature reconstruction strategy is introduced by replacing selected standard convolutions in YOLOv11s with SCConv. The SRU and CRU components suppress redundant spatial and channel responses caused by underwater turbidity, scattering, and color attenuation, thereby improving the representation of degraded underwater targets.
- (2)
- A lightweight Shuffle Attention module is incorporated to strengthen channel–spatial feature interaction. This design helps the network focus on informative target regions and fine-grained texture cues under non-uniform illumination and complex underwater backgrounds.
- (3)
- Focaler-IoU is adopted to improve bounding-box regression for underwater targets with blurred boundaries, small scales, and partial occlusion. By remapping the IoU interval, the regression process can better adapt to difficult underwater samples.
- (4)
- Extensive experiments, including ablation studies, generalization evaluation, challenging-scenario analysis, and underwater robotic tests, are conducted to verify the effectiveness and practical applicability of the proposed framework.
2. Materials and Methods
2.1. Overall Framework of the ROV
2.2. Research on Object Detection Algorithms Based on YOLOv11
2.3. Improvements to the YOLOv11s Object Detection Network
2.3.1. Improvements to the Original Convolution
2.3.2. Improvements to Attention Mechanisms
2.3.3. Improvements to the Loss Function
3. Results
3.1. Data Collection and Experimental Setup
3.2. Evaluation Indicators
3.3. Experimental Results and Analysis
3.3.1. Experiments on the URPC Dataset
3.3.2. Experiments on the RUOD Dataset
3.3.3. Evaluation of a Subset of Challenging Underwater Scenarios
3.3.4. Comparison with Other Models on the URPC and RUOD Datasets
3.3.5. Ablation Experiment
3.3.6. Edge Device Deployment and Power Consumption Analysis
3.3.7. Underwater Robot Prototype Grasping Experiment and Results Analysis
- (1)
- Artificial Pond Experiment
- (2)
- Natural Lake Water Experiment
4. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Peng, W.; Li, Y.Z.; Gao, Y.B. Research on systematic development of marine observation instruments and equipment in China. Chin. J. Sci. Instrum. 2023, 44, 88–100. [Google Scholar]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 580–587. [Google Scholar]
- Wang, N.; Zhi, M. A survey of single-stage general object detection algorithms based on deep learning. J. Front. Comput. Sci. Technol. 2025, 19, 1115–1140. [Google Scholar]
- Anilkumar, S.; Dhanya, P.R.; Balakrishnan, A.A. Algorithm for underwater cable tracking using CLAHE-based enhancement. In Proceedings of the 2019 International Symposium on Ocean Technology (SYMPOL), Ernakulam, India, 11–13 December 2019; pp. 129–137. [Google Scholar]
- Land, E.H. The retinex theory of color vision. Sci. Am. 1977, 237, 108–127. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, K.; Sun, J.; Tang, X. Single image haze removal using dark channel prior. IEEE Trans. Pattern Anal. Mach. Intell. 2010, 33, 2341–2353. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bhatti, U.A.; Yu, Z.; Chanussot, J. Local similarity-based spatial–spectral fusion hyperspectral image classification with deep CNN and Gabor filtering. IEEE Trans. Geosci. Remote Sens. 2021, 60, 5514215. [Google Scholar] [CrossRef] [Scilit]
- Voronin, V.; Semenishchev, E.; Tokareva, S.; Zelenskiy, A.; Agaian, S. Underwater image enhancement algorithm based on logarithmic transform histogram matching with spatial equalization. In Proceedings of the 2018 14th IEEE International Conference on Signal Processing (ICSP), Beijing, China, 12–16 August 2018; pp. 434–438. [Google Scholar] [CrossRef] [Scilit]
- Saleem, A.; Paheding, S.; Rawashdeh, N.; Awad, A.; Kaur, N. A non-reference evaluation of underwater image enhancement methods using a new underwater image dataset. IEEE Access 2023, 11, 10412–10428. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Yang, M.; Yin, G. Self-adversarial generative adversarial network for underwater image enhancement. IEEE J. Ocean. Eng. 2023, 49, 237–248. [Google Scholar] [CrossRef] [Scilit]
- Kumari, L.; Majumder, A. Deep learning based object detection and its application: A review. SN Comput. Sci. 2025, 6, 805. [Google Scholar] [CrossRef] [Scilit]
- Girshick, R. Fast R-CNN. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 7–13 December 2015; pp. 1440–1448. [Google Scholar]
- Kumar, A.; Srivastava, S. Object detection system based on convolution neural networks using single shot multi-box detector. Procedia Comput. Sci. 2020, 171, 2610–2617. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
- Tian, Z.; Shen, C.; Chen, H.; He, T. FCOS: Fully Convolutional One-Stage Object Detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9627–9636. [Google Scholar]
- Zhong, X. CAL-SSD: Lightweight SSD object detection based on coordinated attention. Signal Image Video Process. 2025, 19, 31. [Google Scholar] [CrossRef] [Scilit]
- Archana, V.; Kalaiselvi, S.; Thamaraiselvi, D.; Gomathi, V.; Sowmiya, R. A novel object detection framework using convolutional neural networks (CNN) and RetinaNet. In Proceedings of the 2022 International Conference on Automation, Computing and Renewable Systems (ICACRS), Pudukkottai, India, 13–15 December 2022; pp. 1070–1074. [Google Scholar] [CrossRef] [Scilit]
- Zhao, D.; Yang, B.; Dou, Y.; Guo, X. Underwater fish detection in sonar image based on an improved Faster RCNN. In Proceedings of the 2022 9th International Forum on Electrical Engineering and Automation (IFEEA), Zhuhai, China, 4–6 November 2022; pp. 358–363. [Google Scholar] [CrossRef] [Scilit]
- Hussain, M. YOLOv1 to v8: Unveiling each variant—A comprehensive review of YOLO. IEEE Access 2024, 12, 42816–42833. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Liu, J.; Yu, S.; Wang, K.; Han, Z.; Tang, Y. Underwater object detection based on YOLO-v3 network. In Proceedings of the 2021 IEEE International Conference on Unmanned Systems (ICUS), Beijing, China, 15–17 October 2021; pp. 571–575. [Google Scholar] [CrossRef] [Scilit]
- Huang, K.; Xu, M.; Guo, T.; Chen, F.; Wei, C.-L.; Chu, C. DPT-YOLO: An improved underwater object detection model based on YOLO. In Proceedings of the 2024 Cross Strait Radio Science and Wireless Technology Conference (CSRSWTC), Macao, China, 11–13 October 2024; pp. 1–3. [Google Scholar] [CrossRef] [Scilit]
- Wu Reddy, T.N.; Kumar, N.; Ponnappa, N.P. Intelligent GD&T symbol detection in mechanical drawings: A comparative study of YOLOv11, Faster R-CNN, and RetinaNet for quality assurance. J. Intell. Manuf. 2025. [Google Scholar] [CrossRef] [Scilit]
- Selcuk, B.; Serif, S.; Serif, T. A comparative study on distal radius fracture detection: YOLOv8 and YOLOv11 versus Faster R-CNN. In Mobile Web and Intelligent Information Systems; Younas, M., Awan, I., Martin, L., Wu, H., Eds.; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2026; Volume 16066. [Google Scholar] [CrossRef] [Scilit]
- Hong, X.Y.; Jia, B.W.; Chen, D.Y. An improved YOLOv11-based object detection method for underwater hull cleaning robots. Ship Ocean Eng. 2025, 36, 13–26. [Google Scholar] [CrossRef]
- Chen, H.; Yu, Y.J. LGM-YOLOv11: An underwater object detection model integrating a multi-scale attention mechanism. Comput. Eng. Appl. 2025, 61, 248–263. [Google Scholar]
- Chang, Z.Y.; Du, Y.R.; Qu, M.S. A segmentation and detection algorithm for cableway sheaves and wire ropes based on YOLOv11. Hoist. Convey. Mach. 2026, 1, 42–48. [Google Scholar]
- Gu, H.; Huang, J.; Shen, Y.; Wang, S. YOLOv11-FST: An improved YOLOv11-based model for solar photovoltaic cell defect detection. In Proceedings of the 2025 6th International Conference on Big Data & Artificial Intelligence & Software Engineering (ICBASE), Johor Bahru, Malaysia, 18–20 July 2025; pp. 227–230. [Google Scholar] [CrossRef] [Scilit]
- Kıratlı, R.; Eroğlu, A. Real-time multi-object detection and tracking in UAV systems: Improved YOLOv11-EFAC and optimized tracking algorithms. J. Real-Time Image Process. 2025, 22, 178. [Google Scholar] [CrossRef] [Scilit]



























| Dataset | Total Images | Categories | Training | Validation | Split Ratio |
|---|---|---|---|---|---|
| URPC | 5543 | 4 | 4989 | 554 | 9:1 |
| RUOD | 13,950 | 10 | 12,555 | 1395 | 9:1 |
| Parameter | Value |
|---|---|
| Optimizer | SGD |
| Batch size | 4 |
| Learning rate | 0.01 |
| Confidence threshold | 0.5 |
| Weight decay coefficient | 0.0005 |
| Maximum epochs | 100 |
| Early stopping criterion | Best validation performance |
| Model | AP (%) | mAP@0.5 (%) | FPS | |||
|---|---|---|---|---|---|---|
| Sea Urchin | Sea Cucumber | Star Fish | Scallop | |||
| SSD | 74.7 | 69.9 | 75.2 | 60.2 | 70.0 | 21 |
| YOLOv8s | 88.1 | 73.6 | 87.1 | 86.0 | 83.7 | 92 |
| YOLOv9s | 88.9 | 74.1 | 87.6 | 86.2 | 84.2 | 88 |
| YOLOv10s | 89.2 | 74.5 | 87.9 | 86.6 | 84.6 | 91 |
| RT-DETR-R18 | 85.5 | 73.8 | 87.4 | 86.1 | 84.0 | 74 |
| YOLOv12s | 90.5 | 75.3 | 88.5 | 87.1 | 85.4 | 85 |
| DPT-YOLO | 87.2 | 74.2 | 86.2 | 85.3 | 84.1 | 72 |
| LGM-YOLOv11 | 90.2 | 74.8 | 88.7 | 86.2 | 85.9 | 82 |
| YOLOv11s | 89.8 | 75.1 | 88.3 | 84.9 | 85.2 | 90 |
| Improved YOLOv11s | 92.1 | 76.6 | 89.1 | 87.4 | 88.4 | 87 |
| Model | Precision (%) | Recall (%) | mAP@0.5 (%) | FPS |
|---|---|---|---|---|
| SSD | 72.6 | 68.9 | 70.8 | 22 |
| YOLOv8s | 85.0 | 82.1 | 84.8 | 92 |
| YOLOv9s | 85.7 | 82.6 | 85.2 | 88 |
| YOLOv10s | 86.0 | 82.9 | 85.6 | 91 |
| RT-DETR-R18 | 85.1 | 81.7 | 84.7 | 73 |
| YOLOv12s | 87.0 | 83.8 | 86.6 | 85 |
| DPT-YOLO | 84.8 | 81.6 | 84.3 | 71 |
| LGM-YOLOv11 | 86.5 | 83.4 | 86.2 | 81 |
| YOLOv11s | 86.2 | 83.1 | 86.0 | 90 |
| Improved YOLOv11s | 88.1 | 85.0 | 87.9 | 87 |
| Dataset | Model | Precision (%) | Recall (%) | mAP@0.5 (%) |
|---|---|---|---|---|
| URPC | YOLOv11s | 86.1 ± 0.3 | 82.5 ± 0.3 | 85.2 ± 0.2 |
| URPC | Improved YOLOv11s | 89.0 ± 0.3 | 85.4 ± 0.3 | 88.4 ± 0.2 |
| RUOD | YOLOv11s | 86.2 ± 0.3 | 83.1 ± 0.3 | 86.0 ± 0.2 |
| RUOD | Improved YOLOv11s | 88.1 ± 0.2 | 85.0 ± 0.3 | 87.9 ± 0.2 |
| Aspect | Typical YOLO Improvements | Proposed Method |
|---|---|---|
| Motivation | Accuracy/speed/fusion | Underwater degradation |
| Design logic | Generic module optimization | Degradation-aware adaptation |
| Key focus | Feature enhancement | Redundancy, attention, localization |
| Validation | General comparison | URPC, RUOD, robotic tests |
| Model | B | S | D | F | mAP@0.5 (%) | FLOPs (G) | FPS | Parameter (M) |
|---|---|---|---|---|---|---|---|---|
| 1 | √ | 85.2 | 27.7 | 95.9 | 10.9 | |||
| 2 | √ | √ | 86.1 | 31.0 | 90.2 | 13.4 | ||
| 3 | √ | √ | 86.7 | 31.6 | 89.1 | 13.9 | ||
| 4 | √ | √ | 86.3 | 31.8 | 89.2 | 14.1 | ||
| 5 | √ | √ | √ | 87.1 | 32.6 | 88.7 | 16.4 | |
| 6 | √ | √ | √ | 86.9 | 32.4 | 88.3 | 16.1 | |
| 7 | √ | √ | √ | 87.2 | 32.6 | 87.1 | 16.6 | |
| 8 | √ | √ | √ | √ | 88.4 | 32.9 | 86.7 | 17.6 |
| Model | Inference Speed (FPS) | Average Power (W) |
|---|---|---|
| YOLOv11s | 36.2 | 11.5 |
| Improved YOLOv11s | 32.5 | 12.4 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Shi, Y.; Chen, W.; Wan, D.; Han, L. Underwater Robot Object Detection Algorithm Based on YOLOv11. Sensors 2026, 26, 3611. https://doi.org/10.3390/s26113611
Shi Y, Chen W, Wan D, Han L. Underwater Robot Object Detection Algorithm Based on YOLOv11. Sensors. 2026; 26(11):3611. https://doi.org/10.3390/s26113611
Chicago/Turabian StyleShi, Yongqing, Wei Chen, Duo Wan, and Lu Han. 2026. "Underwater Robot Object Detection Algorithm Based on YOLOv11" Sensors 26, no. 11: 3611. https://doi.org/10.3390/s26113611
APA StyleShi, Y., Chen, W., Wan, D., & Han, L. (2026). Underwater Robot Object Detection Algorithm Based on YOLOv11. Sensors, 26(11), 3611. https://doi.org/10.3390/s26113611
