Towards Practical Object Detection with Limited Data: A Feature Distillation Framework
Abstract
1. Introduction
- (1)
- To mitigate defect sample scarcity and low diversity, we utilize an unsupervised, unpaired style transfer model to adapt images from a large-scale general dataset (e.g., COCO [21]) to the style of the target underwater environment. This approach enriches training data and enhances feature variability without requiring extensive real defect annotations.
- (2)
- To balance the dual requirements of detection accuracy and operational efficiency, we select and adapt a suitable baseline model to serve as the teacher and the student model, thereby establishing a practical foundation for deployment on embedded systems with limited computational resources.
- (3)
- To improve model performance under limited-data conditions, we propose a multi-level feature distillation mechanism, through which robust feature extraction ability from a pre-trained large teacher model (trained on a relevant large-scale dataset) is transferred to the lightweight student network, compensating for the information shortage caused by limited real samples.
2. Related Works
2.1. Image Enhancement/Restoration-Based Pre-Processing Paradigm
2.2. Network Architecture Adaptation Paradigm
2.3. Data and Training Strategy Co-Optimization Paradigm
3. Proposed Method
3.1. Model
3.1.1. Sample Generation
- (1)
- Adversarial Loss:
- (2)
- Structural Similarity Loss:
- (3)
- Identity Consistency Loss:
- (4)
- Total Objective
3.1.2. Baseline Model
- (1)
- Bounding Box Regression Loss:
- (2)
- Classification Loss:
- (3)
- Distribution Focal Loss:
3.1.3. Feature Distillation
3.2. Training
3.2.1. Style Transfer via CUT
3.2.2. Teacher Model Pre-Training
3.2.3. Training for Distillation Model
4. Experiments and Analysis
4.1. Datasets
4.1.1. Style-Transferred Dataset
4.1.2. Operational Environment Dataset
4.2. Component Selection, Analysis, and Validation
4.2.1. Motivation for CUT Adoption
4.2.2. Overview and Selection Rationale for YOLOv10 Variants
4.2.3. Analysis of the Distillation Design
- (1)
- Selection of the Distillation Model
- (2)
- Selection of Distillation Layers
4.2.4. Effectiveness of Individual Components
4.3. Pre- vs. Post-Distillation Performance Analysis
4.3.1. Quantitative Analysis
4.3.2. Qualitative Analysis
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Jian, M.; Yang, N.; Tao, C.; Zhi, H.; Luo, H. Underwater object detection and datasets: A survey. Intell. Mar. Technol. Syst. 2024, 2, 9. [Google Scholar] [CrossRef]
- Hunt, J.D.; Nascimento, A.; Romero, O.J.; Zakeri, B.; Jurasz, J.; Dąbek, P.B.; Strzyżewski, T.; Đurin, B.; Filho, W.L.; Freitas, M.A.V.; et al. Hydrogen Storage with Gravel and Pipes in Lakes and Reservoirs. Nat. Commun. 2024, 15, 7723. [Google Scholar] [CrossRef]
- Ti, Z.; Zhang, M.; Li, Y.; Wei, K. Numerical Study on the Stochastic Response of A Long-Span Sea Crossing Bridge Subjected to Extreme Nonlinear Wave Loads. Eng. Struct. 2019, 196, 109287. [Google Scholar] [CrossRef]
- Wang, Z.; Zhang, X.; Ran, C.; Yu, H.; Wang, S.; Zhang, Q.; Nie, Y.; Zhou, X. Deep Learning-Based Semantic Segmentation and Surface Reconstruction for Point Clouds of Offshore Oil Production Equipment. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5700918. [Google Scholar] [CrossRef]
- Jia, J.; Fu, M.; Liu, X.; Zheng, B. Underwater Object Detection Based on Improved EfficientDet. Remote Sens. 2022, 14, 4487. [Google Scholar] [CrossRef]
- Pramudita, A.A.; Lin, D.-B.; Dhiyani, A.A.; Ryanu, H.H.; Adiprabowo, T.; Yudha, E.A. FMCW Radar for Noncontact Bridge Structure Displacement Estimation. IEEE Trans. Instrum. Meas. 2023, 72, 8504914. [Google Scholar] [CrossRef]
- Yuan, X.; Li, W.; Chen, G.; Yin, X.; Li, X.; Liu, J.; Zhao, J.; Zhao, J. Visual and Intelligent Identification Methods for Defects in Underwater Structure Using Alternating Current Field Measurement Technique. IEEE Trans. Ind. Inform. 2021, 18, 3853–3862. [Google Scholar] [CrossRef]
- Xu, S.; Zhang, M.; Song, W.; Mei, H.; He, Q.; Liotta, A. A Systematic Review and Analysis of Deep Learning-Based Underwater Object Detection. Neurocomputing 2023, 527, 204–232. [Google Scholar] [CrossRef]
- Vidal, G.E.; Hernández Vega, J.D.; Istenič, K.; Carreras, M. Online View Planning for Inspecting Unexplored Underwater Structures. IEEE Robot. Autom. Lett. 2017, 2, 1436–1443. [Google Scholar] [CrossRef]
- Nielsen, P.L.; Muzi, L.; Siderius, M. Seabed Characterization from Ambient Noise Using Short Arrays and Autonomous Vehicles. IEEE J. Ocean. Eng. 2017, 42, 1094–1101. [Google Scholar] [CrossRef]
- Du, H.; Yao, D.; Li, S.; Zhang, Q. Ultrasonic Measurement on the Thickness of Oil Slick Using the Remotely Operated Vehicle (ROV) as a Platform. IEEE Trans. Instrum. Meas. 2023, 72, 7500810. [Google Scholar] [CrossRef]
- Katou, M.; Tara, K.; Saito, S.; Hondori, E.J.; Koshigoe, K.; Asakawa, E. Structural Imaging of Acoustic Survey Using A Deep-Towed Sub-Bottom Profiler and Hydrophone Cable. IEEE J. Ocean. Eng. 2022, 47, 399–416. [Google Scholar] [CrossRef]
- Batmani, Y.; Najafi, S. Event-Triggered H∞ Depth Control of Remotely Operated Underwater Vehicles. IEEE Trans. Syst. Man Cybern. Syst. 2021, 51, 1224–1232. [Google Scholar] [CrossRef]
- Xu, C.; Xie, Z. A Lightweight Underwater Object Detection with Enhanced Detail and Edge-Aware Feature Fusion. Digit. Signal Process. 2025, 167, 105456. [Google Scholar] [CrossRef]
- Chen, L.; Huang, Y.; Dong, J.; Xu, Q.; Kwong, S.; Lu, H.; Lu, H.; Li, C. Underwater Optical Object Detection in the Era of Artificial Intelligence: Current, Challenge, and Future. ACM Comput. Surv. 2025, 58, 62. [Google Scholar] [CrossRef]
- Fayaz, S.; Parah, S.A.; Qureshi, G.J.; Lloret, J.; Del Ser, J.; Muhammad, K. Intelligent Underwater Object Detection and Image Restoration for Autonomous Underwater Vehicles. IEEE Trans. Veh. Technol. 2024, 73, 1726–1735. [Google Scholar] [CrossRef]
- Chen, J.; Er, M.J. Dynamic YOLO for small underwater object detection. Artif. Intell. Rev. 2024, 57, 165. [Google Scholar] [CrossRef]
- Liu, Z.; Wang, B.; Li, Y.; He, J.; Li, Y. UnitModule: A Lightweight Joint Image Enhancement Module for Underwater Object Detection. Pattern Recognit. 2024, 151, 110435. [Google Scholar] [CrossRef]
- Dai, L.; Liu, H.; Song, P.; Liu, M. A Gated Cross-domain Collaborative Network for Underwater Object Detection. Pattern Recognit. 2024, 149, 110222. [Google Scholar] [CrossRef]
- Wang, Y.; Guo, J.; He, W.; Gao, H.; Yue, H.; Zhang, Z.; Li, C. Is Underwater Image Enhancement All Object Detectors Need? IEEE J. Ocean. Eng. 2024, 49, 606–621. [Google Scholar] [CrossRef]
- Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft Coco: Common Objects in Context. In Proceedings of the 2014 13th European Conference on Computer Vision (ECCV), Zurich, Switzerland, 6–12 September 2014; pp. 740–755. [Google Scholar]
- Dakhil, R.A.; Khayeat, A.R.H. Review On Deep Learning Technique for Underwater Object Detection. arXiv 2022, arXiv:2209.10151. [Google Scholar] [CrossRef]
- Wang, H.; Sun, S.; Bai, X.; Wang, J.; Ren, P. A Reinforcement Learning Paradigm of Configuring Visual Enhancement for Object Detection in Underwater Scenes. IEEE J. Ocean. Eng. 2023, 48, 443–461. [Google Scholar] [CrossRef]
- Zhao, L.; Yun, Q.; Yuan, F.; Ren, X.; Jin, J.; Zhu, X. YOLOv7-CHS: An Emerging Model for Underwater Object Detection. J. Mar. Sci. Eng. 2023, 11, 1949. [Google Scholar] [CrossRef]
- Shen, X.; Sun, X.; Wang, H.; Fu, X. Multi-Dimensional, Multi-Functional and Multi-Level Attention in YOLO for Underwater Object Detection. Neural Comput. Appl. 2023, 35, 19935–19960. [Google Scholar] [CrossRef]
- Liang, X.; Song, P. Excavating RoI Attention for Underwater Object Detection. In Proceedings of the 2022 IEEE International Conference on Image Processing (ICIP), Bordeaux, France, 16–19 October 2022; pp. 2651–2655. [Google Scholar]
- Hua, X.; Cui, X.; Xu, X.; Qiu, S.; Liang, Y.; Bao, X.; Li, Z. Underwater Object Detection Algorithm Based on Feature Enhancement and Progressive Dynamic Aggregation Strategy. Pattern Recognit. 2023, 139, 109511. [Google Scholar] [CrossRef]
- Ji, X.; Chen, S.; Hao, L.-Y.; Zhou, J.; Chen, L. FBDPN: CNN-Transformer Hybrid Feature Boosting and Differential Pyramid Network for Underwater Object Detection. Expert Syst. Appl. 2024, 256, 124978. [Google Scholar] [CrossRef]
- Feng, J.; Tao, J. CEH-YOLO: A Composite Enhanced YOLO-Based Model for Underwater Object Detection. Ecol. Inform. 2024, 82, 102758. [Google Scholar] [CrossRef]
- Fu, C.; Fan, X.; Xiao, J.; Yuan, W.; Liu, R.; Luo, Z. Learning Heavily-Degraded Prior for Underwater Object Detection. IEEE Trans. Circuits Syst. Video Technol. 2023, 33, 6887–6896. [Google Scholar] [CrossRef]
- Yeh, C.-H.; Lin, C.-H.; Kang, L.-W.; Huang, C.-H.; Lin, M.-H.; Chang, C.-Y.; Wang, C.-C. Lightweight Deep Neural Network for Joint Learning of Underwater Object Detection and Color Conversion. IEEE Trans. Neural Netw. Learn. Syst. 2022, 33, 6129–6143. [Google Scholar] [CrossRef]
- Wang, B.; Wang, Z.; Guo, W.; Wang, Y. A Dual-Branch Joint Learning Network for Underwater Object Detection. Knowl.-Based Syst. 2024, 293, 111672. [Google Scholar] [CrossRef]
- Song, P.; Li, P.; Dai, L.; Wang, T.; Chen, Z. Boosting R-CNN: Reweighting R-CNN samples by RPN’s error for underwater object detection. Neurocomputing 2023, 530, 150–164. [Google Scholar] [CrossRef]
- Cai, S.; Li, G.; Shan, Y. Underwater Object Detection Using Collaborative Weakly Supervision. Comput. Electr. Eng. 2022, 102, 108159. [Google Scholar] [CrossRef]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J. YOLOv10: Real-Time End-to-End Object Detection. In Proceedings of the 38th International Conference on Neural Information Processing Systems (NIPS), Vancouver, BC, Canada, 10–15 December 2024; pp. 107984–108011. [Google Scholar]
- Park, T.; Efros, A.A.; Zhang, R.; Zhu, J.-Y. Contrastive Learning for Unpaired Image-to-Image Translation. In Proceedings of the 2020 European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 319–345. [Google Scholar]
- Yang, Z.; Li, Z.; Shao, M.; Shi, D.; Yuan, Z.; Yuan, C. Masked Generative Distillation. In Proceedings of the 2022 European Conference on Computer Vision (ECCV), Tel Aviv, Israel, 23–27 October 2022; pp. 53–69. [Google Scholar]
- Shu, C.; Liu, Y.; Gao, J.; Yan, Z.; Shen, C. Channel-wise Knowledge Distillation for Dense Prediction. In Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 5291–5300. [Google Scholar]
- Li, Q.; Jin, S.; Yan, J. Mimicking Very Efficient Network for Object Detection. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 7341–7349. [Google Scholar]










| Variant | Parameters (M) | GFLOPs (640 × 640) | mAP@[0.5~0.95] (%) | Primary Target Scenario |
|---|---|---|---|---|
| YOLOv10-N | 2.3–2.7 | 6–8 | 37.0–38.5 | Extreme-edge deployment: mobile devices, embedded systems, low-power real-time detection |
| YOLOv10-S | 7.0–8.0 | 20–25 | 43.0–44.5 | Lightweight real-time detection: mobile apps, UAVs, surveillance camera video streams |
| YOLOv10-M | 20.0–22.0 | 55–65 | 46.5–48.0 | General-purpose real-time detection: industrial inspection, autonomous driving perception, generic object recognition |
| YOLOv10-B | 28.0–31.0 | 85–95 | 48.5–50.0 | High-performance student model: high-accuracy real-time systems, strong student in knowledge distillation |
| YOLOv10-L | 55.0–60.0 | 160–170 | 50.0–51.5 | High-accuracy server-side inference: cloud-based analysis, offline detection, medium-scale teacher model |
| YOLOv10-X | 90.0–95.0 | 230–240 | 51.0–52.0 | Ultimate-accuracy teacher model: research benchmarking, high-value image analysis, strong teacher in distillation |
| Distillation Strategies | mAP@0.5 (%) | mAP@[0.5~0.95] (%) |
|---|---|---|
| CWD | 82.5 | 53.1 |
| MGD | 83.2 | 53.5 |
| mimic | 81.1 | 51.0 |
| Components | Style-Transfer | Feature Distillation | mAP@0.5 (%) | mAP@[0.5~0.95] (%) |
|---|---|---|---|---|
| – | – | 79.3 | 49.7 | |
| √ | – | 70.5 | 45.7 | |
| – | √ | 80.8 | 51.9 | |
| √ | √ | 83.4 | 53.7 |
| Frameworks | Input Resolution | Epoch | mAP@0.5 (%) | mAP@[0.5~0.95] (%) |
|---|---|---|---|---|
| YOLOv10-N | 640 × 640 | 100 | 55.8 ± 0.3 | 33.1 ± 0.2 |
| 200 | 70.3 ± 0.2 | 46.0 ± 0.1 | ||
| 300 | 76.6 ± 0.1 | 49.5 ± 0.2 | ||
| 400 | 79.5 ± 0.3 | 50.2 ± 0.3 | ||
| 500 | 79.5 ± 0.2 | 49.5 ± 0.1 | ||
| YOLOv10-S | 640 × 640 | 100 | 64.8 ± 0.3 | 38.4 ± 0.3 |
| 200 | 77.0 ± 0.3 | 50.0 ± 0.1 | ||
| 300 | 79.5 ± 0.1 | 52.9 ± 0.2 | ||
| 400 | 80.4 ± 0.3 | 51.5 ± 0.2 | ||
| 500 | 80.2 ± 0.2 | 52.0 ± 0.1 | ||
| YOLOv10-M | 640 × 640 | 100 | 71.5 ± 0.3 | 45.1 ± 0.3 |
| 200 | 79.4 ± 0.3 | 50.4 ± 0.1 | ||
| 300 | 80.7 ± 0.2 | 52.5 ± 0.3 | ||
| 400 | 80.7 ± 0.1 | 52.5 ± 0.1 | ||
| 500 | 81.4 ± 0.3 | 52.6 ± 0.2 | ||
| YOLOv10-B | 640 × 640 | 100 | 70.4 ± 0.1 | 43.1 ± 0.3 |
| 200 | 79.2 ± 0.2 | 49.7 ± 0.1 | ||
| 300 | 80.9 ± 0.2 | 52.0 ± 0.2 | ||
| 400 | 82.5 ± 0.3 | 52.5 ± 0.1 | ||
| 500 | 81.6 ± 0.1 | 52.7 ± 0.1 | ||
| YOLOv10-L | 640 × 640 | 100 | 68.7 ± 0.3 | 40.1 ± 0.3 |
| 200 | 80.4 ± 0.1 | 51.3 ± 0.2 | ||
| 300 | 82.4 ± 0.2 | 52.8 ± 0.1 | ||
| 400 | 81.7 ± 0.2 | 52.5 ± 0.1 | ||
| 500 | 81.6 ± 0.1 | 52.4 ± 0.1 | ||
| YOLOv10-X | 640 × 640 | 100 | 72.1 ± 0.2 | 45.6 ± 0.2 |
| 200 | 78.0 ± 0.1 | 50.0 ± 0.2 | ||
| 300 | 79.0 ± 0.1 | 51.2 ± 0.1 | ||
| 400 | 80.8 ± 0.2 | 51.6 ± 0.2 | ||
| 500 | 82.4 ± 0.1 | 53.4 ± 0.1 |
| Teacher-Model | Student-Model | Epoch | mAP@0.5 (%) | mAP@[0.5~0.95] (%) |
|---|---|---|---|---|
| YOLOv10-X | YOLOv10-N | 200 | 79.5 ± 0.3 | 50.2 ± 0.1 |
| YOLOv10-S | 200 | 81.2 ± 0.1 | 52.2 ± 0.2 | |
| YOLOv10-M | 200 | 81.8 ± 0.2 | 53.2 ± 0.1 | |
| YOLOv10-B | 200 | 83.4 ± 0.2 | 53.6 ± 0.2 | |
| YOLOv10-L | 200 | 82.6 ± 0.2 | 52.9 ± 0.1 | |
| YOLOv10-X | 200 | 83.1 ± 0.1 | 53.9 ± 0.1 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liu, W.; Zhang, S.; Zhang, S. Towards Practical Object Detection with Limited Data: A Feature Distillation Framework. J. Mar. Sci. Eng. 2026, 14, 289. https://doi.org/10.3390/jmse14030289
Liu W, Zhang S, Zhang S. Towards Practical Object Detection with Limited Data: A Feature Distillation Framework. Journal of Marine Science and Engineering. 2026; 14(3):289. https://doi.org/10.3390/jmse14030289
Chicago/Turabian StyleLiu, Wei, Shi Zhang, and Shouxu Zhang. 2026. "Towards Practical Object Detection with Limited Data: A Feature Distillation Framework" Journal of Marine Science and Engineering 14, no. 3: 289. https://doi.org/10.3390/jmse14030289
APA StyleLiu, W., Zhang, S., & Zhang, S. (2026). Towards Practical Object Detection with Limited Data: A Feature Distillation Framework. Journal of Marine Science and Engineering, 14(3), 289. https://doi.org/10.3390/jmse14030289

