State of the Art: Analysis of Deep Learning Techniques in Images Acquired in an Aquatic Environment
Abstract
1. Introduction
2. Methodology
2.1. Introduction
- RQ1: What advanced deep learning architectures have been applied to the analysis of marine species imagery captured both underwater and above water?
- RQ2: What datasets and image acquisition methods are most frequently used in these studies?
- RQ3: What evaluation metrics are reported, and how do performance levels vary across methods and application contexts?
- RQ4: What limitations, challenges, and future research directions are identified in the literature regarding the use of deep learning for marine species image analysis?
2.2. Screening
2.3. Inclusion
2.4. Data Extraction and Synthesis
3. Background
4. Types of Problems in Images and DL Applications
5. Results
5.1. Image Processing, Enhancement and Restoration
5.2. Underwater Species Detection and Classification
5.2.1. Aerial, Plankton, Aquatic Plants and Coral Detection and Classification
5.2.2. Classification of Underwater Animal Species
5.2.3. Object Detection and Segmentation Applied on Aquatic Images
5.2.4. Multimodal Understanding and Reporting
6. Discussion
7. Conclusions
Author Contributions
Funding
Acknowledgments
Conflicts of Interest
Abbreviations
| AASNet | Agricultural Aqua Segmentation Network |
| ACE | Automatic Colour Enhancement |
| ACP | Average Class Precision |
| AdaIN | Adaptive Instance Normalization |
| ADANSE ViT | Amended Dual Attention oN Self-locale and External Visual Transformer |
| AG | Average Gradient |
| AI | Artificial Intelligence |
| AIT-YOLO | Advanced imaging technique YOLO |
| ALSA | Amended Locale Self Attention |
| ANN | Artificial Neural Network |
| ASFF | Adaptively Spatial Feature Fusion |
| AUC | Area Under the Curve |
| AUROC | Receiver Operating Characteristic Area Under the Curve |
| AWS | Amazon Web Services |
| B3DO | The Berkeley 3D Object Dataset |
| BC | Blind Contrast Restoration Assessment (BCRA) |
| BC(e) | BCRA to assess the increase in edge visibility |
| BC(r) | BCRA to assess the increase in edge pixel gradient values |
| BPNN | Back Propagation Neural Network |
| C2f-D-LKA | Deformable Large Kernel Attention integrated into a C2f module (Cross-Stage Partial with full concatenation) |
| CAIT | Conv-Attentional Image Transformer |
| CBAM | Convolutional Block Attention Module |
| CCAE | Classification Convolution Autoencoder |
| CCF | Colourfulness Contrast Fog density index |
| CCL-Net | A tailored UIE approach based on cascaded contrastive learning (CCL) |
| CDR | Channel Dynamic Range |
| CE | Calibration Error |
| CFAN | Co-scale Feature Attention Network |
| CFNet | Correlation Filter Network |
| cGAN | Conditional Generative Adversarial Network |
| CIFAR10 | Canadian Institute For Advanced Research dataset |
| CLIP | Contrastive Language–Image Pretraining |
| CNN | Convolutional Neural Network |
| CNT | Convolutional Network based Tracker |
| CoaT | Co-Scale Conv-Attentional Image Transformer |
| CSA | Category-visual Semantic Alignment |
| CV | Computer Vision |
| CVAE | Conditional Variational Autoencoder |
| CycleGAN | Cycle Generative Adversarial Network |
| DAFL | Dynamic Adaptative Focal Loss |
| DAMNet | Dual Attention Mechanism based network |
| DBN | Deep Belief Network |
| DCS | Distributed Computing System |
| DDEYOLOv9 | An enhanced high-precision detection algorithm based on YOLOv9 |
| DeCAF | Deep Convolutional Activation Feature |
| DeltaE | Colour difference metric |
| DINO | Distillation with No Labels |
| DIP | Digital Image Processing |
| DL | Deep Learning |
| DNN | Deep Neural Network |
| DNnet | Dinamic range and Normalization framework |
| DPM | Deformable Part Model |
| EUPG | End-to-End Underwater Prompt Generator |
| EUVP | Enhancement of Underwater Visual Perception |
| FAN | Fast Average Normalization |
| FCN | Fully Convolutional Network |
| FCNT | Faster Convolutional Network based Tracker |
| FDCNet | Filtering Deep Convolutional Network |
| FE | Fusion Enhance |
| FLOPS | Floating-point Operations |
| FUO | Fixed Underwater Observatory |
| GAN | Generative Adversarial Networks |
| GAT | Graph Attention Network |
| GBR | Great Barrier Reef |
| GELU | Gaussian Error Linear Unit |
| GFLOPS | Giga FLOPS |
| GMG | Geometric guided Visual Mask Generator |
| GMSD | Gradient Magnitude Similarity Deviation |
| HR-Net | Haze Removal Network |
| ILSVRC | Large Scale Visual Recognition Challenge dataset |
| IMViT | Improved Vision Transformer |
| IoU | Intersection over Union |
| J&Fmean | Jaccard Index & Dice Coefficient mean |
| JAMSTEC | Japan Agency for Marine-Earth Science and Technology |
| LCA | Linear Correlation Attention |
| LifeCLEF | Life-based Conference and Labs of the Evaluation Forum |
| LLaVA | Large Language and Vision Assistant |
| LLM | Large Language Model |
| LoVe | Lofoten – Vesterålen |
| LPIPS | Learned Perceptual Image Patch Similarity |
| LSUI | Large Scale Underwater Image dataset |
| MAD | Most Apparent Distortion |
| mAP | Mean Average Precision |
| MaxViT | Multi-axis Vision Transformer |
| MBARI | Monterey Bay Aquarium Research Institute |
| MDGAN | Multiscale Dense Generative Adversarial Network |
| MDP | Markov Decision Process |
| MDPM | Mixed-Domain Periodic Motion |
| MFRF | Multiscale feature Residual Fusion module |
| MG-UKD | Mask GAT-based Underwater Knowledge Distillation |
| mIoU | Mean Intersection over Union |
| ML | Machine Learning |
| MLC | Moorea Labelled Coral dataset |
| MLLM | Multimodal Large Language Model |
| MLP | Multilayer Perceptron |
| MLR-VGG16 | Multi-Level Residual VGG16 |
| MODIS | Moderate Resolution Spectroradiometer |
| MOS | Mean Opinion Score |
| MOTA | Multi Object Tracking Accuracy |
| mRES-uNet | Multi Resolution U-net architecture network |
| MS-COCO | Microsoft Common Objects in Context dataset |
| MSE | Mean Squared Error |
| MSFN | Multi-Scale Fusion Feed-forward Network |
| MSRB | Marine Snow Removal Benchmarking Dataset |
| MSSDLite | An SSD with MobileNetV2 as its backbone |
| MSTB | Modified Swin Transformer Block |
| MUSIQ | Multi-scale Image Quality Transformer |
| NFA | Nonlinear Frequency-aware Attention |
| NIQE | Natural Image Quality Evaluator |
| NLP | Natural Language Processing |
| NMS | Non-maximum suppression |
| NOAA | U.S. National Oceanic and Atmospheric Administration |
| NWD | Normalized Gaussian Wasserstein Distance |
| NYU Depth | New York University Depth dataset |
| OFDM | Orthogonal Frequency Division Multiplexing |
| OHEM | Online Hard Example Mining |
| PAdaIN | Probabilistic Adaptive Instance Normalization |
| PAFE | Feature Pyramid Convolutional Architecture |
| PASCAL VOC | Pattern Analysis, Statistical Modelling, and Computational Learning Visual Object Classes dataset |
| PCA | Principal Component Analysis |
| PCQI | Perception-based Colour Quality Index |
| PGG-YOLO | P2P5hGhostConv-G_Ghostbottleneck-YOLO |
| PhISH-Net | Physics Inspired Network |
| PHISMID | Physics-Inspired Synthesized Marine Snow Image Dataset |
| PRISMA | Preferred Reporting Items for Systematic Reviews and Meta-Analyses |
| PSNR | Peak Signal-to-Noise Ratio |
| PSPNet | Pyramid Scene Parsing Network |
| PUIE-Net | Probabilistic Network for Underwater Image Enhancement |
| PVANet | Lightweight Deep Neural Networks for Real-time Object Detection |
| QUT | Queensland University of Technology dataset |
| RB | Retinex-based (approach) |
| R-CNN | Region-based Convolutional Network |
| RDN | Residual Dense Network |
| ReLU | Rectified Linear Unit |
| ResNet | Residual Network |
| RFEM | Residual Feature Enhancement Module |
| RGB | Red Green Blue |
| RoI | Region of Interest |
| ROV | Remotely operated Vehicle |
| RPN | Region Proposal Network |
| RUIE | Real-world Underwater Image Enhancement dataset |
| SAM | Segment Anything Models |
| SAR | Synthetic Aperture Radar |
| SAUD | Subjectively Annotated UIE benchmark dataset |
| SCCE | Sparse Categorical Cross Entropy |
| SeaCLEF | Sea Content-based Conference and Labs of the Evaluation Forum |
| SEAMAPD21 | Southeast Area Monitoring and Assessment Program Dataset 2021 |
| SegFormer | Segmentation framework which unifies Transformers with lightweight MLP decoders |
| SFD-YOLO | Seafloor-Debris-YOLO |
| SGD | Stochastic Gradient Descent |
| SGMCSS | Statistically Guided Multicolour Space Stretch module |
| SIUM | Segmentation of Underwater Imagery |
| SPL | Subaqueous Perceptual Loss |
| SRR | Super-Resolution Reconstruction |
| SSD | Single Shot Detector |
| SSIM | Structural Similarity Index Measure |
| StructureRSMAS | Structure Rosenstiel School of Marine and Atmospheric Science dataset |
| SUIM | Segmentation of Underwater Imagery dataset |
| SUN | Scene Understanding database |
| SVD | Singular Value Decomposition |
| SVM | Support Vector Machine |
| TLD | Tracking Learning Detection |
| U45 | Public underwater test dataset with 45 images |
| UAV | Unmanned Aerial Vehicle |
| UCIQE | Underwater Colour Image Quality Evaluation |
| UCM | Unsupervised Colour Correction |
| UDAE | Underwater Denoising Autoencoder |
| UDCP | Underwater Dark Channel Prior |
| UDnet | Uncertainty Distribution Network |
| UFO-120 | Dataset for Simultaneous Enhancement and Super-Resolution (SESR) of underwater imagery |
| UGAN | Underwater Generative Adversarial Network |
| UGAN-P | UGAN with Gradient Difference Loss |
| UICM | Underwater Image Colourfulness Measure |
| UIConM | Underwater Image Contrast Measure |
| UIEB | Underwater Image Enhancement Benchmark |
| UIEVUS | Underwater Image Enhancement method designed for Various Underwater Scenes |
| UIIS | Underwater Image Instance Segmentation dataset |
| UIIS10K | Underwater Image Instance Segmentation 10K dataset |
| UIQM | Underwater Image Quality Measure |
| UISD | Underwater Images Segmentation dataset |
| UISFormer | Underwater Image Segmentation Transformer model |
| UMAP | Uniform Manifold Approximation and Projection |
| U-Net-scSE | U-net with Spatial Channel Squeeze-and-Excitation module |
| UOVSBench | Underwater Open-Vocabulary Segmentation Benchmark |
| URanker | Ranking-based underwater image quality assessment |
| URSCT-SESR | U-Net-based reinforced Swin-Convs Transformer for simultaneous enhancement and superresolution |
| USIS10K | Underwater Salient Instance Segmentation 10K dataset |
| UVEB | Underwater Video Enhancement Benchmark |
| UVLM | Underwater Video Language Model |
| UW RGB-D Object | Underwater RGB-Depth Object dataset |
| UWFormer | Multi-scale Transformer-based Network |
| UWSAM | Segment Anything Model Guided Underwater Instance Segmentation |
| VGG | Visual Geometry Group model |
| VIAME | Video and Image Analysis for Marine Environments |
| VidLM | Video-Language Model |
| ViT | Vision Transformers |
| VLM | Vision-Language Model |
| WaterGan | Water Generative Adversarial Network |
| WSCT | Weakly Supervised Colour Transfer |
| YOLO | You Only Look Once |
| YOLOv8-TF | Transformer-enhanced YOLOv8 |
| ZF | Zeiler and Fergus model |
References
- Aguzzi, J.; Chatzievangelou, D.; Marini, S.; Fanelli, E.; Danovaro, R.; Flögel, S.; Lebris, N.; Juanes, F.; De Leo, F.C.; Del Rio, J.; et al. New High-Tech Flexible Networks for the Monitoring of Deep-Sea Ecosystems. Environ. Sci. Technol. 2019, 53, 6616–6631. [Google Scholar] [CrossRef]
- Favali, P.; Beranzoli, L.; De Santis, A. SEAFLOOR OBSERVATORIES: A New Vision of the Earth from the Abyss; Springer Science & Business Media: New York, NY, USA, 2015. [Google Scholar]
- Schoening, T.; Bergmann, M.; Ontrup, J.; Taylor, J.; Dannheim, J.; Gutt, J.; Purser, A.; Nattkemper, T.W. Semi-Automated Image Analysis for the Assessment of Megafaunal Densities at the Arctic Deep-Sea Observatory HAUSGARTEN. PLoS ONE 2012, 7, e38179. [Google Scholar] [CrossRef] [PubMed]
- Aguzzi, J.; Doya, C.; Tecchio, S.; De Leo, F.; Azzurro, E.; Costa, C.; Sbragaglia, V.; Del Río, J.; Navarro, J.; Ruhl, H.; et al. Coastal Observatories for Monitoring of Fish Behaviour and Their Responses to Environmental Changes. Rev. Fish Biol. Fish. 2015, 25, 463–483. [Google Scholar] [CrossRef]
- Widder, E.; Robison, B.H.; Reisenbichler, K.; Haddock, S. Using Red Light for In Situ Observations of Deep-Sea Fishes. Deep Sea Res. Part I Oceanogr. Res. Pap. 2005, 52, 2077–2085. [Google Scholar] [CrossRef]
- Chauvet, P.; Metaxas, A.; Hay, A.E.; Matabos, M. Annual and Seasonal Dynamics of Deep-Sea Megafaunal Epibenthic Communities in Barkley Canyon (British Columbia, Canada): A Response to Climatology, Surface Productivity and Benthic Boundary Layer Variation. In Proceedings of the Progress in Oceanography; Elsevier: Amsterdam, The Netherlands, 2018; Volume 169, pp. 89–105. [Google Scholar]
- Leo, F.D.; Ogata, B.; Sastri, A.R.; Heesemann, M.; Mihály, S.; Galbraith, M.; Morley, M. High-Frequency Observations from a Deep-Sea Cabled Observatory Reveal Seasonal Overwintering of Neocalanus Spp. in Barkley Canyon, NE Pacific: Insights into Particulate Organic Carbon Flux. Prog. Oceanogr. 2018, 169, 120–137. [Google Scholar] [CrossRef]
- Aguzzi, J.; Costa, C.; Matabos, M.; Azzurro, E.; Lázaro, A.; Menesatti, P.; Sarda, F.; Canals, M.; Delory, E.; Cline, D.; et al. Challenges to the Assessment of Benthic Populations and Biodiversity as a Result of Rhythmic Behaviour: Video Solutions from Cabled Observatories. Oceanogr. Mar. Biol. 2012, 50, 235–286. [Google Scholar]
- Bicknell, A.W.; Godley, B.J.; Sheehan, E.V.; Votier, S.C.; Witt, M.J. Camera Technology for Monitoring Marine Biodiversity and Human Impact. Front. Ecol. Environ. 2016, 14, 424–432. [Google Scholar] [CrossRef]
- Danovaro, R.; Aguzzi, J.; Fanelli, E.; Billett, D.; Gjerde, K.; Jamieson, A.; Ramirez-Llodra, E.; Smith, C.; Snelgrove, P.; Thomsen, L.; et al. An Ecosystem-Based Deep-Ocean Strategy. Science 2017, 355, 452–454. [Google Scholar] [CrossRef] [PubMed]
- Szeliski, R. Computer Vision: Algorithms and Applications; Springer Science & Business Media: New York, NY, USA, 2010. [Google Scholar]
- Garcia, R.; Nicosevici, T.; Cufí, X. On the Way to Solve Lighting Problems in Underwater Imaging. In Proceedings of the OCEANS’02 MTS/IEEE; IEEE: New York, NY, USA, 2002; Volume 2, pp. 1018–1024. [Google Scholar]
- Prabhakar, C.; Kumar, P. An Image Based Technique for Enhancement of Underwater Images. arXiv 2012, arXiv:1212.0291. [Google Scholar] [CrossRef]
- Raj, M.V.; Murugan, S.S. Underwater Image Classification Using Machine Learning Technique. In Proceedings of the 2019 International Symposium on Ocean Technology (SYMPOL); IEEE: New York, NY, USA, 2019; pp. 166–173. [Google Scholar]
- Lippmann, R.P. Pattern Classification Using Neural Networks. IEEE Commun. Mag. 1989, 27, 47–50. [Google Scholar] [CrossRef]
- Liu, H.; Ma, X.; Yu, Y.; Wang, L.; Hao, L. Application of Deep Learning-Based Object Detection Techniques in Fish Aquaculture: A Review. J. Mar. Sci. Eng. 2023, 11, 867. [Google Scholar] [CrossRef]
- Zhao, S.; Zhang, S.; Liu, J.; Wang, H.; Zhu, J.; Li, D.; Zhao, R. Application of Machine Learning in Intelligent Fish Aquaculture: A Review. Aquaculture 2021, 540, 736724. [Google Scholar] [CrossRef]
- Alsmadi, M.K.; Almarashdeh, I. A Survey on Fish Classification Techniques. J. King Saud Univ.-Comput. Inf. Sci. 2022, 34, 1625–1638. [Google Scholar] [CrossRef]
- Xu, Z.; Wang, T.; Skidmore, A.K.; Lamprey, R. A Review of Deep Learning Techniques for Detecting Animals in Aerial and Satellite Images. Int. J. Appl. Earth Obs. Geoinf. 2024, 128, 103732. [Google Scholar] [CrossRef]
- Arsad, T.; Awalludin, E.; Bachok, Z.; Yussof, W.; Hitam, M. A Review of Coral Reef Classification Study Using Deep Learning Approach. In Proceedings of the AIP Conference Proceedings; AIP Publishing: New York, NY, USA, 2023; Volume 2484. [Google Scholar]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. BMJ 2021, 372, 71. [Google Scholar] [CrossRef]
- Tricco, A.C.; Lillie, E.; Zarin, W.; O’Brien, K.K.; Colquhoun, H.; Levac, D.; Moher, D.; Peters, M.D.; Horsley, T.; Weeks, L.; et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Ann. Intern. Med. 2018, 169, 467–473. [Google Scholar] [CrossRef]
- McCulloch, W.S.; Pitts, W. A Logical Calculus of the Ideas Immanent in Nervous Activity. Bull. Math. Biophys. 1943, 5, 115–133. [Google Scholar] [CrossRef]
- Yegnanarayana, B. Artificial Neural Networks; PHI Learning Pvt. Ltd.: Delhi, India, 2009. [Google Scholar]
- Rosenblatt, F. The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain. Psychol. Rev. 1958, 65, 386. [Google Scholar] [CrossRef]
- Hopfield, J.J. Neural Networks and Physical Systems with Emergent Collective Computational Abilities. Proc. Natl. Acad. Sci. USA 1982, 79, 2554–2558. [Google Scholar] [CrossRef]
- Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning Representations by Back-Propagating Errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef]
- Ciregan, D.; Meier, U.; Schmidhuber, J. Multi-Column Deep Neural Networks for Image Classification. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2012; pp. 3642–3649. [Google Scholar]
- Chung, C.-L.; Huang, K.-J.; Chen, S.-Y.; Lai, M.-H.; Chen, Y.-C.; Kuo, Y.-F. Detecting Bakanae Disease in Rice Seedlings by Machine Vision. Comput. Electron. Agric. 2016, 121, 404–411. [Google Scholar] [CrossRef]
- Nguyen, T.T.; Hoang, T.D.; Pham, M.T.; Vu, T.T.; Nguyen, T.H.; Huynh, Q.-T.; Jo, J. Monitoring Agriculture Areas with Satellite Images and Deep Learning. Appl. Soft Comput. 2020, 95, 106565. [Google Scholar] [CrossRef]
- Yamamoto, K.; Guo, W.; Yoshioka, Y.; Ninomiya, S. On Plant Detection of Intact Tomato Fruits Using Image Analysis and Machine Learning Methods. Sensors 2014, 14, 12191–12206. [Google Scholar] [CrossRef] [PubMed]
- Haug, S.; Ostermann, J. A Crop/Weed Field Image Dataset for the Evaluation of Computer Vision Based Precision Agriculture Tasks. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2014; pp. 105–116. [Google Scholar]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer: Berlin/Heidelberg, Germany, 2015; pp. 234–241. [Google Scholar]
- Shi, F.; Wang, J.; Shi, J.; Wu, Z.; Wang, Q.; Tang, Z.; He, K.; Shi, Y.; Shen, D. Review of Artificial Intelligence Techniques in Imaging Data Acquisition, Segmentation and Diagnosis for COVID-19. IEEE Rev. Biomed. Eng. 2020, 14, 4–15. [Google Scholar] [CrossRef]
- Liu, S.; Liu, S.; Cai, W.; Pujol, S.; Kikinis, R.; Feng, D. Early Diagnosis of Alzheimer’s Disease with Deep Learning. In Proceedings of the 2014 IEEE 11th International Symposium on Biomedical Imaging (ISBI); IEEE: New York, NY, USA, 2014; pp. 1015–1018. [Google Scholar]
- Criminisi, A. Machine Learning for Medical Images Analysis. Med. Image Anal. 2016, 33, 91–93. [Google Scholar] [CrossRef]
- Jang, H.S.; Bae, K.Y.; Park, H.-S.; Sung, D.K. Solar Power Prediction Based on Satellite Images and Support Vector Machine. IEEE Trans. Sustain. Energy 2016, 7, 1255–1263. [Google Scholar] [CrossRef]
- Taravat, A.; Del Frate, F.; Cornaro, C.; Vergari, S. Neural Networks and Support Vector Machine Algorithms for Automatic Cloud Classification of Whole-Sky Ground-Based Images. IEEE Geosci. Remote Sens. Lett. 2014, 12, 666–670. [Google Scholar] [CrossRef]
- Wardah, T.; Bakar, S.A.; Bardossy, A.; Maznorizan, M. Use of Geostationary Meteorological Satellite Images in Convective Rain Estimation for Flash-Flood Forecasting. J. Hydrol. 2008, 356, 283–298. [Google Scholar] [CrossRef]
- Schalkoff, R.J. Digital Image Processing and Computer Vision; Wiley: New York, NY, USA, 1989; Volume 286. [Google Scholar]
- Lu, D.; Weng, Q. A Survey of Image Classification Methods and Techniques for Improving Classification Performance. Int. J. Remote Sens. 2007, 28, 823–870. [Google Scholar] [CrossRef]
- Viola, P.; Jones, M. Rapid Object Detection Using a Boosted Cascade of Simple Features. In Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition; CVPR 2001; IEEE: New York, NY, USA, 2001; Volume 1, p. I–I. [Google Scholar]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2014; pp. 580–587. [Google Scholar]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2016; pp. 779–788. [Google Scholar]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. Ssd: Single Shot Multibox Detector. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2016; pp. 21–37. [Google Scholar]
- Gordan, M.; Dancea, O.; Stoian, I.; Georgakis, A.; Tsatos, O. A New SVM-Based Architecture for Object Recognition in Color Underwater Images with Classification Refinement by Shape Descriptors. In Proceedings of the 2006 IEEE International Conference on Automation, Quality and Testing, Robotics; IEEE: New York, NY, USA, 2006; Volume 2, pp. 327–332. [Google Scholar]
- Bishop, C.M. Pattern Recognition and Machine Learning; Information Science and Statistics; Springer: Secaucus, NJ, USA, 2006. [Google Scholar]
- Zion, B. The Use of Computer Vision Technologies in Aquaculture—A Review. Comput. Electron. Agric. 2012, 88, 125–132. [Google Scholar] [CrossRef]
- Osterloff, J.; Nilssen, I.; Nattkemper, T.W. Computational Coral Feature Monitoring for the Fixed Underwater Observatory LoVe. In Proceedings of the OCEANS 2016 MTS/IEEE Monterey; IEEE: New York, NY, USA, 2016; pp. 1–5. [Google Scholar]
- Villon, S.; Chaumont, M.; Subsol, G.; Villéger, S.; Claverie, T.; Mouillot, D. Coral Reef Fish Detection and Recognition in Underwater Videos by Supervised Machine Learning: Comparison between Deep Learning and HOG+ SVM Methods. In Proceedings of the International Conference on Advanced Concepts for Intelligent Vision Systems; Springer: Berlin/Heidelberg, Germany, 2016; pp. 160–171. [Google Scholar]
- Kitasato, A.; Miyazaki, T.; Sugaya, Y.; Omachi, S. Automatic Discrimination between Scomber Japonicus and Scomber Australasicus by Geometric and Texture Features. Fishes 2018, 3, 26. [Google Scholar] [CrossRef]
- Saberioon, M.; Císař, P.; Labbé, L.; Souček, P.; Pelissier, P.; Kerneis, T. Comparative Performance Analysis of Support Vector Machine, Random Forest, Logistic Regression and k-Nearest Neighbours in Rainbow Trout (Oncorhynchus Mykiss) Classification Using Image-Based Features. Sensors 2018, 18, 1027. [Google Scholar] [CrossRef]
- Freitas, U.; Gonçalves, W.N.; Matsubara, E.T.; Sabino, J.; Borth, M.R.; Pistori, H. Using Color for Fish Species Classification. In Proceedings of the Workshop of Industry Applications (WIA), SIBGRAPI, São José dos Campos, Brazil, 4–7 October 2016. [Google Scholar]
- Moniruzzaman, M.; Islam, S.M.S.; Bennamoun, M.; Lavery, P. Deep Learning on Underwater Marine Object Detection: A Survey. In Proceedings of the International Conference on Advanced Concepts for Intelligent Vision Systems; Springer: Berlin/Heidelberg, Germany, 2017; pp. 150–160. [Google Scholar]
- Sun, M.; Yang, X.; Xie, Y. Deep Learning in Aquaculture: A Review. J. Comput. 2020, 31, 294–319. [Google Scholar]
- Dong, G.; Ma, Y.; Basu, A. Feature-Guided CNN for Denoising Images from Portable Ultrasound Devices. IEEE Access 2021, 9, 28272–28281. [Google Scholar] [CrossRef]
- Li, C.; Guo, C.; Ren, W.; Cong, R.; Hou, J.; Kwong, S.; Tao, D. An Underwater Image Enhancement Benchmark Dataset and Beyond. IEEE Trans. Image Process. 2019, 29, 4376–4389. [Google Scholar] [CrossRef]
- Xie, Y.; Kong, L.; Chen, K.; Zheng, Z.; Yu, X.; Yu, Z.; Zheng, B. Uveb: A Large-Scale Benchmark and Baseline towards Real-World Underwater Video Enhancement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2024; pp. 22358–22367. [Google Scholar]
- Li, H.; Li, J.; Wang, W. A Fusion Adversarial Underwater Image Enhancement Network with a Public Test Dataset. arXiv 2019, arXiv:1906.06819. [Google Scholar] [CrossRef]
- Islam, M.J.; Xia, Y.; Sattar, J. Fast Underwater Image Enhancement for Improved Visual Perception. IEEE Robot. Autom. Lett. 2020, 5, 3227–3234. [Google Scholar] [CrossRef]
- Perez, J.; Attanasio, A.C.; Nechyporenko, N.; Sanz, P.J. A Deep Learning Approach for Underwater Image Enhancement. In Proceedings of the International Work-Conference on the Interplay Between Natural and Artificial Computation; Springer: Berlin/Heidelberg, Germany, 2017; pp. 183–192. [Google Scholar]
- Getreuer, P. Automatic Color Enhancement (ACE) and Its Fast Implementation. Image Process. Line 2012, 2, 266–277. [Google Scholar] [CrossRef]
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Nets. In Proceedings of the Advances in Neural Information Processing Systems, Montréal, QC, Canada, 8–12 December 2014; pp. 2672–2680. [Google Scholar]
- Yang, M.; Hu, K.; Du, Y.; Wei, Z.; Sheng, Z.; Hu, J. Underwater Image Enhancement Based on Conditional Generative Adversarial Network. Signal Process. Image Commun. 2020, 81, 115723. [Google Scholar] [CrossRef]
- Fabbri, C.; Islam, M.J.; Sattar, J. Enhancing Underwater Imagery Using Generative Adversarial Networks. In Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2018; pp. 7159–7165. [Google Scholar]
- Zhu, J.-Y.; Park, T.; Isola, P.; Efros, A.A. Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 2223–2232. [Google Scholar]
- Islam, M.J.; Sattar, J. Mixed-Domain Biological Motion Tracking for Underwater Human-Robot Interaction. In Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2017; pp. 4457–4464. [Google Scholar]
- Li, J.; Skinner, K.A.; Eustice, R.M.; Johnson-Roberson, M. WaterGAN: Unsupervised Generative Network to Enable Real-Time Color Correction of Monocular Underwater Images. IEEE Robot. Autom. Lett. 2017, 3, 387–394. [Google Scholar] [CrossRef]
- Shin, Y.-S.; Cho, Y.; Pandey, G.; Kim, A. Estimation of Ambient Light and Transmission Map with Common Convolutional Architecture. In Proceedings of the OCEANS 2016 MTS/IEEE Monterey; IEEE: New York, NY, USA, 2016; pp. 1–7. [Google Scholar]
- Schmidhuber, J. Deep Learning in Neural Networks: An Overview. Neural Netw. 2015, 61, 85–117. [Google Scholar] [CrossRef]
- Hashisho, Y.; Albadawi, M.; Krause, T.; von Lukas, U.F. Underwater Color Restoration Using U-Net Denoising Autoencoder. In Proceedings of the 2019 11th International Symposium on Image and Signal Processing and Analysis (ISPA); IEEE: New York, NY, USA, 2019; pp. 117–122. [Google Scholar]
- Li, Y.; Lu, H.; Li, J.; Li, X.; Li, Y.; Serikawa, S. Underwater Image De-Scattering and Classification by Deep Neural Network. Comput. Electr. Eng. 2016, 54, 68–77. [Google Scholar] [CrossRef]
- Sheikh, H.R.; Bovik, A.C. 8.4—Information Theoretic Approaches to Image Quality Assessment. In Handbook of Image and Video Processing, 2nd ed.; BOVIK, A., Ed.; Communications, Networking and Multimedia; Academic Press: Burlington, MA, USA, 2005; pp. 975–989. ISBN 978-0-12-119792-6. [Google Scholar]
- Fu, Z.; Wang, W.; Huang, Y.; Ding, X.; Ma, K.-K. Uncertainty Inspired Underwater Image Enhancement. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 465–482. [Google Scholar]
- Rezende, D.J.; Mohamed, S.; Wierstra, D. Stochastic Backpropagation and Approximate Inference in Deep Generative Models. In Proceedings of the International Conference on Machine Learning, Beijing, China, 21–26 June 2014; pp. 1278–1286. [Google Scholar]
- Huang, X.; Belongie, S. Arbitrary Style Transfer in Real-Time with Adaptive Instance Normalization. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 1501–1510. [Google Scholar]
- Liu, R.; Fan, X.; Zhu, M.; Hou, M.; Luo, Z. Real-World Underwater Enhancement: Challenges, Benchmarks, and Solutions under Natural Light. IEEE Trans. Circuits Syst. Video Technol. 2020, 30, 4861–4875. [Google Scholar] [CrossRef]
- Mittal, A.; Soundararajan, R.; Bovik, A.C. Making a “Completely Blind” Image Quality Analyzer. IEEE Signal Process. Lett. 2012, 20, 209–212. [Google Scholar] [CrossRef]
- Sun, S.; Wang, H.; Zhang, H.; Li, M.; Xiang, M.; Luo, C.; Ren, P. Underwater Image Enhancement with Reinforcement Learning. IEEE J. Ocean. Eng. 2022, 49, 249–261. [Google Scholar] [CrossRef]
- Yang, M.; Sowmya, A. An Underwater Color Image Quality Evaluation Metric. IEEE Trans. Image Process. 2015, 24, 6062–6071. [Google Scholar] [CrossRef]
- Panetta, K.; Gao, C.; Agaian, S. Human-Visual-System-Inspired Underwater Image Quality Measures. IEEE J. Ocean. Eng. 2016, 41, 541–551. [Google Scholar] [CrossRef]
- Guo, Y.; Li, H.; Zhuang, P. Underwater Image Enhancement Using a Multiscale Dense Generative Adversarial Network. IEEE J. Ocean. Eng. 2020, 45, 862–870. [Google Scholar] [CrossRef]
- Cao, T.; Yu, Z.; Zheng, B. DNnet: A Lightweight Network for Real-Time 4K Underwater Image Enhancement Using Dynamic Range and Average Normalization. Expert Syst. Appl. 2025, 270, 126561. [Google Scholar] [CrossRef]
- Ren, S.; Bao, X.; Wang, T.; Xu, X.; Ma, T.; Yu, K. UIEVUS: An Underwater Image Enhancement Method for Various Underwater Scenes. Signal Process. Image Commun. 2025, 270, 117264. [Google Scholar] [CrossRef]
- Liu, Y.; Jiang, Q.; Wang, X.; Luo, T.; Zhou, J. Underwater Image Enhancement with Cascaded Contrastive Learning. IEEE Trans. Multimed. 2025, 27, 1512–1525. [Google Scholar] [CrossRef]
- Zhu, S.; Geng, Z.; Xie, Y.; Zhang, Z.; Yan, H.; Zhou, X.; Jin, H.; Fan, X. New Underwater Image Enhancement Algorithm Based on Improved U-Net. Water 2025, 17, 808. [Google Scholar] [CrossRef]
- Saleh, A.; Sheaves, M.; Jerry, D.; Azghadi, M.R. Adaptive Deep Learning Framework for Robust Unsupervised Underwater Image Enhancement. Expert Syst. Appl. 2025, 268, 126314. [Google Scholar] [CrossRef]
- Vaswani, A. Attention Is All You Need. In Advances in Neural Information Processing Systems 30 (NeurIPS 2017); Curran Associates, Inc.: Red Hook, NY, USA, 2017; pp. 5998–6008. [Google Scholar]
- Dosovitskiy, A. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 10012–10022. [Google Scholar]
- Ji, J.; Man, J. UNet–Transformer Hybrid Architecture for Enhanced Underwater Image Processing and Restoration. Mathematics 2025, 13, 2535. [Google Scholar] [CrossRef]
- Chen, W.; Lei, Y.; Luo, S.; Zhou, Z.; Li, M.; Pun, C.-M. Uwformer: Underwater Image Enhancement via a Semi-Supervised Multi-Scale Transformer. In Proceedings of the 2024 International Joint Conference on Neural Networks (IJCNN); IEEE: New York, NY, USA, 2024; pp. 1–8. [Google Scholar]
- Wang, H.; Köser, K.; Ren, P. Large Foundation Model Empowered Discriminative Underwater Image Enhancement. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5609317. [Google Scholar] [CrossRef]
- Janoch, A.; Karayev, S.; Jia, Y.; Barron, J.T.; Fritz, M.; Saenko, K.; Darrell, T. A Category-Level 3d Object Dataset: Putting the Kinect to Work. In Consumer Depth Cameras for Computer Vision; Springer: Berlin/Heidelberg, Germany, 2013; pp. 141–165. [Google Scholar]
- Lai, K.; Bo, L.; Fox, D. Unsupervised Feature Learning for 3d Scene Labeling. In Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2014; pp. 3050–3057. [Google Scholar]
- Silberman, N.; Fergus, R. Indoor Scene Segmentation Using a Structured Light Sensor. In Proceedings of the 2011 IEEE International Conference on Computer Vision Workshops (ICCV Workshops); IEEE: New York, NY, USA, 2011; pp. 601–608. [Google Scholar]
- Shotton, J.; Glocker, B.; Zach, C.; Izadi, S.; Criminisi, A.; Fitzgibbon, A. Scene Coordinate Regression Forests for Camera Relocalization in RGB-D Images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2013; pp. 2930–2937. [Google Scholar]
- Xiao, J.; Hays, J.; Ehinger, K.A.; Oliva, A.; Torralba, A. Sun Database: Large-Scale Scene Recognition from Abbey to Zoo. In Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2010; pp. 3485–3492. [Google Scholar]
- Peng, L.; Zhu, C.; Bian, L. U-Shape Transformer for Underwater Image Enhancement. IEEE Trans. Image Process. 2023, 32, 3066–3079. [Google Scholar] [CrossRef]
- Saleh, A.; Laradji, I.H.; Konovalov, D.A.; Bradley, M.; Vazquez, D.; Sheaves, M. A Realistic Fish-Habitat Dataset to Evaluate Algorithms for Underwater Visual Analysis. Sci. Rep. 2020, 10, 14671. [Google Scholar] [CrossRef]
- Islam, M.J.; Edge, C.; Xiao, Y.; Luo, P.; Mehtaz, M.; Morse, C.; Enan, S.S.; Sattar, J. Semantic Segmentation of Underwater Imagery: Dataset and Benchmark. In Proceedings of the 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2020; pp. 1769–1776. [Google Scholar]
- Berman, D.; Levy, D.; Avidan, S.; Treibitz, T. Underwater Single Image Color Restoration Using Haze-Lines and a New Quantitative Dataset. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 43, 2822–2837. [Google Scholar] [CrossRef] [PubMed]
- Kratzert, F.; Mader, H. Fish Species Classification in Underwater Video Monitoring Using Convolutional Neural Networks. Earth ArXiv 2018. [Google Scholar] [CrossRef]
- Liang, J.; Fu, Z.; Lei, X.; Dai, X.; Lv, B. Recognition and Classification of Ornamental Fish Image Based on Machine Vision. In Proceedings of the 2020 International Conference on Intelligent Transportation, Big Data Smart City (ICITBS); IEEE Computer Society: Washington, DC, USA, 2020; pp. 910–913. [Google Scholar]
- Cao, Z.; Principe, J.C.; Ouyang, B.; Dalgleish, F.; Vuorenkoski, A. Marine Animal Classification Using Combined CNN and Hand-Designed Image Features. In Proceedings of the OCEANS’15 MTS/IEEE Washington; IEEE: New York, NY, USA, 2015; pp. 1–6. [Google Scholar]
- Mahmood, A.; Bennamoun, M.; An, S.; Sohel, F.; Boussaid, F.; Hovey, R.; Kendrick, G.; Fisher, R.B. Automatic Annotation of Coral Reefs Using Deep Learning. In Proceedings of the Oceans 2016 MTS/IEEE Monterey; IEEE: New York, NY, USA, 2016; pp. 1–5. [Google Scholar]
- Meng, L.; Hirayama, T.; Oyanagi, S. Underwater-Drone with Panoramic Camera for Automatic Fish Recognition Based on Deep Learning. IEEE Access 2018, 6, 17880–17886. [Google Scholar] [CrossRef]
- Rahmat, B.; Waluyo, M.; Rachmanto, T.A.; Afandi, M.I.; Widyantara, H.; Harianto. Video-Based Tancho Koi Fish Tracking System Using CSK, DFT, and LOT. J. Phys. Conf. Ser. 2020, 1569, 022036. [Google Scholar] [CrossRef]
- Rossi, F.; Benso, A.; Carlo, S.D.; Politano, G.; Savino, A.; Acutis, P.L. FishAPP: A Mobile App to Detect Fish Falsification through Image Processing and Machine Learning Techniques. In Proceedings of the 20th IEEE International Conference on Automation, Quality and Testing, Robotics (AQTR 2016), Cluj-Napoca, Romania, 19–21 May 2016; pp. 1–6. [Google Scholar]
- Kutlu, Y.; Altan, G.; İççimen, B.; Doğdu, S.A.; Turan, C. Recognition of Species of Triglidae Family Using Deep Learning. J. Black Sea/Mediterr. Environ. 2017, 23, 56–65. [Google Scholar]
- Banan, A.; Nasiri, A.; Taheri-Garavand, A. Deep Learning-Based Appearance Features Extraction for Automated Carp Species Identification. Aquac. Eng. 2020, 89, 102053. [Google Scholar] [CrossRef]
- Pudaruth, S.; Nazurally, N.; Appadoo, C.; Kishnah, S.; Vinayaganidhi, M.; Mohammoodally, I.; Ally, Y.A.; Chady, F. SuperFish: A Mobile Application for Fish Species Recognition Using Image Processing Techniques and Deep Learning. Int. J. Comput. Digit. Syst. 2020, 10, 1157–1165. [Google Scholar] [CrossRef]
- Cao, S.; Zhao, D.; Liu, X.; Sun, Y. Real-Time Robust Detector for Underwater Live Crabs Based on Deep Learning. Comput. Electron. Agric. 2020, 172, 105339. [Google Scholar] [CrossRef]
- Naddaf-Sh, M.; Myler, H.; Zargarzadeh, H. Design and Implementation of an Assistive Real-Time Red Lionfish Detection System for AUV/ROVs. Complexity 2018, 2018, 5298294. [Google Scholar] [CrossRef]
- Dawkins, M.; Sherrill, L.; Fieldhouse, K.; Hoogs, A.; Richards, B.; Zhang, D.; Prasad, L.; Williams, K.; Lauffenburger, N.; Wang, G. An Open-Source Platform for Underwater Image and Video Analytics. In Proceedings of the Applications of Computer Vision (WACV), 2017 IEEE Winter Conference on Applications of Computer Vision (WACV); IEEE: New York, NY, USA, 2017; pp. 898–906. [Google Scholar]
- Mandal, R.; Connolly, R.M.; Schlacherz, T.A.; Stantic, B. Assessing Fish Abundance from Underwater Video Using Deep Neural Networks. arXiv 2018, arXiv:1807.05838. [Google Scholar] [CrossRef]
- Labao, A.B.; Naval, P.C. Weakly-Labelled Semantic Segmentation of Fish Objects in Underwater Videos Using a Deep Residual Network. In Proceedings of the Asian Conference on Intelligent Information and Database Systems; Springer: Berlin/Heidelberg, Germany, 2017; pp. 255–265. [Google Scholar]
- Martin-Abadal, M.; Ruiz-Frau, A.; Hinz, H.; Gonzalez-Cid, Y. Jellytoring: Real-Time Jellyfish Monitoring Based on Deep Learning Object Detection. Sensors 2020, 20, 1708. [Google Scholar] [CrossRef] [PubMed]
- Lu, H.; Li, Y.; Uemura, T.; Ge, Z.; Xu, X.; He, L.; Serikawa, S.; Kim, H. FDCNet: Filtering Deep Convolutional Network for Marine Organism Classification. Multimed. Tools Appl. 2017, 77, 21847–21860. [Google Scholar] [CrossRef]
- Rimavicius, T.; Gelzinis, A. A Comparison of the Deep Learning Methods for Solving Seafloor Image Classification Task. In Proceedings of the International Conference on Information and Software Technologies; Springer: Berlin/Heidelberg, Germany, 2017; pp. 442–453. [Google Scholar]
- Osterloff, J.; Nilssen, I.; Nattkemper, T.W. A Computer Vision Approach for Monitoring the Spatial and Temporal Shrimp Distribution at the LoVe Observatory. Methods Oceanogr. 2016, 15, 114–128. [Google Scholar] [CrossRef]
- Møller, T.; Nillsen, I.; Nattkemper, T.W. Active Learning for the Classification of Species in Underwater Images from a Fixed Observatory. In Proceedings of the IEEE International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2017. [Google Scholar]
- Chen, C.-H.; Liu, K.-H. Stingray Detection of Aerial Images with Region-Based Convolution Neural Network. In Proceedings of the Consumer Electronics-Taiwan (ICCE-TW), 2017 IEEE International Conference on Consumer Electronics-Taiwan (ICCE-TW); IEEE: New York, NY, USA, 2017; pp. 175–176. [Google Scholar]
- Levy, D.; Belfer, Y.; Osherov, E.; Bigal, E.; Scheinin, A.P.; Nativ, H.; Tchernov, D.; Treibitz, T. Automated Analysis of Marine Video with Limited Data. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops; IEEE: New York, NY, USA, 2018; pp. 1385–1393. [Google Scholar]
- Guirado, E.; Tabik, S.; Rivas, M.L.; Alcaraz-Segura, D.; Herrera, F. Whale Counting in Satellite and Aerial Images with Deep Learning. Sci. Rep. 2019, 9, 14259. [Google Scholar] [CrossRef]
- Dujon, A.M.; Ierodiaconou, D.; Geeson, J.J.; Arnould, J.P.; Allan, B.M.; Katselidis, K.A.; Schofield, G. Machine Learning to Detect Marine Animals in UAV Imagery: Effect of Morphology, Spacing, Behaviour and Habitat. Remote Sens. Ecol. Conserv. 2021, 7, 341–354. [Google Scholar] [CrossRef]
- Boulent, J.; Charry, B.; Kennedy, M.M.; Tissier, E.; Fan, R.; Marcoux, M.; Watt, C.A.; Gagné-Turcotte, A. Scaling Whale Monitoring Using Deep Learning: A Human-in-the-Loop Solution for Analyzing Aerial Datasets. Front. Mar. Sci. 2023, 10, 1099479. [Google Scholar] [CrossRef]
- Patel, M.; Chen, X.; Xu, L.; Cantu, F.J.P.; Turnes, J.N.; Brubacher, N.C.; Clausi, D.A.; Scott, K.A. The Influence of Input Image Scale on Deep Learning-Based Beluga Whale Detection from Aerial Remote Sensing Imagery. In Proceedings of the IGARSS 2023—2023 IEEE International Geoscience and Remote Sensing Symposium; IEEE: New York, NY, USA, 2023; pp. 5732–5734. [Google Scholar]
- Kong, M.; Liu, Y.; Li, B.; Duan, Q. A Lightweight Method for Detecting Turned White Belly Fish in Ponds Using Unmanned Aerial Vehicle Imagery. Eng. Appl. Artif. Intell. 2025, 144, 110111. [Google Scholar] [CrossRef]
- Lee, H.; Park, M.; Kim, J. Plankton Classification on Imbalanced Large Scale Database via Convolutional Neural Networks with Transfer Learning. In Proceedings of the 2016 IEEE International Conference on Image Processing (ICIP); IEEE: New York, NY, USA, 2016; pp. 3713–3717. [Google Scholar]
- Orenstein, E.C.; Beijbom, O.; Peacock, E.E.; Sosik, H.M. Whoi-Plankton-a Large Scale Fine Grained Visual Recognition Benchmark Dataset for Plankton Classification. arXiv 2015, arXiv:1510.00745. [Google Scholar]
- Py, O.; Hong, H.; Zhongzhi, S. Plankton Classification with Deep Convolutional Neural Networks. In Proceedings of the 2016 IEEE Information Technology, Networking, Electronic and Automation Control Conference; IEEE: New York, NY, USA, 2016; pp. 132–136. [Google Scholar]
- Cowen, R.K.; Sponaugle, S.; Robinson, K.; Luo, J. Planktonset 1.0: Plankton Imagery Data Collected from Fg Walton Smith in Straits of Florida from 2014–06-03 to 2014–06-06 and Used in the 2015 National Data Science Bowl (Ncei Accession 0127422) 2015. Available online: https://github.com/Planktos/PlanktonSet-1.0 (accessed on 5 March 2026).
- Dai, J.; Wang, R.; Zheng, H.; Ji, G.; Qiao, X. Zooplanktonet: Deep Convolutional Network for Zooplankton Classification. In Proceedings of the OCEANS 2016-Shanghai; IEEE: New York, NY, USA, 2016; pp. 1–6. [Google Scholar]
- Lumini, A.; Nanni, L. Deep Learning and Transfer Learning Features for Plankton Classification. Ecol. Inform. 2019, 51, 33–43. [Google Scholar] [CrossRef]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. Imagenet Classification with Deep Convolutional Neural Networks. In Proceedings of the Advances in Neural Information Processing Systems 25 (NeurIPS 2012), Lake Tahoe, NV, USA, 3–6 December 2012; pp. 1097–1105. [Google Scholar]
- Li, Y.; Guo, J.; Guo, X.; Hu, Z.; Tian, Y. Plankton Detection with Adversarial Learning and a Densely Connected Deep Learning Model for Class Imbalanced Distribution. J. Mar. Sci. Eng. 2021, 9, 636. [Google Scholar] [CrossRef]
- Yue, J.; Chen, Z.; Long, Y.; Cheng, K.; Bi, H.; Cheng, X. Toward Efficient Deep Learning System for In-Situ Plankton Image Recognition. Front. Mar. Sci. 2023, 10, 1186343. [Google Scholar] [CrossRef]
- Sömek, B.; Yuksel, S.E. Plankton Classification with Deep Learning. In Proceedings of the 2023 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA); IEEE: New York, NY, USA, 2023; pp. 118–123. [Google Scholar]
- Gonzalez-Cid, Y.; Burguera, A.; Bonin-Font, F.; Matamoros, A. Machine Learning and Deep Learning Strategies to Identify Posidonia Meadows in Underwater Images. In Proceedings of the OCEANS 2017-Aberdeen; IEEE: New York, NY, USA, 2017; pp. 1–5. [Google Scholar]
- Martin-Abadal, M.; Guerrero-Font, E.; Bonin-Font, F.; Gonzalez-Cid, Y. Deep Semantic Segmentation in an AUV for Online Posidonia Oceanica Meadows Identification. IEEE Access 2018, 6, 60956–60967. [Google Scholar] [CrossRef]
- Long, J.; Shelhamer, E.; Darrell, T. Fully Convolutional Networks for Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2015; pp. 3431–3440. [Google Scholar]
- Sabato, G.; Scardino, G.; Kushabaha, A.; Chirivì, M.; Luparelli, A.; Scicchitano, G. Deep Learning-Based Segmentation Techniques for Coastal Monitoring and Seagrass Banquette Detection. In Proceedings of the 2023 IEEE International Workshop on Metrology for the Sea; Learning to Measure Sea Health Parameters (MetroSea); IEEE: New York, NY, USA, 2023; pp. 524–527. [Google Scholar]
- Liu, L.; Bao, Z.; Liang, Y.; Deng, H.; Zhang, X.; Cao, T.; Zhou, C.; Zhang, Z. Unsupervised Learning for Lake Underwater Vegetation Classification: Constructing High-Precision, Large-Scale Aquatic Ecological Datasets. Sci. Total Environ. 2025, 958, 177895. [Google Scholar] [CrossRef]
- Gao, L.; Li, X.; Kong, F.; Yu, R.; Guo, Y.; Ren, Y. AlgaeNet: A Deep-Learning Framework to Detect Floating Green Algae from Optical and SAR Imagery. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 2782–2796. [Google Scholar] [CrossRef]
- Elsäßer, J.; Weihl, L.; Cheplygina, V.; Nielsen, L.T. SeagrassFinder: Deep Learning for Eelgrass Detection and Coverage Estimation in the Wild. Ecol. Inform. 2025, 90, 103200. [Google Scholar] [CrossRef]
- Zhao, F.; Huang, B.; Wang, J.; Shao, X.; Wu, Q.; Xi, D.; Liu, Y.; Chen, Y.; Zhang, G.; Ren, Z.; et al. Seafloor Debris Detection Using Underwater Images and Deep Learning-Driven Image Restoration: A Case Study from Koh Tao, Thailand. Mar. Pollut. Bull. 2025, 214, 117710. [Google Scholar] [CrossRef]
- Mahmood, A.; Bennamoun, M.; An, S.; Sohel, F.; Boussaid, F.; Hovey, R.; Kendrick, G.; Fisher, R.B. Coral Classification with Hybrid Feature Representations. In Proceedings of the 2016 IEEE International Conference on Image Processing (ICIP); IEEE: New York, NY, USA, 2016; pp. 519–523. [Google Scholar]
- Bewley, M.; Friedman, A.; Ferrari, R.; Hill, N.; Hovey, R.; Barrett, N.; Marzinelli, E.M.; Pizarro, O.; Figueira, W.; Meyer, L.; et al. Australian Sea-Floor Survey Data, with Images and Expert Annotations. Sci. Data 2015, 2, 150057. [Google Scholar] [CrossRef]
- Osterloff, J.; Nilssen, I.; Järnegren, J.; Buhl-Mortensen, P.; Nattkemper, T.W. Polyp Activity Estimation and Monitoring for Cold Water Corals with a Deep Learning Approach. In Proceedings of the Computer Vision for Analysis of Underwater Imagery (CVAUI), 2016 ICPR 2nd Workshop on Computer Vision for Analysis of Underwater Imagery (CVAUI); IEEE: New York, NY, USA, 2016; pp. 1–6. [Google Scholar]
- Giles, A.B.; Ren, K.; Davies, J.E.; Abrego, D.; Kelaher, B. Combining Drones and Deep Learning to Automate Coral Reef Assessment with RGB Imagery. Remote Sens. 2023, 15, 2238. [Google Scholar] [CrossRef]
- Anwarul, S.; Tanwar, R. CoralClassify: Advancing Coral Health Monitoring Using Deep Learning. Procedia Comput. Sci. 2025, 259, 88–97. [Google Scholar] [CrossRef]
- Zheng, Z.; Liang, H.; Wut, F.H.; Wong, Y.H.; Chui, A.P.-Y.; Yeung, S.-K. Hkcoral: Benchmark for Dense Coral Growth Form Segmentation in the Wild. IEEE J. Ocean. Eng. 2025, 50, 697–713. [Google Scholar] [CrossRef]
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep Learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef]
- LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-Based Learning Applied to Document Recognition. Proc. IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef]
- LeCun, Y.; Jackel, L.D.; Boser, B.; Denker, J.S.; Graf, H.P.; Guyon, I.; Henderson, D.; Howard, R.E.; Hubbard, W. Handwritten Digit Recognition: Applications of Neural Network Chips and Automatic Learning. IEEE Commun. Mag. 1989, 27, 41–46. [Google Scholar] [CrossRef]
- Qiu, C.; Zhang, S.; Wang, C.; Yu, Z.; Zheng, H.; Zheng, B. Improving Transfer Learning and Squeeze-and-Excitation Networks for Small-Scale Fine-Grained Fish Image Classification. IEEE Access 2018, 6, 78503–78512. [Google Scholar] [CrossRef]
- Lin, T.-Y.; RoyChowdhury, A.; Maji, S. Bilinear Cnn Models for Fine-Grained Visual Recognition. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2015; pp. 1449–1457. [Google Scholar]
- Jäger, J.; Simon, M.; Denzler, J.; Wolff, V.; Fricke-Neuderth, K.; Kruschel, C. Croatian Fish Dataset: Fine-Grained Classification of Fish Species in Their Natural Habitat. Swans. Bmvc 2015, 2, 6.1–6.7. [Google Scholar]
- Anantharajah, K.; Ge, Z.; McCool, C.; Denman, S.; Fookes, C.; Corke, P.; Tjondronegoro, D.; Sridharan, S. Local Inter-Session Variability Modelling for Object Classification. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2014; pp. 309–316. [Google Scholar]
- Ding, G.; Song, Y.; Guo, J.; Feng, C.; Li, G.; He, B.; Yan, T. Fish Recognition Using Convolutional Neural Network. In Proceedings of the OCEANS–Anchorage; IEEE: New York, NY, USA, 2017; pp. 1–4. [Google Scholar]
- Qin, H.; Li, X.; Yang, Z.; Shang, M. When Underwater Imagery Analysis Meets Deep Learning: A Solution at the Age of Big Visual Data. In Proceedings of the OCEANS’15 MTS/IEEE Washington; IEEE: New York, NY, USA, 2015; pp. 1–5. [Google Scholar]
- Qin, H.; Li, X.; Liang, J.; Peng, Y.; Zhang, C. DeepFish: Accurate Underwater Live Fish Recognition with a Deep Architecture. Neurocomputing 2016, 187, 49–58. [Google Scholar] [CrossRef]
- Qin, H.; Peng, Y.; Li, X. Foreground Extraction of Underwater Videos via Sparse and Low-Rank Matrix Decomposition. In Proceedings of the 2014 ICPR Workshop on Computer Vision for Analysis of Underwater Imagery; IEEE: New York, NY, USA, 2014; pp. 65–72. [Google Scholar]
- Rathi, D.; Jain, S.; Indu, D.S. Underwater Fish Species Classification Using Convolutional Neural Network and Deep Learning. arXiv 2018, arXiv:1805.10106. [Google Scholar] [CrossRef]
- Han, F.; Zhu, J.; Liu, B.; Zhang, B.; Xie, F. Fish Shoals Behavior Detection Based on Convolutional Neural Network and Spatiotemporal Information. IEEE Access 2020, 8, 126907–126926. [Google Scholar] [CrossRef]
- Jovanović, V.; Svendsen, E.; Risojević, V.; Babić, Z. Splash Detection in Fish Plants Surveillance Videos Using Deep Learning. In Proceedings of the 2018 14th Symposium on Neural Networks and Applications (NEUREL); IEEE: New York, NY, USA, 2018; pp. 1–5. [Google Scholar]
- Huang, R.-J.; Lai, Y.-C.; Tsao, C.-Y.; Kuo, Y.-P.; Wang, J.-H.; Chang, C.-C. Applying Convolutional Networks to Underwater Tracking without Training. In Proceedings of the 2018 IEEE International Conference on Applied System Invention (ICASI); IEEE: New York, NY, USA, 2018; pp. 342–345. [Google Scholar]
- Cui, Y.; Pan, T.; Chen, S.; Zou, X. A Gender Classification Method for Chinese Mitten Crab Using Deep Convolutional Neural Network. Multimed. Tools Appl. 2020, 79, 7669–7684. [Google Scholar] [CrossRef]
- Simonyan, K.; Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv 2014, arXiv:1409.1556. [Google Scholar]
- Taheri-Garavand, A.; Nasiri, A.; Banan, A.; Zhang, Y.-D. Smart Deep Learning-Based Approach for Non-Destructive Freshness Diagnosis of Common Carp Fish. J. Food Eng. 2020, 278, 109930. [Google Scholar] [CrossRef]
- Becken, S.; Connolly, R.; Stantic, B.; Scott, N.; Mandal, R.; Le, D. Monitoring Aesthetic Value of the Great Barrier Reef by Using Innovative Technologies and Artificial Intelligence; Griffith Institute for Tourism Research Report; Griffith Institute for Tourism, Griffith University: Brisbane, Australia, 2018. [Google Scholar]
- Mader, H.; Kratzert, F. The Fishcam Migration Monitoring System for Fish Passes. In Proceedings of the 11th International Symposium on Ecohydraulics; Webb, J., Costelloe, J., Casas-Mulet, R., Lyon, J., Stewardson, M., Eds.; University of Melbourne: Melbourne, Australia, 2016. [Google Scholar]
- Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; Fei-Fei, L. Imagenet: A Large-Scale Hierarchical Image Database. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2009; pp. 248–255. [Google Scholar]
- Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. Imagenet Large Scale Visual Recognition Challenge. Int. J. Comput. Vis. 2015, 115, 211–252. [Google Scholar] [CrossRef]
- Chen, G.; Sun, P.; Shang, Y. Automatic Fish Classification System Using Deep Learning. In Proceedings of the Tools with Artificial Intelligence (ICTAI), 2017 IEEE 29th International Conference on Tools with Artificial Intelligence (ICTAI); IEEE: New York, NY, USA, 2017; pp. 24–29. [Google Scholar]
- Ali-Gombe, A.; Elyan, E.; Jayne, C. Fish Classification in Context of Noisy Images. In Proceedings of the International Conference on Engineering Applications of Neural Networks; Springer: Berlin/Heidelberg, Germany, 2017; pp. 216–226. [Google Scholar]
- Sirigineedi, M.; Jagan Mohan, R.; Sahu, B. Improving Fish Image Detection Speed with Hybrid VGG16 and Darknet. Multimed. Tools Appl. 2024, 84, 10551–10566. [Google Scholar] [CrossRef]
- Thorat, P.; Tongaonkar, R.; Jagtap, V. Towards Designing the Best Model for Classification of Fish Species Using Deep Neural Networks. In Proceedings of the International Conference on Computational Science and Applications: ICCSA 2019; Springer: Berlin/Heidelberg, Germany, 2020; pp. 343–351. [Google Scholar]
- Prasetyo, E.; Suciati, N.; Fatichah, C. Multi-Level Residual Network VGGNet for Fish Species Classification. J. King Saud Univ.-Comput. Inf. Sci. 2022, 34, 5286–5295. [Google Scholar] [CrossRef]
- Prasetyo, E.; Suciati, N.; Fatichah, C. Fish-Gres Dataset for Fish Species Classification. Mendeley Data 2020, 10, 12. [Google Scholar] [CrossRef]
- Mamun, M.R.I.; Rahman, U.S.; Akter, T.; Azim, M.A. Fish Disease Detection Using Deep Learning and Machine Learning. Int. J. Comput. Appl. 2023, 975, 8887. [Google Scholar] [CrossRef]
- Choi, S. Fish Identification in Underwater Video with Deep Convolutional Neural Network: SNUMedinfo at LifeCLEF Fish Task 2015. Available online: https://ceur-ws.org/Vol-1391/110-CR.pdf (accessed on 5 March 2026).
- Jäger, J.; Rodner, E.; Denzler, J.; Wolff, V.; Fricke-Neuderth, K. SeaCLEF 2016: Object Proposal Classification for Fish Detection in Underwater Videos. In Proceedings of the CLEF 2016 Working Notes, Évora, Portugal, 5–8 September 2016; CEUR-WS.org: Aachen, Germany, 2016; pp. 481–489. [Google Scholar]
- Mothes, O.; Denzler, J. Anatomical Landmark Tracking by One-Shot Learned Priors for Augmented Active Appearance Models. In Proceedings of the VISIGRAPP (6: VISAPP), Porto, Portugal, 27 February–1 March 2017; pp. 246–254. [Google Scholar]
- Jäger, J.; Wolff, V.; Fricke-Neuderth, K.; Mothes, O.; Denzler, J. Visual Fish Tracking: Combining a Two-Stage Graph Approach with CNN-Features. In Proceedings of the OCEANS 2017-Aberdeen; IEEE: New York, NY, USA, 2017; pp. 1–6. [Google Scholar]
- Pelletier, S.; Montacir, A.; Zakari, H.; Akhloufi, M. Deep Learning for Marine Resources Classification in Non-Structured Scenarios: Training vs. Transfer Learning. In Proceedings of the 2018 IEEE Canadian Conference on Electrical & Computer Engineering (CCECE); IEEE: New York, NY, USA, 2018; pp. 1–4. [Google Scholar]
- Qu, P.; Li, T.; Zhou, L.; Jin, S.; Liang, Z.; Zhao, W.; Zhang, W. DAMNet: Dual Attention Mechanism Deep Neural Network for Underwater Biological Image Classification. IEEE Access 2022, 11, 6000–6009. [Google Scholar] [CrossRef]
- Mehrunnisa; Leszczuk, M.; Juszka, D.; Zhang, Y. Improved Binary Classification of Underwater Images Using a Modified ResNet-18 Model. Electronics 2025, 14, 2954. [Google Scholar] [CrossRef]
- Jiang, Q.; Gu, Y.; Li, C.; Cong, R.; Shao, F. Underwater Image Enhancement Quality Evaluation: Benchmark Dataset and Objective Metric. IEEE Trans. Circuits Syst. Video Technol. 2022, 32, 5959–5974. [Google Scholar] [CrossRef]
- Kaneko, R.; Ueda, T.; Higashi, H.; Tanaka, Y. PHISWID: Physics-Inspired Underwater Image Dataset Synthesized from RGB-D Images. APSIPA Trans. Signal Inf. Process. 2025, 15, 1–25. [Google Scholar] [CrossRef]
- Kaneko, R.; Sato, Y.; Ueda, T.; Higashi, H.; Tanaka, Y. Marine Snow Removal Benchmarking Dataset. In Proceedings of the 2023 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC); IEEE: New York, NY, USA, 2023; pp. 771–778. [Google Scholar]
- Girshick, R. Fast R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV 2015), Santiago, Chile, 7–13 December 2015; pp. 1440–1448. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. In Advances in Neural Information Processing Systems 28 (NeurIPS 2015); Curran Associates, Inc.: Red Hook, NY, USA, 2015; pp. 91–99. [Google Scholar]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV 2017), Venice, Italy, 22–29 October 2017; pp. 2980–2988. [Google Scholar]
- Li, X.; Tang, Y.; Gao, T. Deep but Lightweight Neural Networks for Fish Detection. In Proceedings of the OCEANS 2017-Aberdeen; IEEE: New York, NY, USA, 2017; pp. 1–5. [Google Scholar]
- Li, X.; Shang, M.; Qin, H.; Chen, L. Fast Accurate Fish Detection and Recognition of Underwater Images with Fast R-CNN. In Proceedings of the OCEANS’15 MTS/IEEE Washington; IEEE: New York, NY, USA, 2015; pp. 1–5. [Google Scholar]
- Wang, H.; Xiao, N. Underwater Object Detection Method Based on Improved Faster RCNN. Appl. Sci. 2023, 13, 2746. [Google Scholar] [CrossRef]
- Ben Tamou, A.; Benzinou, A.; Nasreddine, K. Multi-Stream Fish Detection in Unconstrained Underwater Videos by the Fusion of Two Convolutional Neural Network Detectors. Appl. Intell. 2021, 51, 5809–5821. [Google Scholar] [CrossRef]
- Zeiler, M.D.; Fergus, R. Visualizing and Understanding Convolutional Networks. In Proceedings of the Computer Vision—ECCV 2014, Zurich, Switzerland, 6–12 September 2014; pp. 818–833. [Google Scholar]
- Chatfield, K.; Simonyan, K.; Vedaldi, A.; Zisserman, A. Return of the Devil in the Details: Delving Deep into Convolutional Nets. arXiv 2014, arXiv:1405.3531. [Google Scholar] [CrossRef]
- Xu, W.; Zhu, Z.; Ge, F.; Han, Z.; Li, J. Analysis of Behavior Trajectory Based on Deep Learning in Ammonia Environment for Fish. Sensors 2020, 20, 4425. [Google Scholar] [CrossRef]
- Isa, I.S.; Norzrin, N.N.; Sulaiman, S.N.; Hamzaid, N.A.; Maruzuki, M.I.F. CNN Transfer Learning of Shrimp Detection for Underwater Vision System. In Proceedings of the 2020 1st International Conference on Information Technology, Advanced Mechanical and Electrical Engineering (ICITAMEE); IEEE: New York, NY, USA, 2020; pp. 226–231. [Google Scholar]
- Baletaud, F.; Villon, S.; Gilbert, A.; Côme, J.-M.; Fiat, S.; Iovan, C.; Vigliola, L. Automatic Detection, Identification and Counting of Deep-Water Snappers on Underwater Baited Video Using Deep Learning. Front. Mar. Sci. 2025, 12, 1476616. [Google Scholar] [CrossRef]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Wu, Z.; Shen, C.; van den Hengel, A. High-Performance Semantic Segmentation Using Very Deep Fully Convolutional Networks. arXiv 2016, arXiv:1604.04339. [Google Scholar] [CrossRef]
- Zivkovic, Z. Improved Adaptive Gaussian Mixture Model for Background Subtraction. In Proceedings of the 17th International Conference on Pattern Recognition (ICPR 2004), Cambridge, UK, 23–26 August 2004; pp. 28–31. [Google Scholar]
- Redmon, J.; Farhadi, A. YOLO9000: Better, Faster, Stronger. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 7263–7271. [Google Scholar]
- Bochkovskiy, A. YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv 2020, arXiv:2004.10934. [Google Scholar] [CrossRef]
- Wang, C.-Y.; Yeh, I.-H.; Liao, H.-Y.M. Yolov9: Learning What You Want to Learn Using Programmable Gradient Information. arXiv 2024, arXiv:2402.13616. [Google Scholar] [CrossRef]
- Lu, H.; Uemura, T.; Wang, D.; Zhu, J.; Huang, Z.; Kim, H. Deep-Sea Organisms Tracking Using Dehazing and Deep Learning. Mob. Netw. Appl. 2018, 25, 1008–1015. [Google Scholar] [CrossRef]
- Kalal, Z.; Mikolajczyk, K.; Matas, J. Tracking-Learning-Detection. IEEE Trans. Pattern Anal. Mach. Intell. 2011, 34, 1409–1422. [Google Scholar] [CrossRef]
- Kalal, Z.; Mikolajczyk, K.; Matas, J. Forward-Backward Error: Automatic Detection of Tracking Failures. In Proceedings of the 2010 20th International Conference on Pattern Recognition; IEEE: New York, NY, USA, 2010; pp. 2756–2759. [Google Scholar]
- Hofmann, A.C.; Bergmann, S.M.; Reulke, R. Analysis of Motion Patterns in Video Streams for Automatic Health Monitoring in Koi Ponds. In Proceedings of the Image and Video Technology: 9th Pacific-Rim Symposium, PSIVT 2019, Sydney, Australia, 18–22 November 2019; p. 27. [Google Scholar]
- Al Muksit, A.; Hasan, F.; Emon, M.F.H.B.; Haque, M.R.; Anwary, A.R.; Shatabda, S. YOLO-Fish: A Robust Fish Detection Model to Detect Fish in Realistic Underwater Environment. Ecol. Inform. 2022, 72, 101847. [Google Scholar] [CrossRef]
- Mahmood, A.; Bennamoun, M.; An, S.; Sohel, F.; Boussaid, F.; Hovey, R.; Kendrick, G. Automatic Detection of Western Rock Lobster Using Synthetic Data. ICES J. Mar. Sci. 2019, 77, 1308–1317. [Google Scholar] [CrossRef]
- Zhang, M.; Xu, S.; Song, W.; He, Q.; Wei, Q. Lightweight Underwater Object Detection Based on YOLO v4 and Multi-Scale Attentional Feature Fusion. Remote Sens. 2021, 13, 4706. [Google Scholar] [CrossRef]
- Pedersen, M.; Bruslund Haurum, J.; Gade, R.; Moeslund, T.B. Detection of Marine Animals in a New Underwater Dataset with Varying Visibility. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops; IEEE: New York, NY, USA, 2019; pp. 18–26. [Google Scholar]
- Li, L.; Shi, G.; Jiang, T. Fish Detection Method Based on Improved YOLOv5. Aquac. Int. 2023, 31, 2513–2530. [Google Scholar] [CrossRef]
- Gao, M.; Li, S.; Wang, K.; Bai, Y.; Ding, Y.; Zhang, B.; Guan, N.; Wang, P. Real-Time Jellyfish Classification and Detection Algorithm Based on Improved YOLOv4-Tiny and Improved Underwater Image Enhancement Algorithm. Sci. Rep. 2023, 13, 12989. [Google Scholar] [CrossRef] [PubMed]
- Woo, S.; Park, J.; Lee, J.-Y.; Kweon, I.S. Cbam: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
- Zhang, W.; Rui, F.; Xiao, C.; Li, H.; Li, Y. JF-YOLO: The Jellyfish Bloom Detector Based on Deep Learning. Multimed. Tools Appl. 2024, 83, 7097–7117. [Google Scholar] [CrossRef]
- Chen, V.Y.; Wu, Y.-W.; Hu, C.-W.; Han, Y.-S. Enhancing Green Sea Turtle (Chelonia Mydas) Conservation for Tourists at Little Liuqiu Island, Taiwan: Application of Deep Learning Algorithms. Ocean Coast. Manag. 2024, 252, 107111. [Google Scholar] [CrossRef]
- Bukas, C.; Albrecht, F.; Ur-Rehman, M.S.; Popek, D.; Patalan, M.; Pawłowski, J.; Wecker, B.; Landsch, K.; Golan, T.; Kowalczyk, T.; et al. Robust Deep Learning Based Shrimp Counting in an Industrial Farm Setting. J. Clean. Prod. 2024, 468, 143024. [Google Scholar] [CrossRef]
- Cai, Y.; Yao, Z.; Jiang, H.; Qin, W.; Xiao, J.; Huang, X.; Pan, J.; Feng, H. Rapid Detection of Fish with SVC Symptoms Based on Machine Vision Combined with a NAM-YOLO v7 Hybrid Model. Aquaculture 2024, 582, 740558. [Google Scholar] [CrossRef]
- Liu, Y.; Shao, Z.; Teng, Y.; Hoffmann, N. NAM: Normalization-Based Attention Module. arXiv 2021, arXiv:2111.12419. [Google Scholar] [CrossRef]
- Jin, Y.; Xiao, X.; Pan, Y.; Zhou, X.; Hu, K.; Wang, H.; Zou, X. A Novel Method for the Object Detection and Weight Prediction of Chinese Softshell Turtles Based on Computer Vision and Deep Learning. Animals 2024, 14, 1368. [Google Scholar] [CrossRef]
- Pachaiyappan, P.; Chidambaram, G.; Jahid, A.; Alsharif, M.H. Enhancing Underwater Object Detection and Classification Using Advanced Imaging Techniques: A Novel Approach with Diffusion Models. Sustainability 2024, 16, 7488. [Google Scholar] [CrossRef]
- Bajpai, A.; Tiwari, N.; Yadav, A.; Chaurasia, D.; Kumar, M. Enhancing Underwater Object Detection: Leveraging YOLOv8m for Improved Subaquatic Monitoring. SN Comput. Sci. 2024, 5, 793. [Google Scholar] [CrossRef]
- Shah, C.; Nabi, M.; Alaba, S.Y.; Ebu, I.A.; Prior, J.; Campbell, M.D.; Caillouet, R.; Grossi, M.D.; Rowell, T.; Wallace, F.; et al. Yolov8-Tf: Transformer-Enhanced Yolov8 for Underwater Fish Species Recognition with Class Imbalance Handling. Sensors 2025, 25, 1846. [Google Scholar] [CrossRef]
- Li, Y.; Hu, Z.; Zhang, Y.; Liu, J.; Tu, W.; Yu, H. DDEYOLOv9: Network for Detecting and Counting Abnormal Fish Behaviors in Complex Water Environments. Fishes 2024, 9, 242. [Google Scholar] [CrossRef]
- Huang, T.-W.; Hwang, J.-N.; Romain, S.; Wallace, F. Fish Tracking and Segmentation from Stereo Videos on the Wild Sea Surface for Electronic Monitoring of Rail Fishing. IEEE Trans. Circuits Syst. Video Technol. 2018, 29, 3146–3158. [Google Scholar] [CrossRef]
- Li, M.; Li, X.; Chen, S.; Huang, H. A High-Precision and Lightweight Underwater Fish Detection and Recognition Approach Based on the Improved YOLOX-Nano Algorithm. Aquac. Eng. 2025, 110, 102533. [Google Scholar] [CrossRef]
- Wang, D.; Wu, M.; Zhu, X.; Qin, Q.; Wang, S.; Ye, H.; Guo, K.; Wu, C.; Shi, Y. Real-Time Detection and Identification of Fish Skin Health in the Underwater Environment Based on Improved YOLOv10 Model. Aquac. Rep. 2025, 42, 102723. [Google Scholar] [CrossRef]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 2980–2988. [Google Scholar]
- Cutter, G.; Stierhoff, K.; Zeng, J. Automated Detection of Rockfish in Unconstrained Underwater Videos Using Haar Cascades and a New Image Dataset: Labeled Fishes in the Wild. In Proceedings of the Applications and Computer Vision Workshops (WACVW), 2015 IEEE Winter; IEEE: New York, NY, USA, 2015; pp. 57–62. [Google Scholar]
- OzFish Dataset—Machine Learning Dataset for Baited Remote Underwater Video Stations. Available online: https://doi.org/10.25845/5e28f062c5097 (accessed on 5 March 2026).
- Manikandan, D.L.; Santhanam, S.M. Underwater Species Classification Using Deep Learning Technique. Rev. Română Informatică Autom. 2024, 34, 7–20. [Google Scholar] [CrossRef]
- Manikandan, D.L.; Santhanam, S.M. Parallel Desires: Unifying Local and Semantic Feature Representations in Marine Species Images for Classification. Mar. Geophys. Res. 2024, 45, 16. [Google Scholar] [CrossRef]
- Ji, D.; Hussain, A.F.; Hussain, S.; Ogbonnaya, S.G.; Zhu, S.; Wang, X. Fish Detection and Classification Based on Improved ViT. In Proceedings of the 2023 2nd International Conference on Automation, Robotics and Computer Engineering (ICARCE); IEEE: New York, NY, USA, 2023; pp. 1–5. [Google Scholar]
- Ulucan, O.; Karakaya, D.; Turkan, M. A Large-Scale Dataset for Fish Segmentation and Classification. In Proceedings of the 2020 Innovations in Intelligent Systems and Applications Conference (ASYU); IEEE: New York, NY, USA, 2020; pp. 1–5. [Google Scholar]
- Irfan, M.; Zheng, J.; Iqbal, M.; Arif, M.H. A Novel Feature Extraction Model to Enhance Underwater Image Classification. In Proceedings of the International Symposium on Intelligent Computing Systems; Springer: Berlin/Heidelberg, Germany, 2020; pp. 78–91. [Google Scholar]
- O’Byrne, M.; Pakrashi, V.; Schoefs, F.; Ghosh, B. Semantic Segmentation of Underwater Imagery Using Deep Networks Trained on Synthetic Imagery. J. Mar. Sci. Eng. 2018, 6, 93. [Google Scholar] [CrossRef]
- Fernandes, A.F.A.; Turra, E.M.; de Alvarenga, É.R.; Passafaro, T.L.; Lopes, F.B.; Alves, G.F.O.; Singh, V.; Rosa, G.J.M. Deep Learning Image Segmentation for Extraction of Fish Body Measurements and Prediction of Body Weight and Carcass Traits in Nile Tilapia. Comput. Electron. Agric. 2020, 170, 105274. [Google Scholar] [CrossRef]
- Liu, F.; Fang, M. Semantic Segmentation of Underwater Images Based on Improved Deeplab. J. Mar. Sci. Eng. 2020, 8, 188. [Google Scholar] [CrossRef]
- Kareem, H.H.; Daway, H.G.; Daway, E.G. Underwater Image Enhancement Using Colour Restoration Based on YCbCr Colour Model. In Proceedings of the IOP Conference Series: Materials Science and Engineering; IOP Publishing: Bristol, UK, 2019; Volume 571, p. 012125. [Google Scholar]
- Fan, Z.; Xia, W.; Liu, X.; Li, H. Detection and Segmentation of Underwater Objects from Forward-Looking Sonar Based on a Modified Mask RCNN. Signal Image Video Process. 2021, 15, 1135–1143. [Google Scholar] [CrossRef]
- Jahanbakht, M.; Xiang, W.; Waltham, N.J.; Azghadi, M.R. Distributed Deep Learning and Energy-Efficient Real-Time Image Processing at the Edge for Fish Segmentation in Underwater Videos. IEEE Access 2022, 10, 117796–117807. [Google Scholar] [CrossRef]
- Lin, H.-Y.; Tseng, S.-L.; Li, J.-Y. SUR-Net: A Deep Network for Fish Detection and Segmentation with Limited Training Data. IEEE Sens. J. 2022, 22, 18035–18044. [Google Scholar] [CrossRef]
- Møller, T.; Nilssen, I.; Nattkemper, T.W. Tracking Sponge Size and Behaviour with Fixed Underwater Observatories. In Proceedings of the International Conference on Pattern Recognition; Springer: Berlin/Heidelberg, Germany, 2018; pp. 45–54. [Google Scholar]
- Conrady, C.R.; Er, Ş.; Attwood, C.G.; Roberson, L.A.; de Vos, L. Automated Detection and Classification of Southern African Roman Seabream Using Mask R-CNN. Ecol. Inform. 2022, 69, 101593. [Google Scholar] [CrossRef]
- Chicchon, M.; Bedon, H.; Del-Blanco, C.R.; Sipiran, I. Semantic Segmentation of Fish and Underwater Environments Using Deep Convolutional Neural Networks and Learned Active Contours. IEEE Access 2023, 11, 33652–33665. [Google Scholar] [CrossRef]
- Yang, G.; Yang, J.; Fan, W.; Yang, D. Neural Network for Underwater Fish Image Segmentation Using an Enhanced Feature Pyramid Convolutional Architecture. J. Mar. Sci. Eng. 2025, 13, 238. [Google Scholar] [CrossRef]
- Kong, J.; Tang, S.; Feng, J.; Mo, L.; Jin, X. AASNet: A Novel Image Instance Segmentation Framework for Fine-Grained Fish Recognition via Linear Correlation Attention and Dynamic Adaptive Focal Loss. Appl. Sci. 2025, 15, 3986. [Google Scholar] [CrossRef]
- Lian, S.; Li, H.; Cong, R.; Li, S.; Zhang, W.; Kwong, S. Watermask: Instance Segmentation for Underwater Imagery. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2023; pp. 1305–1315. [Google Scholar]
- Lian, S.; Zhang, Z.; Li, H.; Li, W.; Yang, L.T.; Kwong, S.; Cong, R. Diving into Underwater: Segment Anything Model Guided Underwater Salient Instance Segmentation and a Large-Scale Dataset. arXiv 2024, arXiv:2406.06039. [Google Scholar] [CrossRef]
- Saleh, A.; Sheaves, M.; Jerry, D.; Azghadi, M.R. Overcoming Annotation Bottlenecks in Underwater Fish Segmentation: A Robust Self-Supervised Learning Approach. Signal Image Video Process. 2025, 19, 270. [Google Scholar] [CrossRef]
- Ditria, E.M.; Connolly, R.M.; Jinks, E.L.; Lopez-Marcano, S. Annotated Video Footage for Automated Identification and Counting of Fish in Unconstrained Seagrass Habitats. Front. Mar. Sci. 2021, 8, 629485. [Google Scholar] [CrossRef]
- Xu, N.; Yang, L.; Fan, Y.; Yue, D.; Liang, Y.; Yang, J.; Huang, T. Youtube-Vos: A Large-Scale Video Object Segmentation Benchmark. arXiv 2018, arXiv:1809.03327. [Google Scholar]
- Pavithra, S.; Cicil Melbin Denny, J. An Efficient Approach to Detect and Segment Underwater Images Using Swin Transformer. Results Eng. 2024, 23, 102460. [Google Scholar] [CrossRef]
- Li, D.; Zhao, S.; Hu, J.; Yang, Y.; Ding, J. An Underwater Image Segmentation Model for Complex Scenes in Aquaculture Using Vision Transformer. Comput. Electron. Agric. 2025, 238, 110764. [Google Scholar] [CrossRef]
- Li, H.; Lian, S.; Li, Z.; Cong, R.; Li, C. Taming SAM for Underwater Instance Segmentation and Beyond. arXiv 2025. [Google Scholar] [CrossRef]
- Xue, X.; Zhou, Y.; Yan, D.; Tao, L.; Li, J.; Li, Y.; Zhang, H.; Xiao, R. UVLM: Benchmarking Video Language Model for Underwater World Understanding. arXiv 2025, arXiv:2507.02373. [Google Scholar] [CrossRef]
- Khan, F.F.; Radwan, Y.; Abdelrahman, E.; Felemban, A.; Mir, A.; Michiels, N.K.; Temple, A.J.; Berumen, M.L.; Elhoseiny, M. FishNet++: Analyzing the Capabilities of Multimodal Large Language Models in Marine Biology. arXiv 2025, arXiv:2509.25564. [Google Scholar]
- Li, W.; Zhang, F. Real-Time Vision–Language Analysis for Autonomous Underwater Drones: A Cloud–Edge Framework Using Qwen2. 5-VL. Drones 2025, 9, 605. [Google Scholar] [CrossRef]
- Tian, B.; Zhao, L.; Chen, B.; Zheng, H.; Yang, J.; Wu, M.; Vasisht, D.; Nahrstedt, K. AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models. arXiv 2025, arXiv:2510.21722. [Google Scholar] [CrossRef]
- Li, B.; Huo, T.; Zhang, D.; Zhao, Z.; Gao, J.; Li, X. Exploring the Underwater World Segmentation without Extra Training. arXiv 2025, arXiv:2511.07923. [Google Scholar]
- Zhu, J.; Yin, S.; Liu, X.; Wang, X.; Yang, Y.-H. FishDetectLLM: Multimodal Instruction Tuning with Large Language Models for Fish Detection. Knowl.-Based Syst. 2025, 318, 113418. [Google Scholar] [CrossRef]
- Han, H.; Wang, W.; Zhang, G.; Li, M.; Wang, Y. Enhancing Vision-Language Models with Morphological and Taxonomic Knowledge: Towards Coral Recognition for Ocean Health. Proc. AAAI Conf. Artif. Intell. 2025, 39, 28052–28060. [Google Scholar] [CrossRef]







| Criterion Category | Inclusion Criteria | Exclusion Criteria |
|---|---|---|
| Publication type | Peer-reviewed journal articles, conference papers, and peer-reviewed preprints from recognized repositories (e.g., arXiv). | Abstract-only publications, posters, editorials, commentaries, theses, or non-indexed reports. |
| Methodological scope | Application of AI techniques (deep learning, CNNs, transformers, hybrid approaches) to the analysis of marine species images. | AI methods without application to marine species images. |
| Domain relevance | Imagery of marine species captured underwater (e.g., coral reefs, open ocean) or above water (e.g., shoreline, aerial surveys). | Studies focusing exclusively on non-marine species or terrestrial ecosystems. |
| Methodological transparency | Detailed description of model architecture, datasets, and training/testing procedures. | Insufficient methodological detail to enable replication. |
| Temporal coverage | Published between January 2015 and August 2025. | Publications outside the specified date range. |
| Language and availability | Written in English (or accessible language for review team) with full-text available. | Languages not accessible to the review team or full text unavailable. |
| Data modality | Image-based studies (still images or video frames) of marine species. | Studies using only non-visual modalities (e.g., acoustic, sonar-only, textual). |
| Complexity of methods | Use of advanced deep learning methods (e.g., CNNs, Vision Transformers, hybrid architectures). | Use of only traditional machine learning methods (e.g., SVM, random forest, k-NN) without deep learning models. |
| Access status | Articles available as open access or through freely accessible repositories. | Articles requiring paid subscription or inaccessible through institutional or open channels. |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [61] | Images obtained by an autonomous underwater vehicle | Underwater image enhancement | CNN | Visual inspection and error rate in posterior classification | 14.1% |
| [64] | UIEB dataset and synthetic underwater image dataset | Underwater image enhancement | Conditional generative adversarial network (cGAN) | PSNR/SSIM (Structural Similarity Index) | 17.715 dB/0.6552 |
| [65] | Subsets from Imagenet, YouTube videos and images from Flicr™ | Underwater image enhancement | GAN (U-Net) called UGAN and UGAN-P (with Gradient Difference Loss) | Gradient difference loss metrics in red, blue, green and orange patch/distances in image space | 9.39, 5.50, 3.25, 5.79/94.91 (mean) |
| [68] | B3DO [94], UW RGB-D Object [95], NYU Depth [96] and Microsoft 7- scenes [97] datasets and 3 underwater datasets collected in a laboratory, Jamaica and Australia | Underwater image generation and restoration | WaterGAN | Euclidean distance/variance of intensity-normalized colour in RGB-space (red, green, blue) | (0.0484, 0.2132, 0.1431)/(0.0005, 0.0007, 0.0006) |
| [71] | Images from the Internet (aquariums, underwater images with artificial light and processed images) | Underwater image enhancement and restoration | Underwater Denoising Autoencoder (UDAE) | MSE/SSIM | 0.0028/0.9653 |
| [72] | JAMSTEC database | Underwater image enhancement | Contrast enhancement method | Index Qu/SSIM | 0.7929/0.6027 |
| [74] | Modifed UIEB dataset and RUIE dataset | Underwater image enhancement | U-Net, CVAE, AdaIN | PSNR/SSIM/DeltaE/NIQE/MOS | UIEB dataset: 21.86 dB/0.870/9.556/3.626/4.2 |
| [79] | UIEB dataset | Underwater image enhancement | MDP and deep Q network | NIQE/UCIQE/UIQM | 41.334, 0.6412, 4.5935 |
| [82] | Images from ImageNet and SUN (Scene Understanding) database [98] | Underwater image enhancement | Multiscale dense generative adversarial network (GAN) | UCIQE/UIQM | 0.6028 ± 0.0282/5.0973 ± 0.4163 |
| [83] | UIEB dataset, LSUI [99] and UVEB | Underwater image enhancement | DNnet | PSNR/SSIM/MSE/UIQM | UIEB dataset 26.335 dB/0.910/0.368 × /2.994 |
| [84] | UIEB and LSUI datasets for training, UIEB190, LSUI850, OceanEx (full-reference); C60, RUIE dataset Color90, UPoor200, U45 (non-reference). | Underwater image enhancement | UIEVUS Framework: Integrates Retinex decomposition with GAN-based enhancement | Full-reference: PSNR/SSIM. Non-reference: UIQM | Full-reference: UIEB190 = 23.55 dB/0.90 Non-reference: U45 = 3.26 |
| [85] | UIEB-T90 (90 images from UIEB), UIEB-C60 (60 challenging images from UIEB), EUVP-T515 (515 images from EUVP), SQUID-T16 (16 images from SQUID), UIE-T78 (78 images from RUIE dataset). | Underwater image enhancement | CCL-Net | PSNR, SSIM, UIQM, UCIQE | RUIE-T78: UIQM = 3.168, UCIQE = 0.447 UIEB-T90: PSNR = 20.181 dB, SSIM = 0.866, UIQM = 3.021, UCIQE = 0.464 |
| [86] | EUVP and UIEB datasets | Underwater image enhancement | Improved U-Net | NIQE, UCIQE, PSNR, SSIM | 4.393/0.430/21.565 dB/0.879 |
| [87] | UIEBD (Unsupervised, no ground truth). For testing: EUVP, UFO, UIEBD (paired); DeepFish [100], FISHTRAC, FishID, RUIE dataset, SUIM (Segmentation of Underwater Imagery dataset) (unpaired). | Underwater image enhancement | UDNet framework integrates SGMCSS Module, CVAE Module and PAdaIN Block. | Full-Reference: PSNR, SSIM, MAD (Most Apparent Distortion), GMSD (Gradient Magnitude Similarity Deviation). No-Reference: UIQM, MUSIQ (Multi-scale Image Quality Transformer), NIQE, UCIQE. Comparison: PSNR, SSIM, UIQM, UCIQE. | Full-Reference: EUVP: PSNR = 22.96 dB, SSIM = 0.771, UIQM = 3.265, UCIQE = 0.749 No-Reference: UCCS: UIQM = 3.974, UCIQE = 0.713 |
| [91] | UIEB, EUVP and UFO-120 | Underwater image enhancement and restoration | Hybrid UNet-MaxViT | PSNR, SSIM, PCQI (Perception-based Colour Quality Index), UCIQE, UIQM, UICM, UIConM (Underwater Image Contrast Measure), CCF (Colourfulness Contrast Fog density index). | UIEB: 22.91 dB/23.5286/0.9341/0.6460/1.602/9.6405/1.1865/37.9325. |
| [92] | UIEB, EUVP, EUVPUN, RUIE dataset, U45 | Underwater Image Enhancement | UWFormer | PSNR, SSIM, LPIPS, UIQM, UCIQE | EUVP: PSNR = 24.40, SSIM = 0645, LPIPS = 0.129, UCIQE = 0.431 U45: UIQM = 3.227/UCIQE = 0.440 |
| [93] | U45, SIUM (Segmentation of Underwater Imagery) [101], UIEB, SQUID [102] | Underwater image enhancement | SAM | UIQM, CCF, AG, BC(e) (Blind Contrast Restoration Assessment in edge visibility), BC(r) (Blind Contrast Restoration Assessment in edge pixel gradient values), URanker (Ranking-based underwater image quality assessment) | UIEB: 3.544/48.508/27.165/1.168/3.889/1.964 |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [123] | Aerial images | Stingray detection and classification | Two different Faster R-CNN models (one with a ZF and other VGG based) | Average precision | 99.7% |
| [124] | Underwater images and Aerial images | Marine species classification | RetinaNet as the object detector/Simple Online Realtime Tracker (SORT) for tracking | Average precision | Aerial: 0.4 for Ray, 0.25 for Diver and 0.75 for Shark |
| [125] | Three datasets combining images from Google Earth38, free Arkive83, NOAA Photo Library84, and NWPU-RESISC45 dataset | Whale detection and counting | GoogleNet Inception v3 CNN architecture for detection and Faster R-CNN based on Inception-Resnet v2 CNN architecture for counting | F1-measure | Detection: 81%, Counting: 94% |
| [126] | Aerial images | Gannets, seals sea turtles | CNN | Average Recall/average precision | 0.826/0.403 |
| [127] | Aerial images | Cetacean detection and classification | U-Net with EfficientNet-b3 | Accuracy/f1-score/recall/precision | 91.37%/95.49%/98.96%/92.26% |
| [128] | Aerial images | Beluga whale detection | Faster-RCNN | Intersection over Union | 0.79 |
| [129] | Custom dataset built from UAV imagery | Detection and counting of turned white belly fish | PGG-YOLO based on YOLOv8 | Precision/Recall/F1-score/mAP50 | 99.52%/97.66%/98.58%/99.4% |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [130] | WHOI-Plankton database | Plankton classification | CNN and transfer learning (training with CIFAR10) | Accuracy | 92.80% |
| [132] | Planktonset | Plankton classification | Deep convolutional neural network | Loss | 60% |
| [134] | Zooplankton dataset | Zooplankton classification | ZooplanktoNet | Accuracy | 93.7% |
| [135] | WHOI, ZooScan and one by Kaggle | Plankton classification | Their ensemble method compared to AlexNet, VGG16 or ResNet50 and more | F-measure | 0.953, 0.897 and 0.926 |
| [137] | Datasets generated from the WHOI-Plankton database | Plankton classification | YOLOV3-dense | mAP | 97.21% |
| [138] | Dataset collected by PlanktonScope in the coastal area of Guangdong | Plankton classification | Transformers (Swin-T, ViT-B, and Swin-B) and CNNs (ResNet50, ResNet101, ResNet152, MobileNet V2, ShuffleNet) | Precision/recall | 92.38%, 91.73% |
| [139] | Dataset captured by Oregon State University’s Hatfield Marine Science Centre | Plankton classification | Inceptionv3, InceptionResNetv2, DenseNet, ResNet, and VGG-16 | Precision/recall/f1-score/accuracy/loss | 92%, 92%, 92%, 0.93, 0.31 |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [140] | Dataset obtained in the coastal areas of Mallorca | Detecting and identifying underwater plants (Posidonia) | SVM and ANN, with Gabor filters, texture descriptors and co-occurrence matrix, and on the other hand, a CNN | Mean of the detection hit ratio | Dataset1: 96.37%, Dataset2: 93.35% |
| [141] | Self-created with an AUV in Palma Bay, Cala Blava and Valdemossa port | Identify and segment Posidonia Oceanica semantically | Encoder and decoder, called VGG16-FCN8 | AUC | 95% |
| [143] | Images of Torre Canne beach in Puglia | Posidonia detection | Mask R-CNN, Detectron2, custom CNN | Accuracy/loss/IoU | 90.85%/0.18/0.68 |
| [144] | DeepSeagrass datasets and two private datasets: Erhai Lake and Wuhan East Lake datasets (China) | Underwater vegetation classification | ConvNeXt. UMAP and clustering (K-Means, Agglomerative Clustering and Birch) | Accuracy/precision | DeepSeagrass: 97.32 ± 0.36%/90%; Erhai Lake 92.43 ± 0.97%/93%; Wuhan East Lake: 96.15 ± 0.60%/95% |
| [145] | SAR and MODIS images | Ulva prolifera detection | AlgaeNet (U-Net based model) | Accuracy/precision/recall/f1-score/IoU | 99.83%/95.46%/92.32%/93.86%/88.43% |
| [146] | Custom seafloor debris dataset from Koh Tao (Thailand), COCO and TrashCan dataset | Automated seafloor debris detection and classification | SFD-YOLO (enhanced YOLOv8) | mAP@0.5 | 91.2% (TrashCan pretraining) |
| [147] | Custom dataset from Lynetteholm project (Denmark) | Presence/absence classification of eelgrass and temporal coverage estimation in underwater videos | Different versions of ResNet, InceptionV3, DenseNet and ViT | Accuracy, AUROC, Calibration Error (CE) | ViT: 0.902/0.959/0.087 |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [106] | A subset of the Australian benthic Benthoz15 dataset | Detect coral reefs in images | A CNN based on VGGnet | Accuracy | 92% |
| [148] | MLC (Moorea Labelled Coral) dataset | Coral classification | For feature extraction: Pre-trained VGGnet; for classification: a two-layer Multilayer Perceptron (MLP) | Accuracy | 84.5% |
| [150] | LoVe Observatory dataset | Analysis of the activity of cold-water coral polyps in a certain period of time | CNN | Accuracy | 96% |
| [151] | RGB drone images captured over the North Bay coral reef on Lord Howe Island, Australia | Coral classification | mRES-uNet | Overall accuracy/average Jaccard index | 85.74%/0.56 |
| [152] | Dataset created with images taken from Flickr, StructureRSMAS and ReefBase | Coral health classification | CoralClassify framework (Modified ResNet50) | Accuracy/Precision/ Recall/F1-Score | 87.6%/87.99%/87.06%/87.37% |
| [153] | HKCoral Dataset (collected mostly in Hong Kong) | Coral growth form segmentation | Complementary Architecture (based on a CNN): Fuses original and enhanced underwater images to improve segmentation. | mIoU | 73.54 |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [104] | Images from the internet | Ornamental fish classification | CNN model | Accuracy | 98.5% |
| [105] | Taiwan sea fish and Monterey Bay Aquarium Research Institute (MBARI) benthic animal | Underwater species classification | The DeCAF (Deep Convolutional Activation Feature) framework (CNN) | Overall error | 2.09 ± 0.51%/0.83 ± 1.78% |
| [111] | Self-constructed | Carp species classification | CNN-based method | Accuracy | 100% |
| [120] | Data collected in the Norwegian sea by a remotely operated vehicle (ROV) | Classify Norwegian seabed species | DNN, DBN and CNN | Accuracy | 88.74% by region and 92.78% by pixel |
| [157] | Train on ImageNet and F4K, Test on the Croatian fish dataset and QUT fish | Fish classification | Three types of Bilinear CNNs | Accuracy | 83.92% and 71.80% |
| [161] | SeaCLEF | Fish species classification | CNN | Accuracy | 96% |
| [162] | Fish4Knowledge | Fish recognition | CNN with SGD as the optimizer | accuracy | 98.57% |
| [163] | Fish4Knowledge | Live fish recognition | Deep architecture for live fish recognition composed principally of a ConvNet and a linear SVM | Accuracy | 98.57% |
| [165] | Fish4Knowledge | Fish classification | CNNs and several preprocessing techniques like Gaussian blurring, morphological operations and Otsu’s thresholding | Accuracy | 96.29% |
| [166] | Own dataset (water tank) | Fish behaviour classification | CNN | Accuracy | 82.5% |
| [167] | Videos collected in a water tank | Anomaly detection in carp and koi ponds by analyzing behaviours | Method based on CNN | Accuracy | 99.9% |
| [168] | Videos obtained at the culture pond of the Department of Aquaculture of National Taiwan Ocean University | Detect stressful situations when fishing fish in industry | An improved version of convolutional network-based tracker CNT (Fast-CNT2) | - | - |
| [169] | Self-created dataset | Chinese mitten crab’s gender classification | Custom CNN | Accuracy | 98.90% |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [103] | Self-constructed dataset from Austrian river | River fish classification | Pretrained VGG-16 | Accuracy | 89.4% |
| [171] | Dataset obtained from a fish farm (Khorramabad, Lorestan province, Iran) | Freshness diagnosis of common carp | Method based on CNN (VGG-16) | Accuracy | 98.21% |
| [172] | Images from Flickr | Aesthetic value of underwater images and species recognition | ZF, CNN-M and VGG-16 Conditional generative adversarial network (cGAN) | mAP | 82.4% |
| [176] | Kaggle fishing boat dataset | Fish detection, object pose estimation and horizontal alignment before class prediction | SSD and YOLOv2 as detectors, VGG-16 as pose estimation, VGG-16 and Inception V3 as classification | Loss | 60.4% |
| [177] | Kaggle fishing boat dataset | Fish classification | Two VGG-16 networks (one with transfer learning) | Accuracy | 99.38% |
| [178] | Fish Image dataset from roboflow.ai | Fish classification | VGG-16 and Darknet | Precision/recall | 70%/80% |
| [179] | Fish4Knowledge | Fish classification | VGG-8 and VGG-16 | Micro average: Precision/recall/f1-score | 99%/99%/99% |
| [180] | Fish4Knowledge and Fish-gres dataset (out of water fish images) | Fish classification | MLR-VGG16 and MLR-VGG19 | Accuracy | Fish-gres: 98.46%/Fish4Knowledge: 97.09% |
| [182] | Dataset collected from various public sources | Fish disease classification | ML algorithms, VGG16, VGG19, ResNet-50, VGG16+VGG19 and VGG16+Inception V3 | Accuracy | 99.64% |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [107] | Images from Google Search (86.400 with data augmentation) | Marine species classification | AlexNet, GoogleNet and LeNet | Accuracy | 87% |
| [112] | Dataset collected mostly from open fish markets | Identification of fish species which are found in Mauritian waters | InceptionV3 model | Accuracy | 98% |
| [119] | Kyutech10K dataset (10,728 images and 1489 videos) | Underwater species classification | A modified GoogleNet | Accuracy | 92% |
| [183] | Pretrained on ImageNet, SeaCLEF | Fish species classification | GoogleNet | Precision | 81% |
| [184] | SeaCLEF 2015 | Fish species classification | AlexNet, SVM | Precision | 66% |
| [186] | SeaCLEF 2016 | Fish detection and classification | A method based on the refined method of Mothes and Denzler [185] | MOTA value | 87.6% |
| [187] | Kaggle fishing boat dataset—The Nature conservancy | Fish classification | AlexNet and GoogleNet | Accuracy | +96% (with transfer learning) |
| [188] | Dataset created from multiple sources: OceanDark, RUIE, UIEB, UFO-120, EUVP | Multiclass classification: Fish, Turtles, Sea Urchins, Sea Cucumbers, Corals, Humans, Wreckage | DAMNet | Overall Accuracy/Precision/Recall/F1-Score/Loss | 96.93%/96.70%/96.78%/96.74%/ 0.1860 |
| [189] | SAUD [190], PHISMID [191], MSRB [192] | Binary underwater image classification (raw vs. enhanced) | Modified ResNet-18 | Accuracy, Precision, F1-score, AUC-ROC | SAUD: 96%/99%/95%/96% |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [114] | OpenROV images | Marine species detection | R-CNN | True positive detection rate | 93% |
| [115] | Data from Monterey Bay Aquarium Research Institute | Detect and classify fishes or other organisms | Video and Image Analytics for Marine Environments (VIAME) open-source computer vision library | ROC (AUC) | 70% |
| [116] | Videos obtained in marine waters of beaches and estuaries across southeast Queensland | Object detection and fish abundance estimation | Faster R-CNN | mAP | 82.4% |
| [117] | Six videos recorded at various locations within the Verde Island Passage, Philippines | Fish segmentation and identification | Fully Convolutional Residual Network (ResNet-FCN) | Precision/accuracy/recall | 65.91%/70%/84% |
| [118] | Self-created with a GoPro at the seafloor | Jellyfish detection | Faster R-CNN-based implementation of the Inception ResNet v2 | F1-score | 93.8% |
| [196] | ImageCLEF | Fish detection and classification | PVANet to DPM, R-CNN, Fast R-CNN and Faster R-CNN | mAP | 90% |
| [197] | SeaCLEF 2014 | Fish detection and classification | Fast R-CNN based on an AlexNet | mAP | 81.4% |
| [198] | Self-created underwater dataset | Underwater marine species detection and classificarion | Faster RCNN with Res2Net101 | Average precision/mAP/f1-score | 43%/71.7%/55.3% |
| [199] | LifeClef 2015 | Fish detection and segmentation | Two multi-stream fusion approaches based on Faster R-CNN | F1-score/mAP | 83.16%/73.69% |
| [202] | Self-created in a tank at laboratory | Fish detection and trajectory analysis | Faster R-CNN and YOLO-V3 | Accuracy/proportion of lost points | 98.13%/1.87% |
| [203] | Images downloaded randomly from various sources | Shrimp detection | Faster R-CNN InceptionV2 | Precision/recall/f1-score/accuracy | 97%/97%/96%/96% |
| [204] | Underwater baited video footage from New Caledonia (South Pacific) | Automated detection, species identification, and counting of deep-water snapper species | Faster R-CNN with Inception-ResNet V2 backbone | Recall/Precision/F-measure | Best specie (Etelis coruscans): 0.91/0.84/0.87 |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [113] | Self-created by the cruise muddy underwater video monitoring system at Changzhou | Live crabs’ detector | Faster MSSDLite | Average precision/f-score | 99.01%/98.94% |
| [211] | Kyutech10K | Real-time recognition and tracking of four different organisms | YOLO, Tracking Learning Detection (TLD) and MedianFlow | AUC | 53.1% |
| [124] | Underwater images and Aerial images | Marine species classification | RetinaNet as the object detector/Simple Online Realtime Tracker (SORT) for tracking | Average precision | Aerial: 40% for Ray, 25% for Diver and 75% for Shark |
| [176] | Kaggle fishing boat dataset | Fish detection, object pose estimation and horizontal alignment before class prediction | SSD and YOLOv2 as detectors, VGG-16 as pose estimation, VGG-16 and Inception V3 as classification | Loss | 60.4% |
| [214] | Self-created dataset in a water tank | Analysis of motion patterns for koi fish health monitoring | YOLOv3 | Heatmap visualization of fish locations and different plots comparison | - |
| [215] | DeepFish and OzFish [237] datasets | Fish detection | YOLO-Fish-1 and YOLO-Fish-2 (based on YOLOv3) | Precision/recall/f1-score/AP | 95%/96%/94%/96.15% |
| [216] | Images of lobsters captured in Western Australia using AUV and synthetic images | Western rock lobster detection | YOLOv3 | mAP | 46.9% |
| [217] | PASCAL VOC, Brackish and URPC Dataset | Underwater animal detection | A model based on MobileNet v2 and YOLOv4 | mAP | 92.65% |
| [219] | Self-created underwater dataset in a lake, in a laboratory tank and gathered from the internet | Fish detection | A model based on Res2Net and YOLOv5 | Precision/recall/mAP | 95.7%/88%/95.4% |
| [220] | Self-created with crawler technology and on laboratory | Jellyfish species and fish detection | An improved version of the YOLOv4-tiny | Precision/recall/mAP/f1-score | 92.62%/89.69%/95.01%/0.91 |
| [222] | Images from videos about jellyfish through the web crawler | Jellyfish detection | JF-YOLO (based on YOLOv4) | mAP/recall | 92.67%/85.74% |
| [223] | Self-created dataset collected by UAVs and images gathered from Facebook | Turtle detection | YOLOv3, YOLOv5s, and YOLOv5l | Precision/recall/f1-score | 97.21%/97.82%/97.51% |
| [224] | Images of shrimp in Recirculating Aquaculture System (RAS) culture tanks | Shrimp detection | Faster RCNN and YOLOv5m6 | MAPE | 5.48 |
| [225] | Self-created dataset in a tank | Healthy and sick fish with SVC detection | NAM-YOLOv7 | Precision/recall | 97.3%/93.8% |
| [227] | Self-created; images obtained in a controlled environment (out of water) | Chinese soft-shelled turtle detection | YOLOv7-SS | Precision/recall/mAP | 95.38%/94.68%/89.82% |
| [228] | TrashCan dataset | Underwater object detection and classification (marine debris, biological organisms and submerged artefacts) | AIT-YOLOv7 | mAP@0.5 | 81.40% |
| [229] | Not specified | Fish (sharks, jellyfish and other species) detection | YOLOv8m | Precision/recall/mAP/f1-score | 68.6%/61.24%/66.7%/64.31% |
| [230] | Pascal VOC, SEAMAPD21 and MS COCO datasets | Identification of underwater fish Species | YOLOv8-TF | mAP50/mAP50:95 | Pascal VOC: 94.60/SEAMAPD21: 61.2% |
| [231] | Self-created underwater dataset in a tank in a laboratory | Abnormal fish behaviour detection and classification | A model based on YOLOv9 | Precision/recall/mAP | 91.7%/90.4%/94.1% |
| [232] | Self-created dataset on a fishing vessel | Fish tracking and segmentation | A deep convolutional neural network and an object detector with a Kalman filter in 3D + SSD/YOLOv2 | Multiple Object Tracking Accuracy (MOTA) | 96.3% |
| [233] | Custom underwater fish dataset | Fish detection and classification | Foc_YOLOXn_ASFF (base model YOLOX-nano) | AP50/AP75/AP50-95/AR | 98.2%/93.3%/84.6%/87.0% |
| [234] | Self-created dataset: images captured in a semi-submerged underwater cage | Fish disease detection | DCW-YOLO (base model YOLOv10) | Precision/recall/mAP50/mAP50:95 | 95.46%/90.12%/96.87%/75.04% |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [238] | Proprietary Dataset (4 categories) and WildFish Dataset | Underwater Species Classification (specifically fish and shrimp species) | ADANSE ViT | Accuracy | Propietary: 92.3%/WildFish: 93.9% |
| [239] | Self-collected dataset, Croatian Fish dataset and Blue Bot dataset | Self-collected dataset | ADANSE-TL (ADANSE ViT and DenseNet-169) | Accuracy/Loss/ Precision/ Recall/ F1-Score | 96.21%/ 0.174/96.32%/96.17%/96.20% |
| [240] | A Large-Scale Dataset for Segmentation and Classification [241] | Fish detection and fish species classification | Improved Vision Transformer (IMViT) | Accuracy/ Precision/ Recall/ F1-Score | 95.73%/ 95.31%/95.14%/94.92% |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [242] | ImageNet and Fish4Knowledge | Fish classification | Classification convolution autoencoder (CCAE) | Accuracy | 73.75% and 99.28% |
| [243] | Images containing virtual underwater scenery | Underwater object semantic segmentation | Deep encoder–decoder network called SegNet | MIoU/mean accuracy | 87%/94% |
| [244] | Self-created | Semantic segmentation of tilapia fish body parts | Encoder–decoder based on the SegNet approach | MIoU | 60-85% |
| [245] | Self-made underwater image dataset | Underwater species Semantic segmentation | DeepLabv3 + | MIoU | 64.65% |
| [247] | MS-COCO dataset for pre-training and images acquired by sonar (self-created) | Underwater object detection and semantic segmentation | Mask RCNN | For detection: Precision/recall/mAP For segmentation: AP/ | For detection: 95.73%/97.21%/96.97% For segmentation: 53.23%/95.15% |
| [248] | DeepFish dataset | Underwater fish segmentation | A modified version of the U-Net | SCCE loss/SCF loss/Sparse Categorical Crossentropy Accuracy (SCCA)/IoU | 0.053/3.48/98.81%/87.6% |
| [249] | “Coralfish” and “cavefish” dataset obtained from YouTube videos | Fish detection and classification | U-Net | F1-score/mIoU | 95.04%/88.19% |
| [250] | LoVe Observatory dataset | Sponge size and behaviour tracking | U-Net | Pearson’s r | 0.98 |
| [251] | Dataset collected using submerged action cameras mounted on a baited underwater remote video (BRUV) rig. | Roman seabream detection and tracking | Mask R-CNN | 81.45%/80.28% | |
| [252] | Images collected from multiple existing datasets | Bottom water, seafloor/obstacles, and fish segmentation | DeepLabV3+ and U-Net-based model variations | IoU/Hausdorff distance (HD) | 78.76%/19.17 |
| [253] | Fish4Knowledge and Real deep-sea fish images | Underwater fish segmentation | ResNet50 backbone with an enhanced Feature Pyramid Convolutional Architecture (PAFE) | F1-score/MioU/Pixel Accuracy (PA) | 90.1%/95.1%/92.1% |
| [254] | UIIS and USIS10K | Underwater Fish Instance Segmentation for smart fisheries | AASNet framework (GELAN Backbone from YOLOv9) | mAP/AP/ | 31.7%/49.5%/35.1% |
| [257] | DeepFish (training), Seagrass and YouTube-VOS dataset | Underwater fish segmentation | CoaT Transformer backbone: Combines Conv-Attentional Image Transformer (CAIT) and Co-Scale Feature Attention Network (CFAN) | // | YouTube-VOS: 63.3%/63.9%/74.0%/62.7%/69.6% |
| [260] | SUIM dataset | Underwater Image Semantic Segmentation | SwinConvMixerUNet | mIoU | 84.83% |
| [261] | FishData, UISD, Large-scale fish dataset | Underwater image segmentation | UISFormer | mIoU/Accuracy/CPA (Category Pixel Accuracy)/F1-score | Large-scale Fish Data: 96.7%/99.34%/98.97%/98.32% |
| [262] | UIIS10K (proposed), UIIS, USIS10K | Underwater instance segmentation | UWSAM Framework | Bounding Box metrics (//) Mask metrics (//) | UWSAM-Teacher on USIS10K: = 45.8%/ = 64.1%/ = 55.1%) Mask metrics ( = 46.0%/ = 61.7%/ = 51.7%) |
| Ref | Dataset | Problem Type | Algorithms | Metric | Value (Highest) |
|---|---|---|---|---|---|
| [263] | UVLM (proposed), WebUOT (re-annotated) | Underwater video-language understanding | General VidLMs with supervised fine-tuning (VideoLLaMA3, Qwen2.5VL) | Overall accuracy | Qwen2.5VL-72B: 75.49% VideoLLaMA3-7B fine-tuned: 73.04% (+10.34) |
| [264] | FishNet++ (proposed), FishNet | Fine-grained fish species recognition (open-vocabulary classification), detection, keypoint localization | CLIP, BioCLIP, SigLIP, Qwen2.5-VL, Gemma-3, GPT-4o, YOLO-based baselines | Accuracy/IoU50/IoU90 | GPT-4o: 17.9%/YOLO-12: 95.2%; Qwen2.5-VL: 91.5%/YOLO-12: 35.2%; Qwen2.5-VL: 26.7% (Frequent Species) |
| [265] | Simulated underwater videos | Real-time VLM semantic scene analysis for AUVs | Qwen2.5-VL (72B) | Object Detection Recall (ODR)/Spatial Relationship Accuracy (SRA)/Hallucination Rate /Output Accuracy | 0.94/0.91/0.04/0.88 |
| [266] | Five public scuba diving videos captured in different locations and simulated sensor data | Context-aware message generation & recovery for diver communication (mobile VLM) | MobileVLM-3B | Purpose-Align Rate /Semantic Similarity (received vs. original message) | 80%/90% |
| [267] | AquaOV255, UOVSBench (AquaOV255 + USIS16K, SUIM, MAS3K, USIS10K, DUT-USEG) | Training-free open-vocabulary segmentation in underwater scenes | Earth2Ocean (GMG + CSA; transfers terrestrial VLMs to underwater) | mIoU | 55.24 |
| [268] | FishNet + LLaVA-1.5 datasets; extra 1100 DeepFish images | LLM-based detection and classification | FishDetectLLM (TinyLLaVA + SigLIP + StableLM-2; instruction conversations) | Accuracy | FishNet: 99.09% |
| [269] | HSCR16K | Fine-grained coral recognition (species/genera); zero-/few-shot VLM adaptation | CORAL-Adapter (morphological + taxonomic adapters on CLIP) | Accuracy | 77.27% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Lopez-Vazquez, V.; Satama-Bermeo, G.; Raheem, H.I.; Lopez-Guede, J.M. State of the Art: Analysis of Deep Learning Techniques in Images Acquired in an Aquatic Environment. Mach. Learn. Knowl. Extr. 2026, 8, 131. https://doi.org/10.3390/make8050131
Lopez-Vazquez V, Satama-Bermeo G, Raheem HI, Lopez-Guede JM. State of the Art: Analysis of Deep Learning Techniques in Images Acquired in an Aquatic Environment. Machine Learning and Knowledge Extraction. 2026; 8(5):131. https://doi.org/10.3390/make8050131
Chicago/Turabian StyleLopez-Vazquez, Vanesa, Geovanny Satama-Bermeo, Hasan Issa Raheem, and Jose Manuel Lopez-Guede. 2026. "State of the Art: Analysis of Deep Learning Techniques in Images Acquired in an Aquatic Environment" Machine Learning and Knowledge Extraction 8, no. 5: 131. https://doi.org/10.3390/make8050131
APA StyleLopez-Vazquez, V., Satama-Bermeo, G., Raheem, H. I., & Lopez-Guede, J. M. (2026). State of the Art: Analysis of Deep Learning Techniques in Images Acquired in an Aquatic Environment. Machine Learning and Knowledge Extraction, 8(5), 131. https://doi.org/10.3390/make8050131

