Figure 1.
Implementation-faithful overview of the enhanced YOLO11n architecture and matched experimental protocol. (a) Selected RFA- and RCSA-enhanced feature stages, multi-scale neck, and unchanged detection head; arrows denote forward feature flow, orange blocks denote RFA, green blocks denote RCSA-enabled processing, blue blocks denote input/context operations, and the purple block denotes the unchanged detection head. (b) Provider split, split-integrity audit, common 100-epoch training protocol, and reported evaluation outputs. Abbreviations: RFA, Receptive-Field Aggregation; RCSA, Residual Channel-Spatial Recalibration; DRG, Dynamic Residual Group; SPPF, Spatial Pyramid Pooling-Fast; C2PSA, C2 Position-Sensitive Attention; P, precision; R, recall; mAP, mean average precision; GFLOPs, giga floating-point operations.
Figure 1.
Implementation-faithful overview of the enhanced YOLO11n architecture and matched experimental protocol. (a) Selected RFA- and RCSA-enhanced feature stages, multi-scale neck, and unchanged detection head; arrows denote forward feature flow, orange blocks denote RFA, green blocks denote RCSA-enabled processing, blue blocks denote input/context operations, and the purple block denotes the unchanged detection head. (b) Provider split, split-integrity audit, common 100-epoch training protocol, and reported evaluation outputs. Abbreviations: RFA, Receptive-Field Aggregation; RCSA, Residual Channel-Spatial Recalibration; DRG, Dynamic Residual Group; SPPF, Spatial Pyramid Pooling-Fast; C2PSA, C2 Position-Sensitive Attention; P, precision; R, recall; mAP, mean average precision; GFLOPs, giga floating-point operations.
Figure 2.
Commercially available open-source ROV platform involved in related experiments in which the authors participated: (a) underwater operating view; (b) frontal view of the camera-equipped platform. The vehicle was not designed or manufactured by the authors. The photographs illustrate the application context and are not embedded detector benchmarks.
Figure 2.
Commercially available open-source ROV platform involved in related experiments in which the authors participated: (a) underwater operating view; (b) frontal view of the camera-equipped platform. The vehicle was not designed or manufactured by the authors. The photographs illustrate the application context and are not embedded detector benchmarks.
Figure 3.
Representative samples from the public underwater marine-debris dataset used in this study.
Figure 3.
Representative samples from the public underwater marine-debris dataset used in this study.
Figure 4.
Original YOLO11n network structure. Arrows indicate forward feature propagation; colors distinguish convolution, C3k2, spatial-pyramid/attention, concatenation/upsampling, and detection-stage blocks.
Figure 4.
Original YOLO11n network structure. Arrows indicate forward feature propagation; colors distinguish convolution, C3k2, spatial-pyramid/attention, concatenation/upsampling, and detection-stage blocks.
Figure 5.
Improved YOLO11n network with RFA and RCSA. Solid arrows indicate forward feature propagation. Blue blocks denote input/context operations, orange blocks denote RFA-based context aggregation, green blocks denote RCSA-enabled recalibration and multi-scale fusion, and the purple block denotes the unchanged detection head.
Figure 5.
Improved YOLO11n network with RFA and RCSA. Solid arrows indicate forward feature propagation. Blue blocks denote input/context operations, orange blocks denote RFA-based context aggregation, green blocks denote RCSA-enabled recalibration and multi-scale fusion, and the purple block denotes the unchanged detection head.
Figure 6.
Receptive-Field Aggregation (RFA) module. (
a) Channel-partitioned aggregation with progressively coupled local operators. (
b) Progressive local operators with effective receptive fields of approximately
,
, and
. Arrows denote feature propagation, and colors distinguish input, local-operator, fusion, and intermediate-operation blocks. The schematic was redrawn based on the receptive-field attention concept in Ref. [
23].
Figure 6.
Receptive-Field Aggregation (RFA) module. (
a) Channel-partitioned aggregation with progressively coupled local operators. (
b) Progressive local operators with effective receptive fields of approximately
,
, and
. Arrows denote feature propagation, and colors distinguish input, local-operator, fusion, and intermediate-operation blocks. The schematic was redrawn based on the receptive-field attention concept in Ref. [
23].
Figure 7.
RCSA implementation. The dynamic residual-group topology is adapted from DRPCA-Net [
24] and integrated into YOLO11n for underwater feature recalibration.
Figure 7.
RCSA implementation. The dynamic residual-group topology is adapted from DRPCA-Net [
24] and integrated into YOLO11n for underwater feature recalibration.
Figure 8.
Duration-matched experiment design. All four architectures use a common 100-epoch schedule, seeds 42, 2026, and 3407, and the same source-split training protocol; final performance is evaluated on the audited holdout.
Figure 8.
Duration-matched experiment design. All four architectures use a common 100-epoch schedule, seeds 42, 2026, and 3407, and the same source-split training protocol; final performance is evaluated on the audited holdout.
Figure 9.
F1 score as a function of the confidence threshold for (a) YOLO11n and (b) YOLO11n + RFA + RCSA. Each colored curve represents the F1–confidence relationship for an individual debris class, while the thick blue curve represents the aggregated performance across all 15 classes. The annotated point on the overall curve indicates the confidence threshold at which the highest overall F1 score is obtained.
Figure 9.
F1 score as a function of the confidence threshold for (a) YOLO11n and (b) YOLO11n + RFA + RCSA. Each colored curve represents the F1–confidence relationship for an individual debris class, while the thick blue curve represents the aggregated performance across all 15 classes. The annotated point on the overall curve indicates the confidence threshold at which the highest overall F1 score is obtained.
Figure 10.
Visual comparison of mAP@0.5 and mAP@0.5:0.95 for the single-module screening experiments.
Figure 10.
Visual comparison of mAP@0.5 and mAP@0.5:0.95 for the single-module screening experiments.
Figure 11.
RFA-centered 100-epoch screening. RFA + RCSA is the three-seed mean; the other enhancement pairings are seed-42 screening runs.
Figure 11.
RFA-centered 100-epoch screening. RFA + RCSA is the three-seed mean; the other enhancement pairings are seed-42 screening runs.
Figure 12.
Four-model 100-epoch comparison requested in review. All error bars show standard deviation across seeds 42, 2026, and 3407.
Figure 12.
Four-model 100-epoch comparison requested in review. All error bars show standard deviation across seeds 42, 2026, and 3407.
Figure 13.
Accuracy–complexity comparison of YOLOv8n, YOLO11n, and YOLO11n + RFA + RCSA. The panels report only static model complexity and validation accuracy; no embedded-speed, latency, memory, power, or energy quantity is reported because none was measured.
Figure 13.
Accuracy–complexity comparison of YOLOv8n, YOLO11n, and YOLO11n + RFA + RCSA. The panels report only static model complexity and validation accuracy; no embedded-speed, latency, memory, power, or energy quantity is reported because none was measured.
Figure 14.
Class-wise audited-holdout AP@0.5 for RFA + RCSA. Error bars show the between-seed standard deviation. Orange bars identify rare classes with fewer than 10 holdout instances, blue bars identify all other classes, and the dashed green line denotes the all-class mean AP@0.5.
Figure 14.
Class-wise audited-holdout AP@0.5 for RFA + RCSA. Error bars show the between-seed standard deviation. Orange bars identify rare classes with fewer than 10 holdout instances, blue bars identify all other classes, and the dashed green line denotes the all-class mean AP@0.5.
Table 1.
Comparison of representative underwater-vision studies and the scope of the present work.
Table 1.
Comparison of representative underwater-vision studies and the scope of the present work.
| Study | Task and Data | Representative Method | Recurring Limitation/Relevance |
|---|
| Fulton et al. [25] | Marine-litter detection | Multiple deep object
detectors | Dataset- and platform-specific evaluation |
| Hong et al. [26] | TrashCan detection and segmentation | Faster
R-CNN; Mask R-CNN | Benchmark is not edge optimized |
| Islam et al. [27] | SUIM semantic segmentation | FCN, U-Net, SegNet,
PSPNet, DeepLab-v3, SUIM-Net | Segmentation task differs from debris
detection |
| Zhao et al. [28] | DUO and URPC2020 detection | FEB-YOLOv8 | Different classes and splits |
| Cai et al. [29] | URPC2020 detection | AGW-YOLOv8 | Dataset- and
hardware-dependent evidence |
| Huang et al. [30] | Underwater garbage detection | YOLO-MES | Closest lightweight study; unmatched protocol |
| Luo and Eljamal [31] | Underwater debris detection | SPyramidLightNet | Different evaluation protocol |
| This study | 15-class debris detection | YOLO11n + RFA + RCSA | Audited
holdout; Jetson validation remains open |
Table 2.
Class distribution in the original provider training and validation partitions.
Table 2.
Class distribution in the original provider training and validation partitions.
| Class | Train Images | Val Images | Total Images | Train Instances | Val
Instances | Total Instances |
|---|
| mask | 1367 | 77 | 1444 | 4310 | 90 | 4400 |
| can | 265 | 18 | 283 | 370 | 20 | 390 |
| cellphone | 702 | 61 | 763 | 804 | 71 | 875 |
| electronics | 291 | 27 | 318 | 411 | 40 | 451 |
| gbottle | 480 | 36 | 516 | 1017 | 82 | 1099 |
| glove | 1265 | 37 | 1302 | 3501 | 55 | 3556 |
| metal | 90 | 10 | 100 | 168 | 22 | 190 |
| misc | 510 | 49 | 559 | 519 | 52 | 571 |
| net | 1503 | 146 | 1649 | 1596 | 148 | 1744 |
| pbag | 2908 | 290 | 3198 | 3397 | 330 | 3727 |
| pbottle | 1440 | 122 | 1562 | 2790 | 284 | 3074 |
| plastic | 411 | 51 | 462 | 530 | 59 | 589 |
| rod | 51 | 7 | 58 | 78 | 9 | 87 |
| sunglasses | 42 | 3 | 45 | 42 | 3 | 45 |
| tire | 1464 | 143 | 1607 | 6816 | 627 | 7443 |
Table 3.
Implementation roles of the proposed network modifications.
Table 3.
Implementation roles of the proposed network modifications.
| Component | Network Location | Operation and
Output Handling |
|---|
| C3k2_RFA | Backbone feature-extraction stage | Multi-branch local
operators aggregate responses with different effective receptive fields
before 1 × 1 fusion. The host-stage spatial resolution and channel width
are retained. |
| DRG–RCSAB (RCSA) | Selected backbone and neck blocks | Channel and
spatial attention are embedded in a residual group to recalibrate fused
features while retaining shortcut information. Identity or projection
shortcut P(·) aligns dimensions, and the detection head is unchanged. |
Table 4.
Experimental environment and training configuration.
Table 4.
Experimental environment and training configuration.
| Item | Configuration |
|---|
| Hardware | Intel Xeon Gold 5220 CPU (Intel Corporation, Santa Clara, CA, USA); 32 GB memory; NVIDIA GeForce RTX 5070 12 GB GPU (NVIDIA Corporation, Santa Clara, CA, USA) |
| Software | Windows 11 (Microsoft Corporation, Redmond, WA, USA); Python 3.11.14 (https://www.python.org, accessed on 29 July 2026); PyTorch 2.9.0 (https://pytorch.org, accessed on 29 July 2026); CUDA 12.8 and NVIDIA Driver 595.79 (NVIDIA Corporation) |
| Framework | Ultralytics YOLO11 training pipeline implemented in PyTorch
[32] |
| Input and batch | Input size 640 × 640; batch size 30 |
| Preprocessing and augmentation | Resize/normalize; hsv_h = 0.015,
hsv_s = 0.7, hsv_v = 0.4; translate = 0.1; scale = 0.5; horizontal flip = 0.5;
Mosaic = 1.0; degrees = 0; vertical flip = 0; MixUp = 0 |
| Common architecture schedule | 100 epochs for YOLO11n, YOLO11n + RFA,
YOLO11n + RCSA, and YOLO11n + RFA + RCSA |
| Repeated-seed protocol | Seeds 42, 2026, and 3407 for YOLO11n, YOLO11n +
RFA, YOLO11n + RCSA, and YOLO11n + RFA + RCSA; deterministic mode
enabled |
| Optimization | Ultralytics 8.3.163 optimizer = auto selected Nesterov SGD;
initial lr = 0.01; momentum = 0.9; nominal weight decay = 0.0005 (effective
0.00046875 for batch 30, accumulation 2, nbs = 64); warm-up = 3 epochs;
patience = 100 |
| Initialization | Pretrained yolo11n.pt weights; AMP enabled |
| Learning-rate schedule | Linear decay (cos_lr = false); final ten epochs
without Mosaic (close_mosaic = 10) |
| Holdout evaluation | Group-disjoint test split; 498 images, 963
instances; not used for checkpoint selection |
Table 5.
Duration-matched 100-epoch validation comparison. All four models are reported as mean ± standard deviation over seeds 42, 2026, and 3407.
Table 5.
Duration-matched 100-epoch validation comparison. All four models are reported as mean ± standard deviation over seeds 42, 2026, and 3407.
| Model | Epochs | Runs/Seeds | Precision | Recall | mAP@0.5 | mAP@0.5:0.95 |
|---|
| YOLO11n | 100 | 3: 42, 2026, 3407 | 0.851 ± 0.021 | 0.769 ± 0.012 | 0.819 ± 0.008 | 0.506 ± 0.005 |
| YOLO11n + RFA | 100 | 3: 42, 2026, 3407 | 0.832 ± 0.006 | 0.788 ±
0.014 | 0.831 ± 0.002 | 0.509 ± 0.005 |
| YOLO11n + RCSA | 100 | 3: 42, 2026, 3407 | 0.830 ± 0.015 | 0.781 ±
0.012 | 0.810 ± 0.005 | 0.509 ± 0.004 |
| YOLO11n + RFA + RCSA | 100 | 3: 42, 2026, 3407 | 0.861 ±
0.016 | 0.801 ± 0.012 | 0.848 ± 0.001 | 0.511 ± 0.002 |
Table 6.
Unified single-module and RFA-centered screening under the common 100-epoch schedule.
Table 6.
Unified single-module and RFA-centered screening under the common 100-epoch schedule.
| Group | Model | Epochs | Runs | Precision | Recall | mAP@0.5 | mAP@0.5:0.95 |
|---|
| Single module (seed 42) | YOLO11n | 100 | 1 | 0.836 | 0.770 | 0.810 | 0.502 |
| | YOLO11n + RFA | 100 | 1 | 0.833 | 0.787 | 0.832 | 0.515 |
| | YOLO11n + CSA_ConvBlock | 100 | 1 | 0.858 | 0.771 | 0.814 | 0.510 |
| | YOLO11n + Di_SpAM | 100 | 1 | 0.801 | 0.723 | 0.773 | 0.478 |
| | YOLO11n + RCSA | 100 | 1 | 0.830 | 0.781 | 0.810 | 0.509 |
| | YOLO11n + CoordAtt | 100 | 1 | 0.798 | 0.765 | 0.799 | 0.493 |
| | YOLO11n + iRMB | 100 | 1 | 0.837 | 0.725 | 0.783 | 0.482 |
| RFA pairing (seed 42) | RFA + CoordAtt | 100 | 1 | 0.853 | 0.740 | 0.805 | 0.494 |
| | RFA + DySample + ECA | 100 | 1 | 0.864 | 0.765 | 0.796 | 0.485 |
| | RFA + Di_SpAM | 100 | 1 | 0.756 | 0.696 | 0.747 | 0.473 |
| | RFA + CSA_ConvBlock | 100 | 1 | 0.826 | 0.748 | 0.815 | 0.500 |
| Selected model | RFA + RCSA | 100 | 3 | 0.861 ± 0.016 | 0.801 ± 0.012 | 0.848 ± 0.001 | 0.511 ±
0.002 |
Table 7.
Computational complexity and embedded-deployment measurement status of the evaluated models.
Table 7.
Computational complexity and embedded-deployment measurement status of the evaluated models.
| Model | Params/M | GFLOPs | Weights/MB | Embedded Measurement Status | mAP@0.5 | mAP@0.5:0.95 |
|---|
| YOLO11n | 2.59 | 6.46 | 5.23 | Not measured | 0.819 ± 0.008 | 0.506
± 0.005 |
| YOLO11n + RFA | 2.60 | 6.63 | 5.25 | Not measured | 0.831 ± 0.002 | 0.509 ± 0.005 |
| YOLO11n + RFA + RCSA | 4.19 | 8.91 | 8.43 | Not measured | 0.848
± 0.001 | 0.511 ± 0.002 |
Table 8.
Same-dataset validation comparison with YOLOv8n. Runs are identified; embedded FPS, latency, peak memory, power, and energy per frame were not measured.
Table 8.
Same-dataset validation comparison with YOLOv8n. Runs are identified; embedded FPS, latency, peak memory, power, and energy per frame were not measured.
| Model | Runs | Params/M | GFLOPs | Embedded Measurement Status | Precision | Recall | mAP@0.5 | mAP@0.5:0.95 |
|---|
| YOLOv8n | 1 | 3.20 | 8.70 | Not measured | 0.842 | 0.779 | 0.821 | 0.508 |
| YOLO11n | 3 | 2.59 | 6.46 | Not measured | 0.851 ±
0.021 | 0.769 ± 0.012 | 0.819 ± 0.008 | 0.506 ± 0.005 |
| YOLO11n + RFA + RCSA | 3 | 4.19 | 8.91 | Not measured | 0.861
± 0.016 | 0.801 ± 0.012 | 0.848 ± 0.001 | 0.511 ± 0.002 |
Table 9.
Class-wise audited-holdout AP@0.5 (the bold row is the all-class aggregate), reported as mean ± standard deviation and a between-seed 95% t interval across seeds 42, 2026, and 3407.
Table 9.
Class-wise audited-holdout AP@0.5 (the bold row is the all-class aggregate), reported as mean ± standard deviation and a between-seed 95% t interval across seeds 42, 2026, and 3407.
| Class | Test Inst. | AP@0.5 Mean ± SD [95% CI] | Class | Test Inst. | AP@0.5 Mean ± SD [95% CI] |
|---|
| mask | 34 | 0.730 ± 0.040 [0.630, 0.831] | net | 65 | 0.923 ±
0.019 [0.876, 0.969] |
| can | 19 | 0.730 ± 0.013 [0.697, 0.762] | pbag | 166 | 0.957 ±
0.010 [0.931, 0.982] |
| cellphone | 46 | 0.995 ± 0.000 [0.995, 0.995] | pbottle | 126 | 0.862 ± 0.007 [0.844, 0.881] |
| electronics | 19 | 0.835 ± 0.017 [0.793, 0.877] | plastic | 40 | 0.678 ± 0.060 [0.529, 0.827] |
| gbottle | 63 | 0.643 ± 0.010 [0.618, 0.667] | rod | 2 | 0.414 ±
0.198 [0.000, 0.907] |
| glove | 34 | 0.871 ± 0.016 [0.830, 0.911] | sunglasses | 1 | 0.995
± 0.000 [0.995, 0.995] |
| metal | 5 | 0.437 ± 0.103 [0.182, 0.692] | tire | 310 | 0.796 ±
0.014 [0.760, 0.831] |
| misc | 33 | 0.814 ± 0.011 [0.788, 0.841] | all classes | 963 | 0.779 ± 0.018 [0.734, 0.823]
|
Table 10.
Additional measured evaluations. Values are mean ± standard deviation over the three completed RFA + RCSA checkpoints; no values in this table are simulated.
Table 10.
Additional measured evaluations. Values are mean ± standard deviation over the three completed RFA + RCSA checkpoints; no values in this table are simulated.
| Evaluation | Data/Protocol | mAP@0.5 | mAP@0.5:0.95 | Interpretation |
|---|
| Reference holdout | Audited group-disjoint holdout; 498 images/963
instances | 0.779 ± 0.018 | 0.467 ± 0.007 | Unmodified reference
evaluation. |
| Brightness 0.70 | Same 498-image/963-instance holdout; test-time
transform | 0.776 ± 0.019 | 0.469 ± 0.002 | Within the observed seed
variation. |
| Brightness 1.30 | Same 498-image/963-instance holdout; test-time
transform | 0.773 ± 0.022 | 0.466 ± 0.006 | Small average change. |
| Contrast 0.70 | Same 498-image/963-instance holdout; test-time
transform | 0.644 ± 0.025 | 0.397 ± 0.014 | Substantial degradation
under reduced contrast. |
| Contrast 1.30 | Same 498-image/963-instance holdout; test-time
transform | 0.728 ± 0.019 | 0.427 ± 0.004 | Moderate degradation. |
| rare-class oversampling | Original training partition (10,884 unique images); sampling frequency only;
no new or synthetic images; metal, rod, and sunglasses sampled | 0.772 ± 0.006 | 0.464 ± 0.007 | No overall holdout improvement over the reference
protocol. |
| External material-domain test | TrashCan Material metal/plastic subset;
1204 images, 599 mapped target instances | 0.031 ± 0.023 | 0.017 ±
0.012 | Severe taxonomy and domain shift; not evidence of cross-domain
robustness. |