3.1.2. Ablation Study on the SE Modules
- (1)
Ablation study on SE insertion positions
Table 3 compares the detection performance of the SE module at different insertion positions around the SPPF layer. With the SE structure and compression ratio fixed at 16, the detection metrics of YOLOv8s with the SE module inserted at different positions were all higher than those of the YOLOv8s baseline. The SPPF-after position achieved AP50 of 97.8% and 96.1% in the single-seedling and four-seedling scenarios, respectively, which were 0.7 and 0.9 percentage points higher than those at the SPPF-before position and 1.4 and 1.5 percentage points higher than those at the Neck P5 position. The AP50-95 reached 71.3% and 68.6%, respectively, also higher than the other two positions. These results indicate that under the same SE structure and training conditions, the insertion position of the SE module affects the detection performance, and the SPPF-after position achieves the best performance in both test scenarios. After multi-scale spatial feature aggregation by SPPF, the SE module recalibrates the high-level semantic features through channel-wise weighting and passes the enhanced features to the Neck for subsequent multi-scale fusion, which is beneficial for improving the representation of informative seedling features. Based on the ablation results and the network structural characteristics, the SPPF-after position was selected as the insertion location for the SE module.
- (2)
Comparison of different attention mechanisms
Table 4 shows that different attention mechanisms exhibited varying detection performance in rice seedling detection. SE-YOLOv8s obtained the highest values of the reported detection metrics in both single-seedling and four-seedling scenarios. In the single-seedling scenario, its AP50, AP75, and AP50-95 reached 97.8%, 88.9%, and 71.3%, respectively; in the four-seedling scenario, they reached 96.1%, 86.5%, and 68.6%, respectively. ECA-YOLOv8s and EMA-YOLOv8s achieved performance close to that of SE-YOLOv8s. Specifically, ECA-YOLOv8s achieved AP50 of 97.2% and 95.4% in the single-seedling and four-seedling scenarios, which were only 0.6 and 0.7 percentage points lower than SE-YOLOv8s, and its AP50-95 was 0.4 percentage points lower in both scenarios. EMA-YOLOv8s also achieved good overall performance, but was slightly inferior to SE and ECA in AP50, AP75, and AP50-95. In contrast, CBAM and CA showed unstable improvements on some metrics, indicating that different attention mechanisms have different effects on feature representation of rice seedlings. Because each attention configuration was trained only once, these differences are descriptive rather than statistically supported; the margins of SE over ECA (0.6 and 0.7 percentage points in AP50, and 0.4 percentage points in AP50-95) are comparable to the run-to-run variation reported in
Table 5, where the standard deviation of recall reached 4.8 percentage points. Overall, SE performed comparably to, and marginally above, ECA under this single-run comparison, and the SE attention mechanism demonstrated good applicability on the current dataset and network structure. Combined with the insertion position ablation results, the SE-YOLOv8s architecture with the SE module embedded after SPPF was finally adopted.
- (3)
Analysis of the SE compression ratio
The effects of the SE reduction ratio on parameter count were calculated theoretically for
r = 8, 16, and 32, with
C = 512 input channels and no bias terms in the two fully connected layers. The intermediate channel number is
C/
r, and the additional parameters introduced by the SE module are Δ
N = 2
C2/
r. The parameter increases for the three settings are 65,536, 32,768, and 16,384, respectively (
Table 6).
- (4)
Performance comparison of YOLOv8s and SE-YOLOv8s
The detection performance of YOLOv8s and SE-YOLOv8s in single-seedling and four-seedling scenarios is presented in
Table 5.
In the single-seedling scenario, SE-YOLOv8s achieved a precision of 94.5 ± 1.3%, recall of 94.9 ± 2.2%, mAP@0.5 of 97.8 ± 1.2%, and F1-score of 94.7 ± 1.0%. These values exceeded those of YOLOv8s by 2.2, 2.9, 2.9, and 2.6 percentage points, respectively. The mAP@0.5:0.95 was 71.3 ± 1.9%, 3.4 percentage points higher than the baseline value of 67.9 ± 3.5%, and AP@0.75 was 88.9 ± 2.5%, 4.3 percentage points higher than the baseline value of 84.6 ± 2.5%. These results indicate that the SE module improved both detection sensitivity and localization performance on this test set.
In the four-seedling scenario, SE-YOLOv8s achieved a precision of 92.8 ± 1.3%, recall of 92.0 ± 3%, mAP@0.5 of 96.1 ± 2.2%, and F1-score of 92.4 ± 1.8%, exceeding the corresponding YOLOv8s values by 3.1, 1.7, 3.7, and 2.5 percentage points, respectively. The mAP@0.5:0.95 and AP@0.75 were 68.6 ± 3.4% and 86.5 ± 2.3%, respectively, representing increases of 3.8 and 4.4 percentage points over the baseline. The performance was lower than in the single-seedling scenario, which is consistent with the greater leaf overlap and boundary ambiguity in images containing four seedlings.
On the PC platform, SE-YOLOv8s processed the single-seedling and four-seedling scenarios at 29.8 and 28.8 FPS, respectively, which were 0.5 and 1.7 FPS higher than those of YOLOv8s. The parameter counts, GFLOPs, and file sizes were calculated consistently across models as described for
Table 1, with weights saved in FP32 format and without optimizer states. These FPS results do not represent the processing speed on the Raspberry Pi platform.
Paired
t-tests were performed on the five independent runs for each metric. In the single-seedling scenario, the improvements in mAP@0.5 (
p = 0.015) and precision (
p = 0.046) were statistically significant (
p < 0.05), while the improvement in recall (
p = 0.089) was not statistically significant. In the four-seedling scenario, the improvement in mAP@0.5 (
p = 0.009) was statistically significant (
p < 0.05), while the improvements in precision (
p = 0.087) and recall (
p = 0.071) were not statistically significant. Typical rice images with leaf overlap, water surface background, and local target interference were selected for qualitative comparison, as shown in
Figure 12, to further visually analyze the detection differences before and after the introduction of the SE module.
Figure 12 presents typical detection results of YOLOv8s and SE-YOLOv8s on the same rice images. The yellow dashed boxes indicate the ground-truth annotations, the blue solid boxes indicate the predicted boxes that successfully match the ground-truth targets, and the pink solid boxes indicate the unmatched predictions.
In Example 1, YOLOv8s successfully matched 3 targets, with 2 unmatched predictions and 4 missed detections. SE-YOLOv8s successfully matched 5 targets, with unmatched predictions reduced to 1 and missed detections reduced to 2. The zoomed-in views show that for the same seedling with obvious leaf overlap, SE-YOLOv8s achieved better alignment between the predicted boxes and the ground-truth annotations.
In Example 2, both models successfully matched 2 targets, but YOLOv8s generated 2 unmatched predictions, while SE-YOLOv8s generated none. The zoomed-in views show that YOLOv8s generated extra detections around the leaf regions near the targets, whereas SE-YOLOv8s reduced such interference.
Therefore, the introduction of the SE module improved overall detection metrics, including precision, recall, and AP, and achieved better detection and localization performance under leaf overlap and complex water surface conditions, while reducing false positives in background regions.
3.1.3. Comparison with Representative Detection Models
To further evaluate the comprehensive performance of SE-YOLOv8s, it was compared with Faster R-CNN, YOLOv5s, and YOLOv7-tiny under identical data partitioning, input resolution, and test protocols. The results are shown in
Table 7. YOLOv8s results are given in
Table 5 and are therefore not repeated here.
Table 7 shows that SE-YOLOv8s achieved the highest detection metrics among all compared models in both single-seedling and four-seedling scenarios. In the single-seedling scenario, its precision of 94.5% exceeded those of Faster R-CNN, YOLOv5s, and YOLOv7-tiny by 4.4, 1.4, and 3.8 percentage points, respectively. Its recall of 94.9% exceeded the corresponding values by 7.8, 4.9, and 6.3 percentage points. Its mAP@0.5 of 97.8% was 7.6, 3.7, and 5.1 percentage points higher. The above comparisons are based on the test conditions of this study.
In the four-seedling scenario, SE-YOLOv8s achieved an mAP@0.5 of 96.1%, 7.9, 4.1, and 5.4 percentage points higher than Faster R-CNN, YOLOv5s, and YOLOv7-tiny, respectively. Its mAP@0.5:0.95 was 68.6%, 8.7, 4.4, and 6.0 percentage points higher, and its AP75 was 86.5%, 10.9, 5.4, and 7.4 percentage points higher. These results indicate that SE-YOLOv8s achieved better detection and localization performance among the compared models under the tested scenarios.
In terms of resource metrics, SE-YOLOv8s has 11.169 M parameters, 28.6474 GFLOPs, and a weight file size of 45.020 MB in FP32 precision. Compared to YOLOv8s, it introduces 32,768 additional parameters, with an increase of approximately 0.294%. Its computational cost is lower than that of Faster R-CNN (128.0 GFLOPs), but higher than those of YOLOv5s (15.8 GFLOPs) and YOLOv7-tiny (13.2 GFLOPs). SE-YOLOv8s achieved 29.8 and 28.8 FPS on the PC platform in the single-seedling and four-seedling scenarios, respectively, which were 0.5 and 1.7 FPS higher than YOLOv8s, and 0.3 and 0.9 FPS higher than YOLOv7-tiny. These FPS values were measured only on a PC platform and do not represent the actual processing performance on the Raspberry Pi.