Figure 1.
Overall architecture of RIF-YOLO-N. The proposed RIF module redesigns the fine-scale P3 pathway by fusing backbone and neck P3 features to generate , while preserving the original P4/P5 pathways and multi-scale anchor-free detection head.
Figure 1.
Overall architecture of RIF-YOLO-N. The proposed RIF module redesigns the fine-scale P3 pathway by fusing backbone and neck P3 features to generate , while preserving the original P4/P5 pathways and multi-scale anchor-free detection head.
Figure 2.
Internal structure of the proposed RIF module. Backbone and neck P3 features are projected into a common feature space, adaptively fused through a gating mechanism, and refined using an identity-safe residual branch to generate .
Figure 2.
Internal structure of the proposed RIF module. Backbone and neck P3 features are projected into a common feature space, adaptively fused through a gating mechanism, and refined using an identity-safe residual branch to generate .
Figure 3.
RGFT workflow for RIF-YOLO-N. The detector is first trained at base resolution to obtain the best checkpoint , and is then fine-tuned at a higher resolution using the same architecture to produce the final checkpoint .
Figure 3.
RGFT workflow for RIF-YOLO-N. The detector is first trained at base resolution to obtain the best checkpoint , and is then fine-tuned at a higher resolution using the same architecture to produce the final checkpoint .
Figure 4.
Representative sample images from the three evaluation datasets: (a) VisDrone, (b) VEDAI, and (c) UAVDT, illustrating variations in scene density, object scale, viewpoint, and background complexity.
Figure 4.
Representative sample images from the three evaluation datasets: (a) VisDrone, (b) VEDAI, and (c) UAVDT, illustrating variations in scene density, object scale, viewpoint, and background complexity.
Figure 5.
Controlled mAP comparison between YOLOv8n and RIF-YOLO-N on VisDrone at .
Figure 5.
Controlled mAP comparison between YOLOv8n and RIF-YOLO-N on VisDrone at .
Figure 6.
Controlled mAP comparison between YOLOv8n and RIF-YOLO-N on UAVDT at .
Figure 6.
Controlled mAP comparison between YOLOv8n and RIF-YOLO-N on UAVDT at .
Figure 7.
Controlled mAP comparison between YOLOv8n and RIF-YOLO-N on VEDAI at .
Figure 7.
Controlled mAP comparison between YOLOv8n and RIF-YOLO-N on VEDAI at .
Figure 8.
Size-stratified comparison between YOLOv8n and RIF-YOLO-N on the VisDrone validation set.
Figure 8.
Size-stratified comparison between YOLOv8n and RIF-YOLO-N on the VisDrone validation set.
Figure 9.
Class-wise performance change of RIF-YOLO-N relative to YOLOv8n on VisDrone. (a) AP50 change and (b) AP50:95 change. Positive values indicate class-wise improvement, while negative values indicate degradation.
Figure 9.
Class-wise performance change of RIF-YOLO-N relative to YOLOv8n on VisDrone. (a) AP50 change and (b) AP50:95 change. Positive values indicate class-wise improvement, while negative values indicate degradation.
Figure 10.
Detection performance of RIF-YOLO-N on the VisDrone validation set at : (a) precision–recall curves for individual object classes and the overall detector; (b) F1-score versus confidence-threshold curves for individual object classes and the overall detector.
Figure 10.
Detection performance of RIF-YOLO-N on the VisDrone validation set at : (a) precision–recall curves for individual object classes and the overall detector; (b) F1-score versus confidence-threshold curves for individual object classes and the overall detector.
Figure 11.
Normalized confusion matrix of RIF-YOLO-N on the VisDrone validation set at .
Figure 11.
Normalized confusion matrix of RIF-YOLO-N on the VisDrone validation set at .
Figure 12.
Qualitative prediction comparison on a representative VisDrone validation sample at .
Figure 12.
Qualitative prediction comparison on a representative VisDrone validation sample at .
Figure 13.
RIF-YOLO-N ablation results on the VisDrone validation set at : (a) detection accuracy in terms of ; (b) computational complexity in terms of GFLOPs.
Figure 13.
RIF-YOLO-N ablation results on the VisDrone validation set at : (a) detection accuracy in terms of ; (b) computational complexity in terms of GFLOPs.
Figure 14.
(a) absolute detection performance of RIF-YOLO-N and the alternative architectures in terms of and ; (b) relative and loss of each alternative architecture with respect to RIF-YOLO-N.
Figure 14.
(a) absolute detection performance of RIF-YOLO-N and the alternative architectures in terms of and ; (b) relative and loss of each alternative architecture with respect to RIF-YOLO-N.
Figure 15.
Summary of observed performance trends. (a) VisDrone gains over YOLOv8n; (b) across input resolutions; (c) size-stratified gains; (d) class-wise changes; (e) UAVDT external-validation performance; and (f) accuracy–complexity trade-off across architectural variants.
Figure 15.
Summary of observed performance trends. (a) VisDrone gains over YOLOv8n; (b) across input resolutions; (c) size-stratified gains; (d) class-wise changes; (e) UAVDT external-validation performance; and (f) accuracy–complexity trade-off across architectural variants.
Table 1.
Qualitative comparison of representative method families for tiny-object detection.
Table 1.
Qualitative comparison of representative method families for tiny-object detection.
| Method Family | Fine-Scale Enhancement | Added P2/High-Res. Scale | Attention/Context | Lightweight Focus | Broader Detector Redesign | Controlled P3 Fusion |
|---|
| Feature-pyramid enhancement [9,19] | ✓ | △ | △ | △ | △ | × |
| High-resolution/P2 methods [10,12,20,27] | ✓ | ✓ | △ | △ | ✓ | × |
| Attention/context methods [24,25,26] | △ | × | ✓ | △ | △ | × |
| Lightweight detector redesigns [13,14,15] | △ | △ | △ | ✓ | ✓ | × |
| RIF-YOLO-N | ✓ | × | × | ✓ | × | ✓ |
Table 2.
Conceptual operator-level comparison of feature-fusion formulations and RIF-YOLO-N.
Table 2.
Conceptual operator-level comparison of feature-fusion formulations and RIF-YOLO-N.
| Property | Cross-Layer Feature Fusion | Residual Attention | RIF-YOLO-N |
|---|
| Feature relationship | Features from different layers; spatial alignment may be required | Attention applied to a main or transformed feature stream | Backbone and neck P3 features at the same stride-8 scale |
| Fusion role | Constructs a combined feature representation | Reweights feature responses | Introduces a gated residual correction to the preserved P3 representation |
| Spatial alignment | Required when feature resolutions differ | Not inherent to the formulation | Not required for the selected P3 inputs |
| Residual initialization | No explicit identity constraint in the generic formulation | Architecture dependent | Zero-initialized residual correction (); identity shortcut in the final channel-compatible P3 configuration |
| Prediction hierarchy | Architecture dependent | Architecture dependent | Original P3/P4/P5 hierarchy retained |
| Design objective | Cross-layer or multi-scale feature aggregation | Adaptive feature emphasis | Controlled fine-scale P3 enhancement |
Table 3.
Theoretical comparison of the existing YOLOv8n prediction levels for tiny-object representation.
Table 3.
Theoretical comparison of the existing YOLOv8n prediction levels for tiny-object representation.
| Level | Stride | Map at 1024 | Relative Object Support | Role in RIF-YOLO-N |
|---|
| P3 | 8 | | 1 | Enhanced for fine-scale localization |
| P4 | 16 | | of P3 | Retained for intermediate representation |
| P5 | 32 | | of P3 | Retained for high-level semantic context |
Table 4.
Summary of object detection datasets used for performance evaluation.
Table 4.
Summary of object detection datasets used for performance evaluation.
| Dataset | Domain | Classes | Used Splits | Primary Evaluation Size | Main Characteristics |
|---|
| VisDrone [28] | Drone-based urban aerial scenes | 10 | Official train and validation splits | 548 images, 38,759 instances | Dense scenes, many tiny and small objects, strong scale variation, occlusion, and background clutter. |
| UAVDT [29] | UAV-based vehicle detection scenes | 3 | Train, validation, and test splits | 271 images, 7046 instances | Vehicle-focused aerial scenes, strong class imbalance, moving-camera viewpoints, and moderate variation in object density. |
| VEDAI [30] | Overhead aerial vehicle imagery | 9 | Train, validation, and test splits | 121 test images, 365 instances | High-resolution aerial imagery, multiple vehicle categories, small object instances, and substantial variation in object appearance and orientation. |
Table 5.
Hardware and software environment used for model training and validation.
Table 5.
Hardware and software environment used for model training and validation.
| Component | Specification |
|---|
| Operating system | Windows 10.0.26200 |
| Python version | Python 3.10.20 |
| Deep learning framework | PyTorch 2.11.0+cu128 |
| Detection framework | Ultralytics YOLO 8.4.78 |
| CUDA support | CUDA enabled through PyTorch cu128 build |
| GPU | NVIDIA RTX 2000 Ada Generation Laptop GPU, 8188 MiB memory |
| CPU | Intel Core Ultra 7 155H, 22 logical CPUs |
| RAM | 31.51 GB |
| Storage | Local SSD storage |
| Main libraries | PyTorch, Ultralytics, OpenCV, NumPy, Pandas, Matplotlib 3.11.2 |
| Primary random seed | 42 |
| Deterministic training | Enabled where supported by the framework |
Table 6.
Training, fine-tuning, and validation configuration used for reproducibility.
Table 6.
Training, fine-tuning, and validation configuration used for reproducibility.
| Setting | Value |
|---|
| Baseline detector | YOLOv8n |
| Proposed detector | RIF-YOLO-N |
| Pretrained weights | yolov8n.pt for YOLOv8n-compatible layers |
| Base training resolution | |
| Base training epochs | 100 epochs |
| Resolution-guided fine-tuning | Fine-tuning from the best base-stage checkpoint at a higher resolution |
| RGFT target resolution | for VisDrone and UAVDT; resolution-matched
and configurations for the VEDAI resolution analysis, with used for the final common-resolution comparison. |
| Fine-tuning epochs | 50 epochs for resolution-guided fine-tuning |
| Validation resolutions | , , , , and , depending on the experiment |
| Batch size | Selected according to GPU memory while keeping the baseline and proposed model comparison matched within each dataset |
| Device | CUDA GPU, device 0 |
| Workers | 2 workers for the final controlled experiments unless otherwise stated |
| Random seed | 42 for the reported deterministic runs |
| Additional repeatability seeds | , to be used only when repeated-run results are reported |
| Optimizer and scheduler | Ultralytics YOLO default optimizer and learning-rate schedule under version 8.4.78 |
| Data augmentation | Ultralytics YOLO default augmentation policy under version 8.4.78 |
| Checkpoint selection | Best checkpoint selected according to validation performance during training |
| Label format | YOLO normalized bounding-box format |
| Metrics | Precision, recall, , , parameters, GFLOPs, and inference time |
Table 7.
Performance of the proposed RIF-YOLO-N with RGFT at .
Table 7.
Performance of the proposed RIF-YOLO-N with RGFT at .
| Dataset | Evaluation Split | P | R | | |
|---|
| VisDrone | Validation | 0.520 | 0.424 | 0.415 | 0.243 |
| UAVDT | Validation | 0.832 | 0.838 | 0.862 | 0.543 |
| VEDAI | Held-out test | 0.799 | 0.586 | 0.669 | 0.419 |
Table 8.
Three-seed robustness evaluation of the proposed RIF-YOLO-N framework. Values are reported as mean ± sample standard deviation across seeds .
Table 8.
Three-seed robustness evaluation of the proposed RIF-YOLO-N framework. Values are reported as mean ± sample standard deviation across seeds .
| Dataset | Precision | Recall | | |
|---|
| VisDrone | | | | |
| UAVDT | | | | |
| VEDAI | | | | |
Table 9.
Matched three-seed reproducibility analysis on VisDrone. Values are reported as mean ± sample standard deviation across seeds .
Table 9.
Matched three-seed reproducibility analysis on VisDrone. Values are reported as mean ± sample standard deviation across seeds .
| Configuration | Precision | Recall | | |
|---|
| YOLOv8n | | | | |
| RIF-YOLO-N w/o RGFT | | | | |
| RIF-YOLO-N + RGFT | | | | |
Table 10.
Latency is measured at using the same controlled batch-size-one runtime protocol for both models.
Table 10.
Latency is measured at using the same controlled batch-size-one runtime protocol for both models.
| Model | Params | GFLOPs | Inf. Time | Main Design |
|---|
| YOLOv8n | 3.008M | 8.1 | ms | Standard anchor-free P3/P4/P5 detector |
| RIF-YOLO-N | 3.061M | 8.8 | ms | Identity-safe P3 enhancement + RGFT |
| +0.053M | +0.7 | +0.763 ms | – |
Table 11.
Controlled resolution-sensitivity comparison on the VisDrone validation set.
Table 11.
Controlled resolution-sensitivity comparison on the VisDrone validation set.
| Model | Resolution | P | R | | |
|---|
| YOLOv8n | | 0.442 | 0.350 | 0.326 | 0.184 |
| YOLOv8n | | 0.479 | 0.379 | 0.362 | 0.207 |
| YOLOv8n | | 0.490 | 0.387 | 0.374 | 0.216 |
| YOLOv8n | | 0.497 | 0.408 | 0.389 | 0.226 |
| YOLOv8n | | 0.507 | 0.411 | 0.401 | 0.235 |
| YOLOv8n + RGFT | | 0.511 | 0.415 | 0.405 | 0.238 |
| RIF-YOLO-N w/o RGFT | | 0.437 | 0.356 | 0.330 | 0.186 |
| RIF-YOLO-N w/o RGFT | | 0.481 | 0.381 | 0.363 | 0.207 |
| RIF-YOLO-N w/o RGFT | | 0.494 | 0.388 | 0.377 | 0.215 |
| RIF-YOLO-N w/o RGFT | | 0.491 | 0.409 | 0.388 | 0.223 |
| RIF-YOLO-N w/o RGFT | | 0.496 | 0.422 | 0.402 | 0.232 |
| RIF-YOLO-N + RGFT | | 0.464 | 0.349 | 0.328 | 0.185 |
| RIF-YOLO-N + RGFT | | 0.485 | 0.382 | 0.367 | 0.212 |
| RIF-YOLO-N + RGFT | | 0.505 | 0.391 | 0.381 | 0.221 |
| RIF-YOLO-N + RGFT | | 0.517 | 0.407 | 0.401 | 0.234 |
| RIF-YOLO-N + RGFT | | 0.520 | 0.424 | 0.415 | 0.243 |
Table 12.
Controlled comparison on the UAVDT validation set.
Table 12.
Controlled comparison on the UAVDT validation set.
| Model | Input Resolution | P | R | | |
|---|
| YOLOv8n | | 0.871 | 0.791 | 0.853 | 0.523 |
| YOLOv8n | | 0.850 | 0.811 | 0.844 | 0.527 |
| YOLOv8n | | 0.845 | 0.814 | 0.833 | 0.519 |
| RIF-YOLO-N w/o RGFT | | 0.844 | 0.775 | 0.850 | 0.500 |
| RIF-YOLO-N w/o RGFT | | 0.840 | 0.810 | 0.826 | 0.492 |
| RIF-YOLO-N w/o RGFT | | 0.812 | 0.820 | 0.814 | 0.471 |
| RIF-YOLO-N + RGFT | | 0.882 | 0.814 | 0.870 | 0.513 |
| RIF-YOLO-N + RGFT | | 0.848 | 0.841 | 0.865 | 0.534 |
| RIF-YOLO-N + RGFT | | 0.832 | 0.838 | 0.862 | 0.543 |
Table 13.
Controlled VEDAI comparison at , showing the mAP gains of RIF-YOLO-N + RGFT over YOLOv8n.
Table 13.
Controlled VEDAI comparison at , showing the mAP gains of RIF-YOLO-N + RGFT over YOLOv8n.
| Model | Input Resolution | P | R | | |
|---|
| YOLOv8n | | 0.668 | 0.499 | 0.558 | 0.356 |
| YOLOv8n | | 0.582 | 0.525 | 0.573 | 0.355 |
| YOLOv8n | | 0.387 | 0.479 | 0.409 | 0.252 |
| RIF-YOLO-N w/o RGFT | | 0.628 | 0.684 | 0.711 | 0.379 |
| RIF-YOLO-N w/o RGFT | | 0.752 | 0.484 | 0.633 | 0.359 |
| RIF-YOLO-N w/o RGFT | | 0.640 | 0.391 | 0.414 | 0.232 |
| RIF-YOLO-N + RGFT | | 0.394 | 0.370 | 0.413 | 0.201 |
| RIF-YOLO-N + RGFT | | 0.788 | 0.558 | 0.716 | 0.427 |
| RIF-YOLO-N + RGFT | | 0.799 | 0.586 | 0.669 | 0.419 |
Table 14.
Structural comparison of representative aerial small-object detectors. The notation P2–P5 indicates the feature levels or prediction scales emphasized by each method when reported in the corresponding source.
Table 14.
Structural comparison of representative aerial small-object detectors. The notation P2–P5 indicates the feature levels or prediction scales emphasized by each method when reported in the corresponding source.
| Method | Detector Family | Main Feature Levels/Strategy | P2 Use | Dataset/Split | Main Design Characteristic |
|---|
| LEAF-YOLO-N [31] | YOLOv7-tiny style | P2/P3/P4/P5 | Yes | VisDrone2019-DET-val | Lightweight edge-oriented model using multi-scale high-resolution detection for small aerial objects. |
| LEAF-YOLO [31] | YOLOv7-tiny style | P2/P3/P4/P5 | Yes | VisDrone2019-DET-val | Larger LEAF variant with stronger multi-scale feature aggregation and higher accuracy. |
| YOLC [32] | CenterNet-based | Cluster-region zooming + CenterNet head | Not YOLO-scale based | VisDrone2019-DET-val | Detects object clusters, zooms into dense regions, and refines tiny-object localization. |
| TPH-YOLOv5 [33] | YOLOv5-based | P3/P4/P5 + extra small-object head | Yes | VisDrone2021 | Adds a transformer prediction head and attention modules for drone-captured dense scenes. |
| EdgeYOLO [34] | YOLO-based | Anchor-free multi-scale YOLO head | Not explicitly fixed in this table | VisDrone2019-DET | Edge-real-time anchor-free detector with lightweight decoupled head and small-object-oriented loss design. |
| RIF-YOLO-N | YOLOv8n-based | P3/P4/P5 with enhanced P3 | No | VisDrone validation | Identity-safe same-scale P3 enhancement with resolution-guided fine-tuning; no additional detection head. |
Table 15.
Reported complexity and inference information of representative aerial object detectors. The values are taken from the corresponding studies and are provided for contextual comparison only, as the input resolutions, hardware platforms, and timing protocols do not match.
Table 15.
Reported complexity and inference information of representative aerial object detectors. The values are taken from the corresponding studies and are provided for contextual comparison only, as the input resolutions, hardware platforms, and timing protocols do not match.
| Method | Params | GFLOPs | Reported Inference/Speed |
|---|
| LEAF-YOLO-N [31] | 1.20M | 5.6 | 16.2 ms |
| LEAF-YOLO [31] | 4.28M | 20.9 | 21.7 ms |
| YOLC [32] | NA | NA | 441 ms |
| TPH-YOLOv5 [33] | NA | 315.4 | 7.36 FPS |
| EdgeYOLO-T [34] | 5.50M | 27.24 | 29.93 ms |
| RIF-YOLO-N (ours) | 3.061M | 8.8 | 19.035 ms |
Table 16.
Literature-reported detection performance of representative object detectors across VisDrone, UAVDT, and VEDAI. AP denotes and denotes mAP at IoU 0.5. NA indicates that the corresponding result was not available in the cited source. The results are not protocol matched because the methods use different input resolutions, dataset splits, training schedules, augmentation strategies, and evaluation settings; therefore, the values are provided only for contextual positioning and no direct superiority claim is inferred from this table.
Table 16.
Literature-reported detection performance of representative object detectors across VisDrone, UAVDT, and VEDAI. AP denotes and denotes mAP at IoU 0.5. NA indicates that the corresponding result was not available in the cited source. The results are not protocol matched because the methods use different input resolutions, dataset splits, training schedules, augmentation strategies, and evaluation settings; therefore, the values are provided only for contextual positioning and no direct superiority claim is inferred from this table.
| Method | VisDrone | UAVDT | VEDAI |
|---|
| Input | AP (%) | (%) | Input | AP (%) | (%) | Input | AP (%) | (%) |
|---|
| LEAF-YOLO-N [31] | 640 | 21.9 | 39.7 | NA | NA | NA | NA | NA | NA |
| LEAF-YOLO [31] | 640 | 28.2 | 48.3 | NA | NA | NA | NA | NA | NA |
| YOLC [32] | 1024 | 31.8 | 55.0 | 1024 | 19.3 | 30.9 | NA | NA | NA |
| TPH-YOLOv5 [33] | 1536 | 34.0 | 53.2 | 1024 | 26.9 | 41.3 | NA | NA | NA |
| EdgeYOLO-T [34] | 640 | 21.8 | 38.5 | NA | NA | NA | NA | NA | NA |
| RIF-YOLO-N + RGFT (ours) | 1024 | 24.3 | 41.5 | 1024 | 54.3 | 86.2 | 1024 | 41.9 | 66.9 |
Table 17.
Computational-efficiency context combining the controlled YOLOv8n–RIF-YOLO-N benchmark with literature-reported values for representative tiny- and small-object detectors. External values are reproduced from the corresponding studies and are not protocol-matched because input resolution, hardware platform, batch size, and runtime measurement procedures may differ. Only YOLOv8n and RIF-YOLO-N are measured under the same controlled protocol and are therefore used for direct latency and throughput comparison.
Table 17.
Computational-efficiency context combining the controlled YOLOv8n–RIF-YOLO-N benchmark with literature-reported values for representative tiny- and small-object detectors. External values are reproduced from the corresponding studies and are not protocol-matched because input resolution, hardware platform, batch size, and runtime measurement procedures may differ. Only YOLOv8n and RIF-YOLO-N are measured under the same controlled protocol and are therefore used for direct latency and throughput comparison.
| Method | Input | Params | GFLOPs | Latency (ms) | |
|---|
| LEAF-YOLO-N [31] | 640 | 1.20M | 5.6 | 16.2 | 56 |
| LEAF-YOLO [31] | 640 | 4.28M | 20.9 | 21 | 32 |
| EdgeYOLO-T [34] | 640 | 5.50M | 27.24 | – | 27 |
| YOLC [32] | 1024 | – | 151.0 | 441 | – |
| TPH-YOLOv5 [33] | 1536 | – | 315.4 | – | 7.36 |
| YOLOv8n | 1024 | 3.008M | 8.1 | | 54.73 |
| RIF-YOLO-N | 1024 | 3.061M | 8.8 | | 52.54 |
Table 18.
Object-size distribution in the VisDrone validation set.
Table 18.
Object-size distribution in the VisDrone validation set.
| Object Size | Area Range | Instances | Ratio |
|---|
| Tiny | | 25,967 | 67.00% |
| Small | | 11,894 | 30.69% |
| Medium | | 871 | 2.25% |
| Large | | 27 | 0.07% |
Table 19.
Size-stratified comparison on the VisDrone validation set.
Table 19.
Size-stratified comparison on the VisDrone validation set.
| Object Size | Instances | YOLOv8n | RIF-YOLO-N | Gain |
|---|
| Tiny | 25,967 | 0.1298 | 0.1346 | +0.0048 |
| Small | 11,894 | 0.3771 | 0.3852 | +0.0081 |
Table 20.
Class-wise AP comparison between YOLOv8n and RIF-YOLO-N on the VisDrone validation set at .
Table 20.
Class-wise AP comparison between YOLOv8n and RIF-YOLO-N on the VisDrone validation set at .
| Class | Instances | YOLOv8n | RIF-YOLO-N | | YOLOv8n | RIF-YOLO-N | |
|---|
| Pedestrian | 8844 | 0.474 | 0.493 | +0.019 | 0.212 | 0.223 | +0.011 |
| People | 5125 | 0.334 | 0.361 | +0.027 | 0.121 | 0.134 | +0.013 |
| Bicycle | 1287 | 0.124 | 0.155 | +0.031 | 0.051 | 0.066 | +0.014 |
| Car | 14,064 | 0.823 | 0.829 | +0.006 | 0.568 | 0.575 | +0.007 |
| Van | 1975 | 0.454 | 0.486 | +0.032 | 0.318 | 0.342 | +0.025 |
| Truck | 750 | 0.342 | 0.372 | +0.031 | 0.222 | 0.248 | +0.026 |
| Tricycle | 1045 | 0.277 | 0.265 | −0.012 | 0.155 | 0.149 | −0.006 |
| Awning-tricycle | 532 | 0.156 | 0.137 | −0.019 | 0.101 | 0.086 | −0.015 |
| Bus | 251 | 0.552 | 0.559 | +0.008 | 0.395 | 0.396 | +0.001 |
| Motor | 4886 | 0.477 | 0.492 | +0.015 | 0.203 | 0.216 | +0.012 |
Table 21.
Detection diagnostic outputs used for qualitative and error-pattern analysis.
Table 21.
Detection diagnostic outputs used for qualitative and error-pattern analysis.
| Diagnostic Output | Purpose | Relevance to This Study |
|---|
| Precision–recall curve | Measures the relation between precision and recall across confidence thresholds. | Indicates whether the final detector preserves detection reliability while improving object recovery. |
| F1-confidence curve | Shows the confidence region where precision and recall are balanced. | Helps identify whether the model maintains stable confidence behavior under dense tiny-object conditions. |
| Normalized confusion matrix | Reports class-level prediction patterns and background confusion. | Explains remaining errors for visually similar or low-frequency classes such as tricycle and awning-tricycle. |
| Qualitative prediction samples | Shows predicted bounding boxes on validation images. | Verifies whether numerical improvements correspond to visible recovery of small and dense objects. |
Table 22.
Ablation study of RIF-YOLO-N on the VisDrone validation set at . Each architecture-level variant removes only the component identified in its name while keeping the remaining RIF components unchanged.
Table 22.
Ablation study of RIF-YOLO-N on the VisDrone validation set at . Each architecture-level variant removes only the component identified in its name while keeping the remaining RIF components unchanged.
| Variant | Params | GFLOPs | P | R | | |
|---|
| YOLOv8n baseline | 3.008M | 8.1 | 0.507 | 0.411 | 0.401 | 0.235 |
| Complete RIF w/o RGFT | 3.061M | 8.8 | 0.496 | 0.422 | 0.402 | 0.232 |
| RIF w/o projection | 3.058M | 8.8 | 0.503 | 0.411 | 0.396 | 0.228 |
| RIF w/o gate | 3.026M | 8.4 | 0.496 | 0.409 | 0.394 | 0.228 |
| RIF w/o refinement | 3.029M | 8.4 | 0.487 | 0.416 | 0.391 | 0.226 |
| RIF w/o residual scale | 3.034M | 8.5 | 0.485 | 0.416 | 0.394 | 0.227 |
| RIF + RGFT (final) | 3.061M | 8.8 | 0.520 | 0.424 | 0.415 | 0.243 |
Table 23.
Design-space comparison of the architecture variants considered during RIF-YOLO-N development.
Table 23.
Design-space comparison of the architecture variants considered during RIF-YOLO-N development.
| Architecture Variant | Detection Scales | P2 Use | P5 Retained | Enhancement Strategy | Design Purpose |
|---|
| RIF-YOLO-N final | P3/P4/P5 | No | Yes | Same-scale identity-safe P3 enhancement | Preserve the lightweight YOLOv8n hierarchy while improving the P3 representation. |
| P2/P3/P4 without P5 | P2/P3/P4 | Yes | No | High-resolution detection without the P5 semantic branch | Test whether adding high-resolution detection while removing high-level semantics benefits tiny objects. |
| Heavy LEAF-style V3 | P2/P3/P4/P5 | Yes | Yes | High-resolution context and stronger feature aggregation | Test whether heavier high-resolution processing improves AP enough to justify the added complexity. |
| V3-Efficient | P3/P4/P5 | Internal P2 guidance | Yes | Lightweight P2-guided enhancement | Test whether P2 guidance can improve accuracy without a large computational increase. |
Table 24.
Architecture selection results on the VisDrone validation set at .
Table 24.
Architecture selection results on the VisDrone validation set at .
| Architecture Variant | Params | GFLOPs | P | R | | | Selection Outcome |
|---|
| RIF-YOLO-N final | 3.061M | 8.8 | 0.520 | 0.424 | 0.415 | 0.243 | Selected because it provides the best accuracy–complexity balance. |
| P2/P3/P4 without P5 | 2.099M | 12.3 | 0.463 | 0.386 | 0.360 | 0.209 | Not selected because removing P5 reduced semantic context and AP. |
| Heavy LEAF-style V3 | 5.159M | 38.6 | 0.510 | 0.410 | 0.404 | 0.236 | Not selected because stronger high-resolution and context operations increased complexity but did not exceed the final RIF-YOLO-N result. |
| V3-Efficient | 3.064M | 8.4 | 0.472 | 0.408 | 0.383 | 0.221 | Not selected because lightweight P2-guided enhancement produced lower AP. |
Table 25.
Relative difference of alternative architectures compared with the final RIF-YOLO-N. Negative AP values indicate performance loss relative to the final model.
Table 25.
Relative difference of alternative architectures compared with the final RIF-YOLO-N. Negative AP values indicate performance loss relative to the final model.
| Architecture Variant | ΔParams | ΔGFLOPs | ΔP | ΔR | | |
|---|
| P2/P3/P4 without P5 | −0.962M | +3.5 | −0.057 | −0.038 | −0.055 | −0.034 |
| Heavy LEAF-style V3 | +2.098M | +29.8 | −0.010 | −0.014 | −0.011 | −0.007 |
| V3-Efficient | +0.003M | −0.4 | −0.048 | −0.016 | −0.032 | −0.022 |