DUST-YOLO: A Deployable UAV Swin Transformer YOLO with Multi-Dimensional Pruning and Mixed-Precision Quantization for End-to-End Video Object Detection
Abstract
1. Introduction
- A multi-dimensional structured pruning strategy for lightweight UAV detection: We design an asymmetric channel pruning scheme for deep convolutional layers and feature-fusion modules to remove redundant channels while preserving the multi-scale feature extraction capability required for aerial small-object detection. In addition, we compress the Swin Transformer prediction heads and reduce the number of bottleneck stacks, substantially decreasing model parameters and computational complexity while incurring only a limited accuracy loss.
- A hardware-aware mixed-precision QAT framework for hybrid CNN–Transformer detectors: We introduce quantization-aware fine-tuning by dynamically inserting fake-quantization nodes during training, enabling the model to adapt to quantization noise before deployment. Furthermore, we map computation-intensive backbone layers to INT8 while retaining Transformer-related modules in FP16, thereby improving inference efficiency on edge hardware while maintaining operator fusion compatibility and detection accuracy.
- An efficient DeepStream-integrated end-to-end deployment system on Jetson Orin NX: We perform TensorRT explicit quantization compilation to precompute weights, fuse operators across layers, and exploit high-throughput INT8 kernels for convolutional computation. The resulting engine is integrated into an asynchronous DeepStream video pipeline that leverages hardware units such as NVDEC and VIC to reduce decoding, preprocessing, and memory-transfer overheads, significantly improving end-to-end throughput and power efficiency.
2. Related Work
2.1. Lightweight Optimization and Model Pruning
2.2. Quantization Strategies and Transformer Sensitivity
2.3. Edge Deployment and Acceleration
3. System Overview
3.1. Network Architecture of DUST-YOLO
3.2. System Module Design
- The first module: Algorithmic Compression Layer. This layer performs deployment-oriented model lightweighting on the server side through multi-dimensional structured pruning and mixed-precision QAT. The pruning strategy is customized for the heterogeneous structure of DUST-YOLO: redundant channels in convolutional layers and feature-fusion modules are removed, bottleneck stacks in deep C3 modules are simplified, and Transformer-related branches in C3STR blocks are compressed under scale-aware constraints. The pruning process further incorporates Jetson-oriented channel alignment and preserves the coupling among the embedding dimension, the number of attention heads, and the per-head dimension in Transformer-related modules. After accuracy recovery through fine-tuning, QAT is applied to make the model adapt to quantization noise before edge-side compilation.
- The second module: Hardware Compilation Layer. This layer converts the optimized model from the ONNX intermediate format into a hardware-specific TensorRT inference engine. During compilation, convolution-dominated paths are mapped to INT8 execution where appropriate, while Transformer-related attention subgraphs are preserved in higher precision to reduce quantization sensitivity and maintain operator-fusion compatibility. This precision allocation follows the numerical sensitivity of heterogeneous CNN–Transformer modules and the fusion behavior of TensorRT. TensorRT further performs graph optimization and operator fusion to reduce fragmented kernel execution and unnecessary memory access on the Jetson platform.
- The third module: System Deployment Layer. This layer integrates the compiled TensorRT engine into an end-to-end DeepStream video analytics pipeline. The input UAV video stream is decoded by dedicated hardware units, such as NVDEC, while resizing, color-space conversion, stream batching, inference, and post-processing are handled by DeepStream plugins. This asynchronous heterogeneous pipeline reduces non-inference overheads during neural network execution, including decoding, preprocessing, memory transfer, and post-processing. By coordinating decoding, preprocessing, inference, and output rendering in this pipeline, the system reduces CPU–GPU communication overhead and better utilizes the heterogeneous computing resources of the edge platform.
4. Methodology
4.1. Structured Pruning and Model Compression Strategy
4.1.1. Asymmetric Channel Pruning and Hardware Alignment for Convolutional Layers
4.1.2. Pruning and Parameter Compression of Transformer Modules
4.1.3. Depth Simplification of Feature Cascades
4.1.4. Progressive Pruning Schedule and Target Sparsity
4.2. Mixed-Precision QAT and Edge-Side Mixed-Precision Quantized Deployment
4.2.1. QAT and Attention Fusion Protection
4.2.2. TensorRT Explicit Quantization Compilation and Continuous-Domain Optimization
4.3. DeepStream-Based Multi-Object Detection System Deployment
4.3.1. System Architecture Design
4.3.2. Hardware Acceleration and Parallel Mechanisms
- Heterogeneous Computing Decoupling: As shown in Figure 8, the front-end encoded video is decoded into raw frames using the dedicated hardware decoder (NVDEC). Subsequently, the Video Image Compositor (VIC) performs scaling and color-space conversion in VRAM, outputting RGB tensors for inference. By offloading these compute-intensive preprocessing tasks to specific hardware units, data flow remains strictly resident in the VRAM, entirely circumventing the memory copy overhead typical of traditional architectures.
- Tensor Batching: To process high-resolution UAV streams, the system utilizes the nvstreammux plugin to aggregate discrete frames into a four-dimensional inference tensor (Figure 9). This batching strategy maximizes the occupancy of GPU parallel compute units. By reducing the CUDA Kernel launch frequency from N to 1, system-level scheduling latency is significantly mitigated, guaranteeing real-time inference under high throughput.
- Highly Concurrent Pipeline: Leveraging an asynchronous pipeline architecture, the system achieves four-stage parallelism (decoding, preprocessing, inference, and post-processing). Specifically, the implementation of CUDA Streams enables the overlapping execution of the current frame’s inference with the previous frame’s post-processing. This design effectively masks I/O wait times, maximizing overall system throughput.
5. Experimental Results
5.1. Experimental Settings
- Hardware and Software Environments: The experimental evaluation of the proposed framework was conducted across two distinct environments to simulate both high-performance training and resource-constrained edge deployment. The server-side operations, encompassing model training, structured pruning, fine-tuning, and QAT, were performed on a workstation equipped with an AMD Ryzen 9 9950X CPU (Advanced Micro Devices, Inc., Santa Clara, CA, USA) and an NVIDIA GeForce RTX 4090 D GPU (NVIDIA Corporation, Santa Clara, CA, USA; 24 GB VRAM), running on Ubuntu 22.04.5 LTS (Canonical Ltd., London, UK) with CUDA 12.4 (NVIDIA Corporation, Santa Clara, CA, USA), Python 3.8.18 (Python Software Foundation, Wilmington, DE, USA), and PyTorch 2.4.1 (Meta Platforms, Inc., Menlo Park, CA, USA). For edge-side deployment and inference benchmarking, we utilized the NVIDIA Jetson Orin NX 16GB platform (NVIDIA Corporation, Santa Clara, CA, USA), which features an 8-core Arm Cortex-A78AE CPU (Arm Ltd., Cambridge, UK) and an Ampere-architecture GPU. This edge platform operated under a software environment comprising CUDA 12.6 (NVIDIA Corporation, Santa Clara, CA, USA), Python 3.10.18 (Python Software Foundation, Wilmington, DE, USA), PyTorch 2.6.0 (Meta Platforms, Inc., Menlo Park, CA, USA), TensorRT 10.3.0 (NVIDIA Corporation, Santa Clara, CA, USA), and NVIDIA DeepStream SDK 7.1 (NVIDIA Corporation, Santa Clara, CA, USA).
- Dataset Justification: To validate the effectiveness of the DUST-YOLO in UAV vision, the VisDrone2019-DET dataset and the UAVDT dataset were employed for training and precision assessment [33,34]. The VisDrone dataset is characterized as a standard and challenging benchmark, containing 10 object categories: Pedestrian, Motor, Van, Person, Truck, Car, Bus, Awning, Bicycle, and Tricycle. These images include various angles and heights, captured both during the day and at night, in different resolutions. The VisDrone2019-DET dataset is divided into three parts, consisting of 6471 training images, 548 validation images, and 1610 test images. The UAVDT dataset is collected by UAVs across a range of complex environments. It comprises images recorded in different cities, under varying scenes and altitudes, with annotations for Cars, Trucks, and Buses, containing 10 h of raw video data with approximately 80,000 video frames. For our experiments, considering the high inter-frame redundancy in the original UAVDT videos, where object locations change only slightly between adjacent frames, we uniformly subsampled all sequences at a 1/4 frame rate by extracting one frame every four frames. The resulting UAVDT subset was then partitioned into a training set and a validation set at an approximate 6:4 ratio, finally yielding 6046 training images and 4155 validation images.
- Training Details: All models, including the proposed DUST-YOLO and the selected comparative baselines, were implemented and trained under identical configurations to ensure a fair comparison. Specifically, all algorithms used a standardized input resolution of pixels and were optimized using the Adam optimizer for 300 epochs with a batch size of 4. All models were evaluated under the Jetson Orin NX platform, and inference protocol, and were converted to ONNX format and compiled into TensorRT inference engines to ensure consistent edge-side evaluation conditions.
- Pruning and Fine-Tuning Details: The structured pruning process was performed in three progressive stages, and each pruning stage was followed by an independent fine-tuning process. The best checkpoint obtained after each fine-tuning stage was retained, and the best checkpoints from the first two stages were used as the initialization for the subsequent pruning stages. For all three post-pruning fine-tuning stages, the pruned model was fine-tuned on the VisDrone training set for 160 epochs with an input resolution of and a batch size of 16. Adam was used as the optimizer. The initial learning rate during fine-tuning was set to 0.1 times the initial learning rate used in baseline training and was decayed using a one-cycle cosine learning-rate schedule. The regularization settings followed the baseline hyperparameter configuration: weight decay was applied only to ordinary weight parameters, while BatchNorm parameters and bias terms were excluded from weight decay. Label smoothing was set to 0.0. Exponential moving average (EMA) was enabled, and the checkpoint with the best validation fitness was retained after each fine-tuning stage.
- Practical UAV-View Edge Evaluation Context and Metrics: To reflect practical UAV-view video analytics on edge platforms, the deployment evaluation was conducted on continuous aerial video streams rather than isolated static images. The Jetson Orin NX was locked in MAXN mode for stable benchmarking. Detection accuracy was evaluated on the VisDrone2019-DET validation set using mAP@0.5 and mAP@0.5:0.95. For efficiency evaluation, pure inference latency was measured using the native TensorRT C++ API to isolate model-level acceleration, while end-to-end latency and throughput were measured over the complete video detection pipeline, including video decoding, preprocessing, inference, and post-processing. The end-to-end test used a 211-s UAV-view aerial video at 30 FPS, totaling more than 6300 frames. All tests were conducted after warm-up and repeated five times, with the average end to end latency results reported. For the comparative baselines in Table 4, the end-to-end metrics were measured using the conventional TensorRT-based video detection pipeline, while the final DUST-YOLO system was evaluated with the proposed TensorRT–DeepStream asynchronous pipeline under the same end-to-end measurement scope.
5.2. Comparative Experiments
5.3. Ablation Study
5.4. Deployment-Oriented Robustness and Failure Analysis
5.4.1. Scale-Wise Accuracy Preservation
5.4.2. Cross-Dataset Generalization Performance
5.4.3. Visual Robustness and Failure Case Analysis
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Echchidmi, M.; Bouayad, A. TinyML for sustainable edge intelligence: Practical optimization under extreme resource constraints. Technologies 2026, 14, 215. [Google Scholar] [CrossRef]
- Shih, W.-C.; Wang, Z.-Y.; Kristiani, E.; Hsieh, Y.-J.; Sung, Y.-H.; Li, C.-H.; Yang, C.-T. The construction of a stream service application with DeepStream and simple realtime server using containerization for edge computing. Sensors 2025, 25, 259. [Google Scholar] [CrossRef]
- Gong, J.; Yuan, Z.; Li, W.; Li, W.; Guo, Y.; Guo, B. A Lightweight Upsampling and Cross-Modal Feature Fusion-Based Algorithm for Small-Object Detection in UAV Imagery. Electronics 2026, 15, 298. [Google Scholar] [CrossRef]
- Jiang, Z.; Li, C.; Qu, T.; He, C.; Wang, D. MSQuant: Efficient post-training quantization for object detection via migration scale search. Electronics 2025, 14, 504. [Google Scholar] [CrossRef]
- Tian, L.; Wang, P. An effective mixed-precision quantization method for joint image deblurring and edge detection. Electronics 2025, 14, 1767. [Google Scholar] [CrossRef]
- Sun, J.; Gao, H.; Yan, Z.; Qi, X.; Yu, J.; Ju, Z. Lightweight UAV Object-Detection Method Based on Efficient Multidimensional Global Feature Adaptive Fusion and Knowledge Distillation. Electronics 2024, 13, 1558. [Google Scholar] [CrossRef]
- Yang, R.; Li, W.; Shang, X.; Zhu, D.; Man, X. KPE-YOLOv5: An improved small target detection algorithm based on YOLOv5. Electronics 2023, 12, 817. [Google Scholar] [CrossRef]
- Wang, K.; Zhou, H.; Wu, H.; Yuan, G. RN-YOLO: A Small Target Detection Model for Aerial Remote-Sensing Images. Electronics 2024, 13, 2383. [Google Scholar] [CrossRef]
- Li, Z.; Xiao, J.; Yang, L.; Gu, Q. RepQ-ViT: Scale Reparameterization for Post-Training Quantization of Vision Transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2–6 October 2023; pp. 17181–17190. [Google Scholar] [CrossRef]
- Zhu, W.; Chen, K. Real-time object detection for unmanned aerial vehicles based on vision transformer and edge computing. Sci. Rep. 2026, 16, 6814. [Google Scholar] [CrossRef] [PubMed]
- Jocher, G. Ultralytics YOLOv5. 2020. Available online: https://github.com/ultralytics/yolov5 (accessed on 15 April 2026).
- Zhu, X.; Lyu, S.; Wang, X.; Zhao, Q. TPH-YOLOv5: Improved YOLOv5 based on transformer prediction head for object detection on drone-captured scenarios. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, QC, Canada, 11–17 October 2021; pp. 2778–2788. [Google Scholar] [CrossRef]
- Zhao, Q.; Liu, B.; Lyu, S.; Wang, C.; Zhang, H. TPH-YOLOv5++: Boosting Object Detection on Drone-Captured Scenarios with Cross-Layer Asymmetric Transformer. Remote Sens. 2023, 15, 1687. [Google Scholar] [CrossRef]
- Shi, H.; Cheng, X.; Mao, W.; Wang, Z. P2-ViT: Power-of-two post-training quantization and acceleration for fully quantized vision transformer. IEEE Trans. Very Large Scale Integr. (VLSI) Syst. 2024, 32, 1704–1717. [Google Scholar] [CrossRef]
- Aljami, H.M.; Alrowais, N.A.; AlAwajy, A.M.; Alhrgan, S.O.; Aldwaani, R.A.; Alsawadi, M.S.; Saqib, N.U.; Alam, S.S.; Alsubaie, R. Benchmarking YOLOv8 Variants for Object Detection Efficiency on Jetson Orin NX for Edge Computing Applications. Computers 2026, 15, 74. [Google Scholar] [CrossRef]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the International Conference on Learning Representations (ICLR), Vienna, Austria, 4 May 2021; Available online: https://openreview.net/forum?id=YicbFdNTTy (accessed on 15 April 2026).
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 11–17 October 2021; pp. 10012–10022. [Google Scholar] [CrossRef]
- Hakani, R.; Rawat, A. Edge computing-driven real-time drone detection using YOLOv9 and NVIDIA Jetson Nano. Drones 2024, 8, 680. [Google Scholar] [CrossRef]
- Xue, C.; Xia, Y.; Wu, M.; Chen, Z.; Cheng, F.; Yun, L. EL-YOLO: An efficient and lightweight low-altitude aerial objects detector for onboard applications. Expert Syst. Appl. 2024, 256, 124848. [Google Scholar] [CrossRef]
- Wu, J.; Meng, H.; Yuan, M.; Liu, C.; Lu, Z. Enhanced feature representation for real time UAV image object detection using contextual information and adaptive fusion. Sci. Rep. 2025, 15, 33711. [Google Scholar] [CrossRef]
- Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; Belongie, S. Feature Pyramid Networks for Object Detection. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 936–944. [Google Scholar] [CrossRef]
- Zhao, Z.; Liu, X.; He, P. PSO-YOLO: A contextual feature enhancement method for small object detection in UAV aerial images. Earth Sci. Inform. 2025, 18, 258. [Google Scholar] [CrossRef]
- Cai, S.; Wu, Z.; Liu, K.; Zhang, T.; Weng, W.; Zheng, X. LSOD-YOLO: A visual object detection method for AGV perception systems based on a lightweight backbone and detection head. Technologies 2026, 14, 173. [Google Scholar] [CrossRef]
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path aggregation network for instance segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 8759–8768. [Google Scholar] [CrossRef]
- Mi, Q.; Chao, J.; Chen, A.; Zhang, K.; Lai, J. YOLO11s-UAV: An Advanced Algorithm for Small Object Detection in UAV Aerial Imagery. J. Imaging 2026, 12, 69. [Google Scholar] [CrossRef]
- Fan, Q.; Li, Y.; Deveci, M.; Zhong, K.; Kadry, S. LUD-YOLO: A novel lightweight object detection network for unmanned aerial vehicle. Inf. Sci. 2025, 686, 121366. [Google Scholar] [CrossRef]
- Huang, M.; Mi, W.; Wang, Y. EDGS-YOLOv8: An improved YOLOv8 lightweight UAV detection model. Drones 2024, 8, 337. [Google Scholar] [CrossRef]
- Xie, S.; Deng, G.; Lin, B.; Jing, W.; Li, Y.; Zhao, X. Real-time object detection from UAV inspection videos by combining YOLOv5s and DeepStream. Sensors 2024, 24, 3862. [Google Scholar] [CrossRef]
- Barthelemy, J.; Iqbal, U.; Qian, Y.; Amirghasemi, M.; Perez, P. Safety after dark: A privacy compliant and real-time edge computing intelligent video analytics for safer public transportation. Sensors 2024, 24, 8102. [Google Scholar] [CrossRef]
- Yue, M.; Zhang, L.; Huang, J.; Zhang, H. Lightweight and efficient tiny-object detection based on improved YOLOv8n for UAV aerial images. Drones 2024, 8, 276. [Google Scholar] [CrossRef]
- Liu, C.; Gao, G.; Huang, Z.; Hu, Z.; Liu, Q.; Wang, Y. YOLC: You only look clusters for tiny object detection in aerial images. IEEE Trans. Intell. Transp. Syst. 2024, 25, 13863–13875. [Google Scholar] [CrossRef]
- Ma, C.; Fu, Y.Y.; Wang, D.; Guo, R.; Zhao, X.; Fang, J. YOLO-UAV: Object detection method of unmanned aerial vehicle imagery based on efficient multi-scale feature fusion. IEEE Access 2023, 11, 126857–126878. [Google Scholar] [CrossRef]
- Cao, Y.; He, Z.; Wang, L.; Wang, W.; Yuan, Y.; Zhang, D.; Zhang, J.; Zhu, P.; Van Gool, L.; Han, J.; et al. VisDrone-DET2021: The vision meets drone object detection challenge results. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, QC, Canada, 11–17 October 2021; pp. 2847–2854. [Google Scholar] [CrossRef]
- Du, D.; Qi, Y.; Yu, H.; Yang, Y.; Duan, K.; Li, G.; Zhang, W.; Huang, Q.; Tian, Q. The Unmanned Aerial Vehicle Benchmark: Object Detection and Tracking. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 370–386. [Google Scholar] [CrossRef]
- Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLOv8. 2023. Available online: https://docs.ultralytics.com/models/yolov8/ (accessed on 15 April 2026).
- Wang, C.Y.; Yeh, I.H.; Liao, H.Y.M. YOLOv9: Learning what you want to learn using programmable gradient information. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 1–21. [Google Scholar] [CrossRef]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. Adv. Neural Inf. Process. Syst. 2024, 37, 107984–108011. [Google Scholar]
- Jocher, G.; Qiu, J. Ultralytics YOLO11. 2024. Available online: https://docs.ultralytics.com/models/yolo11/ (accessed on 15 April 2026).
- Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-Centric Real-Time Object Detectors. Adv. Neural Inf. Process. Syst. 2026, 38, 78433–78457. [Google Scholar]
- Gao, F.; An, J.; Zhang, M.; Chen, X.; Zhao, Q. FE-YOLO: A Traffic Target Detection Network Based on YOLOv11n. In Proceedings of the 2025 44th Chinese Control Conference (CCC), Chongqing, China, 28–30 July 2025; pp. 8845–8850. [Google Scholar] [CrossRef]
- Jiang, P.; Ergu, D.; Liu, F.; Cai, Y.; Ma, B. A review of YOLO algorithm developments. Procedia Comput. Sci. 2022, 199, 1066–1073. [Google Scholar] [CrossRef]
- Terven, J.; Cordova-Esparza, D. A comprehensive review of YOLO architectures in computer vision: From YOLOv1 to YOLOv8 and YOLO-NAS. Mach. Learn. Knowl. Extr. 2023, 5, 1680–1716. [Google Scholar] [CrossRef]












| Research Direction | Representative Works/Tools | Key Observation |
|---|---|---|
| Lightweight Optimization | [19,20,24,25,26] | Improve detection efficiency, but the accuracy–complexity trade-off remains challenging. |
| Quantization Strategies | [27] | Low-bit deployment is effective, but Transformer-related modules remain quantization-sensitive. |
| Deployment Frameworks | ONNX/TensorRT/DeepStream [28,29,30,31,32] | Accelerate inference or video processing, but end-to-end pipeline overhead is still important. |
| DUST-YOLO | Proposed | Combines structured pruning, mixed-precision QAT, TensorRT, and DeepStream for edge deployment. |
| Branch | Feature Size | D | Heads | W-MSA | SW-MSA |
|---|---|---|---|---|---|
| P2/xsmall | 64 | 2 | |||
| P3/small | 128 | 4 | |||
| P4/medium | 256 | 8 | |||
| P5/large | 512 | 16 |
| Stage | Pruned Component | Before Pruning | After Pruning | Target Sparsity |
|---|---|---|---|---|
| Stage I | Selected deep Conv/C3/SPPF modules | 256/512/1024 channels | 128/256/512 channels | 50% channel pruning |
| Stage I | C3STR branch P4 | output 512, , | output 256, , | 50% channel/embedding pruning |
| Stage I | C3STR branch P5 | output 1024, , | output 512, , | 50% channel/embedding pruning |
| Stage II | C3STR branch P4 embedding path | , | , | 50% additional embedding pruning |
| Stage II | C3STR branch P5 embedding path | , | , | 50% additional embedding pruning |
| Stage III | C3STR branch P3 embedding path | , | , | 50% embedding pruning |
| Stage III | FFN in C3STR branches P2–P5 | expansion ratio | expansion ratio | 50% hidden-dimension pruning |
| Stage III | Deep C3 bottleneck stack | 9 bottleneck blocks | 6 bottleneck blocks | 33.3% depth pruning |
| Model | Resolution(pixels) | mAP@0.5 (%) | mAP@0.5:0.95 (%) | Inference Time (ms) | End-to-End Latency (ms) | End-to-End FPS | Speedup Ratio |
|---|---|---|---|---|---|---|---|
| YOLOv8s [35] | 38.6 | 23.1 | 33.7 | 94.3 | 10.6 | ||
| YOLOv9s [36] | 39.5 | 23.5 | 26.0 | 87.7 | 11.4 | ||
| YOLOv10s [37] | 38.2 | 22.9 | 26.8 | 86.4 | 11.6 | ||
| YOLOv11s [38] | 38.2 | 22.7 | 21.8 | 84.1 | 11.9 | ||
| YOLOv12s [39] | 38.2 | 22.8 | 27.5 | 86.2 | 11.6 | ||
| YOLOv5l [11] | 40.6 | 24.0 | 40.4 | 104.8 | 9.5 | ||
| YOLOv8l [35] | 43.0 | 26.5 | 52.6 | 113.7 | 8.8 | ||
| YOLOv10l [37] | 42.3 | 25.8 | 50.8 | 108.4 | 9.2 | ||
| FE-YOLO [40] | 34.9 | - | 37.3 | 96.9 | 10.3 | ||
| YOLO-UD-s [20] | 45.5 | - | 24.6 | 84.2 | 11.9 | ||
| DUST-YOLO (Ours) | 43.7 | 25.9 | 18.9 | 36.3 | 27.5 |
| Method | FP16 Quant. | INT8 (QAT) | Structured Pruning | DS Pipeline | mAP@0.5 (%) | mAP @0.5:0.95 (%) | Inference Time (ms) | End-to-End FPS | Speedup Ratio |
|---|---|---|---|---|---|---|---|---|---|
| Baseline | 44.9 | 27.2 | 62.7 | 8.4 | |||||
| Opt. 1 | ✓ | 44.9 | 27.2 | 28.9 | 11.5 | ||||
| Opt. 3 | ✓ | 45.1 | 27.3 | 48.4 | 9.3 | ||||
| Opt. 4 | ✓ | 44.9 | 27.2 | 62.7 | 13.0 | ||||
| Opt. 1+3 | ✓ | ✓ | 45.1 | 27.3 | 23.7 | 12.2 | |||
| Opt. 1+2 | ✓ | ✓ | 43.6 | 25.9 | 23.3 | 12.2 | |||
| Opt. 1+2+3 | ✓ | ✓ | ✓ | 43.7 | 25.9 | 18.9 | 14.1 | ||
| Opt. 1+4 | ✓ | ✓ | 44.9 | 27.2 | 28.9 | 23.2 | |||
| Opt. 3+4 | ✓ | ✓ | 45.1 | 27.3 | 48.4 | 15.1 | |||
| Opt. 1+3+4 | ✓ | ✓ | ✓ | 45.1 | 27.3 | 23.7 | 24.6 | ||
| Opt. 1+2+4 | ✓ | ✓ | ✓ | 43.6 | 25.9 | 23.3 | 26.1 | ||
| Ours (Opt. 1+2+3+4) | ✓ | ✓ | ✓ | ✓ | 43.7 | 25.9 | 18.9 | 27.5 |
| Model | AP (%) | AP50 (%) | AP75 (%) | APs (%) | APm (%) | APl (%) |
|---|---|---|---|---|---|---|
| Baseline | 25.7 | 42.8 | 26.5 | 16.2 | 36.9 | 43.8 |
| DUST-YOLO (Ours) | 24.4 | 41.5 | 24.7 | 15.3 | 35.0 | 43.1 |
| Model | mAP@0.5 (%) | mAP@0.5:0.95 (%) |
|---|---|---|
| Baseline | 33.5 | 18.5 |
| DUST-YOLO (Ours) | 32.7 | 19.2 |
| Model | AP (%) | AP50 (%) | AP75 (%) | APs (%) | APm (%) | APl (%) |
|---|---|---|---|---|---|---|
| Baseline | 16.4 | 32.2 | 14.6 | 12.4 | 25.7 | 20.4 |
| DUST-YOLO (Ours) | 17.3 | 31.6 | 17.2 | 11.8 | 28.5 | 19.0 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Lin, G.; Jiang, J.; Cai, J.; Luo, X.; Wang, Z.; Sun, H.; Pu, Z. DUST-YOLO: A Deployable UAV Swin Transformer YOLO with Multi-Dimensional Pruning and Mixed-Precision Quantization for End-to-End Video Object Detection. Electronics 2026, 15, 2579. https://doi.org/10.3390/electronics15122579
Lin G, Jiang J, Cai J, Luo X, Wang Z, Sun H, Pu Z. DUST-YOLO: A Deployable UAV Swin Transformer YOLO with Multi-Dimensional Pruning and Mixed-Precision Quantization for End-to-End Video Object Detection. Electronics. 2026; 15(12):2579. https://doi.org/10.3390/electronics15122579
Chicago/Turabian StyleLin, Gongxun, Jincheng Jiang, Jiaheng Cai, Xingjian Luo, Zihao Wang, Hao Sun, and Ziyuan Pu. 2026. "DUST-YOLO: A Deployable UAV Swin Transformer YOLO with Multi-Dimensional Pruning and Mixed-Precision Quantization for End-to-End Video Object Detection" Electronics 15, no. 12: 2579. https://doi.org/10.3390/electronics15122579
APA StyleLin, G., Jiang, J., Cai, J., Luo, X., Wang, Z., Sun, H., & Pu, Z. (2026). DUST-YOLO: A Deployable UAV Swin Transformer YOLO with Multi-Dimensional Pruning and Mixed-Precision Quantization for End-to-End Video Object Detection. Electronics, 15(12), 2579. https://doi.org/10.3390/electronics15122579
