PDGV-DETR: Object Detection for Secure On-Site Weapon and Personnel Location Based on Dynamic Convolution and Cross-Scale Semantic Fusion
Abstract
1. Introduction
- (1)
- A bidirectional hybrid feature Pyramid network with channel attention (DWH-FPN) is reconstructed, which realizes the bidirectional interaction between high-level semantic features and low-level detail features through transposed convolution, and combines channel attention to generate dynamic weights. It reduces redundant calculations and strengthens the binding of details and semantics in prohibited weapon detection, the ability to separate foreground and background in personnel object localization, and the ability to integrate global and local in scene understanding, which lays a foundation for multi-scale security-related object detection in security scenes.
- (2)
- Adapt the Dynamic Hierarchical Channel Interaction Convolution Module (BasicPCock) to replace the traditional modules at the end of the backbone network. Through the dual-dimensional collaboration of fine channel convolution in the main path and adaptive channel mapping in the shortcut path, it reduces the computational load while enhancing the robustness of incomplete objects, avoiding missed detections of occluded objects, partial limbs, etc., and improving the integrity of tool and behavior recognition.
- (3)
- The Global Semantic Weaving and Elastic Feature Alignment Network (GSWFN) is introduced and repurposed. Through the cross-scale semantic correlation mechanism, it significantly alleviates the problem of blurred object and background features, enhances the model’s scene understanding ability, effectively reduces object misjudgment and false positives in complex scenarios, and is suitable for various types of security monitoring scenarios.
- (4)
- The systematic verification of multiple datasets and multiple models is carried out. Based on four security object detection datasets with multi-scene personnel and prohibited weapons, under the same experimental configuration, the performance of PDGV-DETR and 13 mainstream general object detection models in security-related object detection and localization tasks is tested, and the defects of the existing models with large accuracy fluctuations across data sets are revealed. The advantages of PDGV-DETR in positioning accuracy, missed detection rate and generalization are verified to support the implementation of the project.
2. Related Work
2.1. Research Status of Object Detection Technology in Security Scenarios
2.1.1. Evolution of Violence Detection Technology
2.1.2. The Core Requirements and Technical Boundaries of Object Detection in Security Scenarios
2.2. Universal Object Detection Technology Has Insufficient Adaptability for Security-Related Scenarios Involving Weapons and Personnel Detection
2.3. The Development of the DETR Model and the Bottleneck in Adapting It to Security-Related Weapon and Personnel Object Detection Scenarios
2.4. Shortcomings of Existing Research and Positioning of This Paper
- (1)
- Insufficient ability of multi-scale feature fusion: the feature fusion of general object detection relies on fixed scale weighting (such as EfficientDet [27]) or simple cross-layer splicing (such as YOLO series [28,29,30,31,32,33]), which cannot take into account the detailed characteristics of small weapons such as knives and pistols and the global characteristics of large personnel objects. It hinders the synchronous and high-precision detection of multi-scale security objects.
- (2)
- Insufficient robustness of incomplete object detection: existing models only rely on a single convolution or sparse attention (such as Deformable-DETR [36]) for feature extraction of occluded and blurred incomplete objects, and lack a hierarchical feature compensation mechanism, which leads to serious missed detection and false alarm of occluded weapons and overlapping personnel objects.
- (3)
- Low object discrimination under complex background: global semantic modeling mostly relies on a single attention mechanism (such as the RT-DETR decoder [40]), which does not combine the collaborative calibration of context and spatial features. In security scenes with dense human flow and cluttered environment, it is difficult to effectively distinguish foreground objects from complex backgrounds.
3. Method
3.1. Backbone Network
3.2. Encoder
3.2.1. Bidirectional Hybrid Feature Pyramid (DWH-FPN)
3.2.2. Global Semantic Weaving and Elastic Feature Alignment Network (GSWFN)
3.3. Decoder
4. Experiments and Datasets
4.1. Experiment Settings
4.2. Data Set
4.3. Evaluation Index
4.4. Contrast Experiment
4.5. Ablation Experiment
4.5.1. Discussion on Module Precision
4.5.2. Cost–Benefit Analysis of Computational Overhead and Demonstration of Deployment Adaptability
Quantitative Analysis of Core Indicators for Deployment of the Abandonment Model
Visualization of Accuracy–Delay Trade-Off and Marginal Benefit Analysis
4.5.3. Robustness Supplementary Verification Experiment
4.6. Architectural Distinctiveness and Complexity Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Rojas-Andrade, R.; Lopez Leiva, V.; Varela, J.J.; Soto García, P.; Álvarez, J.P.; Ramirez, M.T. Feasibility, acceptability, and appropriability of a national whole-school program for reducing school violence and improving school coexistence. Front. Psychol. 2024, 15, 1395990. [Google Scholar] [CrossRef]
- Onat, I.; Bastug, M.F.; Guler, A.; Kula, S. Fears of cyberterrorism, terrorism, and terrorist attacks: An empirical comparison. Behav. Sci. Terror. Polit. Aggress. 2024, 16, 149–165. [Google Scholar] [CrossRef]
- Rajan, S.; Buttar, N.; Ladhani, Z.; Caruso, J.; Allegrante, J.P.; Branas, C.J. School violence exposure as an adverse childhood experience: Protocol for a nationwide study of secondary public schools. JMIR Res. Protoc. 2024, 13, e56249. [Google Scholar] [CrossRef]
- Xu, W.; Zhu, D.; Deng, R.; Yung, K.L.; Ip, A.W.H. Violence-YOLO: Enhanced GELAN algorithm for violence detection. Appl. Sci. 2024, 14, 6712. [Google Scholar] [CrossRef]
- Omarov, B.; Narynov, S.; Zhumanov, Z.; Gumar, A.; Khassanova, M. State-of-the-art violence detection techniques in video surveillance security systems: A systematic review. PeerJ Comput. Sci. 2022, 8, e920. [Google Scholar] [CrossRef] [PubMed]
- Zahra, A.; Ghafoor, M.; Munir, K.; Ullah, A.; Ul Abideen, Z. Application of region-based video surveillance in smart cities using deep learning. Multimed. Tools Appl. 2024, 83, 15313–15338. [Google Scholar] [CrossRef] [PubMed]
- Zhao, X.; Wang, L.; Zhang, Y.; Han, X.; Deveci, M.; Parmar, M. A review of convolutional neural networks in computer vision. Artif. Intell. Rev. 2024, 57, 99. [Google Scholar] [CrossRef]
- You, Y. The impact of deep learning on computer vision: From image classification to scene understanding. Int. J. Sci. Res. Manag. 2024, 12, 5843–5856. [Google Scholar]
- Karim, M.; Khalid, S.; Aleryani, A.; Khan, J.; Ullah, I.; Ali, Z. Human action recognition systems: A review of the trends and state-of-the-art. IEEE Access 2024, 12, 36372–36390. [Google Scholar] [CrossRef]
- LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef]
- Alzubaidi, L.; Zhang, J.; Humaidi, A.J.; Al-Dujaili, A.; Duan, Y.; Al-Shamma, O.; Santamaría, J.; Fadhel, M.A.; Al-Amidie, M.; Farhan, L. Review of deep learning: Concepts, CNN architectures, challenges, applications, future directions. J. Big Data 2021, 8, 53. [Google Scholar] [CrossRef] [PubMed]
- Li, Z.; Liu, F.; Yang, W.; Peng, S.; Zhou, J. A survey of convolutional neural networks: Analysis, applications, and prospects. IEEE Trans. Neural Netw. Learn. Syst. 2021, 32, 6999–7019. [Google Scholar] [CrossRef] [PubMed]
- Negre, P.; Alonso, R.S.; González-Briones, A.; Prieto, J.; Rodríguez-González, S. Literature review of deep-learning-based detection of violence in video. Sensors 2024, 24, 4016. [Google Scholar] [CrossRef] [PubMed]
- Graves, A. Long short-term memory. In Supervised Sequence Labelling with Recurrent Neural Networks; Springer: Berlin/Heidelberg, Germany, 2012; pp. 37–45. [Google Scholar]
- Li, N.; Bai, X.; Shen, X.; Xin, P.; Tian, J.; Chai, T.; Wang, Z. Dense pedestrian detection based on GR-YOLO. Sensors 2024, 24, 4747. [Google Scholar] [CrossRef]
- Zhang, Z.; Xu, S.; Lu, S.; Chen, L. Advancing precision pig behavior recognition through real-time detection transformer. In Proceedings of the 2024 4th International Conference on Neural Networks, Information and Communication Engineering (NNICE), Guangzhou, China, 24–26 May 2024; pp. 707–710. [Google Scholar]
- Santos, F.; Durães, D.; Marcondes, F.S.; Lange, S.; Machado, J.; Novais, P. Efficient violence detection using transfer learning. In Proceedings of the International Conference on Practical Applications of Agents and Multi-Agent Systems, Salamanca, Spain, 6–8 October 2021; pp. 65–75. [Google Scholar]
- Magdy, M.; Fakhr, M.W.; Maghraby, F.A. Violence 4D: Violence detection in surveillance using 4D convolutional neural networks. IET Comput. Vis. 2023, 17, 282–294. [Google Scholar] [CrossRef]
- Liang, Q.; Cheng, C.; Li, Y.; Yang, K.; Chen, B. Fusion and visualization design of violence detection and geographic video. In Proceedings of the 39th National Conference of Theoretical Computer Science (NCTCS 2021), Yinchuan, China, 23–25 July 2021; pp. 33–46. [Google Scholar]
- Qu, W.; Zhu, T.; Liu, J.; Li, J. A time sequence location method of long video violence based on improved C3D network. J. Supercomput. 2022, 78, 19545–19565. [Google Scholar] [CrossRef]
- Aktı, Ş.; Ofli, F.; Imran, M.; Ekenel, H.K. Fight detection from still images in the wild. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 3–8 January 2022; pp. 550–559. [Google Scholar]
- Ehsan, T.Z.; Nahvi, M.; Mohtavipour, S.M. An accurate violence detection framework using unsupervised spatial–temporal action translation network. Vis. Comput. 2024, 40, 1515–1535. [Google Scholar] [CrossRef]
- Kumar, A.; Shetty, A.; Sagar, A.; Charushree, A.; Kanwal, P. Indoor violence detection using lightweight transformer model. In Proceedings of the 2023 4th International Conference for Emerging Technology (INCET), Belgaum, India, 26–28 May 2023; pp. 1–6. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 39, 1137–1149. [Google Scholar] [CrossRef]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2961–2969. [Google Scholar]
- Tian, Z.; Shen, C.; Chen, H.; He, T. FCOS: Fully convolutional one-stage object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9627–9636. [Google Scholar]
- Tan, M.; Pang, R.; Le, Q.V. EfficientDet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 10781–10790. [Google Scholar]
- Jocher, G.; Chaurasia, A.; Stoken, A.; Borovec, J.; Kwon, Y.; Michael, K.; Fang, J.; Yifu, Z.; Wang, C.; Montes, D.; et al. Ultralytics/Yolov5: V7.0-YOLOv5 SOTA Realtime Instance Segmentation; Zenodo: Geneva, Switzerland, 2022. [Google Scholar] [CrossRef]
- Li, C.; Li, L.; Jiang, H.; Weng, K.; Geng, Y.; Li, L.; Ke, Z.; Li, Q.; Cheng, M.; Nie, W.; et al. YOLOv6: A single-stage object detection framework for industrial applications. arXiv 2022, arXiv:2209.02976. [Google Scholar] [CrossRef]
- Terven, J.; Córdova-Esparza, D.-M.; Romero-González, J.-A. A comprehensive review of YOLO architectures in computer vision: From YOLOv1 to YOLOv8 and YOLO-NAS. Mach. Learn. Knowl. Extr. 2023, 5, 1680–1716. [Google Scholar] [CrossRef]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-time end-to-end object detection. arXiv 2024, arXiv:2405.14458. [Google Scholar]
- Khanam, R.; Hussain, M. YOLOv11: An overview of the key architectural enhancements. arXiv 2024, arXiv:2410.17725. [Google Scholar] [CrossRef]
- Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-centric real-time object detectors. arXiv 2025, arXiv:2502.12524. [Google Scholar]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK, 23–28 August 2020; pp. 213–229. [Google Scholar]
- Liu, S.; Li, F.; Zhang, H.; Yang, X.; Qi, X.; Su, H.; Zhu, J.; Zhang, L. DAB-DETR: Dynamic anchor boxes are better queries for DETR. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual Event, 25–29 April 2022. [Google Scholar]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable DETR: Deformable transformers for end-to-end object detection. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual Event, 3–7 May 2021. [Google Scholar]
- Li, F.; Zhang, H.; Liu, S.; Guo, J.; Ni, L.M.; Zhang, L. DN-DETR: Accelerate DETR training by introducing query denoising. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 13619–13627. [Google Scholar]
- Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Ni, L.M.; Shum, H.-Y. DINO: DETR with improved denoising anchor boxes for end-to-end object detection. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual Event, 25–29 April 2022. [Google Scholar]
- Jia, D.; Yuan, Y.; He, H.; Wu, X.; Yu, H.; Lin, W.; Sun, L.; Zhang, C.; Hu, H. DETRs with hybrid matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 19702–19712. [Google Scholar]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 16965–16974. [Google Scholar]
- Chen, Y.; Zhang, C.; Chen, B.; Huang, Y.; Sun, Y.; Wang, C.; Fu, X.; Dai, Y.; Qin, F.; Peng, Y.; et al. Accurate leukocyte detection based on deformable-DETR and multi-level feature fusion for aiding diagnosis of blood diseases. Comput. Biol. Med. 2024, 170, 107917. [Google Scholar] [CrossRef]
- Li, K.; Geng, Q.; Wan, M.; Cao, X.; Zhou, Z. Context and spatial feature calibration for real-time semantic segmentation. IEEE Trans. Image Process. 2023, 32, 4078–4091. [Google Scholar] [CrossRef]
- Zhang, P. Violence-Image-Dataset. GitHub Repository. 2023. Available online: https://github.com/ChinaZhangPeng/Violence-Image-Dataset (accessed on 22 February 2026).
- Olmos, R.; Tabik, S.; Herrera, F. Automatic handgun detection alarm in videos using deep learning. Neurocomputing 2018, 275, 66–72. [Google Scholar] [CrossRef]
- Castillo, A.; Tabik, S.; Pérez, F.; Olmos, R.; Herrera, F. Brightness guided preprocessing for automatic cold steel weapon detection in surveillance videos with deep learning. Neurocomputing 2019, 330, 151–161. [Google Scholar] [CrossRef]
- School Moslk. People Dataset. Roboflow Universe. 2024. Available online: https://universe.roboflow.com/school-moslk/people-jinpn (accessed on 22 February 2026).












| Method | Violence-Image-Dataset Pu | Pistol | Knife | People Dataset |
|---|---|---|---|---|
| YOLOv5 | 0.802 | 0.844 | 0.773 | 0.784 |
| YOLOv6 | 0.812 | 0.837 | 0.824 | 0.801 |
| YOLOv8m | 0.803 | 0.844 | 0.848 | 0.813 |
| YOLOv10m | 0.789 | 0.849 | 0.733 | 0.813 |
| YOLOv11m | 0.832 | 0.856 | 0.885 | 0.811 |
| YOLOv12m | 0.818 | 0.843 | 0.826 | 0.817 |
| EfficientDet-D7 | 0.702 | 0.538 | 0.573 | 0.715 |
| Deformable-DETR | 0.708 | 0.614 | 0.886 | 0.813 |
| DINO | 0.795 | 0.886 | 0.899 | 0.808 |
| FCOS | 0.712 | 0.755 | 0.646 | 0.814 |
| Faster R-CNN | 0.651 | 0.859 | 0.941 | 0.645 |
| Mask-rcnn | 0.696 | 0.699 | 0.949 | 0.735 |
| RT-DETR | 0.847 | 0.830 | 0.908 | 0.796 |
| PDGV-DETR (ours) | 0.859 | 0.867 | 0.930 | 0.824 |
| Random Seed Number | RT-DETR | PDGV-DETR | Single Increment Amount |
|---|---|---|---|
| 42 | 0.844 | 0.864 | 2.0% |
| 100 | 0.845 | 0.854 | 0.9% |
| 58 | 0.835 | 0.858 | 2.3% |
| 2024 | 0.847 | 0.854 | 0.7% |
| 99 | 0.831 | 0.859 | 2.8% |
| Mean ± Standard Deviation | 0.840 ± 0.007 | 0.858 ± 0.004 | An average of 1.8% |
| Method | DWH-FPN | BasicPCock | GSWFN | Box (P) | R | mAP50 | mAP50-95 | mAP50 (Violence) |
|---|---|---|---|---|---|---|---|---|
| RT-DETR+ | ✕ | ✕ | ✕ | 0.822 | 0.802 | 0.847 | 0.631 | 0.950 |
| RT-DETR+ | ✓ | ✕ | ✕ | 0.817 | 0.823 | 0.854 | 0.632 | 0.966 |
| RT-DETR+ | ✕ | ✓ | ✕ | 0.832 | 0.810 | 0.853 | 0.630 | 0.961 |
| RT-DETR+ | ✕ | ✕ | ✓ | 0.805 | 0.832 | 0.854 | 0.644 | 0.952 |
| RT-DETR+ | ✓ | ✓ | ✕ | 0.820 | 0.827 | 0.856 | 0.630 | 0.962 |
| PDGV-DETR | ✓ | ✓ | ✓ | 0.838 | 0.822 | 0.859 | 0.643 | 0.965 |
| Method | DWH-FPN | BasicPCock | GSWFN | Total Layers | Params | GFLOPs | Inference Time (ms) | End-to-End Time (ms) | mAP50 |
|---|---|---|---|---|---|---|---|---|---|
| RT-DETR+ | ✕ | ✕ | ✕ | 299 | 19.88 M | 56.9 | 6.9 | 7.4 | 0.847 |
| RT-DETR+ | ✓ | ✕ | ✕ | 361 | 20.64 M | 58.1 | 7.4 | 7.8 | 0.854 |
| RT-DETR+ | ✕ | ✓ | ✕ | 307 | 14.35 M | 49.9 | 6.1 | 6.7 | 0.853 |
| RT-DETR+ | ✕ | ✕ | ✓ | 323 | 21.11 M | 65.5 | 7.8 | 8.3 | 0.854 |
| RT-DETR+ | ✓ | ✓ | ✕ | 374 | 15.02 M | 51.0 | 7.4 | 7.9 | 0.856 |
| PDGV-DETR | ✓ | ✓ | ✓ | 393 | 16.34 M | 59.5 | 7.5 | 8.0 | 0.859 |
| Input Degradation Type | RT-DETR | PDGV-DETR |
|---|---|---|
| Gaussian Noise σ = 5 | 0.835 | 0.847 |
| Gaussian Noise σ = 10 | 0.783 | 0.796 |
| Gaussian Noise σ = 15 | 0.675 | 0.708 |
| Grayscale Input | 0.835 | 0.842 |
| 320 × 320 | 0.842 | 0.856 |
| 160 × 160 | 0.826 | 0.832 |
| Improvements Used | Original Basic Component | Core Structure Adjustment and Differences | Impact of Computational Complexity |
|---|---|---|---|
| BasicPCock | PConv | Rather than simply introducing partial convolution (PConv), it is designed as a dual-path collaborative downsampling module, aiming to replace the BasicBlock at the end of the main network. It introduces an adaptive shortcut branch beside the fine convolution in the main path (when stride = 2, it dynamically switches to AvgPool + 1 × 1 Conv), ensuring spatial feature alignment during downsampling and preventing the feature breakage that could occlude the object. | As the cornerstone for lightweighting the model, compared with the baseline, it reduced the number of parameters in the backbone network by 27.8% (from 19.88 M to 14.35 M), and the GFLOPs decrease by 12.3% (from 56.9 to 49.9), effectively offsetting the cost of subsequent modules. |
| DWH-FPN | HS-FPN | The existing BiFPN relies on fixed weights, while the standard HS-FPN only supports unidirectional semantic transmission from top to bottom. DWH-FPN restructures it into a bidirectional hybrid architecture. It introduces transposed convolution (Transpose Convolution) to achieve high-fidelity bidirectional mapping, and combines channel attention mechanism with RepC3 module for feature stabilization. | Using this module alone slightly increased the GFLOPs to 58.1, but it yielded significant benefits in preserving edge details and suppressing background noise, solving the semantic fusion problem of simultaneous detection of small-scale weapons and macro personnel. |
| GSWFN | Original context module | These two modules were originally used for feature calibration in real-time semantic segmentation tasks. In this study, the architecture was repositioned and restructured to the end of the encoder (Encoder) in the object detection framework. Through multi-scale pooling and spatial elastic alignment mechanisms, sampling grids were explicitly generated to calibrate the spatial deviations of high- and low-level features (P5 and A3). | Although the module itself increases the computational cost, due to the lightweight compensation of BasicPCock, the overall end-to-end inference time of the model only increased by 0.6 ms compared to the baseline, and at the same time, it significantly reduced the background false alarm rate in complex security scenarios. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, N.; Xin, P.; Tian, J.; Bai, X.; Ding, H.; Xiao, Z.; Liu, Q. PDGV-DETR: Object Detection for Secure On-Site Weapon and Personnel Location Based on Dynamic Convolution and Cross-Scale Semantic Fusion. Sensors 2026, 26, 1542. https://doi.org/10.3390/s26051542
Li N, Xin P, Tian J, Bai X, Ding H, Xiao Z, Liu Q. PDGV-DETR: Object Detection for Secure On-Site Weapon and Personnel Location Based on Dynamic Convolution and Cross-Scale Semantic Fusion. Sensors. 2026; 26(5):1542. https://doi.org/10.3390/s26051542
Chicago/Turabian StyleLi, Nianfeng, Peizeng Xin, Jia Tian, Xinlu Bai, Hongjie Ding, Zhiguo Xiao, and Qian Liu. 2026. "PDGV-DETR: Object Detection for Secure On-Site Weapon and Personnel Location Based on Dynamic Convolution and Cross-Scale Semantic Fusion" Sensors 26, no. 5: 1542. https://doi.org/10.3390/s26051542
APA StyleLi, N., Xin, P., Tian, J., Bai, X., Ding, H., Xiao, Z., & Liu, Q. (2026). PDGV-DETR: Object Detection for Secure On-Site Weapon and Personnel Location Based on Dynamic Convolution and Cross-Scale Semantic Fusion. Sensors, 26(5), 1542. https://doi.org/10.3390/s26051542

